SOC-Based Multimodal Audio Processing Method and System

By frame processing and fractional Fourier transform multi-channel audio signals on the SOC chip, and using DSP and AI accelerators for processing, the problems of insufficient computing resources and large delays in processing multiple audio signals are solved, and efficient and continuous audio signal processing and resource utilization are achieved.

CN119601028BActive Publication Date: 2025-05-27HANK ELECTRONICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510145298.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-27
Estimated Expiration
2045-02-10

Smart Images

  • Figure CN119601028B_ABST
    Figure CN119601028B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of audio processing, and discloses a multi-modal audio processing method and system based on an SOC. The method includes: inputting multiple audio input signals into the ARM processor of the SOC chip for processing to generate a preprocessed audio data stream; parallelly inputting the preprocessed audio data stream into a basic audio processing stream executed by a digital signal processor (DSP) and a multi-modal feature extraction stream executed by an artificial intelligence (AI) accelerator; performing adaptive noise reduction processing and multi-channel audio signal mixing processing in the basic audio processing stream to generate a basic processing feature stream; performing multi-dimensional feature extraction and scene type recognition in the multi-modal feature extraction stream to generate a high-level feature stream; establishing a feature mapping matrix and performing audio signal reconstruction to output target audio data. The present invention improves the processing ability of the SOC chip for different types of audio signals, ensures the continuity and smoothness of the reconstructed audio signal, and effectively avoids signal distortion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and in particular to a multi-modal audio processing method and system based on SOC. Background Art

[0002] With the popularization of smart devices and the development of multimodal human-computer interaction technology, audio processing systems need to simultaneously process signals from multiple audio input sources such as microphone arrays, Bluetooth audio, and linear input, and perform complex operations such as real-time noise reduction, mixing, and sound effect processing on these audio signals. The traditional single processor architecture faces the problems of insufficient computing resources and large processing delays when processing multi-channel audio signals.

[0003] In order to meet the needs of complex multimodal audio processing, modern SOC chips adopt an integrated design, usually including an ARM processor core, a dedicated audio DSP, and an AI accelerator for multimodal processing. However, how to reasonably schedule these heterogeneous computing resources so that each processing unit can fully utilize its advantages has become an urgent problem to be solved. Summary of the invention

[0004] The present invention provides a multimodal audio processing method and system based on SOC, which improves the processing capability of the SOC chip for different types of audio signals, ensures the continuity and smoothness of the reconstructed audio signals, and effectively avoids signal distortion.

[0005] In a first aspect, the present invention provides a multimodal audio processing method based on SOC, and the multimodal audio processing method based on SOC includes:

[0006] Input multiple audio input signals into the ARM processor of the SOC chip for frame processing and fractional Fourier transform to generate pre-processed audio data stream;

[0007] Establish a task scheduling strategy table based on the pre-processed audio data stream, and input the pre-processed audio data stream in parallel to a basic audio processing stream executed by a digital signal processor DSP and a multimodal feature extraction stream executed by an artificial intelligence AI accelerator;

[0008] In the basic audio processing stream, the pre-processed audio data stream is subjected to adaptive noise reduction processing and multi-channel audio signal mixing processing to generate a basic processing feature stream;

[0009] In the multimodal feature extraction stream, the preprocessed audio data stream is input into a deep neural network model for multi-dimensional feature extraction and scene type recognition to generate a high-level feature stream;

[0010] A feature mapping matrix is ​​established according to the basic processing feature stream and the high-level feature stream, and the audio signal is reconstructed through a cubic spline interpolation algorithm to output target audio data.

[0011] In a second aspect, the present invention provides a multimodal audio processing system based on SOC, and the multimodal audio processing system based on SOC includes:

[0012] A transformation module is used to input multiple audio input signals into the ARM processor of the SOC chip for frame processing and fractional Fourier transform to generate a pre-processed audio data stream;

[0013] An establishment module is used to establish a task scheduling strategy table based on the pre-processed audio data stream, and input the pre-processed audio data stream in parallel to a basic audio processing stream executed by a digital signal processor DSP and a multimodal feature extraction stream executed by an artificial intelligence AI accelerator;

[0014] A processing module, configured to perform adaptive noise reduction processing and multi-channel audio signal mixing processing on the pre-processed audio data stream in the basic audio processing stream to generate a basic processing feature stream;

[0015] A recognition module, used for inputting the pre-processed audio data stream into a deep neural network model in the multimodal feature extraction stream to perform multi-dimensional feature extraction and scene type recognition, and generate a high-level feature stream;

[0016] The reconstruction module is used to establish a feature mapping matrix according to the basic processing feature stream and the high-level feature stream, and reconstruct the audio signal through a cubic spline interpolation algorithm to output target audio data.

[0017] A third aspect of the present invention provides a SOC-based multimodal audio processing device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the SOC-based multimodal audio processing device executes the above-mentioned SOC-based multimodal audio processing method.

[0018] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned SOC-based multimodal audio processing method.

[0019] In the technical solution provided by the present invention, by dividing the audio processing into a basic audio processing stream and a multimodal feature extraction stream, and using DSP and AI accelerators for processing respectively, a significant improvement in processing efficiency is achieved, and the processing speed is increased by 4 times compared with a single processor solution. A dynamic task scheduling strategy based on resource status monitoring is adopted, and task allocation can be adaptively adjusted according to real-time load conditions, so that the system can stably support parallel processing of up to 260 audio streams. Through the multi-level feature extraction and fusion mechanism of the deep neural network model, the semantic features and scene features in the audio signal are effectively extracted, and the system's processing ability for different types of audio signals is improved. The audio signal reconstruction method based on the cubic spline interpolation algorithm ensures the continuity and smoothness of the reconstructed audio signal and effectively avoids signal distortion. The manually optimized task scheduling achieves a 2-fold performance improvement compared to the automatic scheduling solution, significantly improving the resource utilization efficiency of the system. The noise modeling method of polynomial fitting and the noise reduction processing of the subband filter group are adopted to improve the noise reduction effect of the system in a complex noise environment. The effective fusion of acoustic features and semantic features is achieved through the feature mapping matrix, which enhances the system's understanding and processing capabilities of audio signals. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0021] Figure 1 A schematic diagram of a flow chart of a multimodal audio processing method based on SOC provided in an embodiment of the present application;

[0022] Figure 2 A schematic block diagram of the structure of a multimodal audio processing system based on SOC provided in an embodiment of the present application;

[0023] Figure 3 A schematic block diagram of the structure of a SOC-based multimodal audio processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0024] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0025] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change based on actual conditions.

[0026] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0027] It should be further understood that the term “and / or” used in the specification and appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0028] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0029] See also Figure 1 , Figure 1 A flow chart of a multimodal audio processing method based on SOC provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the SOC-based multimodal audio processing method provided in the embodiment of the present application includes steps S100 to S600.

[0030] Step S100, inputting multiple audio input signals into the ARM processor of the SOC chip for frame processing and fractional Fourier transform to generate a pre-processed audio data stream;

[0031] It is understandable that the execution subject of the present invention may be a multimodal audio processing system based on SOC, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.

[0032] Specifically, multiple audio input signals are input into the single instruction multiple data (SIMD) unit of the ARM processor in the SOC chip for parallel sampling rate conversion, and the multiple input signals are uniformly converted into a standardized audio data format to ensure the uniformity and accuracy of subsequent processing. Through the parallel computing capability of the SIMD unit, a large number of sampling rate conversion operations can be quickly completed, thereby effectively improving the processing efficiency. The standardized audio data is framed and segmented according to the preset fixed frame length and frame shift. The frame segmentation process is implemented by the ARM processor, and the continuous audio signal is divided into a series of audio frame sequences with a fixed time length to form an audio frame sequence. In order to reduce the spectrum leakage problem caused by the frame boundary, these audio frames are weighted by the Hanning window function through the floating point operation unit of the ARM processor. The Hanning window function weighted operation applies a weighting factor to the signal of each frame, so that the edge part of the frame gradually attenuates, thereby improving the accuracy of the spectrum analysis. After weighting, the time domain characteristics of the audio frame are optimized to form a weighted audio frame. The weighted audio frame is subjected to a Fourier transform based on fractional order α. The fractional Fourier transform can flexibly control the time-frequency resolution of the signal by adjusting the order α, which is suitable for the analysis of non-stationary signals. After the transformation, the generated α-order frequency domain feature matrix contains rich spectrum information of the audio signal, which is a characterization of the time-frequency characteristics of the audio signal. The frequency domain feature matrix is ​​input into the vector operation unit of the ARM processor for spectrum energy calculation. With its efficient parallel computing capability, the vector operation unit can quickly calculate the energy distribution of the spectrum. On this basis, the spectrum energy is analyzed based on the threshold decision mechanism to generate a noise suppression gain coefficient matrix. The noise suppression gain coefficient matrix is ​​obtained by analyzing the energy distribution of the spectrum and combining it with the preset noise threshold value, and is dynamically calculated to guide the subsequent noise reduction processing. The calculation process of the gain coefficient matrix uses the computing power of the ARM processor to ensure that the dynamic adjustment of the noise reduction coefficient is real-time. The α-order frequency domain feature matrix is ​​multiplied with the noise suppression gain coefficient matrix, and the original frequency domain feature data is subjected to denoising to obtain the denoised frequency domain feature data. The denoised frequency domain feature data is input into the ARM processor for inverse discrete Fourier transform processing, and the frequency domain data is converted back to the time domain. Through the inverse transform, the time domain representation of the audio signal after denoising is restored to form a time domain audio signal frame. According to the overlapping part of the time domain audio signal frame, the transition effect between the audio frames is smoothed and optimized through the weighted average superposition processing method to generate a preprocessed audio data stream.

[0033] Step S200: establishing a task scheduling strategy table based on the pre-processed audio data stream, and inputting the pre-processed audio data stream in parallel to the basic audio processing stream executed by the digital signal processor DSP and the multimodal feature extraction stream executed by the artificial intelligence AI accelerator;

[0034] Specifically, the pre-processed audio data stream is subjected to data volume statistics and frame length analysis to extract audio data processing load parameters, including the data volume per frame and frame processing time, which respectively reflect the storage requirements and computational complexity of the audio frame. Based on the audio data processing load parameters, the computing power of the digital signal processor DSP and AI accelerator of the SOC chip is evaluated to determine the resource upper limit of each processor, including the number of parallel processing units of the DSP and the number of computing cores of the AI ​​accelerator. These processor resource parameters directly affect the task allocation strategy and determine the degree of task concurrency that the SOC can process simultaneously and the speed at which the task is completed. By combining the load parameters and resource parameters, the task allocation weight coefficient is calculated, which is used to determine the computing resource allocation ratio of the basic audio processing flow and the multimodal feature extraction flow. Based on the calculation results, an initial task allocation scheme is generated to provide a preliminary allocation framework for processing resources between the two main task flows. The initial task allocation scheme is input into the scheduling optimization module for task dependency analysis. The interdependencies and timing constraints of the audio processing tasks are analyzed to construct a task execution sequence table. The task execution sequence table clearly specifies the execution order of each processing task and describes their timing relationship to ensure that the tasks can run efficiently within the SOC in the expected order. In order to optimize task scheduling, the task nodes in the task execution sequence table are sorted according to priority and finely allocated according to the resource status of the processor to generate a task scheduling strategy table. The pre-processed audio data stream is data-sliced ​​according to the strategy table. The data slicing operation divides the pre-processed audio data into two parts: basic processing data slicing and feature extraction data slicing. These slicings correspond to the data inputs of the basic audio processing flow and the multimodal feature extraction flow, respectively. The basic processing data slicing is allocated to the parallel processing unit of the DSP to establish the data channel of the basic audio processing flow. The DSP uses its parallel processing capability to load the audio data with tasks and provides optimized basic feature data for subsequent processing by performing operations such as noise reduction, mixing, and signal enhancement. At the same time, the feature extraction data slicing is allocated to the computing core of the AI ​​accelerator to establish the data channel of the multimodal feature extraction flow. In the multimodal feature extraction flow, the AI ​​accelerator extracts high-dimensional features of the audio signal through a deep learning model with its high-performance computing core. These feature extraction tasks include complex operations such as audio classification, scene recognition, and sentiment analysis.

[0035] Step S300: In the basic audio processing stream, adaptive noise reduction processing and multi-channel audio signal mixing processing are performed on the pre-processed audio data stream to generate a basic processing feature stream;

[0036] Specifically, a short-time spectrum analysis is performed on the preprocessed audio data stream, and the continuous time domain signal is converted into a spectrum energy distribution matrix to describe the time-varying characteristics of each audio signal in the frequency domain and reveal the energy distribution law of the time-frequency component. Based on the spectrum energy distribution matrix, the characteristics of the background noise are extracted. In order to effectively capture the time-frequency characteristics of the background noise, the background noise energy function is established by statistically analyzing the spectrum characteristics of the audio signal in the noise interval. The background noise energy function is approximated by using polynomial fitting technology to obtain the noise characteristic coefficient. A noise suppression subband filter group is constructed according to the noise characteristic coefficient. The filter group works in the form of a multi-bandpass filter, and its design goal is to selectively suppress the noise energy of each frequency band while retaining the integrity of the target signal as much as possible. By applying the subband filter group to the preprocessed audio data stream, the noise component is efficiently removed to generate the denoised audio data. The denoised audio data is subjected to energy envelope extraction and normalization processing. The energy envelope extraction is achieved by calculating the short-time energy of the signal, and the normalization processing is used to eliminate the energy difference between different signals to obtain the energy envelope feature vector. According to the energy envelope feature vector, the mixing weight coefficient of each audio signal is calculated. The mixing weight coefficient is obtained by analyzing the energy and time domain characteristics of the signal, and can dynamically adjust the contribution ratio of each signal in the mixing process. The noise-reduced audio data is weighted and superimposed using the mixing weight coefficient to generate mixed audio data. The core goal of mixing processing is to fuse multiple audio signals so that they can form a consistent auditory perception effect while retaining their respective characteristics. After completing the weighted superposition, the mixed audio data is input into the digital signal processor (DSP) for dynamic range compression to control the peak amplitude of the audio signal. Dynamic range compression enhances the overall balance of the audio signal by reducing the amplitude dynamic range of the signal, making it more suitable for subsequent sound effect processing. The compressed audio data is subjected to reverberation effect processing and equalizer parameter adjustment to generate sound effect processing data. The reverberation effect makes the audio signal more immersive by increasing the sense of space and environment, while the equalizer parameter adjustment is used to modify the frequency response characteristics of the audio signal to make it more in line with the target auditory needs. The sound effect processing data is combined with the noise characteristic coefficient to form a complete basic processing feature flow. The basic processing feature stream includes noise reduction parameters, mixing parameters and sound effect parameters, which is a comprehensive representation of the entire audio processing flow.

[0037] Step S400: In the multimodal feature extraction flow, the pre-processed audio data stream is input into the deep neural network model for multi-dimensional feature extraction and scene type recognition to generate a high-level feature stream;

[0038] Specifically, the preprocessed audio data stream is input into a feature extraction network containing five convolutional layers in a deep neural network model for processing. The feature extraction network extracts low-level features of the audio signal and generates an initial feature map through layer-by-layer convolution operations. The convolution layer is designed in a small receptive field and deep stacking manner. The local patterns of the audio signal in time and frequency are captured by the convolution kernel. At the same time, batch normalization and ReLU activation functions are combined to improve the stability and nonlinear expression ability of the network. After being processed by the feature extraction network, the initial feature map not only contains the spectral characteristics of the audio signal, but also retains its local structural information. The initial feature map is input into an attention network containing three multi-head attention modules in the deep neural network model. The multi-head attention module can independently pay attention to different parts of the input feature map by introducing multiple groups of independent attention heads, thereby capturing the multi-dimensional characteristics of the signal. In the attention network, the attention weight matrix is ​​calculated through the self-attention mechanism, which describes the importance and mutual relationship of each position in the feature map. The global attention representation is obtained by weighted summing the outputs of multiple attention heads. The attention weight matrix is ​​multiplied with the initial feature map to generate a weighted feature map, which reflects the optimization result of the initial feature map under the action of the attention mechanism. In order to improve the expressive power of the features, the weighted feature map is fused with the initial feature map through jump connections to generate a fused feature map. The fused feature map is input into the time series feature extraction network containing three bidirectional LSTM layers in the deep neural network model for processing. The bidirectional LSTM layer models the time series in both the forward and backward directions, and can capture the long-range dependencies of the audio signal in the time dimension. Through the three-layer stacked bidirectional LSTM layer, the network can effectively model complex time series dynamic characteristics and generate a time series feature vector with a time context relationship, which contains the time series information and spectrum evolution characteristics of the audio signal. The time series feature vector is input into the semantic feature extraction network containing four layers of fully connected layers in the deep neural network model. The semantic feature extraction network gradually extracts high-dimensional semantic information from the time series features through layer-by-layer processing of the fully connected layers. These semantic information include the core semantic features of the audio signal, such as abstract features related to specific scenes or categories. After the regularization operation of the activation function and the Dropout layer, the semantic feature vector is output. The semantic feature vector is input into the K-means clusterer with 256 cluster centers in the deep neural network model. The K-means clusterer performs Euclidean distance calculation on the semantic feature vector and assigns the feature vector to the closest cluster center to complete the preliminary clustering of the scene. The clustering result is represented in the form of a scene category vector, where each vector dimension corresponds to a cluster center and the value represents the similarity between the semantic feature vector and the center. In order to improve the accuracy of scene classification, the scene category vector is input into the scene classification network with 3 fully connected layers in the deep neural network model.The scene classification network performs nonlinear transformation on the scene category vector and outputs the scene type probability distribution of the audio signal, reflecting the possibility of the input audio signal in different scene types. The semantic feature vector and the scene type probability distribution are channel-concatenated to generate a high-level feature stream. The high-level feature stream contains 64-dimensional semantic features and 32-dimensional scene features, which represent the high-dimensional characteristics of the audio signal and the scene classification results.

[0039] Step S500: Establish a feature mapping matrix according to the basic processing feature stream and the high-level feature stream, reconstruct the audio signal through the cubic spline interpolation algorithm, and output the target audio data.

[0040] Specifically, the noise reduction parameters, mixing parameters and sound effect parameters of the basic processing feature stream are time-series aligned and feature spliced. These parameters in the basic processing feature stream describe the specific changes in the acoustic characteristics of the audio signal, including the dynamic adjustment of noise reduction, the weight configuration of mixing and the adjustment effect of sound effects. These parameters are related to time, so they need to be accurately aligned along the time axis to obtain the acoustic parameter matrix. At the same time, the 64-dimensional semantic features and 32-dimensional scene features in the high-level feature stream are normalized to eliminate the dimensional differences between different feature dimensions. The normalized semantic and scene features are input into the matrix transformation module, and the two are unified into a 128-dimensional feature space by linear projection to generate a semantic scene feature matrix. The feature mapping matrix is ​​established by calculating the feature correlation coefficient between the acoustic parameter matrix and the semantic scene feature matrix. The feature correlation coefficient quantifies the strength of the relationship between the acoustic features and the semantic scene features, indicating their degree of coupling at a specific time point. Based on these correlation coefficients, the constructed feature mapping matrix can accurately represent the mapping relationship between acoustic characteristics and semantic scene information, where each element of the matrix reflects the strength of the correlation between a specific acoustic feature and a semantic scene feature. Cubic spline node extraction is performed for each temporal position in the feature mapping matrix. The four adjacent feature points are used as control points, and a sequence of interpolation control points is generated by constructing a fourth-order control vertex matrix. The selection and sorting of control points ensure the temporal consistency and local continuity of the reconstructed signal. According to the interpolation control point sequence, the coefficients of the cubic spline basis function are calculated, and the spline coefficient matrix is ​​obtained by solving the piecewise cubic polynomial equation group. Each coefficient matrix contains the weight parameters of the four control points, which are used to determine the shape and smoothness of the spline interpolation. The spline coefficient matrix is ​​used to perform piecewise function interpolation calculations. In each time interval, the audio sampling points are interpolated and reconstructed to ensure that the waveforms between the sampling points transition evenly in time and amplitude. In order to ensure the smoothness of the reconstructed signal, the continuity constraints of the first-order derivative and the second-order derivative are imposed during the interpolation process, so that the waveform has no obvious breakpoints or peaks at the interval boundary, and the reconstructed audio sequence is obtained. The reconstructed audio sequence is input into the digital signal processor (DSP) of the SOC chip for signal shaping. The waveform is smoothed by the Kaiser window function, which effectively reduces spectral leakage with its flexible parameter adjustment ability while retaining the main characteristics of the signal. After the smoothing process is completed, the amplitude normalization operation is performed to limit the dynamic range of the signal to the target range, thereby improving the overall consistency of the signal. The standardized audio waveform adjusts the sampling rate to the target sampling rate through the resampling operation, and finally generates an analog audio signal through the digital-to-analog conversion module to output the target audio data.

[0041] The number of audio streams currently being processed in the SOC chip is counted to obtain the current system workload level. While counting the number of audio streams, the current processing speed of the basic audio processing stream and the multimodal feature extraction stream is combined to calculate the number of audio frame processing that the system can complete per unit time, reflecting the processing capacity of the SOC chip in the current working state and forming processing load statistics. The processing load statistics are input into the resource evaluation unit, and the number of parallel processing units of the digital signal processor (DSP) and the number of computing cores of the AI ​​accelerator are dynamically monitored using the resource monitoring mechanism to obtain real-time resource occupancy data. By dynamically monitoring resource usage, the system can effectively capture the uneven load or resource bottleneck. The task execution time difference of the basic audio processing stream and the multimodal feature extraction stream is calculated based on the real-time resource occupancy data, and a task processing delay model is established based on this. The model generates a task delay parameter matrix by analyzing the difference between the actual execution time and the theoretical execution time of the task. The elements in the matrix describe the delay of each processing unit in the task execution. Based on the task delay parameter matrix, the task scheduling strategy table is updated. By analyzing the delay parameters, when the task processing delay of some processing units exceeds the preset threshold, its task load is reduced to the preset target value to reduce the task pressure of the bottleneck unit. This adjustment generates a task volume adjustment plan, which clarifies the specific goals of load distribution optimization. According to the task volume adjustment plan, the new task allocation weights are calculated, and the noise reduction and mixing processing tasks in the basic audio processing flow are redistributed among the parallel processing units of the digital signal processor (DSP) to obtain a basic task allocation table. This redistribution process fully considers the parallel computing power of the DSP, aiming to maximize resource utilization efficiency while ensuring that the task load between each unit is more balanced. At the same time, the feature extraction tasks and scene recognition tasks in the multimodal feature extraction flow are redistributed among the computing cores of the AI ​​accelerator according to the computational complexity to generate a feature task allocation table. During the redistribution process, tasks with higher computational complexity are preferentially allocated to cores with stronger computing power, while low-complexity tasks are allocated to suboptimal cores to maximize the utilization of computing resources. In this way, the AI ​​accelerator completes more feature extraction and scene recognition tasks without increasing latency. After generating the basic task allocation table and the feature task allocation table, the two are merged to build a comprehensive task scheduling strategy table that supports parallel processing of 260 audio streams. The comprehensive task scheduling strategy table coordinates the task allocation of the DSP and AI accelerators, and optimizes the dynamic scheduling of resources, so that the SOC chip can efficiently cope with the high concurrency requirements of multimodal audio processing. The updated task allocation scheme is written into the task scheduler of the SOC chip to generate the target audio processing scheme. This scheme not only includes the optimized task allocation strategy, but also can dynamically adapt to changes in the number of audio streams and resource usage, thereby ensuring stable operation and efficient processing of the system under high load conditions.

[0042] In the embodiment of the present invention, by dividing the audio processing into a basic audio processing stream and a multimodal feature extraction stream, and using DSP and AI accelerators for processing respectively, a significant improvement in processing efficiency is achieved, and the processing speed is increased by 4 times compared with a single processor solution. A dynamic task scheduling strategy based on resource status monitoring can be used to adaptively adjust task allocation according to real-time load conditions, so that the system can stably support parallel processing of up to 260 audio streams. Through the multi-level feature extraction and fusion mechanism of the deep neural network model, the semantic features and scene features in the audio signal are effectively extracted, and the system's processing ability for different types of audio signals is improved. The audio signal reconstruction method based on the cubic spline interpolation algorithm ensures the continuity and smoothness of the reconstructed audio signal and effectively avoids signal distortion. Manually optimized task scheduling achieves a 2-fold performance improvement compared to the automatic scheduling solution, significantly improving the resource utilization efficiency of the system. The noise modeling method of polynomial fitting and the noise reduction processing of the subband filter group are used to improve the noise reduction effect of the system in a complex noise environment. The effective fusion of acoustic features and semantic features is achieved through the feature mapping matrix, which enhances the system's understanding and processing capabilities of audio signals.

[0043] In a specific embodiment, the process of executing step S100 may specifically include the following steps:

[0044] Input multiple audio input signals into the SIMD unit of the ARM processor in the SOC chip for parallel sampling rate conversion to obtain standardized audio data;

[0045] The standardized audio data is divided into frames according to a preset fixed frame length and frame shift to obtain an audio frame sequence, and the audio frame sequence is input into a floating point operation unit of an ARM processor to perform a weighted calculation of a Hanning window function to obtain a weighted audio frame;

[0046] Performing a Fourier transform process based on fractional order α on the weighted audio frame to obtain an α-order frequency domain feature matrix, and inputting the α-order frequency domain feature matrix into a vector operation unit of an ARM processor to perform spectral energy calculation and threshold decision processing to obtain a noise suppression gain coefficient matrix;

[0047] Perform matrix multiplication operation on the α-order frequency domain feature matrix and the noise suppression gain coefficient matrix to obtain the frequency domain feature data after noise reduction;

[0048] The frequency domain feature data after noise reduction is input into the ARM processor for inverse discrete Fourier transform processing to obtain a time domain audio signal frame, and a weighted average superposition processing is performed according to the overlapping parts of the time domain audio signal frame to generate a preprocessed audio data stream.

[0049] Specifically, the multi-channel audio input signals are input into the SIMD (Single Instruction Multiple Data) unit of the ARM processor in the SOC chip to perform parallel sampling rate conversion to generate standardized audio data. Assume that the original audio input signal is ,in Indicates The goal of sampling rate conversion is to unify the sampling rates of each audio signal to the target sampling rate. Through the parallel processing capability of the SIMD unit, multiple signals are interpolated and resampled at the same time. The formula is expressed as:

[0050] ;

[0051] in, is the signal after sampling rate conversion, and are the original and target sampling intervals, respectively, is the interpolation filter function. Through this operation, standardized audio data with a uniform sampling rate is generated. According to the preset fixed frame length and frame shift The standardized audio data is divided into frames to form an audio frame sequence. The frame operation divides the continuous time domain signal into multiple overlapping short-time signals, and each signal segment is represented as:

[0052] ;

[0053] in, is the frame index, Represents the sample index within the frame. Frame segmentation effectively captures the short-term characteristics of the audio signal. In order to reduce the spectrum leakage effect, the framed signal is weighted by the Hanning window function. The expression of the Hanning window function is:

[0054] ;

[0055] After weighting, the processing formula for each frame signal is:

[0056] ;

[0057] This operation is completed through the floating-point unit of the ARM processor, which significantly improves the accuracy of the signal's frequency domain analysis. Fourier transform (FRFT). FRFT is a generalized form of Fourier transform, which uses a fractional order parameter Flexible control of time-frequency resolution. Its mathematical definition is:

[0058] ;

[0059] in, is the fractional order kernel function, is an imaginary unit. The result of FRFT is a The 1-order frequency domain feature matrix is ​​used to capture the fine-grained spectral information of the audio signal. The frequency domain feature matrix is ​​input into the vector operation unit of the ARM processor for spectrum energy calculation and threshold decision processing. The calculation formula of spectrum energy is:

[0060] ;

[0061] The threshold decision is dynamically adjusted according to the noise level. , then it is set as the noise frequency band. Through this process, the noise suppression gain coefficient matrix is ​​generated , and its calculation formula is:

[0062] ;

[0063] in, represents the noise spectrum energy, is the noise threshold. Order frequency domain characteristic matrix and gain coefficient matrix Perform element-by-element matrix multiplication to obtain the denoised frequency domain feature data:

[0064] ;

[0065] This result effectively suppresses the noise while retaining the main spectral components of the signal. Perform an inverse discrete Fourier transform on the denoised frequency domain feature data to restore the time domain signal. The formula for the inverse transform is:

[0066] ;

[0067] In order to ensure a smooth transition between frames, the overlapping weighted average method is used to superimpose the time domain audio signal frames. The formula is:

[0068] ;

[0069] The preprocessed audio data stream generated by this step has the advantages of noise reduction, smooth transition and enhanced spectral characteristics.

[0070] In a specific embodiment, the process of executing step S200 may specifically include the following steps:

[0071] Performing data volume statistics and frame length analysis on the pre-processed audio data stream to obtain audio data processing load parameters, wherein the audio data processing load parameters include the data volume per frame and the frame processing time;

[0072] Based on the audio data processing load parameters, the computing capabilities of the digital signal processor DSP and the AI ​​accelerator of the SOC chip are evaluated to obtain the processor resource parameters, where the processor resource parameters include the number of parallel processing units of the digital signal processor DSP and the number of computing cores of the AI ​​accelerator;

[0073] Calculating a task allocation weight coefficient according to the audio data processing load parameter and the processor resource parameter to obtain an initial task allocation scheme, wherein the task allocation weight coefficient is used to determine the computing resource allocation ratio of the basic audio processing flow and the multimodal feature extraction flow;

[0074] Input the initial task allocation plan into the scheduling optimization module to analyze the task dependency and obtain the task execution sequence table, where the task execution sequence table specifies the execution order and timing relationship of each processing task;

[0075] Prioritize and allocate resources for the task nodes in the task execution sequence table to obtain a task scheduling strategy table, and segment the pre-processed audio data stream according to the task scheduling strategy table to obtain basic processing data segments and feature extraction data segments;

[0076] The basic processing data slices are allocated to the parallel processing units of the digital signal processor DSP for task loading, and the data channels of the basic audio processing stream are established. The feature extraction data slices are allocated to the computing cores of the AI ​​accelerator for task loading, and the data channels of the multimodal feature extraction stream are established.

[0077] Specifically, the data volume statistics and frame length analysis are performed on the pre-processed audio data stream to obtain the audio data processing load parameters. Assume that the sampling rate of the audio data stream is , the frame length is , the frame shift is , the amount of data per frame is calculated by the following formula:

[0078] ;

[0079] in, is the amount of data per frame, is the bit depth per sample. The frame processing time is then expressed as:

[0080] ;

[0081] By statistics and The value of describes the load characteristics of audio data in time and space. Based on the audio data processing load parameters, the computing power of the digital signal processor (DSP) and AI accelerator of the SOC chip is evaluated to obtain the processor resource parameters. Assume that there is parallel processing units, each with a computing power of , then the total computing power of DSP is:

[0082] ;

[0083] Similarly, for AI accelerators, assuming there is computing cores, each with a computing power of , then the total computing power of the AI ​​accelerator is:

[0084] ;

[0085] These resource parameters are used to describe the available computing resources in the SOC. Based on the audio data processing load parameters and the processor resource parameters, the task allocation weight coefficient is calculated to determine the resource allocation ratio of the basic audio processing flow and the multimodal feature extraction flow. Assume that the computing requirement of the basic audio processing flow is , the computational requirements of the multimodal feature extraction flow are , then the task allocation weight coefficient is expressed as:

[0086] ;

[0087] in, and are the proportion of computing resources allocated to the basic audio processing flow and the multimodal feature extraction flow, respectively. Using these weights, an initial task allocation plan is generated. After the initial task allocation plan is input into the scheduling optimization module, the task dependencies are analyzed to generate a task execution sequence table. The task execution sequence table describes the timing relationship between tasks. Assume For the tasks, whose start time is , and the end time is , then the task dependencies meet the following conditions:

[0088] ,like Depends on ;

[0089] Based on the task dependencies, the task nodes are prioritized, and the priority is determined by the computational complexity and resource requirements of the task. For example, suppose the task The computational complexity is , the resources required are , then the task priority It is expressed as:

[0090] ;

[0091] After the priorities are sorted, the tasks are matched with the resources to generate a task scheduling strategy table. According to the task scheduling strategy table, the pre-processed audio data stream is divided into basic processing data slices and feature extraction data slices. The basic processing data slices include noise reduction and mixing related tasks, which are assigned to the parallel processing unit of the DSP for task loading. By establishing a data channel for the basic audio processing stream, it is ensured that these tasks can be completed efficiently. Similarly, the feature extraction data slices include multimodal feature extraction and scene recognition tasks, which are assigned to the computing core of the AI ​​accelerator for loading, thereby establishing a data channel for the multimodal feature extraction stream.

[0092] In a specific embodiment, the process of executing step S300 may specifically include the following steps:

[0093] Performing short-time spectrum analysis on the preprocessed audio data stream to obtain a spectrum energy distribution matrix, wherein the spectrum energy distribution matrix includes time-frequency components of each audio signal;

[0094] A background noise energy function is established based on the spectrum energy distribution matrix, and a polynomial fitting is performed on the background noise energy function to obtain a noise characteristic coefficient;

[0095] A noise suppression subband filter group is constructed according to the noise characteristic coefficient, and subband filtering is performed on the preprocessed audio data stream to obtain the denoised audio data;

[0096] Performing energy envelope extraction and normalization processing on the denoised audio data to obtain an energy envelope feature vector, wherein the energy envelope feature vector is used to characterize the loudness characteristics of each audio signal;

[0097] Calculating the mixing weight coefficient of each audio signal according to the energy envelope feature vector, and performing weighted superposition on the denoised audio data to obtain mixed audio data;

[0098] Inputting the mixed audio data into a digital signal processor DSP for dynamic range compression to obtain dynamic compressed audio data, wherein the dynamic range compression is used to control the peak amplitude of the audio signal;

[0099] The dynamic compressed audio data is subjected to reverberation effect processing and equalizer parameter adjustment to obtain sound effect processing data, and the sound effect processing data and the noise characteristic coefficient are combined to generate a basic processing characteristic stream, which includes noise reduction parameters, mixing parameters and sound effect parameters.

[0100] Specifically, short-time spectrum analysis is performed on the preprocessed audio data stream to obtain a spectrum energy distribution matrix. Assume that the audio signal is ,in Indicates The audio signal is divided into multiple frames of fixed length, and each frame signal is used Represents. Through short-time Fourier transform, the time domain signal is converted into a time-frequency domain signal, the formula is:

[0101] ;

[0102] in, It is Road signal Frame, The spectrum value of the frequency point, is the frame length, It is frame shift. is a weighted window function (such as a Hanning window). By calculating , and get the spectrum energy distribution matrix , which describes the energy distribution characteristics of the audio signal in the time-frequency domain. Based on the spectrum energy distribution matrix Establishing the background noise energy function The background noise is extracted by time averaging method, the formula is:

[0103] ;

[0104] in, is the total number of sampling frames. In order to model the background noise energy function, the polynomial fitting technique is used to solve the polynomial coefficients by the least squares method. The fitting formula is:

[0105] ;

[0106] in, is the polynomial order, is the fitting coefficient, which is optimized by The best fit result is obtained. These coefficients constitute the noise characteristic coefficients According to the noise characteristic coefficient, a noise suppression subband filter bank is constructed. Each filter is designed for a different frequency band, and its gain coefficient Determined by the following formula:

[0107] ;

[0108] By applying the filter bank to the spectral energy distribution matrix, the denoised frequency domain signal is obtained :

[0109] ;

[0110] The denoised frequency domain signal is converted back to the time domain through inverse short-time Fourier transform to obtain the denoised audio data Extract the energy envelope of the denoised audio data and calculate the short-time energy of the signal :

[0111] ;

[0112] Then the energy envelope is normalized, the formula is:

[0113] ;

[0114] Normalized energy envelope eigenvector Used to characterize the loudness characteristics of the signal. Calculate the mixing weight coefficients of each audio signal based on the energy envelope eigenvector , with weights proportional to loudness:

[0115] ;

[0116] Then the denoised audio data is weighted and superimposed to generate mixed audio data. :

[0117] ;

[0118] The mixed audio data is input into the DSP for dynamic range compression to control the peak amplitude of the audio signal. The gain function of dynamic range compression Defined as:

[0119] ;

[0120] in, is the compression threshold, and the audio signal after dynamic range compression is:

[0121] ;

[0122] The compressed audio data is processed with reverberation effects and the equalizer parameters are adjusted. The reverberation effect is achieved through convolution operation, and the formula is:

[0123] ;

[0124] in, is the reverberation impulse response. The equalizer parameter adjustment is achieved through frequency domain gain adjustment, the formula is:

[0125] ;

[0126] in, is the gain function of the equalizer. The processed sound effect data is combined with the noise characteristic coefficient to generate the basic processing characteristic flow , which contains the noise reduction parameters , Mixing parameters and sound parameters .

[0127] In a specific embodiment, the process of executing step S400 may specifically include the following steps:

[0128] In the multimodal feature extraction flow, the preprocessed audio data stream is input into the feature extraction network containing 5 convolutional layers in the deep neural network model for processing to obtain the initial feature map;

[0129] The initial feature map is input into the attention network containing three multi-head attention modules in the deep neural network model for processing to obtain the attention weight matrix;

[0130] Perform matrix multiplication on the attention weight matrix and the initial feature map to obtain a weighted feature map, and fuse the weighted feature map with the initial feature map through a jump connection to obtain a fused feature map;

[0131] The fused feature map is input into the time series feature extraction network containing 3 bidirectional LSTM layers in the deep neural network model for processing to obtain the time series feature vector;

[0132] The time series feature vector is input into the semantic feature extraction network containing 4 fully connected layers in the deep neural network model for processing to obtain the semantic feature vector;

[0133] The semantic feature vector is input into the K-means clusterer with 256 cluster centers in the deep neural network model for scene clustering, and the scene category vector is obtained by Euclidean distance calculation;

[0134] The scene category vector is input into the scene classification network containing 3 fully connected layers in the deep neural network model for processing, and the probability distribution of the scene type is output;

[0135] The semantic feature vector and the scene type probability distribution are channel-concatenated to generate a high-level feature stream, which contains 64-dimensional semantic features and 32-dimensional scene features.

[0136] Specifically, the preprocessed audio data stream is input into the feature extraction network containing 5 convolutional layers in the deep neural network model for processing to generate the initial feature map. Assume that the input audio data is ,in Indicates the number of time frames, Represents the frequency feature dimension of each frame. The convolution operation extracts the local features of the audio signal through multiple convolution kernels. The specific formula is:

[0137] ;

[0138] in, For the Layer convolution output, It is Tier convolution kernels, is the bias term, is the activation function (such as ReLU), is the number of channels in the previous layer, and * indicates the convolution operation. After 5 layers of convolution, the initial feature map generated ,in and are the dimensionality reduction results of time and frequency respectively, is the number of channels in the last layer. Input is an attention network containing 3 multi-head attention modules. The multi-head attention mechanism generates an attention weight matrix by capturing the global dependencies between different positions in the feature map. The attention calculation formula is:

[0139] ;

[0140] in, is the linear projection matrix, is the dimension of the attention head. The attention weight matrix With the initial feature map Multiply to generate a weighted feature map:

[0141] ;

[0142] The weighted feature map is connected through skip connections With the initial feature map To perform feature fusion, the formula is:

[0143] ;

[0144] Fusion feature map It contains global context information and retains the local details of the initial features. The fused feature map is input into a temporal feature extraction network consisting of three bidirectional LSTM layers. The bidirectional LSTM captures the dynamic changes of the audio signal in the time dimension by modeling the forward and backward time dependencies. The hidden state update formula of LSTM is:

[0145] ;

[0146] The bidirectional operation forwards the hidden state and the backward hidden state Concatenate into time series feature vectors:

[0147] ;

[0148] After 3 layers of bidirectional LSTM processing, a time series feature vector is generated ,in is the dimension of each hidden unit. The time series feature vector is input into the semantic feature extraction network containing 4 fully connected layers, and high-dimensional semantic information is extracted through layer-by-layer nonlinear transformation. Let the output of each layer be , and its calculation formula is:

[0149] ;

[0150] in, For the The layer weight matrix, is the bias term, is the activation function. The final output semantic feature vector is The semantic feature vector is input into the K-means clusterer with 256 cluster centers for scene clustering. Let the cluster center be , , the scene category vector is calculated by Euclidean distance:

[0151] ;

[0152] right Normalize and generate scene category vector The scene category vector is input into the scene classification network containing 3 layers of fully connected layers. Through feature extraction and classification, the probability distribution of the scene type is output. Each serving Indicates that the audio signal belongs to a scene The probability of semantic feature vector And the probability distribution of scene types Perform channel splicing to generate high-level feature stream , where the first 64 dimensions are semantic features and the last 32 dimensions are scene features.

[0153] In a specific embodiment, the process of executing step S500 may specifically include the following steps:

[0154] The noise reduction parameters, mixing parameters and sound effect parameters of the basic processing feature stream are time-series aligned and feature-joined to obtain an acoustic parameter matrix;

[0155] The 64-dimensional semantic features and 32-dimensional scene features in the high-level feature stream are normalized, and the two feature spaces are unified into a 128-dimensional feature space through matrix transformation to obtain a semantic scene feature matrix;

[0156] The feature correlation coefficient is calculated according to the acoustic parameter matrix and the semantic scene feature matrix, and a feature mapping matrix is ​​constructed based on the feature correlation coefficient, wherein the matrix elements in the feature mapping matrix represent the mapping strength between the acoustic features and the semantic scene features;

[0157] Perform cubic spline node extraction on each time-series position in the feature mapping matrix, take four adjacent feature points as control points, construct a fourth-order control vertex matrix, and obtain an interpolation control point sequence;

[0158] The cubic spline basis function coefficients are calculated according to the interpolation control point sequence, and the spline coefficient matrix is ​​obtained by solving the piecewise cubic polynomial equation system, where each coefficient matrix contains the weight parameters of four control points;

[0159] Perform piecewise function interpolation calculation on the spline coefficient matrix, reconstruct the audio sampling points in each time interval, and ensure smooth transition of the waveform through continuity constraints of the first-order derivative and the second-order derivative to obtain a reconstructed audio sequence;

[0160] The reconstructed audio sequence is input into the digital signal processor DSP of the SOC chip for signal shaping, the waveform is smoothed by the Kaiser window function, and the amplitude normalization operation is performed to obtain a standardized audio waveform. The standardized audio waveform is then resampled and converted into digital-to-analog format to output the target audio data.

[0161] Specifically, the noise reduction parameters, mixing parameters and sound effect parameters in the basic processing feature stream are time-series aligned and feature-joined to generate an acoustic parameter matrix. Assume that the basic processing feature stream is the time series of three sets of parameters: , respectively represent noise reduction, mixing and sound effect parameters, where Represents the time frame index. These parameters are time-series aligned to ensure that they are consistent on the same time scale. The formula for feature splicing is:

[0162] ;

[0163] in, is the feature vector of each frame of the acoustic parameter matrix, is the concatenated feature dimension. For the 64-dimensional semantic features in the high-level feature stream and 32-dimensional scene features Normalization is performed to eliminate the dimensional differences between feature dimensions. The normalization formula is:

[0164] ;

[0165] in, and are the mean and standard deviation of semantic features and scene features respectively. The normalized features are mapped to a unified 128-dimensional feature space through linear matrix transformation, and the formula is:

[0166] ;

[0167] in, and is the mapping matrix, is the feature vector of each frame of the semantic scene feature matrix. Based on the acoustic parameter matrix and the semantic scene feature matrix , by calculating the feature correlation coefficient, the feature mapping matrix is ​​constructed. The feature correlation coefficient is measured by calculating the dot product between two sets of features, and the formula is:

[0168] ;

[0169] in, is in the time frame and The correlation on Represents the bi-norm of the eigenvector. Eigenmapping matrix Represents the mapping strength between acoustic features and semantic scene features. In the feature mapping matrix At each time position, perform cubic spline node extraction. Take the four adjacent feature points as control points to construct a fourth-order control vertex matrix , where each column is a control point:

[0170] ;

[0171] According to the control point sequence , calculate the cubic spline basis function coefficients. The cubic spline basis function is expressed as:

[0172] ;

[0173] in, is the normalized time, are coefficients, which are solved by the following system of equations:

[0174] ;

[0175] The obtained spline coefficient matrix is ​​used for piecewise interpolation calculation. In each time interval, the audio sampling points are reconstructed, and the formula is:

[0176] ;

[0177] Through the continuity constraints of the first-order and second-order derivatives, the waveform is ensured to transition smoothly and form a reconstructed audio sequence. The reconstructed audio sequence is input into the digital signal processor (DSP) of the SOC chip for signal shaping, and the waveform is smoothed by the Kaiser window function. The formula is:

[0178] ;

[0179] in, is the zero-order Bessel function, is the shape parameter of the window function. Amplitude normalization is performed on the smoothed waveform, and the formula is:

[0180] ;

[0181] Resample the normalized audio waveform to the target sampling rate , and output the target audio data through digital-to-analog conversion.

[0182] In a specific embodiment, the above-mentioned SOC-based multimodal audio processing method further includes the following steps:

[0183] The number of audio streams being processed in the SOC chip is counted, and the number of audio frames processed per unit time is calculated based on the current processing speed of the basic audio processing stream and the multimodal feature extraction stream to obtain processing load statistics;

[0184] Input the processing load statistics into the resource evaluation unit, dynamically monitor the number of parallel processing units of the digital signal processor DSP and the number of computing cores of the AI ​​accelerator, and obtain real-time resource occupancy data;

[0185] The task execution time difference between the basic audio processing flow and the multimodal feature extraction flow is calculated based on the real-time resource occupancy data, and a task processing delay model is established to obtain the task delay parameter matrix.

[0186] The task scheduling strategy table is updated based on the task delay parameter matrix, and the task load of the processing unit whose task processing delay exceeds the preset threshold is reduced by a preset target value to obtain a task load adjustment plan;

[0187] Calculate new task allocation weights according to the task volume adjustment plan, reallocate the noise reduction processing and mixing processing tasks in the basic audio processing flow among the parallel processing units of the digital signal processor DSP, and obtain a basic task allocation table;

[0188] The feature extraction tasks and scene recognition tasks in the multimodal feature extraction flow are reallocated among the computing cores of the AI ​​accelerator according to the computational complexity to obtain a feature task allocation table;

[0189] The basic task allocation table and the feature task allocation table are merged to build a comprehensive task scheduling strategy table that supports parallel processing of 260 audio streams, obtain an updated task allocation plan, and write the updated task allocation plan into the task scheduler of the SOC chip to generate a target audio processing plan.

[0190] Specifically, the number of audio streams being processed is counted, and the number of audio frames processed per unit time is calculated based on the current processing speed of the basic audio processing stream and the multimodal feature extraction stream to generate processing load statistics. Assume that the number of audio streams currently processed by the SOC chip is , the sampling rate of each audio stream is , the frame length is , the frame shift is . Processing time per frame It is expressed as:

[0191] ;

[0192] At the same time, the computational complexity of each frame processing is determined by the sum of the task complexity of the basic audio processing flow and the multimodal feature extraction flow, which are denoted as and The number of audio frames processed per unit time. for:

[0193] ;

[0194] By counting the current frame processing rate, the processing load statistics of the entire SOC chip are obtained. The processing load statistics are input into the resource evaluation unit to dynamically monitor the digital signal processor (DSP) and AI accelerator of the SOC chip. Assume that the DSP contains Parallel processing units, each with a computing power of , AI accelerator includes computing cores, and the computing power of each core is The resource utilization of the SOC chip is calculated through real-time monitoring and is expressed as:

[0195] ;

[0196] in, and are the resource occupancy rates of the DSP and AI accelerator respectively. Based on the real-time resource occupancy data, the difference in task execution time between the basic audio processing flow and the multimodal feature extraction flow is calculated. Assume that the actual execution time of each task is , the theoretical execution time is , then the task execution time difference is:

[0197] ;

[0198] Based on this, a task processing delay model is established, and the delay matrix Each element of is:

[0199] ;

[0200] in, Represents the task index, Indicates the audio stream index. The delay matrix reflects the processing delay of each task on different audio streams. Based on the task delay parameter matrix, the task scheduling strategy table is updated. If the task delay of a processing unit exceeds the preset threshold , then its task load is reduced to the preset target value The adjusted task volume is:

[0201] ;

[0202] Thus, a task volume adjustment plan is generated. According to the task volume adjustment plan, the task allocation weights of the basic audio processing flow and the multimodal feature extraction flow are recalculated. Assume that the task weight of the basic audio processing flow is , the task weight of the multimodal feature extraction flow is ,but:

[0203] ;

[0204] The noise reduction and mixing tasks in the basic audio processing stream are reallocated among the parallel processing units of the DSP according to the new task weights to generate a basic task allocation table. At the same time, the feature extraction tasks and scene recognition tasks in the multimodal feature extraction stream are reallocated among the computing cores of the Al accelerator according to the computational complexity to generate a feature task allocation table. The basic task allocation table and the feature task allocation table are merged to construct a comprehensive task scheduling strategy table that supports parallel processing of 260 audio streams. The comprehensive task scheduling strategy table ensures maximum resource utilization of the DSP and AI accelerator and keeps latency within an acceptable range by dynamically adjusting the task allocation plan. The updated task allocation plan is written into the task scheduler of the SOC chip to generate the target audio processing plan.

[0205] See also Figure 2 , Figure 2 The structure schematic block diagram of the multimodal audio processing system 200 based on SOC provided in the embodiment of the present application is as follows: Figure 2 As shown, the SOC-based multimodal audio processing system 200 includes:

[0206] The transformation module 210 is used to input the multi-channel audio input signals into the ARM processor of the SOC chip for frame processing and fractional Fourier transform to generate a pre-processed audio data stream;

[0207] Establishing module 220, for establishing a task scheduling strategy table based on the pre-processed audio data stream, and inputting the pre-processed audio data stream in parallel to a basic audio processing stream executed by a digital signal processor DSP and a multimodal feature extraction stream executed by an artificial intelligence AI accelerator;

[0208] The processing module 230 is used to perform adaptive noise reduction processing and multi-channel audio signal mixing processing on the pre-processed audio data stream in the basic audio processing stream to generate a basic processing feature stream;

[0209] The recognition module 240 is used to input the pre-processed audio data stream into the deep neural network model to perform multi-dimensional feature extraction and scene type recognition in the multi-modal feature extraction stream to generate a high-level feature stream;

[0210] The reconstruction module 250 is used to establish a feature mapping matrix according to the basic processing feature stream and the high-level feature stream, and reconstruct the audio signal through a cubic spline interpolation algorithm to output target audio data.

[0211] Through the collaboration of the above components, the processing efficiency is significantly improved by dividing the audio processing into the basic audio processing stream and the multimodal feature extraction stream, and using DSP and AI accelerators to process them respectively. Compared with the single processor solution, the processing speed is increased by 4 times. The dynamic task scheduling strategy based on resource status monitoring can adaptively adjust the task allocation according to the real-time load situation, so that the system can stably support the parallel processing of up to 260 audio streams. Through the multi-level feature extraction and fusion mechanism of the deep neural network model, the semantic features and scene features in the audio signal are effectively extracted, and the system's processing ability for different types of audio signals is improved. The audio signal reconstruction method based on the cubic spline interpolation algorithm ensures the continuity and smoothness of the reconstructed audio signal and effectively avoids signal distortion. The manually optimized task scheduling achieves a 2-fold performance improvement compared to the automatic scheduling solution, significantly improving the resource utilization efficiency of the system. The noise modeling method of polynomial fitting and the noise reduction processing of the subband filter group are used to improve the noise reduction effect of the system in complex noise environments. The effective fusion of acoustic features and semantic features is achieved through the feature mapping matrix, which enhances the system's understanding and processing capabilities of audio signals.

[0212] See also Figure 3 , Figure 3 The present invention provides a schematic block diagram of the structure of a SOC-based multimodal audio processing device 300 according to an embodiment of the present application. The SOC-based multimodal audio processing device 300 includes a processor 301 and a memory 302. The processor 301 and the memory 302 are connected via a system bus 303, wherein the memory 302 may include a non-volatile storage medium and an internal memory.

[0213] The non-volatile storage medium may store a computer program. The computer program includes program instructions, and when the program instructions are executed by the processor 301, the processor 301 may execute any of the above-mentioned SOC-based multimodal audio processing methods.

[0214] The processor 301 is used to provide computing and control capabilities to support the operation of the entire SOC-based multimodal audio processing device 300.

[0215] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor 301, the processor 301 can execute any of the above-mentioned SOC-based multimodal audio processing methods.

[0216] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a partial structure related to the present application scheme, and does not constitute a limitation on the SOC-based multimodal audio processing device 300 involved in the present application scheme. The specific SOC-based multimodal audio processing device 300 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0217] It should be understood that the processor 301 may be a central processing unit (CPU), and the processor 301 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0218] It should be noted that technical personnel in the relevant field can clearly understand that, for the convenience and conciseness of description, the specific working process of the SOC-based multimodal audio processing device 300 described above can refer to the corresponding process of the aforementioned SOC-based multimodal audio processing method, and will not be repeated here.

[0219] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by one or more processors, the one or more processors implement the SOC-based multimodal audio processing method provided in the embodiment of the present application.

[0220] The computer-readable storage medium may be an internal storage unit of the SOC-based multimodal audio processing device 300 of the aforementioned embodiment, such as a hard disk or memory of the SOC-based multimodal audio processing device 300. The computer-readable storage medium may also be an external storage device of the SOC-based multimodal audio processing device 300, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc., equipped with the SOC-based multimodal audio processing device 300.

[0221] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0222] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

[0223] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A multimodal audio processing method based on SOC, characterized in that: include: Input multiple audio input signals into the ARM processor of the SOC chip for frame processing and fractional Fourier transform to generate pre-processed audio data stream; Establish a task scheduling strategy table based on the pre-processed audio data stream, and input the pre-processed audio data stream in parallel to a basic audio processing stream executed by a digital signal processor DSP and a multimodal feature extraction stream executed by an artificial intelligence AI accelerator; In the basic audio processing stream, the pre-processed audio data stream is subjected to adaptive noise reduction processing and multi-channel audio signal mixing processing to generate a basic processing feature stream; In the multimodal feature extraction stream, the preprocessed audio data stream is input into a deep neural network model for multi-dimensional feature extraction and scene type recognition to generate a high-level feature stream; A feature mapping matrix is ​​established according to the basic processing feature stream and the high-level feature stream, and the audio signal is reconstructed through a cubic spline interpolation algorithm to output target audio data.

2. The multimodal audio processing method based on SOC according to claim 1, characterized in that: The method of inputting the multi-channel audio input signals into the ARM processor of the SOC chip for frame processing and fractional Fourier transform to generate a pre-processed audio data stream includes: Input multiple audio input signals into the SIMD unit of the ARM processor in the SOC chip for parallel sampling rate conversion to obtain standardized audio data; The standardized audio data is frame-segmented according to a preset fixed frame length and frame shift to obtain an audio frame sequence, and the audio frame sequence is input into a floating point operation unit of the ARM processor to perform a Hanning window function weighted calculation to obtain a weighted audio frame; Performing a Fourier transform process based on fractional order α on the weighted audio frame to obtain an α-order frequency domain feature matrix, and inputting the α-order frequency domain feature matrix into a vector operation unit of the ARM processor to perform spectral energy calculation and threshold decision processing to obtain a noise suppression gain coefficient matrix; Performing a matrix multiplication operation on the α-order frequency domain feature matrix and the noise suppression gain coefficient matrix to obtain frequency domain feature data after noise reduction; The denoised frequency domain feature data is input into the ARM processor for inverse discrete Fourier transform processing to obtain a time domain audio signal frame, and a weighted average superposition processing is performed according to the overlapping parts of the time domain audio signal frame to generate a preprocessed audio data stream.

3. The multimodal audio processing method based on SOC according to claim 2, characterized in that: The task scheduling strategy table is established based on the pre-processed audio data stream, and the pre-processed audio data stream is input in parallel to the basic audio processing stream executed by the digital signal processor DSP and the multimodal feature extraction stream executed by the artificial intelligence AI accelerator, including: Performing data volume statistics and frame length analysis on the pre-processed audio data stream to obtain audio data processing load parameters, wherein the audio data processing load parameters include the data volume per frame and the frame processing time; Based on the audio data processing load parameter, the computing capabilities of the digital signal processor DSP and the AI ​​accelerator of the SOC chip are evaluated to obtain processor resource parameters, wherein the processor resource parameters include the number of parallel processing units of the digital signal processor DSP and the number of computing cores of the AI ​​accelerator; Calculating a task allocation weight coefficient according to the audio data processing load parameter and the processor resource parameter to obtain an initial task allocation scheme, wherein the task allocation weight coefficient is used to determine a computing resource allocation ratio of a basic audio processing flow and a multimodal feature extraction flow; Inputting the initial task allocation plan into the scheduling optimization module to perform task dependency analysis to obtain a task execution sequence table, wherein the task execution sequence table specifies the execution order and timing relationship of each processing task; Prioritizing and allocating resources for the task nodes in the task execution sequence table to obtain a task scheduling strategy table, and slicing the pre-processed audio data stream according to the task scheduling strategy table to obtain basic processing data slices and feature extraction data slices; The basic processing data slices are allocated to the parallel processing units of the digital signal processor DSP for task loading, and a data channel for the basic audio processing stream is established; and the feature extraction data slices are allocated to the computing core of the AI ​​accelerator for task loading, and a data channel for the multimodal feature extraction stream is established.

4. The multimodal audio processing method based on SOC according to claim 3, characterized in that: In the basic audio processing stream, the pre-processed audio data stream is subjected to adaptive noise reduction processing and multi-channel audio signal mixing processing to generate a basic processing feature stream, including: Performing short-time spectrum analysis on the preprocessed audio data stream to obtain a spectrum energy distribution matrix, wherein the spectrum energy distribution matrix includes time-frequency components of each audio signal; Establishing a background noise energy function based on the spectrum energy distribution matrix, and performing polynomial fitting on the background noise energy function to obtain a noise characteristic coefficient; Constructing a noise suppression sub-band filter group according to the noise characteristic coefficients, and performing sub-band filtering on the pre-processed audio data stream to obtain denoised audio data; Performing energy envelope extraction and normalization processing on the denoised audio data to obtain an energy envelope feature vector, wherein the energy envelope feature vector is used to characterize the loudness characteristics of each audio signal; Calculating mixing weight coefficients of the audio signals of each channel according to the energy envelope feature vector, and performing weighted superposition on the denoised audio data to obtain mixed audio data; Inputting the mixed audio data into a digital signal processor DSP for dynamic range compression to obtain dynamic compressed audio data, wherein the dynamic range compression is used to control the peak amplitude of the audio signal; The dynamically compressed audio data is subjected to reverberation effect processing and equalizer parameter adjustment to obtain sound effect processing data, and the sound effect processing data is combined with the noise characteristic coefficient to generate a basic processing characteristic stream, wherein the basic processing characteristic stream includes noise reduction parameters, mixing parameters and sound effect parameters.

5. The multimodal audio processing method based on SOC according to claim 4, characterized in that: In the multimodal feature extraction flow, the pre-processed audio data stream is input into a deep neural network model for multi-dimensional feature extraction and scene type recognition to generate a high-level feature stream, including: In the multimodal feature extraction stream, the preprocessed audio data stream is input into a feature extraction network including 5 convolutional layers in a deep neural network model for processing to obtain an initial feature map; Inputting the initial feature map into the attention network including three multi-head attention modules in the deep neural network model for processing to obtain an attention weight matrix; Performing a matrix multiplication operation on the attention weight matrix and the initial feature map to obtain a weighted feature map, and performing feature fusion on the weighted feature map and the initial feature map through a skip connection to obtain a fused feature map; Input the fused feature map into the time series feature extraction network including 3 bidirectional LSTM layers in the deep neural network model for processing to obtain a time series feature vector; Inputting the time series feature vector into a semantic feature extraction network including 4 fully connected layers in the deep neural network model for processing to obtain a semantic feature vector; Inputting the semantic feature vector into a K-means clusterer containing 256 cluster centers in the deep neural network model to perform scene clustering, and obtaining a scene category vector by Euclidean distance calculation; Input the scene category vector into the scene classification network including 3 fully connected layers in the deep neural network model for processing, and output the scene type probability distribution; The semantic feature vector and the scene type probability distribution are subjected to channel splicing processing to generate a high-level feature stream, wherein the high-level feature stream includes 64-dimensional semantic features and 32-dimensional scene features.

6. The multimodal audio processing method based on SOC according to claim 5, characterized in that: The step of establishing a feature mapping matrix according to the basic processing feature stream and the high-level feature stream, reconstructing the audio signal by a cubic spline interpolation algorithm, and outputting target audio data includes: Performing time sequence alignment and feature splicing on the noise reduction parameters, mixing parameters and sound effect parameters of the basic processing feature stream to obtain an acoustic parameter matrix; Normalizing the 64-dimensional semantic features and the 32-dimensional scene features in the high-level feature stream, and unifying the two feature spaces into a 128-dimensional feature space through matrix transformation to obtain a semantic scene feature matrix; Calculating a feature correlation coefficient according to the acoustic parameter matrix and the semantic scene feature matrix, and constructing a feature mapping matrix based on the feature correlation coefficient, wherein the matrix elements in the feature mapping matrix represent the mapping strength between the acoustic features and the semantic scene features; Performing cubic spline node extraction on each time-series position in the feature mapping matrix, taking four adjacent feature points as control points, constructing a fourth-order control vertex matrix, and obtaining an interpolation control point sequence; Calculating cubic spline basis function coefficients according to the interpolation control point sequence, and obtaining a spline coefficient matrix by solving a piecewise cubic polynomial equation group, wherein each coefficient matrix contains weight parameters of four control points; Perform piecewise function interpolation calculation on the spline coefficient matrix, reconstruct the audio sampling points in each time interval, and ensure smooth transition of the waveform through continuity constraints of the first-order derivative and the second-order derivative to obtain a reconstructed audio sequence; The reconstructed audio sequence is input into the digital signal processor DSP of the SOC chip for signal shaping, the waveform is smoothed by the Kaiser window function, and an amplitude normalization operation is performed to obtain a standardized audio waveform, and the standardized audio waveform is resampled and digital-to-analog converted to output the target audio data.

7. The multimodal audio processing method based on SOC according to claim 6, characterized in that: The SOC-based multimodal audio processing method further includes: Counting the number of audio streams being processed in the SOC chip, and calculating the number of audio frames processed per unit time according to the current processing speeds of the basic audio processing stream and the multimodal feature extraction stream, to obtain processing load statistics; Inputting the processing load statistical data into a resource evaluation unit, dynamically monitoring the number of parallel processing units of the digital signal processor DSP and the number of computing cores of the AI ​​accelerator, and obtaining real-time resource occupancy data; Calculating the task execution time difference between the basic audio processing flow and the multimodal feature extraction flow according to the real-time resource occupancy rate data, and establishing a task processing delay model to obtain a task delay parameter matrix; The task scheduling strategy table is updated based on the task delay parameter matrix, and the task load of the processing unit whose task processing delay exceeds the preset threshold is reduced by a preset target value to obtain a task load adjustment plan; Calculate new task allocation weights according to the task volume adjustment scheme, and reallocate the noise reduction processing and mixing processing tasks in the basic audio processing stream among the parallel processing units of the digital signal processor DSP to obtain a basic task allocation table; Redistributing the feature extraction tasks and scene recognition tasks in the multimodal feature extraction flow among the computing cores of the AI ​​accelerator according to computational complexity to obtain a feature task allocation table; The basic task allocation table and the characteristic task allocation table are merged to construct a comprehensive task scheduling strategy table that supports parallel processing of 260 audio streams, obtain an updated task allocation plan, and write the updated task allocation plan into the task scheduler of the SOC chip to generate a target audio processing plan.

8. A multimodal audio processing system based on SOC, characterized in that: The method for performing a multimodal audio processing method based on SOC as claimed in any one of claims 1 to 7 comprises: A transformation module is used to input multiple audio input signals into the ARM processor of the SOC chip for frame processing and fractional Fourier transform to generate a pre-processed audio data stream; An establishment module is used to establish a task scheduling strategy table based on the pre-processed audio data stream, and input the pre-processed audio data stream in parallel to a basic audio processing stream executed by a digital signal processor DSP and a multimodal feature extraction stream executed by an artificial intelligence AI accelerator; A processing module, configured to perform adaptive noise reduction processing and multi-channel audio signal mixing processing on the pre-processed audio data stream in the basic audio processing stream to generate a basic processing feature stream; A recognition module, used for inputting the pre-processed audio data stream into a deep neural network model in the multimodal feature extraction stream to perform multi-dimensional feature extraction and scene type recognition, and generate a high-level feature stream; The reconstruction module is used to establish a feature mapping matrix according to the basic processing feature stream and the high-level feature stream, and reconstruct the audio signal through a cubic spline interpolation algorithm to output target audio data.

9. A multimodal audio processing device based on SOC, characterized in that: The SOC-based multimodal audio processing device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory to enable the SOC-based multimodal audio processing device to perform the SOC-based multimodal audio processing method according to any one of claims 1-7.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the SOC-based multimodal audio processing method as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Domain speech recognition method and system based on RAG

    CN119296516A

  • Voice synthesis method, device and equipment based on AI large model in multi-language scene

    CN119314466A