Audio signal processing apparatus and method based on feature segmentation and feature combination, and computer program
By combining feature segmentation and low-complexity neural network processing with power-law compression technology, the high complexity of deep learning audio processing methods on resource-constrained devices is solved, achieving efficient audio processing on embedded devices.
Patent Information
- Application Number
- CN202480023695.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-29
- Filing Date
- 2024-03-26
- Publication Date
- 2025-12-16
AI Technical Summary
Existing deep learning-based audio processing methods are too computationally complex and memory-intensive, making them difficult to deploy on resource-constrained embedded devices.
The input feature set is divided into multiple subsets by a feature segmenter and processed separately by a low-complexity neural network. The neural network complexity of each subset decreases. Finally, the results are combined by a feature combiner. Power-law compression technique is used to limit the dynamic range to improve the learning and generalization ability of the neural network.
It reduces the computational complexity and resource requirements of neural networks, improves deployment capabilities on embedded devices, and enhances audio processing performance.
Smart Images

Figure CN121153079A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information signal processing, and in particular to information signal processing technology based on neural networks. Background Technology
[0002] Deep learning (DL) based audio processing solutions have become increasingly popular in recent years, outperforming traditional signal processing methods. However, existing methods generally suffer from high computational complexity and memory requirements, making them difficult to deploy on embedded devices that are typically limited by memory and computing resources.
[0003] Figure 2 This demonstrates a typical audio processing workflow based on deep neural networks (DNNs). The input to such methods is typically an audio waveform (also known as a time-domain audio signal). The first block in the workflow is the input preprocessing block (block P1), which mainly uses audio waveform feature extraction techniques and dimensionality adjustment mechanisms to make the input conform to the dimensionality requirements of the DNN model. Feature extraction in this block can employ deterministic time-frequency transforms (such as Short-Time Fourier Transform, STFT) or learn feature sets specific to the application scenario through backpropagation. After the feature extraction step, a multidimensional (usually 2D or 3D) tensor is obtained, which is then provided as input to the main DNN processing block (block D1).
[0004] Block D1 is the main DNN processing block that operates on the provided input to perform a predefined task. For example, in a mask-based denoising / speech enhancement scenario [References], D1 can serve as a DNN mask estimator that estimates the mask to obtain an enhanced representation of the audio input features—preserving only the speech component features while eliminating non-speech components.
[0005] The output of D1 is then provided as input to the output post-processor (block P2). The output post-processor block primarily contains final processing operations to obtain the final feature representation of the output audio signal, and the inverse operations of the transform / feature extraction applied in P1 to obtain the final audio output. For example, in the noise reduction / speech enhancement application scenario described above, block P2 would contain a masking operation, where the mask estimated from D1 is applied to the original feature representation of the input signal to obtain an enhanced feature representation, and a final inverse transform operation is performed on it to generate the output audio waveform.
[0006] Input preprocessing block P1 is labeled 10, DNN processing block D1 is labeled 20, and output postprocessing block P2 is labeled 30. An excessively high dynamic range of input features can negatively impact the learning process of a deep learning system. For example, in DNN-based audio processing systems, this effect is visible as the simultaneous occurrence of strong and weak sounds in the audio signal. In DNN-based image processing systems, this effect is, for example, visible as the simultaneous occurrence of bright and dark areas in the same image. Various input signal normalization techniques / forms have been applied in existing literature to mitigate this problem. A specific input value compression method called power-law compression has also been used. Generally, power-law compression can be applied to any type of input signal (such as audio signals, image signals, radar signals).
[0007] The following text will use a general signal model to represent input signals of any type, and its expression is:
[0008]
[0009] Where s is a vector containing multiple K input signals, which can be represented by the real and imaginary signal components si respectively. re and s im , or amplitude and phase components s abs and s ph For example, when representing the audio signal of one or more microphones in the time-frequency domain, s contains K elements, corresponding to the K microphones respectively, where the real and imaginary parts (or amplitude and phase) are the real and imaginary components (or amplitude and phase components) obtained by time-frequency transformation (such as short-time Fourier transform) of the microphone signal. Taking image representation as an example, vector s may contain 3 elements for each pixel, such as representing RGB values. In this case, the imaginary part (or phase component) is usually zero unless the image has undergone processing such as Fourier transform. Based on this signal model, the i-th element of k can be represented as
[0010]
[0011] In existing research, power-law compression is mainly applied only, for example, to the amplitude spectrum of calculated frequency characteristics, i.e.
[0012] C k =|S k | α =|S k,re +jS k,im | α =S k,abs α
[0013] Where S k Let be the k-th element of s, α be the power-law factor, and C be the k-th element of s. kThis is the compressed signal corresponding to the k-th input signal. Typically, the compressed signal vector c and its elements Ck are... k It will use DNN post-processing to generate a compressed signal vector d and its elements D. k Subsequently, just before the final feature inversion step, power-law compression is reversed to obtain the final output feature representation, i.e., vector t, whose expression is:
[0014] T k =|D k | 1 / α
[0015] Where T k It is the k-th element of the final output feature vector t. Note that in the above equation, only the amplitude of the input (and output) signals is affected by compression (and expansion), while the power-law factor remains constant.
[0016] In this study, power-law compression is used to compress / limit the dynamic range of each independent component of the input feature values, aiming to improve the learning and generalization ability of DNNs. Specifically, power-law compression is applied to the complete input feature representation, rather than just a portion of the input (such as amplitude components). Mathematically, the compressed signal after applying this power-law compression scheme can be expressed as:
[0017]
[0018] Where C k,re C is the compressed real part of the k-th input signal. k,im To compress the imaginary part, C k,abs For compression amplitude, C k,ph This is a compressed phase.
[0019] Technical issues
[0020] Overall, although neural network applications are becoming increasingly prevalent in this field, the complexity of their processing flow still plays a significant role not only in embedded systems and dedicated hardware systems, but also in non-embedded systems. This is because the complexity of neural networks translates into high processing resource requirements, high energy consumption, and related issues, such as processor resource configuration, heat dissipation requirements for electronic processors, and, especially in scenarios with limited computing resources, the need to meet strict timing requirements.
[0021] Technical solutions
[0022] The present invention aims to provide an improved information signal processing scheme.
[0023] This objective is achieved by the following methods: the information signal processing apparatus according to claim 1, the information signal processing method according to claim 15, or the computer program according to claim 16, wherein preferred embodiments are claimed in the dependent claims.
[0024] According to a first aspect of the invention, an apparatus for processing information signals such as audio signals, image signals, or radar signals includes a feature extractor for extracting a feature set having a first dimension from the signal. The apparatus further includes a feature segmenter for segmenting the feature set into a first feature subset and a second feature subset, wherein the first feature subset has a second dimension, the second feature subset has a third dimension, and both the second and third dimensions are lower than the first dimension. Furthermore, the first and second feature subsets overlap, such that one or more features in the feature set are simultaneously included in both the first and second subsets.
[0025] The device also includes a neural network processor configured to process a first subset of features using a first neural network to obtain a first result, and to process a second subset of features using a second neural network to obtain a second result. The results are combined by a feature combiner, which combines the first and second results using a third neural network, wherein the complexity of the third neural network is lower than the first complexity of the first neural network or the second complexity of the second neural network, and the output of the feature combiner is a feature result set whose dimension is greater than the second or third dimension, and preferably equal to the first dimension. Furthermore, an output post-processor is provided to post-process the feature result set to obtain a processed information signal.
[0026] According to a first aspect of the present invention, a processing block for an audio processing method based on a low-complexity deep neural network (DNN) is obtained. In a preferred embodiment, computational complexity is given priority, i.e., computational complexity is prioritized over storage complexity. Furthermore, in a preferred embodiment, the apparatus, method, and computer program are presented within the context of neural network-based noise reduction solutions. However, the innovative process according to the first aspect (and the second and third aspects described below) is also applicable to other neural networks or deep neural network-based solutions that require reconstruction of the information signal at the output when processing audio or other information signal processing problems. For audio signal applications, such applications may include dereverberation processing, echo cancellation processing, speech or audio coding applications, bandwidth extension applications, etc. In the field of image signal processing, the processes of the first, second, and / or third aspects of the present invention can improve the performance of image enhancement, edge sharpening enhancement, image recognition, and other image processing applications. Furthermore, radar signal analysis (such as target detection and localization) also relies on the analysis of the characteristics of different frequency bands of the signal.
[0027] Typically, the complexity or computational complexity of a neural network can be quantified by various metrics, such as floating-point operations (FLOPS), where a high number of FLOPS indicates high complexity; execution time on one or more target hardware, where a high execution time indicates high complexity; power consumption of a specific device, where a high power consumption indicates high complexity; or the number of MAC (multiply-accumulate) operations, where a high number of MAC operations indicates high complexity.
[0028] According to a second aspect of the invention, the focus is on the manner in which the extracted features are presented to a signal processor (e.g., comprising one or more neural networks). The information signal processing apparatus of the second aspect of the invention includes a feature extractor for extracting a set of features from an information signal. The feature extractor includes a raw feature calculator for calculating raw feature results, wherein each raw feature result contains at least two raw feature components. Preferred raw feature components include, on the one hand, amplitude and phase, and on the other hand, the real and imaginary parts of the feature or feature result. The feature extractor also includes a raw feature compressor for compressing the dynamic range of at least two raw feature components, thereby obtaining at least two compressed raw feature components for each raw feature result. Thus, in one embodiment, the raw feature compressor compresses not only the amplitude of the raw feature result but also its phase. Alternatively, in another embodiment, the raw feature compressor compresses not only the real part of the raw feature result but also its imaginary part. At least two raw feature components of each raw feature result are provided to a signal processor for processing the feature set with the two compressed components to obtain a processed information signal, wherein the information signal may be an audio signal, an image signal, or a radar signal, etc.
[0029] According to a second aspect of the invention, a power-law compression method is preferably employed to compress / limit the dynamic range of each independent component of the input feature value, aiming to improve the learning and generalization efficiency and performance of the deep neural network. Therefore, according to the second aspect, the preferred power-law compression should apply to the complete input feature representation, i.e., to all multiple components rather than only a portion of the input (e.g., only the amplitude component, i.e., a subset of multiple components).
[0030] This fully dynamic compression ensures that every representation of the feature set can be processed, i.e., real / imaginary representations can be processed when necessary, or amplitude / phase representations can be processed on demand. This allows for the simultaneous processing of two representations, and experiments show that the neural network processor achieves better complexity and performance compared to compressing only the amplitude without processing the phase.
[0031] According to a third aspect of the invention, multi-level processing is performed in a neural network processor. The information signal processing apparatus of the third aspect of the invention includes a feature extractor for extracting a feature set from an information signal, wherein each feature in the feature set comprises at least two feature components, and a first feature component is more important than a second feature component. Furthermore, the feature set comprises a first subset having the first feature components and a second subset having secondary feature components. The apparatus also includes a neural network processor comprising a first neural network for receiving the first subset as input and outputting a processed first subset. The neural network processor further includes a combiner for combining the processed first subset with the second subset to obtain a combined subset. This combined subset is preferably input to a second neural network that receives the combined subset as input and outputs a processed combined result. The processed combined result represents a processed information signal, or the apparatus is configured to compute a processed information signal using the processed combined result. In particular, the complexity of the first neural network processing the first set of important features is greater than the complexity of the second neural network processing the second set of secondary feature components.
[0032] Therefore, according to the third aspect concerning various applications related to deep neural network-based signal processing, input features are decomposed or divided into multiple components. Furthermore, not all components are equally important in the context of the target task. Therefore, the DNN-based processing flow (such as a neural network processor) is decomposed into multi-stage processing, where each stage processes components of decreasing importance sequentially. Specifically, each subsequent stage processes new but less relevant or less important input information, resulting in a decreasing trend in the computational complexity of the neural network or DNN block in this design across stages.
[0033] Preferably, all aspects can be combined with each other, but they can also be implemented independently, or only two aspects can be combined depending on the specific circumstances. When the first and second aspects are combined, power-law compression will be applied to the feature set of the input feature segmenter. Therefore, in the audio signal processing apparatus according to the first aspect, the feature set further processed by the feature segmenter and subsequent elements has its feature components compressed not only for a single dynamic range, but also for two or more components (such as amplitude and phase, real part and imaginary part) through multiple compression.
[0034] In this case, the feature decompression operation needs to be performed at any stage before or after the feature combiner processing. Finally, the output post-processor performs secondary processing on the decompressed feature result set to generate the processed information signal.
[0035] In a scheme combining the first and second aspects, the third aspect can also be implemented in either the first or second neural network that processes the first and second feature subsets.
[0036] Therefore, the present invention implements a three-pronged approach: the feature segmenter employs compressed features, and at least one of the first or second neural networks operates in a multi-level low-complexity mode—in which the neural networks with progressively decreasing complexity process features of progressively decreasing importance.
[0037] However, the first aspect of the present invention can also be implemented using the multi-level processing method of the third aspect, but it is not necessary to follow the second aspect to apply power-law compression to both components simultaneously.
[0038] In other embodiments, the second aspect of the present invention may be implemented in conjunction with the third aspect or independently, but does not include the first aspect involving overlapping feature segmentation and the corresponding feature combination of a low-complexity neural network. In such a process, the feature set can still be segmented along the channel direction, but overlapping segmentation is not performed. Of course, segmentation along the channel direction can also adopt an overlapping segmentation method.
[0039] Furthermore, the third aspect regarding multi-level low-complexity neural network processing can also be implemented without feature segmentation, segmentation along the channel direction, or even power-law compression (depending on the specific situation). However, for applications processing medium-sized feature sets, it is still recommended to combine the second and first aspects. For example, in audio applications, when the sampling rate is relatively low, the feature set size is usually small, i.e., it falls into the medium-sized category. A typical implementation is feature extraction using time-frequency decomposition, which relies on real-valued or more preferably complex-valued filter banks (such as QMF or FFT-based filter banks) to generate time-frequency representations of information signals such as audio signals or radar signals. In image signals, feature extraction can employ spatial filter banks or spatial Fourier transforms to generate a one-dimensional or two-dimensional representation of the image using cosine basis functions represented by amplitude and phase (or real and imaginary parts). Attached Figure Description
[0040] Preferred embodiments of the present invention will be disclosed subsequently in conjunction with the accompanying drawings, wherein:
[0041] Figure 1a It is an apparatus or method for processing information signals according to the first aspect;
[0042] Figure 1b It is based on the information signal processing apparatus or method of the second aspect;
[0043] Figure 1c shows an information signal processing apparatus or method according to the third aspect;
[0044] Figure 2 This is a schematic diagram of an audio processing flow based on a typical deep neural network (DNN);
[0045] Figure 3a This is an embodiment of the feature extractor and feature segmenter based on the first aspect;
[0046] Figure 3b This is an embodiment of the feature combiner and output post-processor based on the first aspect;
[0047] Figure 4 The feature segmenter and subsequent channel-level subband segmentation are implemented according to the preferred embodiment of the first aspect;
[0048] Figure 5 This is a preferred embodiment of the neural network processor according to the first aspect;
[0049] Figure 6 A schematic diagram illustrating the process of dividing the feature set into a first feature subset and a second feature subset;
[0050] Figure 7 This demonstrates the state of the second feature subset before and after it is segmented into multiple segments along the channel dimension;
[0051] Figure 8 An implementation scheme for a feature combiner is shown, including a stacker for performing channel-level stacking and subsequent frequency subband combining;
[0052] Figure 9 Showing Figure 1a The processing results of the second neural network before and after segmentation and stacking;
[0053] Figure 10 Showing Figure 10 The first and second feature subsets after stacking the channel-level subbands on the left, and Figure 10 The input on the right is a combination of the first and second results when fed into the third neural network;
[0054] Figure 11 This is a preferred embodiment of the neural network processor according to the third aspect;
[0055] Figure 12 This is another preferred embodiment of the neural network processor according to the third aspect;
[0056] Figure 13 It is Figure 1c. Figure 11 or Figure 12 A preferred embodiment of the combiner shown;
[0057] Figure 14a It is based on the second aspect Figure 1b A preferred embodiment of the original feature compressor;
[0058] Figure 14b It is based on the second aspect Figure 1b A preferred embodiment of the signal processor section;
[0059] Figure 15a This is a preferred embodiment of the invention combining the first and second aspects;
[0060] Figure 15b This is an explanation Figure 15a Examples and Figure 16 A table showing the complexity of the examples;
[0061] Figure 16 This is a preferred embodiment of the second aspect or a combination of the second and third aspects of the present invention;
[0062] Figure 17 This is another preferred embodiment implemented according to the second aspect or the second and third aspects of the present invention; and
[0063] Figure 18 To display in a table Figure 16 Implementation Examples and Figure 17 The MFLops complexity of the two embodiments.
[0064] Embodiments of the present invention
[0065] Figure 1a An audio signal processing apparatus or method according to a first aspect is illustrated. The apparatus includes a feature extractor 100 for extracting a feature set having a first dimension from an information signal (such as an audio signal). The feature set having the first dimension is input to a feature segmenter 200 for segmenting the feature set into a first feature subset and a second feature subset, wherein the first feature subset has a second dimension, the second feature subset has a third dimension, and both the second and third dimensions are lower than the first dimension. More subsets (such as three or four groups) can also be used. Furthermore, the first and second feature subsets overlap, such that one or more features in the feature set are simultaneously contained in both subsets. The first and second subsets of features are input to a neural network processor 300. This processor is configured to process the first feature subset through a first neural network 310 to obtain a first result, and to process the second feature subset through a second neural network 320 to obtain a second result.
[0066] The first network and the second network can be two different or independent networks, or they can constitute different parts of the same overall neural network. These parts can be independent of each other or share one or more parts of the overall neural network, provided that there are parts in the overall neural network that are dedicated only to the first neural network 310 or the second neural network 320.
[0067] Both the first and second results are input into feature combiner 400, which... Figure 8The third neural network 420 shown integrates the two. The complexity of this third neural network is lower than that of the first or second neural network. The result processed by the feature combiner 400 is a feature result set, the dimension of which is greater than the second or third dimension, and preferably equal to the dimension of the feature set obtained by the feature extractor 100. This feature result set is input to the post-processor 500 for post-processing to obtain the processed audio signal.
[0068] In a preferred embodiment, the feature extractor 100 includes a time-frequency decomposer for generating a time-frequency decomposed signal from a time-domain represented audio signal or radar signal, the signal comprising a series of time frames, each time frame having several frequency bands. Furthermore, the feature extractor 100 includes a feature set builder for constructing a feature set from the time-frequency representation. Figure 6 The left side shows a time-frequency representation containing only six time frames and 12 exemplary frequency bands, although more frequency bands may be used in real-world applications (see discussion below). Figure 6 The example presents the state after processing by the feature builder, as it extracts six time-frame sequences from a larger time-frequency representation to obtain the feature set. Through the operation of the feature segmenter, the following can be obtained: Figure 6 The first feature subset is shown in the lower right part, and the second feature subset is shown in the upper right part. Figure 6 The overlapping range shown includes two frequency bands between 8kHz and 12kHz, but it should be emphasized that this is only an example. In practical applications, a sampling rate of 48kHz can generate frequency features up to 24kHz, in which case a single time frame will contain approximately 769 frequency features.
[0069] To achieve the reverse processing of feature extraction, a frequency-time synthesizer is used as... Figure 1a The function block of the output post-processor 500 can synthesize a processed time-domain representation information signal from the input time-frequency representation feature set. In one embodiment, the input feature set is a feature result set; or alternatively, the feature set input to the frequency-time synthesizer / frequency-time converter is composed of the feature result set and the feature set (i.e., the input to...). Figure 1a Features of feature extractor 100 are jointly derived.
[0070] For example, when this invention is used to calculate the mask or time-frequency mask required for noise reduction or speech enhancement, the feature input set to the time-frequency synthesizer is not the mask itself, but rather the result of applying the mask to the original time-frequency representation generated by the feature extractor 100. The time-frequency mask can be applied as a spectral gain to the input signal of the feature extractor 100, thereby achieving noise reduction or speech enhancement. Furthermore, the time-frequency mask can also be used to estimate the power spectral density matrix of the noise or speech signal, and then calculate one or more intelligent spatial filters and apply them to the input signal.
[0071] However, when processing for purposes such as bandwidth extension, the result of neural network processing—that is, the feature set obtained by the feature combiner—can serve as an extended audio signal within the bandwidth extension range, or a complete audio signal (containing both baseband and extended frequency bands), while still maintaining a time-frequency representation. In this case, the output post-processor 500 only needs to perform frequency-to-time conversion of the bandwidth-extended audio signal. However, when neural network processing is used to derive bandwidth extension parameters from the input signal, these parameters need to be applied to the input signal to ultimately generate a bandwidth-extended output signal.
[0072] Besides noise reduction, speech enhancement, and bandwidth expansion, other applications include source separation processing. For example, when the input signal contains information from multiple sound sources at different locations, the input signal can be composed of multiple microphone signals. In this context, the feature result set can still contain frequency masks: these can be in binary form—that is, the time-frequency mask corresponding to the dominant sound source has a value of 1 at a specific frequency point, while the masks for the other sound sources have a value of 0; or they can be in soft mask form—each entry in the time-frequency mask at a specific frequency point represents the probability that the corresponding sound source is dominant at that frequency point. As mentioned earlier, the time-frequency mask can be applied as a spectral gain to a microphone signal to achieve sound source separation. However, the time-frequency mask can also be used to estimate the power spectral density (PSD) matrix of different sound source signals, and then calculate the intelligent spatial filter for blind source separation. Therefore, depending on the specific implementation scheme, the feature result set can be applied to the information signal and then converted to the time domain; or the feature result set itself is a frequency domain processing signal, in which case the output post-processor only performs the conversion processing from the frequency domain representation to the time domain representation.
[0073] To implement masking, the output post-processor 500 includes a masking processor that can apply a time-frequency mask as a spectral gain mask to a feature set in a time-frequency representation to obtain an input feature set; or calculate a processing filter using a time-frequency mask and apply the filter to an audio signal or feature set to obtain an input feature set.
[0074] In another embodiment, the first neural network is more complex than the second neural network, and the first feature subset contains information about the lower frequency bands of the audio or radar signal. Furthermore, the second feature subset contains information about the higher frequency bands of the audio or radar signal. Therefore, the more complex neural network is used to process the audio or radar signal portion (i.e., the low-frequency range), which is generally more important for the perception or processing intent; while for the relatively less important portions (e.g., the high-frequency range of the audio signal), a less complex neural network is used. Figure 1aThe second neural network 320, with lower complexity, is used for processing. For radar signals, the processing approach may differ when a specific target to be identified or located corresponds to a specific frequency range within the radar signal. For example, a target with a specific velocity will generate Doppler frequencies in a specific frequency band, which can be processed with higher complexity compared to other frequency bands (corresponding to velocities not of primary concern to the surveillance mission).
[0075] In a preferred embodiment, the feature segmenter 200 is configured to generate a first feature subset and a second feature subset, such that the second dimension of the first feature subset is lower than the third dimension of the second feature subset. Therefore, by processing the low-dimensional first subset through a more complex first neural network 310, a large amount of processing power is concentrated on a specific part of the information signal, thereby optimizing the enhancement effect of a specific frequency band (such as a low-frequency band) as much as possible under limited resources.
[0076] like Figure 6 The feature segmentation shown can be enhanced by further segmenting the second feature subset into multiple sub-segments. Furthermore, as... Figure 7 As shown on the left, the second multi-segmentation obtained through three segmentation windows 1, 2, and 3 can be arranged along the channel dimension (e.g., Figure 7 (As shown on the right), this increases the number of input channels to the second neural network 320. However, it should be noted that while the number of channels increases, the frequency or feature input dimension will decrease accordingly.
[0077] In another embodiment, not only is a second subset covering the high-frequency portion of the audio signal segmented, but the first feature subset is also segmented into multiple initial segments, and these initial segments are arranged along the channel direction, thereby similarly increasing the number of input set channels fed into the first neural network. Figure 10 As shown in the illustrated embodiment, the first subset is divided into two segments, and the second subset is divided into three segments. Therefore, the number of second multi-segments is three, and the number of first multi-segments is two. In other embodiments, depending on the specific circumstances, only the second subset of features (i.e., Figure 7 Performing channel-level band segmentation on the high-frequency range shown, while not performing such segmentation on the first subset (e.g., the low-frequency range of the audio signal), is also of practical value. In this case, Figure 10 The lower area will appear as a single shade of green. Preferably, such as Figure 7 As shown, the segmentation windows that divide the feature subset into corresponding first or second multiple segments should preferably adopt an overlapping design, especially a 50% overlap rate, but a smaller overlap width can also be used. Figure 6 The frequency subband segmentation 200 shown also uses overlapping processing, and its overlapping range accounts for only one-third of the total width of the subset. Figure 6The preferred scheme is specifically shown: the first feature set covers the frequency band of approximately 0 to 12 kHz (i.e., approximately 12 kHz bandwidth), while the second feature subset (covering a higher frequency band) covers the frequency band of 8 kHz to 24 kHz (i.e., 16 kHz bandwidth).
[0078] like Figure 4 As shown, the corresponding output processing of channel-level frequency subband segmentation 301 or 302 is... Figure 8 The channel-level frequency subband stacks 401 and 402 are shown. Furthermore, when channel-level frequency subband segmentation and frequency subband segmentation 200 are applied simultaneously, the two types of features need to be restored at the output. Therefore, the first embodiment of this invention employs a third neural network, whose complexity is lower than the initial complexity of the first neural network or the secondary complexity of the second neural network, and preferably, the complexity of the third neural network is even lower than... Figure 5 The first neural network 310 and the second neural network 320 are shown.
[0079] According to the first aspect, the merging operation is achieved by simply stacking the first and second results, and if necessary, stacking the individual channels of the results to obtain a set of stacked features. Due to the overlapping range of frequency subband segmentation and the preferential overlapping range of channel-level frequency subband segmentation, the dimension of the input to the third neural network will be significantly higher than the dimension of the feature set, i.e., the first dimension output by the feature extractor 100. However, since the third neural network only needs to eliminate the dimensionality increase, and only needs to focus on processing features within the overlapping frequency range for the feature set (features within the non-overlapping frequency range usually do not require deep processing by the network), the network can still maintain low complexity even if the dimension of the input features increases. Therefore, not only is the complexity of the third neural network preferably lower than that of the first and second neural networks, but more preferably, the number of layers of the third neural network is less than that of the first and second neural networks, especially its number of layers does not exceed 1 / 3 of the number of layers of the first / second neural networks, more preferably only 1 / 4, and most preferably only 1 / 10.
[0080] A preferred embodiment in this regard is to introduce a feature redirection block in P1 after the feature calculation step. Figure 3a The CR1 error in the code is not found. (Reference source not found.) Computational complexity is reduced by employing efficient DNN model block hyperparameter settings.
[0081] After the DNN processing block (D1) processes the input signal, the feature redirection step is executed in reverse to restore the DNN output to the original input feature dimensions (e.g., ...). Figure 3b As shown in CR2), the output audio waveform is finally obtained through the reverse feature extraction step.
[0082] As mentioned above and Figure 2As shown, preprocessing block P1 contains a feature extraction / computation step, where the audio input signal is converted into a representable form (e.g., a two-dimensional or three-dimensional real tensor). For example, when using STFT feature extraction, the feature is represented as a two-dimensional complex tensor, which can be represented as a three-dimensional real tensor of size N×K×2, where N corresponds to the number of frames, K corresponds to the number of frequency bands (bins), and 2 corresponds to the real-imaginary or amplitude-phase components. If the feature extraction step uses a learned filter bank / feature, the feature representation is typically a two-dimensional real tensor of size N×K, where K corresponds to the feature dimension and N corresponds to the number of time frames.
[0083] To learn the local structure of feature representations, deep neural network models for audio processing typically employ convolutional layers with small kernels in their initial stages. Due to the use of multiple such layers, the computational complexity of the model is primarily influenced by the spatial dimensions N and K of the input features. Alternatively, other types of layers, such as fully connected layers or recurrent layers, can also be used in the initial stages of DNN models; in these cases, the main factors affecting complexity are still N and K.
[0084] The number of time frames N depends on the window / frame length set by the transform / filter bank / feature extraction used in signal processing, as well as the length of the input audio signal. In the real-time audio application focused on in this paper, the audio segment processed each time is typically less than 50 milliseconds, so the value of N for each processing is very small. Therefore, this invention mainly focuses on the frequency / feature dimension of the input features. The size of the feature / frequency dimension K depends on the required frequency / feature resolution and the sampling rate of the audio signal. For speech or audio processing tasks, the feature / frequency dimension can be chosen as a power of 2 and greater than 128. For higher sampling rates such as 32 or 48 kHz, the feature dimension is usually very high (e.g., 512, 1024, etc.). When the feature / frequency dimension is too high, in order to reduce the computational complexity of the main deep neural network model, this invention combines two feature redirection techniques to achieve complexity optimization (e.g., Figure 4 (As shown). These techniques will be described below using deterministic transforms / filter banks (such as short-time Fourier transforms) as feature extraction steps. Although the following explanation uses the frequency dimension as the feature dimension, these methods are equally applicable to any feature representation.
[0085] Subband splitting
[0086] The first technique corresponds to the frequency subband segmentation method, where the input features are segmented along the frequency dimension K into multiple overlapping segments [s1, s2, ..., s]. M Some academic studies have proposed basic non-overlapping segmentation approaches to improve the performance of speech enhancement methods based on deep neural networks [1, 2]. This report proposes to use segmentation methods as a means to reduce complexity and extend them to overlapping segmentation schemes.
[0087] After the original input features are segmented, each feature segment is fed as an independent input into its corresponding DNN model for processing. Therefore, block D1 in Figure 1 is no longer a single DNN model, but is composed of multiple independently running DNN models (e.g., ...). Figure 5 As shown, different frequency subband segments generated by the subband segmentation technique are processed respectively.
[0088] Figure 6 This paper demonstrates a segmentation technique and an embodiment for generating corresponding independent DNN input segments. This embodiment uses an audio signal with a sampling frequency of 48kHz. Calculating the STFT features yields a feature representation, where the frequency dimension represents signal content up to the Nyquist frequency of 24kHz. Using this segmentation method, the frequency dimension is divided into two parts: a low-frequency feature representation (0-12kHz) and a high-frequency feature representation (8-24kHz). In this example, these two frequency bands are respectively provided as input to two different DNN models in block D1.
[0089] • Sub-band segmentation can be arbitrarily selected. Parameters can be flexibly set according to the required frequency dimension size of each DNN input, the number of independent DNNs to be used, the number of segments M, and the frequency dimension size of each segment.
[0090] • The frequency dimensions of different frequency bands do not need to be the same. The specific size of each frequency band can be determined according to the audio processing task. For example, in a speech processing task, if the Short Time Fourier Transform (STFT) is used as the feature representation, the size of the first frequency band can be determined based on the typical frequency range of speech activity. The sizes of the remaining segments can be determined based on the number of deployable DNNs. Alternatively, equivalent rectangular bandwidth (ERB), Mel scale, or similar perceptual scales can be used for frequency band segmentation, depending on the task requirements of the method.
[0091] • The overlapping design between segments can avoid artifacts in the boundary regions of the segmented feature representation when the final audio output is merged.
[0092] • Partitioning the input along the frequency dimension facilitates the use of independent DNNs with lower computational complexity compared to large, single DNN models. In a single model, the massive input space and extremely high computational cost per layer (especially the initial layer) result in a persistently high overall system complexity. Partitioning the input along the frequency dimension keeps the computational complexity of the initial layer sufficiently low, thus maintaining a lower overall system complexity. This flexible partitioning strategy makes the overall complexity of independent DNN models lower than that of single DNN models in the overall system design.
[0093] • On hardware platforms that support parallel processing, the overall system computation time can be further reduced by processing each independent DNN in parallel.
[0094] Channel-level sub-band segmentation
[0095] This method assumes that the initial layer of each DNN model ( Figure 5 Convolutional layers with small kernels are used to learn the local structure of audio input feature representations. This characteristic is generally applicable to most audio processing DNN models—in various tasks, utilizing / learning the local structure of audio input feature representations is crucial for improving performance.
[0096] In this method, the independent input segments obtained after subband segmentation are further segmented into multiple overlapping segments with equivalent spatial dimensions, arranged along the channel dimension. The idea of channel dimension segmentation has been proposed earlier in [3, 4], and this method can achieve more efficient feature extraction compared to performing convolution operations on the entire input space. This invention combines this concept with the subband segmentation method as a second-stage process to reduce complexity. As mentioned earlier, subband segmentation keeps the computational complexity of the early convolutional layers low by reducing the input space (NxK→NxK1, NxK2, ..., where K1, K2....<K). Channel segmentation further extends this advantage by reducing the spatial dimension (NxK1, NxK2, ...) of the input space for each independent input.
[0097] by Figure 6 Taking the segmentation block on the right as an example: the segmentation block corresponding to the high-frequency feature can be further divided into three uniformly overlapping sub-segments. After each sub-segment is arranged along the channel dimension, the final input feature representation of the independent deep neural network can be obtained.
[0098] The overlap factor for segmentation can be set arbitrarily, but the spatial dimensions (number of frames and frequency band size) of each segmented block must be consistent to ensure they can be arranged along the channel dimension. In typical convolutional layer operations, small-sized filters are first applied to each channel, and then the output is generated by weighted combination of channel dimensions. Therefore, channel dimension segmentation not only helps to integrate cross-frequency band information in the initial layer of the DNN, but also reduces the spatial dimension of the input, making it easier to traverse small-sized filters, thereby reducing the computational complexity of the initial layer.
[0099] - The overlap factor can be determined based on the required reduction in the dimensionality of the input space and the number of channels. For example Figure 7 In order to reduce the number of input channels of a DNN, the input can be divided into three segments with equal spatial dimensions, thereby reducing the computational complexity of the first convolutional layer.
[0100] Merging methods
[0101] After each DNN processes the input, the output of each DNN in D1 yields an enhanced / modified representation of the input features. Subsequent tasks require recombination / merging of these outputs to obtain the original dimensional output features, which are then used as input for the feature inversion step (inverse STFT in this example), ultimately generating the output audio signal.
[0102] In most previous studies employing subband splitting [1, 2] or channel-level subband splitting [3, 4, 5], the merging method typically involves concatenating the outputs of different subbands along the frequency axis to obtain the final output. This is generally feasible because these methods do not consider overlapping splitting strategies. In one study [5], a learnable approach was adopted: the concatenated outputs were processed through a two-stage enhancement model using independent speech enhancement models. In this case, the complexity of the second enhancement model is comparable to that of the first stage.
[0103] This invention employs a simple concatenation process based on the frequency dimension, followed by learning the optimal combination / merging scheme for cross-frequency information through a small neural network, while also considering the overlap between segments.
[0104] Two merge blocks in the overall merge method, such as Figure 8 As shown.
[0105] Channel-level subband merging
[0106] • The output dimension of each deep neural network is unrestricted; its spatial dimension and channel dimension can be arbitrarily set and are all different.
[0107] • The output channels are stacked along the feature / frequency dimension to form the input for subsequent merged blocks. Note: The overlap relationships considered in the segmentation stage are not copied in the stacking step.
[0108] ·Reference Figure 7 The segmentation and merging steps are as follows: Figure 9 As shown. Note that the frequency dimension of the output after merging is always large.
[0109] =or equal to the original input ( Figure 7 The frequency dimension was not considered during merging because the overlap between segments was not taken into account.
[0110] Frequency subband merging
[0111] In the frequency subband merging step, the outputs of each channel after the previous merging step are stacked, and the operation is similar to that described in the channel-level subband merging step.
[0112] Figure 10 The process of merging the output signals of the two sub-bands is demonstrated.
[0113] After the initial merging is completed, a small neural network is used as the final merging step to generate the output feature representation of the target dimension—which can be restored to the original dimension of the system input or converted to a different dimension (e.g., a dimension smaller than the original dimension).
[0114] • The neural network used is not limited to a specific architecture block. It can be a multilayer perceptron (MLP) network with only a few fully connected layers, or a neural network with a combination of convolutional layers, recurrent layers and fully connected layers.
[0115] The only limitation of neural network design is that its final layer must rearrange the frequency dimension of the output signal to be equivalent to the frequency dimension of the original input to the audio processing system. This operation is crucial for the inverse transform step. Alternatively, when a classifier is present, the final layer can adjust the output spectral dimension to a preset value (e.g., smaller than the spectral dimension of the original input to the audio processing system). Example
[0116] 1. In the first embodiment, all the processing blocks described in this invention can be used to design an audio processing system based on a low-complexity deep neural network.
[0117] 2. In the second embodiment, all processing blocks except for the subband segmentation block can be used to design an audio processing system based on a low-complexity deep neural network. When the feature dimension is small (e.g., less than 512), the subband segmentation step can be ignored.
[0118] 3. Other embodiments may include all possible combinations of the various processing blocks described herein.
[0119] Subsequently combined Figure 1b and Figure 14a , 14b The preferred embodiment of the second technical solution is described, for example, using power-law compression to generate compressed components.
[0120] An apparatus for processing information signals, which may be audio signals, image signals, radar signals, etc., and its basic components include: Figures 1a to 10 The feature extractor 100 is shown. Specifically, the feature extractor 100 is configured to extract a feature set from the information signal, which includes a raw feature calculator 120 for calculating the raw feature results—as follows. Figure 1b As shown below block 110, each original feature result contains at least two original feature components. One original feature component of the original feature result can be the amplitude, and the other original feature component can be the phase of a complex amplitude (e.g., a cosine function). Alternatively, the first component can be the real part, and the second component can be the imaginary part of the corresponding original feature result.
[0121] These two components (e.g., amplitude and phase, or real and imaginary parts) are input into the raw feature compressor 120 to compress at least two raw feature components, thereby generating at least two compressed raw feature components for each raw feature result, such as... Figure 1b As shown by the two arrows below block 120. Therefore, the feature set output by the feature extractor contains the compressed original feature components corresponding to each original feature result.
[0122] These raw feature components are input to the signal processor 600 in compressed form to process the feature set containing the compressed raw feature components in each raw feature result, thereby obtaining the processed information signal. The signal processor 600 may include, for example: Figure 1a or Figure 5 One or more neural networks as described in the first aspect may also include, for example: Figure 2 The single neural network 20 shown, or containing elements as shown in Figure 1c, or Figures 11 to 13 The cascaded neural network described in the third aspect is shown. The signal processor may also include a feature segmenter, a feature combiner, and / or a post-processor, such as... Figure 1a As shown in the relevant description.
[0123] on the other hand, Figure 1a Feature extractor 100 shown Figure 3a The feature extractor shown can also be pressed... Figure 1b The method shown is the same as the solution described in the second aspect.
[0124] The original feature calculator 110 is preferably configured to calculate complex values as original feature results, the original feature results including real part and imaginary part or amplitude and phase as at least two original feature components; the original feature compression unit is configured to perform compression operation on real part and imaginary part or amplitude and phase to obtain compressed original feature components.
[0125] The raw feature compressor 120 is further configured to apply a first compression function 121 to the first raw feature component and a second compression function 122 to the second raw feature component, wherein the second compression function may be different from the first compression function. Preferably, the raw feature compressor is configured to employ power-law compression, using a power value 1 / α, where α is a fixed value or a specific value learned in a particular application, and α is in particular a real number or integer greater than 1. Alternatively, the power value is α. In this case, the power value α is a fixed value or a value learned in a particular application, and α is in particular a real number or integer less than 1.
[0126] Preferably, the value α is different for each original feature component, and the value of α is particularly large for original feature components (such as amplitude and phase) that are expected to have a higher numerical range. Preferably, the compression function has a stronger compression effect on amplitude than on phase; however, when the real and imaginary parts are expected to have similar numerical ranges, similar compression function strengths can be applied, thereby using similar α values in parallel compression.
[0127] Preferably, the original feature calculator 110 is configured to calculate the absolute value of the original feature components and their corresponding signs, and the original feature compression unit 112 is configured to apply the corresponding compression function to the absolute value while retaining the sign of the original feature components.
[0128] Preferably, the information signal is an audio signal, and the original feature calculator 110 includes a time-frequency decomposer for decomposing the audio signal into a complex time-frequency representation. Alternatively, the information signal is an image signal, and the original feature calculator 110 includes a spatial transformer for decomposing the image signal into a complex spatial frequency representation.
[0129] Preferably, Figure 1b The signal processor 600 shown includes, for example: Figure 1a , 5 or Figure 2 The neural network processor 300, labeled 20, is used to acquire the processed feature set. Furthermore, as... Figure 14b As shown, a feature decompressor 350 is configured to perform decompression matching with the compression operation performed by the original feature compressor 120, thereby obtaining a feature result set. A post-processor 500 is further configured to post-process the feature result set, ultimately generating a processed information signal.
[0130] Specifically, it depends on the specific implementation method, especially on whether it is applied. Figure 1a The feature segmentation 200 and feature combination 400 shown are illustrated. Feature decompression 350 is preferably performed after feature combination 400, that is, the compressed features generated by the neural network processor 300 are processed by a low-complexity third neural network. However, in other embodiments, feature combination can also be performed in the uncompressed domain or the real / imaginary part domain as described. In this case, feature decomposition 350 will be placed between the neural network processor 300 and the feature combiner 400. To illustrate different schemes, the first feature decomposer 350 is marked as an optional component by dashed lines, and the feature decompressor 350 can also be implemented via dashed paths.
[0131] In another preferred embodiment (see below for details) Figure 16 and Figure 17 The functions of original feature compression 120 and subsequent power-law decompression 350 are integrated into a single deep neural network 20 for processing, such as... Figure 2 As shown. Furthermore, in Figure 17 In China, according to Figure 1b The power-law compression feature shown in the second aspect also combines the functions of channel-level subband segmentation 301 and channel-level subband merging / stacking 401 (such as...). Figure 4 and Figure 8 The first aspect of the function is omitted, namely, the subband segmentation 200 and the frequency subband merging function performed in the small post-processing neural network 420 are omitted.
[0132] Furthermore, for audio signals with a sampling rate below 24kHz, a frequency dimension feature set dimension between 200 and 500, and a channel count of 2, a single neural network is preferred over... Figure 5 , Figure 1a or Figure 15a The parallel dual neural network scheme shown is preferred in this case. Figure 2 The neural network 20 shown, or the third aspect of Figure 1c, is as follows: Figures 11 to 13 The aforementioned cascaded neural network.
[0133] Furthermore, in addition to this aspect, it is also recommended to use 2 to 4 frequency bands for channel segmentation in other aspects, so that the number of channels is between 4 and 8, and the overlap range of the channel segmentation frequency bands should be between 30 and 90 frequency bands, preferably 60 to 70 frequency bands. This specification uses power-law compression technology to compress / limit the dynamic range of each independent component of the input feature value, aiming to improve the learning and generalization ability of deep neural networks—that is, to apply power-law compression to the complete input feature representation, rather than only acting on local features such as amplitude components. Mathematically, the signal processed by power-law compression can be expressed as:
[0134]
[0135] Where C k,re C is the compressed real part of the k-th input signal. k,im To compress the imaginary part, C k,abs For compression amplitude, C k,ph This is a compressed phase.
[0136] Typically, different compression functions can be used for each component and the k-th input channel. The compression of a specific component with channel k can be expressed as:
[0137] C k,(·) =f k,(·) (S k,(·) )
[0138] Where f represents any compression function, and the subscript (·) identifies different signal components (such as the real part re or the imaginary part im). After DNN processing, the components of the final output feature representation can be restored through inverse power-law compression, i.e.
[0139]
[0140] Where f -1 This represents the inverse power-law function (extended function). The final output feature representation of the k-channel can be obtained by combining the components, for example:
[0141]
[0142] Where T k It is the k-th element of vector t.
[0143] As an exemplary implementation of the above compression steps, a signal representation of real and imaginary parts can be considered. The compressed components can then be obtained in the following way:
[0144]
[0145] Where (±) represents the eigenvalue S k,(·) The original symbol. As shown in the figure, in this embodiment, the power-law function f k,(·) This corresponds to taking the absolute value of a specific signal component while preserving its symbolic information. Furthermore, power-law compression is applicable to the complete signal representation, meaning it acts on both the real and imaginary parts simultaneously. After processing by a deep neural network, the eigenvalues of the power-law compressed signal can be recovered through corresponding inverse operations, thus obtaining the components of the final output feature representation.
[0146]
[0147] Another exemplary implementation is to use amplitude and phase representation. In this case, the compressed signal components can be obtained as follows:
[0148]
[0149] C k,ph =S k,ph ·α k,ph
[0150] Similar to the previous example, although the power-law compression function also acts on the complete signal representation, here f k,(·) The function handles the two components differently. After processing by a deep neural network, the power-law compression needs to be reversed accordingly, i.e.:
[0151]
[0152] T k,ph =D k,ph / α k,ph
[0153] Finally, the output feature representation is obtained.
[0154] In the proposed power-law compression scheme
[0155] The power-law function f (or power-law factor α) can be:
[0156] ○ Assign different values to each component (e.g., re, im, abs, ph) and element k of the input feature representation.
[0157] ○α This parameter can be set fixedly or learned through training for specific signal processing tasks.
[0158] The third aspect of the present invention will then be described with reference to FIG1c. FIG1c illustrates an apparatus or method for processing information signals, the apparatus including a feature extractor 100 for extracting a feature set from the information signals. Each feature in the feature set comprises at least two feature components, wherein a first feature component is more important than a second feature component, and the feature set comprises a first subset having the first feature component and a second subset having the second feature component.
[0159] For example, a more important component may be the amplitude of the spectral value obtained by time-frequency decomposition of information signals (such as audio signals, image signals, radar signals, or other information signals from which features can be extracted), an extraction process that will produce more important and less important components in a single feature.
[0160] The device also includes a neural network processor 300 equipped with a first neural network 340 for receiving a first subset 332 as input and outputting a processed first subset. Furthermore, the neural network processor includes a combiner 350 for combining the processed first subset with a second subset 331 to obtain a combined subset. This combined subset, now containing processed first and second components for each feature, is input to a second neural network 360. The network outputs a processed combined output, which may optionally be input to an output stage 700. Specifically, the combined output represents a processed information signal, or the device is configured to use the output stage 700 to calculate the processed information signal based on the combined output.
[0161] Furthermore, the complexity of the first neural network 340 is higher than that of the second neural network 360. Therefore, the more important components are input into the more complex neural network 340, while the less important components 331 are input into the less complex second neural network after combination.
[0162] In image processing, the foreground region of an image represents important features, while the background region provides secondary features. Feature extractors process these features, for example, calculating important and secondary features for a specific image range.
[0163] In preferred embodiments (e.g.) Figure 13As shown, combiner 350 concatenates the processed first subset and second subset along the channel direction. For example, the processed first subset (representing the amplitude mask) and the secondary second subset (representing the phase component) are juxtaposed along the channel direction and input together into the second neural network. This network differs from the first neural network in that it uses dual-channel input for computation, while the first neural network 340 only processes single-channel input.
[0164] when Figure 12 DNN in M When performing overlapping segmentation and concatenation along the channel dimension, network 340 needs to ensure that the overlapping structure also exists in the second subset of the input combiner 350, so that the stacking result of network 340 can be combined with the corresponding stacking result of the second subset in combiner 350. Another approach is to process the output of network 350 to eliminate the overlap caused by channel-direction segmentation, matching the second subset of the input combiner 350 with the output of network 350, thereby achieving... Figure 13 or Figure 12 The combined operations shown.
[0165] Typically, later-order neural networks process more channels than lower-order neural networks (i.e., neural networks that process earlier stages in a cascade). Figure 11 The general architecture of the neural network processor 300 is demonstrated, which includes three neural networks: 340, 360, and 380. 340 is the most complex neural network, 360 is of medium complexity, and 380 is of the lowest complexity. Furthermore, when combiners 350 and 370 are always cascaded along the channels (e.g....), Figure 13 As shown, the first neural network 340 processes fewer channels than the second neural network 360, while the last neural network 380 processes the most channels.
[0166] Figure 12 An embodiment containing only two neural networks 340 and 360 is shown: the output of the second neural network 360 is a representation composed of real and imaginary parts, while the combined signal input to the neural network 360 is the amplitude and phase arranged along the channel direction (e.g., Figure 13 (As shown).
[0167] for Figure 16 and Figure 17 The implementation scheme shown prefers to use two cascaded neural networks of different complexities. Figure 15a The subband segmentation 200 and the corresponding frequency subband merging 410 (requiring two parallel neural networks 310 and 320) shown are not processed here. However, as Figure 17 The additional channel-level subband splitting 301 and the corresponding channel-level subband merging or stacking 401 shown are more preferably with Figure 11 and Figure 12 The sequential or cascaded processing shown combines to make Figure 17 The main DNN330 in the middle was able to pass Figure 11 and Figure 12 The method shown is used, that is, the processing flow marked 340 to 380 or 340 to 360.
[0168] In signal processing applications based on deep neural networks (DNNs), input features can be decomposed / divided into multiple components. Furthermore, not all components are equally important in the target task. In such cases, the DNN-based processing flow can be decomposed into a multi-stage process, with each stage processing components whose relevance to the task decreases sequentially. If the relevance / importance of new input information processed in subsequent stages decreases, the computational complexity of the DNN block in this design can also be reduced accordingly.
[0169] Figure 11 A flowchart illustrating this processing flow is provided. The input features can be decomposed into N components, numbered according to task relevance / importance. As shown, the first deep neural network only receives the most important feature component as input, generates an intermediate output, and combines it with the next most important feature. This intermediate output is then passed as input to the next layer of the neural network, which has lower computational complexity. This processing chain iterates through all N components of the input features, and the final output is the result of the Nth layer of the neural network.
[0170] Specifically, noise reduction techniques can be considered, along with short-time Fourier transform (STFT) as the feature extraction step. This approach is also applicable to other signal processing tasks—when the extracted features can be decomposed into two or more independent components, such as in image processing, some enhancement tasks may be better suited to processing the foreground and background of the image separately in two stages.
[0171] Taking the STFT of complex transform as an example, its characteristic representation can be decomposed into two components: amplitude and phase, namely:
[0172]
[0173] Among them, X m and X p These are the amplitude and phase components represented by the complex short-time Fourier transform, respectively. In specific noise reduction scenarios, enhancing the amplitude component is usually more critical. This characteristic is used to further reduce the overall system complexity.
[0174] Given the importance of amplitude components, in Figure 12 The first processing stage shown on the left employs a deep neural network model (DNN). M An amplitude mask is estimated to enhance the amplitude components of the input features. This model has a significantly lower design complexity than models that simultaneously process the complex components of the input features. The calculated amplitude mask is then combined with the phase components of the input features at this stage.
[0175] If the magnitude mask and the phase component have the same dimension, then it can be done as follows: Figure 13 As shown, the two are simply concatenated along the channel dimension. Another approach is to combine the phase and amplitude with the real and imaginary parts respectively, for example:
[0176] Y R =M M *cos(X P ),
[0177] Y I =M M *sin(X P ),
[0178] Where Y R and Y I M represents the real and imaginary parts of the intermediate output, respectively. M Through DNN M The obtained amplitude mask, X P It is the phase component of the original input features.
[0179] After the calculation is completed, the real and imaginary parts can be as follows: Figure 13 As shown, the inputs are connected along the channel dimension to form the input for the next stage.
[0180] In the subsequent processing stage (see...) Figure 12 ), through another DNN model (DNN MP The amplitude mask and phase components of the intermediate combination are processed to obtain the real and imaginary parts of the final enhanced output signal. This signal is then subjected to inverse transform processing to generate the enhanced audio signal. Since the main task of the second stage is to enhance the phase (in conjunction with enhancing the amplitude mask), a process with lower complexity than DNN can be used. M The deep neural network model further reduces the overall system computational complexity.
[0181] Figure 16 The implementation scheme of the processing flow conforming to the second embodiment is shown, including power-law compression 120 and power-law decompression 350. The DNN block can be implemented with reference to any schematic diagram of this application, or it can be implemented using a simplified neural network that does not include subband segmentation, channel-level subband segmentation, and corresponding merging methods (such as channel-level subband merging and frequency subband merging). Figure 16 As shown, this type of DNN employing sequence processing can contain convolutional layers, LSTM, and convolutional transpose layers for noise reduction and related mask estimation in tasks with a 48kHz sampling rate. Specifically, the frequency dimension is 769 per frame, and the number of channels C equals 2, because the input features are complex features, containing real and imaginary parts or amplitude and phase.
[0182] Research found that, depending on the batch size B and the number of selectable time intervals T, this deep neural network has a computational complexity of 499.5 million floating-point operations (MFlops) with a parameter size of 2.67 million. The computational complexity of this DNN architecture mainly stems from the convolutional layers and the transposed convolutional layers—when the convolutional kernel needs to move along a high-dimensional frequency axis (such as...). Figure 16 In the example, F = 769), the computational complexity increases significantly. In this context, the technical solution of the first aspect of the present invention has significant advantages. The subband segmentation method reduces the frequency axis dimension by dividing the frequency axis feature into two overlapping segments; while the channel-level subband segmentation method further uses fixed overlapping segments and stacks them in the channel dimension, significantly reducing the frequency axis dimension of the segmented features. For example... Figure 15a The device shown is obtained through the corresponding frequency and channel number. Figure 15a The display shows that the frequency dimension of the first subset decreases from 769 to 385, and the frequency dimension of the second subset decreases from 769 to 513. It is important to emphasize that the first sub-band has a smaller dimension than the second sub-band, and the sum of the frequency dimensions of the two (385 + 513) reaches 898 frequency bands, far exceeding... Figure 15a The 769 frequency bands shown in block 120. This difference stems from the setting of the overlap range.
[0183] In addition, such as Figure 15a As shown: the low-frequency subband undergoes two subband divisions, while the high-frequency subband undergoes three subband divisions. Figure 15a In the left branch, the number of channels increases from 2 to 4 (corresponding to two segments); in the right branch, the number of channels increases from 2 to 6 (corresponding to three segments).
[0184] The corresponding merging in block 401 produces 257 × 2 = 514 frequency values, while the right side produces 257 × 3 = 771 frequency values.
[0185] After frequency subband merging, the frequency dimension is as high as 1285 and the number of channels is 2. Then, after dimensionality reduction by a small post-processing neural network 422, the final input dimension is 769 (dual channels), which is consistent with block 120.
[0186] The power-law compression in block 350 does not change the dimension, so the output features that can be input to the postprocessor have 769 channel dimensions, corresponding to a two-channel structure of real / imaginary or amplitude / phase representation.
[0187] In the systems of DNN branches 1 and 2 mentioned above, for ease of comparison, the same as... Figure 16 The same DNN architecture as the sequential system. However, Figure 15a The parallel system shown has a computational complexity of 399.5 MFlops, which is greater than... Figure 16The system reduced its MFlops by 100. This is mainly due to the subband segmentation and channel-level subband partitioning methods, which reduced the frequency axis dimension of the input features—a dimension that was a performance bottleneck in sequential systems.
[0188] In practical applications, DNN branch 1 and DNN branch 2 can adopt a ratio of Figure 15a A simpler, independent neural network. This is because, after segmentation, the neural network needs to process far more localized overlapping features than... Figure 16 The neural network in branch 2 can efficiently complete the task even with a simplified architecture. Another reason is that in practical applications, high-frequency bands have lower energy, and compared to the higher-energy low-frequency bands, the contribution of high-frequency bands to the subjective / objective improvement of metrics such as mask estimation in noise reduction tasks is limited. Therefore, the DNN complexity of branch 2 can be lower than that of branch 1.
[0189] Therefore, DNN branch 1 and DNN branch 2 each employ two different DNN models. Their computational complexities are 112.5 MFlops and 96.3 MFlops, respectively, corresponding to 2.57 million and 2.43 million parameters. After processing using a merging method, a smaller DNN model with a computational complexity of 25 MFlops and 3.18 million parameters is ultimately adopted.
[0190] The system's overall computational complexity is 233.83 million floating-point operations, with over 8.17 million parameters. Compared to... Figure 16 The sequential system shown reduces computational complexity by 265.67 million floating-point operations—even though the number of parameters is almost three times that of the original system.
[0191] The system's performance is significantly better than, both subjectively and objectively, that of other systems. Figure 16 The system shown clearly demonstrates that the aforementioned innovations enable the system to achieve extremely high computational efficiency without sacrificing performance.
[0192] Figure 15b Showing Figure 16 Examples and Figure 16 The complexity comparison of the embodiments, measured in millions of floating-point operations, shows a significant reduction in complexity.
[0193] Figure 17 Another preferred embodiment is shown, applicable to applications with reduced sampling rates. Although Figure 15a The application scenarios of a 48kHz sampling rate audio processing system are described, but Figure 17 A solution suitable for a 16kHz sampling rate is demonstrated—the computational complexity at low sampling rates is reduced by using channel-level subband splitting 301 and channel-level subband merging 401. Figure 16 The same diagram illustrates a 16kHz sampling rate, but Figure 17 The implementation scheme is as follows Figure 18 As shown in the table, additional complexity reduction is achieved. Specifically, through the cubic subband segmentation of block 301, the frequency dimension of 257 is reduced, thereby increasing the number of channels from 2 to 6; the frequency dimension of each channel 129 reaches 387, a value greater than 257, due to the overlap processing applied in the channel-level subband segmentation 301. The channel-level subband merging operation shown in block 401 reduces the number of channels by 3 and increases the frequency dimension to the maximum value of the frequency input features, 387. The small post-processing neural network 420 reduces the dimension from 387 back to the original value of 257 (as shown in block 120), a value that remains unchanged during the power-law decompression process of block 350.
[0194] The original input features of this system are F=257 and C=2. Channel-level subband segmentation uses a segmentation window of 129 F-band frequencies and 64 overlapping bands, therefore the output shape of this block is (B,T,129,6). The computational complexity of this system is 48.23 MFlops. If the architecture is adopted but the segmentation and merging steps are omitted, the computational complexity will reach at least 150 MFlops.
[0195] The computational complexity can be further reduced by narrowing the segmentation window and the overlap range. The model with the lowest computational complexity under this architecture (with only fine-tuning of the RNN layer) uses a segmentation window of 65 frequency bands and an overlap range of 17 frequency bands, with the output shape of the channel-level subband segmentation block being (B,T,65,10).
[0196] Overall, for the above embodiments and other implementations, the segmented overlap between 3 to 50 frequency bands has proven to have significant advantages.
[0197] The computational complexity of the system is reduced to 24.93 MFlops, with only a slight loss in subjective performance.
[0198] This section will explain the design decisions, selection criteria, and complexity reduction effects of multi-level processing. The channel-level subband segmentation method described above has limitations in reducing computational complexity: the smaller the segmentation window, the more channels there are, making it difficult for the main deep neural network (DNN) to estimate intermediate results.
[0199] To avoid processing the real and imaginary parts separately, the input signal is segmented into amplitude and phase features. Given the higher subjective importance of amplitude features in denoising tasks, an amplitude-based channel-level subband segmentation and merging method is adopted, and a relatively small-scale DNN is used. M Process (e.g.) Figure 11 (As shown). Since the amplitude is a one-dimensional feature, it is input into the DNN. M The feature dimension is (B,T,257,1). After reducing the input dimension to 1 dimension, a smaller segmentation window α can be introduced.
[0200] In practical applications, two independent DNNs are used for two different systems. M Model. One application scenario uses a 48-band segmentation window, with 18 overlapping bands. After channel-level subband segmentation, the feature dimensions become (B, T, 48, 8). The system uses a DNN. M The computational complexity is 4.62 MFlops, and the parameter size is 0.684 M. In another case, a DNN was used. M The computational complexity is 2.74 MFlops, and the parameter size is 0.362 M. The system uses a 40-band segmentation window (16 overlapping bands), and the processed feature dimension is (B, T, 40, 10).
[0201] A more streamlined DNN is used in both cases. Mp The phase information is processed and fused with the intermediate output. This DNN Mp The computational complexity is 1.73 MFlops, with 3.4K parameters. Therefore, the total computational complexity of the first system is 6.36 MFlops, and that of the latter is 4.48 MFlops.
[0202] The overall performance degradation after adopting this method remains within an acceptable range, but the computational complexity is reduced by nearly 30 times. This breakthrough enables the algorithm to run on embedded devices at a 16kHz sampling rate.
[0203] Subsequently, other embodiments of the invention (e.g., those relating to the first aspect) are summarized as examples, wherein the reference numbers in parentheses do not constitute any limitation on the basic principles.
[0204] 1. An information signal processing device, comprising:
[0205] A feature extractor (100) is used to extract a set of features from an information signal, the feature set having a first dimension;
[0206] Feature segmenter (200) is used to segment a feature set into a first feature subset and a second feature subset, wherein the first feature subset has a second dimension, the second feature subset has a third dimension, and both the second and third dimensions are lower than the first dimension. At the same time, the first feature subset and the second feature subset have overlapping ranges, such that one or more features in the feature set exist simultaneously in the first feature subset and the second feature subset.
[0207] A neural network processor (300) is configured to process a first subset using a first neural network (310) to obtain a first result, and to process a second subset of features using a second neural network (320) to obtain a second result;
[0208] A feature combiner (400) is used to combine the first result and the second result through a third neural network (420), the third neural network having a third complexity lower than the first complexity of the first neural network (310) or the second complexity of the second neural network (320), thereby obtaining a feature result set with result dimensions; and
[0209] The output post-processor (500) is used to post-process the feature result set to obtain the processed information signal.
[0210] 2. The apparatus as described in Embodiment 1, wherein the feature extractor (100) comprises: a time-frequency decomposer for generating a time-frequency decomposed signal from a time-domain represented information signal, the signal comprising a series of time frames, each time frame having a plurality of frequency bands; and a feature set builder for constructing a feature set from the time-frequency decomposed signal;
[0211] The output post-processor (500) includes a frequency-time synthesizer for synthesizing a processed information signal in the time domain from an input time-frequency representation feature set, wherein the input feature set is a feature result set or is derived from the feature result set and the feature set.
[0212] 3. The apparatus of Example 2, wherein the feature result set is a time-frequency mask, and the output post-processor (500) includes a mask processor: for applying the time-frequency mask to the feature set in the time-frequency representation to obtain the input feature set of the frequency-time synthesizer, or for calculating a processing filter from the time-frequency mask and applying the processing filter to the information signal or feature set to obtain the input feature set of the frequency-time synthesizer.
[0213] 4. The apparatus as described in any of the foregoing embodiments,
[0214] The first neural network (310) is more complex than the second neural network (320), or the first feature subset contains information in the lower frequency range of the information signal, while the second feature subset contains information in the higher frequency range of the information signal.
[0215] 5. The apparatus as described in Example 4, wherein the feature segmenter (200) is configured to generate a first feature subset and a second feature subset such that the second dimension is lower than the third dimension.
[0216] 6. The apparatus of any of the foregoing embodiments, wherein the neural network processor (300) is configured to segment (302) the second feature subset into a plurality of second segments and arrange the plurality of second segments along the channel dimension to increase the number of channels of the input set input to the second neural network (320).
[0217] 7. The apparatus as described in any of the preceding embodiments, wherein the neural network processor (300) is configured to segment a first subset of features into a first plurality of segments and arrange the first plurality of segments along the channel direction to increase the number of channels of the input set input to the first neural network (310).
[0218] 8. The apparatus as described in Example 6 or 7, wherein the number of the second plurality of segments is greater than the number of the first plurality of segments, or wherein the segment size of the second plurality of segments is the same in the second plurality of segments, or wherein the segment size of the first plurality of segments is the same in the first plurality of segments, or wherein only the second subset is segmented and the subset is not segmented.
[0219] 9. An apparatus as shown in any of Examples 6 to 8, wherein the neural network processor (300) is configured to segment a second subset or a first subset into overlapping segments.
[0220] 10. The apparatus as described in Example 9, wherein the overlap of the segments is between 1 / 5 and 4 / 5 of the segment width.
[0221] 11. The apparatus of any of the foregoing embodiments, wherein the feature combiner (400) includes a stacker (402) for stacking a second result of a second neural network (320) or stacking (101) a first result of a first neural network (310) to obtain a stacked feature set; and
[0222] The third neural network is configured to receive a stacked feature set as input and output a feature result set, wherein the dimension of the stacked feature set is higher than that of the result set.
[0223] 12. The apparatus of embodiment 11, wherein the stacker (401, 402) is configured to generate a stacked feature set by sorting the input feature set, and the dimension of the stacked feature set is greater than the dimension of the feature set obtained by the feature extractor (100).
[0224] The third neural network (420) is configured to generate a feature result set with a dimension lower than that of the stacked set.
[0225] 13. The apparatus of any of the foregoing embodiments, wherein the apparatus is configured as an embedded device, or the apparatus is included in an embedded device, or the first neural network (310) and the second neural network (320) are configured to run in parallel in a hardware implementation, or the result dimension is greater than the second dimension or the third dimension.
[0226] 14. The apparatus as described in any of the preceding embodiments, wherein the sampling rate of the information signal input to the feature extractor is greater than 16 kHz, wherein the feature extractor (100) includes a time-frequency decomposer that generates at least 300 frequency bands for each time frame, or wherein the overlapping range includes at least 40 frequency bands.
[0227] 15. An information signal processing method, comprising:
[0228] Extract a feature set from the information signal, which has a first dimension;
[0229] The feature set is divided into a first feature subset and a second feature subset, wherein the first feature subset has a second dimension and the second feature subset has a third dimension, and both the second and third dimensions are lower than the first dimension; at the same time, the first feature subset and the second feature subset have overlapping ranges, such that one or more features in the feature set exist in both the first feature subset and the second feature subset.
[0230] The first feature subset is processed using a first neural network (310) to obtain a first result, and the second feature subset is processed using a second neural network (320) to obtain a second result;
[0231] A third neural network (420) is used to combine the first and second results, having a third complexity lower than the first complexity of the first neural network (310) or the second complexity of the second neural network (320), to obtain a feature result set with result dimensions; and
[0232] Post-processing is performed on the feature result set to obtain the processed information signal.
[0233] 16. A computer program for performing the method described in Example 15 when run on a computer or processor.
[0234] Subsequently, other embodiments of the invention (e.g., embodiments relating to the third aspect) are summarized as examples, wherein the reference numbers in parentheses do not constitute any limitation on the basic principles.
[0235] 1. An information signal processing device, comprising:
[0236] A feature extractor (100) for extracting a feature set from an information signal, wherein each feature of the feature set contains at least two feature components, and the feature set contains a first subset having a first feature component and a second subset having a second feature component; and a neural network processor (300) comprising:
[0237] A first neural network (340) is used to receive a first subset as input and output the processed first subset;
[0238] Combiner (350) is used to combine the processed first subset and the second subset to obtain a combined subset; and
[0239] A second neural network (360) is configured to receive a combined subset as input and output a processed combined output, wherein the processed combined output represents a processed information signal, or the device is configured to use the processed combined output to calculate the processed information signal, and
[0240] The complexity of the first neural network (340) is greater than that of the second neural network.
[0241] 2. The apparatus as described in Embodiment 1, wherein the information signal is an audio signal, and the feature extractor (100) includes a time-frequency decomposer for calculating a time-frequency domain representation of the audio signal, wherein a first subset contains amplitude values of the time-frequency representation, a second subset contains phase values of the time-frequency representation, or wherein a first feature component of at least two feature components is more important than a second feature component of at least two feature components.
[0242] 3. The apparatus as described in any of the preceding embodiments, wherein the combiner (350) is configured to connect the processed first subset and the second subset along the channel direction, wherein the first neural network is configured to process the first subset with a channel number of 1 or greater than 1, and the second neural network (360) is configured to process the combined subset with a second channel number, wherein the second channel number is greater than the first channel number.
[0243] 4. The apparatus as described in Example 3, wherein the number of second channels is one more than the number of first channels.
[0244] 5. The apparatus described in any of the foregoing embodiments, wherein the first subset includes amplitude values, the second subset includes phase values, and wherein the processed first subset includes amplitude values.
[0245] The combiner (350) is configured to: for each component of the processed first subset and second subset, calculate the real component and the imaginary component, and...
[0246] The second neural network (360) is configured to receive a subset containing real and imaginary components as a combined subset.
[0247] 6. The apparatus as described in Embodiment 5, wherein the first subset comprises the amplitude values of the feature set, and the feature set comprises a complex number of frequency band entries.
[0248] The first neural network (340) is configured to compute the first subset after processing, such that the first subset after processing represents a time-frequency mask that has only the amplitude value of the frequency band;
[0249] The combiner (350) is configured to combine the amplitude mask value of the frequency band with the phase value of the complex frequency band entries of the corresponding frequency band, so that the amplitude mask value of the frequency band is combined with the phase value of the complex frequency band entries of the same frequency band.
[0250] 7. The apparatus as described in any of the foregoing embodiments, wherein the neural network processor (300) is configured to
[0251] Perform channel-direction segmentation (301, 302), dividing the first subset into M segments.
[0252] The first neural network (340) is configured to receive a first subset of M channels connected in series along the channel direction, and output the processed first subset; and
[0253] The neural network processor (300) is configured to combine a first processing subset (420) connected in series along the channel direction into a single-channel representation as the first processing subset.
[0254] 8. The apparatus described in Example 7, wherein the M multiple segments include two, three or more segments, wherein the three or more segments are overlapping segments, and the overlap range between the overlapping segments is greater than 3 frequency bands and less than 50 frequency bands.
[0255] 9. The apparatus of embodiment 7, wherein the neural network processor (300) is configured to perform channel-level subband segmentation (301, 302), dividing a first subset into M overlapping segments, and the processed first subset having M channels;
[0256] The neural network processor (300) is configured to perform channel-level subband segmentation of the second subset, dividing it into M overlapping segments, and arranging the second subset into M channels, and
[0257] The combiner (350) is configured to process the first subset and the second subset such that the combined subset includes twice the number of M channels.
[0258] 10. The apparatus of Example 9, wherein the second neural network (360) is configured to process a combined subset having twice the number of M channels to obtain the output of the second neural network (360), and
[0259] The neural network processor (300) is configured to stack (401, 402) and combine the stacked feature sets to form the processed combined output, the dimension of which is similar to the dimension of the feature set extracted by the feature extractor (100).
[0260] 11. The apparatus as described in any of the foregoing embodiments, wherein the complexity of the first neural network (340) and the second neural network (360) is measured in floating-point operations (FLOPS), wherein a higher number of FLOPS indicates higher complexity; or in execution time on one or more target hardware, wherein a longer execution time indicates higher complexity; or in power consumption of a specific device, wherein higher power consumption indicates higher complexity; or in the number of MAC (multiply-accumulate) operations, wherein a higher number of MAC operations indicates higher complexity.
[0261] 12. The apparatus of any of the foregoing embodiments, wherein the feature extractor (100) comprises:
[0262] A raw feature calculator (110) is used to calculate raw feature results, wherein each raw feature result contains at least two raw feature components; and
[0263] A raw feature compressor (120) is used to compress at least two raw feature components to obtain at least two compressed raw feature components for each raw feature result, wherein a first subset contains a first feature of at least two raw feature components in the raw feature result, and a second subset contains a second feature of at least two raw feature components in the raw feature result.
[0264] 13. The apparatus described in any of the foregoing embodiments, wherein the information signal includes an audio signal, an image signal, or a radar signal.
[0265] 14. The apparatus as described in any of the foregoing embodiments, wherein the apparatus is configured as an embedded device, or wherein the first neural network (340) and the second neural network (360) are configured to operate in series.
[0266] 15. Methods for processing information signals, including:
[0267] Extracting a feature set from an information signal, wherein each feature in the feature set contains at least two feature components, and the feature set contains a first subset having a first feature component and a second subset having a second feature component; and
[0268] The neural network processor (300) is used, which includes:
[0269] A first neural network (340) is used to receive a first subset as input and output the processed first subset;
[0270] Combiner (350) is used to combine the processed first subset with the second subset to obtain a combined subset; and
[0271] A second neural network (360) is configured to receive a combined subset as input and output a processed combined output, wherein the processed combined output represents a processed information signal, or the device is configured to use the processed combined output to calculate the processed information signal, and
[0272] The complexity of the first neural network (340) is greater than that of the second neural network.
[0273] 16. A computer program for performing the method described in Example 15 when run on a computer or processor.
[0274] It should be noted that all the aforementioned alternatives or aspects, as well as all aspects defined in the independent claims in the following claims, can be used independently, i.e., without needing to combine with other alternatives, objects, or independent claims. However, in other embodiments, two or more alternatives, technical solutions, or independent claims can be combined with each other; in other embodiments, all technical solutions or alternatives and all independent claims can be combined with each other.
[0275] Although some aspects are described in the context of apparatus, these aspects also clearly describe the corresponding methods, where blocks or apparatuses correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also describe blocks, items, or features of the corresponding apparatus.
[0276] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. This implementation may employ a digital storage medium (e.g., floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM, or flash memory) to store electronically readable control signals, which works in conjunction with (or has the capability to work in conjunction with) a programmable computer system to execute the corresponding method.
[0277] Some embodiments of the present invention include a data carrier having electronically readable control signals that are capable of working in conjunction with a programmable computer system to perform one of the methods described herein.
[0278] Typically, embodiments of the present invention can be implemented as a computer program product containing program code, which can execute any of the methods when the computer program product runs on a computer. The program code can be stored on a machine-readable medium.
[0279] Other embodiments include a computer program for performing any of the methods described herein, the program being stored on a machine-readable carrier or a non-temporary storage medium.
[0280] In other words, one embodiment of the method of the present invention is: a computer program having program code for executing any of the methods described herein when the computer program is running on a computer.
[0281] Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium, computer-readable medium) having a computer program recorded thereon for performing any of the methods described herein.
[0282] Another embodiment of the method of the present invention is a data stream or signal sequence, representing a computer program for performing any of the methods described herein. This data stream or signal sequence may be configured to be transmitted via a data communication connection, such as via the Internet.
[0283] Another embodiment includes a processing device (e.g., a computer or programmable logic device) configured or adapted to perform any of the methods described herein.
[0284] Another embodiment includes a computer on which a computer program is installed for performing any of the methods described herein.
[0285] In some embodiments, programmable logic devices (e.g., field-programmable gate arrays) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may work in conjunction with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.
[0286] The above embodiments are merely illustrative of the principles of the present invention. Those skilled in the art will understand that modifications and variations to the structures and details described are readily apparent. Therefore, the present invention is limited only by the scope of the pending claims and not by the specific details presented herein through the description and explanation of the embodiments.
[0287] References
[0288] [1] Yu, G., Guan, Y., Meng, W., Zheng, C., & Wang, H. (2022). DMF-Net: Adecoupling-style multi-band fusion model for full-band speech enhancement.
[0289] [2] Zhang Xu, Chen Lianwu, Zheng Xiguang, Ren Xinlei, Zhang Chen, Guo Liang, Yu Bin. Two-step backward compatible full-band speech enhancement system. ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2022): 7762-7766.
[0290] [3] S. Zhao, B. Ma, KN Watcharasupat and W.-S. Gan, "FRCRN: Monophonic Speech Enhancement Based on Frequency Recurrence Enhancement Feature Representation", ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 9281-9285, doi: 10.1109 / ICASSP43922.2022.9747578.
[0291] [4] Liu Haohe, Xie Lei, Wu Jian, Yang Geng. "Optimization of high-resolution music vocal accompaniment separation based on channel-division frequency band input". ArXiv abs / 2008.05216(2020): No page number.
[0292] [5] Lü Shubo, Fu Yihui, Xing Mengtao, Sun Jiayao, Xie Lei, Huang Jun, Wang Yannan, Yu Tao. S-DCCRN: Ultra-wideband DCCRN speech enhancement technology based on learnable complex features. ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2022): 7767-7771.
Claims
1. An apparatus for processing information signals, comprising: Feature extractor (100) for extracting a feature set from an information signal, wherein feature extractor (100) includes: The original feature calculation unit (110) is used to calculate the original feature results, each original feature result containing at least two original feature components; and A raw feature compressor (120) is configured to perform dynamic range compression on at least two raw feature components, thereby obtaining at least two compressed raw feature components for each raw feature result, wherein the feature set contains the compressed raw feature components; and A signal processor (600) is used to process a feature set to obtain a processed information signal, wherein the information signal includes an audio signal, an image signal, or a radar signal.
2. The apparatus of claim 1, wherein the original feature calculator (110) is configured to calculate complex values as original feature results, the original feature results having real and imaginary parts or amplitude and phase as at least two original feature components; and The original compressor (120) is configured to compress the real and imaginary parts or the amplitude and phase to obtain the compressed original feature components.
3. The apparatus of claim 1 or 2, wherein the original feature compressor (120) is configured to apply (121) a first compression function to a first original feature component and (122) a second compression function to a second original feature component, the second compression function being different from the first compression function.
4. The apparatus as described in any of the preceding claims, wherein the compression includes applying (121,122) power-law compression using a power value less than 1.
5. The apparatus of claim 4, wherein the power value of each original feature component is different, or Among the features The extractor (110) is configured to provide one or more different values for a specific information signal or a specific processing task performed by the signal processor (600). The feature extractor (100) is configured to perform a specific processing task for a specific information signal or signal processor (600), providing one or more different values derived from specific training performed on the specific information signal or signal processor (600) for the specific processing task.
6. The apparatus as described in any of the preceding claims, wherein the original feature calculation unit (110) is configured to calculate the original feature components as absolute values and their corresponding signs, and The original feature compressor (120) is configured to apply a compression function to the absolute value and retain the sign of the corresponding component.
7. The apparatus as described in any of the preceding claims, wherein the information signal is an audio signal, and the original feature calculation unit (110) is configured to perform time-frequency decomposition to decompose the audio signal into a time-frequency representation, or The information signal is an image signal, and the original feature calculator (110) includes performing a spatial transformation to decompose the image signal into a spatial frequency representation.
8. The apparatus as described in any of the preceding claims, wherein the signal processor (600) comprises: Neural network processors (20, 300) are used to process the feature set to obtain the processed feature set; A feature decompressor (350) is used to perform decompression matching with the compression operation performed by the original feature compressor (120) to obtain a feature result set; as well as The post-processor (500) is used to post-process the feature result set to obtain the processed information signal.
9. The apparatus of claim 8, wherein the neural network processor (300) is configured to segment the feature set (301, 302) into multiple segments and arrange the multiple segments along the channel dimension to increase the number of channels from at least 2 to a value greater than 2 or an integer multiple of 2, and The neural network processor (300) includes a neural network configured to receive multiple segments as input channels and generate multiple output channels as outputs, wherein the neural network processor is configured to stack (401, 402) output channels to obtain a stacked feature set.
10. The apparatus of claim 9, wherein the neural network processor (300) is configured to perform feature combination (400) on a stacked feature set using an additional neural network (420) with a complexity less than that of the neural network.
11. The apparatus of claim 10, wherein the neural network processor (300) is configured to segment the feature set (301, 302) into overlapping segments, wherein the dimension of the stacked feature set is higher than the dimension of the feature set, and The additional neural network (420) is configured to reduce the dimension of the stacked feature set to the dimension of the feature set or a dimension lower than the dimension of the feature set.
12. The apparatus as described in any of the preceding claims, wherein the information signal is an audio signal. The audio signal sampling rate is below 24kHz. The feature set has a frequency dimension of less than 500 and greater than 200, and the number of channels in the two original feature components is 2, with the number of segments between 2 and 4 and the number of channels between 4 and 8. The overlap between each segment is between 30 and 90 frequency bands.
13. The apparatus as described in any of the preceding claims, The signal processor contains one or more neural networks used to process feature sets to obtain processed information signals.
14. The apparatus as described in any of the preceding claims, The signal processor (600) includes a neural network processor (300) for processing a feature set containing compressed raw feature components, the neural network processor (300) being configured to output a feature result set. The device includes a feature decompressor for performing decompression matching with the compression operation performed by the original feature compressor (120) to obtain a decompressed feature result set. in, The decompressed feature set is a time-frequency mask. The signal processor (600) includes a mask processor, which is used to apply the time-frequency mask as a spectral gain mask to the feature set extracted by the original feature calculator (110) and the representation of the processed information signal, or to calculate a processing filter from the time-frequency mask and apply the processing filter to the audio signal or the feature set calculated by the original feature calculator (110) to obtain the representation of the processed information signal.
15. An information signal processing method, comprising: Extracting feature sets from information signals includes: Calculate the original feature results, each original feature result containing at least two original feature components; and Dynamic range compression is performed on at least two original feature components to obtain at least two compressed original feature components for each original feature result, wherein the feature set contains the compressed original feature components; and The feature set is processed to obtain the processed information signal, which includes audio signals, image signals, or radar signals.
16. A computer program for performing the method of claim 15 when executed on a computer or processor.