Master band level audio enhancement system and method based on generative adversarial network

Through the master-level audio enhancement method based on the generative adversarial network, the problem of poor audio enhancement in the prior art is solved, high-quality master-level audio enhancement is achieved, and the overall quality and detail retention ability of the audio are improved.

CN120299474AInactive Publication Date: 2025-07-11DONGGUAN YUTAI ELECTRONICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510775603.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing master-level audio enhancement technology is difficult to adapt to the dynamic range and frequency response differences of different audio materials, resulting in loss of details, sound quality distortion and excessive dynamic range compression of enhanced audio, which cannot meet the requirements of professional mastering.

Method used

The master-level audio enhancement method based on the generative adversarial network is adopted. By segmenting the audio signal into multiple segments, a generative adversarial network of the main generator, auxiliary generator and triple discriminator is built. Combining the three-dimensional adversarial loss vector and the double nested adaptive feedback regulation system, the audio features are deeply fusion and iteratively optimized to generate high-quality final audio signals.

Benefits of technology

It realizes high-quality master-level audio enhancement, improves the overall quality of the audio, meets the needs of professional production, retains the details and delicate texture of the audio, and reduces noise interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299474A_ABST
    Figure CN120299474A_ABST
Patent Text Reader

Abstract

The invention provides a mother band level audio enhancement system and method based on a generative adversarial network, and relates to the technical field of audio processing. The method comprises the following steps: segmenting a mother band level audio to be enhanced into a plurality of segments according to a preset time window, preprocessing to obtain low, medium and high frequency feature matrixes, then constructing a generative adversarial network, respectively processing the low and high frequency feature matrixes by a main generator and an auxiliary generator, and fusing to obtain a fused feature matrix; converting the audio signals into different signals, inputting the different signals into a discriminator to obtain scores, combining the signals into a three-dimensional adversarial loss vector, optimizing network parameters through a double-nested adaptive feedback regulation and control system, performing secondary enhancement on a fusion feature matrix, generating a preliminary audio signal through wavelet inverse transformation, and finally performing noise reduction to obtain a final audio signal. By implementing the method, audio details can be effectively reserved, the problems of tone quality distortion and excessive compression of a dynamic range are reduced, and particularly, the tone quality performance of a high frequency band can be improved, so that the enhanced audio achieves fine texture and rich details required by a master band level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of digital data processing, and particularly to a master-level audio enhancement system and method based on a generative adversarial network. Background Art

[0002] In the field of audio processing, master-level audio enhancement technology is crucial for improving audio quality. Traditional audio enhancement methods often rely on fixed filter parameters and preset algorithm models, and it is difficult to adapt to the complex differences in dynamic range, frequency response, etc. of different audio materials, resulting in problems such as detail loss, sound quality distortion, and excessive dynamic range compression in the enhanced audio, and unable to meet the strict requirements for audio quality in professional master production.

[0003] To solve the above problems, some technologies adopt a convolutional neural network (CNN) model based on deep learning for audio enhancement. By training the model with a large amount of labeled data, the model learns the mapping relationship between different audio features and enhancement parameters, and can adaptively process different audio materials to a certain extent, improving some defects existing in traditional methods.

[0004] However, when the existing CNN-based audio enhancement technology processes audio, it only generates the enhanced audio in the forward direction, lacking an effective feedback mechanism for the difference between the generated audio and the original high-quality audio, resulting in the model being difficult to accurately capture audio details during the learning process, especially prone to losing high-frequency detail information during the enhancement process, making the enhanced audio perform poorly in the high-frequency band and unable to truly achieve the delicate texture and rich details required for master-level audio. Summary of the Invention

[0005] This application provides a master-level audio enhancement system and method based on a generative adversarial network, which is used to solve the defects existing in traditional audio enhancement methods and CNN-based audio enhancement technologies, and achieve a high-quality master-level audio enhancement effect.

[0006] In a first aspect, the present application provides a master track-level audio enhancement method based on a generative adversarial network. The method includes: segmenting the master track-level audio signal to be enhanced according to a preset time window to obtain a plurality of audio segments; performing preprocessing operations on the plurality of audio segments to obtain a low-frequency feature matrix, a mid-frequency feature matrix, and a high-frequency feature matrix; constructing a generative adversarial network, which consists of a main generator, an auxiliary generator, and a triple discriminator. The main generator adopts a recursive convolutional structure based on a gated recurrent unit, the auxiliary generator is a multi-branch attention mechanism network, and the triple discriminator is a time-domain discriminator, a frequency-domain discriminator, and a phase discriminator respectively; inputting the low-frequency feature matrix into the main generator to generate a low-frequency enhanced feature matrix, and inputting the high-frequency feature matrix into the auxiliary generator to generate a high-frequency enhanced feature matrix; deeply fusing the low-frequency enhanced feature matrix, the mid-frequency feature matrix, and the high-frequency enhanced feature matrix to generate a fused feature matrix; converting the fused feature matrix into a time-domain signal, a frequency-domain signal, and a phase signal respectively, and inputting them into the corresponding triple discriminator to obtain three corresponding discriminant scores; combining the three discriminant scores into a three-dimensional adversarial loss vector including time-domain, frequency-domain, and phase dimensions; combining the three-dimensional adversarial loss vector, and iteratively optimizing the network parameters of the main generator and the auxiliary generator through a double nested adaptive feedback regulation system; according to the optimized parameters, performing secondary enhancement processing on the fused feature matrix to determine an enhanced fused feature matrix; generating a preliminary audio signal by performing inverse wavelet transform on the enhanced fused feature matrix; and performing noise reduction processing on the preliminary audio signal to determine the final audio signal.

[0007] By adopting the above technical solution, first, the master track-level audio signal to be enhanced is segmented into a plurality of audio segments, which can make the subsequent processing focus on finer audio units. Preprocessing the audio segments to obtain low, mid, and high-frequency feature matrices can separate the features of different frequency bands of the audio, facilitating targeted enhancement. Constructing a generative adversarial network consisting of a main generator, an auxiliary generator, and a triple discriminator, the generators with different structures can effectively enhance the low-frequency and high-frequency features respectively, and the triple discriminator discriminates from the time-domain, frequency-domain, and phase dimensions to ensure that the generated audio features are closer to the ideal state. After steps such as deep fusion, iterative optimization, and secondary enhancement, finally, a high-quality and low-noise final audio signal can be generated, improving the overall quality of the audio. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 is a flowchart of a master track-level audio enhancement method based on a generative adversarial network in an embodiment of the present application; Figure 2 is another flowchart of a master track-level audio enhancement method based on a generative adversarial network in an embodiment of the present application; Figure 3It is a schematic structural diagram of an entity device in the audio enhancement system according to an embodiment of the present application. Detailed implementation manners

[0009] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to and includes any or all possible combinations of one or more of the listed items.

[0010] Hereinafter, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0011] For ease of understanding, the method provided in this embodiment will be described in terms of a process below. Please refer to Figure 1 , which is a schematic flowchart of a master tape-level audio enhancement method based on a generative adversarial network in an embodiment of the present application.

[0012] S101. Divide the master tape-level audio signal to be enhanced according to a preset time window to obtain a plurality of audio segments; The audio enhancement system first obtains the master tape-level audio signal to be enhanced from an external storage device or network. This obtaining process depends on the input interface of the system, which can identify and receive files in multiple audio formats, such as common formats like WAV and MP3.

[0013] After obtaining the audio signal, the system divides it according to the preset time window. The length of the preset time window is not set arbitrarily, but is determined by considering multiple factors. If the time window is too short, although the audio segments can be processed more carefully, the computational amount and processing complexity will increase; if the time window is too long, some short-term change features in the audio may be lost.

[0014] In the actual segmentation process, the system uses the slicing algorithm in digital signal processing technology. Taking an audio signal with a duration of T seconds and a sampling rate of fs (unit: Hz) as an example, assuming the preset time window length is t seconds. First, calculate the total number of samples N = T×fs of the audio signal according to the sampling rate. Then, slice the audio signal at intervals of the number of samples n = t×fs corresponding to the time window length. Starting from the starting position of the audio signal, each time a sample segment of length n is intercepted until the entire audio signal is traversed, thus obtaining multiple audio segments with a length of t seconds.

[0015] S102. Perform preprocessing operations on the multiple audio segments to obtain a low-frequency feature matrix, a mid-frequency feature matrix, and a high-frequency feature matrix; After obtaining the multiple audio segments, the audio enhancement system immediately performs preprocessing on them to extract feature matrices of different frequency bands. This process is mainly achieved through the fourth-order discrete wavelet transform and the frequency-domain analysis method based on the Mel scale.

[0016] The system performs the fourth-order discrete wavelet transform on each audio segment. The discrete wavelet transform (DWT) is an effective tool for decomposing a signal into different frequency components. In the fourth-order discrete wavelet transform, the audio signal is decomposed into a low-frequency approximation signal and three layers of high-frequency detail signals, forming multiple signal groups. Taking an audio segment as an example, assuming it is x(n), where n represents the time series, after the fourth-order discrete wavelet transform, a low-frequency approximation signal A4(n) and high-frequency detail signals D1(n), D2(n), D3(n) will be obtained. The low-frequency approximation signal A4(n) retains the main trend and general outline of the audio signal and contains most of the energy; while the high-frequency detail signals D1(n), D2(n), D3(n) correspond to the detail information in different frequency ranges, with the frequencies increasing in turn. In actual calculations, the system uses the fast wavelet transform (FWT) algorithm to improve the calculation efficiency. This algorithm decomposes the audio signal through a specific filter bank, greatly reducing the amount of calculation.

[0017] For each signal group, the system uses the frequency-domain analysis method based on the Mel scale to extract a low-frequency feature matrix, a mid-frequency feature matrix, and a high-frequency feature matrix respectively. The Mel scale is a frequency scale that conforms to the characteristics of human auditory perception. It converts the linear frequency axis into a non-linear scale that is more in line with human ear perception.

[0018] When extracting the low-frequency feature matrix, the system mainly focuses on the basic timbre and rhythm information of the audio. Since the low-frequency part plays a key role in the overall timbre and rhythm of the audio, the system extracts relevant features from the low-frequency approximation signal A4(n). For example, by calculating the power spectral density of the low-frequency signal, information reflecting the distribution of audio energy in the low-frequency band is obtained; combined with the time-domain characteristics of the audio, such as parameters like the zero-crossing rate, a low-frequency feature matrix is comprehensively constructed. Each element in the low-frequency feature matrix may represent information such as the energy intensity within a certain frequency range in the low-frequency band and the frequency change trend.

[0019] For the mid-frequency feature matrix, it characterizes the texture information of the transition frequency band. The system extracts mid-frequency features from the high-frequency detail signal D2(n) and part of the low-frequency approximation signal A4(n). By performing spectral analysis on these signals and calculating spectral features within a specific frequency range, such as the peaks, valleys of the spectrum and their relative position relationships, etc., a mid-frequency feature matrix is constructed. These features can reflect the texture changes of the audio in the mid-frequency transition band, such as the subtle changes in instrument timbre and the formant features in speech.

[0020] The high-frequency feature matrix is used to characterize the detail texture information. The system mainly extracts high-frequency features from the high-frequency detail signal D1(n). By using methods such as the Hilbert transform to obtain the instantaneous frequency and phase information of the audio signal, combined with the energy distribution in the high-frequency band, a high-frequency feature matrix is constructed. The elements in the high-frequency feature matrix can reflect the subtle details in the audio, such as the overtones during instrument performance and the sibilants in speech.

[0021] Through the above preprocessing steps, the system can accurately decompose the audio segment into feature matrices of different frequency bands, providing an accurate data basis for subsequent targeted enhancement.

[0022] S103. Construct a generative adversarial network, which consists of a main generator, an auxiliary generator, and a triple discriminator. The main generator adopts a recursive convolutional structure based on gated recurrent units. The auxiliary generator is a multi-branch attention mechanism network. The triple discriminator is respectively a time-domain discriminator, a frequency-domain discriminator, and a phase discriminator. When constructing a generative adversarial network, the audio enhancement system needs to build a main generator, an auxiliary generator, and a triple discriminator separately to achieve effective enhancement and accurate discrimination of audio features. The main generator adopts a recursive convolution structure based on a gated recurrent unit (GRU). GRU is a special recurrent neural network unit that can effectively handle long-term dependencies in sequence data. In the main generator, GRU is embedded in a recursive convolution architecture. First, the low-frequency feature matrix of the audio enters the main generator as input data. In the recursive convolution process, the convolution layer is responsible for extracting local features in the low-frequency feature matrix, and sliding convolution on the matrix through convolution kernels of different sizes to capture local information such as changes in strong and weak beats in the basic rhythm of the audio and specific frequency combinations in the basic timbre. GRU plays its memory characteristics. After each step of convolution processing, GRU will combine the features extracted by the current convolution and the hidden state information of the previous moment to decide which information needs to be retained and which needs to be updated. This mechanism enables the main generator to learn the long-term dependencies of low-frequency features in time series, such as the coherent changes in a continuous bass melody, thereby generating a more logical and coherent low-frequency enhanced feature matrix.

[0023] The auxiliary generator uses a multi-branch attention mechanism network. The network consists of multiple parallel branches, each of which is responsible for focusing on different aspects of the high-frequency feature matrix. When constructing, the high-frequency feature matrix is ​​first divided into multiple channels, and each channel corresponds to high-frequency detail information in different frequency ranges. For example, one branch focuses on the overtone details of musical instruments in the high-frequency band, and the other branch focuses on details such as sibilance in speech. Then, each branch extracts and reduces the dimensions of the features it focuses on through convolutional layers and pooling layers. The core of the attention mechanism is to calculate the importance weight of each branch feature. Specifically, by calculating the correlation between different branch features and the overall high-frequency features, a normalized weight vector is generated using the Softmax function. These weight vectors determine the contribution of each branch in the final generation of the high-frequency enhancement feature matrix. For example, in an audio containing multiple instruments, if the high-frequency overtone of a certain instrument plays a key role in the overall audio effect, the weight of the branch that focuses on the overtone will be relatively high, thereby highlighting this important detail in the generated high-frequency enhancement feature matrix.

[0024] The triple discriminator includes a time-domain discriminator, a frequency-domain discriminator, and a phase discriminator. The time-domain discriminator is constructed based on a bidirectional long short-term memory network (Bi-LSTM), which can simultaneously learn the features of the audio signal in the forward and reverse time series. When the time-domain signal of the audio is input into the time-domain discriminator, the forward and reverse hidden layers of the Bi-LSTM process the signal respectively, capturing the forward and backward dependencies of the audio at different time steps, such as whether the transition between the start and end parts of the audio is natural. Then, the time-domain continuity is evaluated by calculating the cosine similarity between adjacent time frames. The cosine similarity can measure the angle between the feature vectors of two time frames. The smaller the angle, the higher the similarity, indicating better continuity of the audio in the time domain. The frequency-domain discriminator uses Fourier descriptors to analyze the spectral smoothness. It performs a Fourier transform on the frequency-domain signal to obtain the spectral curve. The rate of change of the curvature of the spectral curve is calculated as the frequency-domain smoothness score. If the rate of change of the curvature of the spectral curve is small, it means the spectrum is relatively smooth, and the change of the audio in the frequency domain is relatively stable without sharp spectral mutations. The phase discriminator evaluates the feature synergy by calculating the phase-amplitude mutual information. Phase information plays an important role in the perception of audio, affecting the stereo and spatial sense of the sound. The phase discriminator measures the degree of synergy between the phase signal and the amplitude signal by analyzing the mutual information between them. The larger the mutual information, the better the synergy between the phase and the amplitude, and the higher the overall quality of the audio.

[0025] S104. Input the low-frequency feature matrix into the main generator to generate a low-frequency enhanced feature matrix, and input the high-frequency feature matrix into the auxiliary generator to generate a high-frequency enhanced feature matrix. The audio enhancement system inputs the low-frequency feature matrix and the high-frequency feature matrix obtained through preprocessing into the corresponding generators respectively to achieve targeted enhancement of the low-frequency and high-frequency parts of the audio.

[0026] When the low-frequency feature matrix is input into the main generator, the recursive convolutional structure based on the gated recurrent unit starts to play a role. First, the low-frequency feature matrix enters the convolutional layer, and the convolutional kernels in the convolutional layer slide and convolve on the low-frequency feature matrix according to the preset stride and padding. Assume the convolutional kernel size is 3×3, the stride is 1, and the padding is 1, which can ensure that the convolutional operation can cover every element of the low-frequency feature matrix and fully extract local features. During the convolution process, the parameters of the convolutional kernel are continuously optimized through network training, and these parameters determine the ability of the convolutional kernel to extract features. For example, after training, the convolutional kernel may be more sensitive to the energy changes in a specific frequency range in the low-frequency band, thus being able to capture the subtle changes in the basic tone of the audio.

[0027] After the convolution operation is completed, the resulting feature map enters the GRU unit. The GRU unit contains two gating mechanisms: the reset gate and the update gate. The reset gate determines how to combine the previous hidden state with the current input, and the update gate controls the contribution ratio of the current input and the previous hidden state to the new hidden state. When processing low-frequency features, the GRU unit can remember the long-term dependencies of low-frequency features in the time series. For example, for a steady rhythm of low-frequency drum beats, the GRU can remember the rhythm pattern of the drum beats and maintain the coherence and stability of the drum beats when generating the low-frequency enhanced feature matrix. As the recursive convolution continues, the GRU unit continuously updates the hidden state and finally outputs the low-frequency enhanced feature matrix. This matrix not only contains the enhanced information of the low-frequency features but also retains the logical relationship of the low-frequency part in the time series, making the enhanced low-frequency audio more natural and smooth.

[0028] After the high-frequency feature matrix is input into the auxiliary generator, the multi-branch attention mechanism network starts to work. In the multi-branch structure, each branch processes different frequency ranges of the high-frequency feature matrix. For example, one branch is responsible for processing the high-frequency details of 8 kHz - 16 kHz, and another branch processes the ultra-high-frequency part of 16 kHz - 20 kHz. Each branch first performs feature extraction through a convolutional layer. The convolutional layer can set different convolutional kernel parameters according to the frequency ranges concerned by different branches. For example, the branch concerned with the ultra-high-frequency part can use a smaller convolutional kernel to capture more subtle high-frequency features.

[0029] After being processed by the convolutional layer, the feature maps of each branch enter the attention mechanism module. The attention mechanism generates a weight vector by calculating the correlation between the features of each branch and the overall high-frequency features. For example, the dot product attention mechanism is adopted to perform a dot product operation between the feature vector of each branch and a learnable query vector, and then normalize it through the Softmax function to obtain the weight of each branch. If the features of a certain branch are significant in the overall high-frequency effect, such as the high-frequency overtones in a string instrument performance, then the weight corresponding to this branch will be larger. Finally, each branch performs weighted fusion according to the weights to generate the high-frequency enhanced feature matrix. This matrix highlights the important details of the high-frequency part, making the enhanced high-frequency audio clearer and brighter and improving the overall layer sense of the audio.

[0030] S105. Deeply fuse the low-frequency enhanced feature matrix, the intermediate-frequency feature matrix, and the high-frequency enhanced feature matrix to generate a fused feature matrix; This process is elaborated in detail in steps S201 - S206 and will not be repeated here. Generally, dimension adaptation and normalization are first performed on the three matrices to make them have the basic conditions for fusion. Then, a fusion weight matrix is constructed to clarify the contribution ratio of each frequency band matrix in the fusion. After that, a weighted fusion operation is carried out to obtain a preliminary fusion matrix, and the matrix is optimized through non - linear transformation and feature enhancement. Finally, post - processing is performed on the matrix, such as re - normalization, smoothing, etc., to generate a fusion feature matrix that comprehensively reflects the advantages of each frequency band.

[0031] S106. After the fusion feature matrix is respectively converted into a time - domain signal, a frequency - domain signal, and a phase signal, it is input into the corresponding triple discriminator to obtain three corresponding discriminant scores. After the audio enhancement system generates the fusion feature matrix, in order to comprehensively evaluate the quality of the audio features represented by this matrix, it will be respectively converted into a time - domain signal, a frequency - domain signal, and a phase signal, and input into the corresponding triple discriminator to obtain discriminant scores. The system will use techniques such as inverse Fourier transform to convert the fusion feature matrix into a time - domain signal. After the conversion, the time - domain signal is input into a time - domain discriminator constructed based on a bidirectional long short - term memory network (Bi - LSTM). Bi - LSTM can learn the features of the audio signal in both the forward and backward time series simultaneously. When the time - domain signal of the audio enters the time - domain discriminator, its forward and backward hidden layers process the signal respectively. Taking a 10 - second music audio as an example, the forward hidden layer starts from the beginning of the music and learns the features of the audio signal in chronological order, such as the start, duration, and end changes of the notes; the backward hidden layer starts from the end of the music and learns the features of the audio signal in reverse. In this way, the forward - backward dependence relationship of the audio at different time steps is captured, such as whether the gradual increase at the beginning of the audio is natural and whether the gradual decrease at the end is smooth.

[0032] To evaluate the continuity of the audio in the time dimension, the system obtains the time - domain continuity score by calculating the cosine similarity of adjacent time frames. Suppose the feature vector of the current time frame is A and the feature vector of the next time frame is B, then the cosine similarity calculation formula is used to calculate each pair of adjacent time frames one by one. For example, every 10 milliseconds is a time frame, and the cosine similarity between the 10 - millisecond and 20 - millisecond, 20 - millisecond and 30 - millisecond adjacent time frames is calculated in turn. The cosine similarities of all adjacent time frames are statistically analyzed, such as calculating the average value, and the result obtained is the time - domain continuity score. If this score is relatively high, close to 1, it means that the change of the audio in time is relatively smooth without obvious jumps or freezes; if the score is relatively low, close to 0, it indicates that there are time - discontinuity problems in the audio, such as sudden interruption of the sound or abnormal jumps.

[0033] The system uses Fourier transform to convert the fused feature matrix into a frequency-domain signal, which reflects the energy distribution of the audio signal at different frequencies. After the conversion is completed, the frequency-domain signal is input into the frequency-domain discriminator. The frequency-domain discriminator analyzes the spectral smoothness using Fourier descriptors, and uses the curvature change rate of the spectral curve as the frequency-domain smoothness score. During the analysis, the frequency-domain discriminator first obtains the spectral curve of the frequency-domain signal, which shows the energy intensity corresponding to different frequencies. Then, a specific algorithm is used to calculate the curvature change rate of the spectral curve. For example, for a piano performance audio, under normal circumstances, the transition of its spectral curve between different frequency bands should be relatively smooth. If there is a sudden large increase or decrease in energy at certain frequency points, the spectral curve will show sharp changes, and the curvature change rate will be large at this time, indicating poor frequency-domain smoothness; while if the spectral curve is relatively flat and the curvature change rate is small, the frequency-domain smoothness is better. The frequency-domain smoothness score can effectively detect whether the audio spectrum is smooth, avoid sharp spectral mutations, ensure that the changes in the audio in the frequency domain conform to natural laws, and improve the auditory effect of the audio.

[0034] The system uses a dedicated phase extraction algorithm to convert the fused feature matrix into a phase signal. Phase information plays an important role in the perception of audio, affecting the stereo and spatial sense of sound. After the conversion is completed, the phase signal is input into the phase discriminator, which evaluates the feature synergy by calculating the phase-amplitude mutual information and obtains the phase synergy score. The phase-amplitude mutual information is used to measure the degree of mutual dependence between the phase signal and the amplitude signal. During the calculation, the phase discriminator analyzes the correlation between the phase signal and the amplitude signal and uses a method based on information theory to calculate the mutual information value. For example, for a symphony audio with rich spatial sense, when the synergy between the phase and the amplitude is good, the mutual information value is large, and the phase information of the audio can cooperate with the amplitude information to enhance the stereo and spatial sense of the sound, enabling the listener to more clearly distinguish the positions of different instruments in space; on the contrary, a small mutual information value indicates poor synergy, which may lead to a blurred and unrealistic spatial sense of the sound. The phase synergy score evaluates the quality of the audio from the perspective of phase and provides an important basis for comprehensively evaluating the generated audio features.

[0035] S107. Combine the three discriminant scores into a three-dimensional adversarial loss vector including the time domain, frequency domain, and phase dimensions; After obtaining the time-domain continuity score, frequency-domain smoothness score, and phase synergy score, the audio enhancement system combines these three scores into a three-dimensional adversarial loss vector including the time domain, frequency domain, and phase dimensions. This step is a key link in the entire audio enhancement process. It integrates the audio quality evaluation results in different dimensions and provides a comprehensive and quantitative basis for the subsequent optimization of network parameters.

[0036] The system will standardize the three discriminant scores. Since the value ranges and physical meanings of the time-domain continuity score, frequency-domain smoothness score, and phase coherence score are different, directly combining them may cause the scores in some dimensions to dominate in subsequent calculations, thus affecting the accuracy of the overall evaluation. Therefore, they need to be standardized. Taking the time-domain continuity score as an example, if its value range is [0, 1], while the value range of the frequency-domain smoothness score is [0, 100], then directly combining them will make the frequency-domain smoothness score have too much influence on the result. Using a normalization method, such as min-max normalization, for the time-domain continuity score , it is normalized through the formula , where and are respectively the minimum and maximum values of this score in the training set; similar processing is also performed on the frequency-domain smoothness score and the phase coherence score to obtain the normalized scores and .

[0037] The system uses the three standardized scores as the three components of a three-dimensional vector to construct a three-dimensional adversarial loss vector , vector , where the first component represents the evaluation result in the time domain dimension, the second component represents the evaluation result in the frequency domain dimension, and the third component represents the evaluation result in the phase dimension. This three-dimensional vector comprehensively reflects the degree of difference between the generated audio features and the ideal state in different dimensions. For example, if the time-domain continuity score is low, that is is small, it indicates that there are discontinuity problems in the audio in the time dimension, and this problem will be reflected in the vector; similarly, the frequency-domain smoothness score and the phase coherence score will respectively reflect the situation of the audio in the frequency domain and phase aspects. Through this three-dimensional adversarial loss vector, the system can clearly understand the performance of the audio in each dimension, providing an intuitive and quantitative data basis for subsequent optimization.

[0038] In practical applications, the three-dimensional adversarial loss vector can also be weighted according to specific requirements. For example, if in a specific audio enhancement scenario, more attention is paid to the frequency-domain performance of the audio, then a higher weight can be assigned to the component corresponding to the frequency-domain smoothness score.

[0039] S108. Combine this three-dimensional adversarial loss vector and iteratively optimize the network parameters of the main generator and the auxiliary generator through a double-nested adaptive feedback regulation system; After obtaining the three-dimensional adversarial loss vector, the audio enhancement system needs to use a double-nested adaptive feedback regulation system to iteratively optimize the network parameters of the main generator and the auxiliary generator to achieve high-quality master-level audio enhancement effects. The double-nested adaptive feedback regulation system consists of an outer feedback loop and an inner feedback loop, and optimizes the network parameters of the main generator and the auxiliary generator as well as the Lorenz system parameters of the three-dimensional feature chaotic fusion mechanism through multiple iterations.

[0040] In each iteration, the outer feedback loop adjusts the network parameters of the main generator and the auxiliary generator based on the magnitude of the three-dimensional adversarial loss vector. The three-dimensional adversarial loss vector , where is the normalized time-domain continuity score, is the normalized frequency-domain smoothness score, is the normalized phase coherence score, and its magnitude is used to measure the comprehensive deviation degree of the generated audio features from the ideal state in the three dimensions of time domain, frequency domain, and phase. The calculation formula is . The larger the magnitude, the greater the gap between the generated audio and the ideal state; the smaller the magnitude, the closer it is to the ideal state.

[0041] To adjust the network parameters, the system uses the Proximal Policy Optimization algorithm (PPO). PPO is an efficient policy optimization algorithm, and its core is to maximize the advantage function of the policy. In this system, the magnitude of the three-dimensional adversarial loss vector is used as part of the optimization target. First, the system calculates the advantage function under the current policy based on the network parameters of the current main generator and auxiliary generator. The advantage function represents the degree of advantage of the current policy relative to a certain baseline policy. Then, optimization methods such as gradient descent are used to update the network parameters according to the advantage function. For example, assume that the magnitude of the current three-dimensional adversarial loss vector is 0.9. After calculating the advantage function of the current policy through the PPO algorithm, the gradient descent method is used to update the network parameters. After one iteration update, the magnitude decreases to 0.8, which means that the generated audio is closer to the ideal state as a whole. Through continuous iteration, the network parameters are gradually optimized, the magnitude continues to decrease, and the quality of the generated audio is improved.

[0042] At each iteration of the inner feedback loop, reinforcement learning is used to dynamically adjust the Lorentz system parameters of the three-dimensional feature chaos fusion mechanism for the local peaks of each dimension of the three-dimensional adversarial loss vector. The local peaks of each dimension of the three-dimensional adversarial loss vector reflect that the audio has a large deviation in the corresponding dimension at a specific moment or frequency range. For example, the local peak of the time domain dimension may indicate that there is a big problem with the time domain continuity of the audio in a certain time segment; the local peak of the frequency domain dimension may indicate that the spectral smoothness of the audio in a certain frequency interval is poor; the local peak of the phase dimension may indicate that the phase coordination of the audio in a certain phase region is poor. Reinforcement learning is a method of learning the optimal strategy through the interaction between the agent and the environment and based on the reward signal. In this system, the agent is the inner feedback loop, the environment is the current state of the audio enhancement system, and the reward signal is related to the local peak of each dimension of the three-dimensional adversarial loss vector. When a local peak is detected in a certain dimension of the three-dimensional adversarial loss vector, the intelligent agent in the inner feedback loop starts to act. The intelligent agent will first try different Lorenz system parameter adjustment strategies according to the current system state. The adjustment of the Lorenz system parameters will affect the three-dimensional feature chaotic fusion mechanism, and then change the fusion method of the feature matrix, and finally affect the generated audio features.

[0043] S109, performing secondary enhancement processing on the fusion feature matrix according to the optimized parameters to determine an enhanced fusion feature matrix; The audio enhancement system iteratively optimizes the network parameters of the main generator and the auxiliary generator through a double nested adaptive feedback control system, and obtains the optimized parameters. Next, the system will use these parameters to perform secondary enhancement processing on the fusion feature matrix to further improve the audio quality.

[0044] During the secondary enhancement process, the system will first reload the network parameters of the optimized main generator and auxiliary generator. These parameters are obtained through multiple iterations of optimization and can better reflect the characteristics of the audio and the ideal enhancement direction. The system will input the fused feature matrix into the optimized network again. At this time, the main generator and the auxiliary generator will process the fused feature matrix based on the new parameters. The main generator adopts a recursive convolution structure based on the gated recurrent unit, which will once again deeply mine and enhance the low-frequency part of the fused feature matrix. The convolution layer will convolve the fused feature matrix with the optimized convolution kernel parameters to further extract the subtle features of the low-frequency part, such as strengthening the stability of the basic rhythm of the audio and the richness of the basic timbre. The gated recurrent unit will use its memory characteristics and combine the long-term dependency of the low-frequency features learned before to enhance the low-frequency features more accurately, making the low-frequency part more solid and full in the audio.

[0045] The auxiliary generator is a multi-branch attention mechanism network, which will further optimize the high-frequency part in the fused feature matrix. The multi-branch structure will more accurately focus on different aspects of the high-frequency feature matrix according to the optimized parameters. For example, for the instrument overtone detail branch in the high-frequency band, feature extraction will be performed with more appropriate convolution kernel parameters to enhance the clarity and richness of the instrument overtones; for the detail branches such as sibilants in speech, targeted enhancement will also be carried out to highlight the detail features of the speech. The attention mechanism will more accurately allocate the weights of each branch when generating the high-frequency enhanced features according to the optimized weight calculation method, so that the important details in the high-frequency part can be more prominently reflected in the enhanced fused feature matrix.

[0046] S110. Generate a preliminary audio signal by performing inverse wavelet transform on the enhanced fused feature matrix; After obtaining the enhanced fused feature matrix, the audio enhancement system needs to convert it into a preliminary audio signal for subsequent processing. This conversion process is achieved through inverse wavelet transform.

[0047] Inverse wavelet transform is the inverse operation of wavelet transform, which can resynthesize the signal decomposed by wavelet transform back into the original signal. In this system, the enhanced fused feature matrix contains the enhanced feature information of the audio in different frequency bands, and through inverse wavelet transform, these feature information can be converted into the audio signal in the time domain. The system will determine the specific operation of the inverse wavelet transform according to the parameters used when performing the fourth-order discrete wavelet transform on the audio segment before. Because the audio signal was decomposed into a low-frequency approximation signal and three layers of high-frequency detail signals during the fourth-order discrete wavelet transform on the audio segment in the preprocessing stage, the synthesis needs to be carried out in the reverse process during the inverse wavelet transform.

[0048] The system will adjust the dimension of the enhanced fused feature matrix to meet the input requirements of the inverse wavelet transform. After the enhanced fused feature matrix undergoes the previous processing, its dimension may not be consistent with the dimension required by the inverse wavelet transform. The system will adjust the dimension of the enhanced fused feature matrix to an appropriate size through operations such as interpolation or downsampling. For example, if the dimension of the enhanced fused feature matrix is too high, the system will use pooling operation for downsampling; if the dimension is too low, methods such as bilinear interpolation will be used for upsampling to ensure that its dimension matches the requirements of the inverse wavelet transform.

[0049] After the dimensionality adjustment is completed, the system will process the enhanced fusion feature matrix using the inverse wavelet transform algorithm. Taking the fourth-order discrete inverse wavelet transform as an example, it will combine the low-frequency approximation signal and the three-layer high-frequency detail signals in the enhanced fusion feature matrix according to a certain weight and order. The low-frequency approximation signal plays a fundamental framework role in the synthesis process, determining the main trend and general outline of the audio. The system will accurately restore the basic timbre and rhythm of the audio based on the characteristic information of the low-frequency part in the enhanced fusion feature matrix, while the three-layer high-frequency detail signals will add rich details to the audio. The frequencies of the high-frequency detail signals increase in sequence, corresponding to different levels of detail information in the audio. The system will precisely synthesize these high-frequency detail signals with the low-frequency approximation signal, enabling the synthesized audio to retain the subtle features of the original audio, such as the overtones during instrument playing and the sibilance in speech.

[0050] During the inverse wavelet transform process, the system will utilize the fast inverse wavelet transform algorithm to improve the calculation efficiency. The fast inverse wavelet transform algorithm performs fast calculations on the enhanced fusion feature matrix through a specific filter bank, greatly reducing the amount of calculation and calculation time. This algorithm can quickly convert the enhanced fusion feature matrix into a preliminary audio signal while ensuring the calculation accuracy. Through the inverse wavelet transform, the system successfully converts the enhanced fusion feature matrix into a preliminary audio signal. This preliminary audio signal already contains the enhanced audio features, showing significant improvements in sound quality, details, and overall quality.

[0051] S111. Perform noise reduction processing on the preliminary audio signal to determine the final audio signal.

[0052] After the audio enhancement system generates the preliminary audio signal, to further improve the audio quality, it is necessary to perform noise reduction processing on it to remove the noise introduced during the enhancement process, thereby obtaining the final high-quality audio signal. The system will obtain the signal-to-noise ratio (SNR) of the preliminary audio signal. The signal-to-noise ratio is an indicator that measures the ratio of the strength of the useful signal to the strength of the noise in the audio signal. The system calculates the signal-to-noise ratio of the preliminary audio signal through a specific algorithm. First, the system will sample the preliminary audio signal to obtain a series of audio sample values. Then, according to the calculation methods of signal power and noise power, the power of the audio signal and the power of the noise are calculated respectively. The signal power can be obtained by calculating the sum of the squares of the audio sample values, while the noise power needs to be estimated through some denoising algorithms or statistical methods.

[0053] The system uses an adaptive noise reduction algorithm based on Wiener filtering to perform noise reduction on the preliminary audio signal. Wiener filtering is an optimal linear filter under the minimum mean square error criterion. It can design the filter coefficients according to the statistical characteristics of the signal and noise, so as to achieve the best noise reduction effect. In this system, the coefficients of the Wiener filter are adjusted according to the signal-to-noise ratio of the preliminary audio signal. When the signal-to-noise ratio is low, it means that the noise accounts for a relatively large proportion in the audio signal. At this time, the Wiener filter will increase the suppression of the noise, that is, adjust the filter coefficients to make the filter attenuate the noise more; when the signal-to-noise ratio is high, it means that the useful signal strength is large and the noise is relatively small. The Wiener filter will appropriately reduce the processing strength of the signal to avoid losing the useful information of the audio signal due to excessive noise reduction.

[0054] In practical applications, the adaptive noise reduction algorithm of Wiener filtering will process the preliminary audio signal frame by frame. The system divides the preliminary audio signal into multiple frames, and each frame contains a certain number of audio samples. For each frame of audio signal, the corresponding Wiener filter coefficients are calculated according to the signal-to-noise ratio of this frame, and then these coefficients are used to filter this frame of audio signal. During the filtering process, the Wiener filter will process signals of different frequencies to different extents according to the spectral characteristics of the signal and noise. For the frequency band where the noise is concentrated, the filter will increase the attenuation; for the frequency band where the useful signal is rich, the filter will try to maintain the integrity of the signal. By filtering each frame of audio signal, the noise in the preliminary audio signal is gradually removed.

[0055] In the embodiments of this application, due to the adoption of technologies such as splitting the audio signal, extracting multi-band features, constructing a specific generative adversarial network, deeply fusing the feature matrix, iteratively optimizing the network parameters, secondary enhancement and noise reduction, etc., the potential of each frequency band of the audio is fully explored. This effectively solves the problems in the prior art such as poor audio enhancement effect, easy loss of details and introduction of noise, and further realizes the generation of high-quality, low-noise master-level audio, improves the overall quality of the audio, and meets the purpose of professional production requirements.

[0056] After combining the above content, the following is a further and more specific process description of the method provided in this embodiment. Please refer to Figure 2 , which is another process schematic diagram of the master-level audio enhancement method based on the generative adversarial network in the embodiments of this application.

[0057] S201. Perform three-dimensional space mapping on the low-frequency enhancement feature matrix, the middle-frequency feature matrix, and the high-frequency enhancement feature matrix, and generate a dynamic fusion path through the Lorenz chaotic system; Before fusing the low-frequency, middle-frequency, and high-frequency enhancement feature matrices, the audio enhancement system will first perform three-dimensional space mapping on them and generate a dynamic fusion path with the help of the Lorenz chaotic system.

[0058] For three-dimensional space mapping, the system designs appropriate linear transformation matrices according to the dimensions and data characteristics of the feature matrices. Let the dimension of the low-frequency enhanced feature matrix L be (m, n, p), the dimension of the middle-frequency enhanced feature matrix M be (q, r, s), and the dimension of the high-frequency enhanced feature matrix H be (t, u, v). The system will obtain the transformation matrices for each matrix through calculation. Taking the low-frequency enhanced feature matrix L as an example, after the transformation , it is mapped into a unified three-dimensional space to make the three matrices consistent in space, facilitating subsequent fusion operations. This mapping is not random but the optimal transformation method obtained based on the research on the distribution law of audio features in different frequency bands and a large number of experiments. When generating the dynamic fusion path, the system uses the Lorenz chaotic system. The Lorenz chaotic system is a deterministic nonlinear system, and its equation is: . The system will set the initial values , , , as well as the system parameters , , . The selection of these initial values and parameters will be determined according to the type and characteristics of the audio and the previous training results. For example, for music audio with a strong rhythm and soothing classical music audio, their initial values and parameters will be different. By iteratively calculating the Lorenz equation, the system obtains a series of chaotic values, which will be mapped into the parameter space of the fusion path to form a dynamically changing fusion path. For example, the x value in the chaotic values is mapped as the position parameter of the fusion path in a certain dimension, and the y value and z value are respectively mapped as the relevant parameters of other dimensions. The dynamically generated fusion path can provide diverse fusion methods for the subsequent fusion of feature matrices, fully considering the complexity and diversity of audio features. 、 、 Taking the low-frequency enhanced feature matrix L as an example, after the transformation , it is mapped into a unified three-dimensional space to make the three matrices consistent in space, facilitating subsequent fusion operations. This mapping is not random but the optimal transformation method obtained based on the research on the distribution law of audio features in different frequency bands and a large number of experiments. When generating the dynamic fusion path, the system uses the Lorenz chaotic system. The Lorenz chaotic system is a deterministic nonlinear system, and its equation is: . The system will set the initial values , , , as well as the system parameters , , . The selection of these initial values and parameters will be determined according to the type and characteristics of the audio and the previous training results. For example, for music audio with a strong rhythm and soothing classical music audio, their initial values and parameters will be different. By iteratively calculating the Lorenz equation, the system obtains a series of chaotic values, which will be mapped into the parameter space of the fusion path to form a dynamically changing fusion path. For example, the x value in the chaotic values is mapped as the position parameter of the fusion path in a certain dimension, and the y value and z value are respectively mapped as the relevant parameters of other dimensions. The dynamically generated fusion path can provide diverse fusion methods for the subsequent fusion of feature matrices, fully considering the complexity and diversity of audio features. The system will set the initial values 、 、 as well as the system parameters 、 、 The selection of these initial values and parameters will be determined according to the type and characteristics of the audio and the previous training results. For example, for music audio with a strong rhythm and soothing classical music audio, their initial values and parameters will be different. By iteratively calculating the Lorenz equation, the system obtains a series of chaotic values, which will be mapped into the parameter space of the fusion path to form a dynamically changing fusion path. For example, the x value in the chaotic values is mapped as the position parameter of the fusion path in a certain dimension, and the y value and z value are respectively mapped as the relevant parameters of other dimensions. The dynamically generated fusion path can provide diverse fusion methods for the subsequent fusion of feature matrices, fully considering the complexity and diversity of audio features.

[0059] S202. Calculate the spectral entropy values of the low-frequency enhanced feature matrix, the middle-frequency feature matrix, and the high-frequency enhanced feature matrix. The spectral entropy value is used to measure the spectral complexity of each feature matrix. The audio enhancement system needs to calculate the spectral entropy values of the low-frequency, middle-frequency, and high-frequency enhanced feature matrices to better determine their weights in the fusion process.

[0060] The system will perform spectral analysis on the low-frequency enhanced feature matrix L. Through Fourier transform, the matrix is converted from the time domain to the frequency domain to obtain its spectral distribution , where f represents frequency. Then, the probability distribution is calculated according to the spectral distribution, that is, . After obtaining the probability distribution, the information entropy formula is used. where f represents frequency. Then, the probability distribution is calculated according to the spectral distribution, that is, . After obtaining the probability distribution, the information entropy formula Calculate the spectral entropy value of the low-frequency enhanced feature matrix , this spectral entropy value reflects the complexity of the spectrum of the low-frequency enhanced feature matrix.

[0061] For the intermediate-frequency enhanced feature matrix M and the high-frequency enhanced feature matrix H, the system uses the same method. First, perform Fourier transform on the intermediate-frequency enhanced feature matrix M to obtain the spectral distribution , calculate the probability distribution , and then obtain the spectral entropy value ; perform the same operation on the high-frequency enhanced feature matrix H to obtain the spectral entropy value . By calculating these three spectral entropy values, the system can quantitatively understand the spectral complexity of each feature matrix. A high spectral entropy value means high diversity and complexity of the features in that frequency band; a low spectral entropy value indicates relatively single features. These spectral entropy values provide an important basis for subsequent dynamic adjustment of the fusion weights, enabling the fusion process to perform more reasonable weight allocation according to the actual situation of the features in each frequency band, thereby improving the fusion effect and making the enhanced audio better retain the details and characteristics of each frequency band.

[0062] S203. According to the current position information on the dynamic fusion path, construct a weight adjustment mapping function, which associates the position parameters of the dynamic fusion path with the preset weight adjustment parameters; After the audio enhancement system obtains the current position information of the dynamic fusion path, it will start to construct a weight adjustment mapping function. Suppose the position of the dynamic fusion path in three-dimensional space is represented by parameters (x, y, z), and these parameters change continuously with the iteration of the Lorenz chaotic system. The system has preset a series of weight adjustment parameters, such as α, β, γ, etc., which are determined based on the experience and research of different audio feature fusions.

[0063] The system uses a neural network architecture to construct the mapping function. Taking a simple feedforward neural network as an example, the input layer receives the position parameters (x, y, z) of the dynamic fusion path, and the neurons in the input layer pass these parameters to the hidden layer. The hidden layer contains multiple neurons, and each neuron is connected to the input layer through weights. The neurons in the hidden layer use activation functions to enhance the expression ability of the function. After being processed by the hidden layer, the neurons in the output layer perform weighted summation on the output of the hidden layer according to the connection weights to obtain the weight adjustment values corresponding to the low-frequency, intermediate-frequency, and high-frequency enhanced feature matrices.

[0064] S204. Input the spectral entropy value into the weight adjustment mapping function, and perform weighted correction on the spectral entropy values of each feature matrix through the mapping function to obtain the corrected spectral entropy values; After the audio enhancement system constructs the weight adjustment mapping function, it will use the spectral entropy values of the low-frequency, intermediate-frequency, and high-frequency enhanced feature matrices , , Input this function for weighted correction. The system combines the spectral entropy value and the dynamic fusion path position information into an input vector , which is used as the input of the weight adjustment mapping function. The neurons in the mapping function process the input vector according to the connection weights. In the hidden layer, each element of the input vector is multiplied by the corresponding weight and summed, and then processed by the activation function. For example, the input of a certain neuron in the hidden layer is , and after being processed by the ReLU activation function, it gets , where is the connection weight, is the element of the input vector, is the bias.

[0065] After a series of processes in the hidden layer, the output layer calculates the corrected spectral entropy value. For the low-frequency enhancement feature matrix, the corrected spectral entropy value is obtained by the weighted sum of the output layer neurons, that is , and are the output layer connection weight and bias respectively. Similarly, the corrected spectral entropy values of the middle-frequency and high-frequency enhancement feature matrices and can be obtained. These corrected spectral entropy values comprehensively consider the spectral complexity of the feature matrix and the position information of the dynamic fusion path, laying a foundation for calculating more reasonable dynamic fusion weights in the subsequent calculation, so that the fusion weights can better reflect the importance of each feature matrix in the current fusion state.

[0066] S205. Based on the corrected spectral entropy value, calculate the corresponding dynamic fusion weights of the low-frequency enhancement feature matrix, the middle-frequency feature matrix, and the high-frequency enhancement feature matrix through the normalization function; After the audio enhancement system obtains the corrected spectral entropy value, it calculates the dynamic fusion weight by the normalization method. The Softmax function is used for normalization. This function can turn different spectral entropy values into weights with a sum of 1. First, calculate the exponential values of the corrected spectral entropy values of the low-frequency, middle-frequency, and high-frequency feature matrices respectively. These exponential values can amplify the differences between the entropy values. Then add these exponential values to get the sum. Divide the exponential value of each feature matrix spectral entropy value by the sum to get the corresponding dynamic fusion weight. For example, the dynamic fusion weight of the low-frequency enhancement feature matrix is the ratio of the exponential value of its spectral entropy value to the sum. In this way, the feature matrix with a larger spectral entropy value has a larger corresponding weight and can play a greater role in the fusion, ensuring a reasonable distribution of the weights of each feature matrix in the fusion process.

[0067] S206. According to the dynamic fusion weights, perform weighted superposition on the low-frequency enhanced feature matrix, the intermediate-frequency feature matrix, and the high-frequency enhanced feature matrix, and combine the chaotic perturbation terms at the corresponding positions on the dynamic fusion path for non-linear combination to generate a fusion feature matrix.

[0068] The audio enhancement system processes the three feature matrices based on the calculated dynamic fusion weights. Multiply the low-frequency, intermediate-frequency, and high-frequency enhanced feature matrices by their respective weights to obtain the weighted matrices, and then stack them up to form a preliminary fusion matrix, which integrates the features of each frequency band.

[0069] To make the fusion effect better, the system obtains the chaotic perturbation term at the current position from the dynamic fusion path generated by the Lorenz chaotic system. Use the Sigmoid function to perform non-linear combination on the preliminary fusion matrix and the chaotic perturbation term. The Sigmoid function maps the input value to the range between 0 and 1, and it performs different degrees of transformation according to the size of the input value. Input the sum of the preliminary fusion matrix and the chaotic perturbation term into the Sigmoid function to obtain the final fusion feature matrix. In this way, the chaotic perturbation term is used to increase the randomness and diversity of the fusion, avoid information loss caused by simple superposition, enable better coordination of the features of each frequency band of the audio, and improve the audio quality.

[0070] In the embodiments of the present application, due to the adoption of technical means such as three-dimensional space mapping, generating a dynamic fusion path by the Lorenz chaotic system, calculating the spectral entropy value, constructing a weight adjustment mapping function, normalizing the calculation of dynamic fusion weights, and combining chaotic perturbation terms for non-linear combination, it is possible to fully consider the characteristics and dynamic changes of the feature matrices of each frequency band of the audio, and dynamically and reasonably adjust the fusion weights. This effectively solves the problems of single feature matrix fusion method and unreasonable weight allocation in the prior art, and further realizes the generation of a more diverse and adaptable fusion feature matrix, improves the coordination of the features of each frequency band of the audio, further optimizes the audio enhancement effect, and enables the enhanced audio to achieve the delicate texture and rich details required for the master tape level.

[0071] The audio enhancement system in the embodiments of the present invention application will be described from the perspective of hardware processing. Please refer to Figure 3 , which is a schematic structural diagram of an entity device of the audio enhancement system in the embodiments of the present application.

[0072] It should be noted that Figure 3 The structure of the audio enhancement system shown is only an example, and should not bring any limitations to the functions and usage scopes of the embodiments of the present invention.

[0073] As Figure 3As shown, the audio enhancement system includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 302 or the program loaded from the storage section 308 into the Random Access Memory (RAM) 303, such as executing the methods described in the above embodiments. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.

[0074] The following components are connected to the I / O interface 305: an input section 306 including an audio input device, button switches, etc.; an output section 307 including a Liquid Crystal Display (LCD), an audio output device, indicator lights, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read from it can be installed into the storage section 308 as needed.

[0075] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by the Central Processing Unit (CPU) 301, various functions defined in the present invention are executed.

[0076] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0077] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings.

[0078] Specifically, the audio enhancement system of this embodiment includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, it implements the master tape-level audio enhancement method based on the generative adversarial network provided in the above-mentioned embodiment.

[0079] On the other hand, the present invention also provides a computer-readable storage medium. This storage medium can be included in the audio enhancement system described in the above-mentioned embodiment; or it can exist separately without being assembled into the audio enhancement system. The above storage medium carries one or more computer programs. When the above one or more computer programs are executed by a processor of an audio enhancement system, the audio enhancement system is enabled to implement the master tape-level audio enhancement method based on the generative adversarial network provided in the above-mentioned embodiment.

[0080] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present application.

[0081] As used in the foregoing embodiments, depending on the context, the term "when" may be construed to mean "if" or "after" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "when determining" or "if detecting (the stated condition or event)" may be construed to mean "if determining" or "in response to determining" or "when detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)".

[0082] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the foregoing embodiments can be implemented by a computer program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the foregoing method embodiments. The foregoing storage media include: various media such as ROM or random access memory RAM, magnetic disks, or optical discs that can store program codes.

Claims

1. A master tape-level audio enhancement method based on a generative adversarial network, characterized in that The method includes: Segmenting the master audio signal to be enhanced according to a preset time window to obtain a plurality of audio segments; Performing preprocessing operations on the plurality of audio segments to obtain a low-frequency feature matrix, a mid-frequency feature matrix, and a high-frequency feature matrix; Constructing a generative adversarial network, which consists of a main generator, an auxiliary generator, and a triple discriminator. The main generator adopts a recursive convolutional structure based on a gated recurrent unit, the auxiliary generator is a multi-branch attention mechanism network, and the triple discriminator is a time-domain discriminator, a frequency-domain discriminator, and a phase discriminator respectively; Inputting the low-frequency feature matrix into the main generator to generate a low-frequency enhanced feature matrix, and inputting the high-frequency feature matrix into the auxiliary generator to generate a high-frequency enhanced feature matrix; Performing deep fusion on the low-frequency enhanced feature matrix, the mid-frequency feature matrix, and the high-frequency enhanced feature matrix to generate a fused feature matrix; After converting the fused feature matrix into a time-domain signal, a frequency-domain signal, and a phase signal respectively, inputting them into the corresponding triple discriminator to obtain three corresponding discriminant scores; Combining the three discriminant scores into a three-dimensional adversarial loss vector including time-domain, frequency-domain, and phase dimensions; Combining the three-dimensional adversarial loss vector, and iteratively optimizing the network parameters of the main generator and the auxiliary generator through a double-nested adaptive feedback regulation system; According to the optimized parameters, performing secondary enhancement processing on the fused feature matrix to determine an enhanced fused feature matrix; Generating a preliminary audio signal by performing inverse wavelet transform on the enhanced fused feature matrix; Performing noise reduction processing on the preliminary audio signal to determine the final audio signal.

2. The method according to claim 1, wherein In the step of performing preprocessing operations on the plurality of audio segments to obtain a low-frequency feature matrix, a mid-frequency feature matrix, and a high-frequency feature matrix, it specifically includes: Performing a fourth-order discrete wavelet transform on the plurality of audio segments to obtain a plurality of signal groups including low-frequency approximation signals and three layers of high-frequency detail signals; For each signal group, adopting a frequency-domain analysis method based on the Mel scale to extract a low-frequency feature matrix, a mid-frequency feature matrix, and a high-frequency feature matrix respectively, where the low-frequency feature matrix represents the basic timbre and rhythm information of the audio, the mid-frequency feature matrix represents the transitional frequency band texture information, and the high-frequency feature matrix represents the detail texture information.

3. The method according to claim 1, characterized in that, In the step of performing deep fusion on the low-frequency enhanced feature matrix, the mid-frequency feature matrix, and the high-frequency enhanced feature matrix to generate a fused feature matrix, it specifically includes: Performing three-dimensional space mapping on the low-frequency enhanced feature matrix, the mid-frequency feature matrix, and the high-frequency enhanced feature matrix, and generating a dynamic fusion path through a Lorenz chaotic system; Combining the dynamic fusion path, dynamically adjusting the fusion weights of the three feature matrices to determine the fused feature matrix.

4. The method according to claim 3, wherein In the step of combining the dynamic fusion path, dynamically adjusting the fusion weights of the three feature matrices to determine the fused feature matrix, it specifically includes: Calculating the spectral entropy values of the low-frequency enhanced feature matrix, the mid-frequency feature matrix, and the high-frequency enhanced feature matrix, where the spectral entropy values are used to measure the spectral complexity of each feature matrix; Construct a weight adjustment mapping function according to the current position information on the dynamic fusion path, where the mapping function associates the position parameter of the dynamic fusion path with a preset weight adjustment parameter; Input the spectral entropy value into the weight adjustment mapping function, and use the mapping function to perform weighted correction on the spectral entropy values of each feature matrix to obtain the corrected spectral entropy values; Based on the corrected spectral entropy values, calculate the dynamic fusion weights corresponding to the low-frequency enhanced feature matrix, the intermediate-frequency feature matrix, and the high-frequency enhanced feature matrix through a normalization function; According to the dynamic fusion weights, perform weighted superposition on the low-frequency enhanced feature matrix, the intermediate-frequency feature matrix, and the high-frequency enhanced feature matrix, and combine the chaotic perturbation term at the corresponding position on the dynamic fusion path for non-linear combination to generate a fused feature matrix.

5. The method according to claim 1, wherein The steps of inputting the fused feature matrix into the corresponding triple discriminator after converting it into a time-domain signal, a frequency-domain signal, and a phase signal respectively to obtain three corresponding discriminant scores specifically include: Input the time-domain signal into the time-domain discriminator, and based on a bidirectional long short-term memory network, obtain the time-domain continuity score by calculating the cosine similarity of adjacent time frames; Input the frequency-domain signal into the frequency-domain discriminator, and use Fourier descriptors to analyze the spectral smoothness, and use the curvature change rate of the spectral curve as the frequency-domain smoothness score; Input the phase signal into the phase discriminator, and evaluate the feature synergy by calculating the phase-amplitude mutual information to obtain the phase synergy score.

6. The method according to claim 1, characterized in that The steps of iteratively optimizing the network parameters of the main generator and the auxiliary generator through a double-nested adaptive feedback control system in combination with the three-dimensional adversarial loss vector specifically include: The double-nested adaptive feedback control system includes an outer feedback loop and an inner feedback loop, and iteratively optimizes the network parameters of the main generator and the auxiliary generator as well as the Lorenz system parameters of the three-dimensional feature chaotic fusion mechanism through multiple iterations; In each iteration of the outer feedback loop, based on the norm value of the three-dimensional adversarial loss vector, adjust the network parameters of the main generator and the auxiliary generator through a proximal policy optimization algorithm, and the norm value is used to measure the comprehensive deviation degree of the generated audio features from the ideal state in the three dimensions of time domain, frequency domain, and phase; In each iteration of the inner feedback loop, for the local peaks of each dimension of the three-dimensional adversarial loss vector, use reinforcement learning to dynamically adjust the Lorenz system parameters of the three-dimensional feature chaotic fusion mechanism.

7. The method according to claim 1, characterized in that, The steps of performing noise reduction processing on the preliminary audio signal to determine the final audio signal specifically include: Obtain the signal-to-noise ratio of the preliminary audio signal; Adopt an adaptive noise reduction algorithm based on Wiener filtering, adjust the filtering coefficient according to the signal-to-noise ratio, and eliminate the noise introduced during the enhancement process to obtain the final audio signal.

8. An audio enhancement system, characterized in that, The audio enhancement system includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the audio enhancement system to execute the method described in any one of claims 1-7.

9. A computer-readable storage medium, comprising instructions, characterized in that, When the instructions run on the audio enhancement system, it causes the audio enhancement system to execute the method described in any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product runs on the audio enhancement system, it causes the audio enhancement system to execute the method described in any one of claims 1-7.