Audio information processing method and apparatus, device, and storage medium
Patent Information
- Application Number
- PCT/CN2026/078571
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-14
- Filing Date
- 2026-02-11
- Publication Date
- 2026-09-17
Smart Images

Figure CN2026078571_17092026_PF_FP_ABST
Abstract
Description
Methods, apparatus, equipment and storage media for processing audio information
[0001] This disclosure claims priority to Chinese Patent Application No. 202510319577.7, filed on March 14, 2025, entitled “Method, Apparatus, Device and Storage Medium for Processing Audio Information”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and storage medium for processing audio information.
[0003] Background of the Invention
[0004] With the rapid development of computer technology, people's demands for audio information quality are constantly increasing. To avoid noise in audio information, filtering and other methods are commonly used to remove noise and improve audio quality. However, noise removal may damage the desired sound or be ineffective. Further improving the quality of audio information is an urgent problem to be solved. Summary of the Invention
[0005] This disclosure provides a method, apparatus, device, and storage medium for processing audio information.
[0006] A method for processing audio information, executed by a computer device, the method comprising:
[0007] Obtain sample audio pairs, which include audio with the same audio content but different audio quality, and audio with sound quality enhancement;
[0008] Feature extraction is performed on the audio with sound quality degradation and the audio with sound quality enhancement respectively to obtain degradation features and enhancement features;
[0009] Audio restoration is performed on the damage features by a restoration network to obtain restored features; the restoration network is trained based on the difference between the restored features and the enhanced features to obtain a trained restoration network.
[0010] The audio decoder performs audio decoding on the impairment feature and the enhancement feature respectively to obtain impairment-decoded audio and enhancement-decoded audio; the audio decoder is trained based on the difference between the impairment-decoded audio and the audio quality impairment audio and the difference between the enhancement-decoded audio and the audio quality enhancement audio to obtain a trained audio decoder;
[0011] The repair network and the audio decoder are provided for audio processing. The repair network is used to repair the input audio, and the audio decoder decodes the audio output by the repair network to obtain the output audio.
[0012] An audio information processing device, comprising:
[0013] The acquisition module is used to acquire sample audio pairs, which include audio with the same audio content but different sound quality, and audio with sound quality enhancement.
[0014] The processing module is used to perform feature extraction on the audio with damaged sound quality and the audio with enhanced sound quality, respectively, to obtain damaged features and enhanced features; the processing module is also used to perform audio restoration on the damaged features through a restoration network to obtain restored features;
[0015] A training module is used to train the repair network based on the differences between the repair features and the enhancement features, to obtain the trained repair network; and
[0016] A module is provided to provide the repair network and the audio decoder for audio processing. The repair network is used to repair the input audio, and the audio decoder decodes the audio output from the repair network to obtain the output audio.
[0017] The processing module is further configured to perform audio decoding on the damage feature and the enhancement feature respectively through an audio decoder to obtain damage-decoded audio and enhancement-decoded audio;
[0018] The training module is further configured to train the audio decoder based on the difference between the damaged decoded audio and the audio with reduced sound quality, and the difference between the enhanced decoded audio and the audio with enhanced sound quality, to obtain the trained audio decoder.
[0019] A computer device includes a processor and a memory, the memory storing at least one instruction, at least one program, a code set, or an instruction set, the at least one instruction, the at least one program, the code set, or the instruction set being loaded and executed by the processor to implement the audio information processing method described above.
[0020] A computer-readable storage medium storing at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the audio information processing method described above.
[0021] A computer program product includes computer instructions stored in a computer-readable storage medium, wherein a processor reads from and executes the computer instructions to implement the audio information processing method described above.
[0022] This disclosure repairs the feature representation of audio information by training a neural network, and converts the feature representation of audio information into audio information by using an audio decoder. By training the repair network and the audio decoder separately, the audio repair capability of the repair network and the decoding capability of the audio decoder are made independent of each other. The repair network and the audio decoder are trained on sample audio pairs from the same data source, which ensures that the processing capabilities of the trained repair network and the trained audio decoder are matched, and ensures the accuracy of the audio decoder in performing audio decoding on the input audio features, thereby ensuring the quality of the predicted audio information.
[0023] Brief description of the attached figures
[0024] To more clearly illustrate the technical solutions in this disclosure, the accompanying drawings used in the example description will be briefly introduced below. Obviously, the accompanying drawings described below are only some examples of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0025] Figure 1 is a schematic diagram of an example computer system of this disclosure;
[0026] Figure 2 is a schematic diagram of an example of an audio information processing method disclosed herein;
[0027] Figure 3 is a flowchart of an example of an audio information processing method disclosed herein;
[0028] Figure 4 is a flowchart of an example of an audio information processing method disclosed herein;
[0029] Figure 5 is a schematic diagram of an example repair network of this disclosure;
[0030] Figure 6 is a schematic diagram of an example repair network of this disclosure;
[0031] Figure 7 is a flowchart of an example of an audio information processing method disclosed herein;
[0032] Figure 8 is a flowchart of an example of an audio information processing method disclosed herein;
[0033] Figure 9 is a schematic diagram of an example audio feature of this disclosure;
[0034] Figure 10 is a structural block diagram of an audio information processing apparatus according to an example of this disclosure;
[0035] Figure 11 is a block diagram of an example server of this disclosure.
[0036] The accompanying drawings, which are incorporated in and form part of this specification, illustrate various examples consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0037] Methods of implementing the present invention
[0038] To make the objectives, technical solutions, and advantages of this disclosure clearer, the embodiments of this disclosure will be described in further detail below with reference to the accompanying drawings.
[0039] Exemplary examples will now be described in detail, as illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary examples do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0040] The terminology used in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting of this disclosure. The singular forms “a,” “the,” and “the” used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0041] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, and displayed data) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the sample audio pairs and network parameters of the repair network involved in this disclosure were obtained with full authorization.
[0042] It should be understood that although the terms first, second, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, a first parameter may also be referred to as a second parameter without departing from the scope of this disclosure, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0043] Figure 1 shows a schematic diagram of an example computer system of this disclosure. This computer system can implement a system architecture for a method of processing audio information. The computer system may include: a terminal 100 and a server 200.
[0044] Terminal 100 can be an electronic device such as a mobile phone, tablet computer, or PC (Personal Computer). Client applications can be installed and run on terminal 100. These applications can be audio information processing applications, application management applications, or other applications that provide audio information processing functions; this disclosure does not limit the specific application. Furthermore, this disclosure does not limit the form of the application, including but not limited to Apps (Applications), applets, etc., installed on terminal 100, and can also be web-based applications (i.e., Web applications).
[0045] Server 200 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services. Server 200 can be the backend server for the aforementioned application, used to provide backend services to the application's clients.
[0046] The audio information processing method provided in this disclosure can be executed by a computer device, which refers to an electronic device with data computing, processing, and storage capabilities. Taking the implementation environment shown in Figure 1 as an example, the audio information processing method can be executed by a terminal 100 (e.g., by a client of an application installed and running on the terminal 100), by a server 200, or by the terminal 100 and the server 200 through interactive cooperation. This disclosure does not limit the scope of the method.
[0047] Furthermore, the technical solution disclosed herein can be combined with blockchain technology. For example, some data involved in the audio information processing method disclosed herein (such as sample audio pairs, model parameters of repair networks, etc.) can be stored on a blockchain. Terminal 100 and server 200 can communicate via a network, such as a wired or wireless network.
[0048] Figure 2 shows a schematic diagram of an example of an audio information processing method disclosed herein.
[0049] Obtain initial sample audio 300, perform sound effect enhancement 312 on initial sample audio 300 to obtain sound quality enhanced audio 302; and perform damage simulation 314 on initial sample audio 300 to obtain sound quality damaged audio 304.
[0050] The enhanced audio 302 obtained through sound effect enhancement and the damaged audio 304 obtained through damage simulation can also be collectively referred to as a sample audio pair. Among them, the audio quality of the enhanced audio 302 is higher than that of the damaged audio 304.
[0051] Feature extraction 320 is performed on the audio with enhanced sound quality 302 and the audio with degraded sound quality 304 respectively, to obtain the enhancement feature 322 corresponding to the audio with enhanced sound quality 302 and the degraded feature 324 corresponding to the audio with degraded sound quality 304. For example, the above two feature information are at least one of the deep features extracted from the audio information, such as Mel spectrum, amplitude features, etc.
[0052] The repair network 330 is invoked to perform audio repair on the damaged feature 324, that is, the damaged feature 324 is used as the input parameter of the repair network 330, and the repair feature 332 predicted by the repair network 330 is output. Based on the difference between the repair feature 332 and the enhanced feature 322, a first loss function 335 is constructed, and the repair network 330 is trained by backpropagation to adjust the network parameters of the repair network 330. By adjusting the network parameters of the repair network 330, the ability of the repair network 330 to perform audio repair on feature information is improved, and the difference between the predicted feature information (such as the repair feature 332) and the desired audio information features (such as the enhanced feature 322) is reduced, resulting in the trained repair network.
[0053] The audio decoder 340 is invoked to perform audio decoding on the enhanced feature 322 and the damaged feature 324 respectively, that is, the enhanced feature 322 and the damaged feature 324 are used as input parameters of the audio decoder 340. The audio decoder 340 outputs the enhanced decoded audio 342 corresponding to the predicted enhanced feature 322 and the damaged decoded audio 344 corresponding to the damaged feature 324. Based on the difference between the enhanced decoded audio 342 and the sound quality enhanced audio 302, and the difference between the damaged decoded audio 344 and the sound quality damaged audio 304, a second loss function 345 is constructed, and backpropagation training is performed on the audio decoder 340 to adjust the network parameters of the audio decoder 340. By adjusting the network parameters of the audio decoder 340, the ability of the audio decoder 340 to decode feature information is improved, the difference between the predicted audio information and the expected audio information is reduced, and the decoding accuracy is improved, resulting in the trained audio decoder.
[0054] Figure 3 shows a flowchart of an example method for processing audio information according to this disclosure. This method can be performed by a computer device. The method includes the following steps.
[0055] Step 510: Obtain sample audio pairs.
[0056] Sample audio pairs are used to train audio decoders and repair networks. These pairs include audio with corresponding audio content (i.e., the same audio content but different sound quality) and audio with enhanced sound quality; the enhanced sound quality is higher than the audio with corresponding audio content. The corresponding audio content means that the audio with corresponding audio content carries the same core information, regardless of sound quality, noise, or other interference factors. For example, for speech audio, both audio with corresponding audio content express the same text, semantics, or spoken content. For instance, the phrase "The meeting is at 10 AM tomorrow" spoken by the same person, whether in a noisy, reverb-filled version or a noise-reduced, reverb-free enhanced version, has the same core semantics. Similarly, for musical audio, audio with corresponding audio content shares the same core musical information such as melody, rhythm, and notes; one may have noise or distortion, while the other is clearer and purer. In other words, the difference between audio with corresponding audio content and enhanced sound quality lies only in sound clarity, noise levels, and distortion.
[0057] For example, degraded audio and enhanced audio correspond to each other in terms of audio content, meaning that they are derived from the same original audio. The methods for obtaining degraded and enhanced audio by performing different audio quality adjustments (such as degraded processing, enhanced processing, etc.) on the source audio (or original audio) will be described with separate examples below. The original audio is the common source signal of the two audio information in the sample audio pair, also known as the initial sample audio.
[0058] Taking a sample audio pair where the audio is speech information as an example, the correspondence between degraded and enhanced speech in terms of audio content means that they have the same semantics. Similarly, taking a sample audio pair where the audio is musical information as an example, the correspondence between degraded and enhanced musical pieces in terms of audio content means that they have the same musical melody. For instance, degraded and enhanced audio have the same audio content.
[0059] For example, the audio with degraded sound quality and the audio with enhanced sound quality correspond to each other in terms of audio content. This includes performing subject extraction on the audio with degraded sound quality and the audio with enhanced sound quality in the sample audio pair, and obtaining the same subject audio information. For example, subject extraction is used to extract audio signals with significant auditory effects from the audio with degraded sound quality and the audio with enhanced sound quality, such as extracting human voice signals from speech information while excluding environmental noise from speech information; or extracting musical melodies from musical information while excluding echoes or metronome cue sounds from musical information.
[0060] Audio quality includes, but is not limited to, at least one of the following: audio quality is positively correlated with audio bitrate; audio quality is inversely correlated with the proportion of carried noise; and audio quality is inversely correlated with the proportion of data with waveform loss.
[0061] Step 520: Perform feature extraction on the audio with sound quality loss and the audio with sound quality enhancement respectively to obtain the loss features and enhancement features.
[0062] Feature extraction is used to extract deep features from audio information. For example, an Artificial Neural Network (ANN) can be used to extract feature information in the hidden space of audio with and without sound quality enhancement. Feature extraction can also extract feature information of audio in the time domain, frequency domain, etc., such as at least one of Mel spectrogram, magnitude feature, and filter bank (FBank) coefficients.
[0063] The feature extraction performed in this disclosure can be achieved by calling an artificial neural network, or by performing calculations on the audio (such as time-domain analysis, frequency-domain analysis, or convolution calculations on audio signals in the time or frequency domains). This disclosure does not limit the specific implementation method of feature extraction.
[0064] In one example, feature extraction for degraded and enhanced audio is performed independently. For instance, the degraded features are extracted from the degraded audio and are unrelated to the enhanced audio. However, it is possible that the feature extraction for degraded and enhanced audio may overlap temporally when performed in parallel.
[0065] Step 530: Perform audio restoration on the damage features using the restoration network to obtain restored features; based on the difference between the restored features and the enhanced features, train the restoration network to obtain the trained restoration network.
[0066] In one example, the repair network can be implemented as an artificial neural network with audio repair capabilities; the network architecture of the repair network includes, but is not limited to, at least one of the following: Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Transformer model, and Long Short-Term Memory (LSTM).
[0067] In some examples, the audio restoration model of the restoration network can be implemented by training the restoration network in a way that adjusts the network parameters. In this step, the input information of the restoration network is the damage features. The restoration network performs calculations on the damage features using the network parameters of the restoration network to perform audio restoration on the damage features and output restored features.
[0068] In some examples, the damage features and restoration features often belong to the same feature type, such as any one of the feature types described above, such as Mel spectrum, amplitude features, and filter bank coefficients. In other examples, the damage features and restoration features are of different types; for instance, the restoration network achieves audio restoration by converting one feature type to another.
[0069] In some examples, backpropagation of the repair network is performed based on the differences between the repair and enhancement features to adjust the network parameters, aiming to minimize these differences. By training the repair network, its audio restoration capabilities are improved.
[0070] Step 540: Perform audio decoding on the damage features and enhancement features respectively using the audio decoder to obtain damaged decoded audio and enhanced decoded audio; perform training on the audio decoder based on the differences between damaged decoded audio and audio quality damage audio, and the differences between enhanced decoded audio and audio quality enhancement audio, to obtain the trained audio decoder.
[0071] In some examples, the audio decoder has the ability to restore the audio information from the feature information extracted from the audio information. The audio decoding of impairment features and enhancement features is independent of each other; for example, impairment-decoded audio is obtained by performing audio decoding on impairment features, and is unrelated to enhancement features. However, it is not impossible that the audio decoding of impairment features and enhancement features may overlap in time when performed in parallel.
[0072] The trained repair network and audio decoder can be provided to at least one electronic device for audio processing, wherein the repair network is used to repair the input audio, and the audio decoder decodes the audio output by the repair network to obtain the output audio.
[0073] This disclosure does not limit the structure or training method of the audio decoder. In some examples, based on the differences between damaged decoded audio and audio with degraded sound quality, and the differences between enhanced decoded audio and audio with enhanced sound quality, backpropagation is performed on the audio decoder to adjust the network parameters of the audio decoder with the aim of minimizing the differences between damaged decoded audio and audio with degraded sound quality, and minimizing the differences between enhanced decoded audio and audio with enhanced sound quality. By training the audio decoder, the ability of the audio decoder to restore audio information from the feature information extracted from the audio information is improved.
[0074] In this example, the training processes for the repair network and the audio decoder are independent of each other. It should be noted that the independent training processes described in this step refer to the differences between different information when training the repair network and the audio decoder. During the training of one network, only the parameters of that network are adjusted; it does not restrict the training of the two network models to using the same sample information pairs.
[0075] In the various examples of this disclosure, there is no restriction on the timing of training the repair network and the audio decoder; training can be performed sequentially or in parallel, or the training of the repair network and the audio decoder can have partial overlap in time. In some examples, step 530 in this example can be performed before, after, or simultaneously with step 540; this disclosure does not impose any restrictions on this.
[0076] In summary, the method presented in this example achieves the restoration of the feature representation of audio information through a restoration network. Training the restoration network based on the difference between the restored and enhanced features improves its audio restoration capabilities. The method also converts the feature representation of audio information into audio information through an audio decoder, and training the decoder improves the accuracy of the decoded audio information. Furthermore, training the restoration network and the audio decoder separately ensures that their audio restoration and decoding capabilities are independent. Training both the restoration network and the audio decoder using sample audio pairs from the same data source guarantees that their audio processing capabilities are matched, ensuring the accuracy of the audio decoder in decoding the input audio features and thus guaranteeing the quality of the predicted audio information.
[0077] Figure 4 shows a flowchart of an audio information processing method provided by an exemplary example of this disclosure. This method can be executed by a computer device. Specifically, in the example shown in Figure 3, step 530 can be implemented as steps 532, 534, 536, and 538:
[0078] Step 532: Perform audio restoration on the damage features using the restoration generator to obtain the restored features and the data source labels of the restored features.
[0079] In this example, the repair network includes a repair generator and a repair discriminator. The repair generator performs audio repair on the input impairment features to obtain repaired features. The repair discriminator determines the source of the audio features, such as whether the audio features input to the repair discriminator were predicted by the repair generator or directly acquired.
[0080] In this step, the data source label indicates whether the restoration features were predicted by the restoration generator. For example, the data source label for the audio features of the original audio corresponding to the audio with sound quality degradation in the sample information pair may differ from the data source label for the restoration features. The data source label for the audio features of the original audio indicates that the audio features were acquired.
[0081] Step 534: Perform prediction on the repaired features using the repair discriminator to obtain the predicted source label of the repaired features.
[0082] As described above, the repair discriminator is used to determine the source of audio features; for example, it can determine whether the audio features input to the repair discriminator are predicted by the repair generator or directly acquired.
[0083] In some examples, the repair discriminator is also used to perform predictions on the audio features of directly acquired audio information to obtain corresponding predicted source labels. The repair discriminator is trained by comparing the predicted source labels with the directly acquired audio information to indicate that the audio features are the true source labels obtained from the acquisition.
[0084] Step 536: Based on the differences between the predicted source label and the data source label, and the differences between the repair features and the enhancement features, train the repair generator to obtain the trained repair generator.
[0085] In some examples, the difference between the predicted source label and the data source label is used to indicate how accurately the inpainting discriminator distinguishes the predicted source label. The smaller the difference, the higher the inpainting generator's ability to repair audio features. An inpainting generator with high repair capability can output feature information that causes the inpainting discriminator to incorrectly identify the data source.
[0086] The difference between the restored and enhanced features indicates the discrepancy between the feature information output by the restoration generator and the desired feature information. The smaller the difference, the higher the restoration generator's ability to restore audio features.
[0087] In some examples, the network parameters of the repair network are adjusted to reduce the difference between the repaired features and the enhanced features, and to increase the difference between the predicted source label and the data source label, so as to improve the repair network's ability to repair audio features.
[0088] For example, this step can be implemented as follows:
[0089] • The repair discriminator is trained based on the difference between the predicted source label and the data source label by fixing the network parameters of the repair generator.
[0090] In some examples, by fixing the network parameters of the repair generator while training the repair discriminator, the repair discriminator is trained separately, thereby improving its ability to discriminate the source of audio features.
[0091] • By fixing the network parameters of the repair discriminator, the repair generator is trained based on the difference between the repair features and the enhancement features to obtain the trained repair generator.
[0092] In some examples, the network parameters of the repair discriminator are fixed during the training of the repair generator. This allows for separate training of the repair generator, improving its ability to repair audio features.
[0093] It should be noted that there are no restrictions on the timing of training the repair discriminator and the repair generator in this example. The repair generator can be trained first, or the repair discriminator can be trained first; in some instances, the training of the repair generator and the repair discriminator can be performed alternately multiple times to achieve alternating improvement of the capabilities of the repair generator and the repair discriminator during the training process.
[0094] In one example, the differences between predicted source labels and data source labels, and the differences between repair features and enhancement features are as follows:
[0095] Where n represents the total number of samples in a batch of data, and E[*] represents the cross-entropy loss function; and Let i represent the i-th enhanced feature (also known as the clean Mel spectrum feature) and the i-th repaired feature (also known as the clean Mel spectrum estimate), respectively; W is the weight of the mean squared error, which may be an empirical parameter.
[0096] It is the mean squared error between enhanced and repaired features; to minimize the generation loss. The training criterion for the repair generator, also known as the loss function of the repair generator, aims to minimize the adversarial loss. As a training criterion for the discriminator, it is also called the loss function for repairing the discriminator.
[0097] Formula 1 is used to calculate the mean squared error loss between the enhanced features and the repaired features of this batch of samples. This loss serves as a representation of the difference between the repaired and enhanced features, and is used to drive the repair generator to produce speech representations that are closer to reality. In Formula 1, The squared difference between the enhanced and repaired features of a single sample is calculated to amplify the numerical deviation between them. Formula 1 sums the squared differences of the above formula for all n samples in a batch, and the average of the summations is taken as the mean squared error loss.
[0098] Formula 2 is used to calculate the total loss of the repair generator, constraining the repaired features output by the generator to be both numerically close to the real features and able to "fool" the discriminator, ultimately generating a high-quality speech representation. In Formula 2, It is the probability of the discriminator judging the repaired feature, that is, the probability that the discriminator considers it to be correct. It represents the probability of the features of the real sample, with a value range of [0,1]. It serves as a representation of the difference between the predicted source label and the data source label. The smaller the difference, the closer this value is to 1. The discriminator believes It represents the probability of generating sample features. Approaching 1, that is A value close to 0 indicates that the discriminator has identified... It generates sample features; Approaching 0, that is A value close to 1 indicates that the discriminator has been "deceived" and is therefore considered to be... These are features of real samples. The logarithm of the above probabilities is used to measure the degree of misclassification of the generated samples by the discriminator. The logarithms of the probabilities of all n samples within a batch are summed, and the average of the sums is used to measure the average misclassification rate of the discriminator for that batch of samples, serving as the adversarial loss. W is used to balance the importance of the adversarial loss and the mean squared error loss; its value can be adjusted based on the actual training effect. A larger W emphasizes making the generated features numerically closer to the real features. In the early stages of training... and The differences are significant. Close to 0, therefore Becoming the dominant loss, the forced generator makes Numerically close As training progresses, the generator... Under the constraints, repair features that approximate the enhancement features are gradually generated. It started to rise. Gradually decrease, It begins to shift towards negative infinity, and the constraint against the loss term gradually strengthens, with... Common constraint generator.
[0099] Formula 3 is used to measure the ability of the repair discriminator to distinguish between real sample features and generated features, so that the discriminator can accurately identify real sample features and generated features, forming a closed loop of adversarial training with the generator. It is the positive sample loss term, used to measure the accuracy of the discriminator in recognizing the features of real samples. It is the negative sample loss term, used to measure the accuracy of the discriminator in recognizing generated features.
[0100] During training, the generator strives to reduce To "deceive" the discriminator, the discriminator tries to reduce... This allows for the accurate differentiation between real sample features and generated features, ultimately achieving a Nash equilibrium between the two, with the generator outputting high-quality repaired features.
[0101] In some examples, the repair generator and repair discriminator update parameters alternately. For example, some examples use a Hanning window with a 16K sampling rate, 32ms frame length, and 20ms frame shift to compute a 100-dimensional Mel spectrum. For instance, the batch size of the data is set to 24, and the initial learning rate of the repair generator is set to 1×102. -4 The initial learning rate of the repair discriminator is set to 4×10. -4 The decay coefficient of the learning rate is 0.99 for all models. The maximum number of training iterations is set to 8,000,000. As described above, in some examples disclosed herein, the first learning rate used to train the repair discriminator is greater than the second learning rate used to train the repair generator. The learning rate is a hyperparameter in deep learning model training that controls the magnitude of each update of network parameters, and its value is between 0 and 1. Setting a smaller learning rate for the repair generator is to enable the repair generator to converge with a smaller gradient update magnitude than the repair discriminator, increasing the convergence stability of adversarial attack learning and avoiding the problem of output distortion and failure to fit real features due to excessively rapid updates.
[0102] Step 538: Determine the trained repair generator as the trained repair network.
[0103] In this example, the trained inpainting network includes a trained inpainting discriminator. In some examples, after training the inpainting network, only the trained inpainting generator is used to restore the audio's features, independent of the inpainting discriminator.
[0104] Figure 5 shows a schematic diagram of an example repair network of this disclosure. The structure of the network model shown in Figure 5 can be the model structure of the repair network described in Figure 3, or the model structure of the repair generator shown in Figure 4.
[0105] In some examples, the repair network is a transformer model used to perform audio repair on audio features; it is also called a conformer network. In some examples, the repair network includes: an initial linear layer 602, a transform network block 604, a first linear bridging layer 606, a second linear bridging layer 608, and a convolutional layer 610. The initial linear layer 602 and the second linear bridging layer 608 receive the damaged input feature 600a. The initial linear layer 602, the first linear bridging layer 606, and the second linear bridging layer 608 perform at least one of the following on the information from the input network layer: linear transformation, mapping, etc. The repaired feature 600b is the repair network's prediction of the damaged input feature 600a.
[0106] The number of transformation network blocks 604 in the repair network can be one or more. The transformation network block 604 includes: feedforward module 6041, multi-attention head module 6042, convolution module 6043, feedforward module 6044, and layer standard 6045.
[0107] In some examples, the repair network in this disclosure has a causal structure. For instance, during the training phase, the queries and keys of the multi-attention module 6042 are masked to remove information related to later frames, retaining only the current frame being processed and all its preceding historical frames. This ensures that the training process of the repair network is similar to the real-time audio information reception process during the use of the repair network, where information about future frames cannot be obtained, thereby improving the audio repair capability in application scenarios where real-time audio information is received.
[0108] As described above, the repair network includes an attention layer; the attention layer is used to mask the data corresponding to the current frame in subsequent frames; in some examples, the audio decoder in this disclosure includes an attention layer; the attention layer is used to mask the data corresponding to the current frame in subsequent frames.
[0109] The current frame is the time frame corresponding to the data currently being computed by the repair network and audio decoder; the data corresponding to the subsequent frames is a subset of at least one of the damage features and enhancement features.
[0110] In some examples, the queries and keys of the multi-attention module 6042 have a fixed-length sliding window during the training phase, retaining a fixed number of historical frames.
[0111] In summary, the method presented in this example achieves the restoration of the feature representation of audio information through a restoration network. Training the restoration network based on the difference between the restored and enhanced features improves its audio restoration capabilities. Alternating training of the restoration generator and the restoration discriminator in the restoration network alternately enhances the audio restoration capabilities of the restoration generator and the source discrimination ability of the restoration discriminator, ensuring the quality of the feature information predicted by the restoration generator. The audio decoder converts the feature representation of audio information into audio information, and training the audio decoder improves the accuracy of the decoded audio information. Training the restoration network and the audio decoder separately ensures that their audio restoration capabilities are independent. Training the restoration network and the audio decoder using sample audio pairs from the same data source ensures that the processed audio information capabilities of the trained restoration network and the trained audio decoder are matched, guaranteeing the accuracy of the audio decoder in decoding the input audio features and thus ensuring the quality of the predicted audio information.
[0112] Next, the training process of the repair network will be described as follows. In an optional design, the training process of the repair network described in step 530 of Figure 3 can be implemented as the following two sub-steps:
[0113] • Based on the differences between repair features and enhancement features, the repair network is trained to obtain the teacher repair network.
[0114] In some examples, the process of training the teacher-repaired network involves learning a noise reduction function D. φ The loss function for training the teacher-repaired network is:
[0115] in, This represents the i-th enhancement feature (e.g., the enhancement feature is a Mel-spectral feature). This represents the speech representation after adding noise at a noise level of t, and cond represents the impairment feature of the input.
[0116] The loss weight corresponding to noise level t. σ data p data The standard deviation of (x).
[0117] Regarding the noise reduction function D φ A similar noise reduction function estimation algorithm is used, such as: D φ (x t ,t,cond)=c skip (t)x t +cout (t)F φ (c in (t)x t ,t,c noise (t)) (Formula 5)
[0118] Where, c noise (t) maps the noise level t to F φ A conditional input; c in (t) represents x t The scale weight, c out (t) represents F φ The size weight, c skip (t) represents adjusting the weight of the skip connection, as shown below:
[0119] • Distillation training is performed on the teacher's repair network to obtain the student's repair network, and the student's repair network is identified as the trained repair network.
[0120] Figure 6 shows a schematic diagram of an example repair network of this disclosure.
[0121] The damage feature 652 is input into the initial teacher network 662 to predict the first restoration feature 662a. Based on the first difference between the enhancement feature 322 and the first restoration feature 662a, the initial teacher network 662 is trained to obtain the teacher restoration network 664. In some examples, the initial teacher network 662 is trained using the first difference between the enhancement feature 322 and the first restoration feature 662a, adjusting the network parameters of the initial teacher network 662 and improving its audio restoration capability. The network obtained after adjusting the network parameters of the initial teacher network 662 is called the teacher restoration network 664.
[0122] Damage feature 652 is input into teacher restoration network 664 to predict second restoration feature 664a; and damage feature 652 is input into initial student network 666 to predict third restoration feature 666a. Based on the second difference between second restoration feature 664a and third restoration feature 666a, initial student network 666 is trained to obtain student restoration network 668. Specifically, by adjusting the network parameters of initial student network 666, its restoration ability is made to approach that of teacher restoration network 664, achieving the audio restoration ability of teacher restoration network 664 with a smaller network size in student restoration network 668.
[0123] As described above, a trained teacher repair network is obtained through backpropagation training; a student model is obtained through consistency distillation, so that the student repair network can generate features through a single step or a few steps of sampling. The network size of the student repair network is smaller than that of the teacher repair network.
[0124] In some examples, noisy speech representations are obtained by randomly sampling n from a uniform distribution U(1,N-1). Then, using the pre-trained D φ To estimate based on The next step is to characterize the noise reduction process. As shown below:
[0125] Among them, the teacher model D is used. φ Initialize the student model D with parameter φ θ The model parameters θ and their exponential moving average (EMA) parameters. As shown below:
[0126] In one instance, the momentum coefficient μ = 0.95.
[0127] In some examples, consistency distillation training (also known as model distillation or parameter distillation) is performed on the teacher repair network to obtain the loss function (Consistency Loss) of the student repair network, as shown below:
[0128] Among them, || || 2 It is the L2 norm calculation, also known as Ridge Regression or weight decay.
[0129] In summary, the method presented in this example achieves the restoration of the feature representation of audio information through a restoration network. Training the restoration network based on the difference between the restored and enhanced features improves its audio restoration capabilities. Distillation training of the teacher restoration network yields a student restoration network, achieving the same audio restoration capabilities as the teacher network with a smaller network size. The audio decoder converts the feature representation of audio information into audio information, and training it improves the accuracy of the decoded audio information. Training the restoration network and audio decoder separately ensures that their audio restoration and decoding capabilities are independent. Training both the restoration network and audio decoder using sample audio pairs from the same data source guarantees a match between their audio processing capabilities, ensuring the accuracy of the audio decoder's decoding of the input audio features and thus guaranteeing the quality of the predicted audio information.
[0130] Figure 7 shows a flowchart of an example of an audio information processing method according to this disclosure. This method can be executed by a computer device. Specifically, in the example shown in Figure 3, step 540 can be implemented as steps 542, 544, 546, and 548:
[0131] Step 542: Perform audio decoding on the damage features and enhancement features respectively using an audio decoder to obtain the damage-decoded audio and the enhancement-decoded audio; and obtain the data source labels for the damage-decoded audio and the enhancement-decoded audio.
[0132] In this example, the audio decoder includes a decoder generator and a decoder discriminator. The decoder generator has the ability to restore the audio information from the feature information extracted from the audio information; this enables audio decoding of the input audio features and outputting the corresponding audio information. The decoder discriminator is used to determine the source of the audio, such as whether the audio input to the decoder was predicted by the restoration generator or directly acquired.
[0133] In this step, the data source label indicates that the audio was predicted by the audio decoder. In an optional implementation, the data source label for the original audio corresponding to the audio with sound quality degradation in the sample information pair is different from the data source labels for the damaged decoded audio and the enhanced decoded audio. The data source label for the original audio indicates that the audio features were acquired.
[0134] For example, the decoder generator includes deep convolutional layers and pointwise convolutional layers; correspondingly, this step can be implemented as follows:
[0135] • By using the deep convolutional layer and the pointwise convolutional layer in the audio decoder, inverse Fourier transforms are performed on the damage features and enhancement features respectively, resulting in damage-decoded audio and enhancement-decoded audio based on the data distribution of Fourier time-frequency features.
[0136] In some examples, the inverse Fourier transforms performed on the damage and enhancement features are independent. The decoder generator includes deep convolutional layers (DW) and point-wise convolutional layers (PW). The deep convolutional layers have a larger receptive field than the point-wise convolutional layers; the larger receptive field of the deep convolutional layers allows them to capture historical temporal information from the input features through dilated convolution. The point-wise convolutional layers further regularize the short-term output features of the deep convolutional layer portion, improving the convergence speed and stability of the decoder generator.
[0137] For example, the audio decoder is implemented as a VOCOS vocoder. In this example, the audio decoder is trained based on a Generative Adversarial Network (GAN), and is also called a VOCOS-GAN vocoder. The audio decoder transforms the predicted short-time Fourier coefficients from the frequency domain to the time domain through a short-time inverse Fourier transform, obtaining audio information based on the data distribution of Fourier time-frequency features. In other words, it uses the data distribution based on short-time Fourier time-frequency features as the decoding generator.
[0138] In some examples, the audio decoder in this example does not use transposed convolution. Instead, it upsamples the frequency domain information using the Inverse Short-Time Fourier Transform (ISTFT). Because the audio decoder does not use transposed convolution, aliasing artifacts are avoided. Furthermore, separable convolution further reduces the computational complexity and number of parameters in the convolutional layers, effectively improving the overall computational efficiency of the model while ensuring good speech generation results.
[0139] In some examples, the audio decoder includes one or more building blocks (ConvNeXt Blocks), each containing a deep convolutional layer and a pointwise convolutional layer.
[0140] Step 544: Perform prediction on the damaged decoded audio and the enhanced decoded audio respectively through the decoding discriminator to obtain the first predicted source label of the damaged decoded audio and the second predicted source label of the enhanced decoded audio.
[0141] As described above, the decoder discriminator is used to determine the source of audio information; for example, it can determine whether the audio information input to the decoder discriminator is predicted by the decoder generator or directly acquired.
[0142] In some examples, the decoder discriminator is also used to perform predictions on directly acquired audio information to obtain corresponding predicted source labels. The decoder discriminator is trained by comparing the predicted source labels with the directly acquired audio information to indicate the difference between the audio features and the acquired true source labels.
[0143] In some examples, the decoder discriminator consists of a multi-period discriminator and a multi-scale frequency domain sub-band discriminator. The multi-period discriminator is composed of several sub-discriminators for signals with different periods. For multi-period signals, the output of the decoder generator or the clean speech waveform signal is sampled at different intervals to transform the one-dimensional waveform data into a two-dimensional signal. Then, it is convolved to obtain the discrimination result for each period and the intermediate layer feature output. The multi-scale frequency domain sub-band discriminator convolves the amplitude spectra of multiple sub-bands obtained after multi-scale short-time Fourier transform to obtain the discrimination result for each Fourier transform scale and the intermediate layer feature output.
[0144] Step 546: Based on the differences between the first predicted source label, the second predicted source label and the data source label, the differences between the damaged decoded audio and the audio with degraded quality, and the differences between the enhanced decoded audio and the audio with enhanced quality, train the decoder generator to obtain the trained decoder generator.
[0145] In some examples, the difference between the first predicted source label, the second predicted source label, and the data source label is used to indicate the accuracy with which the decoder discriminator distinguishes the predicted source labels. As this difference increases, it indicates a higher ability of the decoder generator to reconstruct audio information from feature information. A decoder generator with a higher ability to reconstruct audio information from feature information is more likely to output audio information that causes the decoder discriminator to incorrectly determine the data source.
[0146] The difference between damaged decoded audio and audio quality-damaged audio, and the difference between enhanced decoded audio and audio quality-enhanced audio, are used to indicate the difference between the audio information output by the decoder generator and the audio information expected to be obtained. As the difference increases, it indicates that the decoder generator has a lower ability to reconstruct audio information from feature information.
[0147] In some examples, the network parameters of the audio decoder are adjusted to reduce the difference between damaged decoded audio and audio quality-damaged audio, and the difference between enhanced decoded audio and audio quality-enhanced audio, and to increase the difference between the first predicted source label, the second predicted source label and the data source label, so as to improve the audio decoder's ability to restore feature information to audio information.
[0148] For example, this step can be implemented as follows:
[0149] • The decoding discriminator is trained based on the differences between the first predicted source label, the second predicted source label, and the data source label by fixing the network parameters of the decoding generator.
[0150] In some examples, the network parameters of the decoder generator are fixed while training the decoder discriminator. This allows for separate training of the decoder discriminator, improving its ability to distinguish the source of audio information.
[0151] • By fixing the network parameters of the decoder discriminator, the decoder generator is trained based on the differences between damaged decoded audio and audio with reduced sound quality, and the differences between enhanced decoded audio and audio with enhanced sound quality, to obtain the trained decoder generator.
[0152] In some examples, the network parameters of the decoder discriminator are fixed while training the decoder generator. This allows for separate training of the decoder generator, improving its ability to reconstruct audio information from features extracted from audio information.
[0153] It should be noted that there are no restrictions on the timing of training the decoder generator and decoder discriminator in this example. The decoder generator can be trained first, or the decoder discriminator can be trained first; in some instances, the training of the decoder generator and the decoder discriminator can be performed alternately multiple times to achieve alternating improvement of the capabilities of the decoder generator and decoder discriminator during the training process.
[0154] In some examples, the differences between the predicted source label and the data source label, and the differences between the decoded features and the augmented features are as follows:
[0155] During the training phases of the decoder generator and decoder discriminator, the feature information output by the repair network is used as the input features of the decoder generator. Taking the decoding process of damaged audio as an example, the damage features are input into the decoder generator to obtain the damaged audio.
[0156] The decoded audio output by the calculation decoder generator. Mel's score Mel spectrum of audio x with sound quality impairment in sample audio pairs L1 distance between
[0157] Using the audio from the sample audio pair and the audio signal output from the decoder generator as input samples for the decoder discriminator, the audio output of the decoder discriminator for the sample audio pair is obtained. and the output audio signal for the decoder generator. Then, the final output of the subband spectrum of the decoder at each period or short-time Fourier scale and the corresponding intermediate layer output are calculated respectively. and L1 error; to minimize generation loss The training criterion for the decoder generator, also known as the loss function of the decoder generator, is to minimize the adversarial loss. The training criterion for the decoder discriminator, also known as the loss function of the decoder discriminator, is as follows:
[0158] in, Let L1 distance be the distance between the discriminator and the intermediate layer features of the audio in the sample audio pair and the audio output by the decoder generator, where K and L are the number of sub-discriminators and intermediate layers included in the decoder discriminator, respectively. α and β are the decoding errors of the audio. Matching error with audio source The weighting coefficients.
[0159] In some examples, a batch of data includes multiple data points. During iterative training of each batch of data, the decoder generator and decoder discriminator alternately update their parameters. When updating the decoder generator, the parameter updates of the decoder discriminator are frozen. When updating the parameters of the decoder discriminator, the gradient backpropagation of the decoder generator is blocked.
[0160] Step 548: Determine the trained decoder generator as the trained audio decoder.
[0161] In this example, the trained audio decoder includes a trained decoder generator. In some examples, after training the audio decoder, only the trained decoder generator is used to decode the audio information, independent of the decoder discriminator.
[0162] It should be noted that this example can be combined with Figure 4 to form a new example. For example, step 530 in this example can be implemented as steps 532 to 538 shown in Figure 4. This disclosure does not impose any limiting provisions in this regard.
[0163] In summary, the method presented in this example achieves the restoration of the feature representation of audio information through a restoration network. Training the restoration network based on the difference between the restored and enhanced features improves its audio restoration capabilities. The audio decoder converts the feature representation of audio information into audio information, and training the decoder improves the accuracy of the decoded audio information. Alternating training of the decoder generator and decoder discriminator in the audio decoder alternately enhances the generator's ability to convert feature information into audio information and the discriminator's ability to distinguish the source of audio information, ensuring the quality of the audio information predicted by the decoder. Training the restoration network and the audio decoder separately ensures that their audio restoration capabilities are independent. Training both the restoration network and the audio decoder using sample audio pairs from the same data source ensures that their audio processing capabilities are matched, guaranteeing the accuracy of the audio decoder's audio decoding of the input audio features and thus ensuring the quality of the predicted audio information.
[0164] Figure 8 shows a flowchart of an example method for processing audio information according to this disclosure. This method can be executed by a computer device. Specifically, in the example shown in Figure 3, step 510 can be implemented as steps 512 and 514; and also includes steps 562, 564, and 566.
[0165] Step 512: Obtain the initial sample audio.
[0166] In some examples, the initial sample audio is directly acquired audio information.
[0167] The information in the sample information pair, the initial sample audio is usually audio information with human voice content; also called initial sample speech; correspondingly, the audio with enhanced sound quality obtained based on the initial sample speech is also called sound effect enhanced speech, and the audio with damaged sound quality is also called noise damaged speech.
[0168] However, it cannot be ruled out that there may be other sounds such as musical instruments, background music, or ambient sounds, or that there may be no human voice content.
[0169] Step 514: Perform sound effect enhancement on the initial sample audio to obtain sound quality enhanced audio; and perform damage simulation on the initial sample audio to obtain sound quality damaged audio.
[0170] In some examples, the methods for enhancing the initial sample audio include, but are not limited to, performing noise reduction, reverberation removal, and harmonic enhancement, to reduce the background noise and inter-harmonic noise of the audio information acquisition while ensuring that the audio of the initial sample audio is not distorted.
[0171] For example, performing damage simulation can be achieved as follows:
[0172] • Perform audio impairment on the initial sample audio to obtain initial noisy audio;
[0173] In some examples, performing audio impairments includes, but is not limited to, adding noise, adding reverberation, nonlinear distortion, clipping, and high-pass filtering to the initial sample audio.
[0174] • Perform audio denoising on the initial noisy audio to obtain audio with degraded sound quality.
[0175] In some examples, audio denoising is performed on the initial noisy audio, such as performing at least one of Acoustic Echo Cancellation (AEC), Automatic Noise Suppression (ANS), and Automatic Gain Control (AGC). The audio with degraded sound quality obtained by performing audio denoising carries at least one of the following: spectral harmonic loss or residual noise resulting from the audio denoising process. In some instances, as described above, the various examples in this disclosure are not limited to audio with degraded sound quality directly added; the audio with degraded sound quality is only used to specify that the sound quality is reduced due to factors such as spectral harmonic loss or residual noise.
[0176] In some examples, audio enhancement and audio degradation are two audio files with different audio qualities obtained by performing sound effect editing on the initial sample audio. Taking a speech information with human voice content as an example, audio enhancement and audio degradation have the same audio content and carry the same semantics; however, the audio quality of the two audio files is different, for example, there are differences in dimensions such as noise, background noise, and spectral impairment.
[0177] Step 562: Obtain the input audio and perform feature extraction on the input audio to obtain the input audio features.
[0178] In some examples, the input audio is the audio information to be repaired, and the input audio is usually captured in real time. For example, it can be captured through an instant messaging (IM) application, or through the audio conversation function provided by a game application.
[0179] In some examples, feature extraction is performed on the input audio, and the resulting input audio features are deep features extracted from the input audio, such as feature information in the hidden space, or at least one of the extracted audio features in the time domain, frequency domain, or other spectral domains.
[0180] Step 564: Perform audio inpainting on the input audio features using the trained inpainting network to obtain the predicted inpainted features.
[0181] The trained restoration network possesses audio restoration capabilities. As described above, during the training process of the restoration network, the audio with degraded sound quality carries at least one of the following: spectral harmonic loss or residual noise resulting from audio denoising. Accordingly, the trained restoration network, obtained through backpropagation training, has the ability to reduce spectral harmonic loss and decrease residual noise.
[0182] Figure 9 illustrates a schematic diagram of audio features provided in an exemplary example of this disclosure. In subgraph (a), the spectral information of the input audio has more missing portions, displayed as dark areas, compared to the spectral information of the predicted audio in subgraph (b). The predicted audio in subgraph (b) undergoes audio restoration via a restoration network to fill in the missing spectral information, thus improving audio quality.
[0183] Step 566: Perform audio decoding on the predicted repair features using the trained audio decoder to obtain the predicted audio.
[0184] The trained audio decoder has the ability to restore audio information from the feature information extracted from the audio information. The predicted inpainting features input to the trained audio decoder are predicted by the trained inpainting network, realizing the decoding of the audio features output by the trained inpainting network, and obtaining predicted audio with reduced spectral harmonic loss and reduced noise residue.
[0185] In summary, the method presented in this example achieves the restoration of the feature representation of audio information through a restoration network. Training the restoration network based on the difference between the restored and enhanced features improves its audio restoration capabilities. The method also converts the feature representation of audio information into audio information through an audio decoder, and training the decoder improves the accuracy of the decoded audio information. Furthermore, training the restoration network and the audio decoder separately ensures that their audio restoration and decoding capabilities are independent. Training both the restoration network and the audio decoder using sample audio pairs from the same data source guarantees that their audio processing capabilities are matched, ensuring the accuracy of the audio decoder in decoding the input audio features and thus guaranteeing the quality of the predicted audio information.
[0186] In one application scenario, an instant messaging (IM) application provides audio information processing functionality; for example, the instant messaging application is at least one of a chat application, a dating application, or an online conferencing application. Accordingly, the audio information processing method provided in this disclosure is implemented as follows:
[0187] • Obtain sample chat audio pairs, which include audio with degraded sound quality and audio with enhanced sound quality. Audio with degraded sound quality is audio information that has been degraded in sound quality, while audio with enhanced sound quality has higher audio quality than audio with degraded sound quality.
[0188] • Perform feature extraction on both audio with and without sound quality enhancement to obtain damage features and enhancement features;
[0189] • Audio restoration is performed on the damage features using a restoration network to obtain restored features; based on the difference between the restored features and the enhanced features, the restoration network is trained to obtain the trained restoration network;
[0190] • Perform audio decoding on the damage features and enhancement features respectively using an audio decoder to obtain damaged decoded audio and enhanced decoded audio; train the audio decoder based on the differences between damaged decoded audio and audio quality-damaged audio, and the differences between enhanced decoded audio and audio quality-enhanced audio; obtain the trained audio decoder.
[0191] In some examples, the chat input audio is obtained, and feature extraction is performed on the chat input audio to obtain chat audio features;
[0192] The trained inpainting network performs audio inpainting on the input audio features to obtain the predicted inpainted features;
[0193] The chat repair audio is obtained by performing audio decoding on the predicted repair features using the trained audio decoder.
[0194] Specifically, by calling the trained repair network and the trained audio decoder to perform prediction on chat audio features, chat repair audio is obtained, which can improve the audio quality of the audio conversation function provided by instant messaging applications. This avoids the problem that the audio information quality is reduced due to noisy ambient sound, which makes it impossible for the recipient to distinguish the audio content, thus improving the user experience.
[0195] In some examples, the repair network includes a repair generator and a repair discriminator; accordingly: audio repair is performed on the damage features by the repair generator to obtain repair features, and the data source labels of the repair features are obtained, which are used to indicate that the repair features were predicted by the repair generator;
[0196] The predicted source label of the repaired feature is obtained by performing prediction on the repaired feature through the repair discriminator;
[0197] Based on the differences between predicted source labels and data source labels, and the differences between repair features and enhancement features, the repair generator is trained to obtain the trained repair generator.
[0198] The trained repair generator is identified as the trained repair network.
[0199] In some examples, the first learning rate used to train the repair discriminator is greater than the second learning rate used to train the repair generator.
[0200] In some examples, the repair network is trained based on the differences between the repair features and the enhancement features to obtain the teacher repair network;
[0201] Distillation training is performed on the teacher repair network to obtain the student repair network, and the student repair network is identified as the trained repair network.
[0202] In some examples, the audio decoder includes a decoding generator and a decoding discriminator; accordingly: audio decoding is performed on the impairment features and enhancement features by the audio decoder to obtain impairment-decoded audio and enhancement-decoded audio; and the data source labels of the impairment-decoded audio and enhancement-decoded audio are obtained, the data source labels being used to indicate that the audio was predicted by the audio decoder;
[0203] The decoding discriminator performs predictions on the damaged decoded audio and the enhanced decoded audio respectively, and obtains the first predicted source label of the damaged decoded audio and the second predicted source label of the enhanced decoded audio.
[0204] Based on the differences between the first predicted source label, the second predicted source label and the data source label, the differences between impaired decoded audio and audio quality impaired audio, and the differences between enhanced decoded audio and audio quality enhanced audio, the decoding generator is trained to obtain the trained decoding generator.
[0205] The trained decoder generator is used as the trained audio decoder.
[0206] In some examples, the decoder generator includes deep convolutional layers and pointwise convolutional layers; through the deep convolutional layers and pointwise convolutional layers in the audio decoder, inverse Fourier transforms are performed on the damage features and enhancement features respectively to obtain damaged decoded audio and enhanced decoded audio based on the data distribution of Fourier time-frequency features.
[0207] In some examples, at least one of the repair network and the audio decoder includes an attention layer; the attention layer is used to mask the data corresponding to the current frame in the next frame.
[0208] The current frame is the time frame corresponding to the data currently being computed by the repair network and audio decoder; the data corresponding to the subsequent frames is a subset of at least one of the damage features and enhancement features.
[0209] In another application scenario, a game application that supports audio conversations provides audio information processing capabilities; the game application supports audio conversation functionality to enable real-time audio conversations between the controllers of virtual objects in a virtual game. Accordingly, the audio information processing method provided in this disclosure is implemented as follows:
[0210] • Obtain sample game audio pairs, which include audio with degraded sound quality and audio with enhanced sound quality. Audio with degraded sound quality is audio information that has been degraded in sound quality, while audio with enhanced sound quality has higher audio quality than audio with degraded sound quality.
[0211] • Perform feature extraction on both audio with and without sound quality enhancement to obtain damage features and enhancement features;
[0212] • Audio restoration is performed on the damage features using a restoration network to obtain restored features; based on the difference between the restored features and the enhanced features, the restoration network is trained to obtain the trained restoration network;
[0213] • Perform audio decoding on the damage features and enhancement features respectively using an audio decoder to obtain damaged decoded audio and enhanced decoded audio; train the audio decoder based on the differences between damaged decoded audio and audio quality-damaged audio, and the differences between enhanced decoded audio and audio quality-enhanced audio; obtain the trained audio decoder.
[0214] In some examples, game input audio is acquired, and feature extraction is performed on the game input audio to obtain game audio features;
[0215] The trained inpainting network performs audio inpainting on the input audio features to obtain the predicted inpainted features;
[0216] The game repair audio is obtained by performing audio decoding on the predicted repair features using the trained audio decoder.
[0217] Specifically, by calling the trained repair network and the trained audio decoder to perform prediction on the game audio features, the game repair audio is obtained. This can improve the audio quality of the audio conversation function supported by the game application, and avoid the problem that the audio information quality is reduced due to noisy ambient sounds or excessive operation sounds of human-computer interaction devices when controlling virtual objects, which makes it impossible for the receiver to distinguish the audio content, thus improving the user experience.
[0218] In some examples, the repair network includes a repair generator and a repair discriminator; accordingly: audio repair is performed on the damage features by the repair generator to obtain repair features, and the data source labels of the repair features are obtained, which are used to indicate that the repair features were predicted by the repair generator;
[0219] The predicted source label of the repaired feature is obtained by performing prediction on the repaired feature through the repair discriminator;
[0220] Based on the differences between predicted source labels and data source labels, and the differences between repair features and enhancement features, the repair generator is trained to obtain the trained repair generator.
[0221] The trained repair generator is identified as the trained repair network.
[0222] For example, the first learning rate used to train the repair discriminator is greater than the second learning rate used to train the repair generator.
[0223] In some examples, the repair network is trained based on the differences between the repair features and the enhancement features to obtain the teacher repair network;
[0224] Distillation training is performed on the teacher repair network to obtain the student repair network, and the student repair network is identified as the trained repair network.
[0225] In some examples, the audio decoder includes a decoding generator and a decoding discriminator; accordingly: audio decoding is performed on the impairment features and enhancement features by the audio decoder to obtain impairment-decoded audio and enhancement-decoded audio; and the data source labels of the impairment-decoded audio and enhancement-decoded audio are obtained, the data source labels being used to indicate that the audio was predicted by the audio decoder;
[0226] The decoding discriminator performs predictions on the damaged decoded audio and the enhanced decoded audio respectively, and obtains the first predicted source label of the damaged decoded audio and the second predicted source label of the enhanced decoded audio.
[0227] Based on the differences between the first predicted source label, the second predicted source label and the data source label, the differences between impaired decoded audio and audio quality impaired audio, and the differences between enhanced decoded audio and audio quality enhanced audio, the decoding generator is trained to obtain the trained decoding generator.
[0228] The trained decoder generator is used as the trained audio decoder.
[0229] In some examples, the decoder generator includes deep convolutional layers and pointwise convolutional layers; through the deep convolutional layers and pointwise convolutional layers in the audio decoder, inverse Fourier transforms are performed on the damage features and enhancement features respectively to obtain damaged decoded audio and enhanced decoded audio based on the data distribution of Fourier time-frequency features.
[0230] In some examples, at least one of the repair network and the audio decoder includes an attention layer; the attention layer is used to mask the data corresponding to the current frame in the next frame.
[0231] The current frame is the time frame corresponding to the data currently being computed by the repair network and audio decoder; the data corresponding to the subsequent frames is a subset of at least one of the damage features and enhancement features.
[0232] Those skilled in the art will understand that the above examples can be implemented independently, or the above examples can be freely combined to create new examples that implement the audio information processing method of this disclosure.
[0233] Figure 10 shows a structural block diagram of an audio information processing apparatus according to an example of this disclosure. The apparatus includes:
[0234] The acquisition module 810 is used to acquire sample audio pairs, which include audio with the same audio content but different audio quality, and audio with enhanced audio quality.
[0235] The processing module 820 is used to perform feature extraction on the audio with damaged sound quality and the audio with enhanced sound quality respectively to obtain damaged features and enhanced features; the processing module 820 is also used to perform audio restoration on the damaged features through a restoration network to obtain restored features;
[0236] Training module 830 is used to train the repair network based on the difference between the repair features and the enhancement features to obtain the trained repair network;
[0237] The processing module 820 is further configured to perform audio decoding on the damage feature and the enhancement feature respectively through an audio decoder to obtain damage-decoded audio and enhancement-decoded audio;
[0238] The training module 830 is further configured to train the audio decoder based on the difference between the damaged decoded audio and the audio quality damaged audio, and the difference between the enhanced decoded audio and the audio quality enhanced audio, to obtain a trained audio decoder.
[0239] The device may also include a providing module for providing the repair network and the audio decoder for audio processing, wherein the repair network is used to repair the input audio, and the audio decoder decodes the audio output by the repair network to obtain the output audio.
[0240] In some examples, the repair network includes a repair generator and a repair discriminator; the processing module 820 is also used for:
[0241] The repair generator performs audio repair on the damage features to obtain repaired features and obtains the data source label of the repaired features, the data source label being used to indicate that the repaired features were predicted by the repair generator;
[0242] The repair discriminator performs prediction on the repair features to obtain the predicted source label of the repair features;
[0243] The training module 830 is also used for:
[0244] Based on the differences between the predicted source label and the data source label, and the differences between the repair feature and the enhancement feature, the repair generator is trained to obtain the trained repair generator;
[0245] The trained repair generator is identified as the trained repair network.
[0246] In some examples, the first learning rate used to train the repair discriminator is greater than the second learning rate used to train the repair generator.
[0247] In some examples, the training module 830 is also used for:
[0248] Based on the differences between the repair features and the enhancement features, the repair network is trained to obtain the teacher repair network;
[0249] Distillation training is performed on the teacher repair network to obtain the student repair network, and the student repair network is identified as the trained repair network.
[0250] In some examples, the audio decoder includes a decoder generator and a decoder discriminator; the processing module 820 is also used for:
[0251] The audio decoder performs audio decoding on the impairment feature and the enhancement feature respectively to obtain impairment-decoded audio and enhancement-decoded audio; and obtains the data source tags of the impairment-decoded audio and the enhancement-decoded audio, the data source tags being used to indicate that the audio was predicted by the audio decoder;
[0252] The decoding discriminator performs predictions on the damaged decoded audio and the enhanced decoded audio respectively to obtain a first predicted source label for the damaged decoded audio and a second predicted source label for the enhanced decoded audio.
[0253] The training module 830 is also used for:
[0254] Based on the differences between the first predicted source label, the second predicted source label and the data source label, the differences between the damaged decoded audio and the audio quality damaged audio, and the differences between the enhanced decoded audio and the audio quality enhanced audio, the decoder generator is trained to obtain the trained decoder generator;
[0255] The trained decoder generator is determined as the trained audio decoder.
[0256] In some examples, the decoder generator includes deep convolutional layers and pointwise convolutional layers; the processing module 820 is also used for:
[0257] The inverse Fourier transform is performed on the damage features and the enhancement features respectively through the deep convolutional layer and the pointwise convolutional layer in the audio decoder to obtain the damage-decoded audio and the enhancement-decoded audio based on the data distribution of Fourier time-frequency features.
[0258] In some examples, at least one of the repair network and the audio decoder includes an attention layer; the attention layer is used to mask the data corresponding to the current frame in the subsequent frame;
[0259] Wherein, the current frame is the time frame corresponding to the data currently being calculated by the repair network and the audio decoder; the data corresponding to the subsequent frame is a subset of the sub-features included in at least one of the damage features and the enhancement features.
[0260] In some examples, the acquisition module 810 is also used for:
[0261] Obtain initial sample audio;
[0262] The initial sample audio is subjected to sound effect enhancement to obtain the sound quality enhanced audio; and the initial sample audio is subjected to damage simulation to obtain the sound quality damaged audio.
[0263] In some examples, the acquisition module 810 is also used for:
[0264] Perform audio impairment on the initial sample audio to obtain initial noisy audio;
[0265] Audio denoising is performed on the initial noisy audio to obtain the audio with impaired sound quality, the audio with impaired sound quality carrying at least one of the following: loss of spectral harmonics and residual noise caused by the audio denoising.
[0266] In some examples, the acquisition module 810 is also used for:
[0267] Get the input audio;
[0268] The processing module 820 is further configured to:
[0269] Feature extraction is performed on the input audio to obtain the input audio features;
[0270] The trained repair network is used to perform audio repair on the input audio features to obtain predicted repair features;
[0271] The predicted audio is obtained by performing audio decoding on the predicted repair features using the trained audio decoder.
[0272] It should be noted that the device provided in the above example is only used as an example to illustrate the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules according to actual needs, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0273] Regarding the apparatus in the above examples, the specific ways in which each module performs its operations have been described in detail in the examples related to the method; the technical effects achieved by each module performing its operations are the same as those in the examples related to the method, and will not be elaborated here.
[0274] This disclosure also provides a computer device comprising: a processor and a memory, wherein a computer program is stored in the memory; the processor is configured to execute the computer program in the memory to implement the audio information processing method provided in the above method examples.
[0275] For example, the computer device is a server. Figure 11 is a structural block diagram of a server provided in an exemplary example of this disclosure.
[0276] Typically, server 2300 includes a processor 2301 and memory 2302.
[0277] Processor 2301 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 2301 may be implemented using at least one of the following hardware forms: Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). Processor 2301 may also include a main processor and a coprocessor. The main processor, also known as the Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some examples, processor 2301 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content required to be displayed on the screen. In some examples, processor 2301 may also include an Artificial Intelligence (AI) processor, which handles computational operations related to machine learning.
[0278] Memory 2302 may include one or more computer-readable storage media, which may be non-transitory. Memory 2302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some examples, the non-transitory computer-readable storage media in memory 2302 is used to store at least one instruction for execution by processor 2301 to implement the audio information processing method provided in the method examples of this disclosure.
[0279] In some examples, server 2300 may optionally include an input interface 2303 and an output interface 2304. Processor 2301, memory 2302, and input interfaces 2303 and 2304 can be connected via a bus or signal lines. Various peripheral devices can be connected to input interfaces 2303 and 2304 via buses, signal lines, or circuit boards. Input interfaces 2303 and 2304 can be used to connect at least one input / output (I / O) related peripheral device to processor 2301 and memory 2302. In some examples, processor 2301, memory 2302, and input interfaces 2303 and 2304 are integrated on the same chip or circuit board; in other examples, any one or two of processor 2301, memory 2302, and input interfaces 2303 and 2304 can be implemented on separate chips or circuit boards, which is not limited in this disclosure.
[0280] Those skilled in the art will understand that the structure shown above does not constitute a limitation on server 2300, and may include more or fewer components than shown, or combine certain components, or employ different component arrangements.
[0281] In some examples, a chip is also provided, the chip including programmable logic circuitry and / or program instructions, which, when the chip is run on a computer device, are used to implement the audio information processing method described above.
[0282] In some examples, a computer program product is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions to implement the audio information processing methods provided in the above-described method examples.
[0283] In some examples, a computer-readable storage medium is also provided, in which a computer program is stored, which is loaded and executed by a processor to implement the audio information processing methods provided in the above method examples.
[0284] Those skilled in the art will understand that all or part of the steps of the above examples can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0285] Those skilled in the art will recognize that the functions described in the examples above can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium accessible to a general-purpose or special-purpose computer.
[0286] The technical features in the above examples can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above examples are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0287] The above description is merely an optional example of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for processing audio information, executed by a computer device, the method comprising: Obtain sample audio pairs, which include audio with the same audio content but different audio quality, and audio with sound quality enhancement; By performing feature extraction on the audio with damaged sound quality and the audio with enhanced sound quality respectively, damage features and enhancement features are obtained; Audio restoration is performed on the damage features by a restoration network to obtain restored features; the restoration network is trained based on the difference between the restored features and the enhanced features to obtain a trained restoration network. The audio decoder performs audio decoding on the damage feature and the enhancement feature respectively to obtain damage-decoded audio and enhancement-decoded audio. The audio decoder is trained based on the differences between the damaged decoded audio and the audio with degraded sound quality, as well as the differences between the enhanced decoded audio and the audio with enhanced sound quality, to obtain a trained audio decoder. The repair network and the audio decoder are provided for audio processing. The repair network is used to repair the input audio, and the audio decoder decodes the audio output by the repair network to obtain the output audio.
2. The method according to claim 1, wherein the repair network comprises a repair generator and a repair discriminator; and the audio repair is performed on the damage features through the repair network to obtain repair features; Based on the difference between the repair features and the enhancement features, the repair network is trained to obtain the trained repair network, including: The repair generator performs audio repair on the damage features to obtain repaired features and obtains the data source label of the repaired features, the data source label being used to indicate that the repaired features were predicted by the repair generator; The repair discriminator performs prediction on the repair features to obtain the predicted source label of the repair features; Based on the differences between the predicted source label and the data source label, and the differences between the repair feature and the enhancement feature, the repair generator is trained to obtain the trained repair generator; The trained repair generator is identified as the trained repair network.
3. The method according to claim 2, wherein the first learning rate for training the repair discriminator is greater than the second learning rate for training the repair generator.
4. The method according to claim 1, wherein training the repair network based on the difference between the repair features and the enhancement features to obtain the trained repair network comprises: Based on the differences between the repair features and the enhancement features, the repair network is trained to obtain the teacher repair network; Distillation training is performed on the teacher repair network to obtain the student repair network, and the student repair network is identified as the trained repair network.
5. The method according to any one of claims 1 to 4, wherein the audio decoder comprises a decoding generator and a decoding discriminator; wherein the audio decoder performs audio decoding on the impairment feature and the enhancement feature respectively to obtain impairment-decoded audio and enhancement-decoded audio; The audio decoder is trained based on the differences between the damaged decoded audio and the audio with degraded sound quality, and the differences between the enhanced decoded audio and the audio with enhanced sound quality, to obtain a trained audio decoder, including: The audio decoder performs audio decoding on the impairment feature and the enhancement feature respectively to obtain impairment-decoded audio and enhancement-decoded audio; and obtains the data source tags of the impairment-decoded audio and the enhancement-decoded audio, the data source tags being used to indicate that the audio was predicted by the audio decoder; The decoding discriminator performs predictions on the damaged decoded audio and the enhanced decoded audio respectively to obtain a first predicted source label for the damaged decoded audio and a second predicted source label for the enhanced decoded audio. Based on the differences between the first predicted source label, the second predicted source label and the data source label, the differences between the damaged decoded audio and the audio quality damaged audio, and the differences between the enhanced decoded audio and the audio quality enhanced audio, the decoder generator is trained to obtain the trained decoder generator; The trained decoder generator is determined as the trained audio decoder.
6. The method according to claim 5, wherein the decoding generator comprises a deep convolutional layer and a pointwise convolutional layer; the step of performing audio decoding on the impairment feature and the enhancement feature respectively through the audio decoder to obtain impairment-decoded audio and enhancement-decoded audio comprises: The inverse Fourier transform is performed on the damage features and the enhancement features respectively through the deep convolutional layer and the pointwise convolutional layer in the audio decoder to obtain the damage-decoded audio and the enhancement-decoded audio based on the data distribution of Fourier time-frequency features.
7. The method according to any one of claims 1 to 4, wherein at least one of the repair network and the audio decoder includes an attention layer; the attention layer is used to mask data corresponding to the current frame in a subsequent frame; in, The current frame is the time frame corresponding to the data currently being calculated by the repair network and the audio decoder; the data corresponding to the subsequent frame is a subset of at least one of the damage features and the enhancement features.
8. The method according to any one of claims 1 to 4, wherein acquiring the sample audio pair comprises: Obtain initial sample audio; The initial sample audio is enhanced with sound effects to obtain the enhanced audio. And perform a damage simulation on the initial sample audio to obtain the audio with damaged sound quality.
9. The method according to claim 8, wherein performing impairment simulation on the initial sample audio to obtain the audio with impaired sound quality includes: Perform audio impairment on the initial sample audio to obtain initial noisy audio; Audio denoising is performed on the initial noisy audio to obtain the audio with impaired sound quality, the audio with impaired sound quality carrying at least one of the following: loss of spectral harmonics and residual noise caused by the audio denoising.
10. The method according to any one of claims 1 to 4, further comprising: The input audio is acquired, and feature extraction is performed on the input audio to obtain the input audio features; The trained repair network is used to perform audio repair on the input audio features to obtain predicted repair features; The predicted audio is obtained by performing audio decoding on the predicted repair features using the trained audio decoder.
11. An audio information processing apparatus, comprising: The acquisition module is used to acquire sample audio pairs, which include audio with the same audio content but different sound quality, and audio with sound quality enhancement. The processing module is used to perform feature extraction on the audio with damaged sound quality and the audio with enhanced sound quality, respectively, to obtain damage features and enhancement features; The processing module is further configured to perform audio restoration on the damage features through a repair network to obtain repair features; The training module is used to train the repair network based on the difference between the repair features and the enhancement features to obtain the trained repair network; and A module is provided to provide the repair network and the audio decoder for audio processing. The repair network is used to repair the input audio, and the audio decoder decodes the audio output from the repair network to obtain the output audio. The processing module is further configured to perform audio decoding on the damage feature and the enhancement feature respectively through an audio decoder to obtain damage-decoded audio and enhancement-decoded audio; The training module is further configured to train the audio decoder based on the difference between the damaged decoded audio and the audio with reduced sound quality, and the difference between the enhanced decoded audio and the audio with enhanced sound quality, to obtain the trained audio decoder.
12. The apparatus according to claim 11, wherein the training module is configured to: The repair generator performs audio repair on the damage features to obtain repaired features and obtains the data source label of the repaired features, the data source label being used to indicate that the repaired features were predicted by the repair generator; The repair discriminator performs prediction on the repair features to obtain the predicted source label of the repair features; Based on the differences between the predicted source label and the data source label, and the differences between the repair feature and the enhancement feature, the repair generator is trained to obtain the trained repair generator; The trained repair generator is identified as the trained repair network.
13. The apparatus according to claim 11, wherein the training module is configured to: Based on the differences between the repair features and the enhancement features, the repair network is trained to obtain the teacher repair network; Distillation training is performed on the teacher repair network to obtain the student repair network, and the student repair network is identified as the trained repair network.
14. The apparatus according to any one of claims 11 to 13, wherein the training module is configured to: The audio decoder performs audio decoding on the impairment feature and the enhancement feature respectively to obtain impairment-decoded audio and enhancement-decoded audio; and obtains the data source tags of the impairment-decoded audio and the enhancement-decoded audio, the data source tags being used to indicate that the audio was predicted by the audio decoder; The decoding discriminator performs predictions on the damaged decoded audio and the enhanced decoded audio respectively to obtain a first predicted source label for the damaged decoded audio and a second predicted source label for the enhanced decoded audio. Based on the differences between the first predicted source label, the second predicted source label and the data source label, the differences between the damaged decoded audio and the audio quality damaged audio, and the differences between the enhanced decoded audio and the audio quality enhanced audio, the decoder generator is trained to obtain the trained decoder generator; The trained decoder generator is determined as the trained audio decoder.
15. The apparatus according to any one of claims 11 to 13, wherein at least one of the repair network and the audio decoder includes an attention layer; the attention layer is configured to mask data corresponding to the current frame in a subsequent frame; in, The current frame is the time frame corresponding to the data currently being calculated by the repair network and the audio decoder; the data corresponding to the subsequent frame is a subset of at least one of the damage features and the enhancement features.
16. The apparatus according to any one of claims 11 to 13, wherein, The acquisition module is further configured to acquire input audio, perform feature extraction on the input audio, and obtain input audio features; The processing module is further used for: The trained repair network is used to perform audio repair on the input audio features to obtain predicted repair features; The predicted audio is obtained by performing audio decoding on the predicted repair features using the trained audio decoder.
17. The apparatus according to any one of claims 11 to 13, wherein the acquisition module is configured to: Obtain initial sample audio; The initial sample audio is subjected to sound effect enhancement to obtain the sound quality enhanced audio; and the initial sample audio is subjected to damage simulation to obtain the sound quality damaged audio.
18. A computer device, comprising: A processor and a memory, wherein the memory stores at least one program; The processor is configured to execute the at least one program in the memory to implement the audio information processing method as described in any one of claims 1 to 10.
19. A computer-readable storage medium storing executable instructions, said executable instructions being loaded and executed by a processor to implement the audio information processing method as described in any one of claims 1 to 10.
20. A computer program product comprising computer instructions stored in a computer-readable storage medium, wherein a processor reads from the computer-readable storage medium and executes the computer instructions to implement the audio information processing method as described in any one of claims 1 to 10.