Audio restoration method and apparatus, and electronic device

By performing feature extraction and autoregression prediction on the audio clips and their contexts, the repaired audio clips are generated, and the problem of unnatural audio clips and contexts in the existing technology is solved, achieving better repair effects.

CN120452459APending Publication Date: 2025-08-08VIVO MOBILE COMM HANGZHOU CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510725431.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing audio repair method only processes the audio clips themselves, resulting in the repaired audio clips and context not being natural enough and the repair effect is not good.

Method used

By extracting the feature of the repaired audio clips, the above audio clips and the following audio clips, acoustic feature vectors are obtained, and these feature vectors are input to the audio model for autoregression prediction and decoding, and the repaired audio clips are generated.

Benefits of technology

The fixed audio clips are more natural to connect with the context, and the repair effect is better.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452459A_ABST
    Figure CN120452459A_ABST
Patent Text Reader

Abstract

The invention discloses an audio restoration method and device and electronic equipment, and belongs to the technical field of artificial intelligence. The method comprises the following steps: performing feature extraction on a to-be-repaired first audio clip, a previous-text audio clip and a next-text audio clip of the first audio clip to obtain a first acoustic feature vector of the first audio clip, a second acoustic feature vector of the previous-text audio clip and a third acoustic feature vector of the next-text audio clip; and inputting the first acoustic feature vector, the second acoustic feature vector and the third acoustic feature vector into an audio large model, performing autoregression prediction through the audio large model to obtain a predicted discrete coding vector, and decoding the predicted discrete coding vector to obtain a repaired second audio clip.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to an audio repair method, device and electronic device. Background Art

[0002] Audio or video recording is an important way for users to document their lives. During the recording process, users often encounter various noises, which can affect their enjoyment of the audio or video. In these cases, editing tools can be used to repair the first audio clip. However, current audio repair methods only process the first audio clip itself, resulting in the processed audio clip sounding less natural. Summary of the Invention

[0003] The purpose of the embodiments of the present application is to provide an audio repair method, device, and electronic device, which can make the repaired second audio segment more naturally connected with the context and achieve better repair effect.

[0004] In a first aspect, an embodiment of the present application provides an audio repair method, the method comprising:

[0005] Performing feature extraction on a first audio segment to be repaired, a preceding audio segment of the first audio segment, and a following audio segment of the first audio segment to obtain a first acoustic feature vector of the first audio segment, a second acoustic feature vector of the preceding audio segment, and a third acoustic feature vector of the following audio segment;

[0006] The first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector are input into the audio large model, autoregressive prediction is performed through the audio large model to obtain a predicted discrete coding vector, and the predicted discrete coding vector is decoded to obtain a repaired second audio segment.

[0007] In a second aspect, an embodiment of the present application provides an audio repair device, comprising:

[0008] a feature extraction module, configured to perform feature extraction on a first audio segment to be repaired, a preceding audio segment of the first audio segment, and a following audio segment of the first audio segment, to obtain a first acoustic feature vector of the first audio segment, a second acoustic feature vector of the preceding audio segment, and a third acoustic feature vector of the following audio segment;

[0009] The processing module is used to input the first acoustic eigenvector, the second acoustic eigenvector and the third acoustic eigenvector into the audio large model, perform autoregressive prediction through the audio large model to obtain a predicted discrete coding vector, and decode the predicted discrete coding vector to obtain a repaired second audio segment.

[0010] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.

[0011] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0012] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method described in the first aspect.

[0013] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the method described in the first aspect.

[0014] In an embodiment of the present application, feature extraction can be performed on the first audio segment to be repaired, the previous audio segment of the first audio segment, and the following audio segment to obtain a first acoustic feature vector of the first audio segment, a second acoustic feature vector of the previous audio segment, and a third acoustic feature vector of the following audio segment; the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector are input into the audio large model for autoregressive prediction to obtain a predicted discrete coding vector, and the predicted discrete coding vector is decoded to obtain the repaired second audio segment.

[0015] In this way, when repairing the first audio clip, the previous audio clip and the following audio clip of the first audio clip can also be considered to extract the acoustic feature vectors of the three, so that the acoustic feature vectors of the three are all used as inputs of the audio big model, so that when the audio big model performs autoregressive prediction, it can combine the context information, so that the repaired second audio clip obtained by the audio big model is more naturally connected with the context and the repair effect is better. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a flowchart of an audio repair method provided by some embodiments of the present application;

[0017] Figure 2 is a schematic diagram of data conversion in the audio repair method provided in some embodiments of the present application;

[0018] Figure 3a is a schematic diagram of a repair scenario in an audio repair method provided in some embodiments of the present application;

[0019] Figure 3b is a schematic diagram of a repair scenario in an audio repair method provided in some embodiments of the present application;

[0020] Figure 4 is a schematic structural diagram of an audio word segmenter in an audio repair method provided in some embodiments of the present application;

[0021] Figure 5 is a schematic structural diagram of an autoregressive module in an audio restoration method provided by some embodiments of the present application;

[0022] Figure 6 is a schematic diagram of data transformation in the training of an autoregressive module in the audio restoration method provided in some embodiments of the present application;

[0023] Figure 7 is a schematic diagram of the input and output of the initial model in the audio restoration method provided in some embodiments of the present application;

[0024] Figure 8 is a schematic diagram of data conversion in autoregressive module optimization in the audio restoration method provided by some embodiments of the present application;

[0025] Figure 9 is a schematic diagram of obtaining discrete coding vector sequence samples in the audio restoration method provided in some embodiments of the present application;

[0026] Figure 10 This is one of the schematic diagrams of the client interface in the audio repair method provided in some embodiments of the present application;

[0027] Figure 11 This is a second schematic diagram of the client interface in the audio repair method provided in some embodiments of the present application;

[0028] Figure 12 This is a third schematic diagram of the client interface in the audio repair method provided in some embodiments of the present application;

[0029] Figure 13 is a schematic structural diagram of an audio repair device provided in some embodiments of the present application;

[0030] Figure 14 is a schematic structural diagram of an electronic device provided by some embodiments of the present application;

[0031] Figure 15 This is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application. DETAILED DESCRIPTION

[0032] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0033] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0034] The terms used in the embodiments of this application are only used to explain the specific embodiments of this application and are not intended to limit this application. The following is an explanation of the terms involved in the embodiments of this application.

[0035] Large audio model: This model includes an audio word segmenter and an autoregressive module, with a large number of parameters. The autoregressive module is a decoder-only transformer that uses audio information as the smallest semantic unit (token).

[0036] Decoder-Only Transformer: This is an autoregressive Transformer architecture that uses causal masks to ensure that data at each position can only interact with the information above it during training. The Transformer is a neural network architecture that processes sequential data through a self-attention mechanism, strengthening the interaction of information between different positions.

[0037] Autoregressive (AR): A model structure used to model sequential data. During inference, the current token is predicted using the smallest semantic unit (token) in the previous context. The predicted token is then considered part of the previous context and used to predict subsequent tokens. Tokens can represent words, phrases, frames, and other content.

[0038] Audio Tokenizer: A module that converts input audio data into discrete tokens after preprocessing.

[0039] Codebook: A set of fixed-size vectors representing a dataset.

[0040] First audio clip: An audio clip containing unwanted noise in the user's recorded audio or video. The goal of audio restoration is to remove the noise from the first audio clip, resulting in an overall natural-sounding audio or video.

[0041] Generative Adversarial Networks (GAN): A neural network architecture consisting of two modules: a generator and a discriminator. The generator is responsible for generating data by imitating the original input, while the discriminator is responsible for determining whether the input is the original input or data generated by the generator.

[0042] Vector Quantization-Generative Adversarial Network (VQ-GAN): A vector quantization (VQ) module is introduced between the encoder and decoder in the generator, and a fixed set of codebook vectors, called the codebook, is maintained. The encoder output is mapped to the codebook through the VQ module, and the resulting codebook vectors are fed into the subsequent decoder. The index value of the codebook vector can be used as the discrete code vector output by the generator.

[0043] Logits: The output of a neural network without an activation function, which can be regarded as the unnormalized probability of each category.

[0044] Gumbel Softmax: A sampling method that applies Gumbel-distributed noise to the logits and then applies the Softmax function to obtain a probability distribution vector. Taking the maximum value of the probability distribution vector output by the Gumbel Softmax function is equivalent to sampling the probability distribution vector obtained by normalizing the logits. Compared to sampling directly based on the Softmax output, using Gumbel Softmax sampling allows for gradient propagation of the sampling results, which can be used for neural network training.

[0045] Self-Supervised Learning: A machine learning task paradigm that does not require manual data annotation. It uses a portion of input data to train the model to predict the input of another portion of input data.

[0046] The audio repair method provided in the embodiment of the present application can be applied to scenarios where a user needs to repair an audio segment containing noise. Among them, one specific application scenario can be that a user repairs an audio segment corresponding to a certain video segment in a recorded video file. Another specific application scenario can be that a user repairs an audio segment in a recorded audio file. It will be understood that the above specific application scenarios are only examples, and the application scenarios of the audio repair method of the embodiment of the present application can include but are not limited to the above specific application scenarios.

[0047] The audio repair method provided in the embodiment of the present application is described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.

[0048] Figure 1 : is a flowchart of an audio repair method provided in an embodiment of the present application. The audio repair method may include:

[0049] Step 101 : Feature extraction is performed on a first audio segment to be repaired, a preceding audio segment of the first audio segment, and a following audio segment of the first audio segment to obtain a first acoustic feature vector of the first audio segment, a second acoustic feature vector of the preceding audio segment, and a third acoustic feature vector of the following audio segment.

[0050] In some embodiments of the present application, the first audio segment can be an audio segment to be repaired, extracted from a video file or an audio file by the client in response to a user's operation. The upper audio segment is the first n seconds of the audio segment adjacent to the first audio segment, and the lower audio segment is the last m seconds of the audio segment adjacent to the first audio segment. Where n and m can be set according to actual needs, for example, they can range from 1 to 5. The following description uses n and m both as 5 as an example.

[0051] It is understood that the execution entity of the audio repair method can be set according to actual circumstances. For example, the audio repair method can be executed by a client. Alternatively, the audio repair method can be executed by a server. In this case, the client can upload the intercepted first audio clip, the preceding audio clip of the first audio clip, and the following audio clip to the server for subsequent operations.

[0052] Acoustic feature vectors of the first audio segment, the previous audio segment, and the next audio segment may be extracted respectively to obtain a first acoustic feature vector of the first audio segment, a second acoustic feature vector of the previous audio segment, and a third acoustic feature vector of the next audio segment.

[0053] Among them, acoustic feature vectors may include vectors such as time domain features, frequency domain features, time-frequency domain features and advanced features. For example, time domain features may include short-time energy, zero-crossing rate and amplitude envelope, frequency domain features may include spectrum centroid, spectrum bandwidth and spectrum flatness, time-frequency domain features may include Mel-frequency cepstral coefficients, chromaticity features and spectral contrast, and advanced features may include fundamental frequency, rhythm features and deep learning features.

[0054] In some embodiments of the present application, the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector extracted are all 512-dimensional acoustic feature vectors with a sampling rate of 80 samples / second for description.

[0055] In step 102, the first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector are input into a large audio model, autoregressive prediction is performed through the large audio model to obtain a predicted discrete coding vector, and the predicted discrete coding vector is decoded to obtain a repaired second audio segment.

[0056] In step 102, the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector can be input into a large audio model. The large audio model can include a pre-trained autoregressive module, which is used to perform autoregressive prediction on the input first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector. For example, a first discrete coding vector of the first acoustic feature vector, a second discrete coding vector of the second acoustic feature vector, and a third discrete coding vector of the third acoustic feature vector can be extracted first. The autoregressive module then predicts a predicted discrete coding vector for representing the audio segment after noise removal based on the first discrete coding vector, the second discrete coding vector, and the third discrete coding vector. The predicted discrete coding vector can be associated with the second discrete coding vector and the third discrete coding vector.

[0057] The large audio model may further include an audio word segmenter for decoding the discrete code vectors, and the audio word segmenter may be used to convert the discrete code vectors into audio segments. Based on this, the audio word segmenter may be used to decode the predicted discrete code vectors to obtain a restored second audio segment.

[0058] In an embodiment of the present application, feature extraction can be performed on the first audio segment to be repaired, the previous audio segment of the first audio segment, and the following audio segment to obtain a first acoustic feature vector of the first audio segment, a second acoustic feature vector of the previous audio segment, and a third acoustic feature vector of the following audio segment; the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector are input into the audio large model for autoregressive prediction to obtain a predicted discrete coding vector, and the predicted discrete coding vector is decoded to obtain the repaired second audio segment.

[0059] In this way, when repairing the first audio clip, the previous audio clip and the following audio clip of the first audio clip can also be considered to extract the acoustic feature vectors of the three, so that the acoustic feature vectors of the three are all used as inputs of the audio big model, so that when the audio big model performs autoregressive prediction, it can combine the context information, so that the repaired second audio clip obtained by the audio big model is more naturally connected with the context and the repair effect is better.

[0060] In some embodiments, the audio model may include an audio word segmenter and an autoregressive module; step 102 may further include the following steps:

[0061] Inputting the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector into an audio word segmenter, and extracting a first discrete coding vector of the first acoustic feature vector, a second discrete coding vector of the second acoustic feature vector, and a third discrete coding vector of the third acoustic feature vector through the audio word segmenter;

[0062] Inputting the discrete coding vector sequence into the autoregressive module, performing deep inference on the discrete coding vector sequence through the autoregressive module, and outputting a predicted discrete coding vector; the discrete coding vector sequence includes a first discrete coding vector, a second discrete coding vector, and a third discrete coding vector;

[0063] The predicted discrete coding vector is input into the audio word segmenter, and the predicted discrete coding vector is decoded by the audio word segmenter to obtain a repaired second audio segment.

[0064] In some embodiments of the present application, the first acoustic feature vector, the second acoustic feature vector and the third acoustic feature vector can be input into an audio word segmenter, and the first acoustic feature vector, the second acoustic feature vector and the third acoustic feature vector are encoded by the audio word segmenter, and their discrete representations are further determined to obtain a first discrete encoding vector of the first acoustic feature vector, a second discrete encoding vector of the second acoustic feature vector and a third discrete encoding vector of the third acoustic feature vector.

[0065] For example, the audio word segmenter can downsample the acoustic feature vector to obtain the downsampled acoustic feature vector, compare the downsampled acoustic feature vector with multiple pre-set code book vectors in the code book, calculate the cosine similarity, and assign the corresponding code book vector according to the cosine similarity. The index value of the assigned code book vector in the code book is the discrete coding vector of the acoustic feature vector.

[0066] A discrete code vector sequence can be determined based on the first discrete code vector, the second discrete code vector, and the third discrete code vector. For example, the first discrete code vector, the second discrete code vector, and the third discrete code vector can be concatenated to obtain a discrete code vector sequence. The concatenation order can be set based on actual needs and is not specifically limited here. It is understood that the concatenation order must be consistent with that used during training of the autoregressive module.

[0067] For example, in some embodiments of the present application, the second discrete code vector c1, the third discrete code vector c2 and the first discrete code vector c3 may be concatenated in the order of the second discrete code vector c1, the third discrete code vector c2 and the first discrete code vector c3. When concatenating, a special token, namely the start symbol b, may be added to the beginning and end of each of the second discrete code vector c1, the third discrete code vector c2 and the first discrete code vector c3. i and the terminator e i , and a start symbol b4 corresponding to the predicted discrete code vector c4 representing the second audio segment may be added at the end. The discrete code vector sequence obtained after splicing may be: {b1,c1,e1,b2,c2,e2,b3,c3,e3,b4}.

[0068] The discrete coding vector sequence can be input into the autoregressive module, and the discrete coding vector sequence can be deeply inferred by the autoregressive module to output a predicted discrete coding vector. Among them, the autoregressive module adopts the design of Decoder-Only Transformer, and the pre-trained autoregressive module can infer the predicted discrete coding vector in the form of autoregressive prediction. For example, when the candidate discrete coding vector inferred by the autoregressive module corresponds to the reference terminator e4 corresponding to the predicted discrete coding vector, the candidate discrete coding vector can be used as the predicted discrete coding vector c4. The autoregressive module can output the predicted discrete coding vector c4.

[0069] The predicted discrete code vector c4 can be input into an audio word segmenter, which decodes the predicted discrete code vector. For example, a codebook vector corresponding to the predicted discrete code vector c4 can be matched from a codebook. Based on the codebook vector, an acoustic feature vector corresponding to the second audio segment can be reconstructed. The acoustic feature vector corresponding to the second audio segment can then be converted into a second audio segment. The second audio segment can then be used to replace the first audio segment.

[0070] In one example, if Figure 2As shown, original audio 201 can determine the region corresponding to the first audio segment based on the user's selection operation. Audio is then captured based on this region to obtain the first audio segment a3, as well as the preceding audio segment a1 and the following audio segment a2 adjacent to the first audio segment a3. Acoustic feature vectors are extracted for the first audio segment a3, the preceding audio segment a1, and the following audio segment a2. Discrete code vectors are then extracted based on the acoustic feature vectors to obtain a second discrete code vector c1, a third discrete code vector c2, and a first discrete code vector c3. The second discrete code vector c1, the third discrete code vector c2, and the first discrete code vector c3 are concatenated to obtain a discrete code vector sequence 202: {b1, c1, e1, b2, c2, e2, b3, c3, e3, b4}. Discrete code vector sequence 202 is input into an autoregressive module, which outputs a predicted discrete code vector c4 whose terminator is the reference terminator e4. The predicted discrete code vector c4 is decoded to obtain the repaired second audio segment a4. The restored second audio segment a4 is used to replace the first audio segment of the original audio 201 to obtain the restored audio 203 .

[0071] In this way, when repairing the first audio clip, the previous audio clip and the following audio clip of the first audio clip will be considered, and the acoustic feature vectors of the three will be extracted. Based on the acoustic feature vectors of the three, the discrete coding vectors of the three are obtained to form a discrete coding vector sequence. The discrete coding vector sequence is used as the input of the autoregressive module so that the autoregressive module can combine the context information, so that the repaired second audio clip obtained by the autoregressive module is more naturally connected with the context and the repair effect is better.

[0072] In some embodiments, before inputting the discrete code vector sequence into an autoregressive module, performing deep inference on the discrete code vector sequence through the autoregressive module, and outputting a predicted discrete code vector, the method may further include:

[0073] identifying a vocal processing intent for a first audio segment;

[0074] The discrete coding vector sequence is input into the autoregressive module, and the discrete coding vector sequence is deeply inferred by the autoregressive module to output the predicted discrete coding vector, including:

[0075] In a case where the vocal processing is intended to retain speech information in the first audio segment, the discrete code vector sequence is input into a first autoregressive module, deep inference is performed on the discrete code vector sequence by the first autoregressive module, and a predicted discrete code vector is output;

[0076] In a case where the vocal processing is intended to remove speech information from the first audio segment, the discrete code vector sequence is input into a second autoregressive module, deep inference is performed on the discrete code vector sequence by the second autoregressive module, and a predicted discrete code vector is output;

[0077] Among them, the first autoregressive module is trained based on the first audio sequence samples, and the second autoregressive module is trained based on the second audio sequence samples. The first audio sequence samples and the second audio sequence samples both include main audio segment samples, previous audio segment samples, following audio segment samples, and noisy audio segment samples after the main audio segment samples are denoised; the main audio segment samples in the first audio sequence samples include voice information, and the main audio segment samples in the second audio sequence samples do not include voice information.

[0078] In some embodiments of the present application, the vocal processing intent for the first audio clip can be identified based on the vocal processing settings set by the user during the audio restoration operation, and the intent of the first audio clip can be determined as to whether to retain the voice information in the first audio clip, i.e., retain the voice, or remove the voice information in the first audio clip, i.e., remove the voice. Different autoregressive modules can be selected for autoregressive prediction based on different vocal processing intents.

[0079] For example, a first autoregressive module can be trained based on a first audio sequence sample. The first audio sequence sample can include a main text audio segment sample, a preceding audio segment sample, a following audio segment sample, and a noisy audio segment sample obtained by noisy processing the main text audio segment sample. The main text audio segment sample in the first audio sequence sample includes speech information. Therefore, the first autoregressive module trained based on the first audio sequence sample can be suitable for use in scenarios where the speech information in the first audio segment is retained.

[0080] A second autoregressive module can be trained based on a second audio sequence sample. The second audio sequence sample can include a main text audio segment sample, a preceding audio segment sample, a following audio segment sample, and a noisy audio segment sample obtained by denoising the main text audio segment sample. The main text audio segment sample in the second audio sequence sample does not include speech information. Therefore, the second autoregressive module trained based on the second audio sequence sample can be applied to a scenario where speech information is removed from the first audio segment.

[0081] If the vocal processing is intended to preserve the speech information in the first audio segment, the discrete code vector sequence can be input into a first autoregressive module, which then performs deep inference on the discrete code vector sequence and outputs a predicted discrete code vector. In this way, the second audio segment decoded based on the predicted discrete code vector will retain the vocal information.

[0082] If the vocal processing is intended to remove speech information from the first audio clip, the discrete code vector sequence is input into the second autoregressive module, which performs deep inference on the discrete code vector sequence and outputs a predicted discrete code vector. This process then removes the vocals from the restored second audio clip, decoded based on the predicted discrete code vector.

[0083] Among them, the output of retaining human voice and removing human voice in different scenarios is different.

[0084] like Figure 3a As shown in scenario 1, the first audio segment a3 contains both vocals a and b, along with ambient sound, while the preceding audio segment a1 and the following audio segment a2 contain only vocals a and ambient sound. If the vocals are retained, the audio restoration process retains vocals a and removes vocals b, resulting in a second audio segment a4 containing both vocals a and ambient sound. If the vocals are removed, the restoration process removes both vocals a and b, resulting in a second audio segment a4 containing only ambient sound.

[0085] like Figure 3b As shown in the example, scenario 1 involves the first audio segment a3 containing a human voice a, ambient sound, and a sudden noise, such as a car horn. However, the preceding audio segment a1 and the following audio segment a2 contain only the human voice a and the ambient sound. If the human voice is retained, the audio restoration process retains the human voice a and removes the sudden noise, resulting in a second audio segment a4 containing both the human voice a and the ambient sound. If the human voice is removed, the restoration process removes both the human voice a and the sudden noise, resulting in a second audio segment a4 containing only the ambient sound.

[0086] In this way, the user's vocal processing intention can be identified when repairing the audio, and the corresponding autoregressive module can be selected to perform audio repair according to different vocal processing intentions, so as to obtain a repaired second audio segment that meets the vocal processing intention.

[0087] In some embodiments, the audio word segmenter includes a generator, the generator including an encoder and a vector quantization module;

[0088] Inputting the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector into an audio word segmenter, and extracting a first discrete coding vector of the first acoustic feature vector, a second discrete coding vector of the second acoustic feature vector, and a third discrete coding vector of the third acoustic feature vector through the audio word segmenter may include:

[0089] Inputting the first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector into an encoder, downsampling the first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector through the encoder to obtain a first encoding vector of the first acoustic eigenvector, a second encoding vector of the second acoustic eigenvector, and a third encoding vector of the third acoustic eigenvector;

[0090] Inputting the first coding vector, the second coding vector, and the third coding vector into a vector quantization module, the vector quantization module splitting the first coding vector into N first low-dimensional vectors, splitting the second coding vector into N second low-dimensional vectors, and splitting the third coding vector into N third low-dimensional vectors, where N is an integer greater than 1;

[0091] The vector quantization module calculates a first cosine similarity between each first low-dimensional vector and the plurality of codebook vectors, a second cosine similarity between each second low-dimensional vector and the plurality of codebook vectors, and a third cosine similarity between each third low-dimensional vector and the plurality of codebook vectors;

[0092] The vector quantization module determines a first discrete code vector according to the first cosine similarity, determines a second discrete code vector according to the second cosine similarity, and determines a third discrete code vector according to the third cosine similarity.

[0093] In some embodiments of the present application, Figure 4 As shown, the audio word segmenter 400 may include a generator 410, which includes an encoder 411 and a VQ (Vector Quantization) module 412. Encoder 411 may be designed using any 2D convolutional neural network, and the output of encoder 411 is downsampled relative to the input by a factor of [4, 4] along the time and frequency dimensions. VQ module 412 may include a codebook consisting of 1024 32-dimensional vectors.

[0094] The first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector can be input into the encoder 411, and the first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector can be downsampled by the encoder 411 to obtain a first encoding vector of the first acoustic eigenvector, a second encoding vector of the second acoustic eigenvector, and a third encoding vector of the third acoustic eigenvector. For example, the first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector can all be acoustic eigenvectors with a speed of 80 samples / second and 512 dimensions, and the first encoding vector, the second encoding vector, and the third encoding vector output by the encoder 411 are all encoding vectors with a speed of 20 samples / second and 128 dimensions.

[0095] The first coding vector, the second coding vector and the third coding vector are input into the VQ module 412, and the first coding vector, the second coding vector and the third coding vector are split by the VQ module 412 to obtain N first low-dimensional vectors, N second low-dimensional vectors and N third low-dimensional vectors respectively.

[0096] The VQ module 412 calculates a first cosine similarity between each first low-dimensional vector and multiple codebook vectors, a second cosine similarity between each second low-dimensional vector and multiple codebook vectors, and a third cosine similarity between each third low-dimensional vector and multiple codebook vectors.

[0097] For example, the 128-dimensional encoding vector can be split into four 32-dimensional low-dimensional vectors, and then the cosine similarity between each low-dimensional vector and the 1024 32-dimensional vectors in the code book is calculated, that is, the first cosine similarity, the second cosine similarity and the third cosine similarity can be obtained.

[0098] The calculation of the cosine similarity between the low-dimensional vector and the codebook vector is the same as the conventional method for calculating the cosine similarity between vectors. The calculation formula for the cosine similarity between vectors is shown in formula (1):

[0099] cosθ=(A·B) / (‖A‖·‖B‖) (1)

[0100] Where cosθ is the cosine similarity between vector A and vector B, (A·B) is the dot product of vector A and vector B, and ‖A‖ and ‖B‖ are the moduli of vector A and vector B, respectively.

[0101] The following will take the cosine similarity between two 2-dimensional vectors as an example to introduce the method of calculating cosine similarity between vectors.

[0102] Vector A is [1,2], and vector B is [2,3]. According to formula (1), the dot product of vector A and vector B is: 1*2+2*3=8, and the modulus of vector A is: √(1 2 +2 2 )=√(1+4)=√5≈2.236, the modulus of vector B is:√(2 2 +3 2 )=√(4+9)=√13≈3.606. Then the cosine similarity between vector A and vector B is: 8 / √(5*13)≈0.992.

[0103] Each low-dimensional vector can be assigned to the codebook vector with the highest cosine similarity. That is, each codebook vector can correspond to four codebook vectors. The index values of these four codebook vectors are the discrete representations of the acoustic feature vectors of this frame, namely discrete code vectors. In this way, the first discrete code vector of the first acoustic feature vector, the second discrete code vector of the second acoustic feature vector, and the third discrete code vector of the third acoustic feature vector can be obtained.

[0104] In this way, the first acoustic feature vector, the second acoustic feature vector and the third acoustic feature vector can be discretized through the encoder and VQ module in the audio segmenter to obtain the first encoding vector, the second encoding vector and the third encoding vector as the input of the subsequent autoregressive module, which can reduce the computational complexity of data processing in the autoregressive module and improve the model processing efficiency.

[0105] In some embodiments, the audio word segmenter further comprises a vocoder, and the generator further comprises a decoder;

[0106] Inputting the predicted discrete code vector into the audio word segmenter, decoding the predicted discrete code vector by the audio word segmenter to obtain a repaired second audio segment, may include:

[0107] Inputting the predicted discrete code vector into a decoder, and decoding the predicted discrete code vector through the decoder to obtain a fourth acoustic feature vector;

[0108] The fourth acoustic feature vector is input into a vocoder, and the vocoder converts the fourth acoustic feature vector into a second audio segment.

[0109] In some embodiments of the present application, Figure 4 As shown, the audio word segmenter may further include a vocoder 420, and the generator 410 may further include a decoder 413. Decoder 413 may be designed using any 2D convolutional neural network. The upsampling ratio of the decoder 413 output relative to the input along the time and frequency dimensions is [4, 4]. As a tool for converting acoustic feature vectors into audio, vocoder 420 may use any model structure or existing open source tools, such as the vocoder module of HiFi-GAN.

[0110] The predicted discrete coding vector can be input into the decoder 413, and the predicted discrete coding vector can be decoded by the decoder 413 to obtain a fourth acoustic feature vector. For example, the predicted discrete coding vector can be input into the decoder 413, and the decoder 413 first obtains a codebook vector corresponding to the predicted discrete coding vector based on the index value represented by the predicted discrete coding vector, combines the codebook vectors based on the codebook vectors to obtain a 20 sample / second, 128-dimensional vector, and then upsamples the vector to obtain a fourth acoustic feature vector, which is an 80 sample / second, 512-dimensional acoustic feature vector.

[0111] The fourth acoustic feature vector may be input to the vocoder 420 , and the vocoder 420 may convert the fourth acoustic feature vector into a second audio segment.

[0112] In this way, the predicted discrete coding vector can be decoded by the decoder in the audio segmenter to obtain a fourth acoustic feature vector, and the fourth acoustic feature vector can be transformed by the vocoder to obtain a second audio segment, so that the first audio segment can be replaced based on the second audio segment to obtain audio with a natural repair effect.

[0113] In some embodiments, the audio word segmenter further includes a discriminator; the training method of the audio word segmenter is as follows:

[0114] Obtaining a first acoustic eigenvector sample;

[0115] Inputting the first acoustic feature vector sample into the generator, the encoder downsamples the first acoustic feature vector sample to obtain an encoded vector sample;

[0116] The vector quantization module splits the coded vector sample into N low-dimensional vector samples, calculates the cosine similarity between each low-dimensional vector sample and multiple codebook vectors, and determines N codebook vector samples based on the cosine similarity;

[0117] Inputting N codebook vector samples into a decoder, upsampling the N codebook vector samples through the decoder to obtain a second acoustic feature vector sample;

[0118] Calculating a first loss value based on a pre-constructed first loss function and a second acoustic feature vector sample;

[0119] Inputting the acoustic feature vector sample or the second acoustic feature vector sample into the discriminator to obtain a probability value that the input sample is the acoustic feature vector sample;

[0120] Calculating a second loss value based on a pre-constructed second loss function and a probability value;

[0121] According to the sum of the first loss value and the second loss value, the parameters of the generator and the parameters of the discriminator are updated to obtain the trained audio segmenter.

[0122] In some embodiments of the present application, Figure 4 As shown, the audio word segmenter can also include a discriminator 430. The discriminator 430 can be designed using any 2D convolutional neural network. The discriminator 430 outputs a downsampled output relative to the input along the time and frequency dimensions with a magnification of [1,32], and the last layer outputs logits. The input of the discriminator 430 can be a first acoustic feature vector sample x with 80 samples / second and 512 dimensions, or a second acoustic feature vector sample x with 80 samples / second and 512 dimensions reconstructed by the generator. The output is 16-dimensional logits at 80 samples / second. Logits can be the probability that the unnormalized input sample is a sample of the first acoustic eigenvector.

[0123] A first acoustic feature vector sample may be obtained, wherein the first acoustic feature vector sample is an acoustic feature vector with a frequency of 80 samples / second and 512 dimensions. The first acoustic feature vector sample is input into the generator 410, and the encoder 411 downsamples the first acoustic feature vector sample to obtain a coded vector sample with a frequency of 20 samples / second and 128 dimensions.

[0124] The VQ module 412 splits the coding vector sample into four 32-dimensional low-dimensional vector samples, and then calculates the cosine similarity between each low-dimensional vector sample and the 1024 32-dimensional vectors in the code book, and assigns each low-dimensional vector sample to the code book vector sample with the highest cosine similarity, that is, each coding vector sample can correspond to four code book vector samples.

[0125] The calculation of cosine similarity between vectors is as mentioned above and will not be repeated here.

[0126] The codebook vector samples are upsampled by the decoder 413 to obtain second acoustic feature vector samples of 512 dimensions at 80 samples / second.

[0127] The generator 410 takes the first acoustic feature vector sample x as input and the second acoustic feature vector sample As output, based on the pre-built first loss function and the second acoustic feature vector sample Calculate the first loss value. The first loss value can be expressed as formula (2):

[0128]

[0129] Among them, L VQ is the first loss value; E is the encoder in the generator; x is the first acoustic feature vector sample input to the generator, is the second acoustic feature vector sample output by the generator, z is the discrete encoding vector sample obtained by feeding x into the encoder and VQ module. sg is the stop-gradient operation, which is the identity function during forward propagation and has a partial derivative of 0 during backward propagation.

[0130] Input the first acoustic feature vector sample or the second acoustic feature vector sample into the discriminator, obtain the probability value of the input sample being the first acoustic feature vector sample, and calculate the second loss value based on the pre-constructed second loss function and the probability value. The second loss value can be expressed as formula (3):

[0131]

[0132] Among them, L GAN is the second loss value, x is the first acoustic feature vector sample of the input generator, is the second acoustic feature vector sample output by the generator, and D is the discriminator.

[0133] The parameters of the generator and the discriminator can be updated according to the sum of the first loss value and the second loss value to obtain the trained audio segmenter. For example, the sum of the first loss value and the second loss value L can be calculated first, as shown in formula (4):

[0134] L=L GAN +L VQ (4)

[0135] Among them, L VQ is the first loss value, L GAN is the second loss value.

[0136] The parameters of the generator and the discriminator can be updated so that the sum of the first loss value and the second loss value changes. The parameters of the generator and the discriminator corresponding to the minimum sum of the first loss value and the second loss value are used as the final parameters to obtain the trained audio segmenter.

[0137] In this way, the audio segmenter can be trained based on the first acoustic feature vector sample and combined with the loss function of the generator and the loss function of the discriminator to obtain an audio segmenter with higher accuracy, so that a more accurate discrete coding vector of the audio segment can be extracted as the input of the autoregressive module, thereby improving the accuracy of the output result of the autoregressive module, and thus a more accurate repaired second audio segment can be obtained.

[0138] In some embodiments, the autoregressive module may include an acoustic embedding submodule, a sequence modeling submodule, and an encoding prediction submodule;

[0139] Inputting the discrete coding vector sequence into the autoregressive module, performing deep inference on the discrete coding vector sequence through the autoregressive module, and outputting the predicted discrete coding vector, which may include:

[0140] Input the discrete encoding vector sequence into the acoustic embedding submodule to obtain the latent variable;

[0141] Input the latent variables into the sequence modeling submodule to obtain the latent variable sequence;

[0142] The last frame in the latent variable sequence is input into the encoding prediction submodule, multiple probability distribution vectors are generated by the encoding prediction submodule, and core sampling is performed on each probability distribution vector to obtain multiple candidate discrete encoding vectors;

[0143] Among multiple candidate discrete coding vectors, a candidate discrete coding vector corresponding to a reference terminator is determined as a predicted discrete coding vector, and the predicted discrete coding vector is output.

[0144] In some embodiments of the present application, Figure 5 As shown, the autoregressive module 500 may include an acoustic embedding submodule 510 , a sequence modeling submodule 520 , and a coding prediction submodule 530 .

[0145] Among them, the acoustic embedding submodule 510 is composed of 1032 1536-dimensional embedding vectors, corresponding to 1024 discrete coding vectors and 8 special tokens, namely the start and end symbols corresponding to the first coding vector, the second coding vector, the third coding vector and the predicted coding vector respectively.

[0146] The sequence modeling submodule 520 consists of 24 stacked transformer blocks. Each transformer block consists of a multi-head attention module and a fully connected network connected end to end. The input and output dimensions are both [T, 1536], where T is the length of the time dimension sequence. The multi-head attention module is used to calculate the multi-head self-attention of the sequence. The number of heads is 16 and the dimension is 4096. A causal mask (CausalMask) is used to ensure that the data at each position can only interact with the information above it during training. Rotational Position Encoding (RoPE) is applied to the input sequence before calculation. The fully connected network consists of two layers of linear layers. The first layer first projects each 1536-dimensional vector to 4096 dimensions, applies the ReLU activation function, and then reduces the dimension to 1536 dimensions through the second layer. A normalization layer (LayerNorm) and a skip connection are applied between each layer of transformer blocks to enhance model stability.

[0147] The encoding prediction submodule 530 consists of four independent linear layers, each with an input dimension of 1536 and an output dimension of 1032.

[0148] The discrete code vector sequence can be input into the acoustic embedding submodule 510 to obtain a latent variable. For example, the input is a discrete code vector sequence of [T, 4]. Based on the four discrete code vector sequences in each frame, the acoustic embedding submodule 510 obtains the corresponding four 1536-dimensional embedding vectors, adds them together, and outputs a latent variable x of dimension [T, 1536].

[0149] The latent variables are input into the sequence modeling submodule 520 to obtain a latent variable sequence. For example, the latent vector x is input into the sequence modeling submodule 520, and after passing through 24 transformer blocks, a latent variable sequence with a dimension of [T, 1536] is output.

[0150] The last frame in the latent variable sequence is input to the coding prediction submodule 530. Multiple probability distribution vectors are generated by the coding prediction submodule 530. Core sampling is performed on each probability distribution vector to obtain multiple candidate discrete coding vectors. Among the multiple candidate discrete coding vectors, the candidate corresponding to the reference terminator is determined as the predicted discrete coding vector, and the predicted discrete coding vector is output.

[0151] For example, the last frame of the latent variable sequence, with a dimension of [1, 1536], is fed into the code prediction submodule 530. This is then fed into four linear layers and subjected to a Softmax function, resulting in four 1032-dimensional probability distribution vectors. These four probability distribution vectors are then subjected to core sampling with a value of p = 0.5. This involves randomly selecting a sample from the smallest subset whose sum of probabilities is greater than p. The index values corresponding to the four sampled samples are the autoregressive module's predictions for the discrete code vector for the next frame. The predicted value is used as input for the next moment, and the above steps are repeated. When any of the four predicted discrete code vectors corresponds to the special token "e4," inference concludes. The discrete code vector at this point is the predicted discrete code vector c4, and the predicted discrete code vector c4 is output.

[0152] In this way, the acoustic embedding submodule, sequence modeling submodule, and encoding prediction submodule in the autoregressive module perform deep inference on the input discrete encoding vector sequence and predict the predicted discrete encoding vector corresponding to the repaired second audio segment, which can then be decoded to obtain the repaired second audio segment. This entire process achieves automated audio restoration without requiring the user to understand any technical terminology, making it easy to use.

[0153] In some embodiments, the training method of the autoregressive module is as follows:

[0154] Obtaining discrete coding vector sequence training samples; wherein each discrete coding vector sequence training sample is obtained by splicing the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample, and the fourth discrete coding vector sample, and the splicing position of the fourth discrete coding vector sample is located at the end of the discrete coding vector sequence training sample, and corresponding start symbols are added to the heads of the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample, and the fourth discrete coding vector sample, respectively, and corresponding terminators are added to the tails of the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample, and the fourth discrete coding vector sample, respectively;

[0155] Perform autoregressive training on the model based on the discrete coding vector sequence training samples to obtain the trained initial module;

[0156] Obtain discrete code vector sequence test samples; wherein each discrete code vector sequence test sample is obtained by splicing the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample; corresponding start symbols are added to the heads of the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample; corresponding terminators are added to the tails of the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample; and the start symbol corresponding to the fourth discrete code vector sample is added after the terminator corresponding to the fifth discrete code vector sample; the fifth discrete code vector sample is the discrete code vector sample whose splicing position is located at the tail of the discrete code vector sequence test sample;

[0157] Input the discrete coding vector sequence test sample into the initial module and output the predicted discrete coding vector sample;

[0158] According to the difference value between the predicted discrete coding vector sample and the fourth discrete coding vector sample, the initial module is optimized to obtain an autoregressive module.

[0159] In some embodiments of the present application, a discrete coding vector sequence training sample can be obtained; it can be understood that, among them, the splicing order of the first discrete coding vector sample, the second discrete coding vector sample and the third discrete coding vector sample can be arbitrarily combined, but the splicing position of the fourth discrete coding vector sample must be located at the end of the discrete coding vector sequence training sample. The corresponding start symbols are added to the heads of the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample and the fourth discrete coding vector sample, and the corresponding terminators are added to the tails of the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample and the fourth discrete coding vector sample. Finally, a dimension of An integer vector of , which serves as the input to the autoregressive module.

[0160] For example, taking the case where discrete coding vector sequence training samples are spliced in the order of the second discrete coding vector sample, the third discrete coding vector sample, the first discrete coding vector sample, and the fourth discrete coding vector sample, the spliced discrete coding vector sequence training samples may be: {b1, c'1, e1, b2, c'2, e2, b3, c'3, e3, b4, c'4, e4}.

[0161] The model can be trained autoregressively based on discrete coding vector sequence training samples to obtain the trained initial module.

[0162] For example, Figure 6As shown, after the discrete coding vector sequence training samples are obtained, they can be processed to obtain discrete coding vector sequence training samples 601 as model input and discrete coding vector sequence training samples 602 as model output.

[0163] The discrete coded vector sequence training sample 601 used as model input can be: {b1, c'1, e1, b2, c'2, e2, b3, c'3, e3, b4, c'4}. The discrete coded vector sequence training sample 602 used as model output can be: {c'1, e1, b2, c'2, e2, b3, c'3, e3, b4, c'4, e4}. Based on this model pre-training, the prediction target for each input token is the token at the next moment.

[0164] The autoregressive module 500 receives a discrete code vector sequence training sample 601 of dimension [T, 4] as the model input. After passing through the acoustic embedding submodule 510, the sequence modeling submodule 520, and the coding prediction submodule 530, the logits 603 of dimension [T, 4, 1032] is obtained. After softmax normalization, the cross entropy is calculated between the logits 603 and the discrete code vector sequence training sample 602 of dimension [T, 4] as the model output, as shown in formula (5):

[0165]

[0166] Where T is the length of the sequence training sample, M is the number of discrete coding vector samples at each moment (M=4), and D is the number of categories of discrete coding vector samples (D=1032). d,t,m is the probability value corresponding to the dth dimension of the mth discrete coding vector sample in the tth frame after applying Softmax to logits, y m,t To predict the value of the mth discrete coding vector sample in the tth frame.

[0167] The training generates an initial module that acquires the ability to infer the restored second audio segment. A discrete code vector sequence test sample can be input into the initial module, and a predicted discrete code vector sample can be output.

[0168] For example, a discrete coding vector sequence test sample can be obtained; wherein, each discrete coding vector sequence test sample is obtained by splicing the first discrete coding vector sample, the second discrete coding vector sample and the third discrete coding vector sample, and the splicing order of the first discrete coding vector sample, the second discrete coding vector sample and the third discrete coding vector sample is consistent with the discrete coding vector sequence training sample.

[0169] Corresponding start symbols are added to the heads of the first discrete coding vector sample, the second discrete coding vector sample, and the third discrete coding vector sample, respectively; corresponding terminators are added to the tails of the first discrete coding vector sample, the second discrete coding vector sample, and the third discrete coding vector sample, respectively; and a start symbol corresponding to the fourth discrete coding vector sample is added after the terminator corresponding to the discrete coding vector sample whose splicing position is at the tail of the discrete coding vector sequence test sample.

[0170] For example, Figure 7 As shown, the discrete code vector sequence test sample 701 can be: {b1, c'1, e1, b2, c'2, e2, b3, c'3, e3, b4}.

[0171] Input {b1, c'1, e1, b2, c'2, e2, b3, c'3, e3, b4} into the initial module 700, and the predicted discrete coding vector sample can be output. The reasoning process of the initial module is as mentioned above and will not be repeated here.

[0172] The initial module may be optimized according to the difference between the predicted discrete coding vector sample and the fourth discrete coding vector sample to obtain an autoregressive module.

[0173] For example, a loss function can be constructed based on the difference value between the predicted discrete coding vector sample and the fourth discrete coding vector sample, and the model parameters of the initial module can be updated to minimize the loss function to obtain an optimized autoregressive module.

[0174] In this way, the autoregressive module can be pre-trained to obtain an initial module with reasoning capabilities, and then the initial module can be optimized to obtain an optimized autoregressive module. This can ensure the accuracy of the autoregressive module. The repaired second audio segment obtained based on the autoregressive module reasoning can be closer to the original audio, and the connection with the context is more natural, thereby improving the effect of audio repair.

[0175] In some embodiments, inputting a discrete code vector sequence test sample into an initial module and outputting a predicted discrete code vector sample may include:

[0176] Input the discrete encoding vector sequence test sample into the acoustic embedding submodule in the initial module to obtain the latent variable sample;

[0177] Input the latent variable sample into the sequence modeling submodule in the initial module to obtain the latent variable sequence sample;

[0178] The last frame in the latent variable sequence sample is input into the coding prediction submodule in the initial module, multiple probability distribution vectors are generated by the coding prediction submodule, the discrete coding vector corresponding to the maximum value of the probability distribution vector is determined as the predicted discrete coding vector sample, and the predicted discrete coding vector sample is output.

[0179] In some embodiments of the present application, Figure 7 As shown, the discrete code vector sequence test sample 701 can be: {b1, c'1, e1, b2, c'2, e2, b3, c'3, e3, b4}.

[0180] {b1, c'1, e1, b2, c'2, e2, b3, c'3, e3, b4} can be input into the acoustic embedding submodule in the initial module 700 to obtain a latent variable sample. The latent variable sample is then input into the sequence modeling submodule in the initial module 700 to obtain a latent variable sequence sample. The processing of the acoustic embedding submodule and the sequence modeling submodule is as described above and will not be repeated here.

[0181] The last frame in the latent variable sequence sample can be input into the coding prediction submodule in the initial module 700, and multiple probability distribution vectors are generated by the coding prediction submodule. The discrete coding vector corresponding to the maximum value of the probability distribution vector is determined as the predicted discrete coding vector sample, and the predicted discrete coding vector sample is output.

[0182] For example, take the last frame of the latent variable sequence, which has a dimension of [1,1536], and send it to the coding prediction submodule in the initial module 700. Then, send it to four linear layers and apply the Softmax function to obtain four 1032-dimensional probability distribution vectors. Perform Gumbel-Softmax sampling on these four probability distribution vectors to obtain the index value corresponding to the maximum value of the probability distribution vector, which is the autoregressive module model's prediction of the discrete coding vector sample of the next frame. Use the predicted value as the input of the next moment and repeat the above steps. When any of the four predicted discrete coding vector samples corresponds to the special token "e4", the reasoning ends. At this time, the discrete coding vector sample is the predicted discrete coding vector sample.

[0183] In this way, Gumbel-Softmax sampling is used to ensure that the sampling process can perform gradient backpropagation, so that the parameters of the initial module can be updated, so as to optimize the model of the initial module and obtain an autoregressive module with better accuracy and better effect.

[0184] In some embodiments, performing model optimization on the initial module based on the difference between the predicted discrete code vector sample and the fourth discrete code vector sample to obtain the autoregressive module may include:

[0185] splicing the first discrete coding vector sample, the second discrete coding vector sample, and the third discrete coding vector sample to obtain a first discrete coding vector sequence sample;

[0186] splicing the fourth discrete coding vector sample, the second discrete coding vector sample, and the third discrete coding vector sample to obtain a second discrete coding vector sequence sample;

[0187] splicing the predicted discrete coding vector sample, the second discrete coding vector sample, and the third discrete coding vector sample to obtain a third discrete coding vector sequence sample;

[0188] Inputting the first discrete code vector sequence sample, the second discrete code vector sequence sample, and the third discrete code vector sequence sample into a decoder in the audio word segmentor, obtaining a first reconstructed acoustic feature vector sample of the first discrete code vector sequence sample, a second reconstructed acoustic feature vector sample of the second discrete code vector sequence sample, and a third reconstructed acoustic feature vector sample of the third discrete code vector sequence sample;

[0189] Inputting the first reconstructed acoustic feature vector sample, the second reconstructed acoustic feature vector sample, and the third reconstructed acoustic feature vector sample into the discriminator in the audio word segmenter, obtaining a first discriminant latent variable of the first reconstructed acoustic feature vector sample, a second discriminant latent variable of the second reconstructed acoustic feature vector sample, and a third discriminant latent variable of the third reconstructed acoustic feature vector sample;

[0190] The first discriminant latent variable is used as a positive sample, the second discriminant latent variable is used as a negative sample, and the third discriminant latent variable is used as an anchor sample to calculate the triplet loss;

[0191] According to the triplet loss value, the model parameters of the initial module are updated to obtain the autoregressive module, wherein the triplet loss value corresponding to the model parameters of the autoregressive module is the minimum value.

[0192] In some embodiments of the present application, Figure 8 As shown, the first discrete code vector sample c'3, the second discrete code vector sample c'1 and the third discrete code vector sample c'2 can be concatenated to obtain a first discrete code vector sequence sample. For example, the first discrete code vector sequence sample can be: {c'1, c'3, c'2}.

[0193] The fourth discrete code vector sample c'4, the second discrete code vector sample c'1 and the third discrete code vector sample c'2 may be concatenated to obtain a second discrete code vector sequence sample. For example, the second discrete code vector sequence sample may be: {c'1, c'4, c'2}.

[0194] It is also possible to predict discrete code vector samples The second discrete code vector sample c'1 and the third discrete code vector sample c'2 are concatenated to obtain a third discrete code vector sequence sample. For example, the third discrete code vector sequence sample can be:

[0195] The above splicing process can be executed in a serial manner or in a parallel manner, which is not specifically limited here.

[0196] The first discrete code vector sequence sample, the second discrete code vector sequence sample, and the third discrete code vector sequence sample can be input into the decoder 413 in the audio word segmenter 400 to obtain a first reconstructed acoustic feature vector sample of the first discrete code vector sequence sample, a second reconstructed acoustic feature vector sample of the second discrete code vector sequence sample, and a third reconstructed acoustic feature vector sample of the third discrete code vector sequence sample. The processing process of the decoder 413 is as described above and is not further described here.

[0197] The first reconstructed acoustic feature vector sample, the second reconstructed acoustic feature vector sample and the third reconstructed acoustic feature vector sample can be input into the discriminator 430 in the audio word segmenter 400 to obtain the first discriminant latent variable y of the first reconstructed acoustic feature vector sample. clean , the second discriminant latent variable y of the second reconstructed acoustic eigenvector sample noisy and the third discriminant latent variable y of the third reconstructed acoustic eigenvector sample pred ;

[0198] The first discriminant latent variable y clean As a positive sample, the second discriminant latent variable y noisy As a negative sample, the third discriminant latent variable y pred As anchor samples, the triplet loss value between the three is calculated as the training loss function. The training loss function is shown in formula (6):

[0199]

[0200] The model parameters of the initial module can be updated according to the triplet loss value to obtain the autoregressive module, wherein the triplet loss value corresponding to the model parameters of the autoregressive module is the minimum value. For example, the Triplet Loss can be minimized while reducing y pred with y clean The distance between them increases y pred with y noisyThe distance between the two is calculated, so that the initial module output gradually approaches the original audio and gradually moves away from the noisy frequency, resulting in an optimized autoregressive module. Using a distance metric, rather than calculating the distance between acoustic feature vectors, the similarity of the discriminator output is calculated, thereby optimizing the autoregressive module along the dimension of naturalness.

[0201] This approach uses the similarity of the discriminator's output as a basis for model optimization, optimizing the naturalness of the autoregressive module's contextual cohesion. This ensures that the restored second audio clip, derived from the autoregressive module's inference, seamlessly connects with the context. Furthermore, the reused audio word segmenter discriminates the naturalness of the autoregressive module and optimizes the model through comparative learning. This optimization eliminates the need for a new model structure, reducing training costs.

[0202] In some embodiments, obtaining a discrete code vector sequence sample may include:

[0203] Based on each of the M original audios, K audio segment combinations are randomly generated, where the audio segment combinations include a main audio segment sample, a preceding audio segment sample of the main audio segment sample, and a following audio segment sample; M and K are positive integers;

[0204] Obtain a reference audio segment from the reference audio; the reference audio is any original audio from the M original audios except the original audio;

[0205] Performing audio fusion processing on the main text audio clip sample and the reference audio clip to obtain a noise-added audio clip sample;

[0206] Obtaining an audio sequence sample according to the main text audio segment sample, the preceding audio segment sample, the following audio segment sample, and the noise-added audio segment sample;

[0207] Perform feature extraction on the audio sequence samples to obtain acoustic feature vector sequence samples;

[0208] The acoustic feature vector sequence samples are input into the audio word segmenter, and the discrete coding vector sequence training samples are extracted from the acoustic feature vector sequence samples by the audio word segmenter.

[0209] In some embodiments of the present application, M original audio files may be obtained. For each of the M original audio files, K audio segment combinations may be randomly generated. The audio segment combinations include a main audio segment sample, a preceding audio segment sample, and a following audio segment sample. The values of M and K may be set as required.

[0210] For example, assuming K is 10, for each original audio clip, 10 sets of duration sequences (d1, d3, d2) of preceding audio clip samples, main audio clip samples, and following audio clip samples can be randomly generated. d3 can take a random value between 1 and 20 seconds, and d1 and d2 can each take a random value between 0 and 20 seconds.

[0211] like Figure 9 As shown, for each original audio 901, a continuous segment with a duration of d1+d3+d2 is randomly selected and divided into three parts with a duration of d1, d3, and d2, corresponding to the above audio segment sample a′1, the main text audio segment sample a′4, and the following audio segment sample a′2 respectively.

[0212] A reference audio segment in the reference audio can be obtained, and the main audio segment sample can be noised according to the reference audio segment, that is, the main audio segment sample and the reference audio segment are subjected to audio fusion processing to obtain a noisy audio segment sample.

[0213] For example, a reference audio segment with the same duration d3 is randomly selected from the reference audio different from the original audio 901, a volume gain of 20dB is applied, and then a weighted sum is calculated with the main audio segment sample a′4, where the weights are random, the sum of the weights is 1, and the weight of the reference audio segment is not less than 0.5, to obtain the noisy audio segment sample a′3.

[0214] An audio sequence sample can be obtained based on the main audio segment sample a′4, the preceding audio segment sample a′1, the following audio segment sample a′2, and the noise-added audio segment sample a′3. Repeat this step to obtain 10 audio sequence samples for each original audio segment. Each audio sequence sample is a four-tuple of audio segments (a′1, a′2, a′3, a′4).

[0215] The audio sequence samples can be subjected to feature extraction to obtain acoustic feature vector sequence samples, which are then input into an audio word segmenter. Discrete coding vector sequence training samples are then extracted from the acoustic feature vector sequence samples by the audio word segmenter. For example, the audio word segmenter extracts the main text audio segment sample a′4, the preceding audio segment sample a′1, the following audio segment sample a′2, and the noised audio segment sample a′3 from the acoustic feature vector sequence samples to obtain a fourth discrete coding vector sample c'4, a second discrete coding vector sample c'1, a third discrete coding vector sample c'2, and a first discrete coding vector sample c'3. These samples are then concatenated to obtain discrete coding vector sequence training samples 903. The specific operational procedures are as described above and will not be elaborated upon here.

[0216] In this way, the acquisition of training samples for the autoregressive module does not require manual labeling. Audio sequence samples can be generated based on self-supervised learning, and then discrete coding vector sequence samples for model training can be obtained, saving labeling costs and improving the generalization ability of the autoregressive module.

[0217] In some embodiments, obtaining an audio sequence sample based on the main text audio segment sample, the previous audio segment sample, the following audio segment sample, and the noise-added audio segment sample may include:

[0218] Perform speech detection on the text audio clip samples;

[0219] In the case where the main text audio segment sample includes speech information, a first audio sequence sample is obtained according to the main text audio segment sample, the preceding audio segment sample, the following audio segment sample, and the noise-added audio segment sample; the first audio sequence sample is used to train a first autoregressive module;

[0220] When the main audio segment sample does not include speech information, a second audio sequence sample is obtained based on the main audio segment sample, the previous audio segment sample, the following audio segment sample and the noise-added audio segment sample; the second audio sequence sample is used to train the second autoregressive module.

[0221] In some embodiments of the present application, speech detection can also be performed on the main audio segment samples. For example, a Voice Activity Detection (VAD) algorithm can be applied to the main audio segment sample a′4 to determine whether the main audio segment sample a′4 contains speech information, that is, whether there is a human voice. Based on the VAD results, the constructed audio sequence samples are divided into a first audio sequence sample containing a human voice and a second audio sequence sample without a human voice.

[0222] The first audio sequence samples are used to train a first autoregressive module, which is suitable for audio restoration in a scenario where the speech information in the first audio segment is retained. The second audio sequence samples are used to train a second autoregressive module, which is suitable for audio restoration in a scenario where the speech information in the first audio segment is removed.

[0223] In this way, a voice activity detection label can be added to the audio sequence sample to indicate whether the main audio clip sample contains voice information. Then, based on this label, the training task can be divided into two independent tasks, and two independent autoregressive modules can be trained. During inference, the corresponding autoregressive module is called according to the user's voice processing intention. While reducing the training difficulty, the repaired second audio clip that meets the user's voice processing intention can be obtained.

[0224] In some embodiments, when the method is executed by a server, after inputting the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector into a large audio model, performing autoregressive prediction using the large audio model to obtain a predicted discrete coding vector, and decoding the predicted discrete coding vector to obtain a restored second audio segment, the method further includes:

[0225] The repaired second audio segment is sent to the client, so that the client replaces the first audio segment in the audio with the second audio segment, where the first audio segment is the audio segment intercepted from the audio.

[0226] In some embodiments of the present application, after obtaining the repaired second audio segment, the server may send the repaired second audio segment to the client, so that the client replaces the first audio segment intercepted from the audio with the repaired second audio segment.

[0227] Two examples of audio restoration scenarios using the autoregressive module are given below.

[0228] Suppose a user has a 2-minute original audio clip 901, selects the audio segment from 1:55 to 2:05 as the first audio segment, and clicks the "Keep Vocals" control. Suppose the first audio segment contains vocals a+b, and the preceding and following audio segments 5 seconds before and after the first audio segment only contain vocals a. In this case, the client sends the first audio segment, the preceding audio segment, and the following audio segment to the server. The server extracts discrete coding vectors and concatenates them in the order of "previous audio segment, following audio segment, and first audio segment." Using a first autoregressive module, the server infers the discrete coding vector corresponding to the audio segment containing only vocals a. This is decoded to obtain the repaired second audio segment and sent to the client. The client uses this repaired second audio segment to replace the portion from 1:55 to 2:05, resulting in the repaired audio.

[0229] Suppose a user has a one-minute original audio clip, selects the audio segment from 0:00 to 0:08 as the first audio segment, and clicks the "Remove Vocals" control. Suppose the first audio segment contains burst noise, while the following audio segment 5 seconds after the first audio segment contains only ambient sound. In this case, the client sends the first and following audio segments to the server. The server extracts the discrete encoding vectors and concatenates them in the order of "previous audio segment, following audio segment, and first audio segment." At this point, the preceding audio segment has a length of zero, but still has labels indicating its start and end points. Using the second autoregressive module, the server infers the discrete encoding vector corresponding to the audio segment containing only ambient sound. This is decoded to obtain the repaired second audio segment and sent to the client. The client then uses this repaired second audio segment to replace the segment from 0:00 to 0:08, resulting in the repaired audio.

[0230] In this way, the repaired second audio segment is sent to the client, so that the client replaces the first audio segment with the repaired second audio segment, thereby obtaining audio with more natural context connection and better repair effect.

[0231] In some embodiments, before extracting features from the first audio segment to be repaired, the preceding audio segment of the first audio segment, and the following audio segment to obtain a first acoustic feature vector for the first audio segment, a second acoustic feature vector for the preceding audio segment, and a third acoustic feature vector for the following audio segment, the method may further include:

[0232] Obtaining a first audio segment to be restored, a preceding audio segment of the first audio segment, and a following audio segment uploaded by a client;

[0233] The first audio segment, the preceding audio segment, and the following audio segment are audio segments obtained by intercepting the audio in response to the user's first input of the audio.

[0234] In some embodiments of the present application, Figure 10 As shown, the client can display a preview interface 1000, which can display the video 1001 being previewed and a playback progress bar 1002 corresponding to the video 1001. The preview interface 1000 can also include an "edit" control 1003, and the user can enter the editing settings interface by clicking the "edit" control 1003.

[0235] like Figure 11 As shown, the editing setting interface 1100 may include an editing category menu, which may include an "audio" control 1101. The editing category menu may also include a "video" control, a "crop" control, and a "filter" control, etc. A user may edit the audio in a video file by clicking the "audio" control 1101.

[0236] The editing settings interface 1100 may also include function controls corresponding to the audio, such as a "mute" control, a "sound effect" control, a "noise reduction" control, and a "repair" control 1102. Users can enter the audio editing interface by clicking the "repair" control 1102.

[0237] like Figure 12 As shown, the audio editing interface 1200 may include an audio preview progress control 1201. The user may capture a first audio segment 1202 by dragging the start and end position indicators in the audio preview progress control 1201.

[0238] It is understandable that when the audio repair function is used for the first time, the audio editing interface 1200 may pop up a prompt message to introduce the purpose and operation steps of the audio repair function. During subsequent use, the prompt message can still be reviewed through the relevant controls in the audio editing interface 1200.

[0239] The client can respond to the user's first input of the audio, such as dragging the start and end position indicators in the audio preview progress control 1201, selecting the input of the first audio segment 1202, intercepting the audio, obtaining the first audio segment 1202, and the previous audio segment and the following audio segment within a preset time period before and after the first audio segment 1202, and uploading them to the server.

[0240] The server can obtain the first audio segment, the previous audio segment, and the next audio segment uploaded by the client to carry out the subsequent audio repair process.

[0241] In this way, based on the first audio clip selected by the user on the client, the first audio clip, the previous audio clip and the following audio clip can be uploaded to the server. The server can then automatically repair the first audio clip through the autoregressive module, which simplifies the user's operation process and improves the convenience of audio repair.

[0242] In some embodiments, the method may further include:

[0243] Get the voice processing setting information uploaded by the client;

[0244] identifying a vocal processing intent for the first audio segment based on the vocal processing setting information;

[0245] Among them, the human voice processing intention is used to select the autoregressive module in the audio model. When the human voice processing intention is to retain the voice information in the first audio segment, the predicted discrete coding vector is obtained through the first autoregressive module; when the human voice processing intention is to remove the voice information in the first audio segment, the predicted discrete coding vector is obtained through the second autoregressive module; the first autoregressive module is trained based on the first audio sequence samples, and the second autoregressive module is trained based on the second audio sequence samples. The first audio sequence samples and the second audio sequence samples both include text audio segment samples, previous audio segment samples, following audio segment samples and noisy audio segment samples after the text audio segment samples are denoised; the text audio segment samples in the first audio sequence samples include voice information, and the text audio segment samples in the second audio sequence samples do not include voice information.

[0246] In some embodiments of the present application, the server can also obtain the voice processing setting information uploaded by the client, wherein the voice processing setting information can be obtained by the client based on the voice or text information input by the user, or can be obtained by the client based on the user's historical repair records, or can be obtained by the client based on the user's selection of the corresponding voice setting control.

[0247] For example, Figure 12 As shown, the audio editing interface 1200 may include a "Keep Vocals" control 1203 and a "Remove Vocals" control 1204. User input to either control may be received to determine vocal processing settings. In some examples, if no user input is received, the vocal processing settings may be set to "Keep Vocals" control 1203.

[0248] The user's selections of the "keep vocals" control 1203 and the "remove vocals" control 1204 can be used as vocal processing setting information and uploaded to the server together with the first audio segment, the previous audio segment, and the next audio segment.

[0249] The server can obtain the voice processing setting information uploaded by the client, identify whether the voice processing intention of the first audio clip is to retain the voice information in the first audio clip or to remove the voice information in the first audio clip, and then call the corresponding autoregressive module to infer and obtain the repaired second audio clip.

[0250] In this way, the vocal processing setting information uploaded by the client can be obtained to identify the vocal processing intention of the first audio clip, and then the autoregressive module corresponding to the vocal processing intention can be called for reasoning to obtain the repaired second audio clip that can meet user needs.

[0251] In some embodiments, when the method is executed by a client, after inputting the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector into a large audio model, performing autoregressive prediction using the large audio model to obtain a predicted discrete coding vector, and decoding the predicted discrete coding vector to obtain a restored second audio segment, the method further includes:

[0252] Replaces the first audio clip in the audio with the second audio clip.

[0253] In some embodiments of the present application, if the audio repair method is executed directly on the client, as mentioned above, the client can directly replace the first audio segment intercepted from the audio with the second audio segment to obtain audio with more natural context connection and better repair effect.

[0254] In some embodiments, before extracting features from the first audio segment to be repaired, the preceding audio segment of the first audio segment, and the following audio segment to obtain a first acoustic feature vector of the first audio segment, a second acoustic feature vector of the preceding audio segment, and a third acoustic feature vector of the following audio segment, the method further includes:

[0255] Displays the audio editing interface, which includes an audio preview progress control;

[0256] In response to a second input to the audio preview progress control, determining a segment captured by the second input on the audio preview progress control as a first audio segment to be repaired;

[0257] The preceding audio segment and the following audio segment of the first audio segment are extracted from the audio.

[0258] In some embodiments of the present application, as mentioned above, Figure 12 As shown, the audio editing interface 1200 may include an audio preview progress control 1201. A second input from the user to the audio preview progress control 1201 may be received. For example, the user may drag the start and end position indication marks in the audio preview progress control 1201.

[0259] In response to the second input, the segment captured by the second input on the audio preview progress control 1201 is determined as the first audio segment to be repaired. Based on the first audio segment, the preceding audio segment and the following audio segment adjacent to the first audio segment are captured from the audio.

[0260] In this way, the user only needs to make relevant operations on the audio preview progress control in the audio editing interface to capture the first audio clip, the previous audio clip and the following audio clip to achieve automatic audio repair, which simplifies the user's operation process and improves the convenience of audio repair.

[0261] In some embodiments, the editing interface further includes a vocal setting control, and the method further includes:

[0262] acquiring vocal processing setting information in response to a third input to the vocal setting control;

[0263] identifying a vocal processing intent for the first audio segment based on the vocal processing setting information;

[0264] Among them, the human voice processing intention is used to select the autoregressive module in the audio model. When the human voice processing intention is to retain the voice information in the first audio segment, the predicted discrete coding vector is obtained through the first autoregressive module; when the human voice processing intention is to remove the voice information in the first audio segment, the predicted discrete coding vector is obtained through the second autoregressive module; the first autoregressive module is trained based on the first audio sequence samples, and the second autoregressive module is trained based on the second audio sequence samples. The first audio sequence samples and the second audio sequence samples both include text audio segment samples, previous audio segment samples, following audio segment samples and noisy audio segment samples after the text audio segment samples are denoised; the text audio segment samples in the first audio sequence samples include voice information, and the text audio segment samples in the second audio sequence samples do not include voice information.

[0265] In some embodiments of the present application, as mentioned above, Figure 12 As shown, the audio editing interface 1200 may also include vocal setting controls, which may include a "keep vocals" control 1203 and a "remove vocals" control 1204. User input to either control may be received to determine vocal processing setting information. In some examples, if no user input is received, the vocal processing setting information may be defaulted to selecting the "keep vocals" control 1203.

[0266] Based on the vocal processing setting information, it can be identified whether the vocal processing intention of the first audio segment is to retain the voice information in the first audio segment or to remove the voice information in the first audio segment, so as to call the corresponding autoregressive module for inference to obtain the repaired second audio segment.

[0267] In this way, based on the user's relevant operations on the voice setting controls, the voice processing intention of the first audio clip can be identified, and then the autoregressive module corresponding to the voice processing intention can be called for reasoning to obtain a repaired second audio clip that can meet the user's needs.

[0268] The audio repair method provided in the embodiment of the present application can be executed by an audio repair device. In the embodiment of the present application, the audio repair method performed by the audio repair device is taken as an example to illustrate the audio repair device provided in the embodiment of the present application.

[0269] like Figure 13 As shown, the audio repair apparatus 1300 may include:

[0270] A feature extraction module 1301 is configured to perform feature extraction on a first audio segment to be repaired, a preceding audio segment of the first audio segment, and a following audio segment of the first audio segment, to obtain a first acoustic feature vector of the first audio segment, a second acoustic feature vector of the preceding audio segment, and a third acoustic feature vector of the following audio segment.

[0271] Processing module 1302 is used to input the first acoustic eigenvector, the second acoustic eigenvector and the third acoustic eigenvector into the audio large model, perform autoregressive prediction through the audio large model to obtain a predicted discrete coding vector, and decode the predicted discrete coding vector to obtain a repaired second audio segment.

[0272] In this way, when repairing the first audio clip, the previous audio clip and the following audio clip of the first audio clip can also be considered to extract the acoustic feature vectors of the three, so that the acoustic feature vectors of the three are all used as inputs of the audio big model, so that when the audio big model performs autoregressive prediction, it can combine the context information, so that the repaired second audio clip obtained by the audio big model is more naturally connected with the context and the repair effect is better.

[0273] In some embodiments, the audio large model includes an audio word segmenter and an autoregressive module; the processing module 1302 is further configured to:

[0274] Inputting the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector into an audio word segmenter, and extracting a first discrete coding vector of the first acoustic feature vector, a second discrete coding vector of the second acoustic feature vector, and a third discrete coding vector of the third acoustic feature vector through the audio word segmenter;

[0275] Inputting the discrete coding vector sequence into the autoregressive module, performing deep inference on the discrete coding vector sequence through the autoregressive module, and outputting a predicted discrete coding vector; the discrete coding vector sequence includes a first discrete coding vector, a second discrete coding vector, and a third discrete coding vector;

[0276] The predicted discrete coding vector is input into the audio word segmenter, and the predicted discrete coding vector is decoded by the audio word segmenter to obtain a repaired second audio segment.

[0277] In this way, when repairing the first audio clip, the previous audio clip and the following audio clip of the first audio clip will be considered, and the acoustic feature vectors of the three will be extracted. Based on the acoustic feature vectors of the three, the discrete coding vectors of the three are obtained to form a discrete coding vector sequence. The discrete coding vector sequence is used as the input of the autoregressive module so that the autoregressive module can combine the context information, so that the repaired second audio clip obtained by the autoregressive module is more naturally connected with the context and the repair effect is better.

[0278] In some embodiments, the processing module 1302 may also be configured to:

[0279] Before inputting the discrete coding vector sequence into the autoregressive module, performing deep inference on the discrete coding vector sequence through the autoregressive module, and outputting a predicted discrete coding vector, identifying the vocal processing intent of the first audio segment;

[0280] In a case where the vocal processing is intended to retain speech information in the first audio segment, the discrete code vector sequence is input into a first autoregressive module, deep inference is performed on the discrete code vector sequence by the first autoregressive module, and a predicted discrete code vector is output;

[0281] In a case where the vocal processing is intended to remove speech information from the first audio segment, the discrete code vector sequence is input into a second autoregressive module, deep inference is performed on the discrete code vector sequence by the second autoregressive module, and a predicted discrete code vector is output;

[0282] Among them, the first autoregressive module is trained based on the first audio sequence samples, and the second autoregressive module is trained based on the second audio sequence samples. The first audio sequence samples and the second audio sequence samples both include main audio segment samples, previous audio segment samples, following audio segment samples, and noisy audio segment samples after the main audio segment samples are denoised; the main audio segment samples in the first audio sequence samples include voice information, and the main audio segment samples in the second audio sequence samples do not include voice information.

[0283] In this way, the user's vocal processing intention can be identified when repairing the audio, and the corresponding autoregressive module can be selected to perform audio repair according to different vocal processing intentions, so as to obtain a repaired second audio segment that meets the vocal processing intention.

[0284] In some embodiments, the audio word segmenter includes a generator, the generator including an encoder and a vector quantization module;

[0285] The processing module 1302 is further configured to:

[0286] Inputting the first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector into an encoder, downsampling the first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector through the encoder to obtain a first encoding vector of the first acoustic eigenvector, a second encoding vector of the second acoustic eigenvector, and a third encoding vector of the third acoustic eigenvector;

[0287] Inputting the first coding vector, the second coding vector, and the third coding vector into a vector quantization module, the vector quantization module splitting the first coding vector into N first low-dimensional vectors, splitting the second coding vector into N second low-dimensional vectors, and splitting the third coding vector into N third low-dimensional vectors, where N is an integer greater than 1;

[0288] The vector quantization module calculates a first cosine similarity between each first low-dimensional vector and the plurality of codebook vectors, a second cosine similarity between each second low-dimensional vector and the plurality of codebook vectors, and a third cosine similarity between each third low-dimensional vector and the plurality of codebook vectors;

[0289] The vector quantization module determines a first discrete code vector according to the first cosine similarity, determines a second discrete code vector according to the second cosine similarity, and determines a third discrete code vector according to the third cosine similarity.

[0290] In this way, the first acoustic feature vector, the second acoustic feature vector and the third acoustic feature vector can be discretized through the encoder and VQ module in the audio segmenter to obtain the first encoding vector, the second encoding vector and the third encoding vector as the input of the subsequent autoregressive module, which can reduce the computational complexity of data processing in the autoregressive module and improve the model processing efficiency.

[0291] In some embodiments, the audio word segmenter further comprises a vocoder, and the generator further comprises a decoder;

[0292] The processing module 1302 is further configured to:

[0293] Inputting the predicted discrete code vector into a decoder, and decoding the predicted discrete code vector through the decoder to obtain a fourth acoustic feature vector;

[0294] The fourth acoustic feature vector is input into a vocoder, and the vocoder converts the fourth acoustic feature vector into a second audio segment.

[0295] In this way, the predicted discrete coding vector can be decoded by the decoder in the audio segmenter to obtain a fourth acoustic feature vector, and the fourth acoustic feature vector can be transformed by the vocoder to obtain a second audio segment, so that the first audio segment can be replaced based on the second audio segment to obtain audio with a natural repair effect.

[0296] In some embodiments, the audio word segmenter further includes a discriminator; the audio repair apparatus 1300 further includes an audio word segmenter training module for:

[0297] Obtaining a first acoustic eigenvector sample;

[0298] Inputting the first acoustic feature vector sample into the generator, the encoder downsamples the first acoustic feature vector sample to obtain an encoded vector sample;

[0299] The vector quantization module splits the coded vector sample into N low-dimensional vector samples, calculates the cosine similarity between each low-dimensional vector sample and multiple codebook vectors, and determines N codebook vector samples based on the cosine similarity;

[0300] Inputting N codebook vector samples into a decoder, upsampling the N codebook vector samples through the decoder to obtain a second acoustic feature vector sample;

[0301] Calculating a first loss value based on a pre-constructed first loss function and a second acoustic feature vector sample;

[0302] Inputting the first acoustic feature vector sample or the second acoustic feature vector sample into the discriminator to obtain a probability value that the input sample is the first acoustic feature vector sample;

[0303] Calculating a second loss value based on a pre-constructed second loss function and a probability value;

[0304] According to the sum of the first loss value and the second loss value, the parameters of the generator and the parameters of the discriminator are updated to obtain the trained audio segmenter.

[0305] In this way, the audio segmenter can be trained based on the first acoustic feature vector sample and combined with the loss function of the generator and the loss function of the discriminator to obtain an audio segmenter with higher accuracy, so that a more accurate discrete coding vector of the audio segment can be extracted as the input of the autoregressive module, thereby improving the accuracy of the output result of the autoregressive module, and thus a more accurate repaired second audio segment can be obtained.

[0306] In some embodiments, the autoregressive module includes an acoustic embedding submodule, a sequence modeling submodule, and an encoding prediction submodule;

[0307] The processing module 1302 is further configured to:

[0308] Input the discrete encoding vector sequence into the acoustic embedding submodule to obtain the latent variable;

[0309] Input the latent variables into the sequence modeling submodule to obtain the latent variable sequence;

[0310] The last frame in the latent variable sequence is input into the encoding prediction submodule, multiple probability distribution vectors are generated by the encoding prediction submodule, and core sampling is performed on each probability distribution vector to obtain multiple candidate discrete encoding vectors;

[0311] Among multiple candidate discrete coding vectors, a candidate discrete coding vector corresponding to a reference terminator is determined as a predicted discrete coding vector, and the predicted discrete coding vector is output.

[0312] In this way, the acoustic embedding submodule, sequence modeling submodule, and encoding prediction submodule in the autoregressive module perform deep inference on the input discrete encoding vector sequence and predict the predicted discrete encoding vector corresponding to the repaired second audio segment, which can then be decoded to obtain the repaired second audio segment. This entire process achieves automated audio restoration without requiring the user to understand any technical terminology, making it easy to use.

[0313] In some embodiments, the audio repair apparatus 1300 may further include a model training module for:

[0314] Obtaining discrete coding vector sequence training samples; wherein each discrete coding vector sequence training sample is obtained by splicing the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample, and the fourth discrete coding vector sample, and the splicing position of the fourth discrete coding vector sample is located at the end of the discrete coding vector sequence training sample, and corresponding start symbols are added to the heads of the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample, and the fourth discrete coding vector sample, respectively, and corresponding terminators are added to the tails of the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample, and the fourth discrete coding vector sample, respectively;

[0315] Perform autoregressive training on the model based on the discrete coding vector sequence training samples to obtain the trained initial module;

[0316] Obtain discrete code vector sequence test samples; wherein each discrete code vector sequence test sample is obtained by splicing the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample; corresponding start symbols are added to the heads of the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample; corresponding terminators are added to the tails of the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample; and the start symbol corresponding to the fourth discrete code vector sample is added after the terminator corresponding to the fifth discrete code vector sample; the fifth discrete code vector sample is the discrete code vector sample whose splicing position is located at the tail of the discrete code vector sequence test sample;

[0317] Input the discrete coding vector sequence test sample into the initial module and output the predicted discrete coding vector sample;

[0318] According to the difference value between the predicted discrete coding vector sample and the fourth discrete coding vector sample, the initial module is optimized to obtain an autoregressive module.

[0319] In this way, the autoregressive module can be pre-trained to obtain an initial module with reasoning capabilities, and then the initial module can be optimized to obtain an optimized autoregressive module. This can ensure the accuracy of the autoregressive module. The repaired second audio segment obtained based on the autoregressive module reasoning can be closer to the original audio, and the connection with the context is more natural, thereby improving the effect of audio repair.

[0320] In some embodiments, the model training module may also be used to:

[0321] Input the discrete encoding vector sequence test sample into the acoustic embedding submodule in the initial module to obtain the latent variable sample;

[0322] Input the latent variable sample into the sequence modeling submodule in the initial module to obtain the latent variable sequence sample;

[0323] The last frame in the latent variable sequence sample is input into the coding prediction submodule in the initial module, multiple probability distribution vectors are generated by the coding prediction submodule, the discrete coding vector corresponding to the maximum value of the probability distribution vector is determined as the predicted discrete coding vector sample, and the predicted discrete coding vector sample is output.

[0324] In this way, Gumbel-Softmax sampling is used to ensure that the sampling process can perform gradient backpropagation, so that the parameters of the initial module can be updated, so as to optimize the model of the initial module and obtain an autoregressive module with better accuracy and better effect.

[0325] In some embodiments, the model training module may also be used to:

[0326] splicing the first discrete coding vector sample, the second discrete coding vector sample, and the third discrete coding vector sample to obtain a first discrete coding vector sequence sample;

[0327] splicing the fourth discrete coding vector sample, the second discrete coding vector sample, and the third discrete coding vector sample to obtain a second discrete coding vector sequence sample;

[0328] splicing the predicted discrete coding vector sample, the second discrete coding vector sample, and the third discrete coding vector sample to obtain a third discrete coding vector sequence sample;

[0329] Inputting the first discrete code vector sequence sample, the second discrete code vector sequence sample, and the third discrete code vector sequence sample into a decoder in the audio word segmentor, obtaining a first reconstructed acoustic feature vector sample of the first discrete code vector sequence sample, a second reconstructed acoustic feature vector sample of the second discrete code vector sequence sample, and a third reconstructed acoustic feature vector sample of the third discrete code vector sequence sample;

[0330] Inputting the first reconstructed acoustic feature vector sample, the second reconstructed acoustic feature vector sample, and the third reconstructed acoustic feature vector sample into the discriminator in the audio word segmenter, obtaining a first discriminant latent variable of the first reconstructed acoustic feature vector sample, a second discriminant latent variable of the second reconstructed acoustic feature vector sample, and a third discriminant latent variable of the third reconstructed acoustic feature vector sample;

[0331] The first discriminant latent variable is used as a positive sample, the second discriminant latent variable is used as a negative sample, and the third discriminant latent variable is used as an anchor sample to calculate the triplet loss value;

[0332] According to the triplet loss value, the model parameters of the initial module are updated to obtain the autoregressive module, wherein the triplet loss value corresponding to the model parameters of the autoregressive module is the minimum value.

[0333] This approach uses the similarity of the discriminator's output as a basis for model optimization, optimizing the naturalness of the autoregressive module's contextual cohesion. This ensures that the restored second audio clip, derived from the autoregressive module's inference, seamlessly connects with the context. Furthermore, the reused audio word segmenter discriminates the naturalness of the autoregressive module and optimizes the model through comparative learning. This optimization eliminates the need for a new model structure, reducing training costs.

[0334] In some embodiments, the model training module may also be used to:

[0335] Based on each of the M original audios, K audio segment combinations are randomly generated, where the audio segment combinations include a main audio segment sample, a preceding audio segment sample of the main audio segment sample, and a following audio segment sample; M and K are positive integers;

[0336] Obtain a reference audio segment from the reference audio; the reference audio is any original audio from the M original audios except the original audio;

[0337] Performing audio fusion processing on the main text audio clip sample and the reference audio clip to obtain a noise-added audio clip sample;

[0338] Obtaining an audio sequence sample according to the main text audio segment sample, the preceding audio segment sample, the following audio segment sample, and the noise-added audio segment sample;

[0339] Perform feature extraction on the audio sequence samples to obtain acoustic feature vector sequence samples;

[0340] The acoustic feature vector sequence samples are input into the audio word segmenter, and the discrete coding vector sequence training samples are extracted from the acoustic feature vector sequence samples by the audio word segmenter.

[0341] In this way, the acquisition of training samples for the autoregressive module does not require manual labeling. Audio sequence samples can be generated based on self-supervised learning, and then discrete coding vector sequence samples for model training can be obtained, saving labeling costs and improving the generalization ability of the autoregressive module.

[0342] In some embodiments, the model training module may also be used to:

[0343] Perform speech detection on the text audio clip samples;

[0344] In the case where the main text audio segment sample includes speech information, a first audio sequence sample is obtained according to the main text audio segment sample, the preceding audio segment sample, the following audio segment sample, and the noise-added audio segment sample; the first audio sequence sample is used to train a first autoregressive module;

[0345] When the main audio segment sample does not include speech information, a second audio sequence sample is obtained based on the main audio segment sample, the previous audio segment sample, the following audio segment sample and the noise-added audio segment sample; the second audio sequence sample is used to train the second autoregressive module.

[0346] In this way, a voice activity detection label can be added to the audio sequence sample to indicate whether the main audio clip sample contains voice information. Then, based on this label, the training task can be divided into two independent tasks, and two independent autoregressive modules can be trained. During inference, the corresponding autoregressive module is called according to the user's voice processing intention. While reducing the training difficulty, the repaired second audio clip that meets the user's voice processing intention can be obtained.

[0347] In some embodiments, the audio repair apparatus 1300 may further include:

[0348] The transmission module is used to send the repaired second audio segment to the client, so that the client replaces the first audio segment in the audio with the repaired second audio segment, where the first audio segment is an audio segment intercepted from the audio.

[0349] In this way, the repaired second audio segment is sent to the client, so that the client replaces the first audio segment with the repaired second audio segment, thereby obtaining audio with more natural context connection and better repair effect.

[0350] In some embodiments, the audio repair apparatus 1300 may further include:

[0351] An acquisition module, configured to acquire a first audio segment to be restored, a preceding audio segment of the first audio segment, and a following audio segment of the first audio segment uploaded by a client;

[0352] The first audio segment, the preceding audio segment, and the following audio segment are audio segments obtained by intercepting the audio in response to the user's first input of the audio.

[0353] In this way, based on the first audio clip selected by the user on the client, the first audio clip, the previous audio clip and the following audio clip can be uploaded to the server. The server can then automatically repair the first audio clip through the autoregressive module, which simplifies the user's operation process and improves the convenience of audio repair.

[0354] In some embodiments, the acquisition module may also be used to:

[0355] Get the voice processing setting information uploaded by the client;

[0356] The processing module 1302 may also be used to:

[0357] identifying a vocal processing intent for the first audio segment based on the vocal processing setting information;

[0358] Among them, the human voice processing intention is used to select the autoregressive module in the audio model. When the human voice processing intention is to retain the voice information in the first audio segment, the predicted discrete coding vector is obtained through the first autoregressive module; when the human voice processing intention is to remove the voice information in the first audio segment, the predicted discrete coding vector is obtained through the second autoregressive module; the first autoregressive module is trained based on the first audio sequence samples, and the second autoregressive module is trained based on the second audio sequence samples. The first audio sequence samples and the second audio sequence samples both include text audio segment samples, previous audio segment samples, following audio segment samples and noisy audio segment samples after the text audio segment samples are denoised; the text audio segment samples in the first audio sequence samples include voice information, and the text audio segment samples in the second audio sequence samples do not include voice information.

[0359] In this way, the vocal processing setting information uploaded by the client can be obtained to identify the vocal processing intention of the first audio clip, and then the autoregressive module corresponding to the vocal processing intention can be called for reasoning to obtain the repaired second audio clip that can meet user needs.

[0360] In some embodiments, the processing module 1302 is further configured to:

[0361] Replaces the first audio clip in the audio with the second audio clip.

[0362] In this way, the client can directly replace the first audio segment intercepted from the audio with the second audio segment to obtain audio with more natural context connection and better restoration effect.

[0363] In some embodiments, the audio repair apparatus 1300 may further include:

[0364] A display module, used to display an audio editing interface, the editing interface including an audio preview progress control;

[0365] The processing module 1302 may also be used to:

[0366] In response to a second input to the audio preview progress control, determining a segment captured by the second input on the audio preview progress control as a first audio segment to be repaired;

[0367] The preceding audio segment and the following audio segment of the first audio segment are extracted from the audio.

[0368] In this way, the user only needs to make relevant operations on the audio preview progress control in the audio editing interface to capture the first audio clip, the previous audio clip and the following audio clip to achieve automatic audio repair, which simplifies the user's operation process and improves the convenience of audio repair.

[0369] In some embodiments, the editing interface further includes vocal setting controls, and the audio repair apparatus 1300 may further include:

[0370] an acquisition module, configured to acquire vocal processing setting information in response to a third input to the vocal setting control;

[0371] The processing module can also be used to:

[0372] identifying a vocal processing intent for the first audio segment based on the vocal processing setting information;

[0373] Among them, the human voice processing intention is used to select the autoregressive module in the audio model. When the human voice processing intention is to retain the voice information in the first audio segment, the predicted discrete coding vector is obtained through the first autoregressive module; when the human voice processing intention is to remove the voice information in the first audio segment, the predicted discrete coding vector is obtained through the second autoregressive module; the first autoregressive module is trained based on the first audio sequence samples, and the second autoregressive module is trained based on the second audio sequence samples. The first audio sequence samples and the second audio sequence samples both include text audio segment samples, previous audio segment samples, following audio segment samples and noisy audio segment samples after the text audio segment samples are denoised; the text audio segment samples in the first audio sequence samples include voice information, and the text audio segment samples in the second audio sequence samples do not include voice information.

[0374] In this way, based on the user's relevant operations on the voice setting controls, the voice processing intention of the first audio clip can be identified, and then the autoregressive module corresponding to the voice processing intention can be called for reasoning to obtain a repaired second audio clip that can meet the user's needs.

[0375] The audio repair device in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or chip. The electronic device can be a terminal or other device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., and the embodiment of the present application does not specifically limit it.

[0376] The audio repair device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0377] The audio repair device provided in the embodiment of the present application can achieve Figures 1 to 12 To avoid repetition, the various processes implemented in the method embodiment are not described here.

[0378] Alternatively, as Figure 14 As shown, an embodiment of the present application further provides an electronic device 1400, including a processor 1401 and a memory 1402, wherein the memory 1402 stores a program or instruction that can be run on the processor 1401, and when the program or instruction is executed by the processor 1401, the various steps of the above-mentioned audio repair method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0379] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0380] Figure 15 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application.

[0381] The electronic device 1500 includes but is not limited to components such as a radio frequency unit 1501 , a network module 1502 , an audio output unit 1503 , an input unit 1504 , a sensor 1505 , a display unit 1506 , a user input unit 1507 , an interface unit 1508 , a memory 1509 , and a processor 1510 .

[0382] Those skilled in the art will understand that the electronic device 1500 may also include a power source (such as a battery) to power each component, and the power source may be logically connected to the processor 1510 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system. Figure 15 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be repeated here.

[0383] The processor 1510 may be configured to:

[0384] Performing feature extraction on a first audio segment to be repaired, a preceding audio segment of the first audio segment, and a following audio segment of the first audio segment to obtain a first acoustic feature vector of the first audio segment, a second acoustic feature vector of the preceding audio segment, and a third acoustic feature vector of the following audio segment;

[0385] The first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector are input into the audio large model, autoregressive prediction is performed through the audio large model to obtain a predicted discrete coding vector, and the predicted discrete coding vector is decoded to obtain a repaired second audio segment.

[0386] In this way, when repairing the first audio clip, the previous audio clip and the following audio clip of the first audio clip can also be considered to extract the acoustic feature vectors of the three, so that the acoustic feature vectors of the three are all used as inputs of the audio big model, so that when the audio big model performs autoregressive prediction, it can combine the context information, so that the repaired second audio clip obtained by the audio big model is more naturally connected with the context and the repair effect is better.

[0387] In some embodiments, the audio large model includes an audio word segmenter and an autoregressive module; the processor 1510 may also be configured to:

[0388] Inputting the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector into an audio word segmenter, and extracting a first discrete coding vector of the first acoustic feature vector, a second discrete coding vector of the second acoustic feature vector, and a third discrete coding vector of the third acoustic feature vector through the audio word segmenter;

[0389] Inputting the discrete coding vector sequence into the autoregressive module, performing deep inference on the discrete coding vector sequence through the autoregressive module, and outputting a predicted discrete coding vector; the discrete coding vector sequence includes a first discrete coding vector, a second discrete coding vector, and a third discrete coding vector;

[0390] The predicted discrete coding vector is input into the audio word segmenter, and the predicted discrete coding vector is decoded by the audio word segmenter to obtain a repaired second audio segment.

[0391] In this way, when repairing the first audio clip, the previous audio clip and the following audio clip of the first audio clip will be considered, and the acoustic feature vectors of the three will be extracted. Based on the acoustic feature vectors of the three, the discrete coding vectors of the three are obtained to form a discrete coding vector sequence. The discrete coding vector sequence is used as the input of the autoregressive module so that the autoregressive module can combine the context information, so that the repaired second audio clip obtained by the autoregressive module is more naturally connected with the context and the repair effect is better.

[0392] In some embodiments, the processor 1510 may also be configured to:

[0393] Before inputting the discrete coding vector sequence into the autoregressive module, performing deep inference on the discrete coding vector sequence through the autoregressive module, and outputting a predicted discrete coding vector, identifying the vocal processing intent of the first audio segment;

[0394] In a case where the vocal processing is intended to retain speech information in the first audio segment, the discrete code vector sequence is input into a first autoregressive module, deep inference is performed on the discrete code vector sequence by the first autoregressive module, and a predicted discrete code vector is output;

[0395] In a case where the vocal processing is intended to remove speech information from the first audio segment, the discrete code vector sequence is input into a second autoregressive module, deep inference is performed on the discrete code vector sequence by the second autoregressive module, and a predicted discrete code vector is output;

[0396] Among them, the first autoregressive module is trained based on the first audio sequence samples, and the second autoregressive module is trained based on the second audio sequence samples. The first audio sequence samples and the second audio sequence samples both include main audio segment samples, previous audio segment samples, following audio segment samples, and noisy audio segment samples after the main audio segment samples are denoised; the main audio segment samples in the first audio sequence samples include voice information, and the main audio segment samples in the second audio sequence samples do not include voice information.

[0397] In this way, the user's vocal processing intention can be identified when repairing the audio, and the corresponding autoregressive module can be selected to perform audio repair according to different vocal processing intentions, so as to obtain a repaired second audio segment that meets the vocal processing intention.

[0398] In some embodiments, the audio word segmenter includes a generator, the generator including an encoder and a vector quantization module;

[0399] The processor 1510 may also be configured to:

[0400] Inputting the first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector into an encoder, downsampling the first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector through the encoder to obtain a first encoding vector of the first acoustic eigenvector, a second encoding vector of the second acoustic eigenvector, and a third encoding vector of the third acoustic eigenvector;

[0401] Inputting the first coding vector, the second coding vector, and the third coding vector into a vector quantization module, the vector quantization module splitting the first coding vector into N first low-dimensional vectors, splitting the second coding vector into N second low-dimensional vectors, and splitting the third coding vector into N third low-dimensional vectors, where N is an integer greater than 1;

[0402] The vector quantization module calculates a first cosine similarity between each first low-dimensional vector and the plurality of codebook vectors, a second cosine similarity between each second low-dimensional vector and the plurality of codebook vectors, and a third cosine similarity between each third low-dimensional vector and the plurality of codebook vectors;

[0403] The vector quantization module determines a first discrete code vector according to the first cosine similarity, determines a second discrete code vector according to the second cosine similarity, and determines a third discrete code vector according to the third cosine similarity.

[0404] In this way, the first acoustic feature vector, the second acoustic feature vector and the third acoustic feature vector can be discretized through the encoder and VQ module in the audio segmenter to obtain the first encoding vector, the second encoding vector and the third encoding vector as the input of the subsequent autoregressive module, which can reduce the computational complexity of data processing in the autoregressive module and improve the model processing efficiency.

[0405] In some embodiments, the audio word segmenter further comprises a vocoder, and the generator further comprises a decoder;

[0406] The processor 1510 may also be configured to:

[0407] Inputting the predicted discrete code vector into a decoder, and decoding the predicted discrete code vector through the decoder to obtain a fourth acoustic feature vector;

[0408] The fourth acoustic feature vector is input into a vocoder, and the vocoder converts the fourth acoustic feature vector into a second audio segment.

[0409] In this way, the predicted discrete coding vector can be decoded by the decoder in the audio segmenter to obtain a fourth acoustic feature vector, and the fourth acoustic feature vector can be transformed by the vocoder to obtain a second audio segment, so that the first audio segment can be replaced based on the second audio segment to obtain audio with a natural repair effect.

[0410] In some embodiments, the audio word segmenter further includes a discriminator; the processor 1510 may also be configured to:

[0411] Obtaining a first acoustic eigenvector sample;

[0412] Inputting the first acoustic feature vector sample into the generator, the encoder downsamples the first acoustic feature vector sample to obtain an encoded vector sample;

[0413] The vector quantization module splits the coded vector sample into N low-dimensional vector samples, calculates the cosine similarity between each low-dimensional vector sample and multiple codebook vectors, and determines N codebook vector samples based on the cosine similarity;

[0414] Inputting N codebook vector samples into a decoder, upsampling the N codebook vector samples through the decoder to obtain a second acoustic feature vector sample;

[0415] Calculating a first loss value based on a pre-constructed first loss function and a second acoustic feature vector sample;

[0416] Inputting the first acoustic feature vector sample or the second acoustic feature vector sample into the discriminator to obtain a probability value that the input sample is the first acoustic feature vector sample;

[0417] Calculating a second loss value based on a pre-constructed second loss function and a probability value;

[0418] According to the sum of the first loss value and the second loss value, the parameters of the generator and the parameters of the discriminator are updated to obtain the trained audio segmenter.

[0419] In this way, the audio segmenter can be trained based on the first acoustic feature vector sample and combined with the loss function of the generator and the loss function of the discriminator to obtain an audio segmenter with higher accuracy, so that a more accurate discrete coding vector of the audio segment can be extracted as the input of the autoregressive module, thereby improving the accuracy of the output result of the autoregressive module, and thus a more accurate repaired second audio segment can be obtained.

[0420] In some embodiments, the autoregressive module includes an acoustic embedding submodule, a sequence modeling submodule, and an encoding prediction submodule;

[0421] The processor 1510 may also be configured to:

[0422] Input the discrete encoding vector sequence into the acoustic embedding submodule to obtain the latent variable;

[0423] Input the latent variables into the sequence modeling submodule to obtain the latent variable sequence;

[0424] The last frame in the latent variable sequence is input into the encoding prediction submodule, multiple probability distribution vectors are generated by the encoding prediction submodule, and core sampling is performed on each probability distribution vector to obtain multiple candidate discrete encoding vectors;

[0425] Among multiple candidate discrete coding vectors, a candidate discrete coding vector corresponding to a reference terminator is determined as a predicted discrete coding vector, and the predicted discrete coding vector is output.

[0426] In this way, the acoustic embedding submodule, sequence modeling submodule, and encoding prediction submodule in the autoregressive module perform deep inference on the input discrete encoding vector sequence and predict the predicted discrete encoding vector corresponding to the repaired second audio segment, which can then be decoded to obtain the repaired second audio segment. This entire process achieves automated audio restoration without requiring the user to understand any technical terminology, making it easy to use.

[0427] In some embodiments, the processor 1510 may also be configured to:

[0428] Obtaining discrete coding vector sequence training samples; wherein each discrete coding vector sequence training sample is obtained by splicing the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample, and the fourth discrete coding vector sample, and the splicing position of the fourth discrete coding vector sample is located at the end of the discrete coding vector sequence training sample, and corresponding start symbols are added to the heads of the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample, and the fourth discrete coding vector sample, respectively, and corresponding terminators are added to the tails of the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample, and the fourth discrete coding vector sample, respectively;

[0429] Perform autoregressive training on the model based on the discrete coding vector sequence training samples to obtain the trained initial module;

[0430] Obtain discrete code vector sequence test samples; wherein each discrete code vector sequence test sample is obtained by splicing the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample; corresponding start symbols are added to the heads of the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample; corresponding terminators are added to the tails of the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample; and the start symbol corresponding to the fourth discrete code vector sample is added after the terminator corresponding to the fifth discrete code vector sample; the fifth discrete code vector sample is the discrete code vector sample whose splicing position is located at the tail of the discrete code vector sequence test sample;

[0431] Input the discrete coding vector sequence test sample into the initial module and output the predicted discrete coding vector sample;

[0432] According to the difference value between the predicted discrete coding vector sample and the fourth discrete coding vector sample, the initial module is optimized to obtain an autoregressive module.

[0433] In this way, the autoregressive module can be pre-trained to obtain an initial module with reasoning capabilities, and then the initial module can be optimized to obtain an optimized autoregressive module. This can ensure the accuracy of the autoregressive module. The repaired second audio segment obtained based on the autoregressive module reasoning can be closer to the original audio, and the connection with the context is more natural, thereby improving the effect of audio repair.

[0434] In some embodiments, the processor 1510 may also be configured to:

[0435] Input the discrete encoding vector sequence test sample into the acoustic embedding submodule in the initial module to obtain the latent variable sample;

[0436] Input the latent variable sample into the sequence modeling submodule in the initial module to obtain the latent variable sequence sample;

[0437] The last frame in the latent variable sequence sample is input into the coding prediction submodule in the initial module, multiple probability distribution vectors are generated by the coding prediction submodule, the discrete coding vector corresponding to the maximum value of the probability distribution vector is determined as the predicted discrete coding vector sample, and the predicted discrete coding vector sample is output.

[0438] In this way, Gumbel-Softmax sampling is used to ensure that the sampling process can perform gradient backpropagation, so that the parameters of the initial module can be updated, so as to optimize the model of the initial module and obtain an autoregressive module with better accuracy and better effect.

[0439] In some embodiments, the processor 1510 may also be configured to:

[0440] splicing the first discrete coding vector sample, the second discrete coding vector sample, and the third discrete coding vector sample to obtain a first discrete coding vector sequence sample;

[0441] splicing the fourth discrete coding vector sample, the second discrete coding vector sample, and the third discrete coding vector sample to obtain a second discrete coding vector sequence sample;

[0442] splicing the predicted discrete coding vector sample, the second discrete coding vector sample, and the third discrete coding vector sample to obtain a third discrete coding vector sequence sample;

[0443] Inputting the first discrete code vector sequence sample, the second discrete code vector sequence sample, and the third discrete code vector sequence sample into a decoder in the audio word segmentor, obtaining a first reconstructed acoustic feature vector sample of the first discrete code vector sequence sample, a second reconstructed acoustic feature vector sample of the second discrete code vector sequence sample, and a third reconstructed acoustic feature vector sample of the third discrete code vector sequence sample;

[0444] Inputting the first reconstructed acoustic feature vector sample, the second reconstructed acoustic feature vector sample, and the third reconstructed acoustic feature vector sample into the discriminator in the audio word segmenter, obtaining a first discriminant latent variable of the first reconstructed acoustic feature vector sample, a second discriminant latent variable of the second reconstructed acoustic feature vector sample, and a third discriminant latent variable of the third reconstructed acoustic feature vector sample;

[0445] The first discriminant latent variable is used as a positive sample, the second discriminant latent variable is used as a negative sample, and the third discriminant latent variable is used as an anchor sample to calculate the triplet loss value;

[0446] According to the triplet loss value, the model parameters of the initial module are updated to obtain the autoregressive module, wherein the triplet loss value corresponding to the model parameters of the autoregressive module is the minimum value.

[0447] This approach uses the similarity of the discriminator's output as a basis for model optimization, optimizing the naturalness of the autoregressive module's contextual cohesion. This ensures that the restored second audio clip, derived from the autoregressive module's inference, seamlessly connects with the context. Furthermore, the reused audio word segmenter discriminates the naturalness of the autoregressive module and optimizes the model through comparative learning. This optimization eliminates the need for a new model structure, reducing training costs.

[0448] In some embodiments, the processor 1510 may also be configured to:

[0449] Based on each of the M original audios, K audio segment combinations are randomly generated, where the audio segment combinations include a main audio segment sample, a preceding audio segment sample of the main audio segment sample, and a following audio segment sample; M and K are positive integers;

[0450] Obtain a reference audio segment from the reference audio; the reference audio is any original audio from the M original audios except the original audio;

[0451] Performing audio fusion processing on the main text audio clip sample and the reference audio clip to obtain a noise-added audio clip sample;

[0452] Obtaining an audio sequence sample according to the main text audio segment sample, the preceding audio segment sample, the following audio segment sample, and the noise-added audio segment sample;

[0453] Perform feature extraction on the audio sequence samples to obtain acoustic feature vector sequence samples;

[0454] The acoustic feature vector sequence samples are input into the audio word segmenter, and the discrete coding vector sequence training samples are extracted from the acoustic feature vector sequence samples by the audio word segmenter.

[0455] In this way, the acquisition of training samples for the autoregressive module does not require manual labeling. Audio sequence samples can be generated based on self-supervised learning, and then discrete coding vector sequence samples for model training can be obtained, saving labeling costs and improving the generalization ability of the autoregressive module.

[0456] In some embodiments, the processor 1510 may also be configured to:

[0457] Perform speech detection on the text audio clip samples;

[0458] In the case where the main text audio segment sample includes speech information, a first audio sequence sample is obtained according to the main text audio segment sample, the preceding audio segment sample, the following audio segment sample, and the noise-added audio segment sample; the first audio sequence sample is used to train a first autoregressive module;

[0459] When the main audio segment sample does not include speech information, a second audio sequence sample is obtained based on the main audio segment sample, the previous audio segment sample, the following audio segment sample and the noise-added audio segment sample; the second audio sequence sample is used to train the second autoregressive module.

[0460] In this way, a voice activity detection label can be added to the audio sequence sample to indicate whether the main audio clip sample contains voice information. Then, based on this label, the training task can be divided into two independent tasks, and two independent autoregressive modules can be trained. During inference, the corresponding autoregressive module is called according to the user's voice processing intention. While reducing the training difficulty, the repaired second audio clip that meets the user's voice processing intention can be obtained.

[0461] In some embodiments, the processor 1510 may also be configured to:

[0462] The repaired second audio segment is sent to the client, so that the client replaces the first audio segment in the audio with the repaired second audio segment, where the first audio segment is an audio segment intercepted from the audio.

[0463] In this way, the repaired second audio segment is sent to the client, so that the client replaces the first audio segment with the repaired second audio segment, thereby obtaining audio with more natural context connection and better repair effect.

[0464] In some embodiments, the processor 1510 may also be configured to:

[0465] Obtaining a first audio segment to be restored, a preceding audio segment of the first audio segment, and a following audio segment uploaded by a client;

[0466] The first audio segment, the preceding audio segment, and the following audio segment are audio segments obtained by intercepting the audio in response to the user's first input of the audio.

[0467] In this way, based on the first audio clip selected by the user on the client, the first audio clip, the previous audio clip and the following audio clip can be uploaded to the server. The server can then automatically repair the first audio clip through the autoregressive module, which simplifies the user's operation process and improves the convenience of audio repair.

[0468] In some embodiments, the processor 1510 may also be configured to:

[0469] Get the voice processing setting information uploaded by the client;

[0470] identifying a vocal processing intent for the first audio segment based on the vocal processing setting information;

[0471] Among them, the human voice processing intention is used to select the autoregressive module in the audio model. When the human voice processing intention is to retain the voice information in the first audio segment, the predicted discrete coding vector is obtained through the first autoregressive module; when the human voice processing intention is to remove the voice information in the first audio segment, the predicted discrete coding vector is obtained through the second autoregressive module; the first autoregressive module is trained based on the first audio sequence samples, and the second autoregressive module is trained based on the second audio sequence samples. The first audio sequence samples and the second audio sequence samples both include text audio segment samples, previous audio segment samples, following audio segment samples and noisy audio segment samples after the text audio segment samples are denoised; the text audio segment samples in the first audio sequence samples include voice information, and the text audio segment samples in the second audio sequence samples do not include voice information.

[0472] In this way, the vocal processing setting information uploaded by the client can be obtained to identify the vocal processing intention of the first audio clip, and then the autoregressive module corresponding to the vocal processing intention can be called for reasoning to obtain the repaired second audio clip that can meet user needs.

[0473] In some embodiments, the processor 1510 may also be configured to:

[0474] Replaces the first audio clip in the audio with the second audio clip.

[0475] In this way, the client can directly replace the first audio segment intercepted from the audio with the second audio segment to obtain audio with more natural context connection and better restoration effect.

[0476] In some embodiments, the display unit 1506 can be used to:

[0477] Displays the audio editing interface, which includes an audio preview progress control;

[0478] The processor 1510 may also be used to:

[0479] In response to a second input to the audio preview progress control, determining a segment captured by the second input on the audio preview progress control as a first audio segment to be repaired;

[0480] The preceding audio segment and the following audio segment of the first audio segment are extracted from the audio.

[0481] In this way, the user only needs to make relevant operations on the audio preview progress control in the audio editing interface to capture the first audio clip, the previous audio clip and the following audio clip to achieve automatic audio repair, which simplifies the user's operation process and improves the convenience of audio repair.

[0482] In some embodiments, the editing interface further includes a vocal setting control, and the processor 1510 may further be configured to:

[0483] acquiring vocal processing setting information in response to a third input to the vocal setting control;

[0484] identifying a vocal processing intent for the first audio segment based on the vocal processing setting information;

[0485] Among them, the human voice processing intention is used to select the autoregressive module in the audio model. When the human voice processing intention is to retain the voice information in the first audio segment, the predicted discrete coding vector is obtained through the first autoregressive module; when the human voice processing intention is to remove the voice information in the first audio segment, the predicted discrete coding vector is obtained through the second autoregressive module; the first autoregressive module is trained based on the first audio sequence samples, and the second autoregressive module is trained based on the second audio sequence samples. The first audio sequence samples and the second audio sequence samples both include text audio segment samples, previous audio segment samples, following audio segment samples and noisy audio segment samples after the text audio segment samples are denoised; the text audio segment samples in the first audio sequence samples include voice information, and the text audio segment samples in the second audio sequence samples do not include voice information.

[0486] In this way, based on the user's relevant operations on the voice setting controls, the voice processing intention of the first audio clip can be identified, and then the autoregressive module corresponding to the voice processing intention can be called for reasoning to obtain a repaired second audio clip that can meet the user's needs.

[0487] It should be understood that in an embodiment of the present application, the input unit 1504 may include a graphics processing unit (GPU) 15041 and a microphone 15042, and the graphics processor 15041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1506 may include a display panel 15061, and the display panel 15061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1507 includes a touch panel 15071 and at least one of other input devices 15072. The touch panel 15071 is also called a touch screen. The touch panel 15071 may include two parts: a touch detection device and a touch controller. Other input devices 15072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an operating stick, which will not be repeated here.

[0488] The memory 1509 can be used to store software programs and various data. The memory 1509 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1509 may include a volatile memory or a non-volatile memory, or the memory 1509 may include both volatile and non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 1509 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0489] Processor 1510 may include one or more processing units. Optionally, processor 1510 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 1510.

[0490] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned audio repair method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0491] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0492] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned audio repair method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0493] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0494] An embodiment of the present application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-mentioned audio repair method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0495] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0496] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0497] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. An audio repair method, characterized in that: The method comprises: performing feature extraction on a first audio segment to be repaired, a preceding audio segment, and a following audio segment of the first audio segment to obtain a first acoustic feature vector of the first audio segment, a second acoustic feature vector of the preceding audio segment, and a third acoustic feature vector of the following audio segment; The first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector are input into a large audio model, autoregressive prediction is performed through the large audio model to obtain a predicted discrete coding vector, and the predicted discrete coding vector is decoded to obtain a repaired second audio segment.

2. The method according to claim 1, characterized in that The audio large model includes an audio word segmenter and an autoregressive module; Inputting the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector into a large audio model, performing autoregressive prediction using the large audio model to obtain a predicted discrete coding vector, and decoding the predicted discrete coding vector to obtain a restored second audio segment, includes: Inputting the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector into the audio word segmenter, and extracting a first discrete coding vector of the first acoustic feature vector, a second discrete coding vector of the second acoustic feature vector, and a third discrete coding vector of the third acoustic feature vector through the audio word segmenter; Inputting a discrete coding vector sequence into the autoregressive module, performing deep inference on the discrete coding vector sequence through the autoregressive module, and outputting a predicted discrete coding vector; the discrete coding vector sequence includes the first discrete coding vector, the second discrete coding vector, and the third discrete coding vector; The predicted discrete code vector is input into the audio word segmenter, and the predicted discrete code vector is decoded by the audio word segmenter to obtain a repaired second audio segment.

3. The method according to claim 2, characterized in that Before inputting the discrete code vector sequence into the autoregressive module, performing deep inference on the discrete code vector sequence through the autoregressive module, and outputting the predicted discrete code vector, the method further includes: identifying a vocal processing intent for the first audio segment; The step of inputting the discrete coding vector sequence into an autoregressive module, performing deep inference on the discrete coding vector sequence through the autoregressive module, and outputting a predicted discrete coding vector comprises: In a case where the vocal processing is intended to retain speech information in the first audio segment, inputting the discrete code vector sequence into a first autoregressive module, performing deep inference on the discrete code vector sequence through the first autoregressive module, and outputting a predicted discrete code vector; In a case where the vocal processing is intended to remove speech information from the first audio segment, inputting the discrete code vector sequence into a second autoregressive module, performing deep inference on the discrete code vector sequence through the second autoregressive module, and outputting a predicted discrete code vector; The first autoregressive module is trained based on a first audio sequence sample, and the second autoregressive module is trained based on a second audio sequence sample. The first audio sequence sample and the second audio sequence sample both include a main text audio segment sample, a previous audio segment sample, a following audio segment sample, and a noisy audio segment sample after the main text audio segment sample is denoised. The main text audio segment sample in the first audio sequence sample includes speech information, and the main text audio segment sample in the second audio sequence sample does not include speech information.

4. The method according to claim 2 or 3, characterized in that The audio word segmenter includes a generator, and the generator includes an encoder and a vector quantization module; The step of inputting the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector into an audio word segmenter, and extracting a first discrete coding vector of the first acoustic feature vector, a second discrete coding vector of the second acoustic feature vector, and a third discrete coding vector of the third acoustic feature vector by the audio word segmenter, comprises: Inputting the first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector into the encoder, and downsampling the first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector by the encoder to obtain a first encoding vector of the first acoustic eigenvector, a second encoding vector of the second acoustic eigenvector, and a third encoding vector of the third acoustic eigenvector; Inputting the first encoding vector, the second encoding vector, and the third encoding vector into the vector quantization module, the vector quantization module splitting the first encoding vector into N first low-dimensional vectors, splitting the second encoding vector into N second low-dimensional vectors, and splitting the third encoding vector into N third low-dimensional vectors, where N is an integer greater than 1; The vector quantization module calculates a first cosine similarity between each of the first low-dimensional vectors and a plurality of codebook vectors, a second cosine similarity between each of the second low-dimensional vectors and the plurality of codebook vectors, and a third cosine similarity between each of the third low-dimensional vectors and the plurality of codebook vectors; The vector quantization module determines the first discrete code vector according to the first cosine similarity, determines the second discrete code vector according to the second cosine similarity, and determines the third discrete code vector according to the third cosine similarity.

5. The method according to claim 4, characterized in that The audio word segmenter further includes a vocoder, and the generator further includes a decoder; The step of inputting the predicted discrete code vector into the audio word segmenter and decoding the predicted discrete code vector by the audio word segmenter to obtain a repaired second audio segment includes: Inputting the predicted discrete code vector into the decoder, and decoding the predicted discrete code vector by the decoder to obtain a fourth acoustic feature vector; The fourth acoustic feature vector is input into the vocoder, and the vocoder converts the fourth acoustic feature vector into a second audio segment.

6. The method according to claim 5, characterized in that The audio word segmenter also includes a discriminator; the training method of the audio word segmenter is as follows: Obtaining a first acoustic eigenvector sample; Inputting the first acoustic feature vector sample into the generator, and the encoder downsampling the first acoustic feature vector sample to obtain an encoded vector sample; The vector quantization module splits the coding vector sample into N low-dimensional vector samples, calculates the cosine similarity between each low-dimensional vector sample and a plurality of codebook vectors, and determines N codebook vector samples according to the cosine similarity; Inputting the N codebook vector samples into the decoder, and upsampling the N codebook vector samples by the decoder to obtain second acoustic feature vector samples; Calculating a first loss value based on a pre-constructed first loss function and the second acoustic feature vector sample; Inputting the first acoustic feature vector sample or the second acoustic feature vector sample into the discriminator to obtain a probability value that the input sample is the first acoustic feature vector sample; Calculating a second loss value based on a pre-constructed second loss function and the probability value; According to the sum of the first loss value and the second loss value, the parameters of the generator and the parameters of the discriminator are updated to obtain a trained audio segmenter.

7. The method according to claim 2 or 3, characterized in that The autoregressive module includes an acoustic embedding submodule, a sequence modeling submodule and a coding prediction submodule; The step of inputting the discrete coding vector sequence into an autoregressive module, performing deep inference on the discrete coding vector sequence through the autoregressive module, and outputting a predicted discrete coding vector comprises: Inputting the discrete encoding vector sequence into the acoustic embedding submodule to obtain latent variables; Inputting the latent variables into the sequence modeling submodule to obtain a latent variable sequence; Inputting the last frame in the latent variable sequence into the coding prediction submodule, generating multiple probability distribution vectors through the coding prediction submodule, performing core sampling on each of the probability distribution vectors, and obtaining multiple candidate discrete coding vectors; Among the multiple candidate discrete coding vectors, a candidate discrete coding vector corresponding to a reference terminator is determined as a predicted discrete coding vector, and the predicted discrete coding vector is output.

8. The method according to claim 7, characterized in that The training method of the autoregressive module is as follows: Obtaining discrete coding vector sequence training samples; wherein each discrete coding vector sequence training sample is obtained by splicing a first discrete coding vector sample, a second discrete coding vector sample, a third discrete coding vector sample, and a fourth discrete coding vector sample, and the splicing position of the fourth discrete coding vector sample is located at the end of the discrete coding vector sequence training samples, and corresponding start symbols are added to the heads of the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample, and the fourth discrete coding vector sample, respectively, and corresponding terminators are added to the tails of the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample, and the fourth discrete coding vector sample, respectively; Performing autoregressive training on the model according to the discrete coding vector sequence training samples to obtain a trained initial module; Obtaining discrete code vector sequence test samples; wherein each discrete code vector sequence test sample is obtained by splicing a first discrete code vector sample, a second discrete code vector sample, and a third discrete code vector sample; a corresponding start symbol is added to the head of each of the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample; a corresponding terminator is added to the tail of each of the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample; and a start symbol corresponding to the fourth discrete code vector sample is added after the terminator corresponding to the fifth discrete code vector sample; the fifth discrete code vector sample is a discrete code vector sample whose splicing position is located at the tail of the discrete code vector sequence test sample; Inputting the discrete code vector sequence test sample into the initial module and outputting the predicted discrete code vector sample; According to the difference value between the predicted discrete code vector sample and the fourth discrete code vector sample, the initial module is optimized to obtain an autoregressive module.

9. The method according to claim 8, characterized in that The step of inputting the discrete code vector sequence test sample into the initial module and outputting the predicted discrete code vector sample comprises: Inputting the discrete coding vector sequence test sample into the acoustic embedding submodule in the initial module to obtain a latent variable sample; Inputting the latent variable sample into the sequence modeling submodule in the initial module to obtain a latent variable sequence sample; The last frame in the latent variable sequence sample is input into the coding prediction submodule in the initial module, multiple probability distribution vectors are generated by the coding prediction submodule, the discrete coding vector corresponding to the maximum value of the probability distribution vector is determined as the predicted discrete coding vector sample, and the predicted discrete coding vector sample is output.

10. The method according to claim 9, characterized in that The step of optimizing the initial module according to the difference between the predicted discrete code vector sample and the fourth discrete code vector sample to obtain the autoregressive module includes: splicing the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample to obtain a first discrete code vector sequence sample; splicing the fourth discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample to obtain a second discrete code vector sequence sample; splicing the predicted discrete coding vector sample, the second discrete coding vector sample, and the third discrete coding vector sample to obtain a third discrete coding vector sequence sample; Inputting the first discrete code vector sequence sample, the second discrete code vector sequence sample, and the third discrete code vector sequence sample into a decoder in the audio word segmentor, obtaining a first reconstructed acoustic feature vector sample of the first discrete code vector sequence sample, a second reconstructed acoustic feature vector sample of the second discrete code vector sequence sample, and a third reconstructed acoustic feature vector sample of the third discrete code vector sequence sample; Inputting the first reconstructed acoustic feature vector sample, the second reconstructed acoustic feature vector sample, and the third reconstructed acoustic feature vector sample into the discriminator in the audio word segmenter, obtaining a first discriminant latent variable of the first reconstructed acoustic feature vector sample, a second discriminant latent variable of the second reconstructed acoustic feature vector sample, and a third discriminant latent variable of the third reconstructed acoustic feature vector sample; Taking the first discriminant latent variable as a positive sample, the second discriminant latent variable as a negative sample, and the third discriminant latent variable as an anchor sample, and calculating the triplet loss value; According to the triple loss value, the model parameters of the initial module are updated to obtain an autoregressive module, wherein the triple loss value corresponding to the model parameters of the autoregressive module is a minimum value.

11. The method according to claim 9, characterized in that The obtaining of discrete coding vector sequence training samples includes: Randomly generate K audio segment combinations based on each of the M original audio segments, wherein the audio segment combinations include a main audio segment sample, a preceding audio segment sample of the main audio segment sample, and a following audio segment sample; M and K are positive integers; Obtain a reference audio segment from the reference audio; the reference audio is any original audio from the M original audios except the original audio; Performing audio fusion processing on the text audio segment sample and the reference audio segment to obtain a noise-added audio segment sample; Obtaining an audio sequence sample according to the main text audio segment sample, the preceding audio segment sample, the following audio segment sample, and the noise-added audio segment sample; Performing feature extraction on the audio sequence samples to obtain acoustic feature vector sequence samples; The acoustic feature vector sequence samples are input into the audio word segmenter, and discrete coding vector sequence training samples are extracted from the acoustic feature vector sequence samples by the audio word segmenter.

12. The method according to claim 11, characterized in that The step of obtaining an audio sequence sample according to the main text audio segment sample, the preceding audio segment sample, the following audio segment sample, and the noise-added audio segment sample comprises: Performing voice detection on the text audio clip sample; In a case where the main text audio segment sample includes speech information, a first audio sequence sample is obtained based on the main text audio segment sample, the preceding audio segment sample, the following audio segment sample, and the noise-added audio segment sample; the first audio sequence sample is used to train a first autoregressive module; When the main text audio segment sample does not include speech information, a second audio sequence sample is obtained based on the main text audio segment sample, the previous audio segment sample, the following audio segment sample and the noisy audio segment sample; the second audio sequence sample is used to train a second autoregressive module.

13. The method according to claim 1, wherein When the method is executed by a server, after inputting the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector into a large audio model, performing autoregressive prediction using the large audio model to obtain a predicted discrete coding vector, and decoding the predicted discrete coding vector to obtain a repaired second audio segment, the method further includes: The repaired second audio segment is sent to the client, so that the client replaces the first audio segment in the audio with the second audio segment, where the first audio segment is an audio segment intercepted from the audio.

14. The method according to claim 13, characterized in that Before extracting features from the first audio segment to be repaired, the preceding audio segment, and the following audio segment of the first audio segment to obtain a first acoustic feature vector of the first audio segment, a second acoustic feature vector of the preceding audio segment, and a third acoustic feature vector of the following audio segment, the method further includes: Obtaining a first audio segment to be restored, a preceding audio segment, and a following audio segment of the first audio segment uploaded by the client; The first audio segment, the preceding audio segment, and the following audio segment are audio segments obtained by intercepting the audio by the client in response to the user's first input on the audio.

15. The method according to claim 14, characterized in that The method further comprises: Obtaining voice processing setting information uploaded by the client; identifying a vocal processing intent for the first audio segment according to the vocal processing setting information; In which, the human voice processing intention is used to select the autoregressive module in the audio model. When the human voice processing intention is to retain the speech information in the first audio segment, the predicted discrete coding vector is obtained through the first autoregressive module; when the human voice processing intention is to remove the speech information in the first audio segment, the predicted discrete coding vector is obtained through the second autoregressive module; the first autoregressive module is trained based on the first audio sequence samples, and the second autoregressive module is trained based on the second audio sequence samples. The first audio sequence samples and the second audio sequence samples both include text audio segment samples, previous audio segment samples, following audio segment samples and noisy audio segment samples after the text audio segment samples are denoised; the text audio segment samples in the first audio sequence samples include speech information, and the text audio segment samples in the second audio sequence samples do not include speech information.

16. The method according to claim 1, wherein When the method is executed by a client, after inputting the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector into a large audio model, performing autoregressive prediction using the large audio model to obtain a predicted discrete coding vector, and decoding the predicted discrete coding vector to obtain a restored second audio segment, the method further includes: The first audio segment in the audio is replaced with the second audio segment.

17. The method according to claim 16, characterized in that Before extracting features from the first audio segment to be repaired, the preceding audio segment, and the following audio segment of the first audio segment to obtain a first acoustic feature vector of the first audio segment, a second acoustic feature vector of the preceding audio segment, and a third acoustic feature vector of the following audio segment, the method further includes: Displaying an audio editing interface, wherein the editing interface includes an audio preview progress control; In response to a second input to the audio preview progress control, determining a segment intercepted by the second input on the audio preview progress control as a first audio segment to be repaired; The preceding audio segment and the following audio segment of the first audio segment are extracted from the audio.

18. The method according to claim 17, characterized in that The editing interface further includes a vocal setting control, and the method further includes: acquiring vocal processing setting information in response to a third input to the vocal setting control; identifying a vocal processing intent for the first audio segment according to the vocal processing setting information; In which, the human voice processing intention is used to select the autoregressive module in the audio model. When the human voice processing intention is to retain the speech information in the first audio segment, the predicted discrete coding vector is obtained through the first autoregressive module; when the human voice processing intention is to remove the speech information in the first audio segment, the predicted discrete coding vector is obtained through the second autoregressive module; the first autoregressive module is trained based on the first audio sequence samples, and the second autoregressive module is trained based on the second audio sequence samples. The first audio sequence samples and the second audio sequence samples both include text audio segment samples, previous audio segment samples, following audio segment samples and noisy audio segment samples after the text audio segment samples are denoised; the text audio segment samples in the first audio sequence samples include speech information, and the text audio segment samples in the second audio sequence samples do not include speech information.

19. An audio repair device, characterized in that: The device comprises: a feature extraction module, configured to perform feature extraction on a first audio segment to be repaired, a preceding audio segment, and a following audio segment of the first audio segment, to obtain a first acoustic feature vector of the first audio segment, a second acoustic feature vector of the preceding audio segment, and a third acoustic feature vector of the following audio segment; A processing module is used to input the first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector into an audio large model, perform autoregressive prediction through the audio large model to obtain a predicted discrete coding vector, and decode the predicted discrete coding vector to obtain a repaired second audio segment.

20. The device according to claim 19, characterized in that The audio large model includes an audio word segmenter and an autoregressive module; the processing module is further used to: Inputting the first acoustic feature vector, the second acoustic feature vector, and the third acoustic feature vector into the audio word segmenter, and extracting a first discrete coding vector of the first acoustic feature vector, a second discrete coding vector of the second acoustic feature vector, and a third discrete coding vector of the third acoustic feature vector through the audio word segmenter; Inputting a discrete coding vector sequence into the autoregressive module, performing deep inference on the discrete coding vector sequence through the autoregressive module, and outputting a predicted discrete coding vector; the discrete coding vector sequence includes the first discrete coding vector, the second discrete coding vector, and the third discrete coding vector; The predicted discrete code vector is input into the audio word segmenter, and the predicted discrete code vector is decoded by the audio word segmenter to obtain a repaired second audio segment.

21. The device according to claim 20, characterized in that The processing module is further configured to: Before inputting the discrete code vector sequence into the autoregressive module, performing deep inference on the discrete code vector sequence by the autoregressive module, and outputting a predicted discrete code vector, identifying a vocal processing intention of the first audio segment; In a case where the vocal processing is intended to retain speech information in the first audio segment, inputting the discrete code vector sequence into a first autoregressive module, performing deep inference on the discrete code vector sequence through the first autoregressive module, and outputting a predicted discrete code vector; In a case where the vocal processing is intended to remove speech information from the first audio segment, inputting the discrete code vector sequence into a second autoregressive module, performing deep inference on the discrete code vector sequence through the second autoregressive module, and outputting a predicted discrete code vector; The first autoregressive module is trained based on a first audio sequence sample, and the second autoregressive module is trained based on a second audio sequence sample. The first audio sequence sample and the second audio sequence sample both include a main text audio segment sample, a previous audio segment sample, a following audio segment sample, and a noisy audio segment sample after the main text audio segment sample is denoised. The main text audio segment sample in the first audio sequence sample includes speech information, and the main text audio segment sample in the second audio sequence sample does not include speech information.

22. The device according to claim 20 or 21, characterized in that The audio word segmenter includes a generator, and the generator includes an encoder and a vector quantization module; The processing module is further configured to: Inputting the first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector into the encoder, and downsampling the first acoustic eigenvector, the second acoustic eigenvector, and the third acoustic eigenvector by the encoder to obtain a first encoding vector of the first acoustic eigenvector, a second encoding vector of the second acoustic eigenvector, and a third encoding vector of the third acoustic eigenvector; Inputting the first encoding vector, the second encoding vector, and the third encoding vector into the vector quantization module, the vector quantization module splitting the first encoding vector into N first low-dimensional vectors, splitting the second encoding vector into N second low-dimensional vectors, and splitting the third encoding vector into N third low-dimensional vectors, where N is an integer greater than 1; The vector quantization module calculates a first cosine similarity between each of the first low-dimensional vectors and a plurality of codebook vectors, a second cosine similarity between each of the second low-dimensional vectors and the plurality of codebook vectors, and a third cosine similarity between each of the third low-dimensional vectors and the plurality of codebook vectors; The vector quantization module determines the first discrete code vector according to the first cosine similarity, determines the second discrete code vector according to the second cosine similarity, and determines the third discrete code vector according to the third cosine similarity.

23. The device according to claim 22, characterized in that The audio word segmenter further includes a vocoder, and the generator further includes a decoder; The processing module is further configured to: Inputting the predicted discrete code vector into the decoder, and decoding the predicted discrete code vector by the decoder to obtain a fourth acoustic feature vector; The fourth acoustic feature vector is input into the vocoder, and the vocoder converts the fourth acoustic feature vector into a second audio segment.

24. The device according to claim 23, characterized in that The audio word segmenter further includes a discriminator; the device further includes an audio word segmenter training module, which is used to: Obtaining a first acoustic eigenvector sample; Inputting the first acoustic feature vector sample into the generator, and the encoder downsampling the first acoustic feature vector sample to obtain an encoded vector sample; The vector quantization module splits the coding vector sample into N low-dimensional vector samples, calculates the cosine similarity between each low-dimensional vector sample and a plurality of codebook vectors, and determines N codebook vector samples according to the cosine similarity; Inputting the N codebook vector samples into the decoder, and upsampling the N codebook vector samples by the decoder to obtain second acoustic feature vector samples; Calculating a first loss value based on a pre-constructed first loss function and the second acoustic feature vector sample; Inputting the first acoustic feature vector sample or the second acoustic feature vector sample into the discriminator to obtain a probability value that the input sample is the first acoustic feature vector sample; Calculating a second loss value based on a pre-constructed second loss function and the probability value; According to the sum of the first loss value and the second loss value, the parameters of the generator and the parameters of the discriminator are updated to obtain a trained audio segmenter.

25. The device according to claim 20 or 21, characterized in that The autoregressive module includes an acoustic embedding submodule, a sequence modeling submodule and a coding prediction submodule; The processing module is further configured to: Inputting the discrete encoding vector sequence into the acoustic embedding submodule to obtain latent variables; Inputting the latent variables into the sequence modeling submodule to obtain a latent variable sequence; Inputting the last frame in the latent variable sequence into the coding prediction submodule, generating multiple probability distribution vectors through the coding prediction submodule, performing core sampling on each of the probability distribution vectors, and obtaining multiple candidate discrete coding vectors; Among the multiple candidate discrete coding vectors, a candidate discrete coding vector corresponding to a reference terminator is determined as a predicted discrete coding vector, and the predicted discrete coding vector is output.

26. The device according to claim 25, characterized in that The device also includes a model training module, which is used to: Obtaining discrete coding vector sequence training samples; wherein each discrete coding vector sequence training sample is obtained by splicing a first discrete coding vector sample, a second discrete coding vector sample, a third discrete coding vector sample, and a fourth discrete coding vector sample, and the splicing position of the fourth discrete coding vector sample is located at the end of the discrete coding vector sequence training samples, and corresponding start symbols are added to the heads of the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample, and the fourth discrete coding vector sample, respectively, and corresponding terminators are added to the tails of the first discrete coding vector sample, the second discrete coding vector sample, the third discrete coding vector sample, and the fourth discrete coding vector sample, respectively; Performing autoregressive training on the model according to the discrete coding vector sequence training samples to obtain a trained initial module; Obtaining discrete code vector sequence test samples; wherein each discrete code vector sequence test sample is obtained by splicing a first discrete code vector sample, a second discrete code vector sample, and a third discrete code vector sample; a corresponding start symbol is added to the head of each of the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample; a corresponding terminator is added to the tail of each of the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample; and a start symbol corresponding to the fourth discrete code vector sample is added after the terminator corresponding to the fifth discrete code vector sample; the fifth discrete code vector sample is a discrete code vector sample whose splicing position is located at the tail of the discrete code vector sequence test sample; Inputting the discrete code vector sequence test sample into the initial module and outputting the predicted discrete code vector sample; According to the difference value between the predicted discrete code vector sample and the fourth discrete code vector sample, the initial module is optimized to obtain the autoregressive module.

27. The device according to claim 26, characterized in that The model training module is further used to: Inputting the discrete coding vector sequence test sample into the acoustic embedding submodule in the initial module to obtain a latent variable sample; Inputting the latent variable sample into the sequence modeling submodule in the initial module to obtain a latent variable sequence sample; The last frame in the latent variable sequence sample is input into the coding prediction submodule in the initial module, multiple probability distribution vectors are generated by the coding prediction submodule, the discrete coding vector corresponding to the maximum value of the probability distribution vector is determined as the predicted discrete coding vector sample, and the predicted discrete coding vector sample is output.

28. The device according to claim 26, characterized in that The model training module is further used to: splicing the first discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample to obtain a first discrete code vector sequence sample; splicing the fourth discrete code vector sample, the second discrete code vector sample, and the third discrete code vector sample to obtain a second discrete code vector sequence sample; splicing the predicted discrete coding vector sample, the second discrete coding vector sample, and the third discrete coding vector sample to obtain a third discrete coding vector sequence sample; Inputting the first discrete code vector sequence sample, the second discrete code vector sequence sample, and the third discrete code vector sequence sample into a decoder in the audio word segmentor, obtaining a first reconstructed acoustic feature vector sample of the first discrete code vector sequence sample, a second reconstructed acoustic feature vector sample of the second discrete code vector sequence sample, and a third reconstructed acoustic feature vector sample of the third discrete code vector sequence sample; Inputting the first reconstructed acoustic feature vector sample, the second reconstructed acoustic feature vector sample, and the third reconstructed acoustic feature vector sample into the discriminator in the audio word segmenter, obtaining a first discriminant latent variable of the first reconstructed acoustic feature vector sample, a second discriminant latent variable of the second reconstructed acoustic feature vector sample, and a third discriminant latent variable of the third reconstructed acoustic feature vector sample; Taking the first discriminant latent variable as a positive sample, the second discriminant latent variable as a negative sample, and the third discriminant latent variable as an anchor sample, and calculating the triplet loss value; According to the triple loss value, the model parameters of the initial module are updated to obtain an autoregressive module, wherein the triple loss value corresponding to the model parameters of the autoregressive module is a minimum value.

29. The device according to claim 27, characterized in that The model training module is further used to: Randomly generate K audio segment combinations based on each of the M original audio segments, wherein the audio segment combinations include a main audio segment sample, a preceding audio segment sample of the main audio segment sample, and a following audio segment sample; M and K are positive integers; Obtain a reference audio segment from the reference audio; the reference audio is any original audio from the M original audios except the original audio; Performing audio fusion processing on the text audio segment sample and the reference audio segment to obtain a noise-added audio segment sample; Obtaining an audio sequence sample according to the main text audio segment sample, the preceding audio segment sample, the following audio segment sample, and the noise-added audio segment sample; Performing feature extraction on the audio sequence samples to obtain acoustic feature vector sequence samples; The acoustic feature vector sequence samples are input into the audio word segmenter, and the discrete coding vector sequence training samples are extracted from the acoustic feature vector sequence samples by the audio word segmenter.

30. The device according to claim 29, characterized in that The model training module is further used to: Performing voice detection on the text audio clip sample; In a case where the main text audio segment sample includes speech information, a first audio sequence sample is obtained based on the main text audio segment sample, the preceding audio segment sample, the following audio segment sample, and the noise-added audio segment sample; the first audio sequence sample is used to train a first autoregressive module; When the main text audio segment sample does not include speech information, a second audio sequence sample is obtained based on the main text audio segment sample, the previous audio segment sample, the following audio segment sample and the noisy audio segment sample; the second audio sequence sample is used to train a second autoregressive module.

31. The device according to claim 29, characterized in that The device further comprises: The transmission module is configured to send the repaired second audio segment to the client, so that the client replaces the first audio segment in the audio with the repaired second audio segment, where the first audio segment is an audio segment intercepted from the audio.

32. The device according to claim 31, characterized in that The device further comprises: an acquisition module, configured to acquire a first audio segment to be restored, a preceding audio segment, and a following audio segment of the first audio segment, uploaded by the client; The first audio segment, the preceding audio segment, and the following audio segment are audio segments obtained by intercepting the audio by the client in response to the user's first input on the audio.

33. The device according to claim 32, characterized in that The acquisition module is further used to: Obtaining voice processing setting information uploaded by the client; The processing module is further configured to: identifying a vocal processing intent for the first audio segment according to the vocal processing setting information; Among them, the human voice processing intention is used to select the autoregressive module in the audio model. When the human voice processing intention is to retain the voice information in the first audio segment, the predicted discrete coding vector is obtained by the first autoregressive module; when the human voice processing intention is to remove the voice information in the first audio segment, the predicted discrete coding vector is obtained by the second autoregressive module; the first autoregressive module is trained based on the first audio sequence sample, and the second autoregressive module is trained based on the second audio sequence sample. The first audio sequence sample and the second audio sequence sample both include text audio segment samples, previous audio segment samples, following audio segment samples and noisy audio segment samples after the text audio segment samples are denoised; the text audio segment samples in the first audio sequence samples include voice information, and the text audio segment samples in the second audio sequence samples do not include voice information.

34. The device according to claim 19, wherein The processing module is further configured to: The first audio segment in the audio is replaced with the second audio segment.

35. The device according to claim 34, characterized in that The device further comprises: A display module, configured to display an audio editing interface, wherein the editing interface includes an audio preview progress control; The processing module is further used for: In response to a second input to the audio preview progress control, determining a segment intercepted by the second input on the audio preview progress control as a first audio segment to be repaired; The preceding audio segment and the following audio segment of the first audio segment are extracted from the audio.

36. The device according to claim 35, characterized in that The editing interface further includes a vocal setting control, and the device further includes: an acquisition module, configured to acquire vocal processing setting information in response to a third input to the vocal setting control; The processing module is further configured to: identifying a vocal processing intent for the first audio segment according to the vocal processing setting information; Among them, the human voice processing intention is used to select the autoregressive module in the audio model. When the human voice processing intention is to retain the voice information in the first audio segment, the predicted discrete coding vector is obtained by the first autoregressive module; when the human voice processing intention is to remove the voice information in the first audio segment, the predicted discrete coding vector is obtained by the second autoregressive module; the first autoregressive module is trained based on the first audio sequence sample, and the second autoregressive module is trained based on the second audio sequence sample. The first audio sequence sample and the second audio sequence sample both include text audio segment samples, previous audio segment samples, following audio segment samples and noisy audio segment samples after the text audio segment samples are denoised; the text audio segment samples in the first audio sequence samples include voice information, and the text audio segment samples in the second audio sequence samples do not include voice information.

37. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 18 are implemented.

Citation Information

Cited By

  • Audio generation method and device based on audio processing model

    CN120877703A