Audio processing method and device and electronic equipment
By receiving text editing input and utilizing time mapping relationships and audio generation models, the system automatically edits audio, solving the problem of high barriers to audio editing in existing technologies and enabling non-professional users to process audio quickly and accurately.
Patent Information
- Application Number
- CN202511397385.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-01-09
AI Technical Summary
In existing technologies, audio editing has a high technical threshold, which means that users lacking professional knowledge need to spend a lot of time on audio processing, resulting in low efficiency and low accuracy.
By receiving text editing input, updating the text information of the audio, and processing the audio based on time mapping, including cutting, splicing and inserting audio segments, and using speech recognition and audio generation models to generate supplementary audio segments, automatic audio editing is achieved.
It enables non-professional users to edit audio quickly and accurately, improving the efficiency and accuracy of audio processing and reducing tedious cutting and splicing operations.
Smart Images

Figure CN121306139A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, specifically relating to an audio processing method, apparatus, and electronic device. Background Technology
[0002] Audio data contains rich information such as text, emotions, and gender, making it one of the main mediums for information exchange and sharing. With the upgrading of digital content consumption, the number of users of audio media such as audiobooks and podcasts is constantly increasing. At the same time, more and more people are participating in audio creation. Due to issues such as slips of the tongue, idiomatic expressions, and omissions that may occur during audio recording, users need to repair the audio data through post-processing.
[0003] In related technologies, users mainly use editing software to repair audio data. However, due to the high technical threshold of audio editing, users who lack professional editing knowledge need to spend a lot of time on audio processing, resulting in low efficiency and errors during the editing process, leading to low accuracy. Summary of the Invention
[0004] The purpose of this application is to provide an audio processing method, apparatus, electronic device, storage medium, and program product that can improve the efficiency and accuracy of audio processing.
[0005] In a first aspect, embodiments of this application provide an audio processing method, including:
[0006] Receive text editing input for the first text information of the first audio;
[0007] In response to text editing input, update the first text information to obtain the second text information;
[0008] Based on the second text information, the first audio is processed to obtain the second audio.
[0009] As one possible implementation, before receiving text-edited input of the first text information of the first audio, the method further includes:
[0010] The first audio is input into the speech recognition model, and the speech recognition model performs speech recognition on the first audio to obtain the first text information of the first audio.
[0011] The first audio and the first text information are input into the audio-text alignment model. The audio-text alignment model is used to perform time alignment on the first audio and the first text information to obtain the time mapping relationship between the first text information and the first audio. The time mapping relationship is used to indicate the start and end times of the audio segment corresponding to each character in the first text information in the first audio.
[0012] The audio editing interface is displayed, which includes a timeline, a waveform of the first audio, and first text information. The timeline includes at least one time marker, and each character in the first text information is aligned with the corresponding waveform and the corresponding time marker in the waveform diagram.
[0013] Receive text editing input for the first text information of the first audio, including:
[0014] Receive text editing input for the first text information in the audio editing interface.
[0015] As one possible implementation, the text editing input is an edit input that deletes the first character sequence in the first text information;
[0016] Update the first text information to obtain the second text information, including:
[0017] Delete the first character sequence from the first text information to obtain the second text information.
[0018] As one possible implementation, the first character sequence is a series of consecutive characters in the first text information;
[0019] Based on the second text information, the first audio is processed to obtain the second audio, including:
[0020] The second text information is compared with the first text information to determine the first character sequence that was deleted from the first text information;
[0021] Based on the time mapping relationship, the start time and end time of the first audio segment corresponding to the first character sequence are determined. The start time of the first audio segment is the start time of the audio segment corresponding to the first character in the first character sequence in the first audio, and the end time of the first audio segment is the end time of the audio segment corresponding to the last character in the first character sequence in the first audio.
[0022] Based on the start and end times of the first audio segment, the first audio is cut to obtain the first audio segment;
[0023] Delete the first audio segment from the first audio, and if a second audio segment is obtained, identify the second audio segment as the second audio;
[0024] If at least two second audio segments are obtained, the audio obtained by splicing the at least two second audio segments is determined as the second audio.
[0025] As one possible implementation, the first character sequence comprises at least two non-contiguous character subsequences;
[0026] Based on the second text information, the first audio is processed to obtain the second audio, including:
[0027] The second text information is compared with the first text information to determine at least two character subsequences that have been deleted from the first text information;
[0028] Based on the time mapping relationship, the start time and end time of the third audio segment corresponding to each character subsequence are determined. The start time of the third audio segment corresponding to each character subsequence is the start time of the audio segment corresponding to the first character in the first audio and the end time of the third audio segment corresponding to each character subsequence is the end time of the audio segment corresponding to the last character in the first audio and the last character in the character subsequence.
[0029] Based on the start and end times of at least one character subsequence, the first audio segment is cut to obtain the third audio segment corresponding to each character subsequence;
[0030] Delete at least two third audio segments from the first audio segment. If a fourth audio segment is obtained, identify the fourth audio segment as the second audio segment.
[0031] If at least two fourth audio segments are obtained, the audio obtained by splicing the at least two fourth audio segments is determined as the second audio.
[0032] As one possible implementation, text editing input is an editing input that inserts a second character sequence into the first text information;
[0033] Update the first text information to obtain the second text information, including:
[0034] Insert the second character sequence into the first text information to obtain the second text information.
[0035] As one possible implementation, based on the second text information, audio processing is performed on the first audio to obtain the second audio, including:
[0036] The second text information is compared with the first text information to determine the insertion position of the second character sequence and the second character sequence inserted in the first text information;
[0037] The first audio, the first text information, and the second character sequence are input into the audio generation model. The audio generation model generates a first supplementary audio segment corresponding to the second character sequence.
[0038] Based on the insertion position, the first supplementary audio segment is inserted into the first audio to obtain the second audio.
[0039] As one possible implementation, the insertion position is before the first character of the first text information;
[0040] Based on the insertion position, the first supplementary audio is inserted into the first audio to obtain the second audio, which includes:
[0041] After aligning the end time of the first supplementary audio segment with the start time of the first audio segment, insert the first audio segment to obtain the second audio segment.
[0042] As one possible implementation, the insertion position is after the last character of the first text message;
[0043] Based on the insertion position, the first supplementary audio segment is inserted into the first audio to obtain the second audio, which includes:
[0044] After aligning the start time of the first supplementary audio segment with the end time of the first audio segment, insert the first audio segment to obtain the second audio segment.
[0045] As one possible implementation, the insertion position is located between the first and last characters of the first text information;
[0046] Based on the insertion position, the first supplementary audio is inserted into the first audio to obtain the second audio, which includes:
[0047] Determine the first character to the left of the insertion position and the second character to the right of the insertion position in the first text information;
[0048] Based on the time mapping relationship, determine the end time of the first segment of the audio segment corresponding to the first character and the start time of the first segment of the audio segment corresponding to the second character in the first audio;
[0049] Align the start time of the first supplementary audio segment with the end time of the first segment, and align the end time of the first supplementary audio segment with the start time of the first segment, then insert the first audio to obtain the second audio.
[0050] As one possible implementation, the text editing input is an edit input that modifies the third character sequence in the first text information into a fourth character sequence;
[0051] Update the first text information to obtain the second text information, including:
[0052] The third character sequence in the first text information is replaced with the fourth character sequence to obtain the second text information.
[0053] As one possible implementation, the first audio is updated based on the second text information to obtain the second audio, including:
[0054] The second text information is compared with the first text information to determine the third character sequence that was deleted and the fourth character sequence that was added in the first text information;
[0055] Based on the time mapping relationship, the start time and end time of the fifth audio segment corresponding to the third character sequence in the first audio are determined. The start time of the fifth audio segment is the start time of the audio segment corresponding to the first character in the third character sequence in the first audio, and the end time of the fifth audio segment is the end time of the audio segment corresponding to the last character in the third character sequence in the first audio.
[0056] Based on the start and end times of the fifth audio segment, the first audio segment is cut to obtain the fifth audio segment;
[0057] The first audio, the first text information, and the fourth character sequence are input into the audio generation model. The audio generation model generates a second supplementary audio segment corresponding to the fourth character sequence.
[0058] Replace the fifth audio segment in the first audio with the second supplementary audio segment to obtain the second audio.
[0059] As one possible implementation, the audio generation model includes an encoder, a quantization module, a language model, a decoder, and a vocoder;
[0060] An encoder is used to extract the acoustic features of the first audio audio.
[0061] The quantization module is used to quantize acoustic features to obtain quantized discrete codes;
[0062] A language model is used to predict the audio encoding corresponding to a newly added character sequence in a first text message based on input features. The input features include a first character encoding sequence and a second character encoding sequence. The first character encoding sequence is obtained by word segmentation of the first text message, and the second character encoding sequence is obtained by word segmentation of the newly added character sequence. The newly added character sequence is either the second character sequence inserted into the first text message or a fourth character sequence used to replace the third character sequence in the first text message.
[0063] A decoder is used to convert audio encoding into predictive acoustic features;
[0064] A vocoder is used to generate supplementary audio segments corresponding to newly added character sequences based on predicted acoustic features.
[0065] As one possible implementation, after responding to text editing input, updating the first text information, and obtaining the second text information, the method also includes:
[0066] The second text information is displayed in the audio editing interface;
[0067] The original character sequence and the newly added character sequence in the second text information have different display modes; the original character sequence is the same as the character sequence in the first text information in the second text information, and the newly added character sequence is the second character sequence inserted into the first text information or the fourth character sequence used to replace the third character sequence in the first text information.
[0068] As one possible implementation, the method also includes:
[0069] Based on the sample dataset and loss function, the initial model is trained to obtain the speech recognition model;
[0070] The sample dataset includes multiple audio samples and corresponding text information samples. The multiple audio samples include audio recorded in at least two scenarios, and the audio samples include human voice features.
[0071] As one possible implementation, the loss function is a joint loss function, which includes at least two loss function terms, each of which corresponds to a weight, and one of the at least two loss function terms includes the forward relative entropy.
[0072] Based on the training dataset and loss function, the initial model is trained to obtain a speech recognition model, including:
[0073] Extract acoustic feature samples for each audio sample in the sample dataset;
[0074] The text information sample corresponding to each audio sample is segmented into words to obtain the character encoding sequence sample of each audio sample;
[0075] Based on the acoustic feature samples and character encoding sequence samples of all audio samples in the sample dataset, multiple sets of training data are constructed; each set of training data includes the acoustic feature samples and character encoding sequence samples of one audio sample.
[0076] The initial model is trained based on multiple sets of training data to obtain the predicted character encoding sequence corresponding to each set of training data;
[0077] Having completed at least two rounds of training on the initial model, the forward relative entropy is calculated based on the predicted character encoding sequence obtained in the i-th round of training and the predicted character encoding sequence obtained in the (i-1)-th round of training; where i ≥ 2.
[0078] Based on the forward relative entropy, update the entropy value of the forward relative entropy in the loss function to obtain the updated loss function;
[0079] Based on the updated loss function, calculate the loss value of the initial model;
[0080] If the loss value is greater than or equal to the loss threshold, adjust the weight of each loss function term in the loss function, and continue to train the initial model with adjusted parameters based on acoustic feature samples and text sequence samples.
[0081] Training stops when the loss value is less than the loss threshold, resulting in a trained speech recognition model.
[0082] As one possible implementation, the initial model includes an encoder, a decoder, and a joint network;
[0083] Based on multiple sets of training data, the initial model is trained to obtain the predicted character encoding sequence corresponding to each set of training data, including:
[0084] The acoustic feature samples from each set of training data are embedded into a spatial vector by an encoder to obtain the acoustic feature vector samples of each set of training data.
[0085] The character encoding sequence samples in each training data set are embedded into a spatial vector using a decoder, resulting in character encoding sequence vector samples for each training data set.
[0086] By using a joint network, the acoustic feature vector samples and character encoding sequence vector samples of each training data set are concatenated to obtain the concatenated vector of each training data set. The concatenated vector of each training data set is then mapped to the vocabulary space to obtain the predicted character encoding sequence of each training data set.
[0087] Secondly, embodiments of this application provide an audio processing apparatus, including:
[0088] The receiving module is used to receive text editing input of the first text information of the first audio.
[0089] The text update module is used to update the first text information in response to text editing input, and obtain the second text information;
[0090] The audio processing module is used to process the first audio based on the second text information to obtain the second audio.
[0091] As one possible implementation, the device also includes:
[0092] The speech recognition module is used to input the first audio into the speech recognition model before receiving the text editing input of the first text information of the first audio, and to perform speech recognition on the first audio through the speech recognition model to obtain the first text information of the first audio.
[0093] The alignment module is used to input the first audio and the first text information into the audio-text alignment model. Through the audio-text alignment model, the first audio and the first text information are time-aligned to obtain the time mapping relationship between the first text information and the first audio. The time mapping relationship is used to indicate the start and end times of the audio segment corresponding to each character in the first text information in the first audio.
[0094] The display module is used to display the audio editing page interface. The audio editing page interface includes a timeline, a waveform of the first audio, and first text information. The timeline includes at least one time marker. Each character in the waveform and the first text information is aligned with the corresponding waveform and the corresponding time marker in the waveform.
[0095] The receiving module is specifically used for:
[0096] Receive text editing input for the first text information in the audio editing page interface.
[0097] As one possible implementation, the text editing input is an edit input that deletes the first character sequence in the first text information;
[0098] The text update module is specifically used for:
[0099] Delete the first character sequence from the first text information to obtain the second text information.
[0100] As one possible implementation, the first character sequence is a series of consecutive characters in the first text information;
[0101] The audio processing module is specifically used for:
[0102] The second text information is compared with the first text information to determine the first character sequence that was deleted from the first text information;
[0103] Based on the time mapping relationship, the start time and end time of the first audio segment corresponding to the first character sequence are determined. The start time of the first audio segment is the start time of the audio segment corresponding to the first character in the first character sequence in the first audio, and the end time of the first audio segment is the end time of the audio segment corresponding to the last character in the first character sequence in the first audio.
[0104] Based on the start and end times of the first audio segment, the first audio is cut to obtain the first audio segment;
[0105] Delete the first audio segment from the first audio, and if a second audio segment is obtained, identify the second audio segment as the second audio;
[0106] If at least two second audio segments are obtained, the audio obtained by splicing the at least two second audio segments is determined as the second audio.
[0107] As one possible implementation, the first character sequence comprises at least two non-contiguous character subsequences;
[0108] The audio processing module is specifically used for:
[0109] The second text information is compared with the first text information to determine at least two character subsequences that have been deleted from the first text information;
[0110] Based on the time mapping relationship, the start time and end time of the third audio segment corresponding to each character subsequence are determined. The start time of the third audio segment corresponding to each character subsequence is the start time of the audio segment corresponding to the first character in the first audio and the end time of the third audio segment corresponding to each character subsequence is the end time of the audio segment corresponding to the last character in the first audio and the last character in the character subsequence.
[0111] Based on the start and end times of at least one character subsequence, the first audio segment is cut to obtain the third audio segment corresponding to each character subsequence;
[0112] Delete at least two third audio segments from the first audio segment. If a fourth audio segment is obtained, identify the fourth audio segment as the second audio segment.
[0113] If at least two fourth audio segments are obtained, the audio obtained by splicing the at least two fourth audio segments is determined as the second audio.
[0114] As one possible implementation, text editing input is an editing input that inserts a second character sequence into the first text information;
[0115] The text update module is specifically used for:
[0116] Insert the second character sequence into the first text information to obtain the second text information.
[0117] As one possible implementation, the audio processing module is specifically used for:
[0118] The second text information is compared with the first text information to determine the insertion position of the second character sequence and the second character sequence inserted in the first text information;
[0119] The first audio, the first text information, and the second character sequence are input into the audio generation model. The audio generation model generates a first supplementary audio segment corresponding to the second character sequence.
[0120] Based on the insertion position, the first supplementary audio segment is inserted into the first audio to obtain the second audio.
[0121] As one possible implementation, the insertion position is before the first character of the first text information;
[0122] The audio processing module is specifically used for:
[0123] After aligning the end time of the first supplementary audio segment with the start time of the first audio segment, insert the first audio segment to obtain the second audio segment.
[0124] As one possible implementation, the insertion position is after the last character of the first text message;
[0125] The audio processing module is specifically used for:
[0126] After aligning the start time of the first supplementary audio segment with the end time of the first audio segment, insert the first audio segment to obtain the second audio segment.
[0127] As one possible implementation, the insertion position is located between the first and last characters of the first text information;
[0128] The audio processing module is specifically used for:
[0129] Determine the first character to the left of the insertion position and the second character to the right of the insertion position in the first text information;
[0130] Based on the time mapping relationship, determine the end time of the first segment of the audio segment corresponding to the first character and the start time of the first segment of the audio segment corresponding to the second character in the first audio;
[0131] Align the start time of the first supplementary audio segment with the end time of the first segment, and align the end time of the first supplementary audio segment with the start time of the first segment, then insert the first audio to obtain the second audio.
[0132] As one possible implementation, the text editing input is an edit input that modifies the third character sequence in the first text information into a fourth character sequence;
[0133] The text update module is specifically used for:
[0134] The third character sequence in the first text information is replaced with the fourth character sequence to obtain the second text information.
[0135] As one possible implementation, the audio processing module is specifically used for:
[0136] The second text information is compared with the first text information to determine the third character sequence that was deleted and the fourth character sequence that was added in the first text information;
[0137] Based on the time mapping relationship, the start time and end time of the fifth audio segment corresponding to the third character sequence in the first audio are determined. The start time of the fifth audio segment is the start time of the audio segment corresponding to the first character in the third character sequence in the first audio, and the end time of the fifth audio segment is the end time of the audio segment corresponding to the last character in the third character sequence in the first audio.
[0138] Based on the start and end times of the fifth audio segment, the first audio segment is cut to obtain the fifth audio segment;
[0139] The first audio, the first text information, and the fourth character sequence are input into the audio generation model. The audio generation model generates a second supplementary audio segment corresponding to the fourth character sequence.
[0140] Replace the fifth audio segment in the first audio with the second supplementary audio segment to obtain the second audio.
[0141] As one possible implementation, the audio generation model includes an encoder, a quantization module, a language model, a decoder, and a vocoder;
[0142] An encoder is used to extract the acoustic features of the first audio audio.
[0143] The quantization module is used to quantize acoustic features to obtain quantized discrete codes;
[0144] A language model is used to predict the audio encoding corresponding to a newly added character sequence in a first text message based on input features. The input features include a first character encoding sequence and a second character encoding sequence. The first character encoding sequence is obtained by word segmentation of the first text message, and the second character encoding sequence is obtained by word segmentation of the newly added character sequence. The newly added character sequence is either the second character sequence inserted into the first text message or a fourth character sequence used to replace the third character sequence in the first text message.
[0145] A decoder is used to convert audio encoding into predictive acoustic features;
[0146] A vocoder is used to generate supplementary audio segments corresponding to newly added character sequences based on predicted acoustic features.
[0147] As one possible implementation, the display module is also used for:
[0148] In response to text editing input, the first text information is updated, and after obtaining the second text information, the second text information is displayed in the audio editing interface;
[0149] The original character sequence and the newly added character sequence in the second text information have different display modes; the original character sequence is the same as the character sequence in the first text information in the second text information, and the newly added character sequence is the second character sequence inserted into the first text information or the fourth character sequence used to replace the third character sequence in the first text information.
[0150] As one possible implementation, the device also includes: a model training module, used for:
[0151] Based on the sample dataset and loss function, the initial model is trained to obtain the speech recognition model;
[0152] The sample dataset includes multiple audio samples and corresponding text information samples. The multiple audio samples include audio recorded in at least two scenarios, and the audio samples include human voice features.
[0153] As one possible implementation, the loss function is a joint loss function, which includes at least two loss function terms, each of which corresponds to a weight, and one of the at least two loss function terms includes the forward relative entropy.
[0154] The model training module is specifically used for:
[0155] Extract acoustic feature samples for each audio sample in the sample dataset;
[0156] The text information sample corresponding to each audio sample is segmented into words to obtain the character encoding sequence sample of each audio sample;
[0157] Based on the acoustic feature samples and character encoding sequence samples of all audio samples in the sample dataset, multiple sets of training data are constructed; each set of training data includes the acoustic feature samples and character encoding sequence samples of one audio sample.
[0158] The initial model is trained based on multiple sets of training data to obtain the predicted character encoding sequence corresponding to each set of training data;
[0159] Having completed at least two rounds of training on the initial model, the forward relative entropy is calculated based on the predicted character encoding sequence obtained in the i-th round of training and the predicted character encoding sequence obtained in the (i-1)-th round of training; where i ≥ 2.
[0160] Based on the forward relative entropy, update the entropy value of the forward relative entropy in the loss function to obtain the updated loss function;
[0161] Based on the updated loss function, calculate the loss value of the initial model;
[0162] If the loss value is greater than or equal to the loss threshold, adjust the weight of each loss function term in the loss function, and continue to train the initial model with adjusted parameters based on acoustic feature samples and text sequence samples.
[0163] Training stops when the loss value is less than the loss threshold, resulting in a trained speech recognition model.
[0164] As one possible implementation, the initial model includes an encoder, a decoder, and a joint network;
[0165] The model training module is specifically used for:
[0166] The acoustic feature samples from each set of training data are embedded into a spatial vector by an encoder to obtain the acoustic feature vector samples of each set of training data.
[0167] The character encoding sequence samples in each training data set are embedded into a spatial vector using a decoder, resulting in character encoding sequence vector samples for each training data set.
[0168] By using a joint network, the acoustic feature vector samples and character encoding sequence vector samples of each training data set are concatenated to obtain the concatenated vector of each training data set. The concatenated vector of each training data set is then mapped to the vocabulary space to obtain the predicted character encoding sequence of each training data set.
[0169] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0170] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0171] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0172] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0173] In this embodiment, users can edit audio by directly editing the text information corresponding to the audio, eliminating the need for tedious and complex manual editing operations such as cutting and splicing. This allows users without audio editing skills to quickly edit audio, improving the efficiency of audio processing while ensuring its accuracy. Attached Figure Description
[0174] Figure 1 This is a flowchart illustrating an audio processing method provided in some embodiments of this application;
[0175] Figure 2 These are schematic diagrams of the main interface provided in some embodiments of this application;
[0176] Figure 3 This is a schematic diagram of an audio selection interface provided in some embodiments of this application;
[0177] Figure 4 These are schematic diagrams of the RNN-T architecture provided in some embodiments of this application;
[0178] Figure 5 This is a general framework diagram of a Zipformer encoder provided in some embodiments of this application;
[0179] Figure 6 These are schematic diagrams of the structure of a Zipformer block provided in some embodiments of this application;
[0180] Figure 7 This is a schematic diagram of the structure of a Non-Linear Attention module provided in some embodiments of this application;
[0181] Figure 8 This is a schematic diagram of the decoder architecture provided in some embodiments of this application;
[0182] Figure 9 These are schematic diagrams of the architecture of a federated network provided in some embodiments of this application;
[0183] Figure 10 These are visual illustrations of TextGrid files provided in some embodiments of this application;
[0184] Figure 11 These are schematic diagrams of audio editing interfaces provided in some embodiments of this application;
[0185] Figure 12 These are schematic diagrams of audio editing interfaces provided in some embodiments of this application;
[0186] Figure 13These are schematic diagrams of audio editing interfaces provided in some embodiments of this application;
[0187] Figure 14 These are schematic diagrams of audio editing interfaces provided in some embodiments of this application;
[0188] Figure 15 These are schematic diagrams of audio editing interfaces provided in some embodiments of this application;
[0189] Figure 16 This is a schematic diagram of a waiting interface provided in some embodiments of this application;
[0190] Figure 17 These are waveform diagrams of the first audio provided in some embodiments of this application;
[0191] Figure 18 These are waveform diagrams of the second audio provided in some embodiments of this application;
[0192] Figure 19 This is a schematic diagram of the structure of an audio generation model provided in some embodiments of this application;
[0193] Figure 20 These are schematic diagrams of GAN structures provided in some embodiments of this application;
[0194] Figure 21 This is a schematic diagram of the internal structure of the transformer block provided in some embodiments of this application;
[0195] Figure 22 These are schematic diagrams of audio editing interfaces provided in some embodiments of this application;
[0196] Figure 23 These are schematic diagrams of audio processing apparatuses provided in some embodiments of this application;
[0197] Figure 24 These are block diagrams of electronic devices provided in some embodiments of this application;
[0198] Figure 25 These are schematic diagrams of the structure of electronic devices provided in some embodiments of this application. Detailed Implementation
[0199] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0200] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0201] Before providing a further detailed description of the embodiments of the present invention, the terms and concepts involved in the embodiments of the present invention will be explained. The terms and concepts involved in the embodiments of the present invention are applicable to the following interpretations. The terminology used in the implementation section of this application is only used to explain specific embodiments of this application and is not intended to limit this application. The terminology involved in the embodiments of this application will be explained below.
[0202] Fbank: Filter Bank features are a commonly used audio feature extraction method, widely applied in fields such as speech recognition and voiceprint recognition. It extracts the spectral energy distribution characteristics of audio signals through steps such as pre-emphasis, framing, windowing, Fourier transform, Mel filter bank, and logarithmic operations, thereby better simulating the nonlinear perception of sound by the human ear.
[0203] ASR stands for Automatic Speech Recognition. It establishes a mapping between speech signals (time-domain waveforms) and written symbols through acoustic and language modeling. The output of an ASR model is an audio clip, and the output is its corresponding text.
[0204] MFA stands for Montreal Forced Alignment. Given a speech string and its text, MFA can find the time interval corresponding to each character, syllable, and phoneme. It has wide applications in speech recognition, speech synthesis, and other fields.
[0205] Zero-Shot TTS (Zero-Shot Text-to-Speech) is a text-to-speech synthesis technology. It requires no training samples from the target speaker; only short audio and text prompts are needed to clone the speaker's timbre and style and generate speech from any text.
[0206] CTC Loss: Connectionist Temporal Classification Loss. It is a loss function specifically designed for sequence modeling, primarily used to address issues such as inconsistent input and output sequence lengths and alignment uncertainties.
[0207] RNN-T (Recurrent Neural Network Transducer) is an end-to-end deep learning architecture designed for sequence-to-sequence tasks (such as speech recognition). Its core idea is to jointly optimize the acoustic model and the language model to overcome the limitations of traditional CTC models. This architecture generally consists of three parts: 1) Transcription Network (encoder), which typically uses an RNN or CNN to extract acoustic features from the input sequence (such as audio frames); 2) Prediction Network (decoder), which uses a unidirectional RNN to model the language context and generate the hidden states of historical output symbols; 3) Joint Network, which fuses the features of the encoder and decoder and generates the output probability distribution at the current time step through a feedforward network.
[0208] Dropout: During the forward propagation of training, each neuron is kept in an inactive state with a certain probability p, in order to reduce overfitting.
[0209] KL divergence, also known as relative entropy, measures the degree of difference between two probability distributions. A larger KL divergence indicates a greater degree of difference between the two distributions, while a smaller KL divergence indicates a smaller degree of difference.
[0210] Self-Attention Module: The self-attention module is a component used to enable the model to focus on the relationships between different locations in the input data.
[0211] Multi-Head Self-Attention Module: The multi-head attention module is an extension of Self-Attention. It captures the relationships between elements in the sequence from different perspectives by performing multiple group self-attention calculations in parallel, further improving the model's ability to express complex features.
[0212] The feed-forward module is used to perform non-linear transformations and dimension mappings on the input features.
[0213] The BiasNorm module is a novel normalization method. Its core idea is to apply normalization directly after the convolution operation, rather than the traditional sequence of convolution followed by normalization. This structure effectively avoids normalization bias caused by changes in convolutional layer parameters, thereby improving the stability and effectiveness of model training.
[0214] Non-Linear Attention Module: A novel attention module designed to enhance a model's understanding of complex relationships by introducing non-linear computation. Its core idea is to simulate attention weight allocation through a non-linear energy landscape, enabling the model to learn richer contextual relationships.
[0215] The convolution module is a component that extracts local features through convolution operations.
[0216] Bypass module: also known as bypass model or direct connection structure, refers to a mechanism for direct feature transfer, which allows input data to be directly transferred to the output through a "shortcut" while being processed by complex modules, and then fused with the processed features.
[0217] Linear module: Linear layer, refers to a basic component that processes input features through linear transformation. Its core function is to perform linear mapping on the input vector.
[0218] The tanh module is a hyperbolic tangent function, whose core function is to introduce a non-linear transformation into the model, enabling the model to learn the complex non-linear relationships between input features.
[0219] Single-head Attention Module: The single-head attention module is the basic form of self-attention mechanism and a building block of multi-head attention modules. It achieves the fusion of global information of the sequence by calculating the association weights between each element in the sequence and all other elements. Its core feature is "single-group attention calculation".
[0220] The audio processing method, apparatus, electronic device, storage medium, and program product provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.
[0221] The audio processing method provided in this application can be applied to audio editing scenarios. One specific application scenario is editing the audio corresponding to the text information "coming to Beijing to study media studies". The following will illustrate this with... Figures 1-22The audio processing method provided in the embodiments of this application will be described in detail. It should be noted that the audio processing method provided in the embodiments of this application can be executed by an electronic device, which may include, but is not limited to, mobile phones, tablets, desktop computers, etc. This application embodiment uses an electronic device executing the audio processing method as an example to illustrate the audio processing method provided in the embodiments of this application.
[0222] See Figure 1 The above is a flowchart illustrating some embodiments of the audio processing method provided in this application, such as... Figure 1 As shown, the method includes steps 110-130, which will be explained in detail below.
[0223] Step 110. Receive text editing input for the first text information of the first audio.
[0224] In this embodiment of the application, the first audio is the audio that needs to be processed.
[0225] In some embodiments of this application, the first audio file may be selected by the user from audio stored in the electronic device according to the actual situation. Optionally, before step 110 above, the electronic device may display an audio selection interface, receive the user's selection input for the first audio file in the audio selection interface, and, in response to the selection input, treat the first audio file as the audio to be processed. The audio selection interface may include at least one audio file for the user to select.
[0226] For example, see Figure 2 When the user needs to perform audio processing, the electronic device can display something like this. Figure 2 The main interface 200 shown includes an upload control 201. Upon user click on the upload control 201, the electronic device displays the following... Figure 3 The audio selection interface 300 shown includes at least one audio file stored in the electronic device. The user can select the audio file to be edited from the audio selection interface 300 according to their needs. In response to the user's selection operation in the audio data interface 300, the electronic device designates the selected audio file as the first audio file to be processed. For example, if the user selects Audio1.wav from the audio selection interface 300, then Audio1.wav will be designated as the first audio file. This allows the user to select the first audio file to be processed according to their actual needs, improving the flexibility of audio processing.
[0227] The first text information of the first audio is the textual representation of the text content corresponding to the first audio. For example, if the text content of the first audio is "coming to Beijing to study media", then the first text information is the textual representation of "coming to Beijing to study media".
[0228] The text editing input is used to edit the first text information. When the user needs to edit the first audio, they can edit the first text information through the text editing input, thereby indirectly achieving the editing of the first audio.
[0229] In some embodiments of this application, to enable users to edit the first text information more intuitively, optionally, before step 110 above, the electronic device may first display an audio editing interface. The audio editing interface includes at least the first text information of the first audio. Thus, step 110 above may specifically include: receiving text editing input for the first text information in the audio editing interface. The audio editing interface is an interface used to edit the audio to be processed, allowing users to intuitively edit the first text information.
[0230] Step 120. In response to text editing input, update the first text information to obtain the second text information.
[0231] In this embodiment, the text editing input is input for editing the first text information. Editing the first text information may include, but is not limited to, deletion, addition, and replacement. Exemplarily, the text editing input may include at least one of the following: editing input that deletes a first character sequence from the first text information, editing input that inserts a second character sequence into the first text information, and editing input that replaces a third character sequence in the first text information with a fourth character sequence. The first and third character sequences can be sequences composed of any characters in the first text information, while the second and fourth character sequences can be any character sequences that the user actually needs to input.
[0232] For example, taking the first text message as "came to Beijing to study media," the text editing input could be to delete the word "go" from the first text message. It could also be to add the character sequence "I'm very happy" to the end of the first text message. Furthermore, it could be to replace "Beijing" with "Nanjing" in the first text message.
[0233] When an electronic device receives text editing input from a user for first text information, it edits the first text information based on the text editing input to obtain second text information.
[0234] For example, if the first text message is "came to Beijing to study media", and the text editor deletes the word "go" from the first text message, the resulting second text message will be "came to Beijing to study media".
[0235] For example, if the first text message is "I came to Beijing to study media", and the text editing input is to add the character sequence "I am very happy" to the end of the first text message, the resulting second text message will be "I am very happy to come to Beijing to study media".
[0236] For example, if the first text message is "came to Beijing to study media", and the text editor input is to replace "Beijing" with "Nanjing" in the first text message, the resulting second text message will be "came to Nanjing to study media".
[0237] Step 130. Based on the second text information, perform audio processing on the first audio to obtain the second audio.
[0238] In this embodiment of the application, after obtaining the second text information, the first audio can be adjusted accordingly based on the difference between the first and second text information to obtain the second audio. The second text information is a textual representation of the text content corresponding to the second audio.
[0239] In some embodiments of this application, when the text editing input is an editing input that deletes the first character sequence in the first text information, the first character sequence to be deleted in the first text information can be determined based on the difference between the first text information and the second text information. Based on this, when performing audio processing on the first audio based on the second text information, the audio segment corresponding to the first character sequence in the first audio can be deleted to obtain the second audio.
[0240] In some embodiments of this application, when the text editing input is an editing input that adds a second character sequence to the first text information, the second character sequence added to the first text information can be determined based on the difference between the first text information and the second text information. Based on this, when performing audio processing on the first audio based on the second text information, an audio segment corresponding to the second character sequence can be generated, and the audio segment can be added to the first audio to obtain the second audio.
[0241] In some embodiments of this application, when the text editing input is an editing input that replaces the third character sequence in the first text information with the fourth audio segment, the third character sequence to be replaced in the first text information and the fourth character sequence used to replace the third character sequence can be determined based on the difference between the first text information and the second text information. Based on this, when performing audio processing on the first audio based on the second text information, an audio segment corresponding to the fourth character sequence can be generated, and the audio segment corresponding to the third character sequence in the first audio can be replaced by the audio segment, thereby obtaining the second audio.
[0242] In this embodiment, users can edit audio by directly editing the text information corresponding to the audio, eliminating the need for tedious and complex manual editing operations such as cutting and splicing. This allows users without any audio editing experience to quickly edit audio, improving the efficiency of audio processing while ensuring its accuracy.
[0243] In some embodiments, prior to step 110 above, the electronic device may first perform steps 101-103 as follows.
[0244] Step 101. Input the first audio into the speech recognition model, and use the speech recognition model to perform speech recognition on the first audio to obtain the first text information of the first audio.
[0245] In this embodiment of the application, the first audio is obtained by performing speech recognition on the first audio through a speech recognition model.
[0246] A speech recognition model is an artificial intelligence model used to convert audio into text information. It can recognize the semantic content in audio and output the semantic content in text form. Based on this, after acquiring the first audio, the speech recognition model can recognize the text content of the first audio, generate a text representation of the text content and output it. The output result of the speech recognition model is determined as the first text information corresponding to the first audio.
[0247] In some embodiments of this application, optionally, before step 101 above, an initial model can be trained based on a sample dataset and a loss function to obtain a speech recognition model. The sample dataset includes multiple audio samples and corresponding text information samples. The multiple audio samples include audio recorded in at least two scenarios, and the audio samples include human voice features. Here, audio samples refer to the audio used for model training, and the corresponding text information refers to the text representation of the text content corresponding to the audio sample. The audio samples include human voice features, meaning they contain human speech content, which is necessary to obtain the text information of the audio samples.
[0248] In some embodiments of this application, to improve the recognition accuracy and robustness of the trained speech recognition model, a large amount of diverse audio data can be acquired as audio samples when constructing the sample dataset. For example, audio samples can be constructed from approximately 100,000 hours of audio recordings containing different speakers outputting different text content in different environments. The speakers can cover different ages, genders, accents, speaking speeds, pronunciation habits, etc., and the environments can include, but are not limited to, quiet places, offices, streets, shopping malls, etc. The text content can include, but is not limited to, news releases, novel excerpts, and daily conversations. For example, when acquiring audio samples, a text set can be constructed as needed. Then, multiple different speakers read the text in the text set aloud in different environments and record the corresponding audio. The recorded audio is used as the audio sample, and the text in the text set corresponding to the recorded audio is used as the text information sample corresponding to the audio sample. For example, when acquiring audio samples, audio from different speakers can be recorded in different environments, or audio from different speakers can be downloaded from the internet. The recorded or downloaded audio can be used as audio samples. Then, the text content of the audio samples can be manually identified, and the corresponding text can be recorded. The recorded text can be used as the text information sample of the corresponding audio sample.
[0249] In some embodiments of this application, to improve the robustness and generalization ability of the model, the aforementioned loss function may employ a joint loss function. The joint loss function includes at least two loss function terms, each corresponding to a weight, and one of the at least two loss function terms includes forward relative entropy. Based on this, after obtaining the sample dataset, the initial model can be trained through the following steps 310-390 to obtain the speech recognition model.
[0250] Step 310. Extract acoustic feature samples for each audio sample in the sample dataset.
[0251] Acoustic feature samples refer to acoustic features extracted from audio samples. Acoustic features include quantitative features extracted from the waveform of the audio that characterize the physical properties of the sound. For example, acoustic features may include the F-bank (Filter Bank) features of the audio.
[0252] In some embodiments of this application, optionally, to improve model training efficiency and model generalization ability, before extracting acoustic feature samples from audio samples, the audio samples can be resampled to a preset frequency and then segmented into frames. Then, acoustic feature extraction is performed on the resampled and segmented audio samples to obtain acoustic feature samples. The preset frequency can be set according to actual needs; for example, the preset frequency can be 16kHz, and the preset length can be 25ms, without specific limitations.
[0253] In some embodiments of this application, audio samples can be segmented into frames using window functions according to preset frame shifts and frame lengths. Window functions are used for weighted processing when segmenting continuous signals into frames. Window functions can "weightedly attenuate" the segmented audio signal (frames) to reduce the negative impact of abrupt changes in the signal at frame edges and to reduce problems such as spectral leakage. Exemplarily, window functions may include, but are not limited to, the Povey window, Hamming window, Hanning window, Rectangular window, Sine window, and Blackmann window.
[0254] For example, an audio sample can be resampled to 16kHz and then framed using a Povy window with a frame shift of 10ms and a frame length of 25ms. The 80-dimensional Fbank features of the resampled and framed audio samples can then be extracted and used as the acoustic feature samples of the audio samples.
[0255] Step 320. Perform word segmentation on the text information sample corresponding to each audio sample to obtain the character encoding sequence sample of each audio sample.
[0256] A character encoding sequence sample refers to the character encoding sequence of a text information sample. By performing word segmentation on the text information sample, it is divided into multiple discrete characters, i.e., tokens. A unique code corresponding to each token is determined based on a vocabulary. The sequence of unique codes corresponding to multiple discrete tokens is the character encoding sequence of the text information sample. The vocabulary indicates the predefined mapping relationship between tokens and unique codes. The vocabulary is a lookup table that maps a token to a unique code. For example, the text information sample "Today the weather is nice" can be split into three words: "today," "weather," and "nice." The unique codes corresponding to these three words in the vocabulary are "250," "251," and "252," respectively. Therefore, the character encoding sequence corresponding to this text information sample is [250, 251, 252].
[0257] In some embodiments, text information samples can be segmented using BPE (Byte-Pair Encoding) to obtain corresponding character encoding sequence samples.
[0258] Step 330. Based on the acoustic feature samples and character encoding sequence samples of all audio samples in the sample dataset, construct multiple sets of training data. Each set of training data includes the acoustic feature samples and character encoding sequence samples of one audio sample.
[0259] For each audio sample in the sample dataset, the corresponding acoustic feature sample and character encoding sequence sample are combined into a set of training data. In this way, multiple sets of training data can be constructed based on the sample dataset.
[0260] Step 340. Based on multiple sets of training data, train the initial model to obtain the predicted character encoding sequence corresponding to each set of training data.
[0261] In this embodiment, the initial model refers to the basic model structure at the beginning of the training process. A suitable model can be selected as the initial model according to actual needs.
[0262] In some embodiments, an ASR (Automatic Speech Recognition) model can be used as the initial model; see [link to relevant documentation]. Figure 4 The initial model can adopt an RNN-T (Recurrent Neural Network Transducer) architecture, which includes an encoder 401, a decoder 402, and a joint network 403. Based on this, step 340 above can specifically include the following steps 3401-3403.
[0263] Step 3401. Using an encoder, embed the acoustic feature samples from each set of training data into a spatial vector to obtain the acoustic feature vector samples of each set of training data.
[0264] In this embodiment of the application, the encoder is used to convert the acoustic features in the input into corresponding vector representations, that is, to convert the acoustic features into acoustic feature vectors.
[0265] When training the initial model, for each set of training data, acoustic feature samples from the training data can be input into the encoder. The encoder embeds the acoustic feature samples into a spatial vector, thus obtaining the vector representation of the acoustic feature samples output by the encoder, i.e., the acoustic feature vector sample. Embedding is a technique that converts discrete data (such as text sequences, acoustic features, etc.) into low-dimensional continuous vectors.
[0266] In some embodiments of this application, a Zipformer encoder may be used to achieve multi-scale feature extraction. Unlike traditional techniques that use Conformer to embed acoustic features, the Zipformer encoder adopts a multi-resolution encoding stack structure similar to U-Net, achieving multi-scale feature extraction through hierarchical downsampling and upsampling. In contrast, embedding acoustic features using Conformer can only operate at a fixed frame rate of 25Hz. Conformer is a hybrid network structure that combines the advantages of Convolutional Neural Networks (CNN) and Transformers.
[0267] See Figure 5 This is a general framework diagram of the Zipformer encoder, which includes a convolutional embedding layer (Conv-Embed) and multiple consecutive encoder stacks. The number of encoder stacks can be set according to actual needs and is not specifically limited. Figure 5 This example uses a 6-encoder stack. Conv-Embed is used to downsample the input data (such as acoustic features). The encoder stack is used for temporal modeling of the input data; different encoder stacks can perform temporal modeling at different sampling rates. Temporal modeling refers to learning the temporal features of the input data, including capturing short-term details and long-term dependencies. In the multiple consecutive encoder stacks, except for the first one, all other encoder stacks employ a downsampling structure. Specifically, the first encoder stack includes a feature extraction module, i.e., a Zipformer block structure. Other encoder stacks, in addition to the Zipformer block structure, also include a downsampling module (Downsample) placed before the Zipformer block structure and an upsampling module (Upsample) placed after the Zipformer block structure. Furthermore, a bypass module (Bypass) is set between every two adjacent encoder stacks.
[0268] During the initial model training process, after inputting acoustic feature samples into the Zipformer encoder, the input acoustic feature samples are first downsampled using Conv-Embed to obtain a feature sequence. Then, multiple consecutive encoder stacks are used to perform temporal modeling on the features at different sampling rates, ultimately yielding a vector of acoustic feature samples, i.e., acoustic feature vector samples. Temporal modeling refers to learning the temporal dimension features of the input temporal data, including capturing short-term details and long-term dependencies.
[0269] For example, taking an acoustic feature sample with a frequency of 100Hz as an example, the acoustic feature sample is input as follows: Figure 5 After the Zipformer encoder shown, the input 100Hz acoustic feature samples are downsampled to a 50Hz feature sequence via Conv-Embed. Then, temporal modeling is performed by six consecutive encoder stacks at sampling rates of 50Hz, 25Hz, 12.5Hz, 6.25Hz, 12.5Hz, and 25Hz, respectively. Except for the first encoder stack, all other encoder stacks employ a downsampling structure. Between encoder stacks, the feature sequence sampling rate remains at 50Hz. Different encoder stacks have different embedding dimensions, with the middle encoder stack having a larger embedding dimension. The output of each encoder stack is truncated or padded with zeros to align with the dimension of the next encoder stack. The final output dimension of the Zipformer depends on the encoder stack with the largest embedding dimension.
[0270] See Figure 6 This is a schematic diagram of the Zipformer block, a feature extraction module, as shown below. Figure 6As shown, the Zipformer block includes a Multi-Head Self-Attention module, a Feed-forward module, a Non-Linear Attention module, two module groups, a BiasNorm module, and a Bypass module. Each module group contains a Self-Attention module, a Convolution module, and a Feed-forward module. Data input to the Zipformer block is first fed to the Multi-Head Self-Attention module to calculate attention weights, which are then shared with the Non-Linear Attention module and the two Self-Attention modules. Simultaneously, the data input to the Zipformer block is also fed to the Feed-forward module, followed by the Non-Linear Attention module. This is followed by two consecutive module groups. Finally, a BiasNorm module normalizes the output data of the Zipformer block. In addition to the standard additive residual connections, each Zipformer block uses two Bypass modules to combine the Zipformer block input and the outputs of intermediate modules; these are located in the middle and at the end of the Zipformer block, respectively.
[0271] See Figure 7 This is a schematic diagram of the Non-Linear Attention module. The structure of the Non-Linear Attention module is similar to that of the Self-Attention module. The Non-Linear Attention module utilizes the attention weights calculated by the Multi-Head Self-Attention module and converges the vectors from different frames along the time axis. Figure 7 As shown, the Non-Linear Attention module includes a linear module, a tanh module, and a Single-head Attention module.
[0272] The Zipformer encoder reduces sequence length and computational complexity through hierarchical downsampling, resulting in approximately 30% faster inference speed compared to the Conformer. It enhances adaptability to noisy environments and accent variations through multi-scale feature fusion and weighted attention mechanisms. Furthermore, it supports FP16 training without loss of accuracy, making it suitable for large-scale distributed training.
[0273] Step 3402. Using a decoder, embed the character encoding sequence samples in each group of training data into a spatial vector to obtain the character encoding sequence vector samples of each group of training data.
[0274] In this embodiment of the application, the decoder is used to convert the character encoding sequence in the input into a corresponding vector representation, that is, to convert it into a character encoding sequence vector.
[0275] When training the initial model, the character-encoded sequence samples in each set of training data can be input into the decoder. The decoder embeds each character-encoded sequence sample into a spatial vector, thereby obtaining the vector representation of each character-encoded sequence sample output by the decoder, i.e., the character-encoded sequence vector sample.
[0276] In some embodiments of this application, to reduce the computational cost of model training and improve training and inference speed, the decoder may adopt a stateless architecture, i.e., it does not contain RNN layers and only uses convolutional layers. See also Figure 8 This is a schematic diagram of the decoder architecture. For example... Figure 8 As shown, the decoder includes an embedding layer, a first-channel regularization module, a grouped convolutional module, a non-linear activation function, and a second-channel regularization module. The non-linear activation function can be ReLU (Rectified Linear Unit). Based on this, during the initial model training, after inputting the character-encoded sequence samples into the encoder, the embedding layer maps the input character-encoded sequence samples to continuous vector embeddings. Then, the first-channel regularization module constrains the numerical range of the embeddings to prevent gradient anomalies. The grouped convolutional module captures the local context of the data output by the first-channel regularization module. If the context size is set to 2, the kernel size of the grouped convolutional module is equal to 2. The grouped convolutional module divides the input and output channels into multiple independent groups, effectively reducing the number of parameters and computational complexity. After the data output by the grouped convolutional module is processed by the non-linear activation function, and then the second-channel regularization module constrains the numerical range of the embeddings, the final output, i.e., the character-encoded sequence vector sample, is obtained.
[0277] Step 3403. Through the joint network, the acoustic feature vector samples and character encoding sequence vector samples of each training data are concatenated to obtain the concatenated vector of each training data, and the concatenated vector of each training data is mapped to the vocabulary space to obtain the predicted character encoding sequence of each training data.
[0278] In this embodiment, the role of the joint network is to combine the output of the decoder and the output of the encoder to obtain the predicted character encoding sequence of the training data.
[0279] In some embodiments of this application, see Figure 9 The joint network comprises a first linear projection layer, a second linear projection layer, a nonlinear activation function, and a fully connected output layer. The nonlinear activation function can be the hyperbolic tangent function tanh. The first and second linear projection layers linearly transform the outputs of the decoder and encoder, respectively, unifying them to the same dimension. The nonlinear activation function performs a nonlinear transformation on the sum of the outputs from the two linear projection layers. The fully connected output layer maps the features output by the nonlinear activation function to the vocabulary space. Based on this, during the initial model training, for each set of training data, the acoustic feature vector samples from the encoder and the character encoding sequence vector samples from the decoder are input into the joint network. The first linear projection layer linearly transforms the acoustic feature vector samples to obtain a first transform vector. The second linear projection layer linearly transforms the character encoding sequence vector samples to obtain a second transform vector. The first and second transform vectors are summed to obtain an initial fused transform vector. The nonlinear activation function performs a nonlinear transformation on the initial fused transform vector to obtain a fused vector. The fully connected output layer maps the fused vector to the vocabulary space, thus obtaining the predicted character encoding sequence for that set of training data.
[0280] Step 350. After completing at least two rounds of training on the initial model, calculate the forward relative entropy based on the predicted character encoding sequence obtained in the i-th round of training and the predicted character encoding sequence obtained in the (i-1)-th round of training.
[0281] In the embodiments of this application, i is an integer greater than or equal to 2.
[0282] When training an initial model based on multiple sets of training data, multiple training rounds can be performed to improve the model's accuracy. One training round is called 1 Epoch, which refers to the process by which the model learns all the training data completely once.
[0283] In related technologies, during model training, a loss value is calculated using a loss function after each training round, and the model parameters are adjusted based on the calculated loss value. To prevent overfitting, Dropout is typically used during training. However, Dropout can cause inconsistent outputs for the same input across multiple forward propagations, potentially leading to model instability, especially with smaller datasets or during fine-tuning. Therefore, to prevent overfitting and improve the robustness of the speech recognition model, this embodiment employs a joint loss function to calculate the loss value for the initial model. This joint loss function introduces forward relative entropy, i.e., R-Drop loss (Regularized Dropout Loss), which uses the KL divergence between the distributions of the two outputs after two forward propagations—the forward relative entropy—as an additional regularization term to encourage consistent outputs under Dropout. Based on this, in this embodiment of the application, when training the initial model, after completing at least two rounds of training of the initial model, the forward relative entropy of the model is calculated based on the two predicted character encoding sequences obtained from the i-th round of training and the (i-1)-th round of training. Here, the i-th round of training refers to the latest round of training, and the (i-1)-th round of training refers to the round preceding the i-th round of training.
[0284] In some embodiments of this application, the forward relative entropy of the model can be calculated using the following formula (1).
[0285]
[0286] In the above formula (1), L KL P represents the forward relative entropy, also known as the KL divergence. i =softmax(z i ), z i Let P represent the predicted character encoding sequence obtained in the i-th round of training. i-1 =softmax(z i-1 ), z i-1 Let represent the predicted character encoding sequence obtained in the (i-1)th round of training, softmax(·) represents the normalization exponential function, and KL(·) represents the KL calculation function.
[0287] Step 360. Based on the forward relative entropy, update the entropy value of the forward relative entropy in the loss function to obtain the updated loss function.
[0288] In this embodiment of the application, the loss function is a joint loss function, which includes at least two loss function terms. One of the at least two loss function terms is the forward relative entropy. Based on this, after obtaining the forward relative entropy based on the latest training results through the above step 350, the forward relative entropy in the joint loss function is replaced with the latest calculated forward relative entropy, thereby updating the loss function and obtaining the updated loss function.
[0289] In some embodiments of this application, the loss function terms other than the forward relative entropy in the joint loss function can be adapted to the initial model according to the structure settings of the initial model.
[0290] In some embodiments of this application, when the initial model adopts an RNN-T structure, the first loss function can be used as a loss function term in the joint loss function. Here, the first loss function is the standard RNN-T loss Simpleloss, and its purpose is to maximize the total probability of all valid alignment paths. The acoustic feature samples are denoted as x = (x1, x2, ..., x...). T The character encoded sequence sample is denoted as y = (y1, y2, ..., y3). U The goal of RNN-T is to model the joint probability that satisfies the following formula (2):
[0291]
[0292] In formula (2) above, P(y|x) represents the joint probability, π represents the alignment path with blank, which is all possible ways to align the input acoustic feature samples to the output predicted character encoding sequence, with a length of T+U, B -1 (y) represents the set of all paths that map to the character-encoded sequence sample y.
[0293] The acoustic feature vector sample output by the Zipformer encoder is represented by the following formula (3):
[0294] h=Encoder(x) ∈T×d (3)
[0295] The character encoded sequence vector sample output by the decoder is represented by the following formula (4):
[0296] g=Decoder(y) ∈U×d (4)
[0297] The predicted character encoding sequence output by the joint network is represented by the following formula (5):
[0298] z t,u =Joiner(h t ,g u)∈|V|+1 (5)
[0299] In the above formula (5), V represents the vocabulary and +1 represents blank, i.e., a blank space.
[0300] The simple loss is defined as shown in equation (6):
[0301]
[0302] In the above formula (6), L simple This represents the loss value calculated based on simple loss, and A represents the set of valid paths.
[0303] In the simple loss, each predicted character encoding sequence z t,u All of these require calculating the probability of all words in the sequence, i.e., tokens. This method is too memory-intensive and computationally intensive, especially for large vocabularys and long sequences. In view of this, in order to reduce the memory required to calculate the loss value and improve computational efficiency, the joint loss function can optionally include a second loss function in addition to the first loss function. The second loss function is the pruning loss function, i.e., pruned loss. Pruned loss is obtained by pruning the simple loss. In pruned loss, only the Top-K output tokens are retained at each position (t,u). A subset alignment path graph is constructed and approximate summation is performed. Pruned loss can be expressed as the following formula (7):
[0304]
[0305] In the above formula (7), L pruned A represents the loss value calculated based on pruned loss. pruned ∈A represents the set of valid paths that contain only Top-K tokens.
[0306] In actual training, corresponding weights can be set for simpleloss, prunedloss, and forward relative entropy respectively. Simpleloss, prunedloss, and forward relative entropy are used together based on the weights, and the weights are dynamically adjusted during training. For example, at the beginning of training, simple loss is mainly optimized. In the warm-up phase, the weight of simple loss can be gradually reduced and the weight of pruned loss can be increased. Let t be the current training step number, w be the warm-up step number, and s be the initial weight of simple loss. Then the weight calculation formula of simple loss can be expressed as the following formula (8), and the weight calculation formula of pruned loss can be expressed as the following formula (9):
[0307]
[0308] In the above formula (8), scale simpleloss The scale represents the weights of the simple loss. prunedloss The values of w and s represent the weights of the pruned loss. In actual training, the values of w and s can be set according to actual needs. For example, w can be set to 2000 and s can be set to 0.5.
[0309] When the preset loss functions include simple loss and pruned loss, the loss function updated based on forward relative entropy can be expressed as the following formula (10):
[0310]
[0311] In the above formula (10), L total This represents the final calculated loss value, and λ represents the hyperparameter. The value of λ determines the weight of R-Drop Loss. During model pre-training, the value range of λ can be [0.1, 0.3], and during model fine-tuning, the value range of λ can be [0.2, 0.4].
[0312] Step 370. Calculate the loss value of the initial model based on the updated loss function.
[0313] Step 380. If the loss value is greater than or equal to the loss threshold, adjust the weight of each loss function term in the loss function, and continue to train the initial model with adjusted parameters based on multiple sets of training data.
[0314] In this embodiment, the loss threshold is used to determine whether the trained model meets the accuracy requirements. The specific value of the loss threshold can be set according to actual needs, and there is no specific limitation on it.
[0315] After calculating the model's loss value based on the updated loss function, the loss value is compared with the loss threshold. If the loss value is greater than or equal to the loss threshold, it means that the current accuracy of the model does not meet the accuracy requirements. Therefore, the process returns to step 340 above to continue training the model until the model's loss value is less than the loss threshold.
[0316] Step 390. If the loss value is less than the loss threshold, stop training and obtain the trained speech recognition model.
[0317] If the loss value is less than the loss threshold, it means that the accuracy of the current model meets the accuracy requirements, so training stops and the model trained at this time is determined as the speech recognition model.
[0318] In the training process of the above model, the introduction of R-Drop Loss, namely forward relative entropy, can make the encoder output more stable under dropout, prevent the output of the joint network from being too biased, improve the generalization ability of the model, effectively alleviate the overfitting of the model under small data volume without affecting the inference speed, improve the robustness and generalization ability of the model, and improve the speech recognition accuracy of the model.
[0319] In some embodiments of this application, the trained speech recognition model can be deployed in an electronic device, and the electronic device can directly perform speech recognition on the first audio through its own deployed speech recognition model.
[0320] In some other embodiments of this application, in order to reduce the memory space occupied by the model on the electronic device, the trained speech recognition model can be deployed on a cloud server. Based on this, the electronic device can upload the first audio to the cloud server, so that the cloud server can use the deployed speech recognition model to perform speech recognition on the first audio to obtain the first text information corresponding to the first audio. Then, the cloud server transmits the first text information to the electronic device, so that the electronic device can obtain the first text information.
[0321] Step 102. Input the first audio and the first text information into the audio-text alignment model. Through the audio-text alignment model, perform time alignment on the first audio and the first text information to obtain the time mapping relationship between the first text information and the first audio. The time mapping relationship is used to indicate the start and end times of the audio segment corresponding to each character in the first text information in the first audio.
[0322] In this embodiment, the audio-text alignment model is a model used for time alignment of audio and its corresponding text information. Based on the audio-text alignment model, the input audio and text information can be aligned in time sequence to determine the position information of each character in the text information within the audio, i.e., the start and end times of the audio segment corresponding to each character. Thus, the time mapping relationship between audio and text can be obtained.
[0323] In this embodiment, the first audio and the first text information are input into an audio-text alignment model for time alignment, thereby obtaining the time mapping relationship between the first text information and the first audio. The time mapping relationship represents the start and end times of the audio segment corresponding to each character in the first text information in the first audio.
[0324] In some embodiments of this application, the audio-text alignment model may employ the MFA model. The MFA model is used to align audio and Chinese characters, obtaining the start and end times of the audio segment corresponding to each character in the audio. Inputting the first audio and first text information into the MFA model yields a TextGrid file output by the MFA model corresponding to the first audio. The TextGrid file is a file format used by Praat software to store audio annotations, which are typically used to represent the time boundaries of phonemes, words, phrases, or other speech components. The TextGrid file typically contains two levels: a phoneme layer and a word layer. The phoneme layer annotates the precise time boundaries of each phoneme, with a time accuracy down to the millisecond level, and includes silent segment markers. The word layer annotates word-level alignment information (such as pinyin or Chinese characters), including the start position of each character in the audio.
[0325] For example, taking the text content corresponding to the first audio as "came to Beijing to study media", the visualization diagram of the TextGrid file output by the MFA model is as follows. Figure 10 As shown.
[0326] In some embodiments of this application, the audio-text alignment model can be deployed in an electronic device, and the electronic device can directly perform time alignment of the first audio and the first text information through its own deployed audio-text alignment model.
[0327] In some other embodiments of this application, in order to reduce the memory space occupied by the model on the electronic device, the audio-text alignment model can be deployed on a cloud server. Based on this, the electronic device can upload the first audio and the first text information to the cloud server, so that the cloud server can use the deployed audio-text alignment model to perform time alignment on the first audio and the first text information to obtain the time mapping relationship between the first text information and the first audio. Then, the cloud server transmits the time mapping relationship to the electronic device, so that the electronic device can obtain the time mapping relationship between the first text information and the first audio.
[0328] In some other embodiments of this application, if both the audio text alignment model and the speech recognition model are deployed on a cloud server, the electronic device only needs to transmit the first audio to the cloud server. The cloud server can input the first audio into the speech recognition model, and input the first text information output by the speech recognition model into the audio text alignment model to perform time alignment with the first audio.
[0329] Step 103. Display the audio editing interface, which includes a timeline, a waveform of the first audio, and first text information. The timeline includes at least one time marker, and each character in the first text information is aligned with the corresponding waveform and the corresponding time marker in the waveform diagram.
[0330] In this embodiment, after acquiring the first audio, the first text information, and the time mapping relationship, the electronic device renders and displays an audio editing interface based on the first audio, the first text information, and the time mapping relationship, thereby presenting the first text information and the time-series mapping relationship between the first audio and the first text information to the user through the audio editing interface. The audio editing interface includes a timeline, a waveform of the first audio, and the first text information. The timeline includes at least one time marker, and each character in the first text information is aligned with the corresponding waveform and the corresponding time marker in the waveform diagram.
[0331] For example, such as Figure 11 As shown, the audio editing interface 1100 includes a timeline 1101, a waveform of the first audio 1102, and first text information 1103, with the waveform 1102 and the first text information 1103 aligned on the timeline 1101.
[0332] In some embodiments of this application, the audio editing interface may further include controls for adjusting the content displayed in the audio editing interface. This allows the user to adjust the displayed content in the audio editing interface when the first audio content is long and the waveform cannot be fully displayed at once, enabling the user to view all the complete waveforms and the first text information. For example, as... Figure 11 As shown, the audio editing interface 1100 also includes a control 1104. When the audio editing interface 1100 cannot display all the waveforms and first text information, the user can slide the control 1104 to display the undisplayed waveforms and text information.
[0333] This application provides a novel, intuitive, and efficient interaction method. By using a speech recognition model and an audio-text alignment model, the start and end times of the audio segment corresponding to each character in the first text information of the first audio are obtained in the first audio. Users can directly edit the first text information in the audio editing interface to edit the first audio, thereby improving audio editing efficiency while ensuring accuracy.
[0334] In some embodiments, after the first audio is determined, the electronic device may display, during the execution of steps 101-102, the following: Figure 12The waiting interface 1200 shown includes a loading animation 1201 and a prompt message "Audio parsing in progress, please wait...". By displaying the waiting interface 1200, the waiting anxiety of the user can be alleviated, and the experience of the user during the waiting process can be optimized.
[0335] For different text editing inputs, different strategies can be adopted to obtain the second text information and the second audio. The strategies for obtaining the second text information and the second audio corresponding to each text editing input will be described below.
[0336] In some embodiments, in the case where the text editing input is an editing input for deleting the first character sequence in the first text information, the above step 120 may specifically include: deleting the first character sequence in the first text information to obtain the second text information.
[0337] The first character sequence may be a continuous or discontinuous sequence composed of one or more characters to be deleted in the first text information.
[0338] In some embodiments of the present application, in the case where the first character sequence is a continuous character in the first text information, that is, the first character sequence is a continuous sequence composed of at least one character in the first text information, the above step 130 may specifically include the following steps 1311-step 1313.
[0339] Step 1311. Compare the second text information with the first text information to determine the first character sequence deleted in the first text information.
[0340] As recorded above, in the case where the text editing input is an editing input for deleting the first character sequence in the first text information, in response to the text editing input, the first character sequence is deleted from the first text information to obtain the second text information. Thus, by comparing the first file information and the second text information, the first character sequence deleted in the first text information can be determined.
[0341] Exemplarily, the first text information is "Come to Beijing to study media major", and the second text information is "Come to Beijing study media major". By comparing the two, it can be determined that the first character sequence deleted in the first text information is the character "go".
[0342] Step 1312. Based on the time mapping relationship, determine the segment start time and the segment end time of the first audio segment corresponding to the first character sequence.
[0343] In the embodiments of the present application, the first audio segment is the audio segment in the first audio corresponding to the first character sequence. For example, if the first character sequence is the character "go", the first audio segment is the audio segment in the first audio with the text content "go".
[0344] The start time of the first audio segment is the start time of the audio segment in the first audio corresponding to the first character in the first character sequence, and the end time of the first audio segment is the end time of the audio segment in the first audio corresponding to the last character in the first character sequence. For example, if the first character sequence is the character "去", the start time of the first audio segment is the start time of the audio segment corresponding to the character "去", and the end time of the first audio segment is the end time of the audio segment corresponding to the character "去".
[0345] Step 1313. Cut the first audio based on the start time and end time of the first audio segment to obtain the first audio segment.
[0346] In the embodiments of the present application, after determining the start time and end time of the first audio segment, at least one cut point of the first audio can be determined based on the start time and end time of the first audio segment, and the first audio segment can be cut based on the at least one cut point, so as to obtain the first audio segment.
[0347] In some embodiments of the present application, when the start time of the first audio segment is the start time of the first audio, the end time of the first audio segment can be used as a cut point to cut the first audio, so as to obtain two audio segments. Among them, the audio segment ranked first in chronological order is the first audio segment. For example, if the first text information is "来到北京去学习传媒专业" and the first character sequence is "来到北京", then the end time of the audio segment corresponding to the character "京" is used as a cut point to cut the first audio, so as to obtain two audio segments. The text contents of the two audio segments are "来到北京" and "去学习传媒专业" respectively. Among them, the audio segment ranked first in chronological order is the audio segment with the text content of "来到北京", and this audio segment is the first audio segment corresponding to the first character sequence.
[0348] In some embodiments of the present application, when the segment end time of the first audio segment is the audio end time of the first audio, the segment start time of the first audio segment can be used as a cut point to cut the first audio, thereby obtaining two audio segments. Among them, the audio segment ranked second in chronological order is the first audio segment. For example, if the first text information is "Come to Beijing to study media major", and the first character sequence is "study media major", then the start time of the audio segment corresponding to the character "学" is used as the cut point to cut the first audio, thereby obtaining two audio segments. The text contents of the two audio segments are "Come to Beijing" and "study media major" respectively. Among them, the audio segment ranked second in chronological order is the audio segment with the text content of "study media major", and this audio segment is the first audio segment corresponding to the first character sequence.
[0349] In some embodiments of the present application, when the segment start time of the first audio segment is not the audio start time of the first audio and the segment end time of the first audio segment is not the audio end time of the first audio, the segment start time and the segment end time of the first audio segment can be used as a cut point respectively, and then two cut points are obtained. Based on the two cut points, the first audio is cut, thereby obtaining three audio segments. Among them, the audio segment ranked second in chronological order is the first audio segment. For example, if the first text information is "Come to Beijing to study media major", and the first character sequence is "去", then the start time and the end time of the audio segment corresponding to the character "去" are used as cut points to cut the first audio, thereby obtaining three audio segments. The text contents of the three audio segments are "Come to Beijing", "去", and "study media major" respectively. Among them, the audio segment ranked second in chronological order is the audio segment with the text content of "去", and this audio segment is the first audio segment corresponding to the first character sequence.
[0350] Step 1314. Delete the first audio segment in the first audio. When a second audio segment is obtained, determine the second audio segment as the second audio.
[0351] In an embodiment of the present application, the second audio segment is the remaining audio segment after deleting the first audio segment in the first audio.
[0352] As described above, by cutting the first audio, at least two audio segments can be obtained. Based on this, when two audio segments are obtained by cutting, after deleting the first audio segment, only one audio segment remains. At this time, this remaining audio segment is used as the second audio. For example, if the text content of the first audio is "Come to Beijing to study media major", and the first character sequence is "Come to Beijing", then after deleting the first audio segment, the audio segment corresponding to the remaining text content "Study media major" is used as the second audio. Another example, if the text content of the first audio is "Come to Beijing to study media major", and the first character sequence is "Study media major", then after deleting the first audio segment, the audio segment corresponding to the remaining text content "Come to Beijing" is used as the second audio.
[0353] Exemplarily, refer to Figure 17 , which is the waveform diagram of the first audio. t1 represents the start time of the first audio segment corresponding to the first character sequence "go", and t2 represents the end time of the first audio segment corresponding to the first character sequence "go". The audio segment within the time period from t1 to t2 is the first audio segment. Based on this, t1 and t2 are used as two cutting points to cut the first audio. After cutting, the first audio segment within the time period from t1 to t2 is deleted, and the two remaining audio segments on the left and right of the first audio segment, that is, the second audio segments, are spliced together to obtain the second audio. The waveform diagram of the second audio is as Figure 18 shown.
[0354] Step 1315. When at least two second audio segments are obtained, the audio obtained by splicing the at least two second audio segments is determined as the second audio.
[0355] When at least three audio segments are obtained by cutting, after deleting the first audio segment, the remaining second audio segments are spliced in chronological order to obtain a spliced audio, and this spliced audio is used as the second audio. For example, the first text information is "Come to Beijing to study media major", and the first character sequence is "go". After deleting the first audio segment, the remaining audio segments include the audio segment with the text content "Come to Beijing" and the audio segment with the text content "Study media major". The above two remaining audio segments are spliced in chronological order to obtain a spliced audio with the text content "Come to Beijing to study media major", and this spliced audio is used as the second audio.
[0356] In some embodiments of this application, after obtaining the second audio, the waveform of the second audio and second text information can be displayed in the audio editing interface for easy viewing and further editing by the user. For example, see... Figure 12 After the user deletes the word "go" from the first text message, the result is as follows: Figure 12 The waveform of the second audio signal 1201 and the second text information 1202 are shown.
[0357] Through the above embodiments, it is possible to obtain audio segments corresponding to consecutive characters in the first audio.
[0358] In some embodiments of this application, when the first character sequence includes at least two non-contiguous character subsequences, that is, when the first character sequence is a discrete sequence composed of at least two characters in the first text information, the above step 130 may specifically include the following steps 1321-1325.
[0359] Step 1321. Compare the second text information with the first text information to determine at least two character subsequences that have been deleted from the first text information.
[0360] As described above, when the text editing input is used to delete the first character sequence in the first text information, in response to the text editing input, the first character sequence is deleted from the first text information, resulting in the second text information. Thus, by comparing the first text information and the second text information, the deleted first character sequence in the first text information can be determined, i.e., at least two character subsequences. For example, if the first text information is "coming to Beijing to study media", and the second text information is "coming to Beijing to study media", by comparing the two, it can be determined that the first character sequence includes two character subsequences: "go" and "major".
[0361] Step 1322. Based on the time mapping relationship, determine the start time and end time of the third audio segment corresponding to each character subsequence.
[0362] In this embodiment, for each character subsequence in the first character sequence, the first audio includes an audio segment corresponding to that character subsequence, namely, a third audio segment. Based on this, the start and end times of the third audio segment corresponding to each character subsequence can be determined according to the time mapping relationship.
[0363] For each character subsequence, the start time of the corresponding third audio segment of the character subsequence is the start time of the audio segment in the first audio that corresponds to the first character in the character subsequence, and the end time of the segment of the character subsequence is the end time of the audio segment in the first audio that corresponds to the last character in the character subsequence. For example, for the character subsequence "去", the start time of the corresponding third audio segment is the start time of the audio segment in the first audio that corresponds to the character "去", and the end time of the segment of this character subsequence is the end time of the audio segment in the first audio that corresponds to the character "去". Another example, for the character subsequence "专业", the start time of the corresponding third audio segment is the start time of the audio segment in the first audio that corresponds to the character "专", and the end time of the segment of this character subsequence is the end time of the audio segment in the first audio that corresponds to the character "业".
[0364] Step 1323. Based on the start time and end time of the segment of at least one character subsequence, clip the first audio to obtain the third audio segment corresponding to each character subsequence.
[0365] In the embodiments of the present application, after determining the start time and end time of the segment of each third audio segment, at least two clip points of the first audio can be determined based on the start time and end time of the segment of all third audio segments, and the first audio segment can be clipped based on at least two clip points to obtain each third audio segment.
[0366] The method for determining the clip points of the first audio based on the start time and end time of the segment of the third audio segment is the same as the method for determining the clip points of the first audio based on the start time and end time of the segment of the first audio segment in the above embodiments. To avoid repetition, it will not be elaborated here too much.
[0367] For example, the first text information is "coming to Beijing to study media". The first character sequence includes two subsequences: "go" and "major". Based on the start and end times of the two third audio segments corresponding to these two subsequences, the start and end times of the third audio segments corresponding to the "go" subsequence and the start time of the third audio segments corresponding to the "major" subsequence are used as three cutting points for the first audio. The first audio is then cut based on these two cutting points to obtain four audio segments: an audio segment with the text content "coming to Beijing", an audio segment with the text content "go", an audio segment with the text content "study media", and an audio segment with the text content "major". The audio segment with the text content "go" is the third audio segment corresponding to the "go" subsequence, and the audio segment with the text content "major" is the third audio segment corresponding to the "major" subsequence. In this way, the third audio segment corresponding to each character subsequence can be obtained through cutting.
[0368] Step 1324. Delete at least two third audio segments from the first audio. If a fourth audio segment is obtained, identify the fourth audio segment as the second audio.
[0369] In this embodiment of the application, the fourth audio segment refers to the audio segment remaining after all the third audio segments in the first audio are deleted.
[0370] In this embodiment of the application, at least two cut points of the first audio are determined based on the start time and end time of all third audio segments. The first audio is cut based on the at least two cut points to obtain at least three audio segments. At least two of the at least three audio segments are third audio segments corresponding to character subsequences. Based on this, after deleting all third audio segments of the first audio, at least one audio segment can remain, that is, at least one fourth audio segment remains.
[0371] If there is only one remaining fourth audio segment, then that fourth audio segment can be directly identified as the second audio segment. For example, if the first text information is "coming to Beijing to study media", and the first character sequence includes two character subsequences, namely "come" and "major", after deleting the third audio segment corresponding to the above two character subsequences, the remaining fourth audio segment is the audio segment with the text content "going to Beijing to study media", and this fourth audio segment can be directly identified as the second audio segment.
[0372] Step 1325. If at least two fourth audio segments are obtained, the audio obtained by splicing the at least two fourth audio segments shall be determined as the second audio.
[0373] In the case where the number of remaining fourth audio segments is at least two, at least two fourth audio segments are spliced in chronological order to obtain a spliced audio, and the spliced audio is determined as the second audio. For example, if the first text information is "Come to Beijing to study media major", the first character sequence includes two character subsequences, namely "to" and "major". After deleting the third audio segments corresponding to the above two character subsequences, the remaining fourth audio segments include an audio segment with the text content "Come to Beijing" and an audio segment with the text content "study media". The above two fourth audio segments are spliced in chronological order to obtain a spliced audio with the text content "Come to Beijing to study media", and the spliced audio is determined as the second audio.
[0374] In some embodiments of the present application, after obtaining the second audio, the waveform diagram of the second audio and the second text information can be displayed in the audio editing interface, facilitating the user to view and further edit. Exemplarily, see Figure 13 , after the user deletes the two character subsequences "to" and "major" in the first text information, the waveform diagram 1301 of the second audio and the second text information 1302 as shown in Figure 13 are obtained.
[0375] By the above method, the audio segments corresponding to the discrete characters in the first audio can be deleted.
[0376] In some embodiments, in the case where the text editing input is an editing input for inserting a second character sequence into the first text information, the above step 120 may specifically include: inserting the second character sequence into the first text information to obtain the second text information.
[0377] In some embodiments of the present application, when inserting the second character sequence into the first text information, the position where the cursor stays in the audio editing interface can be determined as the insertion position of the second character sequence.
[0378] Correspondingly, in the case where the text editing input is an editing input for inserting a second character sequence into the first text information, the above step 130 may include the following steps 1331-step 1333.
[0379] Step 1331. Compare the second text information with the first text information to determine the inserted second character sequence and the insertion position of the second character sequence in the first text information.
[0380] As described above, in the case where the text editing input is an editing input for inserting a first character sequence into the first text information, in response to the text editing input, the second character sequence is inserted into the first text information and then deleted to obtain the second text information. Thus, by comparing the first file information and the second text information, the second character sequence inserted into the first text information can be determined. Moreover, based on the position of the character adjacent to the second character sequence in the second text information in the first text information, the insertion position of the second character sequence can be determined.
[0381] For example, the first text information is "Come to Beijing to study media major", and the second text information is "Come to Beijing to study media major. I am very happy". By comparing the two, it can be determined that the second character sequence inserted into the first text information is "I am very happy", and the character adjacent to the second character sequence is the character "major" located in front of, i.e., to the left of, the second character sequence. Based on the position of the character "major" at the end of the first text information, it can be determined that the insertion position of the second character sequence is at the end of the first text information.
[0382] Another example, the first text information is "Come to Beijing to study media major", and the second text information is "I come to Beijing to study media major". By comparing the two, it can be determined that the second character sequence inserted into the first text information is "I", and the character adjacent to the second character sequence is the character "come" located behind, i.e., to the right of, the second character sequence. Based on the position of the character "come" at the beginning of the first text information, it can be determined that the insertion position of the second character sequence is at the beginning of the first text information.
[0383] Another example, the first text information is "Come to Beijing to study media major", and the second text information is "Come to Beijing to study at a university for media major". By comparing the two, it can be determined that the second character sequence inserted into the first text information is "at a university", and the characters adjacent to the second character sequence are the character "to" located in front of the second character sequence and the character "study" located behind the second character sequence. Thus, it can be determined that the insertion position of the second character sequence is between the characters "to" and "study" in the first text information.
[0384] Step 1332. Input the first audio, the first text information, and the second character sequence into the audio generation model, and through the audio generation model, generate a first supplementary audio segment corresponding to the second character sequence.
[0385] In the embodiment of the present application, the audio generation model is a model for generating an audio corresponding to the input text information.
[0386] In related technologies, if a user misses a segment while recording audio, they typically record a replacement segment and then insert it into the corresponding position of the first audio using editing software. This method not only takes more time, but the re-recorded audio is also difficult to match the first audio in terms of emotion, timbre, and background noise, resulting in poor audio quality. Listeners can easily distinguish which parts were added later, affecting their listening experience. Therefore, to improve audio quality, this application embodiment uses an audio generation model to generate supplementary audio with the same timbre and style as the first audio based on the inserted second character sequence, and then inserts the generated supplementary audio into the corresponding position of the first audio.
[0387] In some embodiments of this application, the audio generation model may adopt a Zero-Shot TTS model structure, such as... Figure 19 As shown, the audio generation model includes an encoder, a quantization module, a language model, a decoder, and a vocoder. The encoder extracts the acoustic features of the first audio file; the quantization module quantizes the acoustic features to obtain quantized discrete codes; the language model predicts the audio code corresponding to a newly added character sequence in the first text information based on input features, including a first character code sequence and a second character code sequence. The first character code sequence is obtained by segmenting the first text information into words, and the second character code sequence is obtained by segmenting the newly added character sequence into words. The newly added character sequence is either a second character sequence inserted into the first text information or a fourth character sequence used to replace a third character sequence in the first text information; the decoder converts the audio code into predicted acoustic features; and the vocoder generates supplementary audio segments corresponding to the newly added character sequences based on the predicted acoustic features. Therefore, step 1332 can specifically include steps 13321-13328.
[0388] Step 13321. Extract the acoustic features of the first audio signal using an encoder.
[0389] Step 13322. The acoustic features are quantized using the quantization module to obtain the quantized discrete code.
[0390] Step 13323. Perform word segmentation on the first text information to obtain the first character encoding sequence.
[0391] In this embodiment of the application, the first text information can be segmented using BPE.
[0392] Step 13324. Perform word segmentation on the second character sequence to obtain the second character encoding sequence.
[0393] In this embodiment of the application, the second character sequence can be segmented using BPE.
[0394] Step 13325. Concatenate the first character encoding sequence, the second character encoding sequence, and the discrete code to obtain the input features.
[0395] Step 13326. Using a language model, predict the audio encoding corresponding to the second character sequence based on the input features.
[0396] Step 13327. Convert the audio encoding into predicted acoustic features using a decoder.
[0397] Step 13328. Using a vocoder, generate a supplementary audio segment corresponding to the second character sequence based on the predicted acoustic features.
[0398] In some embodiments of this application, the vocoder in the audio generation model can be any open-source model, such as Parallel WaveGAN, HiFi-GAN, etc.
[0399] In some embodiments of this application, optionally, before step 1332 above, the encoder and quantization module in the audio generation model can be trained through steps 410-420, and then the language model in the audio generation model can be trained based on the trained encoder and quantization module. Finally, the trained audio generation model is obtained based on the trained encoder, quantization module and language model.
[0400] The audio generation model is used to convert a second character sequence into supplementary audio with the same timbre and style as the first audio.
[0401] Step 410. Construct training data.
[0402] In this embodiment of the application, audio samples can be obtained in the same way as audio samples are obtained when training a speech recognition model, and acoustic features can be extracted from the obtained audio samples, and the extracted acoustic features can be used as training data.
[0403] In some embodiments, when extracting acoustic features from audio samples, the audio samples can first be resampled to a preset frequency and then framed. The resampled and framed audio data can then be used to extract acoustic features. For example, the audio samples can be resampled to 24kHz, and a Fourier transform can be performed with a frame length of 1024 and a frame shift of 300 to obtain 512-dimensional acoustic features.
[0404] Step 420. Train the encoder and quantization model based on the training data.
[0405] In this embodiment of the application, the quantization model is the VQ (Vector quantization) module.
[0406] In some embodiments of this application, a GAN (Generative Adversarial Network) structure can be used to train the encoder and quantization model. For example... Figure 20 As shown, the GAN structure includes an encoder, a quantization module, a decoder, and a discriminator. Based on this, when training the encoder and the quantization model, the acoustic features obtained in step 410 can be used as input, the encoder can obtain the embedding, and then the quantization model can perform discrete encoding. The decoder can reconstruct the acoustic features through discrete encoding. The discriminator determines the training loss value. If the training loss value is less than the set threshold, training stops, and the trained encoder and quantization model are obtained. If the training loss value is greater than or equal to the threshold, training continues until the training loss value is less than the threshold.
[0407] Let x be the acoustic features in the training data, and let y be the acoustic features reconstructed by the decoder. Let z be the discrete code obtained through the quantization module. Then the training loss value can be calculated using the following formulas (11)-(13):
[0408]
[0409] L = L GAN +L VQ (13)
[0410] In the above formulas (11)-(13), L represents the training loss value, E represents the encoder, D represents the discriminator, and sg represents the stop-gradient operation, which is an identity function during forward propagation and has a partial derivative of 0 during backward propagation.
[0411] After training the encoder and quantization module, the language model is trained using these modules. In this embodiment, the initial model used for training the speech model can employ a Decoder-Only Transformer architecture similar to Llama2. Considering training and inference costs, the number of transformer blocks is reduced from 32 to 24, and the number of attention heads is reduced from 32 to 16, based on Llama2-7B. LayerNorm (layer normalization) and skip connections are added between transformer blocks to enhance model stability. The internal structure of each transformer block is as follows: Figure 21As shown, N represents N identical stacked structures. RoPE (Rotary Positional Embedding) represents rotational position encoding, used to encode temporal position information into the Q and K parameters of the attention mechanism, enhancing the model's ability to handle long-range dependencies and supporting longer contexts. Furthermore, using Root Mean Square LayerNorm (RMSNorm) instead of the traditional LayerNorm helps accelerate training and improve numerical stability. Each attention mechanism layer uses a residual connection with a Multilayer Perceptron (MLP), and the last layer is linearly mapped to obtain logits.
[0412] When training the speech model, the encoder and quantization model trained in step 420 can be used to convert the audio samples used to train the speech model into discrete codes x. VQ The text corresponding to the audio sample is segmented using BPE to obtain the character encoding sequence t, and x is then processed. VQ The x and t are concatenated together as input to the second initial model, which will learn to reconstruct x. VQ And t, thus outputting the reconstructed discrete code. and character encoding sequence character encoding sequence x VQ and t and The cross-entropy between the input and output features is used as the loss function for model training. The goal of model training is to minimize the cross-entropy loss so that the reconstructed features are as close as possible to the input features. The loss function is expressed by formulas (14) and (15) as follows:
[0413]
[0414] In formulas (14) and (15) above, L text L represents the reconstruction loss on the text. semantic This indicates the reconstruction loss in audio.
[0415] The audio samples used to train the speech model can contain various timbres, acoustic scenes, emotions, etc. In this way, the acoustic codes that the trained speech model can predict also contain relevant information. Therefore, the audio generation model built on the trained speech model can be used for continuation writing tasks that maintain consistency in timbres and styles.
[0416] Step 1333. Based on the insertion position, insert the first supplementary audio segment into the first audio to obtain the second audio.
[0417] In an embodiment of the present application, after obtaining the first supplementary audio segment corresponding to the second character sequence, inserting the first supplementary audio segment into the corresponding insertion position in the first audio can obtain the second audio.
[0418] In some embodiments of the present application, when the insertion position of the second character sequence is before the first character of the first text information, that is, at the head of the first text information, step 1333 may specifically include: after aligning the segment end time of the first supplementary audio segment with the audio start time of the first audio, inserting it into the first audio to obtain the second audio. For example, if the first text information is "Come to Beijing to study media majors", the second text information is "I come to Beijing to study media majors", the second character sequence is "I", and the insertion position of the second character sequence is at the head of the first text information, then after obtaining the first supplementary audio segment, insert the first supplementary audio segment in front of the audio segment corresponding to the character "Come" in the first audio.
[0419] In some embodiments of the present application, when the insertion position of the second character sequence is after the last character of the first text information, that is, at the tail of the first text information, step 1333 may specifically include: after aligning the segment start time of the first supplementary audio segment with the audio end time of the first audio, inserting it into the first audio to obtain the second audio. For example, if the first text information is "Come to Beijing to study media majors", the second text information is "Come to Beijing to study media majors and I am very happy", the second character sequence is "and I am very happy", and the insertion position of the second character sequence is at the tail of the first text information, then after obtaining the first supplementary audio segment, insert the first supplementary audio segment behind the audio segment corresponding to the character "majors" in the first audio.
[0420] In some embodiments of this application, when the insertion position of the second character sequence is located between the first character and the last character of the first text information, step 1333 may specifically include: determining the first character to the left of the insertion position and the second character to the right of the insertion position in the first text information; based on the time mapping relationship, determining the first segment end time of the audio segment corresponding to the first character and the first segment start time of the audio segment corresponding to the second character in the first audio; aligning the segment start time of the first supplementary audio segment with the segment end time of the first segment, and after aligning the segment end time of the first supplementary audio segment with the segment start time of the first segment, inserting the first audio to obtain the second audio. For example, if the first text information is "coming to Beijing to study media", the second text information is "coming to Beijing to study media at university", the second character sequence is "university", and the insertion position of the second character sequence is determined to be between the characters "go" and "study" in the first text information, then after obtaining the first supplementary audio segment, the first supplementary audio segment is inserted between the audio segment corresponding to the character "go" and the audio segment corresponding to the character "study" in the first audio.
[0421] In some embodiments of this application, after obtaining the second audio, the waveform of the second audio and second text information can be displayed in the audio editing interface for easy viewing and further editing by the user. For example, see... Figure 14 After the user inserts the second character sequence "I'm happy" at the end of the first text message, the result is as follows: Figure 14 The waveform diagram of the second audio 1401 and the second text information 1402 are shown.
[0422] In this embodiment of the application, by combining an audio generation model with a structure such as Zero-Shot TTS, a supplementary audio segment with the same timbre and style as the user-recorded first audio can be synthesized and then inserted into the user-specified position to obtain the second audio. In this way, the supplementary audio segment is not easily perceived in the entire second audio, does not affect the overall naturalness, and effectively improves the user's listening experience.
[0423] In some embodiments, when the text editing input is an editing input that modifies the third character sequence in the first text information to a fourth character sequence, the above step 120 specifically includes: replacing the third character sequence in the first text information with the fourth character sequence to obtain the second text information.
[0424] In some embodiments of the present application, in the case where the text editing input is an editing input for replacing the third character sequence in the first text information with the fourth character sequence, in response to the text editing input, the third character sequence is deleted from the first text information, and the fourth character sequence is inserted into the position where the original third character sequence was located in the first text information, thereby implementing the replacement of the character sequence and obtaining the second text information.
[0425] Correspondingly, step 130 described above may specifically include the following steps 1341-step 1345.
[0426] Step 1341. Compare the second text information with the first text information to determine the third character sequence deleted from the first text information and the newly added fourth character sequence.
[0427] In the embodiments of the present application, in the case where the text editing input is an editing input for replacing the third character sequence in the first text information with the fourth character sequence, in response to the text editing input, the third character sequence is deleted from the first text information, and the fourth character sequence is inserted into the corresponding position to obtain the second text information. Thus, by comparing the first file information and the second text information, the third character sequence deleted from the first text information and the inserted fourth character sequence can be determined.
[0428] For example, if the first text information is "Come to Beijing to study media majors", and the second text information is "Come to Nanjing to study media majors", by comparing the two, it can be determined that the third character sequence deleted from the first text information is "Beijing", and the inserted fourth character sequence is "Nanjing".
[0429] Step 1342. Based on the time mapping relationship, determine the start time and end time of the fifth audio segment corresponding to the third character sequence in the first audio.
[0430] In the embodiments of the present application, the fifth audio segment is the audio segment corresponding to the third character sequence in the first audio. For example, if the third character sequence is the character sequence "Beijing", then the fifth audio segment is the audio segment in the first audio with the text content "Beijing".
[0431] The start time of the fifth audio segment is the start time of the audio segment corresponding to the first character in the third character sequence in the first audio, and the end time of the fifth audio segment is the end time of the audio segment corresponding to the last character in the third character sequence in the first audio. For example, if the third character sequence is the word "Beijing", then the start time of the fifth audio segment is the start time of the audio segment corresponding to the character "North", and the end time of the fifth audio segment is the end time of the audio segment corresponding to the character "Jing".
[0432] Step 1343. Based on the start and end times of the fifth audio segment, cut the first audio segment to obtain the fifth audio segment.
[0433] In this embodiment, after determining the start and end times of the fifth audio segment, at least one cut point of the first audio segment can be determined based on the start and end times of the fifth audio segment. The first audio segment is then cut based on the at least one cut point to obtain the fifth audio segment. Here, the method of cutting the first audio segment based on the start and end times of the fifth audio segment to obtain the fifth audio segment is consistent with the method of step 1313 described above, and can be found in the relevant description of step 1313 above. To avoid repetition, it will not be repeated here.
[0434] Step 1344. Input the first audio, the first text information, and the fourth character sequence into the audio generation model, and generate a second supplementary audio segment corresponding to the fourth character sequence through the audio generation model.
[0435] In this embodiment of the application, the second supplementary audio segment is an audio segment generated by an audio generation model, and the text content is the fourth character sequence.
[0436] The method for generating the second supplementary audio segment is the same as the method for generating the first supplementary audio segment in step 1332 above. For details, please refer to the relevant description of step 1332 above. To avoid repetition, it will not be repeated here.
[0437] Step 1345. Replace the fifth audio segment in the first audio with the second supplementary audio segment to obtain the second audio.
[0438] In this embodiment, the insertion position of the fourth character sequence is the position of the third character sequence in the first text information. Based on this, after obtaining the second supplementary audio segment and the fifth audio segment, the second audio can be obtained by directly replacing the fifth audio segment in the first audio with the second supplementary audio segment. For example, if the third character sequence is "Beijing" and the fourth character sequence is "Nanjing", then the fifth audio segment in the first audio with the text content "Beijing" is directly replaced with the second supplementary audio segment with the text content "Nanjing". The audio composed of the second supplementary audio segment and other audio segments in the first audio besides the fifth audio segment is determined as the second audio.
[0439] In some embodiments of this application, after obtaining the second audio, the waveform of the second audio and second text information can be displayed in the audio editing interface for easy viewing and further editing by the user. For example, see... Figure 15 After the user replaces "Beijing" with "Nanjing" in the first text message, the result is as follows: Figure 15The waveform diagram of the second audio is shown in Figure 1501, and the second text information is shown in Figure 1502.
[0440] The above method allows for the modification of certain audio segments within the first audio file.
[0441] In some embodiments, in addition to step 120 described above, the following steps may also be performed:
[0442] In the audio editing interface, a second text information is displayed. The original character sequence and the newly added character sequence in the second text information are displayed in different modes; the original character sequence is the same as the first text information in the second text information, while the newly added character sequence is a character sequence in the second text information that is different from the first text information. For example, the newly added character sequence is either a second character sequence inserted into the first text information or a fourth character sequence used to replace the third character sequence in the first text information.
[0443] In some embodiments of this application, different display modes may include, but are not limited to, different fonts and different font colors.
[0444] For example, the first text information is "I came to Beijing to study media". After the user inserts the second character sequence "I am very happy" at the end of the first text information, the second text information "I am very happy to come to Beijing to study media" can be displayed in the audio editing interface. "I came to Beijing to study media" is the original character sequence, which can be displayed in black, and "I am very happy" is the newly added character sequence, which can be displayed in red.
[0445] The above methods allow users to intuitively understand the adjusted content.
[0446] In some embodiments, to enhance editing flexibility, the audio editing interface may also include a control for undoing adjustments. After adjusting the first text information in response to user input, if the user is not satisfied with the adjusted text, they can undo the previous adjustments by clicking the control for undoing the adjustments. For example, see... Figure 11 This includes a control 1107 for undoing adjustments, which allows users to undo previous adjustments by clicking on the control 1107.
[0447] In some embodiments, the audio editing interface may further include a control for submitting second text information. Before clicking this control, the user can adjust the first text information multiple times through text editing input. After the user determines that the adjusted text meets the requirements, clicking the control for submitting the second text information determines the text obtained from the last adjustment as the second text information. For example, see... Figure 11This includes a control 1108 for submitting second text information. If the user determines that the adjusted text meets the requirements, they can click on the control 1108 to confirm the last adjusted text as the second text information.
[0448] For example, the first text information is "I came to Beijing to study media". The user deleted the word "go" by editing the input, and then inserted "I am very happy" at the end by editing the input. After making the modification, the user clicked the control to submit the second text information. In this way, "I am very happy to come to Beijing to study media" is identified as the second text information.
[0449] Optionally, in some embodiments, to avoid user error leading to inaccurate finalized second text information, after the user clicks the control for submitting the second text information, a prompt message can be further displayed in the audio editing interface to ask the user whether they are sure they want to submit. If the user confirms submission, the last adjusted text is then determined as the second text information. For example, see... Figure 22 After the user clicks the control used to submit the second text information, a pop-up window 2200 will display the prompt message "Confirm to submit current text?" on the audio editing interface. In addition to the prompt message, the pop-up window 2200 also includes a control 2201. When the user confirms to submit the current text, the user can click the control 2201. In response to the user's click input on the control 2201, the user is confirmed to submit the current text, thereby confirming the last adjusted text as the second text information.
[0450] The audio processing method provided in this application can be executed by an audio processing device. This application uses an audio processing device executing the audio processing method as an example to illustrate the audio processing device provided in this application.
[0451] See Figure 23 The above is a schematic diagram of an audio processing apparatus provided in some embodiments of this application, such as... Figure 23 As shown, the device 2300 includes the following modules:
[0452] The receiving module 2301 is used to receive text editing input of the first text information of the first audio.
[0453] The text update module 2302 is used to update the first text information in response to text editing input and obtain the second text information;
[0454] The audio processing module 2303 is used to perform audio processing on the first audio based on the second text information to obtain the second audio.
[0455] In this embodiment, users can edit audio by directly editing the text information corresponding to the audio, eliminating the need for tedious and complex manual editing operations such as cutting and splicing. This allows users without audio editing skills to quickly edit audio, improving the efficiency of audio processing while ensuring its accuracy.
[0456] In some embodiments, the device 2300 further includes:
[0457] The speech recognition module is used to input the first audio into the speech recognition model before receiving the text editing input of the first text information of the first audio, and to perform speech recognition on the first audio through the speech recognition model to obtain the first text information of the first audio.
[0458] The alignment module is used to input the first audio and the first text information into the audio-text alignment model. Through the audio-text alignment model, the first audio and the first text information are time-aligned to obtain the time mapping relationship between the first text information and the first audio. The time mapping relationship is used to indicate the start and end times of the audio segment corresponding to each character in the first text information in the first audio.
[0459] The display module is used to display the audio editing page interface. The audio editing page interface includes a timeline, a waveform of the first audio, and first text information. The timeline includes at least one time marker. Each character in the waveform and the first text information is aligned with the corresponding waveform and the corresponding time marker in the waveform.
[0460] Receiver module 2301 is specifically used for:
[0461] Receive text editing input for the first text information in the audio editing page interface.
[0462] In some embodiments, the text editing input is an editing input that deletes a first character sequence from the first text information;
[0463] Text update module 2302 is specifically used for:
[0464] Delete the first character sequence from the first text information to obtain the second text information.
[0465] In some embodiments, the first character sequence is a series of consecutive characters in the first text information;
[0466] Audio processing module 2303 is specifically used for:
[0467] The second text information is compared with the first text information to determine the first character sequence that was deleted from the first text information;
[0468] Based on the time mapping relationship, the start time and end time of the first audio segment corresponding to the first character sequence are determined. The start time of the first audio segment is the start time of the audio segment corresponding to the first character in the first character sequence in the first audio, and the end time of the first audio segment is the end time of the audio segment corresponding to the last character in the first character sequence in the first audio.
[0469] Based on the start and end times of the first audio segment, the first audio is cut to obtain the first audio segment;
[0470] Delete the first audio segment from the first audio, and if a second audio segment is obtained, identify the second audio segment as the second audio;
[0471] If at least two second audio segments are obtained, the audio obtained by splicing the at least two second audio segments is determined as the second audio.
[0472] In some embodiments, the first character sequence comprises at least two non-contiguous character subsequences;
[0473] Audio processing module 2303 is specifically used for:
[0474] The second text information is compared with the first text information to determine at least two character subsequences that have been deleted from the first text information;
[0475] Based on the time mapping relationship, the start time and end time of the third audio segment corresponding to each character subsequence are determined. The start time of the third audio segment corresponding to each character subsequence is the start time of the audio segment corresponding to the first character in the first audio and the end time of the third audio segment corresponding to each character subsequence is the end time of the audio segment corresponding to the last character in the first audio and the last character in the character subsequence.
[0476] Based on the start and end times of at least one character subsequence, the first audio segment is cut to obtain the third audio segment corresponding to each character subsequence;
[0477] Delete at least two third audio segments from the first audio segment. If a fourth audio segment is obtained, identify the fourth audio segment as the second audio segment.
[0478] If at least two fourth audio segments are obtained, the audio obtained by splicing the at least two fourth audio segments is determined as the second audio.
[0479] In some embodiments, the text editing input is an editing input that inserts a second character sequence into the first text information;
[0480] Text update module 2302 is specifically used for:
[0481] Insert the second character sequence into the first text information to obtain the second text information.
[0482] In some embodiments, the audio processing module 2303 is specifically used for:
[0483] The second text information is compared with the first text information to determine the insertion position of the second character sequence and the second character sequence inserted in the first text information;
[0484] The first audio, the first text information, and the second character sequence are input into the audio generation model. The audio generation model generates a first supplementary audio segment corresponding to the second character sequence.
[0485] Based on the insertion position, the first supplementary audio segment is inserted into the first audio to obtain the second audio.
[0486] In some embodiments, the insertion position is before the first character of the first text information;
[0487] Audio processing module 2303 is specifically used for:
[0488] After aligning the end time of the first supplementary audio segment with the start time of the first audio segment, insert the first audio segment to obtain the second audio segment.
[0489] In some embodiments, the insertion position is after the last character of the first text information;
[0490] Audio processing module 2303 is specifically used for:
[0491] After aligning the start time of the first supplementary audio segment with the end time of the first audio segment, insert the first audio segment to obtain the second audio segment.
[0492] In some embodiments, the insertion position is located between the first character and the last character of the first text information;
[0493] Audio processing module 2303 is specifically used for:
[0494] Determine the first character to the left of the insertion position and the second character to the right of the insertion position in the first text information;
[0495] Based on the time mapping relationship, determine the end time of the first segment of the audio segment corresponding to the first character and the start time of the first segment of the audio segment corresponding to the second character in the first audio;
[0496] Align the start time of the first supplementary audio segment with the end time of the first segment, and align the end time of the first supplementary audio segment with the start time of the first segment, then insert the first audio to obtain the second audio.
[0497] In some embodiments, the text editing input is an editing input that modifies the third character sequence in the first text information into a fourth character sequence;
[0498] Text update module 2302 is specifically used for:
[0499] The third character sequence in the first text information is replaced with the fourth character sequence to obtain the second text information.
[0500] In some embodiments, the audio processing module 2303 is specifically used for:
[0501] The second text information is compared with the first text information to determine the third character sequence that was deleted and the fourth character sequence that was added in the first text information;
[0502] Based on the time mapping relationship, the start time and end time of the fifth audio segment corresponding to the third character sequence in the first audio are determined. The start time of the fifth audio segment is the start time of the audio segment corresponding to the first character in the third character sequence in the first audio, and the end time of the fifth audio segment is the end time of the audio segment corresponding to the last character in the third character sequence in the first audio.
[0503] Based on the start and end times of the fifth audio segment, the first audio segment is cut to obtain the fifth audio segment;
[0504] The first audio, the first text information, and the fourth character sequence are input into the audio generation model. The audio generation model generates a second supplementary audio segment corresponding to the fourth character sequence.
[0505] Replace the fifth audio segment in the first audio with the second supplementary audio segment to obtain the second audio.
[0506] In some embodiments, the audio generation model includes an encoder, a quantization module, a language model, a decoder, and a vocoder;
[0507] An encoder is used to extract the acoustic features of the first audio audio.
[0508] The quantization module is used to quantize acoustic features to obtain quantized discrete codes;
[0509] A language model is used to predict the audio encoding corresponding to a newly added character sequence in a first text message based on input features. The input features include a first character encoding sequence and a second character encoding sequence. The first character encoding sequence is obtained by word segmentation of the first text message, and the second character encoding sequence is obtained by word segmentation of the newly added character sequence. The newly added character sequence is either the second character sequence inserted into the first text message or a fourth character sequence used to replace the third character sequence in the first text message.
[0510] A decoder is used to convert audio encoding into predictive acoustic features;
[0511] A vocoder is used to generate supplementary audio segments corresponding to newly added character sequences based on predicted acoustic features.
[0512] In some embodiments, the display module is further configured to:
[0513] In response to text editing input, the first text information is updated, and after obtaining the second text information, the second text information is displayed in the audio editing interface;
[0514] The original character sequence and the newly added character sequence in the second text information have different display modes; the original character sequence is the same as the character sequence in the first text information in the second text information, and the newly added character sequence is the second character sequence inserted into the first text information or the fourth character sequence used to replace the third character sequence in the first text information.
[0515] In some embodiments, the apparatus 2300 further includes: a model training module, configured to:
[0516] Based on the sample dataset and loss function, the initial model is trained to obtain the speech recognition model;
[0517] The sample dataset includes multiple audio samples and corresponding text information samples. The multiple audio samples include audio recorded in at least two scenarios, and the audio samples include human voice features.
[0518] In some embodiments, the loss function is a joint loss function, wherein the joint loss function includes at least two loss function terms, each loss function term corresponds to a weight, and one of the at least two loss function terms includes forward relative entropy;
[0519] The model training module is specifically used for:
[0520] Extract acoustic feature samples for each audio sample in the sample dataset;
[0521] The text information sample corresponding to each audio sample is segmented into words to obtain the character encoding sequence sample of each audio sample;
[0522] Based on the acoustic feature samples and character encoding sequence samples of all audio samples in the sample dataset, multiple sets of training data are constructed; each set of training data includes the acoustic feature samples and character encoding sequence samples of one audio sample.
[0523] The initial model is trained based on multiple sets of training data to obtain the predicted character encoding sequence corresponding to each set of training data;
[0524] Having completed at least two rounds of training on the initial model, the forward relative entropy is calculated based on the predicted character encoding sequence obtained in the i-th round of training and the predicted character encoding sequence obtained in the (i-1)-th round of training; where i ≥ 2.
[0525] Based on the forward relative entropy, update the entropy value of the forward relative entropy in the loss function to obtain the updated loss function;
[0526] Based on the updated loss function, calculate the loss value of the initial model;
[0527] If the loss value is greater than or equal to the loss threshold, adjust the weight of each loss function term in the loss function, and continue to train the initial model with adjusted parameters based on acoustic feature samples and text sequence samples.
[0528] Training stops when the loss value is less than the loss threshold, resulting in a trained speech recognition model.
[0529] In some embodiments, the initial model includes an encoder, a decoder, and a joint network;
[0530] The model training module is specifically used for:
[0531] The acoustic feature samples from each set of training data are embedded into a spatial vector by an encoder to obtain the acoustic feature vector samples of each set of training data.
[0532] The character encoding sequence samples in each training data set are embedded into a spatial vector using a decoder, resulting in character encoding sequence vector samples for each training data set.
[0533] By using a joint network, the acoustic feature vector samples and character encoding sequence vector samples of each training data set are concatenated to obtain the concatenated vector of each training data set. The concatenated vector of each training data set is then mapped to the vocabulary space to obtain the predicted character encoding sequence of each training data set.
[0534] The audio processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.
[0535] The audio processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0536] The audio processing device provided in this application embodiment can achieve... Figures 1 to 22 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0537] Optionally, such as Figure 24 As shown, this application embodiment also provides an electronic device 2400, including a processor 2401 and a memory 2402. The memory 2402 stores a program or instructions that can run on the processor 2401. When the program or instructions are executed by the processor 2401, they implement the various steps of the above-described audio processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0538] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0539] Figure 25 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0540] The electronic device 2500 includes, but is not limited to, components such as: radio frequency unit 2501, network module 2502, audio output unit 2503, input unit 2504, sensor 2505, display unit 2506, user input unit 2507, interface unit 2508, memory 2509, and processor 2510.
[0541] Those skilled in the art will understand that the electronic device 2500 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 2510 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 25 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0542] The user input unit 2507 is used to receive text editing input of the first text information of the first audio.
[0543] Processor 2510 is used to update first text information and obtain second text information in response to text editing input;
[0544] The processor 2510 is also used to perform audio processing on the first audio based on the second text information to obtain the second audio.
[0545] In this embodiment, users can edit audio by directly editing the text information corresponding to the audio, eliminating the need for tedious and complex manual editing operations such as cutting and splicing. This allows users without audio editing skills to quickly edit audio, improving the efficiency of audio processing while ensuring its accuracy.
[0546] In some embodiments, the processor 2510 is further configured to:
[0547] Before receiving the text editing input of the first audio's first text information, the first audio is input into the speech recognition model. The speech recognition model performs speech recognition on the first audio to obtain the first text information of the first audio.
[0548] The first audio and the first text information are input into the audio-text alignment model. The audio-text alignment model is used to perform time alignment on the first audio and the first text information to obtain the time mapping relationship between the first text information and the first audio. The time mapping relationship is used to indicate the start and end times of the audio segment corresponding to each character in the first text information in the first audio.
[0549] Display unit 2506 is used to display an audio editing page interface. The audio editing page interface includes a timeline, a waveform of the first audio, and first text information. The timeline includes at least one time marker. Each character in the waveform and the first text information is aligned with the corresponding waveform and the corresponding time marker in the waveform.
[0550] User input unit 2507 is specifically used for:
[0551] Receive text editing input for the first text information in the audio editing page interface.
[0552] In some embodiments, the text editing input is an editing input that deletes a first character sequence from the first text information;
[0553] Processor 2510, specifically used for:
[0554] Delete the first character sequence from the first text information to obtain the second text information.
[0555] In some embodiments, the first character sequence is a series of consecutive characters in the first text information;
[0556] Processor 2510, specifically used for:
[0557] The second text information is compared with the first text information to determine the first character sequence that was deleted from the first text information;
[0558] Based on the time mapping relationship, the start time and end time of the first audio segment corresponding to the first character sequence are determined. The start time of the first audio segment is the start time of the audio segment corresponding to the first character in the first character sequence in the first audio, and the end time of the first audio segment is the end time of the audio segment corresponding to the last character in the first character sequence in the first audio.
[0559] Based on the start and end times of the first audio segment, the first audio is cut to obtain the first audio segment;
[0560] Delete the first audio segment from the first audio, and if a second audio segment is obtained, identify the second audio segment as the second audio;
[0561] If at least two second audio segments are obtained, the audio obtained by splicing the at least two second audio segments is determined as the second audio.
[0562] In some embodiments, the first character sequence comprises at least two non-contiguous character subsequences;
[0563] Processor 2510, specifically used for:
[0564] The second text information is compared with the first text information to determine at least two character subsequences that have been deleted from the first text information;
[0565] Based on the time mapping relationship, the start time and end time of the third audio segment corresponding to each character subsequence are determined. The start time of the third audio segment corresponding to each character subsequence is the start time of the audio segment corresponding to the first character in the first audio and the end time of the third audio segment corresponding to each character subsequence is the end time of the audio segment corresponding to the last character in the first audio and the last character in the character subsequence.
[0566] Based on the start and end times of at least one character subsequence, the first audio segment is cut to obtain the third audio segment corresponding to each character subsequence;
[0567] Delete at least two third audio segments from the first audio segment. If a fourth audio segment is obtained, identify the fourth audio segment as the second audio segment.
[0568] If at least two fourth audio segments are obtained, the audio obtained by splicing the at least two fourth audio segments is determined as the second audio.
[0569] In some embodiments, the text editing input is an editing input that inserts a second character sequence into the first text information;
[0570] Processor 2510, specifically used for:
[0571] Insert the second character sequence into the first text information to obtain the second text information.
[0572] In some embodiments, the processor 2510 is specifically used for:
[0573] The second text information is compared with the first text information to determine the insertion position of the second character sequence and the second character sequence inserted in the first text information;
[0574] The first audio, the first text information, and the second character sequence are input into the audio generation model. The audio generation model generates a first supplementary audio segment corresponding to the second character sequence.
[0575] Based on the insertion position, the first supplementary audio segment is inserted into the first audio to obtain the second audio.
[0576] In some embodiments, the insertion position is before the first character of the first text information;
[0577] Processor 2510, specifically used for:
[0578] After aligning the end time of the first supplementary audio segment with the start time of the first audio segment, insert the first audio segment to obtain the second audio segment.
[0579] In some embodiments, the insertion position is after the last character of the first text information;
[0580] Processor 2510, specifically used for:
[0581] After aligning the start time of the first supplementary audio segment with the end time of the first audio segment, insert the first audio segment to obtain the second audio segment.
[0582] In some embodiments, the insertion position is located between the first character and the last character of the first text information;
[0583] Processor 2510, specifically used for:
[0584] Determine the first character to the left of the insertion position and the second character to the right of the insertion position in the first text information;
[0585] Based on the time mapping relationship, determine the end time of the first segment of the audio segment corresponding to the first character and the start time of the first segment of the audio segment corresponding to the second character in the first audio;
[0586] Align the start time of the first supplementary audio segment with the end time of the first segment, and align the end time of the first supplementary audio segment with the start time of the first segment, then insert the first audio to obtain the second audio.
[0587] In some embodiments, the text editing input is an editing input that modifies the third character sequence in the first text information into a fourth character sequence;
[0588] Processor 2510, specifically used for:
[0589] The third character sequence in the first text information is replaced with the fourth character sequence to obtain the second text information.
[0590] In some embodiments, the audio processing module 2303 is specifically used for:
[0591] The second text information is compared with the first text information to determine the third character sequence that was deleted and the fourth character sequence that was added in the first text information;
[0592] Based on the time mapping relationship, the start time and end time of the fifth audio segment corresponding to the third character sequence in the first audio are determined. The start time of the fifth audio segment is the start time of the audio segment corresponding to the first character in the third character sequence in the first audio, and the end time of the fifth audio segment is the end time of the audio segment corresponding to the last character in the third character sequence in the first audio.
[0593] Based on the start and end times of the fifth audio segment, the first audio segment is cut to obtain the fifth audio segment;
[0594] The first audio, the first text information, and the fourth character sequence are input into the audio generation model. The audio generation model generates a second supplementary audio segment corresponding to the fourth character sequence.
[0595] Replace the fifth audio segment in the first audio with the second supplementary audio segment to obtain the second audio.
[0596] In some embodiments, the audio generation model includes an encoder, a quantization module, a language model, a decoder, and a vocoder;
[0597] An encoder is used to extract the acoustic features of the first audio audio.
[0598] The quantization module is used to quantize acoustic features to obtain quantized discrete codes;
[0599] A language model is used to predict the audio encoding corresponding to a newly added character sequence in a first text message based on input features. The input features include a first character encoding sequence and a second character encoding sequence. The first character encoding sequence is obtained by word segmentation of the first text message, and the second character encoding sequence is obtained by word segmentation of the newly added character sequence. The newly added character sequence is either the second character sequence inserted into the first text message or a fourth character sequence used to replace the third character sequence in the first text message.
[0600] A decoder is used to convert audio encoding into predictive acoustic features;
[0601] A vocoder is used to generate supplementary audio segments corresponding to newly added character sequences based on predicted acoustic features.
[0602] In some embodiments, the display unit 2506 is further configured to:
[0603] In response to text editing input, the first text information is updated, and after obtaining the second text information, the second text information is displayed in the audio editing interface;
[0604] The original character sequence and the newly added character sequence in the second text information have different display modes; the original character sequence is the same as the character sequence in the first text information in the second text information, and the newly added character sequence is the second character sequence inserted into the first text information or the fourth character sequence used to replace the third character sequence in the first text information.
[0605] In some embodiments, the processor 2510 is further configured to:
[0606] Based on the sample dataset and loss function, the initial model is trained to obtain the speech recognition model;
[0607] The sample dataset includes multiple audio samples and corresponding text information samples. The multiple audio samples include audio recorded in at least two scenarios, and the audio samples include human voice features.
[0608] In some embodiments, the loss function is a joint loss function, wherein the joint loss function includes at least two loss function terms, each loss function term corresponds to a weight, and one of the at least two loss function terms includes forward relative entropy;
[0609] Processor 2510, specifically used for:
[0610] Extract acoustic feature samples for each audio sample in the sample dataset;
[0611] The text information sample corresponding to each audio sample is segmented into words to obtain the character encoding sequence sample of each audio sample;
[0612] Based on the acoustic feature samples and character encoding sequence samples of all audio samples in the sample dataset, multiple sets of training data are constructed; each set of training data includes the acoustic feature samples and character encoding sequence samples of one audio sample.
[0613] The initial model is trained based on multiple sets of training data to obtain the predicted character encoding sequence corresponding to each set of training data;
[0614] Having completed at least two rounds of training on the initial model, the forward relative entropy is calculated based on the predicted character encoding sequence obtained in the i-th round of training and the predicted character encoding sequence obtained in the (i-1)-th round of training; where i ≥ 2.
[0615] Based on the forward relative entropy, update the entropy value of the forward relative entropy in the loss function to obtain the updated loss function;
[0616] Based on the updated loss function, calculate the loss value of the initial model;
[0617] If the loss value is greater than or equal to the loss threshold, adjust the weight of each loss function term in the loss function, and continue to train the initial model with adjusted parameters based on acoustic feature samples and text sequence samples.
[0618] Training stops when the loss value is less than the loss threshold, resulting in a trained speech recognition model.
[0619] In some embodiments, the initial model includes an encoder, a decoder, and a joint network;
[0620] Processor 2510, specifically used for:
[0621] The acoustic feature samples from each set of training data are embedded into a spatial vector by an encoder to obtain the acoustic feature vector samples of each set of training data.
[0622] The character encoding sequence samples in each training data set are embedded into a spatial vector using a decoder, resulting in character encoding sequence vector samples for each training data set.
[0623] By using a joint network, the acoustic feature vector samples and character encoding sequence vector samples of each training data set are concatenated to obtain the concatenated vector of each training data set. The concatenated vector of each training data set is then mapped to the vocabulary space to obtain the predicted character encoding sequence of each training data set.
[0624] It should be understood that, in this embodiment, the input unit 2504 may include a graphics processing unit (GPU) 25041 and a microphone 25042. The GPU 25041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 2506 may include a display panel 25061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 2507 includes at least one of a touch panel 25071 and other input devices 25072. The touch panel 25071 is also called a touch screen. The touch panel 25071 may include a touch detection device and a touch controller. Other input devices 25072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0625] The memory 2509 can be used to store software programs and various data. The memory 2509 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 2509 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 2509 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0626] Processor 2510 may include one or more processing units; optionally, processor 2510 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 2510.
[0627] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described audio processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0628] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0629] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described audio processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0630] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0631] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the audio processing method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0632] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0633] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0634] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An audio processing method, characterized in that, include: Receive text editing input for the first text information of the first audio; In response to text editing input, the first text information is updated to obtain the second text information; Based on the second text information, the first audio is processed to obtain the second audio.
2. The method according to claim 1, characterized in that, Before receiving text editing input of the first text information of the first audio, the method further includes: The first audio is input into a speech recognition model, and the speech recognition model performs speech recognition on the first audio to obtain the first text information of the first audio. The first audio and the first text information are input into an audio-text alignment model. The first audio and the first text information are time-aligned through the audio-text alignment model to obtain a time mapping relationship between the first text information and the first audio. The time mapping relationship is used to indicate the start and end times of the audio segment corresponding to each character in the first text information in the first audio. The audio editing interface is displayed, which includes a timeline, a waveform of the first audio, and the first text information. The timeline includes at least one time marker, and each character in the first text information is aligned with the corresponding waveform and the corresponding time marker in the waveform diagram. The text editing input for receiving the first text information of the first audio includes: Receive text editing input for the first text information in the audio editing interface.
3. The method according to claim 2, characterized in that, The text editing input is an edit input that deletes the first character sequence in the first text information; The step of updating the first text information to obtain the second text information includes: The first character sequence is deleted from the first text information to obtain the second text information.
4. The method according to claim 3, characterized in that, The first character sequence is a series of consecutive characters in the first text information; The step of processing the first audio based on the second text information to obtain the second audio includes: The second text information is compared with the first text information to determine the first character sequence that was deleted from the first text information; Based on the time mapping relationship, the start time and end time of the first audio segment corresponding to the first character sequence are determined. The start time of the first audio segment is the start time of the audio segment corresponding to the first character in the first character sequence in the first audio, and the end time of the first audio segment is the end time of the audio segment corresponding to the last character in the first character sequence in the first audio. Based on the start and end times of the first audio segment, the first audio is cut to obtain the first audio segment; Delete the first audio segment from the first audio, and if a second audio segment is obtained, determine the second audio segment as the second audio; If at least two second audio segments are obtained, the audio obtained by splicing the at least two second audio segments is determined as the second audio.
5. The method according to claim 3, characterized in that, The first character sequence includes at least two non-contiguous character subsequences; The step of processing the first audio based on the second text information to obtain the second audio includes: The second text information is compared with the first text information to determine the at least two character subsequences that have been deleted from the first text information; Based on the time mapping relationship, the start time and end time of the third audio segment corresponding to each character subsequence are determined. The start time of the third audio segment corresponding to each character subsequence is the start time of the audio segment in the first audio corresponding to the first character in the character subsequence, and the end time of the third audio segment corresponding to each character subsequence is the end time of the audio segment in the first audio corresponding to the last character in the character subsequence. Based on the segment start time and segment end time of the at least one character subsequence, the first audio is cut to obtain a third audio segment corresponding to each character subsequence; Delete at least two third audio segments from the first audio, and if a fourth audio segment is obtained, determine the fourth audio segment as the second audio; If at least two fourth audio segments are obtained, the audio obtained by splicing the at least two fourth audio segments is determined as the second audio.
6. The method according to claim 2, characterized in that, The text editing input is an editing input that inserts a second character sequence into the first text information; The step of updating the first text information to obtain the second text information includes: The second character sequence is inserted into the first text information to obtain the second text information.
7. The method according to claim 6, characterized in that, The step of processing the first audio based on the second text information to obtain the second audio includes: The second text information is compared with the first text information to determine the second character sequence inserted into the first text information and the insertion position of the second character sequence; The first audio, the first text information, and the second character sequence are input into the audio generation model, and the first supplementary audio segment corresponding to the second character sequence is generated through the audio generation model. Based on the insertion position, the first supplementary audio segment is inserted into the first audio to obtain the second audio.
8. The method according to claim 7, characterized in that, The insertion position is before the first character of the first text information; The step of inserting the first supplementary audio into the first audio based on the insertion position to obtain the second audio includes: After aligning the end time of the first supplementary audio segment with the start time of the first audio segment, insert the first audio segment to obtain the second audio segment.
9. The method according to claim 7, characterized in that, The insertion position is after the last character of the first text information; The step of inserting the first supplementary audio segment into the first audio based on the insertion position to obtain the second audio includes: After aligning the start time of the first supplementary audio segment with the end time of the first audio segment, insert the first audio segment to obtain the second audio segment.
10. The method according to claim 7, characterized in that, The insertion position is located between the first character and the last character of the first text information; The step of inserting the first supplementary audio into the first audio based on the insertion position to obtain the second audio includes: Determine the first character located to the left of the insertion position and the second character located to the right of the insertion position in the first text information; Based on the time mapping relationship, the end time of the first segment of the audio segment corresponding to the first character and the start time of the first segment of the audio segment corresponding to the second character in the first audio are determined. Align the start time of the first supplementary audio segment with the end time of the first segment, and align the end time of the first supplementary audio segment with the start time of the first segment, then insert the first audio to obtain the second audio.
11. The method according to claim 3, characterized in that, The text editing input is an editing input that modifies the third character sequence in the first text information into a fourth character sequence; The step of updating the first text information to obtain the second text information includes: The third character sequence in the first text information is replaced with the fourth character sequence to obtain the second text information.
12. The method according to claim 11, characterized in that, The step of updating the first audio based on the second text information to obtain the second audio includes: The second text information is compared with the first text information to determine the deleted third character sequence and the newly added fourth character sequence in the first text information; Based on the time mapping relationship, the start time and end time of the fifth audio segment corresponding to the third character sequence in the first audio are determined. The start time of the fifth audio segment is the start time of the audio segment corresponding to the first character in the third character sequence in the first audio, and the end time of the fifth audio segment is the end time of the audio segment corresponding to the last character in the third character sequence in the first audio. Based on the start and end times of the fifth audio segment, the first audio segment is cut to obtain the fifth audio segment; The first audio, the first text information, and the fourth character sequence are input into the audio generation model, and a second supplementary audio segment corresponding to the fourth character sequence is generated through the audio generation model. The fifth audio segment in the first audio is replaced with the second supplementary audio segment to obtain the second audio.
13. The method according to claim 7 or 12, characterized in that, The audio generation model includes an encoder, a quantization module, a language model, a decoder, and a vocoder; The encoder is used to extract the acoustic features of the first audio. The quantization module is used to quantize the acoustic features to obtain quantized discrete codes; The language model is used to predict the audio encoding corresponding to a newly added character sequence in the first text information based on input features. The input features include a first character encoding sequence and a second character encoding sequence. The first character encoding sequence is obtained by segmenting the first text information into words, and the second character encoding sequence is obtained by segmenting the newly added character sequence into words. The newly added character sequence is either a second character sequence inserted into the first text information or a fourth character sequence used to replace a third character sequence in the first text information. The decoder is used to convert the audio encoding into predicted acoustic features; The vocoder is used to generate supplementary audio segments corresponding to the newly added character sequence based on the predicted acoustic features.
14. The method according to any one of claims 2-12, characterized in that, After responding to the text editing input, updating the first text information, and obtaining the second text information, the method further includes: The second text information is displayed in the audio editing interface; The original character sequence and the newly added character sequence in the second text information have different display modes; the original character sequence is the character sequence in the second text information that is consistent with the first text information, and the newly added character sequence is the second character sequence inserted into the first text information or the fourth character sequence used to replace the third character sequence in the first text information.
15. The method according to any one of claims 2-12, characterized in that, The method further includes: Based on the sample dataset and loss function, the initial model is trained to obtain the speech recognition model; The sample dataset includes multiple audio samples and corresponding text information samples. The multiple audio samples include audio recorded in at least two scenarios, and the audio samples include human voice features.
16. The method according to claim 15, characterized in that, The loss function is a joint loss function, wherein the joint loss function includes at least two loss function terms, each loss function term corresponds to a weight, and one of the at least two loss function terms includes forward relative entropy; The process of training the initial model based on the training dataset and loss function to obtain a speech recognition model includes: Extract acoustic feature samples from each audio sample in the sample dataset; The text information sample corresponding to each audio sample is segmented into words to obtain the character encoding sequence sample of each audio sample; Based on the acoustic feature samples and character encoding sequence samples of all audio samples in the sample dataset, multiple sets of training data are constructed; wherein each set of training data includes the acoustic feature samples and character encoding sequence samples of one audio sample; Based on the multiple sets of training data, the initial model is trained to obtain the predicted character encoding sequence corresponding to each set of training data; Having completed at least two rounds of training on the initial model, the forward relative entropy is calculated based on the predicted character encoding sequence obtained in the i-th round of training and the predicted character encoding sequence obtained in the (i-1)-th round of training; where i ≥ 2; Based on the forward relative entropy, update the entropy value of the forward relative entropy in the loss function to obtain the updated loss function; Based on the updated loss function, calculate the loss value of the initial model; If the loss value is greater than or equal to the loss threshold, the weight of each loss function term in the loss function is adjusted, and the initial model with adjusted parameter values is trained again based on the acoustic feature samples and the text sequence samples. If the loss value is less than the loss threshold, training is stopped, and a trained speech recognition model is obtained.
17. The method according to claim 16, characterized in that, The initial model includes an encoder, a decoder, and a joint network; The step of training the initial model based on the multiple sets of training data to obtain the predicted character encoding sequence corresponding to each set of training data includes: The encoder embeds the acoustic feature samples from each set of training data into a spatial vector to obtain the acoustic feature vector samples of each set of training data. The decoder embeds the character encoding sequence samples from each training data set into a spatial vector to obtain the character encoding sequence vector samples for each training data set. The joint network concatenates the acoustic feature vector samples and character encoding sequence vector samples of each training data set to obtain the concatenated vector of each training data set. The concatenated vector of each training data set is then mapped to the vocabulary space to obtain the predicted character encoding sequence of each training data set.
18. An audio processing apparatus, characterized in that, include: The receiving module is used to receive text editing input of the first text information of the first audio. The text update module is used to update the first text information in response to text editing input to obtain the second text information; An audio processing module is used to process the first audio based on the second text information to obtain a second audio.
19. The apparatus according to claim 18, characterized in that, The device further includes: The speech recognition module is used to input the first audio into the speech recognition model before receiving the text editing input of the first text information of the first audio, and to perform speech recognition on the first audio through the speech recognition model to obtain the first text information of the first audio. The alignment module is used to input the first audio and the first text information into the audio-text alignment model, and to perform time alignment on the first audio and the first text information through the audio-text alignment model to obtain the time mapping relationship between the first text information and the first audio. The time mapping relationship is used to indicate the start time and end time of the audio segment corresponding to each character in the first text information in the first audio. The display module is used to display the audio editing page interface, which includes a timeline, a waveform of the first audio, and the first text information. The timeline includes at least one time marker, and each character in the waveform and the first text information is aligned with the corresponding waveform and the corresponding time marker in the waveform. The receiving module is specifically used for: Receive text editing input for the first text information in the audio editing page interface.
20. The apparatus according to claim 19, characterized in that, The text editing input is an edit input that deletes the first character sequence in the first text information; The text update module is specifically used for: Delete the first character sequence from the first text information to obtain the second text information.
21. The apparatus according to claim 20, characterized in that, The first character sequence is a series of consecutive characters in the first text information; The audio processing module is specifically used for: The second text information is compared with the first text information to determine the first character sequence that was deleted from the first text information; Based on the time mapping relationship, the start time and end time of the first audio segment corresponding to the first character sequence are determined. The start time of the first audio segment is the start time of the audio segment corresponding to the first character in the first audio sequence, and the end time of the first audio segment is the end time of the audio segment corresponding to the last character in the first audio sequence. Based on the start and end times of the first audio segment, the first audio is cut to obtain the first audio segment; Delete the first audio segment from the first audio, and if a second audio segment is obtained, determine the second audio segment as the second audio; If at least two second audio segments are obtained, the audio obtained by splicing the at least two second audio segments is determined as the second audio.
22. The apparatus according to claim 20, characterized in that, The first character sequence includes at least two non-contiguous character subsequences; The audio processing module is specifically used for: The second text information is compared with the first text information to determine the at least two character subsequences that have been deleted from the first text information; Based on the time mapping relationship, the start time and end time of the third audio segment corresponding to each character subsequence are determined. The start time of the third audio segment corresponding to each character subsequence is the start time of the audio segment in the first audio corresponding to the first character in the character subsequence, and the end time of the third audio segment corresponding to each character subsequence is the end time of the audio segment in the first audio corresponding to the last character in the character subsequence. Based on the segment start time and segment end time of the at least one character subsequence, the first audio is cut to obtain a third audio segment corresponding to each character subsequence; Delete at least two third audio segments from the first audio, and if a fourth audio segment is obtained, determine the fourth audio segment as the second audio; If at least two fourth audio segments are obtained, the audio obtained by splicing the at least two fourth audio segments is determined as the second audio.
23. The apparatus according to claim 19, characterized in that, The text editing input is an editing input that inserts a second character sequence into the first text information; The text update module is specifically used for: The second character sequence is inserted into the first text information to obtain the second text information.
24. The apparatus according to claim 23, characterized in that, The audio processing module is specifically used for: The second text information is compared with the first text information to determine the second character sequence inserted into the first text information and the insertion position of the second character sequence; The first audio, the first text information, and the second character sequence are input into the audio generation model, and the first supplementary audio segment corresponding to the second character sequence is generated through the audio generation model. Based on the insertion position, the first supplementary audio segment is inserted into the first audio to obtain the second audio.
25. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the audio processing method as described in any one of claims 1-17.