Audio processing methods, apparatus, equipment and products
By acquiring audio and text, determining phoneme duration and predicting prosodic features, and extracting timbre features to synthesize audio, the problem of long training time in existing technologies is solved, and efficient audio synthesis is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-03-13
AI Technical Summary
Existing audio synthesis techniques require a large amount of training data, resulting in low time efficiency.
By acquiring audio and corresponding text, the phoneme duration in the phoneme sequence is determined, and prosodic features are predicted by combining text and language information. Timbre features are also extracted, and finally, audio is synthesized, reducing the dependence on training data.
It improves the efficiency of audio synthesis, makes the synthesized audio closer to the user's real audio, and reduces training time.
Smart Images

Figure CN119207370B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of audio processing technology, specifically relating to an audio processing method, apparatus, device, and product. Background Technology
[0002] Audio synthesis refers to the process of learning and training an artificial intelligence model on large-scale data to mimic the audio of a specified object, enabling it to be reproduced on electronic devices or media.
[0003] Current audio synthesis technology relies solely on the audio of a specified object, thus requiring a large amount of training data. This results in training sessions lasting several hours or even tens of hours to achieve relatively stable results, thereby enabling audio synthesis through artificial intelligence models.
[0004] It is evident that current audio synthesis technologies are time-consuming and inefficient. Summary of the Invention
[0005] The purpose of this application is to provide an audio processing method, apparatus, device, and product that can solve the technical problem of low efficiency in audio synthesis in related technologies.
[0006] In a first aspect, embodiments of this application provide an audio processing method, including:
[0007] Get the first audio file and the text corresponding to the first audio file;
[0008] Based on the first audio, determine the phoneme duration of each phoneme in the phoneme sequence corresponding to the text; based on the text, the phoneme duration of each phoneme, and the language information corresponding to the text, predict the prosodic features of the first audio.
[0009] The timbre of the first audio is extracted to obtain its timbre characteristics;
[0010] The text, timbre features, and prosodic features are synthesized to obtain the second audio.
[0011] Secondly, embodiments of this application provide an audio processing apparatus, including:
[0012] The acquisition module is used to acquire the first audio and the text corresponding to the first audio.
[0013] The determination module is used to determine the phoneme duration of each phoneme in the phoneme sequence corresponding to the text based on the first audio.
[0014] The prediction module is used to predict the prosodic features of the first audio based on the text, the phoneme duration of each phoneme, and the language information corresponding to the text.
[0015] The extraction module is used to extract the timbre of the first audio and obtain the timbre features of the first audio.
[0016] The synthesis module is used to synthesize text, timbre features, and prosodic features to obtain a second audio.
[0017] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the method as described in the first aspect.
[0018] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method as described in the first aspect.
[0019] Fifthly, embodiments of this application provide a chip, which includes a processor and a communication interface, the communication interface and the processor being coupled together, the processor being used to run programs or instructions to implement the steps of the method as described in the first aspect.
[0020] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the steps of the method as described in the first aspect.
[0021] In this embodiment, a first audio file and the corresponding text are obtained; the phoneme duration of each phoneme in the phoneme sequence corresponding to the text is determined based on the first audio file; the prosodic features of the first audio file are predicted based on the text, the phoneme duration of each phoneme, and the language information corresponding to the text; timbre extraction is performed on the first audio file to obtain its timbre features; and the text, timbre features, and prosodic features are synthesized to obtain a second audio file. This embodiment introduces the corresponding text to the audio file, and based on the audio, the phoneme duration of each phoneme corresponding to the text is obtained. By combining the phoneme-word relationship, the user's speaking habits and prosodic style can be learned, making the resynthesized audio as close as possible to the user's real audio, without requiring a large amount of audio training data, thus improving the efficiency of audio synthesis. Attached Figure Description
[0022] Figure 1 A flowchart illustrating an audio processing method provided in an embodiment of this application;
[0023] Figure 2 This is a schematic diagram of the structure of a global timbre feature extraction model provided in an embodiment of this application;
[0024] Figure 3 This is a schematic diagram of the structure of a decoder provided in an embodiment of this application;
[0025] Figure 4 A flowchart illustrating the training process of a duration prediction model provided in this application embodiment;
[0026] Figure 5 An alignment diagram of a sample audio, sample audio text, and sample phoneme sequence provided for an embodiment of this application;
[0027] Figure 6 A flowchart illustrating another audio processing method provided in this application embodiment;
[0028] Figure 7 A flowchart illustrating another audio processing method provided in this application embodiment;
[0029] Figure 8 A schematic diagram illustrating a prosodic feature prediction process provided in an embodiment of this application;
[0030] Figure 9 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of this application;
[0031] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0032] Figure 11 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0033] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0034] The terms "first," "second," etc., used in this application's specification are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, in the specification, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects have an "or" relationship.
[0035] The audio processing method provided by the embodiments of this application can be applied to scenarios where audio needs to be synthesized, such as voice navigation, novel reading, simultaneous interpretation, consultation broadcasting, short video production, film and television dubbing, virtual human voice generation, etc. Specifically, for a piece of audio, through the method of the embodiments of this application, an audio can be synthesized that is as similar as possible to the tone color, rhythm, etc. of the audio.
[0036] For example, user A enters the audio "Today is really too happy, go on a trip quickly!" in an electronic device. During the audio recording process, "zhen shi" is read as "zhen si" and "chu" is read as "cu". Using the method of the embodiments of this application, the synthesized audio obtained will also read "zhen shi" as "zhen si" and "chu" as "cu", that is, the pronunciation style is the same as the audio entered by the user.
[0037] Another example is that user B said "Today's weather is so bad, I'm a bit unhappy", with a rather sad tone. Then, for this audio, using the method of the embodiments of this application, the tone of the synthesized audio obtained is also sad.
[0038] The execution subject of the audio processing method provided by the embodiments of this application can be an electronic device such as a mobile phone, a tablet computer, a notebook computer, a handheld computer, an in-vehicle electronic device, etc. In some embodiments of this application, taking an electronic device as the execution subject to execute the information display method as an example, the audio processing method provided by the embodiments of this application is described.
[0039] Next, in conjunction with the accompanying drawings, the audio processing method provided by the embodiments of this application will be described in detail through specific embodiments and their application scenarios.
[0040] Figure 1 is a flowchart of an audio processing method provided by the embodiments of this application. As Figure 1 shown, the audio processing method may include the following steps:
[0041] S110. Obtain a first audio and text corresponding to the first audio.
[0042] S120. Determine the phoneme duration of each phoneme in the phoneme sequence corresponding to the text according to the first audio.
[0043] S130. Predict the prosodic features of the first audio according to the text, the phoneme duration of each phoneme, and the language information corresponding to the text.
[0044] S140. Extract the tone color of the first audio to obtain the tone color feature of the first audio.
[0045] S150. Perform synthesis processing on the text, the tone color feature, and the prosodic features to obtain a second audio.
[0046] Based on the audio, the embodiment of the present application introduces the text corresponding to the audio, and obtains the phoneme duration of each phoneme according to the audio and the text, and learns the speaking habits and prosody styles of the user in combination with the relationship between the phonetic characters, so that the re-synthesized audio is as close as possible to the user's real audio, thus improving the synthesis efficiency of the audio.
[0047] The above steps will be described in detail as follows:
[0048] In S110, the first audio here can be an audio segment pre-recorded by the user, or an audio segment intercepted from a video, or an audio of the user collected in real time. The duration of the first audio can be, for example, 5s, 10s, etc. The embodiment of the present application does not limit the language of the first audio. For example, it can be a Chinese audio, an English audio, a Chinese-English audio, or an audio in other languages.
[0049] The text here is the text corresponding to the first audio. Exemplarily, the first audio can be subjected to speech recognition through a speech recognition model to obtain the text corresponding to the first audio. The speech recognition model can adopt, for example, the whisper model, or other models, as long as it can perform speech recognition on the first audio to obtain the text corresponding to the first audio.
[0050] Considering that the first audio may contain white noise or background noise, which causes great interference to the first audio and further affects the subsequent audio synthesis effect, in some embodiments, before recognizing the first audio, the first audio is first subjected to noise reduction processing, and then the noise-reduced first audio is subjected to speech recognition to obtain the corresponding text.
[0051] In S120, exemplarily, each character in the text can be converted into pinyin. For example, the Chinese character "中" can be converted into pinyin "zhong1". If the text contains English, each word can be kept unchanged, that is, the English words are not converted.
[0052] After obtaining the converted text, the phoneme corresponding to each pinyin or word can be determined in combination with the pronunciation dictionary. The pronunciation dictionary is used to store the mapping relationship between the pinyin or word and the phoneme. Table 1 exemplarily lists the mapping relationship between the English word "human" and the Chinese character "中" and the phoneme:
[0053] Table 1 Partial mapping relationship between pinyin or word and phoneme
[0054] Words / Pinyin phoneme human HH Y UW1 M AH0 N zhong1 zh ong1
[0055] Arrange the phonemes in the order of each word or pinyin in the audio text to obtain the phoneme sequence.
[0056] In some embodiments, the text can be processed before being converted into a phoneme sequence. For example, text normalization can be used to normalize expressions, abbreviations, punctuation marks, dates, etc., in the text to obtain normalized text. For example, the date 08 / 23 can be normalized to August 23rd. Then, the normalized text is converted into a phoneme sequence according to the method described in the above embodiments.
[0057] After obtaining the phoneme sequence, the phoneme duration of each phoneme can be determined by combining it with the first audio.
[0058] For example, a duration prediction model can be used to determine the phoneme duration of each phoneme. Specifically, the first audio and phoneme sequence can be input into the duration prediction model, and the phoneme duration of each phoneme can be output.
[0059] For example, an alignment method can be used to align the first audio with the phoneme sequence to obtain the position of each phoneme in the first audio, and then obtain the phoneme duration of each phoneme based on the position.
[0060] In S130, prosodic features are used to characterize changes in the strength, length, tone, and speed of a sound.
[0061] Based on the text, the phoneme duration of each phoneme, and the language information corresponding to the text, the prosodic features of the first audio can be predicted, providing a basis for the subsequent synthesis of new audio. This application does not limit the specific prediction process; for example, the text, the phoneme duration of each phoneme, and the language information corresponding to the text can be input into the prosodic prediction model to predict the prosodic features of the first audio.
[0062] For example, the text, the phoneme duration of each phoneme, and the language information corresponding to the text can be matched with a prosodic template to obtain the prosodic features of the first audio. The prosodic template is used to store the relationship between the text, phoneme duration, language information, and prosodic features.
[0063] In S140, timbre features are used to characterize the quality of a sound, articulation style, or speaking style, etc. The embodiments of this application do not limit the timbre extraction process.
[0064] For example, a global timbre feature extraction model can be used to extract global timbre features of the first audio, which are used to characterize the overall timbre features of the first audio.
[0065] It should be understood that, generally speaking, a user's timbre is globally stable and unchanging, and each user's voice frequency only varies within a narrow and fixed range. Although a user's timbre is generally stable, it is worth noting that in certain situations, timbre may differ due to variations in the structure and position of the oral cavity during vocalization. For example, speaking different languages, experiencing significant emotional fluctuations, or varying speech speed can all lead to changes in timbre due to differences in the position of the mouth, tongue, and lips. Therefore, for example, a local timbre feature extraction model can be used to extract local timbre features of the first audio audio. These local timbre features are used to characterize the timbre of the first audio audio in specific local areas.
[0066] For example, the timbre features of a certain audio segment in the first audio can be extracted using a local timbre feature extraction model to obtain local timbre features.
[0067] For example, both global and local timbre features of the first audio can be extracted simultaneously, capturing both the overall characteristics of the timbre and local timbre variations, thereby making the resynthesized audio more similar to the original audio.
[0068] The structures of global timbre feature extraction models and local timbre feature extraction models can be similar. Taking the global timbre feature extraction model as an example, such as... Figure 2 As shown, the global timbre feature extraction model 200 may include a linear layer 201, a normalization layer 202, and a convolutional layer 203. The convolutional layer 203 may be composed of multiple convolutional neural network (CNN) blocks stacked together. In some embodiments, the convolutional layer 203 may be composed of four CNN blocks stacked together.
[0069] Before inputting the first audio file into the global timbre feature extraction model 200, the first audio file can be dimensionally transformed. For example, the first audio file can be converted into a vector of dimension T*D, where T represents the duration of the first audio file and D represents the feature dimension of the vector. Through the global timbre feature extraction model 200, the first audio file of dimension T*D can be compressed into a 1*D feature vector, which is used to represent the global timbre features.
[0070] In S150, the text, timbre features, and prosodic features are synthesized to obtain the second audio, which is a synthesized audio of the first audio.
[0071] For example, text, timbre features, and prosodic features can be input into an audio synthesis model to output synthesized audio as a second audio.
[0072] For example, the feature vectors corresponding to text, timbre features, and prosodic features can be concatenated to obtain a concatenated feature vector. Then, a decoder can be used to process the concatenated feature vector to convert it into a Melspectrogram. Finally, a vocoder can be used to process the Melspectrogram to convert it into audio as synthesized audio.
[0073] The decoder can be used Figure 3 The structure shown may be used in practice, but other structures may also be used. This application does not limit the specific structure.
[0074] like Figure 3 As shown, the decoder 300 may include multiple stacked decoding modules 301. In some embodiments, the decoder 300 may include four stacked decoding modules 301. For example, the decoding module 301 may include a linear layer 3011, a multi-head attention layer 3012, a convolutional layer 3013, and a normalization layer 3014. By inputting the concatenated feature vector into the decoder 300, the Mel spectrum can be output. The function of each layer can be found in related technologies, and will not be described in detail here.
[0075] The above method can make the timbre and rhythm of the second audio as similar as possible to those of the first audio, thus improving the audio synthesis effect.
[0076] In some embodiments, the above-described S120 may include the following steps:
[0077] The first audio and the phoneme sequence corresponding to the text are input into the second duration prediction model, which outputs the phoneme duration of each phoneme in the phoneme sequence.
[0078] The second duration prediction model is a trained model used to predict the phoneme duration of each phoneme in a phoneme sequence. Different phoneme durations can represent different speaking styles of users. For example, the second duration prediction model can employ a combination of a Gaussian Mixture Model (GMM) and Deep Learning (DL), or other models capable of predicting phoneme duration. Using the second duration prediction model, the phoneme duration of each phoneme can be automatically obtained, simplifying the user's operation and improving the accuracy of the results.
[0079] The training process of the second duration prediction model will be explained below. Figure 4 A flowchart illustrating the training process of a duration prediction model provided in this application embodiment. This training process can be executed before S110, such as... Figure 4 As shown, the training process may include the following steps S410-S440:
[0080] S410: Obtain multiple training samples of different durations.
[0081] Each training sample of a certain duration includes a sample audio, a sample phoneme sequence corresponding to the sample text, and a sample duration corresponding to the sample phoneme sequence. The sample text is the text corresponding to the sample audio.
[0082] For example, training samples of a certain duration can be obtained in the following way:
[0083] Obtain the sample audio and the corresponding sample text;
[0084] The sample text is segmented based on its language information to obtain at least one segmented text.
[0085] Search for each segmented text to obtain at least one phoneme corresponding to each segmented text; then, according to the order of each segmented text in the sample text, concatenate each phoneme corresponding to each segment to obtain the sample phoneme sequence corresponding to the sample text.
[0086] Determine the start and end positions of each phoneme in the sample phoneme sequence from the sample audio; obtain the sample duration of each phoneme based on the time difference between the start and end positions of each phoneme.
[0087] The sample audio, sample phoneme sequence, and sample duration corresponding to the sample phoneme sequence are determined as duration training samples.
[0088] There are multiple ways to obtain sample audio, and the duration of each sample audio can be the same or different.
[0089] During the model training phase, in order to more easily determine the relationship between sounds and characters, the sample text can be segmented based on the language information of the sample text to obtain at least one segmented text.
[0090] For example, if the sample text includes Chinese text, each character in the Chinese text can be identified as a segmented text; if the sample text includes English text, each word in the English text can be identified as a segmented text. For instance, when the segmented text is Chinese characters, the Chinese characters can be further converted into a concatenation; when the segmented text is English, the segmented text can remain unchanged.
[0091] For each segmented text, a pronunciation dictionary can be consulted based on its pinyin or words to obtain the phonemes corresponding to that segmented text, thereby obtaining the sample phoneme sequence corresponding to the sample text.
[0092] The process for determining the sample duration for each phoneme is explained below:
[0093] Exemplarily, the converted sample text and sample audio can be aligned using the Montreal Forced Aligner (MFA), where the converted sample audio text is text containing pinyin or words. After alignment, the start position and end position of each phoneme in the sample audio can be determined by combining the sample phoneme sequence and the sample audio. Based on the time difference between the end position and the start position, the sample duration of each phoneme can be obtained.
[0094] Taking the sample text "哇塞" as an example, where the pinyin of the Chinese character "哇" is "wa5", and the corresponding phoneme is "ua5", and the pinyin of the Chinese character "塞" is "sail", and the corresponding phoneme is "s ail". Figure 5 Exemplarily, an alignment schematic diagram of a sample audio, sample text, and sample phoneme sequence is listed, where 500 is the waveform of the audio "哇", 5,01 is the waveform of the audio "塞", 502 is the spectrum of the audio "哇", 503 is the spectrum of the audio "塞", 504 is the pinyin of the audio "哇", 505 is the pinyin of the audio "塞", 506 is the phoneme of the pinyin "wa5", and 507 is the phoneme of the pinyin "sail". Through alignment, the boundary of each phoneme in the sample audio can be determined, and then the sample duration of each phoneme can be obtained.
[0095] In the embodiments of this application, the alignment relationship between the sample audio and the sample text is obtained from the perspective of pinyin or words, and then the sample duration of each phoneme is obtained by combining the pronunciation dictionary, so that subsequent model training can be performed based on the sample duration of each phoneme, improving the training effect of the model.
[0096] S420. Input the sample phoneme sequence and the sample audio into the first duration prediction model, and output the predicted duration of each phoneme in the sample phoneme sequence.
[0097] In the specific training process, the sample phoneme sequence and the sample audio can be input into the first duration prediction model, and the predicted duration of each sample phoneme in the sample phoneme sequence is output.
[0098] S430. Determine the value of the duration loss function according to the sample duration and the predicted duration.
[0099] Exemplarily, according to the predicted duration and the sample duration, that is, the actual duration of each phoneme, combined with the loss function, the value of the duration loss function can be obtained. The embodiments of this application do not limit the specific loss function. For example, the mean square error loss function, the mean absolute value loss function, etc. can be used.
[0100] S440. Train the first duration prediction model according to the value of the duration loss function to obtain the second duration prediction model.
[0101] The model parameters can be updated based on the duration loss function value until the duration loss function value tends to converge or the number of iterations reaches the threshold, at which point training terminates and the trained duration prediction model can be obtained.
[0102] Before practical application, the embodiments of this application can train a first duration prediction model by combining the relationship between sound and word to obtain a second duration prediction model. This allows the second duration prediction model to be used to predict the phoneme duration of each phoneme corresponding to each text, thereby obtaining the user's speaking habits and improving the accuracy of synthesized audio.
[0103] Figure 6 A flowchart illustrating another audio processing method provided in this application embodiment. Figure 6 and Figure 1 The difference is that, Figure 1 S130 in the text can be further refined into Figure 6 S610-S640 in the series.
[0104] S610. Perform text encoding on the text to obtain the text feature vector corresponding to the text.
[0105] For example, a text encoder can be used to encode the text corresponding to the first audio, obtaining a text feature vector corresponding to that text. For example, the dimension of the text feature vector is T*D. This text feature vector contains the semantic information of the text.
[0106] S620. Vectorize the phoneme duration of each phoneme to obtain the duration feature vector of each phoneme; and vectorize the language information corresponding to the text to obtain the language feature vector corresponding to the language information.
[0107] The vectorization process here involves converting the one-dimensional phoneme duration and language information into vectors of specific dimensions, such as vectors with dimensions T*D. In other words, the text feature vector, duration feature vector, and language feature vector in this embodiment have the same dimension.
[0108] For example, phoneme duration can be input into the embedding layer to obtain the duration feature vector corresponding to each phoneme. Similarly, language information can be input into the embedding layer to obtain the language feature vector corresponding to the language information.
[0109] In practical applications, the execution order of S610 and S620 is not limited in the embodiments of this application. For example, S610 can be executed first and then S620 can be executed, or S620 can be executed first and then S610 can be executed, or S610 and S620 can be executed simultaneously.
[0110] S630. By concatenating the text feature vector, the duration feature vector of each phoneme, and the language feature vector according to the time dimension, a concatenated feature vector is obtained.
[0111] If the dimensions of the text feature vector, duration feature vector, and language feature vector are all T*D, then the dimension of the concatenated feature vector is T*3D.
[0112] S640. Input the concatenated feature vector into the prosody prediction model and output the prosody features of the first audio.
[0113] The prosody prediction model is obtained by training a large language model (LLM) with the sample spliced feature vector as input and the sample prosodic features of the sample audio as output.
[0114] This application embodiment continues training on the basis of a large language model to obtain a prosody prediction model. When predicting the prosodic features of the speaker other than the prosodic training samples, there is no need to retrain the model or fine-tune it. Moreover, since the large language model has been trained in advance using a large amount of corpus, when this application embodiment continues training on this basis, it is only necessary to obtain a short audio clip as a training sample.
[0115] This application's embodiments combine semantic information, phoneme duration information, and language information of the text with a large language model to predict prosodic features, making the prediction results closer to the user's prosodic style.
[0116] The training process of the prosody prediction model is explained below:
[0117] In some embodiments, prior to S110, the audio processing method may further include the following steps:
[0118] Multiple prosodic training samples are obtained. Each prosodic training sample includes a sample concatenated feature vector and sample prosodic features corresponding to the sample audio. The sample concatenated feature vector includes the sample duration feature vector of each phoneme, the sample language feature vector corresponding to the sample text, and the sample text feature vector corresponding to the sample text. The sample text is the text corresponding to the sample audio.
[0119] The first prosodic feature of the sample audio at the first sampling time and the sample concatenation vector corresponding to the second sampling time are input into the large language model, and the second prosodic feature of the sample audio at the second sampling time is output. The first sampling time and the second sampling time are the sampling times within the time period corresponding to the sample audio, and the first sampling time is the time before the second sampling time.
[0120] Based on the second prosodic feature and the sample prosodic feature of the sample audio at the second sampling time, the prosodic loss function value is determined;
[0121] Based on the prosody loss function value, a large language model is trained to obtain a prosody prediction model.
[0122] The process of generating the sample concatenation feature vector can be found in the above embodiments, and will not be repeated here for the sake of brevity. The sample prosodic features are the actual prosodic features contained in the sample audio. In this embodiment, the prosodic features of the sample audio at discrete times can be obtained, and different prosodic features can be represented by different tokens.
[0123] The first sampling time and the second sampling time are discrete sampling times of the time period corresponding to the sample audio, and the first sampling time and the second sampling time are adjacent sampling times. For example, the first sampling time can be the time before the second sampling time. Of course, in practical applications, the first sampling time can also be the time after the second sampling time.
[0124] The first prosodic feature is the prosodic feature of the predicted sample audio at the first sampling time, and the second prosodic feature is the prosodic feature of the predicted sample audio at the second sampling time. That is, the embodiments of this application can predict the prosodic feature of the sample audio at time t based on the prosodic features predicted before time t and the sample splicing vector at time t, that is, predict the prosodic feature through autoregression, thereby improving the accuracy of the prediction results.
[0125] For example, by inputting the concatenated vector of the first prosodic feature and the samples corresponding to the second sampling time into the large language model, the second prosodic feature output by the large language model can be obtained:
[0126] For example, the prosodic features at time t predicted by a large language model can be determined by the following formula:
[0127]
[0128] Where c t u represents the sample concatenation vector at time t. t Let θ be the prosodic token predicted at time t, also known as the second prosodic feature, and θ be the training parameters of the large language model. The main body of the large language model can adopt a decoder-only GPT structure. The prosodic feature prediction process can be found in [reference needed]. Figure 7 This approach uses autoregression to predict prosody, improving the accuracy of the prediction results. Here, s is empty, ci represents the sample concatenation vector at the i-th sampling time (1≤i≤n), n is the number of sampling times, and ui represents the prosody token at the i-th sampling time. Different prosody tokens represent different prosodic features.
[0129] Based on the second prosodic feature output by the large language model and the corresponding actual prosodic feature at the same time, combined with the prosodic loss function, the prosodic loss function value can be determined. Based on the prosodic loss function value, the model parameters θ of the large language model can be further refined until the training termination condition is met, thus obtaining the prosodic prediction model. The prosodic loss function can be a mean squared error loss function, a mean absolute value loss function, etc. The training termination condition could be, for example, that the prosodic loss function value tends to stabilize, or that the number of iterations reaches a pre-set threshold.
[0130] Based on the large language model, this application embodiment trains the model through autoregression, so that when predicting prosody in the future, the contextual information of the audio can be fully considered to obtain a prosodic style that is close to the user.
[0131] Figure 8 A flowchart illustrating another audio processing method provided in this application embodiment. Figure 8 and Figure 1 The difference is that, Figure 1 S140 in the middle can be further refined into Figure 8 S810-S830 in the series.
[0132] S810. Extract the timbre of the first audio to obtain the global timbre features of the first audio.
[0133] For example, a global timbre feature extraction model can be used to extract the global timbre of the first audio, thereby obtaining the global timbre features of the first audio. The structure of the global timbre feature extraction model can be found in the above embodiments, and will not be repeated here for the sake of brevity.
[0134] S820: Extract at least one audio segment from the first audio.
[0135] This application does not limit the duration, number, or position of audio segments within the first audio. For example, multiple audio segments can be extracted from the first audio according to different duration ratios.
[0136] In some embodiments, the above-described S820 may include the following steps:
[0137] Determine the duration of at least one audio segment being extracted;
[0138] Based on the duration of each audio segment, determine the starting position of each extracted audio segment from the first audio.
[0139] Starting from the initial position, extract audio segments from the first audio along the reference direction, each segment having the same duration as the previous audio segment.
[0140] For example, the segment duration may include, but is not limited to, 1 / 3, 1 / 4, 1 / 5, or 1 / 6 of the first audio duration. There may be one or more audio segments with the same segment duration.
[0141] The starting position of each audio segment can be determined based on its duration. Taking a segment length of 1 / 3 of the first audio duration as an example, if the segment is cut forward, its starting position can be any position between the beginning of the first audio segment and the 2 / 3 duration position. If the segment is cut backward, its starting position can be any position between the 2 / 3 duration position and the end position of the first audio segment. Forward cutting means cutting from the beginning to the end of the first audio segment, and backward cutting means cutting from the end to the beginning of the first audio segment. The process for determining the starting positions for other segment lengths is similar.
[0142] The reference direction here can be either forward or reverse. Taking forward cutting as an example, you can start from a determined starting position and cut audio segments of the corresponding duration along the forward direction.
[0143] When there are multiple audio segments corresponding to each segment duration, the starting position of each audio segment can be different.
[0144] The embodiments of this application can predetermine the segment length of the audio segments to be extracted, and then extract audio segments of different lengths from the first audio according to the segment length. This can improve the flexibility of audio segmentation and help capture the timbre changes that occur in the audio at different times, thereby improving the synthesis effect of subsequent audio.
[0145] S830. Extract the timbre of each audio segment to obtain the segment timbre features of each audio segment.
[0146] For each audio segment, a local timbre feature extraction model can be used to extract the timbre. In this way, the segmental timbre features of each audio segment can be obtained, that is, the local timbre features of the first audio segment. The structure of the local timbre feature extraction model can be found in the above embodiment, and will not be repeated here for the sake of brevity.
[0147] Taking the first audio as "The weather is so bad today, I'm a little unhappy" as an example, the extracted global timbre features include the user's fundamental frequency and overall timbre characteristics such as hoarseness, warmth, and brightness. When extracting local timbre features, the above-mentioned audio segmentation and truncation strategy can capture the timbre changes in the local part of the first audio, such as the timbre changes when the user speaks different languages.
[0148] The embodiments of this application employ a global and local timbre extraction scheme, which can not only capture the overall characteristics of the first timbre, but also capture the timbre changes of the first audio in local areas, thereby enabling a more comprehensive learning of the user's timbre characteristics and improving the accuracy of synthesized audio.
[0149] In this embodiment, audio synthesis considers not only the audio itself but also the corresponding text. This allows the system to learn the user's pronunciation habits from the relationship between sound and word. Furthermore, it extracts timbre at multiple scales, both holistically and locally, capturing both the overall characteristics and subtle variations of the user's timbre, resulting in a synthesized audio timbre that more closely resembles the user's actual timbre. In addition, this embodiment utilizes a large language model combined with an autoregressive approach for prosodic prediction, fully considering the contextual information of the audio to generate a prosodic style that closely matches the user's. Therefore, the solution in this embodiment achieves a realistic and lifelike synthesis effect using only short audio clips, such as 5 seconds, while significantly improving synthesis efficiency.
[0150] It should be noted that the audio processing method provided in this application embodiment can be executed by an audio processing device or a processing module within that audio processing device for executing the audio processing method. This application embodiment uses an audio processing device executing the audio processing method as an example to illustrate the audio processing device provided in this application embodiment.
[0151] Figure 9 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of this application.
[0152] like Figure 9 As shown, the audio processing device 900 may include:
[0153] Module 901 is used to acquire the first audio and the text corresponding to the first audio.
[0154] The determining module 902 is used to determine the phoneme duration of each phoneme in the phoneme sequence corresponding to the text based on the first audio.
[0155] The prediction module 903 is used to predict the prosodic features of the first audio based on the text, the phoneme duration of each phoneme, and the language information corresponding to the text.
[0156] Extraction module 904 is used to extract the timbre of the first audio and obtain the timbre features of the first audio.
[0157] Synthesis module 905 is used to synthesize text, timbre features and prosodic features to obtain a second audio.
[0158] In this embodiment, a first audio file and the corresponding text are obtained; the phoneme duration of each phoneme in the phoneme sequence corresponding to the text is determined based on the first audio file; the prosodic features of the first audio file are predicted based on the text, the phoneme duration of each phoneme, and the language information corresponding to the text; timbre extraction is performed on the first audio file to obtain its timbre features; and the text, timbre features, and prosodic features are synthesized to obtain a second audio file. This embodiment introduces the corresponding text to the audio file and, based on the audio, obtains the phoneme duration of each phoneme corresponding to the text. By combining the phoneme-word relationship, it learns the user's speaking habits and prosodic style, making the resynthesized audio as close as possible to the user's real audio, without requiring a large amount of audio training data, thus improving the efficiency of audio synthesis.
[0159] In some possible implementations of the embodiments of this application, the determining module 902 is specifically used for:
[0160] Input the first audio and the phoneme sequence corresponding to the text into the second duration prediction model, and output the phoneme duration of each phoneme in the phoneme sequence;
[0161] The acquisition module 901 is also used to acquire multiple duration training samples before acquiring the first audio and the text corresponding to the first audio. Each duration training sample includes the sample audio, the sample phoneme sequence corresponding to the sample text, and the sample duration corresponding to the sample phoneme sequence. The sample text is the text corresponding to the sample audio.
[0162] The determination module 902 is also used to input the sample phoneme sequence and sample audio into the first duration prediction model, output the predicted duration of each phoneme in the sample phoneme sequence, and determine the duration loss function value based on the sample duration and the predicted duration.
[0163] The audio processing device 900 may further include:
[0164] The training module is used to train the first duration prediction model based on the duration loss function value, and then obtain the second duration prediction model.
[0165] In some possible implementations of the embodiments of this application, the acquisition module 901 is specifically used for:
[0166] Obtain the sample audio and the corresponding sample text;
[0167] The audio processing device 900 may further include:
[0168] The segmentation module is used to segment the sample text according to the language information of the sample text to obtain at least one segmented text.
[0169] The lookup module is used to search for each segmented text to obtain at least one phoneme corresponding to each segmented text; and to concatenate each phoneme corresponding to each segment according to the order of each segmented text in the sample text to obtain the sample phoneme sequence corresponding to the sample text.
[0170] The determination module 902 is also used to determine the start and end positions of each phoneme in the sample phoneme sequence from the sample audio; to obtain the sample duration of each phoneme based on the time difference between the start and end positions of each phoneme; and to determine the sample audio, the sample phoneme sequence, and the sample duration corresponding to the sample phoneme sequence as duration training samples.
[0171] In some possible implementations of the embodiments of this application, the prediction module 903 is specifically used for:
[0172] Text encoding is performed on the text to obtain the corresponding text feature vector;
[0173] The phoneme duration of each phoneme is vectorized to obtain the duration feature vector of each phoneme; and the language information corresponding to the text is vectorized to obtain the language feature vector corresponding to the language information.
[0174] By concatenating the text feature vector, the duration feature vector of each phoneme, and the language feature vector along the time dimension, a concatenated feature vector is obtained.
[0175] Input the concatenated feature vector into the prosody prediction model and output the prosody features of the first audio.
[0176] The prosody prediction model is obtained by training a large language model with the sample splicing feature vector as input and the sample prosody features of the sample audio as output.
[0177] In some possible implementations of the embodiments of this application, the acquisition module 901 is further configured to acquire multiple prosodic training samples before acquiring the first audio and the text corresponding to the first audio. Each prosodic training sample includes a sample splicing feature vector and a sample prosodic feature corresponding to the sample audio. The sample splicing feature vector includes a sample duration feature vector for each phoneme, a sample language feature vector corresponding to the sample text, and a sample text feature vector corresponding to the sample text. The sample text is the text corresponding to the sample audio.
[0178] The training module is also used to input the first prosodic feature of the sample audio at the first sampling time and the sample concatenation vector corresponding to the second sampling time into the large language model, and output the second prosodic feature of the sample audio at the second sampling time. The first sampling time and the second sampling time are the sampling times within the time period corresponding to the sample audio, and the first sampling time is the time before the second sampling time. Based on the second prosodic feature and the sample prosodic feature of the sample audio at the second sampling time, the prosodic loss function value is determined. Based on the prosodic loss function value, the large language model is trained to obtain the prosodic prediction model.
[0179] In some possible implementations of the embodiments of this application, the extraction module 904 is specifically used for:
[0180] The timbre of the first audio is extracted to obtain the global timbre features of the first audio.
[0181] Extract at least one audio segment from the first audio;
[0182] Timbre extraction is performed on each audio segment to obtain the segment timbre features of each audio segment.
[0183] In some possible implementations of the embodiments of this application, the extraction module 904 is specifically used for:
[0184] Determine the duration of at least one audio segment being extracted;
[0185] Based on the duration of each audio segment, determine the starting position of each extracted audio segment from the first audio.
[0186] Starting from the initial position, extract audio segments from the first audio along the reference direction, each segment having the same duration as the previous audio segment.
[0187] In this embodiment, audio synthesis considers not only the audio itself but also the corresponding text. This allows the system to learn the user's pronunciation habits from the relationship between sound and word. Furthermore, it extracts timbre at multiple scales, both holistically and locally, capturing both the overall characteristics and subtle variations of the user's timbre, resulting in a synthesized audio timbre that more closely resembles the user's actual timbre. In addition, this embodiment utilizes a large language model combined with an autoregressive approach for prosodic prediction, fully considering the contextual information of the audio to generate a prosodic style that closely matches the user's. Therefore, the solution in this embodiment achieves a realistic and lifelike synthesis effect using only short audio clips, such as 5 seconds, while significantly improving synthesis efficiency.
[0188] The audio processing device in this application embodiment can be a device or a component in an electronic device, such as an integrated circuit or a chip. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.
[0189] The electronic device in this application embodiment can be a terminal with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0190] The audio processing device provided in this application embodiment can achieve... Figures 1 to 8 The various processes in the audio processing method embodiments can achieve the same technical effect, and will not be described again here to avoid repetition.
[0191] like Figure 10 As shown, this application embodiment also provides an electronic device 1000, including a processor 801 and a memory 1002. The memory 1002 stores programs or instructions that can run on the processor 1001. When the program or instructions are executed by the processor 1001, they implement the various steps of the above-described audio processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0192] It should be noted that the electronic devices in the embodiments of this application include the mobile terminals and non-mobile terminals mentioned above.
[0193] Figure 11 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application.
[0194] The electronic device 1100 includes, but is not limited to, components such as: radio frequency unit 1101, network module 1102, audio output unit 1103, input unit 1104, sensor 1105, display unit 1106, user input unit 1107, interface unit 1108, memory 1109, and processor 1110.
[0195] Those skilled in the art will understand that the electronic device 1100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 11 The structure of the electronic device 1100 shown does not constitute a limitation on the electronic device 1100. The electronic device 1100 may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be described in detail here.
[0196] The processor 1110 is used to acquire the first audio and the text corresponding to the first audio.
[0197] The phoneme duration of each phoneme in the phoneme sequence corresponding to the text is determined based on the first audio.
[0198] Based on the text, the phoneme duration of each phoneme, and the language information of the audio text, predict the prosodic features of the first audio.
[0199] The timbre of the first audio is extracted to obtain its timbre characteristics;
[0200] The text, timbre features, and prosodic features are synthesized to obtain the synthesized audio of the first audio.
[0201] In this embodiment, a first audio file and the corresponding text are obtained; the phoneme duration of each phoneme in the phoneme sequence corresponding to the text is determined based on the first audio file; the prosodic features of the first audio file are predicted based on the text, the phoneme duration of each phoneme, and the language information corresponding to the text; timbre extraction is performed on the first audio file to obtain its timbre features; and the text, timbre features, and prosodic features are synthesized to obtain a second audio file. This embodiment introduces the corresponding text to the audio file and, based on the audio, obtains the phoneme duration of each phoneme corresponding to the text. By combining the phoneme-word relationship, it learns the user's speaking habits and prosodic style, making the resynthesized audio as close as possible to the user's real audio, without requiring a large amount of audio training data, thus improving the efficiency of audio synthesis.
[0202] In some possible implementations of embodiments of this application, the processor 1110 is specifically used for:
[0203] Input the first audio and the phoneme sequence corresponding to the text into the second duration prediction model, and output the phoneme duration of each phoneme in the phoneme sequence;
[0204] Before obtaining the first audio and the text corresponding to the first audio, multiple duration training samples are obtained. Each duration training sample includes the sample audio, the sample phoneme sequence corresponding to the sample text, and the sample duration corresponding to the sample phoneme sequence. The sample text is the text corresponding to the sample audio.
[0205] Input the sample phoneme sequence and sample audio into the first duration prediction model, and output the predicted duration of each phoneme in the sample phoneme sequence;
[0206] The duration loss function value is determined based on the sample duration and the prediction duration;
[0207] Based on the duration loss function value, the first duration prediction model is trained to obtain the second duration prediction model.
[0208] In some possible implementations of embodiments of this application, the processor 1110 is specifically used for:
[0209] Obtain the sample audio and the corresponding sample text;
[0210] The sample text is segmented based on its language information to obtain at least one segmented text.
[0211] Search for each segmented text to obtain at least one phoneme corresponding to each segmented text; then, according to the order of each segmented text in the sample text, concatenate each phoneme corresponding to each segment to obtain the sample phoneme sequence corresponding to the sample text.
[0212] Determine the start and end positions of each phoneme in the sample phoneme sequence from the sample audio; obtain the sample duration of each phoneme based on the time difference between the start and end positions of each phoneme.
[0213] The sample audio, sample phoneme sequence, and sample duration corresponding to the sample phoneme sequence are determined as duration training samples.
[0214] In some possible implementations of embodiments of this application, the processor 1110 is specifically used for:
[0215] Text encoding is performed on the text to obtain the corresponding text feature vector;
[0216] The phoneme duration of each phoneme is vectorized to obtain the duration feature vector of each phoneme; and the language information corresponding to the text is vectorized to obtain the language feature vector corresponding to the language information.
[0217] By concatenating the text feature vector, the duration feature vector of each phoneme, and the language feature vector along the time dimension, a concatenated feature vector is obtained.
[0218] Input the concatenated feature vector into the prosody prediction model and output the prosody features of the first audio.
[0219] The prosody prediction model is obtained by training a large language model with the sample splicing feature vector as input and the sample prosody features of the sample audio as output.
[0220] In some possible implementations of embodiments of this application, the processor 1110 is specifically used for:
[0221] Before obtaining the first audio and the text corresponding to the first audio, multiple prosodic training samples are obtained. Each prosodic training sample includes a sample concatenation feature vector and a sample prosodic feature corresponding to the sample audio. The sample concatenation feature vector includes the sample duration feature vector of each phoneme in the sample text, the sample language feature vector corresponding to the sample text, and the sample text feature vector corresponding to the sample text. The sample text is the text corresponding to the sample audio.
[0222] The first prosodic feature of the sample audio at the first sampling time and the sample concatenation vector corresponding to the second sampling time are input into the large language model, and the second prosodic feature of the sample audio at the second sampling time is output. The first sampling time and the second sampling time are the sampling times within the time period corresponding to the sample audio, and the first sampling time is the time before the second sampling time.
[0223] Based on the second prosodic feature and the sample prosodic feature of the sample audio at the second sampling time, the prosodic loss function value is determined;
[0224] Based on the prosody loss function value, a large language model is trained to obtain a prosody prediction model.
[0225] In some possible implementations of embodiments of this application, the processor 1110 is specifically used for:
[0226] The timbre of the first audio is extracted to obtain the global timbre features of the first audio.
[0227] Extract at least one audio segment from the first audio;
[0228] Timbre extraction is performed on each audio segment to obtain the segment timbre features of each audio segment.
[0229] In some possible implementations of embodiments of this application, the processor 1110 is specifically used for:
[0230] Determine the duration of at least one audio segment being extracted;
[0231] Based on the duration of each audio segment, determine the starting position of each extracted audio segment from the first audio.
[0232] Starting from the initial position, extract audio segments from the first audio along the reference direction, each segment having the same duration as the previous audio segment.
[0233] In this embodiment, audio synthesis considers not only the audio itself but also the corresponding text. This allows the system to learn the user's pronunciation habits from the relationship between sound and word. Furthermore, it extracts timbre at multiple scales, both holistically and locally, capturing both the overall characteristics and subtle variations of the user's timbre, resulting in a synthesized audio timbre that more closely resembles the user's actual timbre. In addition, this embodiment utilizes a large language model combined with an autoregressive approach for prosodic prediction, fully considering the contextual information of the audio to generate a prosodic style that closely matches the user's. Therefore, the solution in this embodiment achieves a realistic and lifelike synthesis effect using only short audio clips, such as 5 seconds, while significantly improving synthesis efficiency.
[0234] It should be understood that, in this embodiment, the input unit 1104 may include a graphics processing unit (GPU) 11041 and a microphone 11042. The GPU 11041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1106 may include a display panel 11061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1107 includes at least one of a touch panel 11071 and other input devices 11072. The touch panel 11071 is also called a touch screen. The touch panel 11071 may include a touch detection device and a touch controller. Other input devices 11072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0235] The memory 1109 can be used to store software programs and various data. The memory 1109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1109 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0236] Processor 1110 may include one or more processing units; optionally, processor 1110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1110.
[0237] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described audio processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0238] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0239] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described audio processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0240] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0241] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the audio processing method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0242] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0243] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0244] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An audio processing method, characterized in that, include: Get the first audio and the text corresponding to the first audio; The phoneme duration of each phoneme in the phoneme sequence corresponding to the text is determined based on the first audio. Based on the text, the phoneme duration of each phoneme, and the language information corresponding to the text, predict the prosodic features of the first audio. The timbre of the first audio is extracted to obtain the timbre features of the first audio. The text, the timbre features, and the prosodic features are synthesized to obtain a second audio. The step of predicting the prosodic features of the first audio based on the text, the phoneme duration of each phoneme, and the language information corresponding to the text includes: The text is encoded to obtain the text feature vector corresponding to the text; The phoneme duration of each phoneme is vectorized to obtain a duration feature vector for each phoneme; and the language information corresponding to the text is vectorized to obtain a language feature vector corresponding to the language information. The concatenated feature vector is obtained by concatenating the text feature vector, the duration feature vector of each phoneme, and the language feature vector along the time dimension. The concatenated feature vector is input into the prosody prediction model, and the prosody features of the first audio are output. The prosody prediction model is obtained by training a large language model with sample splicing feature vectors as input and sample prosody features of sample audio as output.
2. The method according to claim 1, characterized in that, The step of determining the phoneme duration of each phoneme in the phoneme sequence corresponding to the text based on the first audio includes: The first audio and the phoneme sequence corresponding to the text are input into the second duration prediction model, and the phoneme duration of each phoneme in the phoneme sequence is output. Before obtaining the first audio and the text corresponding to the first audio, the method further includes: Multiple training samples of varying durations are obtained. Each training sample of varying duration includes a sample audio, a sample phoneme sequence corresponding to the sample text, and a sample duration corresponding to the sample phoneme sequence. The sample text is the text corresponding to the sample audio. The sample phoneme sequence and the sample audio are input into the first duration prediction model, and the predicted duration of each phoneme in the sample phoneme sequence is output. The duration loss function value is determined based on the sample duration and the predicted duration; Based on the duration loss function value, the first duration prediction model is trained to obtain the second duration prediction model.
3. The method according to claim 2, characterized in that, Obtain training samples of varying durations, including: Obtain the sample audio and the sample text corresponding to the sample audio; The sample text is segmented based on its language information to obtain at least one segmented text. Based on each segmented text, at least one phoneme corresponding to each segmented text is obtained; according to the order of each segmented text in the sample text, each phoneme corresponding to each segment is concatenated to obtain the sample phoneme sequence corresponding to the sample text. The start and end positions of each phoneme in the sample phoneme sequence are determined from the sample audio; the sample duration of each phoneme is obtained based on the time difference between the start and end positions of each phoneme. The sample audio, the sample phoneme sequence, and the sample duration corresponding to the sample phoneme sequence are determined as duration training samples.
4. The method according to claim 1, characterized in that, Before obtaining the first audio and the text corresponding to the first audio, the method further includes: Multiple prosodic training samples are obtained. Each prosodic training sample includes a sample concatenation feature vector and sample prosodic features corresponding to the sample audio. The sample concatenation feature vector includes a sample duration feature vector for each phoneme, a sample language feature vector corresponding to the sample text, and a sample text feature vector corresponding to the sample text. The sample text is the text corresponding to the sample audio. The first prosodic feature of the sample audio at the first sampling time and the sample concatenation vector corresponding to the second sampling time are input into the large language model, and the second prosodic feature of the sample audio at the second sampling time is output. The first sampling time and the second sampling time are the sampling times within the time period corresponding to the sample audio, and the first sampling time is the time before the second sampling time. Based on the second prosodic feature and the sample prosodic feature of the sample audio at the second sampling time, the prosodic loss function value is determined; Based on the prosody loss function value, the large language model is trained to obtain the prosody prediction model.
5. The method according to claim 1, characterized in that, The step of extracting the timbre of the first audio to obtain its timbre features includes: The timbre of the first audio is extracted to obtain the global timbre features of the first audio. Extract at least one audio segment from the first audio; Timbre extraction is performed on each audio segment to obtain the segment timbre features of each audio segment.
6. The method according to claim 5, characterized in that, Extracting at least one audio segment from the first audio includes: Determine the duration of at least one audio segment being extracted; Based on the duration of each audio segment, determine the starting position of each audio segment extracted from the first audio. Starting from the aforementioned starting position, audio segments with the same duration as each audio segment are extracted from the first audio along the reference direction.
7. An audio processing device, characterized in that, include: The acquisition module is used to acquire the first audio and the text corresponding to the first audio. The determining module is used to determine the phoneme duration of each phoneme in the phoneme sequence corresponding to the text based on the first audio. The prediction module is used to predict the prosodic features of the first audio based on the text, the phoneme duration of each phoneme, and the language information corresponding to the text. The extraction module is used to extract the timbre of the first audio and obtain the timbre features of the first audio. A synthesis module is used to synthesize the text, the timbre features, and the prosodic features to obtain a second audio. The prediction module is specifically used for: The text is encoded to obtain the text feature vector corresponding to the text; The phoneme duration of each phoneme is vectorized to obtain a duration feature vector for each phoneme; and the language information corresponding to the text is vectorized to obtain a language feature vector corresponding to the language information. The concatenated feature vector is obtained by concatenating the text feature vector, the duration feature vector of each phoneme, and the language feature vector along the time dimension. The concatenated feature vector is input into the prosody prediction model, and the prosody features of the first audio are output. The prosody prediction model is obtained by training a large language model with sample splicing feature vectors as input and sample prosody features of sample audio as output.
8. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing programs or instructions that can run on the processor, the programs or instructions being executed by the processor to implement the steps of the method as described in any one of claims 1 to 6.
9. A computer program product, characterized in that, The program product is stored in a storage medium and is executed by at least one processor to implement the steps of the method as claimed in any one of claims 1 to 6.
Citation Information
Patent Citations
Cross-language speech synthesis method and device, electronic equipment and storage medium
CN114203153A