Voice editing method and device, electronic equipment and readable storage medium
By combining the masked original audio with the text to be synthesized to generate and convert the edited audio, the problem of insufficient voice editing quality in the existing technology is solved, high-quality voice editing is achieved, and custom character name voice needs in the fields of games, film and television, etc.
Patent Information
- Application Number
- CN202510166411.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-09
AI Technical Summary
The existing technology is difficult to meet the high-quality voice editing needs of customized character names in the fields of games, film and television, and the consistency between the characteristics such as tone and tone and the original audio is insufficient, resulting in the dubbing scheme based on voice synthesis being unacceptable.
By combining the masked original audio with the text to be synthesized, edited audio is generated and converted to obtain the target audio, ensuring that the target audio meets user needs in terms of sound quality, intonation, etc.
It improves the quality of voice editing, enhances the flexibility and personalization of audio editing, makes the target audio more in terms of tone, intonation, etc., and improves the effect of audio editing.
Smart Images

Figure CN119964548A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of voice processing technology, and more specifically, relates to a voice editing method and device, an electronic device, and a readable storage medium. Background Art
[0002] With the continuous advancement of speech synthesis technology, its application in the fields of intelligent assistants, novel dubbing, etc. is becoming more and more extensive. However, in the fields of games, movies and TV shows that require high audio quality and expressiveness, the existing technology is difficult to meet the needs of users for customized character names, and there are deficiencies in the consistency of features such as timbre and tone with the original audio, resulting in the dubbing scheme based on speech synthesis often not being accepted by game manufacturers.
[0003] Therefore, a high-quality speech editing method is urgently needed. Summary of the invention
[0004] The purpose of the present disclosure is to provide a voice editing method and device, an electronic device, and a readable storage medium to improve the quality of voice editing.
[0005] According to a first aspect of the embodiments of the present disclosure, a method for voice editing is provided, comprising: Determine the edited audio based on the first audio and the text to be synthesized, where the first audio is the audio after masking the original audio; Converting the edited audio to obtain target audio; The target audio is output.
[0006] According to a second aspect of the embodiments of the present disclosure, a voice editing device is provided, comprising: An audio editing module, used for determining an edited audio based on a first audio and a text to be synthesized, wherein the first audio is an audio after masking processing is performed on the original audio; An audio conversion module, used for converting the edited audio to obtain a target audio; An audio output module is used to output the target audio.
[0007] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of the above-mentioned voice editing method when executing the computer program.
[0008] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned voice editing method are implemented.
[0009] The beneficial effects of the voice editing method and device, electronic device, and readable storage medium provided by the embodiments of the present disclosure are: The present disclosure combines the masked audio (i.e., the first audio) with the text to be synthesized, which is conducive to the subsequent more accurate synthesis of the edited audio that meets the requirements. At the same time, the present disclosure converts the edited audio to obtain the target audio, which not only retains the information of the text to be synthesized, but also makes the target audio more in line with user needs in terms of sound quality, intonation, etc., thereby improving the flexibility and personalization of audio editing. Therefore, the present disclosure can improve the quality of voice editing. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0011] Figure 1 A flowchart of a voice editing method provided by an embodiment of the present disclosure; Figure 2 A structural block diagram of a voice editing device provided by an embodiment of the present disclosure; Figure 3 A schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0012] In the following description, specific details such as specific system structures and technologies are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present disclosure. However, it should be clear to those skilled in the art that the present disclosure may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obstructing the description of the present disclosure with unnecessary details.
[0013] In order to make the purpose, technical solutions and advantages of the present disclosure more clear, specific embodiments will be described below in conjunction with the accompanying drawings.
[0014] Please refer to Figure 1 , Figure 1 A flowchart of a voice editing method provided by an embodiment of the present disclosure is provided, and the method includes: S101: Determine an edited audio based on a first audio and a text to be synthesized, where the first audio is the audio after masking the original audio.
[0015] In this embodiment, the first audio is the audio obtained after the original audio is masked. The masking process is a preprocessing operation on the original audio, which is used to extract or convert certain key information in the original audio, such as timbre, intonation or speech speed, so that it can be better combined with the text to be synthesized for voice editing. The masking process method can be filtering or transforming the spectrum of the original audio, or data screening or feature extraction based on the acoustic features of the original audio.
[0016] The text to be synthesized is the text content that needs to be converted into speech, including the semantic information to be expressed, grammatical structure, and possible emotions. The text to be synthesized will serve as the basis for speech synthesis and work together with the first audio to generate the edited audio. The text to be synthesized can come from various scenarios, such as the script of the game character, the narration or dialogue text in the film and television drama, the commentary of the short video, etc.
[0017] This embodiment uses the masked original audio (i.e., the first audio) and the text to be synthesized as input, and uses the speech editing model to generate the edited audio. This embodiment needs to comprehensively consider the audio feature information contained in the first audio and the semantics and grammar of the text to be synthesized, so that the generated edited audio can match the text to be synthesized in terms of timbre, intonation, speech speed, etc., and conform to the characteristics of natural speech as much as possible.
[0018] Suppose you are dubbing a game character. The original audio is a basic voice material recorded by a professional voice actor, such as a general voice clip with neutral emotion and medium speaking speed.
[0019] First, the original audio is masked to obtain the first audio. For example, the rhythmic features (such as pitch contour, rhythm pattern, etc.) and some timbre features (such as formant frequency, etc.) in the original audio are extracted through an audio analysis algorithm, and the above features are saved in a data format to form the first audio.
[0020] Then, the text to be synthesized is the lines of the game character in a specific scene, such as "There are enemies ahead, get ready to fight!".
[0021] Next, a deep learning-based speech editing model is used. During the training process, the model learns the correspondence between a large number of different audio features and text semantics, emotions, etc. The first audio and the text to be synthesized are input into the speech editing model. The model determines that this is an urgent and warning sentence based on the semantics of the text, and should be expressed with a higher pitch, faster speaking speed, and stronger tone. At the same time, combined with the rhythm and timbre features in the first audio, the model generates an edited audio, the timbre of which retains some characteristics of the original audio to a certain extent, but is highly consistent with the text to be synthesized in terms of tone, speaking speed, and emotional expression, and sounds like a game character nervously issuing a warning. At this point, the generation process from the first audio and the text to be synthesized to the edited audio is completed.
[0022] S102: Convert the edited audio to obtain the target audio.
[0023] In this embodiment, the edited audio is an audio generated based on the first audio and the text to be synthesized, and has preliminarily possessed the speech content corresponding to the text to be synthesized, but it still needs further optimization in some aspects, such as the consistency of timbre and the naturalness of audio transition, and further processing is needed to improve the audio quality.
[0024] The target audio is the final audio obtained after converting the edited audio, and is also the expected result. At this time, the target audio meets the requirements of the actual application scenario in terms of sound quality, timbre, and naturalness. For example, in the game dubbing scene, the target audio should be able to blend into the game scene naturally and smoothly, so that players can't feel obvious traces of synthesis, and the timbre should be highly matched with the game character setting.
[0025] This embodiment converts the generated edited audio to further optimize the problems existing in the edited audio and make the edited audio reach the expected high-quality audio standard, and finally obtain the target audio that can be directly applied to the actual scene. The conversion process can use an audio processing algorithm to improve the audio quality by adjusting various parameters and features of the audio.
[0026] Taking game dubbing as an example, the edited audio of the game character's lines "There are enemies ahead, prepare for battle!" has been obtained, but it is found that the timbre of the edited audio is somewhat different from the set timbre of the character in the game, and the audio transition at the beginning and end of the sentence is not natural enough.
[0027] At this time, the sound conversion model is used to convert the edited audio. The model first analyzes the timbre characteristics, audio spectrum and other information of the edited audio, and refers to the timbre characteristics of the original audio (or other standard audio samples of the game character), and adjusts the spectral parameters and formant frequency of the edited audio to gradually bring the timbre of the edited audio closer to the target timbre.
[0028] In terms of processing audio transitions, the model uses audio fade-in and fade-out technology and spectrum smoothing of the audio edge, making the audio at the beginning and end of a sentence sound more natural and smooth, without any abruptness. After the above conversion processing, the target audio obtained perfectly matches the game character in terms of timbre, the audio transition is natural, and it can be well integrated into the game scene, becoming a high-quality voice clip that can be used directly in the game, greatly improving the audio experience of the game.
[0029] S103: Output the target audio.
[0030] In this embodiment, the target audio can be sent to a designated location or device so that the target audio can be used by other systems or users, ensuring that the target audio can play a role in practical applications, such as being correctly played in scenes such as games and film and television production. For example, when a game character enters a battle scene, the game engine recognizes that the previously generated target audio "There are enemies ahead, get ready for battle!" needs to be played, and the audio file will be extracted from the resource library and played according to the audio settings of the game (such as volume, channels, etc.), so that players can hear clear, natural voice prompts that match the characteristics of the character, thereby enhancing the immersion and experience of the game.
[0031] After outputting the target audio, it also includes: Performing quality assessment on the target audio to obtain an audio quality assessment result; If the audio quality assessment result meets the audio quality standard, the corresponding speech editing method is retained; If the audio quality assessment result does not meet the audio quality standard, the corresponding voice editing method is adjusted.
[0032] In this embodiment, it is necessary to evaluate the quality of the target audio during the actual application process to determine whether the target audio meets the requirements of the actual application. The quality assessment can use a mathematical model or a machine learning model to detect and analyze various indicators of the target audio to determine the quality of the target audio. The indicators may include the clarity of the audio, the naturalness of the timbre, the stability of the volume, the degree of matching with the text content, and whether there is noise or distortion. The audio quality can be expressed as a quantitative score, grade, or qualitative description (such as excellent, good, qualified, unqualified) to represent the volume of the target audio.
[0033] The volume quality standard is a standard for measuring the target audio quality, that is, a series of specifications and requirements for measuring the target audio quality that are predetermined in actual application scenarios. For example, in film and television production, the audio quality standard can be that the background noise is less than a preset noise threshold, and the synchronization between the target audio and the lip shape is greater than a preset synchronization threshold.
[0034] The speech editing method refers to the model and parameter settings used in the entire process from original audio processing to target audio generation, including the masking processing method of the original audio, the generation algorithm of the edited audio and the parameters of the sound conversion model.
[0035] In response to the audio quality assessment result being a plurality of evaluation indicators, each evaluation indicator corresponds to an evaluation standard; If the evaluation index is the formant frequency of the audio spectrum, the formant frequency is compared with the standard formant frequency; If the evaluation index is the spectral smoothness of the audio edge, the spectral smoothness is compared with the standard spectral smoothness; If the evaluation indicator is the fit of intonation, the fit is compared with the standard fit.
[0036] Therefore, each evaluation index is compared with the corresponding evaluation standard to determine whether the audio quality evaluation result meets the audio quality standard, and then further processing is performed based on the judgment result. In addition, the standard formant frequency, the standard spectrum smoothness and the standard fit are all audio quality standards.
[0037] If the audio quality evaluation result does not meet the audio quality standard, the corresponding voice editing method is adjusted, including: If any of the multiple evaluation indicators does not meet the audio quality standard, the parameters of the speech editing model or the sound conversion model are adjusted.
[0038] The parameters of the speech editing model affect the generation result of the edited audio, and the parameters of the sound conversion model affect the result of the audio conversion. Therefore, adjusting the respective parameters can change the generation result of the target audio, thereby improving the quality of the target audio.
[0039] It can be concluded from the above that the present disclosure combines the masking audio (i.e., the first audio) with the text to be synthesized, which is conducive to the subsequent more accurate synthesis of the edited audio that meets the requirements. At the same time, the present disclosure converts the edited audio to obtain the target audio, which not only retains the information of the text to be synthesized, but also makes the target audio more in line with user needs in terms of sound quality, intonation, etc., thereby improving the flexibility and personalization of audio editing. Therefore, the present disclosure can improve the quality of voice editing.
[0040] In one embodiment of the present disclosure, the voice editing method further includes: Perform masking processing based on the original audio to determine masked audio, and use the masked audio as the first audio; Determine the basic characteristics of the original audio; Audio characteristics of the first audio are determined.
[0041] In this embodiment, the basic features of the original audio are some basic acoustic features of the original audio itself, which may include timbre, pitch, rhythm, and volume, etc. The above features can be analyzed and extracted from the waveform, spectrum, etc. of the original audio. Among them, timbre is the characteristic texture of the sound, which is determined by factors such as the ratio of different frequency costs; pitch is the high and low frequency of the sound; rhythm is the length and strength of the sound in the time dimension; volume is the loudness of the sound.
[0042] The audio features of the first audio are some acoustic features that are different from the original audio or have their own characteristics, and may include not only timbre, pitch, rhythm and volume, but also energy distribution of frequency bands and rhythmic features.
[0043] Suppose you want to dub an animated character. The original audio is a relatively neutral narration voice material recorded by a professional voice actor, which is about 30 seconds long.
[0044] First, this embodiment adopts a masking processing method based on spectrum analysis, converts the original audio from the time domain to the frequency domain through fast Fourier transform, and analyzes its spectrum characteristics. Then, some low-frequency background noise bands and the parts of excessively high frequencies that may cause the sharpness of the audio are filtered out through filters, and some key resonance peak frequency information of the intermediate frequency part is extracted. The resonance peak frequency information can be used to reflect the timbre, and the audio that is converted back to the time domain after the above processing is used as the masked audio (that is, the first audio).
[0045] Then, this embodiment uses an audio analysis algorithm to extract the basic features of the original audio. From the perspective of timbre, the proportional relationship of different frequency components in its spectrum is analyzed, and it is found that the timbre is relatively mellow and soft; in terms of pitch, the average pitch is in a moderate range, without particularly high or low pitch changes; in terms of rhythm, the pauses and stress distribution of sentences are relatively regular, in line with the normal rhythm of language expression; in terms of volume, the average loudness is roughly stable at a suitable level, and the overall sound is relatively comfortable.
[0046] Finally, when analyzing the audio features of the first audio, it was found that after the masking process, the background noise was effectively reduced, and the audio as a whole was purer. In the spectrum, the frequencies of the intermediate-frequency resonance peaks that were screened and retained are more prominent, forming a relatively clearer timbre characteristic performance, which helps to highlight the recognition of the character's voice when generating the edited audio later. Moreover, due to the filtering out of some unnecessary frequency components, the rhythmic sense of the rhythm has also changed slightly. The rhythmic characteristics of the audio have become more concise and clear, which is more conducive to matching the semantic rhythm of the text to be synthesized, thus laying a good foundation for the subsequent generation of edited audio that conforms to the characteristics of the animation character.
[0047] It can be concluded from the above that this embodiment can better combine with the text to be synthesized by masking the original audio, which is conducive to the subsequent more accurate synthesis of edited audio that meets the requirements. Secondly, clarifying the basic characteristics of the original audio provides an important reference for subsequent audio processing and helps to maintain the consistency and accuracy of audio processing. Furthermore, the audio characteristics of the first audio are determined so that the edited audio can retain the core information of the original audio while also having specific audio properties to meet diverse usage requirements. In summary, this method shows significant beneficial effects in improving processing accuracy and meeting diverse needs.
[0048] In one embodiment of the present disclosure, the voice editing method further includes: determining a plurality of original audio segments based on basic features of the original audio; determining a plurality of first audio segments based on the audio feature of the first audio; determining a degree of matching based on each original audio segment and each first audio segment; If the matching degree is greater than or equal to the first matching threshold, taking the first audio segment corresponding to the matching degree as the basic audio segment, and determining the first audio feature of the basic audio segment; If the matching degree is less than the first matching threshold, the first audio segment corresponding to the matching degree is used as the masking audio segment, and the second audio feature of the masking audio segment is determined.
[0049] In this embodiment, the original audio segment is each audio segment obtained by dividing the original audio according to a certain rule or time interval. For example, it can be divided according to time length, voice pause or acoustic feature change, and each original audio segment retains the acoustic features and related information of the corresponding part of the original audio. The first audio segment can be divided according to the same division rule or time interval as the original audio segment, and each first audio segment retains the audio feature information of the first audio.
[0050] The matching degree is an indicator used to measure the similarity or closeness of the relationship between each original audio segment and each first audio segment. It can be calculated by comparing the differences between the two in multiple audio feature dimensions, such as the similarity of timbre, the consistency of pitch change patterns, and the fit of rhythm. The matching degree value can be a quantitative value, and the higher the value, the more matched the two are.
[0051] The first matching threshold is a pre-set critical numerical standard for judging the degree of matching. The first matching threshold is dynamically adjusted according to the actual scenario. When the matching degree between a basic audio segment and a certain original audio segment and the corresponding first audio segment is greater than or equal to the first matching threshold, the first audio segment is used as the basic audio segment, indicating that the first audio segment and the corresponding part of the original audio segment are relatively consistent in various aspects and can be used as the basis for subsequent further processing. At this time, the audio features of the basic audio segment, such as timbre features, rhythm features, etc., can be extracted as the first audio features to provide feature basis for subsequent voice editing.
[0052] When the matching degree between a certain original audio segment and the corresponding first audio segment is less than a first matching threshold, the first audio segment is used as a masked audio segment, indicating that the first audio segment has certain differences from the original audio segment in the corresponding part and has its own characteristics. Afterwards, feature extraction is performed on the masked audio segment to obtain a second audio feature.
[0053] The formula for calculating the matching degree is:
[0054]
[0055]
[0056]
[0057] in, Indicates the timbre similarity, represents the pitch similarity, Indicates rhythm similarity; Represents the first Mel-Frequency Cepstral Coefficients (MFCC), Indicates the first audio segment Mel frequency cepstral coefficients (MFCC), which can effectively characterize the timbre, is the number of MFCCs selected, is an adjustment coefficient, so that the larger the difference value, the smaller the contribution of this item. The exponential form can strengthen the positive impact of the similarity coefficient on the similarity; Indicates the original audio segment in The pitch value at a time point, Indicates that the first audio segment is in The pitch value at a time point, is the total number of selected time points; By comparing the ratio of the minimum pitch to the maximum pitch at each time point, the degree of pitch matching is reflected. The closer it is to 1, the more similar the pitches are. Indicates the original audio segment in The beat intensity vector within the beat interval, Indicates the original audio segment in The beat intensity vector within the beat interval, The function calculates the correlation between two vectors. is the total number of beat intervals; The numerator of represents the sum of the correlations between the beat intensity vectors of the two audio segments. The denominator of is the normalization term, The closer the value of is to 1, the more similar the rhythms are. Represents weight, satisfying , the emphasis on timbre, pitch and rhythm can be adjusted according to different application scenarios.
[0058] It can be concluded from the above that this embodiment achieves accurate recognition and classification of audio content by carefully dividing the original audio and the first audio into multiple audio segments and calculating the matching degree between them. When the matching degree reaches or exceeds the preset first matching threshold, the basic audio segment and its features can be accurately extracted, which helps to retain key information. For the masked audio segment with a lower matching degree, by identifying its second audio feature, subsequent replacement, repair or enhancement operations are facilitated. This embodiment not only improves the precision and efficiency of audio editing, but also enhances the flexibility and applicability of audio processing.
[0059] In one embodiment of the present disclosure, determining the edited audio based on the first audio and the text to be synthesized includes: Determine text features of the text to be synthesized; The first audio feature, the second audio feature and the text feature are input into a speech editing model to obtain edited audio; the speech editing model is obtained by training a generative adversarial network based on historical editing data, and the historical editing data includes historical data of the first audio feature, historical data of the second audio feature and historical data of the text feature.
[0060] In one embodiment of the present disclosure, a generative adversarial network includes: introducing a multimodal attention module to a discriminator and a generator.
[0061] In this embodiment, the text to be synthesized contains many aspects of characteristics and attributes, and the quantifiable or describable attributes are summarized as text features. For example, the semantic information, grammatical structure, emotional tendency, part-of-speech distribution, and language style of the text can all be used as text features. The above features help the voice editing model to better understand the text content and generate edited audio. Among them, semantic information is the specific meaning of expression, grammatical structure is the part-of-speech composition and sentence pattern of the sentence, emotional tendency is positive, negative or neutral emotion, part-of-speech distribution is the proportion of various parts of speech such as nouns and verbs, and language style is formal, colloquial or literary.
[0062] The speech editing model is a model that generates edited audio based on input audio features and text features. It can be obtained by training the Generative Adversarial Network (GAN) through historical editing data, that is, learning the association between audio features and text features. The historical editing data is the relevant data accumulated in the past speech editing process, including the first audio features, the second audio features and the corresponding text features that have been processed. The above data is used to train the speech editing model, so that the model can mine the patterns and rules of how to generate high-quality edited audio under different feature combinations.
[0063] GAN is a deep learning model architecture consisting of two parts: the generator and the discriminator. The generator is used to generate data, that is, to generate edited audio, trying to generate audio that is as realistic and meets the requirements as possible; the discriminator is to determine whether the input data is real or generated by the generator. The two compete with each other and through continuous adjustment and optimization, the generator can eventually generate high-quality edited audio that is difficult for the discriminator to distinguish between true and false.
[0064] The multimodal attention module enables the GAN model to allocate "attention" to data of different modes (i.e., audio mode and text mode) when processing data, that is, to focus more on the information parts that are more critical to generating high-quality edited audio. For example, higher attention weights are given to important keywords and semantic core parts in the text, as well as feature parts that play a key role in timbre and rhythm in the audio, so that the audio generated by the generator can better fit the text content, and the discriminator can also more accurately judge the degree of match between audio and text.
[0065] For the generator, in the process of generating edited audio, the attention mechanism is used to more accurately combine text and audio features. The objective function of the generator is:
[0066] in, represents the first audio feature, represents the second audio feature, Represents text features, represents a generator, represents the discriminator, Represents the prior distribution The noise vector sampled from represents the attention function; The objective function of the discriminator is:
[0067] in, represents the first audio feature, represents the second audio feature, Represents text features, represents a generator, represents the discriminator, From real data The distribution obtained in Represents the prior distribution The noise vector sampled from Represents the attention function for audio features, which is used to and Perform weighted processing, Represents the attention function for text features, which is used to Perform weighted processing.
[0068] In this embodiment, through the multimodal attention module, the discriminator can more effectively focus on the key audio and text information of the input, so as to better judge the authenticity of the input data and the quality of the generator output.
[0069] Specifically, the steps of this embodiment are as follows: First, this embodiment analyzes the text to be synthesized, extracts text features, and then uses the previously acquired first audio features, second audio features, and these text features as inputs to the speech editing model. The model has learned the ability to generate appropriate edited audio under different feature combinations by training the generative adversarial network with a large amount of historical editing data, and finally outputs edited audio that meets the requirements. The historical editing data covers the first audio features, second audio features, and corresponding text features in various situations in the past, providing rich learning materials for model training, enabling it to better respond to new input data to generate edited audio.
[0070] Second, when constructing the generative adversarial network, multimodal attention modules are added to both the discriminator and generator parts. The purpose is to enable the discriminator and generator to more effectively focus on key information when processing data of different modalities such as audio features and text features, thereby improving the quality of generated audio and the accuracy of the discriminator's judgment, and further optimizing the effect of the entire speech editing.
[0071] From the above, it can be concluded that this embodiment, by introducing a speech editing model constructed by a generative adversarial network including a multimodal attention module, can comprehensively consider the first audio feature, the second audio feature and the text feature of the text to be synthesized, and realize the accurate generation of the edited audio. The introduction of the multimodal attention module enhances the relevance and accuracy of the model when processing complex audio and text information, making the generated edited audio more consistent with the content and emotion of the text to be synthesized. At the same time, the use of historical editing data for training ensures the stability and reliability of the model.
[0072] In one embodiment of the present disclosure, converting the edited audio to obtain the target audio includes: Determine the tone characteristics and expression phoneme characteristics of the edited audio; The basic features of the original audio, the tone features and the expression phoneme features of the edited audio are input into the sound conversion model to obtain the target audio; the sound conversion model is obtained by training the retrieval-based speech conversion model according to the historical sound data, and the historical sound data includes the historical data of the original audio and the historical data of the edited audio.
[0073] In one embodiment of the present disclosure, the voice editing method further includes: Determine an audio feature vector and a text feature vector based on the edited audio; Determine similarity based on audio feature vector, text feature vector and attention weight; Determine a retrieval-based speech conversion model based on similarity.
[0074] In this embodiment, the tone feature of the edited audio can reflect the characteristics of the character's attitude, emotional intensity, etc. when expressing the content, such as declarative, interrogative, exclamatory tone, or tone expressions of different emotional levels such as gentle, serious, and excited. In the audio, it can be reflected by factors such as changes in the pitch of the audio, changes in the volume, and changes in the speed of the rhythm. For example, an upward tone and a slightly louder volume may indicate a questioning or excited tone, while a steady tone and a moderate volume are mostly declarative.
[0075] Phonemes are the smallest pronunciation units in speech. Expression phoneme features refer to those phoneme-related features that can convey information such as emotions and attitudes similar to human expressions. For example, certain pronunciation methods, the duration of phonemes, and the transition characteristics between phonemes will be different in different emotional states. When you are happy, the pronunciation may be lighter, and when you are sad, the pronunciation may be more dragged and heavy. These special phoneme expression characteristics together constitute the expression phoneme features, which are an important part of reflecting the emotional color of audio from a micro level.
[0076] The sound conversion model is a model that adjusts or converts the features of the input edited audio to generate the target audio. It can be obtained by training the retrieval-based speech conversion model based on a large amount of historical sound data. The purpose is to learn how to reasonably change the audio according to different input features so that it can achieve the desired target audio effect.
[0077] Historical sound data is a collection of sound-related data accumulated in the past, including historical data of original audio and historical data of edited audio. It covers audio samples of various scenarios and features and corresponding related information, providing rich materials for the training of sound conversion models and helping the models grasp the conversion rules and relationships between different audio features.
[0078] The audio feature vector is a vector representation that is extracted by analyzing the edited audio and can characterize various aspects of the audio. Audio processing algorithms can be used to quantify the audio's timbre, pitch, rhythm and other features, and integrate them into a vector space to facilitate subsequent calculations, comparisons and use as model input. For example, different dimensions in the vector can correspond to the energy of different frequency bands, the average value of the pitch, specific parameters of the rhythm, etc.
[0079] The text feature vector is a vector representation extracted from the text content associated with the edited audio. The semantics, grammar, sentiment, part-of-speech distribution and other features of the text are quantified through a specific natural language processing method and converted into a vector form so that it can be used together with the audio feature vector in subsequent steps to measure the correlation and similarity between them.
[0080] Attention weight is a weight coefficient set in order to highlight the importance of different feature dimensions or different elements in the process of calculating similarity.
[0081] Similarity is an indicator used to measure the similarity between audio feature vectors and text feature vectors. It can be used to determine the degree of match between audio and text, and further provide a basis for related operations of the retrieval-based speech conversion model, such as determining the scope of retrieval and selecting appropriate reference audio.
[0082] The similarity calculation formula is:
[0083] in, Indicates text feature vectors, Indicates audio feature vectors, Indicates Attention weights. Attention weights can be learned from prior knowledge of audio and text. For example, higher weights are given to keywords in text or key rhythmic parts in audio, so that the model pays more attention to these important information during retrieval, improving the accuracy and relevance of retrieval. Attention weights can also be adjusted dynamically based on actual conditions.
[0084] First, this embodiment analyzes the edited audio and extracts the tone features and expression phoneme features from it. The above features reflect the characteristics of the edited audio in terms of emotional expression and attitude communication. Then, the basic features of the original audio itself, together with the above two features of the edited audio, are input into the sound conversion model. The sound conversion model is a retrieval-based speech conversion model that has been trained with a large amount of historical sound data. The historical sound data covers relevant information of various original audio and edited audio in the past. By learning these data, the model has mastered how to perform reasonable audio conversion based on the input features and finally outputs the target audio that meets the requirements.
[0085] First, for the edited audio, the corresponding audio processing technology and natural language processing technology are used to extract the vector that can represent its audio features and the feature vector of the related text. Then, combined with the dynamically changing attention weight, the similarity between the two vectors is determined. This similarity reflects the degree of match between the audio and the text. Finally, based on the value of this similarity and other conditions, the retrieval-based speech conversion model to be used is further determined, so as to better perform accurate sound conversion based on the actual relationship between the audio and the text.
[0086] It can be concluded from the above that this embodiment can generate target audio that is highly consistent with the editing intention by accurately capturing the tone characteristics and expression phoneme characteristics of the edited audio, combining the basic characteristics of the original audio, and using the sound conversion model trained with historical sound data, which significantly improves the naturalness and realism of the audio conversion. At the same time, the introduction of similarity based on audio feature vectors, text feature vectors and attention weights further optimizes the conversion effect and enhances the flexibility and accuracy of audio editing.
[0087] Corresponding to the voice editing method of the above embodiment, Figure 2 This is a structural block diagram of a voice editing device provided by an embodiment of the present disclosure. For ease of explanation, only the parts related to the embodiment of the present disclosure are shown. Figure 2 The voice editing device 20 includes: an audio editing module 21, an audio conversion module 22 and an audio output module 23.
[0088] The audio editing module 21 is used to determine the edited audio based on the first audio and the text to be synthesized, where the first audio is the audio after masking the original audio; The audio conversion module 22 is used to convert the edited audio to obtain the target audio; The audio output module 23 is used to output target audio.
[0089] In one embodiment of the present disclosure, the speech editing device 20 further includes: a feature extraction module; A feature extraction module, used for performing masking processing on the original audio to determine masked audio, and using the masked audio as the first audio; Determine the basic characteristics of the original audio; Audio characteristics of the first audio are determined.
[0090] In one embodiment of the present disclosure, the speech editing device 20 further includes: a feature matching module; A feature matching module, for determining a plurality of original audio segments based on basic features of the original audio; determining a plurality of first audio segments based on the audio feature of the first audio; determining a degree of matching based on each original audio segment and each first audio segment; If the matching degree is greater than or equal to the first matching threshold, taking the first audio segment corresponding to the matching degree as the basic audio segment, and determining the first audio feature of the basic audio segment; If the matching degree is less than the first matching threshold, the first audio segment corresponding to the matching degree is used as the masking audio segment, and the second audio feature of the masking audio segment is determined.
[0091] In one embodiment of the present disclosure, the audio editing module 21 is specifically used to: Determine text features of the text to be synthesized; The first audio feature, the second audio feature and the text feature are input into a speech editing model to obtain edited audio; the speech editing model is obtained by training a generative adversarial network based on historical editing data, and the historical editing data includes historical data of the first audio feature, historical data of the second audio feature and historical data of the text feature.
[0092] In one embodiment of the present disclosure, a generative adversarial network includes: introducing a multimodal attention module to a discriminator and a generator.
[0093] In one embodiment of the present disclosure, the audio conversion module 22 is specifically configured to: Determine the tone characteristics and expression phoneme characteristics of the edited audio; The basic features of the original audio, the tone features and the expression phoneme features of the edited audio are input into the sound conversion model to obtain the target audio; the sound conversion model is obtained by training the retrieval-based speech conversion model according to the historical sound data, and the historical sound data includes the historical data of the original audio and the historical data of the edited audio.
[0094] In one embodiment of the present disclosure, the speech editing device 20 further includes: a speech conversion model building module; A speech conversion model building module for determining an audio feature vector and a text feature vector based on the edited audio; Determine similarity based on audio feature vector, text feature vector and attention weight; Determine a retrieval-based speech conversion model based on similarity.
[0095] See also Figure 3 , Figure 3 A schematic block diagram of an electronic device provided by an embodiment of the present disclosure. Figure 3 The electronic device 300 in the embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303 and one or more memories 304. The processors 301, input devices 302, output devices 303 and memories 304 communicate with each other via a communication bus 305. The memory 304 is used to store computer programs, which include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. The processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above-mentioned device embodiments, such as Figure 2 The functions of modules 21 to 23 are shown.
[0096] It should be understood that in the embodiment of the present disclosure, the processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0097] The input device 302 may include a touch panel, a fingerprint collection sensor (for collecting the user's fingerprint information and fingerprint direction information), a microphone, etc., and the output device 303 may include a display (LCD, etc.), a speaker, etc.
[0098] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A portion of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.
[0099] In a specific implementation, the processor 301, input device 302, and output device 303 described in the embodiments of the present disclosure can execute the implementation methods described in the first and second embodiments of the voice editing method provided in the embodiments of the present disclosure, and can also execute the implementation methods of the electronic device described in the embodiments of the present disclosure, which will not be repeated here.
[0100] In another embodiment of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by the processor, all or part of the processes in the above-mentioned embodiment method are implemented, and the computer program can also be completed by instructing the relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, the steps of each of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc.
[0101] The computer-readable storage medium may be an internal storage unit of the electronic device of any of the aforementioned embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of the electronic device and an external storage device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0102] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this disclosure.
[0103] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the electronic devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0104] In the several embodiments provided in the present application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces or units, or it can be an electrical, mechanical or other form of connection.
[0105] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present disclosure.
[0106] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0107] The above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present disclosure, and these modifications or replacements should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.
Claims
1. A voice editing method, characterized in that: include: Determine the edited audio based on the first audio and the text to be synthesized, where the first audio is the audio after masking the original audio; Converting the edited audio to obtain target audio; The target audio is output.
2. The voice editing method according to claim 1, characterized in that: Also includes: Perform masking processing on the original audio to determine masked audio, and use the masked audio as the first audio; Determining basic characteristics of the original audio; Audio characteristics of the first audio are determined.
3. The voice editing method according to claim 2, characterized in that: Also includes: determining a plurality of original audio segments based on basic features of the original audio; determining a plurality of first audio segments based on the audio features of the first audio; determining a degree of matching based on each original audio segment and each first audio segment; If the matching degree is greater than or equal to a first matching threshold, taking a first audio segment corresponding to the matching degree as a basic audio segment, and determining a first audio feature of the basic audio segment; If the matching degree is less than a first matching threshold, the first audio segment corresponding to the matching degree is used as a masking audio segment, and a second audio feature of the masking audio segment is determined.
4. The voice editing method according to claim 3, characterized in that: The step of determining the edited audio based on the first audio and the text to be synthesized includes: Determining text features of the text to be synthesized; The first audio feature, the second audio feature and the text feature are input into a speech editing model to obtain the edited audio; the speech editing model is obtained by training a generative adversarial network based on historical editing data, and the historical editing data includes historical data of the first audio feature, historical data of the second audio feature and historical data of the text feature.
5. The voice editing method according to claim 4, characterized in that: The generative adversarial network includes: introducing a multimodal attention module to the discriminator and the generator.
6. The voice editing method according to claim 2, characterized in that: The step of converting the edited audio to obtain the target audio comprises: Determining the tone characteristics and expression phoneme characteristics of the edited audio; The basic features of the original audio, the tone features and the expression phoneme features of the edited audio are input into a sound conversion model to obtain the target audio; the sound conversion model is obtained by training a retrieval-based speech conversion model according to historical sound data, and the historical sound data includes the historical data of the original audio and the historical data of the edited audio.
7. The voice editing method according to claim 6, characterized in that: Also includes: Determine an audio feature vector and a text feature vector based on the edited audio; Determining similarity based on the audio feature vector, the text feature vector, and the attention weight; The retrieval-based speech conversion model is determined based on the similarity.
8. A voice editing device, characterized in that: include: An audio editing module, used for determining an edited audio based on a first audio and a text to be synthesized, wherein the first audio is an audio after masking processing is performed on the original audio; An audio conversion module, used for converting the edited audio to obtain a target audio; An audio output module is used to output the target audio.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.