Audio editing method and device, computing equipment and medium

By obtaining the voice signal and text from the audio to be edited, adjusting the speech recognition model, and generating the edited audio, the problems of user missed tongue and forgetting words during recording are solved, and fast editing and high-quality output are achieved.

CN120164482APending Publication Date: 2025-06-17FACE CUTE CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311734457.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

During the UGC/PUGC creation process, users often experience problems such as erroneous tongue, forgetting words, and unstandard pronunciation when recording audio. Re-recording is expensive and will lead to a lack of environmental sounds, affecting the quality of the content.

Method used

By obtaining the voice signal and corresponding text from the audio to be edited, users can edit the text in real time, adjust the pre-trained voice recognition model based on the voice signal, and generate the edited audio to keep it consistent with the original audio tone, style and background sound.

Benefits of technology

It realizes users to edit audio content quickly and in real time, keeps the other characteristics of the audio consistent, improves user creation efficiency and content quality, and avoids the cost of re-recording and the problem of ambient sound loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164482A_ABST
    Figure CN120164482A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio editing method and device, computing equipment and a medium. The method comprises the following steps: acquiring a voice signal and a text corresponding to the voice signal from a to-be-edited audio; modifying the text in response to a user editing operation; adjusting a pre-trained speech recognition model based on the speech signal; and generating edited audio based on the modified text using the adjusted speech recognition model. In this way, the user can quickly edit the content of the audio in real time, and other features of the modified audio are kept consistent with the original features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, an apparatus, a computing device, a computer-readable storage medium, and a computer program product for editing audio. Background Art

[0002] With the development of Internet technologies, more and more users participate in the creation of user-generated content (UGC) and professional user-generated content (PUGC). Users can display their original content through Internet platforms or provide it to other users. During the UGC / PUGC creation process, users inevitably make mistakes, forget words, or have inaccurate pronunciations when recording audio or video. Although re-recording can solve these problems, it is costly and there will be a problem of missing ambient sound. This also brings great troubles to users. Summary of the Invention

[0003] In view of this, the present disclosure provides a technical solution for editing audio, enabling users to quickly and real-time edit the content of audio, and other features (such as timbre, ambient background sound, etc.) of the modified audio to remain the same as the original.

[0004] According to a first aspect of the present disclosure, a method for editing audio is provided. The method includes: obtaining a voice signal and text corresponding to the voice signal from the audio to be edited; modifying the text in response to a user editing operation; adjusting a pre-trained speech recognition model based on the voice signal; and using the adjusted speech recognition model to generate an edited audio based on the modified text.

[0005] In a second aspect of the present disclosure, an apparatus for editing audio is provided. The apparatus includes: a content recognition unit configured to obtain a voice signal and text corresponding to the voice signal from the audio to be edited; a content modification unit configured to modify the text in response to a user editing operation; a model adjustment unit configured to adjust a pre-trained speech recognition model based on the voice signal; and an audio generation unit configured to use the adjusted speech recognition model to generate an edited audio based on the modified text.

[0006] In a third aspect of the present disclosure, a computing device is provided. The computing device includes: at least one processing unit and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit causing the computing device to perform the method according to the first aspect of the present disclosure.

[0007] In a fourth aspect of the present disclosure, there is provided a non-transitory computer storage medium including machine-executable instructions that, when executed by a device, cause the device to perform the method according to the first aspect of the present disclosure.

[0008] In a fifth aspect of the present disclosure, there is provided a computer program product including machine-executable instructions that, when executed by a device, cause the device to perform the method according to the first aspect of the present disclosure.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0011] Figure 1 A block diagram showing an environment capable of implementing multiple embodiments of the present disclosure;

[0012] Figure 2 A block diagram showing an example structure of an audio editing system according to an embodiment of the present disclosure;

[0013] Figures 3A to 3C An example of a user interface for editing audio according to an embodiment of the present disclosure;

[0014] Figure 4 A schematic diagram showing a process for editing audio according to an embodiment of the present disclosure;

[0015] Figure 5 A schematic diagram showing some components of a speech recognition model according to an embodiment of the present disclosure;

[0016] Figure 6 A schematic flowchart showing a method for editing audio according to an embodiment of the present disclosure;

[0017] Figure 7 A schematic block diagram showing a device for editing audio according to an embodiment of the present disclosure; and

[0018] Figure 8 A block diagram showing a computing device capable of implementing some implementations of the present disclosure.

[0019] In all the drawings, the same or similar reference numerals represent the same or similar elements. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The exemplary embodiments of the present disclosure will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0021] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions below.

[0022] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0023] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.

[0024] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0025] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0026] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0027] A "neural network" is a machine learning network based on deep learning. A neural network can process inputs and provide corresponding outputs, and it usually includes an input layer and an output layer as well as one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications usually include many hidden layers, thus increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also called processing nodes or neurons), and each node processes the input from the previous layer.

[0028] Generally, machine learning can roughly include three stages, namely the training stage, the testing stage, and the usage stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously iteratively updated until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also called the input-to-output mapping) from the training data. The parameter values of the trained model are determined. In the testing stage, the test inputs are applied to the trained model to test whether the model can provide the correct output, so as to determine the performance of the model. In some implementations, the testing stage can be omitted. In the usage stage, the model can be used to process actual inputs based on the trained parameter values to determine the corresponding outputs.

[0029] As mentioned above, during the UGC and PUGC creation processes, when users record videos, there will inevitably be situations such as slips of the tongue, forgetting lines, or incorrect pronunciations, resulting in the need to re-record. However, re-recording has the problems of high cost and missing ambient sounds. Some traditional solutions can intercept the content that needs to be modified and let the user re-record a small section of content to cover it. This method may bring obvious modification traces, resulting in a decline in the quality of the created content.

[0030] In view of this, embodiments of the present disclosure provide a solution for quickly editing audio. In this solution, speech signals and corresponding text content are identified from the audio to be edited (e.g., a recording that needs to be modified). The user can perform editing operations on the text content to be modified (e.g., deleting words or replacing them with new words). The speech signals can be provided to a speech recognition model so that the speech recognition model is fine-tuned based on the received speech signals. The fine-tuned speech recognition model stores the feature information of the speech signals. Then, using the adjusted speech recognition model, the edited audio is generated based on the modified text. According to the implementation of the present disclosure, the user can modify the audio as simply as editing text content, and at the same time, the edited audio still maintains the same features as the original audio, such as the speaking tone color, style, background sound, etc., thereby improving the user's creation efficiency and content quality.

[0031] Figure 1 A block diagram of an example environment 100 capable of implementing multiple embodiments of the present disclosure is shown. In environment 100, an audio editing system 110 is provided for performing audio editing operations of user 101 using terminal device 120. In Figure 1 this case, the audio editing system 110 can be any system with computing capabilities, and the terminal device 120 can be any device with user interaction capabilities. It should be understood that Figure 1 the components and arrangements in the illustrated environment are only examples, and the computing systems suitable for implementing the example implementations described in the present disclosure may include one or more different components, other components, and / or different arrangements.

[0032] In operation, the audio editing system 110 obtains the audio to be edited (sometimes also referred to as "original audio" herein, and the two can be used interchangeably) 102 from the terminal device 120, which represents the audio content that user 101 wants to modify. For example, it is a user recording with a slip of the tongue, inappropriate words, etc. In some implementations, the audio to be edited 102 can be separate audio content. Alternatively or additionally, the audio to be edited 102 can be audio embedded in a video. In some implementations, the audio to be edited 102 can also be obtained by the audio editing system 110 from other sources, such as local or other servers.

[0033] User 101 can also operate the terminal device 120 such that the terminal device 120 provides editing data 104 to the audio editing system 110. The editing data 104 indicates the modified text, that is, the content modification that the user expects to make to the audio 102 to be edited. In some examples, the editing data 104 may include operations such as addition, deletion, modification, or replacement of the speech content in the audio 102 to be edited. For example, the terminal device 120 can play the audio 102 and present the text content in the audio 102. The user 101 can select some words of the text content to perform deletion or replacement operations, or the user 101 can also input more text content, thereby generating new text content. Then, the terminal device 102 provides the new content as the editing data 104 to the audio editing system 110.

[0034] The audio editing system 110 can generate a new audio based on the received audio 102 to be edited and the editing data 103, and send the generated new audio as the edited audio 115 to the terminal device 120. The user 101 can continue to browse the edited audio 115 and repeat the above process if further modification is needed until the editing is completed. In some implementations, the audio editing system 110 can generate the edited audio 115 based on a neural network model. By using the neural network model, the timbre information of the user in the original audio 102 can be retained so that the generated audio has the same timbre as the original audio. In some implementations, the audio editing system 110 can also retain the speech style of the user and the ambient background sound in the original audio 102. Therefore, except for the modified text content, the edited audio 115 is basically the same as the original audio 102. It should be noted that the audio editing system 110 according to the embodiments of the present disclosure can quickly and real-time provide the edited audio to the user 101, thereby significantly improving the efficiency and experience of the user creating content.

[0035] To facilitate the understanding of the embodiments of the present disclosure, the following gives exemplary application scenarios. In some scenarios, the embodiments of the present disclosure can be used to correct slips of the tongue in a dubbing or voice-over program. Whether it is dubbing or voice-over, according to the user's needs, it is as simple as modifying text on the text. In some scenarios, to make the user's speech fluent, the entire unsmooth part of the speech can be regenerated to make the speech fluent. In some scenarios, it is also possible to automatically detect grammar errors in the speech script and achieve quick correction to make the speech content more accurate and professional. In some scenarios, it is also possible to implement swear word filtering. Currently, beep sounds are used to cover swear words, and it is possible to detect and replace swear words. In some scenarios, it is also possible to implement cross-language editing, which can change a pure Chinese sentence into a Chinese-English mixture, or change a Chinese-English mixture into pure Chinese or English, and even directly implement Chinese-English translation. In some scenarios, it is also possible to implement secondary creation of film and television plots to increase the fun of interaction among users. It can be understood that the above scenarios are only exemplary, and the embodiments of the present disclosure are also applicable to any other application scenarios. Some specific implementations of the audio editing system in the present disclosure will be discussed in more detail below.

[0036] Figure 2 FIG. shows a block diagram of an exemplary structure of an audio editing system according to an embodiment of the present disclosure. As Figure 2 shown, the audio editing system 110 may include a voice separation module 210, an automatic speech recognition module 220, a speech recognition model 230, and an audio synthesis module 240. It can be understood that Figure 2 the shown audio editing system is only exemplary, and some of the modules may be omitted or modified, and other modules may also be included.

[0037] The voice separation module 210 is used to extract the user's voice signal and background sound signal from the received audio 102 to be edited. In some implementations, based on spectral analysis (such as independent component analysis method), filters can be used to generate separated voice signals and background sound signals from the audio 102 to be edited. The spatial filtering method can also be used. The sound source signal is collected through a microphone array, and then the beamforming and filtering algorithms are used to process the mixed signal to achieve the separation of the voice signal and the background sound. Other methods can also be used to achieve the separation of the voice signal and the background sound, and the present disclosure does not limit this aspect. If the audio to be edited has no background sound or the background sound is small, the voice separation module 210 can be omitted or not used.

[0038] The automatic speech recognition module 220 is used to recognize the content of the audio to be edited. The automatic speech recognition module 220 obtains the audio to be edited 102 or the speech signal output by the speech separation module 210, and generates the text corresponding to the speech signal. In some implementations, the automatic speech recognition module may include a feature extraction module, a statistical acoustic model, and a language model. The feature extraction module extracts features from the input speech signal for acoustic model modeling and the decoding process. The acoustic model is used to model basic acoustic units such as words, syllables, and phonemes to generate an acoustic model. The language model is used to model the language to be recognized by the system at the word level. The text content generated by the automatic speech recognition module 220 may include a set of words, and each word has corresponding timestamp information. The audio editing system 110 can provide the text content 103 to the terminal device 120 for the user to view. Additionally, the automatic speech recognition module 220 may have the ability to recognize multi-lingual texts, including Chinese, English, Japanese, etc., and can output mixed-language texts.

[0039] The user 101 views the text content 103 on the terminal device 120. If the user 101 notices that some text content in the audio needs to be modified, the user 101 can operate on the terminal device 120 so that the terminal device 120 provides the editing data 104 to the audio editing system 110. In some implementations, the editing data 104 may include the words in the text content 103 that need to be modified, replaced, or deleted, and may also include newly added words. The words in the modified text may have corresponding timestamp information.

[0040] In response to receiving the editing data 104, the audio editing system 110 may determine that the audio 101 needs to be edited. As shown in the figure, the audio editing system 110 uses the speech recognition model 230 to perform the editing operation. The speech recognition model 230 obtains the audio 102 to be edited or the speech signal in the audio 102 to be edited and the editing data 104, and generates a target speech signal. The target speech signal can be synthesized with the background sound signal to form the edited audio 115 and provided to the terminal device 115. In some implementations, the speech recognition model 230 is used to decouple the content, style, and background of the audio. For example, the speech recognition model 230 may include a timbre encoder, a style encoder, and a background sound encoder. These three encoders are respectively used to learn the timbre information, style information, and background sound information in the input audio, so that the generated audio can maintain the same timbre, style, and background sound as the original audio. The background sound may refer to other sounds other than the speech in the audio 102. For example, accompaniment instruments, environmental noise, etc. It can be understood that the speech recognition model 230 may include more or fewer encoders, and is not limited to the above encoders. It should be noted that the speech recognition model 230 may be pre-trained, that is, it already has knowledge or information about timbre, style, or background sound. Once it receives the audio 102 or speech signal to be edited, it can quickly adjust its parameters (also known as "fine-tuning") and learn the timbre, style, or background sound information of the input audio. For example, the fine-tuning operation can be completed in just a few seconds.

[0041] In response to the completion of the fine-tuning of the speech recognition model 230, the speech recognition model 230 can switch to the inference stage. At this time, the speech recognition model 230 acts as a decoder and generates a new speech signal according to the received editing data 104, which includes the modified text content indicated by the editing data 104. In some implementations, the speech recognition model 230 can also generate a new background sound based on the fine-tuned background sound encoder.

[0042] The audio synthesis module 240 is used to generate the edited audio from the new voice signal. It can be a separate module or embedded in the speech recognition model 230. In some implementations, the audio synthesis module 240 can merge the generated new voice signal with the original background sound signal from the voice separation module 210 to form the edited audio 115. Alternatively, the audio synthesis module can also merge the new voice signal with the new background sound from the speech recognition model 230 to form the edited audio 115. Then, the audio editing system 110 sends the edited audio 115 to the terminal device 120. Since the speech recognition model 230 has learned information such as the timbre, style, and background sound of the original audio during the fine-tuning stage, the edited audio 115 can retain other features of the original audio 102 while modifying the content of the original audio, just like a real recording.

[0043] Figures 3A to 3C Shows an example of a user interface for editing audio according to an embodiment of the present disclosure. In Figure 3A In the user interface 300A, the user can watch the video they created and check whether there are any problems with the audio in the video and whether editing is required. In the user interface 300A, there are a video playback window 305, a text content display window 310 for the audio, editing icons 315, a video playback progress bar 316, and other controls 318. The text content window 310 can display the text content 103 recognized by the automatic speech recognition module 220 in a scrolling manner. In some embodiments, the words in the text content 103 can have associated timestamp information, so that the terminal device 120 can display the currently playing text content 311 (I wanna sing) and adjacent text before or after according to the playback progress in the window 310.

[0044] An exemplary text content 103 is as follows, where the words in the text content 103 have information such as their start time and end time in the video or audio.

[0045]

[0046]

[0047] Suppose the user wants to modify the text content currently displayed in the window 310 and hopes to change "sing" to "dance". The user 101 can click the editing control 315. Accordingly, the terminal device 120 can enter Figure 3B The audio editing interface 300B shown.

[0048] In interface 300B, in addition to the video playback window 305 and the text content display window 310, interface elements 322 to 325 related to the editing operation are also shown. As shown in the figure, when the user clicks on the word "sing" in window 310, it means that user 101 selects "sing" as the object to be edited. Accordingly, the word "sing" can be automatically input into interface element 322. Then, the user can input the word "dance" in interface element 323 to replace "sing". If the user also wishes to edit other words, the user can continue to click on control 324. At this time, interface 300B can add new lines similar to interface elements 322 and 323 for the user to input the word to be replaced and the desired word. The user's multiple editing operations can be temporarily saved for combined submission. After completing the editing operation, the user can click on the preview 325 control to generate the editing data 104. Thus, the terminal device 120 can provide the editing data 104 to the audio editing system 110, which generates new audio therefrom. The following shows an exemplary editing data 104, which indicates the desire to replace the word numbered 7 (it can be understood that the audio editing system has the numbered information for each word), that is, "sing", with the word "dance".

[0049]

[0050] Next, based on the received editing data 104 and the original text content, the audio recognition system 110 generates the modified text. In the modified text, "dance" can have the same timestamp as "sing". However, the audio recognition system 110 can generate the edited audio based on the modified text and send the audio to the terminal device 120. Accordingly, the terminal device 120 presents the modified audio and its text content in the Figure 3C shown user interface 300C. As Figure 3C shown, the text content 331 has been modified from "sing" to "dance". At the same time, the terminal device 120 can play the modified audio, and the time when the word "dance" appears in the modified audio is the same as that of "sing" in the original audio.

[0051] Figure 4 FIG. shows a schematic diagram of a process 400 for editing audio according to an embodiment of the present disclosure. Process 400 can be implemented at the Figure 1 and Figure 2 shown audio editing system 110. For ease of understanding, process 400 will be described with reference to Figure 2 .

[0052] The audio to be edited 102 can be provided to the speech recognition model 230 of the audio editing system 110. The audio to be edited 102 can include a speech signal and background noise. In some implementations, if there is no separate speech separation module, the audio to be edited can be directly provided to the speech recognition model 230; if there is a speech separation module, the separated speech signal can be provided to the speech recognition module 230.

[0053] As shown in the figure, the speech recognition model 230 can include a pre-trained timbre encoder 432, which has been trained by training data to be able to recognize the timbre of the speech in the audio. The timbre encoder 432 can be set in a fine-tuning mode and adjusted by changing a part of its parameters in response to receiving the speech signal of the audio to be edited, so that the adjusted timbre encoder 432 can generate an encoded representation of the timbre of the speech of the audio to be edited, that is, can recognize the timbre of the speech in the current audio.

[0054] Additionally or optionally, the speech recognition model 230 can also include a pre-trained style encoder 434, which has been trained by training data to be able to recognize the style of the speech in the audio. Similarly, the style encoder 434 can be set in a fine-tuning mode and adjusted by changing a part of its parameters in response to receiving the speech signal of the audio to be edited, so that the adjusted style encoder 434 can generate an encoded representation of the style of the speech of the audio to be edited, that is, can recognize the style of the speech in the current audio.

[0055] Additionally or optionally, the speech recognition model 230 can also include a pre-trained background encoder 436, which has been trained by training data to be able to recognize the characteristics of the background noise in the audio. Similarly, the background encoder 436 can be set in a fine-tuning mode and adjusted by changing a part of its parameters in response to receiving the background noise signal of the audio to be edited, so that the adjusted background encoder 436 can generate an encoded representation of the background noise of the audio to be edited, that is, can recognize the background noise in the current audio.

[0056] In process 400, the audio editing system 110 can also generate text content 103 based on the audio to be edited 102, which can be provided to the terminal device 120 to help the user determine whether the audio needs to be edited. For example, the audio editing system 110 can use the automatic speech recognition module 220 to generate the text content 103. The audio editing system 110 can receive the editing data 104 generated in response to the user's editing operation from the terminal device. Then, the audio editing system 110 can modify the original text content 103 based on the editing data 104 to obtain the modified text. The process of generating the modified text can be completed instantaneously, while the process of adjusting the encoder of the speech recognition model may take a slightly longer time, such as a few seconds, but it has been greatly reduced compared to the existing model adjustment time because the speech recognition model 430 decouples the speech content and requires fewer parameters to be adjusted.

[0057] Next, using the adjusted speech recognition model 235, a target speech signal 440 is generated based on the modified text. Specifically, the adjusted speech recognition model 235 can be set to the inference mode. The speech recognition model 235 in the inference mode includes a decoder 438, which can generate the target speech signal 440 based on the encoded representations of the text and the timbre. As mentioned above, the parameters of the fine-tuned timbre encoder have been changed to generate the encoded representation of the timbre of the original audio. Thus, the timbre of the target speech signal 440 generated by the speech recognition model 235 is the same as that of the original audio 102, but the content has changed.

[0058] In some embodiments, the speech signal of the audio to be edited 102 can also be used to adjust the style encoder 434 to generate the encoded representation of the corresponding style. Correspondingly, the adjusted speech recognition model 235 can also generate the target speech signal 440 based on the encoded representation of the timbre, the encoded representation of the style, and the modified text. Thus, both the timbre and the style of the target speech signal 440 generated by the speech recognition model 235 are the same as those of the original audio 102, but the content has changed.

[0059] Then, the audio editing system 110 can merge the target speech signal 440 and the background sound of the original audio 102 to generate the edited audio 115. In some embodiments, if there is no separated background sound, the audio editing system 110 can regenerate the background sound. For example, a new background sound is generated based on the encoded representation of the background sound generated by the background encoder, and the new background sound is merged with the target speech signal 440 to generate the edited audio 115.

[0060] It can be understood that Figure 4 only the example process is given, and there can be other ways to detect errors and perform pattern matching from the signal sequence.

[0061] Figure 5 A schematic diagram showing some components of a speech recognition model according to an embodiment of the present disclosure. Any one of the encoders 432, 434, and 436 of the speech recognition model 230 may have Figure 5 the exemplary structure shown.

[0062] As shown in the figure, the exemplary structure includes a first feed-forward module 502, a multi-head self-attention module 504, a convolutional module 506, and a second feed-forward module 508 connected in sequence, where the input and output of each module can be combined together through a combiner 510 as the input of the next module. The combiner 510 can be a vector addition or a concatenation operation. In some implementations, the multi-head self-attention module 504 may include a transformer module for extracting long sequence dependencies. The convolutional module 506 is used to extract local features. Through the combination of the multi-head self-attention module 504 and the convolutional module 506, the performance of the speech recognition model in both long-term sequences and local features is improved simultaneously.

[0063] Figure 6 A flowchart showing an example method 600 according to some implementations of the present disclosure. The method 600 may be implemented at the audio editing system 110.

[0064] At block 610, the audio editing system 110 obtains a speech signal and text corresponding to the speech signal from the audio to be edited. At block 620, the audio editing system 110 modifies the text in response to a user editing operation. At block 630, the audio editing system 110 adjusts a pre-trained speech recognition model based on the speech signal. At block 640, the audio editing system 110 uses the adjusted speech recognition model to generate an edited audio based on the modified text.

[0065] In some embodiments, the pre-trained speech recognition model may include a pre-trained timbre encoder, where adjusting the pre-trained speech recognition model may include: adjusting the pre-trained timbre encoder using the speech signal such that the adjusted timbre encoder generates a coded representation of the timbre in the speech signal.

[0066] In some embodiments, generating the edited audio may include: generating a target speech signal based on the coded representation of the timbre and the modified text.

[0067] In some embodiments, the pre-trained speech recognition model may include a style encoder, where adjusting the pre-trained speech recognition model may further include: adjusting the pre-trained style encoder using the speech signal such that the adjusted style encoder generates a coded representation of the style in the speech signal.

[0068] In some embodiments, generating the target voice signal may further include: generating the target voice signal based on the encoded representation of the timbre, the encoded representation of the style, and the modified text.

[0069] In some embodiments, method 600 may further include: obtaining a background sound signal from the audio to be edited; and synthesizing the edited audio based on the background sound signal and the target voice signal.

[0070] In some embodiments, the pre-trained speech recognition model may further include a pre-trained background encoder, and adjusting the pre-trained speech recognition model may further include: using the background sound signal to adjust the pre-trained background encoder such that the adjusted background encoder generates an encoded representation of the background sound of the audio to be edited.

[0071] In some embodiments, the obtained text may include a set of words, each word having an associated timestamp, and modifying the text in response to a user editing operation may include: receiving a user editing operation that specifies a first word in the set of words and a second word for replacing the first word; and obtaining the modified text by replacing the first word with the second word, where in the modified text, the second word has the timestamp of the first word.

[0072] In some embodiments, generating the edited audio based on the modified text may include: generating the edited audio based on the modified text with the timestamps of the words, where the appearance time of the words in the edited audio is the same as that in the audio to be edited.

[0073] In some embodiments, the speech recognition model includes at least one encoder, and the at least one encoder includes a first feed-forward module, a multi-head self-attention module, a convolutional module, and a second feed-forward module connected in sequence.

[0074] Figure 7 FIG. shows a schematic block diagram of a device 700 for editing audio according to an embodiment of the present disclosure. As Figure 7 shown, device 700 includes: a content recognition unit 702, a content modification unit 704, a model adjustment unit 706, and an audio generation unit 708.

[0075] The content recognition unit 702 is configured to obtain a voice signal and the text corresponding to the voice signal from the audio to be edited. The content modification unit 704 is configured to modify the text in response to a user editing operation. The model adjustment unit 706 is configured to adjust the pre-trained speech recognition model based on the voice signal. The audio generation unit 708 is configured to use the adjusted speech recognition model to generate the edited audio based on the modified text.

[0076] It should be noted that with reference to Figures 1 to 6The additional actions or steps described can be implemented by Figure 7 the apparatus 700 shown in. For example, the apparatus 700 may include additional modules or units to implement the actions or steps described above, or Figure 7 some of the units or modules shown can be further configured to implement the actions or steps described above. This will not be elaborated here again.

[0077] The foregoing has referred to Figures 2 to 7 exemplary embodiments of the present disclosure. Compared with existing solutions, in the technical solution provided by the present disclosure, a user can quickly and in real time edit the content of an audio, and other features (such as timbre, ambient background sound, etc.) of the edited audio remain the same as the original. This ensures the consistency of the user's timbre and ambient background sound, and there will be no editing traces or "abruptness".

[0078] Figure 8 FIG. shows a block diagram of a computing device 800 capable of implementing multiple implementations of the present disclosure. It should be understood that Figure 8 the computing device 800 shown is merely exemplary and should not constitute any limitation to the functions and scope of the implementations described in the present disclosure. The computing device 800 can be used to implement the audio editing system 110.

[0079] As Figure 8 shown, the computing device 800 includes a computing device 800 in the form of a general-purpose computing device. The components of the computing device 800 may include, but are not limited to, one or more processors or processing units 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860.

[0080] In some implementations, the computing device 800 can be implemented as various user terminals or service terminals with computing capabilities. The service terminal can be a server, a large computing device, etc. provided by various service providers. The user terminal is, for example, any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablets, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is also foreseeable that the computing device 800 can support any type of user interface (such as a "wearable" circuit, etc.).

[0081] The processing unit 810 can be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing ability of the computing device 800. The processing unit 810 can also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0082] The computing device 800 generally includes multiple computer storage media. Such media can be any available media accessible to the computing device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 can be a volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The memory 820 can include an audio editing module 822, and these program modules are configured to perform the functions of various implementations described herein. The audio editing module 822 can be accessed and run by the processing unit 810 to implement the corresponding functions.

[0083] The storage device 830 can be a removable or non-removable medium and can include a machine-readable medium that can be used to store information and / or data and can be accessed within the computing device 800. The computing device 800 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 8 a disk drive for reading from or writing to a removable, non-volatile disk and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces.

[0084] The communication unit 840 enables communication with other computing devices through a communication medium. Additionally, the functions of the components of the computing device 800 can be implemented in a single computing cluster or multiple computer machines that can communicate through a communication connection. Thus, the computing device 800 can operate in a networked environment using a logical connection with one or more other servers, personal computers (PCs), or another general network node.

[0085] The input device 850 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. The output device 860 can be one or more output devices, such as a display, speaker, printer, etc. The computing device 800 can also communicate with one or more external devices (not shown) as needed via the communication unit 840. The external devices such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the computing device 800, or communicate with any device that enables the computing device 800 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0086] In some implementations, in addition to being integrated on a single device, some or all of the various components of the computing device 1200 can also be arranged in the form of a cloud computing architecture. In a cloud computing architecture, these components can be remotely located and can work together to implement the functions described in this disclosure. In some implementations, cloud computing provides computing, software, data access, and storage services, which do not require an end user to be aware of the physical location or configuration of the system or hardware providing these services. In various implementations, cloud computing uses appropriate protocols to provide services over a wide area network such as the Internet. For example, a cloud computing provider provides applications over a wide area network, and they can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be consolidated at a remote data center location or they can be distributed. The cloud computing infrastructure can provide services through a shared data center, even though they appear as a single access point for users. Thus, the components and functions described herein can be provided from a service provider at a remote location using a cloud computing architecture. Alternatively, they can also be provided from a conventional server, or they can be directly or otherwise installed on a client device.

[0087] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0088] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0089] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0090] In addition, although the operations are depicted in a particular order, this should be understood to require that the operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, the various features described in the context of a single implementation can also be implemented separately or in any suitable sub-combination in multiple implementations.

[0091] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for editing audio, comprising: Obtain a speech signal and text corresponding to the speech signal from the audio to be edited; Modify the text in response to a user editing operation; Adjust a pre-trained speech recognition model based on the speech signal; And Use the adjusted speech recognition model to generate edited audio based on the modified text.

2. The method according to claim 1, wherein the pre-trained speech recognition model comprises a pre-trained voice encoder, and wherein adjusting the pre-trained speech recognition model comprises: Use the speech signal to adjust the pre-trained timbre encoder such that the adjusted timbre encoder generates a coded representation of the timbre in the speech signal.

3. The method according to claim 2, wherein generating the edited audio comprises: Generate a target speech signal based on the coded representation of the timbre and the modified text.

4. The method according to claim 3, wherein the pre-trained speech recognition model comprises a style encoder, and wherein adjusting the pre-trained speech recognition model further comprises: Use the speech signal to adjust the pre-trained style encoder such that the adjusted style encoder generates a coded representation of the style in the speech signal.

5. The method according to claim 4, wherein generating the target speech signal further comprises: Generate the target speech signal based on the coded representation of the timbre, the coded representation of the style, and the modified text.

6. The method according to claim 3, further comprising: Obtain a background sound signal from the audio to be edited; And Synthesize the edited audio based on the background sound signal and the target speech signal.

7. The method according to claim 6, wherein the pre-trained speech recognition model further comprises a pre-trained background encoder, and wherein adjusting the pre-trained speech recognition model further comprises: Use the background sound signal to adjust the pre-trained background encoder such that the adjusted background encoder generates a coded representation of the background sound of the audio to be edited.

8. The method according to claim 1, wherein the obtained text comprises a set of words, each word having an associated timestamp, and wherein modifying the text in response to a user editing operation comprises: Receive the user editing operation, the user editing operation specifying a first word in the set of words and a second word for replacing the first word; And Obtain the modified text by replacing the first word with the second word, in which the second word has the time stamp of the first word in the modified text.

9. The method according to claim 8, wherein generating the edited audio based on the modified text comprises: Generate edited audio based on the modified text with time stamps of words, in which the occurrence time of words in the edited audio is consistent with the audio to be edited.

10. The method according to claim 1, wherein the speech recognition model comprises at least one encoder, and the at least one encoder comprises a first feed-forward module, a multi-head self-attention module, a convolutional module, and a second feed-forward module connected in sequence.

11. An apparatus for editing audio, comprising: A content recognition unit configured to obtain a speech signal and text corresponding to the speech signal from the audio to be edited; A content modification unit configured to modify the text in response to a user editing operation; A model adjustment unit configured to adjust a pre-trained speech recognition model based on the speech signal; And An audio generation unit configured to use the adjusted speech recognition model to generate edited audio based on the modified text.

12. A computing device, comprising: At least one processing unit; And At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the computing device to perform the method according to any one of claims 1 to 10.

13. A non-transitory computer storage medium, comprising machine-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 10.

14. A computer program product, comprising machine-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Self-evolution method and system of speech recognition model

    CN120808760A