Method and apparatus for editing audio, computing device, and medium

By obtaining the voice signal and text from the audio to be edited, adjusting the speech recognition model and generating the edited audio, problems such as tongue errors and word forgetting in user recordings are solved, and efficient and high-quality audio editing is achieved.

WO2025123869A1PCT designated stage expired Publication Date: 2025-06-19LEMON INC(GB) +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/121518
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-14
Filing Date
2024-09-26
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

During the UGC/PUGC creation process, users often experience problems such as erroneous tongue, forgetting words, and unstandard pronunciation when recording audio. Re-recording is expensive and will lead to a lack of environmental sounds, causing trouble to users.

Method used

By obtaining the voice signal and the corresponding text from the audio to be edited, modifying the text in response to user editing operations, and adjusting the pre-trained speech recognition model based on the voice signal, and generating the edited audio using the adjusted model.

Benefits of technology

It enables users to quickly and in real time to edit audio content, while maintaining the sound, style and background sound consistent with the edited audio and original audio, improving user creative efficiency and content quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024121518_19062025_PF_FP_ABST
    Figure CN2024121518_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a method and apparatus for editing an audio, a computing device, and a medium. The method comprises: from an audio to be edited, acquiring a voice signal and a text corresponding to the voice signal; in response to a user editing operation, modifying the text; adjusting a pre-trained voice recognition model on the basis of the voice signal; and using the adjusted voice recognition model to generate an edited audio on the basis of the modified text.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, computing device and medium for editing audio

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed on December 14, 2023, with application number 202311734457.0 and invention name “Method, apparatus, computing device and medium for editing audio”, the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present disclosure relates to the field of computer technology, and in particular to a method, apparatus, computing device, computer-readable storage medium, and computer program product for editing audio. Background Art

[0004] With the development of internet technology, more and more users are participating in the creation of user-generated content (UGC) and professional user-generated content (PUGC). Users can display their original content through internet platforms or make it available to other users. During the UGC / PUGC creation process, users are bound to make mistakes, forget words, or mispronounce words when recording audio or video. While re-recording can solve these problems, it is costly and can result in a loss of ambient sound, which can be a significant problem for users.

[0005] Summary of the Invention

[0006] In view of this, the present disclosure provides a technical solution for editing audio.

[0007] According to a first aspect of the present disclosure, a method for editing audio is provided. The method includes: acquiring a speech signal and text corresponding to the speech signal from the audio to be edited; modifying the text in response to a user editing operation; adjusting a pre-trained speech recognition model based on the speech signal; and generating edited audio based on the modified text using the adjusted speech recognition model.

[0008] In a second aspect of the present disclosure, a device for editing audio is provided. The device includes: a content recognition unit configured to obtain a speech signal and text corresponding to the speech signal from the audio to be edited; a content modification unit configured to modify the text in response to a user editing operation; a model adjustment unit configured to adjust a pre-trained speech recognition model based on the speech signal; and an audio generation unit configured to generate edited audio based on the modified text using the adjusted speech recognition model.

[0009] In a third aspect of the present disclosure, a computing device is provided. The computing device includes at least one processing unit and at least one memory, wherein the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit, and when the instructions are executed by the at least one processing unit, the computing device performs the method according to the first aspect of the present disclosure.

[0010] In a fourth aspect of the present disclosure, a non-transitory computer storage medium is provided, comprising machine-executable instructions, which, when executed by a device, cause the device to perform the method according to the first aspect of the present disclosure.

[0011] In a fifth aspect of the present disclosure, a computer program product is provided, comprising machine-executable instructions, which, when executed by a device, cause the device to perform the method according to the first aspect of the present disclosure.

[0012] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0014] FIG1 illustrates a block diagram of an environment in which various embodiments of the present disclosure can be implemented;

[0015] FIG2 is a block diagram showing an example structure of an audio editing system according to an embodiment of the present disclosure;

[0016] 3A to 3C illustrate examples of a user interface for editing audio according to an embodiment of the present disclosure;

[0017] FIG4 shows a schematic diagram of a process for editing audio according to an embodiment of the present disclosure;

[0018] FIG5 shows a schematic diagram of some components of a speech recognition model according to an embodiment of the present disclosure;

[0019] FIG6 shows a schematic flow chart of a method for editing audio according to an embodiment of the present disclosure;

[0020] FIG7 shows a schematic block diagram of an apparatus for editing audio according to an embodiment of the present disclosure; and

[0021] FIG8 illustrates a block diagram of a computing device capable of implementing some implementations of the present disclosure.

[0022] Throughout the drawings, the same or similar reference numbers denote the same or similar elements. DETAILED DESCRIPTION

[0023] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0024] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The terms "first," "second," etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0025] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0026] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0027] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0028] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0029] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, a "model" may also be referred to as a "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0030] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs. It typically includes an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence so that the output of the previous layer is provided as input to the next layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes the input from the previous layer.

[0031] Generally, machine learning can be roughly divided into three stages, namely the training stage, the testing stage and the use stage (also called the reasoning stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values ​​are continuously updated iteratively until the model can obtain consistent reasoning that meets the expected goals from the training data. Through training, the model can be considered to be able to learn the association between input and output (also called input-to-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. In some implementations, the testing stage can be omitted. In the use stage, the model can be used to process the actual input based on the parameter values ​​obtained through training to determine the corresponding output.

[0032] As mentioned above, during the creation of user-generated content (UGC) and user-generated content (PUGC), users inevitably make mistakes, forget words, or mispronounce words, necessitating re-recording. However, re-recording is costly and can also result in the loss of ambient sound. Some traditional solutions capture the content that needs to be modified and have the user re-record a short segment to overwrite it. This approach can leave noticeable traces of modification, leading to a decrease in the quality of the content.

[0033] In view of this, an embodiment of the present disclosure provides a solution for quickly editing audio, in which a voice signal and corresponding text content are identified from the audio to be edited (e.g., a recording that needs to be modified), and the user can edit the text content that he wants to modify (e.g., delete words or replace them with new words). The voice signal can be provided to a voice recognition model so that the voice recognition model is fine-tuned based on the received voice signal. The fine-tuned voice recognition model stores the feature information of the voice signal. Then, the adjusted voice recognition model is used to generate the edited audio based on the modified text. According to the implementation of the present disclosure, the user can modify the audio as easily as editing the text content, and the edited audio also maintains features consistent with the original audio, such as the timbre, style, background sound, etc. of the speech, thereby improving the user's creative efficiency and content quality.

[0034] FIG1 shows a block diagram of an example environment 100 in which multiple embodiments of the present disclosure can be implemented. In the environment 100, an audio editing system 110 is provided for performing audio editing operations by a user 101 using a terminal device 120. In FIG1 , the audio editing system 110 can be any system with computing capabilities, and the terminal device 120 can be any device with user interaction capabilities. It should be understood that the components and arrangements in the environment shown in FIG1 are merely examples, and a computing system suitable for implementing the example implementations described in the present disclosure may include one or more different components, other components, and / or different arrangements.

[0035] In operation, the audio editing system 110 obtains audio to be edited (sometimes referred to herein as "original audio," the two being used interchangeably) 102 from a terminal device 120, which represents audio content that a user 101 wants to modify, such as a user recording containing slips of the tongue, inappropriate words, etc. In some implementations, the audio to be edited 102 may be separate audio content. Alternatively or additionally, the audio to be edited 102 may be audio embedded in a video. In some implementations, the audio to be edited 102 may also be obtained by the audio editing system 110 from other sources, such as a local or other server.

[0036] The user 101 can also operate the terminal device 120 so that the terminal device 120 provides the editing data 104 to the audio editing system 110. The editing data 104 indicates the modified text, that is, the content modification that the user desires to make to the audio 102 to be edited. In some examples, the editing data 104 may include operations such as adding, deleting, modifying, or replacing the voice content in the audio 102 to be edited. For example, the terminal device 120 can play the audio 102 and present the text content in the audio 102. The user 101 can select some words in the text content to perform deletion or replacement operations, or the user 101 can also enter more text content to generate new text content. The terminal device 102 then provides the new content as the editing data 104 to the audio editing system 110.

[0037] The audio editing system 110 can generate new audio based on the received audio to be edited 102 and the editing data 103, and send the generated new audio as edited audio 115 to the terminal device 120. The user 101 can continue to browse the edited audio 115 and repeat the above process if further modifications are required until the editing is completed. In some implementations, the audio editing system 110 can generate the edited audio 115 based on a neural network model. Using the neural network model, the user's timbre information in the original audio 102 can be retained so that the generated audio has a consistent timbre with the original audio. In some implementations, the audio editing system 110 can also retain the user's speaking style and environmental background sounds in the original audio 102. Therefore, except for the modified text content, the edited audio 115 is basically consistent with the original audio 102. It should be noted that the audio editing system 110 according to the embodiment of the present disclosure can quickly and in real time provide the edited audio to the user 101, thereby significantly improving the efficiency and experience of the user's content creation.

[0038] To facilitate understanding of the embodiments of the present disclosure, the following are exemplary application scenarios. In some scenarios, the embodiments of the present disclosure can be used to correct verbal errors in heavily spoken programs, whether dubbing or spoken, as easily as modifying text in a text file. In some scenarios, the embodiments of the present disclosure can be used to improve user speech fluency by regenerating incoherent parts of the entire speech to achieve smoother speech. In some scenarios, grammatical errors in speech drafts can be automatically detected, enabling quick corrections and making the speech more accurate and professional. In some scenarios, banned word filtering can be implemented. Currently, banned words are covered by a beep sound, but this can now be detected and replaced. In some scenarios, cross-language editing can be implemented, changing a pure Chinese sentence into a mixture of Chinese and English, or modifying a mixture of Chinese and English to pure Chinese or English, or even directly translating between Chinese and English. In some scenarios, secondary creation of film and television plots can be implemented, increasing the fun of user interaction. It should be understood that the above scenarios are merely exemplary, and the embodiments of the present disclosure are also applicable to any other application scenarios. The following will discuss in more detail some specific implementations of the audio editing system of the present disclosure.

[0039] FIG2 is a block diagram illustrating an example structure of an audio editing system according to an embodiment of the present disclosure. As shown in FIG2 , the audio editing system 110 may include a speech separation module 210, an automated speech recognition module 220, a speech recognition model 230, and an audio synthesis module 240. It will be appreciated that the audio editing system shown in FIG2 is merely exemplary, and some of the modules therein may be omitted or modified, and other modules may also be included.

[0040] The speech separation module 210 is used to extract the user's speech signal and background sound signal from the received audio to be edited 102. In some implementations, a filter can be used to generate separated speech signals and background sound signals from the audio to be edited 102 based on spectrum analysis (e.g., independent component analysis method). A spatial filtering method can also be used to collect the sound source signal through a microphone array, and then use beamforming and filtering algorithms to process the mixed signal to achieve separation of speech signals and background sounds. Other methods can also be used to achieve separation of speech signals and background sounds, and the present disclosure is not limited in this respect. If the audio to be edited has no background sound or the background sound is small, the speech separation module 210 can be omitted or not used.

[0041] The automated speech recognition module 220 is used to identify the content of the audio to be edited. The automated speech recognition module 220 acquires the audio to be edited 102 or the speech signal output by the speech separation module 210 and generates text corresponding to the speech signal. In some implementations, the automated speech recognition module may include a feature extraction module, a statistical acoustic model, and a language model. The feature extraction module extracts features from the input speech signal for use in acoustic modeling and decoding. The acoustic model is used to model basic acoustic units such as words, syllables, and phonemes to generate an acoustic model. The language model is used to model the language the system needs to recognize at the word level. The text content generated by the automated speech recognition module 220 may include a set of words, each with corresponding timestamp information. The audio editing system 110 may provide the text content 103 to the terminal device 120 for user browsing. Furthermore, the automated speech recognition module 220 may have the ability to recognize multiple languages, including Chinese, English, Japanese, and other languages, and can output mixed-language text.

[0042] User 101 browses text content 103 on terminal device 120. If user 101 notices that some text content in the audio needs to be modified, user 101 can operate terminal device 120, causing terminal device 120 to provide editing data 104 to audio editing system 110. In some implementations, editing data 104 may include words in text content 103 that need to be modified, replaced, or deleted, and may also include newly added words. The modified words in the text may have corresponding timestamp information.

[0043] In response to receiving the editing data 104, the audio editing system 110 can determine that the audio 101 needs to be edited. As shown in the figure, the audio editing system 110 uses a speech recognition model 230 to perform the editing operation. The speech recognition model 230 obtains the audio to be edited 102 or the speech signal and editing data 104 in the audio to be edited 102, and generates a target speech signal. The target speech signal can be synthesized with the background sound signal to form the edited audio 115 and provided to the terminal device 115. In some implementations, the speech recognition model 230 is used to realize the decoupling of the content, style, and background of the audio. For example, the speech recognition model 230 can include a timbre encoder, a style encoder, and a background sound encoder. These three decoders are respectively used to learn the timbre information, style information, and background sound information in the input audio, so that the generated audio can maintain the timbre, style, and background sound consistent with the original audio. Background sound can refer to other sounds other than the speech in the audio 102, for example, accompaniment instrumental music, environmental noise, etc. It is understood that the speech recognition model 230 may include more or fewer encoders and is not limited to the above-mentioned encoders. It should be noted that the speech recognition model 230 may be pre-trained, that is, it already has knowledge or information about the timbre, style, or background sound. Once it receives the audio 102 or speech signal to be edited, it can quickly adjust its parameters (also known as "fine-tune") to learn the timbre, style, or background sound information of the input audio. For example, the fine-tuning operation can be completed in just a few seconds.

[0044] Upon completion of fine-tuning of the speech recognition model 230, the speech recognition model 230 may switch to the inference phase. At this point, the speech recognition model 230 functions as a decoder, generating a new speech signal based on the received edited data 104, which includes the modified text content indicated by the edited data 104. In some implementations, the speech recognition model 230 may also generate new background sound based on the fine-tuned background sound encoder.

[0045] The audio synthesis module 240 is used to generate edited audio from the new speech signal. It can be a separate module or embedded in the speech recognition model 230. In some implementations, the audio synthesis module 240 can merge the generated new speech signal with the original background sound signal from the speech separation module 210 to form the edited audio 115. Alternatively, the audio synthesis module can also merge the new speech signal with the new background sound from the speech recognition model 230 to form the edited audio 115. Then, the audio editing system 110 sends the edited audio 115 to the terminal device 120. Since the speech recognition model 230 has learned the timbre, style, background sound and other information of the original audio during the fine-tuning stage, the edited audio 115 can modify the original audio content while retaining other features of the original audio 102, just like a real recording.

[0046] Figures 3A to 3C show examples of user interfaces for editing audio according to embodiments of the present disclosure. In the user interface 300A of Figure 3A, users can watch the videos they create and check whether the audio in the video has problems and whether it needs to be edited. In the user interface 300A, a video playback window 305, a text content display window 310 of the audio, an edit icon 315, a video playback progress bar 316, and other controls 318 are included. The text content window 310 can display the text content 103 identified by the automated speech recognition module 220 in a scrolling manner. In certain embodiments, the words in the text content 103 can have associated timestamp information, so that the terminal device 120 can display the text content 311 (I wanna sing) currently being played and the adjacent text before or after it in the window 310 according to the playback progress.

[0047] An exemplary text content 103 is shown below, where the words in the text content 103 have information such as the start time and end time of the words in the video or audio.

[0048] Assuming that the user wants to modify the text content currently displayed in the current window 310, and hopes to change "sing" to "dance", the user 101 can click the editing control 315. Accordingly, the terminal device 120 can enter the audio editing interface 300B shown in Figure 3B.

[0049] In addition to the video playback window 305 and the text content display window 310, interface 300B also shows interface elements 322 to 325 related to editing operations. As shown, when a user clicks the word "sing" in window 310, it indicates that the user 101 has selected "sing" as the object to be edited. Accordingly, the word "sing" is automatically entered into interface element 322. The user can then enter the word "dance" in interface element 323 to replace "sing." If the user wishes to edit other words, the user can continue to click control 324. At this point, interface 300B can add new lines similar to interface elements 322 and 323 for the user to enter the words to be replaced and the desired words. The user's multiple editing operations can be temporarily saved for merging and submission. After completing the editing operation, the user can click the preview control 325 to generate the edited data 104. The terminal device 120 can then provide the edited data 104 to the audio editing system 110, which then generates the new audio. The following shows exemplary editing data 104, which indicates that it is desired to replace the word "dance" with the word numbered 7 (it can be understood that the audio editing system has the number information of each word), ie, "sing".

[0050] Next, the audio recognition system 110 generates a modified text based on the received editing data 104 and the original text content. In the modified text, "dance" can have the same timestamp as "sing". However, the audio recognition system 110 can generate edited audio based on the modified text and send the audio to the terminal device 120. Accordingly, the terminal device 120 presents the modified audio and its text content in the user interface 300C shown in Figure 3C. As shown in Figure 3C, the text content 331 has been modified from "sing" to "dance". At the same time, the terminal device 120 can play the modified audio, and the time when the word "dance" appears in the modified audio is consistent with the "sing" of the original audio.

[0051] FIG4 shows a schematic diagram of a process 400 for editing audio according to an embodiment of the present disclosure. The process 400 may be implemented at the audio editing system 110 shown in FIG1 and FIG2. For ease of understanding, the process 400 will be described with reference to FIG2.

[0052] The audio to be edited 102 may be provided to the speech recognition model 230 of the audio editing system 110. The audio to be edited 102 may include a speech signal and background sound. In some implementations, if there is no separate speech separation module, the audio to be edited may be provided directly to the speech recognition model 230; if there is a speech separation module, the separated speech signal may be provided to the speech recognition module 230.

[0053] As shown, the speech recognition model 230 may include a pre-trained timbre encoder 432, which has been previously trained using training data to recognize the timbre of speech in audio. The timbre encoder 432 may be set in a fine-tuning mode and, in response to receiving a speech signal of the audio to be edited, may be adjusted by changing a portion of its parameters. This adjusted timbre encoder 432 is capable of generating an encoded representation of the timbre of the speech in the audio to be edited, i.e., is capable of recognizing the timbre of the speech in the current audio.

[0054] Additionally or alternatively, the speech recognition model 230 may further include a pre-trained style encoder 434, which has been previously trained using training data to be able to recognize the style of speech in the audio. Similarly, the style encoder 434 may be set in a fine-tuning mode and, in response to receiving a speech signal of the audio to be edited, is adjusted by changing a portion of its parameters so that the adjusted style encoder 434 is able to generate an encoded representation of the style of the speech of the audio to be edited, i.e., is able to recognize the style of the speech of the current audio.

[0055] Additionally or alternatively, the speech recognition model 230 may further include a pre-trained background encoder 436, which has been previously trained using training data to identify features of background sounds in the audio. Similarly, the background encoder 436 may be set in a fine-tuning mode and, in response to receiving a background sound signal of the audio to be edited, may be adjusted by changing a portion of its parameters, such that the adjusted background encoder 436 is capable of generating an encoded representation of the background sound of the audio to be edited, i.e., capable of identifying the background sound of the current audio.

[0056] In process 400, the audio editing system 110 can also generate text content 103 based on the audio to be edited 102, which can be provided to the terminal device 120 to help the user determine whether the audio needs to be edited. For example, the audio editing system 110 can use the automated speech recognition module 220 to generate text content 103. The audio editing system 110 can receive the editing data 104 generated in response to the user's editing operation from the terminal device. Then, the audio editing system 110 can modify the original text content 103 based on the editing data 104 to obtain the modified text. The process of generating the modified text can be completed instantly, and the process of adjusting the encoder of the speech recognition model may take a slightly longer time, such as a few seconds, but compared to the existing model adjustment time has been greatly reduced, because the speech recognition model 430 realizes the decoupling of the speech content, and the number of parameters that need to be adjusted is less.

[0057] Next, the adjusted speech recognition model 235 is used to generate a target speech signal 440 based on the modified text. Specifically, the adjusted speech recognition model 235 can be set to inference mode. The speech recognition model 235 in inference mode includes a decoder 438, which is capable of generating the target speech signal 440 based on the encoded representation of the text and timbre. As mentioned above, the parameters of the fine-tuned timbre encoder have been changed to produce an encoded representation of the timbre of the original audio. As a result, the timbre of the target speech signal 440 generated by the speech recognition model 235 is consistent with the original audio 102, but the content has been changed.

[0058] In some embodiments, the speech signal of the audio to be edited 102 can also be used to adjust the style encoder 434 to generate a corresponding encoded representation of the style. Accordingly, the adjusted speech recognition model 235 can also generate a target speech signal 440 based on the encoded representation of timbre, the encoded representation of style, and the modified text. Thus, the timbre and style of the target speech signal 440 generated by the speech recognition model 235 are consistent with those of the original audio 102, but the content has been modified.

[0059] The audio editing system 110 can then merge the target speech signal 440 with the background sound of the original audio 102 to generate the edited audio 115. In some embodiments, if the separated background sound is not available, the audio editing system 110 can regenerate the background sound, for example, by generating a new background sound based on the encoded representation of the background sound generated by the background encoder, and merge the new background sound with the target speech signal 440 to generate the edited audio 115.

[0060] It can be understood that FIG4 only shows an example process, and there may be other ways to detect errors and perform pattern matching from a signal sequence.

[0061] FIG5 is a schematic diagram showing some components of a speech recognition model according to an embodiment of the present disclosure. Any one of the encoders 432 , 434 , and 436 of the speech recognition model 230 may have the exemplary structure shown in FIG5 .

[0062] As shown in the figure, the exemplary structure includes a first feedforward module 502, a multi-head self-attention module 504, a convolution module 506 and a second feedforward module 508 connected in sequence, wherein the input and output of each module can be combined together by a combiner 510 as the input of the next module. The combiner 510 can be a vector addition or splicing operation. In some implementations, the multi-head self-attention module 504 can include a transformer module for extracting long sequence dependencies. The convolution module 506 is used to extract local features. Through the combination of the multi-head self-attention module 504 and the convolution module 506, the effect of the speech recognition model on long-term sequences and local features is improved simultaneously.

[0063] 6 shows a flow diagram of an example method 600 according to some implementations of the present disclosure. The method 600 may be implemented at the audio editing system 110.

[0064] At block 610, the audio editing system 110 obtains a speech signal and text corresponding to the speech signal from the audio to be edited. At block 620, the audio editing system 110 modifies the text in response to the user's editing operation. At block 630, the audio editing system 110 adjusts a pre-trained speech recognition model based on the speech signal. At block 640, the audio editing system 110 generates edited audio based on the modified text using the adjusted speech recognition model.

[0065] In some embodiments, the pretrained speech recognition model may include a pretrained timbre encoder, wherein adjusting the pretrained speech recognition model may include: adjusting the pretrained timbre encoder using the speech signal so that the adjusted timbre encoder produces an encoded representation of the timbre in the speech signal.

[0066] In some embodiments, generating the edited audio may include generating a target speech signal based on the encoded representation of the timbre and the modified text.

[0067] In some embodiments, the pre-trained speech recognition model may include a style encoder, wherein adjusting the pre-trained speech recognition model may further include: adjusting the pre-trained style encoder using the speech signal so that the adjusted style encoder produces an encoded representation of the style in the speech signal.

[0068] In some embodiments, generating the target speech signal may further include: generating the target speech signal based on the encoded representation of timbre, the encoded representation of style, and the modified text.

[0069] In some embodiments, the method 600 may further include: acquiring a background sound signal from the audio to be edited; and synthesizing the edited audio based on the background sound signal and the target speech signal.

[0070] In some embodiments, the pretrained speech recognition model may also include a pretrained background encoder, wherein adjusting the pretrained speech recognition model may also include: adjusting the pretrained background encoder using the background sound signal so that the adjusted background encoder generates an encoded representation of the background sound of the audio to be edited.

[0071] In some embodiments, the acquired text may include a set of words, each word having an associated timestamp, and modifying the text in response to a user editing operation may include: receiving a user editing operation, the user editing operation specifying a first word in the set of words and a second word to replace the first word; and acquiring a modified text by replacing the first word with the second word, in which the second word has the timestamp of the first word.

[0072] In some embodiments, generating the edited audio based on the modified text may include: generating the edited audio based on the modified text with the timestamp of the words, wherein the appearance time of the words in the edited audio is consistent with the audio to be edited.

[0073] In some embodiments, the speech recognition model includes at least one encoder, which includes a first feedforward module, a multi-head self-attention module, a convolution module and a second feedforward module connected in sequence.

[0074] Fig. 7 shows a schematic block diagram of an apparatus 700 for editing audio according to an embodiment of the present disclosure. As shown in Fig. 7 , the apparatus 700 includes: a content identification unit 702, a content modification unit 704, a model adjustment unit 706, and an audio generation unit 708.

[0075] The content recognition unit 702 is configured to obtain a speech signal and text corresponding to the speech signal from the audio to be edited. The content modification unit 704 is configured to modify the text in response to a user editing operation. The model adjustment unit 706 is configured to adjust a pre-trained speech recognition model based on the speech signal. The audio generation unit 708 is configured to use the adjusted speech recognition model to generate edited audio based on the modified text.

[0076] It should be noted that more actions or steps described with reference to Figures 1 to 6 can be implemented by the apparatus 700 shown in Figure 7. For example, the apparatus 700 may include more modules or units to implement the actions or steps described above, or some units or modules shown in Figure 7 may be further configured to implement the actions or steps described above. This will not be repeated here.

[0077] The exemplary embodiments of the present disclosure have been described above with reference to Figures 2 to 7. Compared to existing solutions, the technical solution provided by the present disclosure allows users to quickly and in real time edit audio content, while maintaining the consistency of other features of the modified audio (such as timbre and ambient background sound). This ensures that the user's timbre and ambient background sound are consistent, eliminating any editing artifacts or unnatural effects.

[0078] FIG8 shows a block diagram of a computing device 800 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 800 shown in FIG8 is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described herein. The computing device 800 can be used to implement the audio editing system 110.

[0079] 8 , computing device 800 comprises a computing device in the form of a general-purpose computing device 800. Components of computing device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage devices 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860.

[0080] In some implementations, the computing device 800 can be implemented as various user terminals or service terminals with computing capabilities. The service terminal can be a server, a large computing device, etc. provided by various service providers. The user terminal is such as a mobile terminal, a fixed terminal, or a portable terminal of any type, including a mobile phone, a site, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is also foreseeable that the computing device 800 can support any type of interface for the user (such as a "wearable" circuit, etc.).

[0081] Processing unit 810 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device 800. Processing unit 810 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0082] The computing device 800 typically includes a plurality of computer storage media. Such media can be any available media accessible to the computing device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The memory 820 can include an audio editing module 822, which are program modules configured to perform the functions of the various implementations described herein. The audio editing module 822 can be accessed and executed by the processing unit 810 to implement the corresponding functions.

[0083] The storage device 830 can be a removable or non-removable medium and can include machine-readable media that can be used to store information and / or data and can be accessed within the computing device 800. The computing device 800 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 8, a disk drive for reading or writing from a removable, non-volatile disk and an optical drive for reading or writing from a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces.

[0084] The communication unit 840 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 800 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 800 can operate in a networked environment using logical connections to one or more other servers, personal computers (PCs), or another general network node.

[0085] Input device 850 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 860 may be one or more output devices, such as a display, speaker, printer, etc. Computing device 800 may also communicate with one or more external devices (not shown) via communication unit 840 as needed, such as storage devices, display devices, etc., with one or more devices that allow a user to interact with computing device 800, or with any device that allows computing device 800 to communicate with one or more other computing devices (e.g., a network card, modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0086] In some implementations, in addition to being integrated on a single device, some or all of the various components of computing device 1200 may be configured in the form of a cloud computing architecture. In a cloud computing architecture, these components may be remotely located and may work together to implement the functionality described herein. In some implementations, cloud computing provides computing, software, data access, and storage services that do not require the end user to be aware of the physical location or configuration of the systems or hardware providing these services. In various implementations, cloud computing provides services over a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides applications over a wide area network, and these applications can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on servers at a remote location. Computing resources in a cloud computing environment may be consolidated at a remote data center location or they may be dispersed. Cloud computing infrastructure may provide services through a shared data center, even though they appear to be a single access point for users. Therefore, the components and functionality described herein may be provided by a service provider at a remote location using a cloud computing architecture. Alternatively, they may be provided from a conventional server, or they may be installed directly or otherwise on a client device.

[0087] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), and the like.

[0088] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0089] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0090] In addition, although each operation is described in a specific order, this should be understood as requiring such operation to be performed in the specific order shown or in a sequential order, or requiring that all illustrated operations should be performed to obtain the desired result. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate implementation can also be implemented in a single implementation in combination. On the contrary, the various features described in the context of a single implementation can also be implemented in multiple implementations individually or in any suitable sub-combination mode.

[0091] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for editing audio, comprising: Acquire a speech signal and text corresponding to the speech signal from the audio to be edited; In response to a user editing operation, modifying the text; adjusting a pre-trained speech recognition model based on the speech signal; as well as Using the adjusted speech recognition model, edited audio is generated based on the modified text.

2. The method of claim 1, wherein the pre-trained speech recognition model comprises a pre-trained timbre encoder, wherein adjusting the pre-trained speech recognition model comprises: The pre-trained timbre encoder is tuned using the speech signal such that the tuned timbre encoder produces an encoded representation of timbre in the speech signal.

3. The method of claim 2, wherein generating the edited audio comprises: A target speech signal is generated based on the encoded representation of the timbre and the modified text.

4. The method of claim 3, wherein the pre-trained speech recognition model comprises a style encoder, wherein adjusting the pre-trained speech recognition model further comprises: The pre-trained style encoder is tuned using the speech signal such that the tuned style encoder produces an encoded representation of the style in the speech signal.

5. The method according to claim 4, wherein generating the target speech signal further comprises: The target speech signal is generated based on the encoded representation of the timbre, the encoded representation of the style, and the modified text.

6. The method according to claim 3, further comprising: Acquire a background sound signal from the audio to be edited; as well as The edited audio is synthesized based on the background sound signal and the target voice signal.

7. The method of claim 6, wherein the pre-trained speech recognition model further comprises a pre-trained background encoder, wherein adjusting the pre-trained speech recognition model further comprises: The pre-trained background encoder is adjusted using the background sound signal so that the adjusted background encoder generates an encoded representation of the background sound of the audio to be edited.

8. The method of claim 1 , wherein the acquired text comprises a set of words, each word having an associated timestamp, and modifying the text in response to a user editing operation comprises: receiving the user editing operation, the user editing operation specifying a first word in the group of words and a second word for replacing the first word; as well as A modified text is obtained by replacing the first word with the second word, in which the second word has a timestamp of the first word.

9. The method of claim 8, wherein generating edited audio based on the modified text comprises: Based on the modified text with the timestamps of the words, an edited audio is generated, in which the appearance time of the words in the edited audio is consistent with the audio to be edited.

10. The method according to claim 1, wherein the speech recognition model comprises at least one encoder, and the at least one encoder comprises a first feedforward module, a multi-head self-attention module, a convolution module and a second feedforward module connected in sequence.

11. A device for editing audio, comprising: A content recognition unit configured to obtain a voice signal and a text corresponding to the voice signal from the audio to be edited; a content modification unit configured to modify the text in response to a user editing operation; A model adjustment unit, configured to adjust a pre-trained speech recognition model based on the speech signal; as well as The audio generation unit is configured to generate edited audio based on the modified text using the adjusted speech recognition model.

12. A computing device comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the computing device to perform the method as claimed in any one of claims 1 to 10.

13. A non-transitory computer storage medium comprising machine executable instructions which, when executed by a device, cause the device to perform the method of any one of claims 1 to 10.

14. A computer program product comprising machine executable instructions which, when executed by a device, cause the device to perform the method of any one of claims 1 to 10.

Citation Information

Patent Citations

  • Audio processing method and electronic device

    CN106971749A

  • Audio editing method and device, electronic equipment and storage medium

    CN113724686A

  • Speech synthesis model training method, speech synthesis method and speech synthesis device

    CN114141228A

  • Voice editing method and device, storage medium and electronic device

    CN116434731A