Audio editing method and system, terminal and storage medium

By converting the original audio into text and editing to generate new audio, the problems of audio editing complexity and low editing efficiency in the existing technology are solved, and efficient and fine audio editing is achieved.

CN120148518APending Publication Date: 2025-06-13GUANGDONG-HONG KONG-MACAO GREATER BAY AREA DIGITAL ECONOMY RESEARCH INSTITUTE (INTERNATIONAL ADVANCED TECHNOLOGY APPLICATION PROMOTION CENTER (SHENZHEN) +1

Patent Information

Application Number
CN202510616027.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art audio editing system is complex for non-professional users, and the lightweight online editor cannot achieve refined local editing, and the editing efficiency and quality are difficult to meet the requirements.

Method used

By determining the original text corresponding to the original audio, and obtaining the target text modified to some text, deleting the audio band corresponding to some text in the original audio to obtain the reference audio, and finally generating the target audio based on the original text, the target text and the reference audio.

Benefits of technology

It provides an intelligent voice editing processing method, which is simple and easy to operate, suitable for ordinary users, improves the efficiency and quality of voice editing and reduces the computing burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148518A_ABST
    Figure CN120148518A_ABST
Patent Text Reader

Abstract

The invention discloses an audio editing method and system, a terminal and a storage medium, and the method comprises the steps: determining an original text corresponding to an original audio, and obtaining a target text after part of the original text is modified; deleting an audio segment corresponding to the partial text in the original audio to obtain a reference audio; and based on the original text, the target text and the reference audio, generating a target audio corresponding to the target text. The invention provides an intelligent voice editing processing method which does not depend on a locally installed voice editing system and is high in usability. Moreover, the reference audio is generated after the audio segments corresponding to the partial text are deleted, and then the target audio is generated based on the reference audio, the original text and the target text after the partial text is modified, so that the calculation burden is reduced, and the voice editing effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and in particular, to an audio editing method, system, terminal, and storage medium. Background Art

[0002] Currently, voice editing systems are mainly divided into two categories: professional editing software and lightweight online editors. Professional editing software requires users to install local applications and relies on high-performance GPUs and CPUs to support computing resources, which is often difficult for non-professional users to quickly get started, and the audio editing process is cumbersome and the operation threshold is relatively high. While lightweight online editors are simple and easy to use for ordinary users, they cannot achieve refined local editing, and it is difficult to meet the requirements in terms of editing efficiency and editing quality.

[0003] Therefore, there are still defects in the prior art. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an audio editing method, system, terminal, and storage medium in view of the above-mentioned defects of the prior art. The technical solutions adopted by the present invention are as follows: In a first aspect, the present invention provides an audio editing method, wherein the method includes: Determine the original text corresponding to the original audio, and obtain a target text obtained by modifying a part of the text in the original text; Delete the audio segment corresponding to the part of the text in the original audio to obtain a reference audio; Generate a target audio corresponding to the target text based on the original text, the target text, and the reference audio.

[0005] In one implementation, the obtaining of the target text obtained by modifying a part of the text in the original text includes: Proofread and realign the original text; Determine a part of the text in the realigned original text, and modify the part of the text in the realigned original text to the target content to obtain the target text.

[0006] In one implementation, the proofreading and realigning of the original text includes: Proofread the original text based on the actual speech content of the original audio; Realign the proofread original text with the actual content of the original audio to obtain the start and end times of each word in the original text.

[0007] In one implementation, the deleting of the audio segment corresponding to the part of the text in the original audio to obtain a reference audio includes: Match the start and end times of each word with a partial text in the original audio to obtain the start and end times of the partial text; Obtain the reference audio based on the start and end times of the partial text.

[0008] In one implementation, the obtaining the reference audio based on the start and end times of the partial text includes: Determine an audio segment corresponding to the partial text based on the start and end times of the partial text; Delete the audio segment corresponding to the partial text in the original audio, and splice the remaining audio intervals to obtain the reference audio.

[0009] In one implementation, the generating the target audio corresponding to the target text based on the original text, the target text, and the reference audio includes: After replacing the partial text in the original text with the target text, obtain the edited complete text; Input the edited complete text and the reference audio into a preset speech editing model to output the target audio.

[0010] In one implementation, the method further includes: Splice the target audio and the reference audio according to the time stamp; Perform smoothing processing on the splicing boundary to obtain the edited audio.

[0011] In one implementation, the splicing the target audio and the reference audio according to the time stamp includes: Upsample the resolution of the target audio by a preset value; If the preset value is inconsistent with the resolution of the original audio, adjust the resolution of the target audio to be consistent with the resolution of the original audio; Splice the adjusted target audio and the reference audio according to the time stamp.

[0012] In one implementation, the method further includes: Obtain the speech part and the background sound part in the original audio, and delete the background sound part.

[0013] In one implementation, the method further includes: Obtain the background sound part of the original video, and synthesize the background sound part of the original video and the edited audio to restore the background sound.

[0014] In one implementation, the synthesizing the background sound part of the original video and the edited audio includes: If the duration of the edited audio is less than the duration of the original audio, truncate the background sound part of the original video; Stitch the truncated background sound parts together, perform smoothing processing on the stitching boundaries, and synthesize the stitched background sound part with the edited audio, where the duration of the stitched background sound part is equal to the duration of the edited audio.

[0015] In one implementation, synthesizing the background sound part of the original video with the edited audio includes: If the duration of the edited audio is greater than the duration of the original audio, expand the background sound part of the original video so that the duration of the expanded background sound part is equal to the duration of the edited audio; Synthesize the expanded background sound part with the edited audio.

[0016] In one implementation, the method further includes: Obtain a reference text and match the reference text with the edited audio; If the reference text fails to match the edited audio, re-edit the original audio.

[0017] In one implementation, re-editing the original audio includes: Modify the target text again, and based on the re-modified target text and the reference audio, obtain a new edited audio.

[0018] In one implementation, re-editing the original audio further includes: Based on the edited audio, re-determine part of the text in the edited audio, and based on the re-determined part of the text, re-determine the reference audio and re-determine the target text; Based on the re-determined reference audio and the re-determined target text, re-determine the target audio.

[0019] In a second aspect, an embodiment of the present invention further provides an audio editing system, where the system is used to implement the steps of the audio editing method described in any one of the above, and the system includes: An original audio processing module, configured to determine the original text corresponding to the original audio and obtain a target text obtained by modifying part of the text in the original text; A target text determination module, configured to delete the audio segment corresponding to the part of the text in the original audio to obtain a reference audio; A target audio generation module, configured to generate a target audio corresponding to the target text based on the original text, the target text, and the reference audio.

[0020] In a third aspect, an embodiment of the present invention further provides a terminal, where the terminal includes a memory, a processor, and an audio editing program stored in the memory and executable on the processor. When the processor executes the audio editing program, the steps of the audio editing method in any one of the above solutions are implemented.

[0021] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, where an audio editing program is stored on the computer-readable storage medium, and the audio editing program implements the steps of the audio editing method in any one of the above solutions on the computer-readable storage medium.

[0022] Beneficial effects: Compared with the prior art, the present invention provides an audio editing method. First, the original text corresponding to the original audio is determined, and a target text obtained by modifying a part of the text in the original text is acquired. Then, the audio segment corresponding to the part of the text in the original audio is deleted to obtain a reference audio. Finally, based on the original text, the target text, and the reference audio, the target audio corresponding to the target text is generated. The present invention provides an intelligent voice editing and processing method, which does not depend on a locally installed voice editing system, is simple and easy to operate, and has strong usability. Moreover, since the present invention only deletes the audio segment corresponding to a part of the text to generate a reference audio, and then generates the target audio based on the reference audio and the target text obtained by modifying a part of the text, the computational burden is reduced, which is convenient for improving the voice editing quality. Description of the Drawings

[0023] Figure 1 It is a flowchart of a preferred embodiment of the audio editing method provided by an embodiment of the present invention.

[0024] Figure 2 It is a specific application flowchart of the audio editing method provided by an embodiment of the present invention.

[0025] Figure 3 It is a schematic diagram of determining the start and end times of a part of the text in the audio editing method provided by an embodiment of the present invention.

[0026] Figure 4 It is a principle illustration diagram of voice editing in the audio editing method provided by an embodiment of the present invention.

[0027] Figure 5 It is a schematic diagram of the architecture of the audio editing system provided by an embodiment of the present invention.

[0028] Figure 6 It is a principle block diagram of the terminal provided by an embodiment of the present invention. Detailed Embodiments

[0029] To make the objectives, technical solutions and effects of the present invention clearer and more explicit, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not intended to limit the present invention.

[0030] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the content, operations or steps, nor do they necessarily need to be executed in the described order. For example, some operations or steps can be decomposed, combined or partially merged, so the actual execution order may be changed according to the actual situation.

[0031] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms. It should be understood that, in order to facilitate a clear description of the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. For example, the first control information and the second control information are only used to distinguish different control information, and do not limit their sequence. Those skilled in the art can understand that the terms such as "first" and "second" do not limit the quantity and execution order, and the terms such as "first" and "second" do not necessarily mean different. It should also be understood that the term "and / or" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0032] In the prior art, professional editing software requires users to install local applications and relies on high-performance GPUs and CPUs to support computing resources, which has high requirements for the hardware configuration of devices. If the hardware conditions of users do not meet the requirements, they may not be able to use these tools smoothly. Moreover, due to the waveform-based editing method and complex operation options adopted by professional editing software, non-professional users often find it difficult to get started quickly, and the audio editing process is cumbersome and the operation threshold is relatively high. On the other hand, lightweight online editors are generally implemented relying on small programs or web pages, which are simple and easy to use and suitable for ordinary users to quickly perform basic editing. However, these tools have relatively simple functions and cannot achieve refined local editing. In addition, due to certain errors in the current automatic speech recognition and speech synthesis technologies, the generated results often do not meet the expectations of users, resulting in users having to generate repeatedly, which affects the editing efficiency and quality. Furthermore, when processing long audio, most current speech editing tools are prone to problems such as freezing, crashing, or out-of-memory issues, especially when dealing with the efficient processing of a large amount of data, and the optimization effect is not good.

[0033] To address these problems, this embodiment provides an audio editing method. This audio editing method does not need to rely on a locally installed speech editing system or software, and optimizes the audio processing efficiency and editing accuracy. Specifically, in practical application, this embodiment first determines the original text corresponding to the original audio, and obtains the target text obtained by modifying part of the text in the original text. Then, the audio segment corresponding to the part of the text in the original audio is deleted to obtain a reference audio. Finally, based on the original text, the target text, and the reference audio, the target audio corresponding to the target text is generated. The audio editing method of this embodiment is simple and easy to operate, and has strong usability. Moreover, since this embodiment deletes the audio segment corresponding to part of the text to generate a reference audio, and then generates the target audio based on the reference audio, the original text, and the target text obtained by modifying part of the text, the computational burden is reduced, which is convenient for improving the quality of speech editing.

[0034] The audio editing method of this embodiment can be applied to a terminal or the cloud. The terminal can be an intelligent product terminal such as a mobile phone, a computer, or a smart TV. Whether applied to a terminal or the cloud, the audio editing method of this embodiment does not require the installation of any speech editing software or system, such as Figure 1 As shown in Step S100: Determine the original text corresponding to the original audio, and obtain the target text obtained by modifying part of the text in the original text.

[0035] Specifically, in combination with Figure 2As shown, in this embodiment, the original audio is first received. The original audio can be a recording uploaded by the user through the system or other audio files, and the original audio reflects the audio data and the total duration. In one implementation, after obtaining the original audio, this embodiment can perform audio noise reduction processing on the original audio. Specifically, in practical applications, this embodiment can use background sound separation technology to separate the speech part and the background noise part from the audio, reducing the interference of environmental noise (such as wind noise, vehicle noise, etc.) or background music on the speech. This embodiment can load a noise reduction model (such as DeepFilterNet, an efficient full-band audio noise reduction framework using deep filtering) through a noise reduction module to identify and remove the background audio that does not belong to the speech. If the background sound part in the original audio is too noisy, the user can independently start the noise reduction model to improve the clarity and intelligibility of the audio. The noise-reduced audio can improve the accuracy of recognition, alignment, and synthesis in subsequent steps. The audio noise reduction process in this embodiment is an optional process, and in practical applications, it can be determined whether to perform noise reduction based on the actual speech content of the original audio.

[0036] Furthermore, this embodiment can use speech recognition technology to perform speech-to-text operation on the original audio to obtain the original text corresponding to the original audio. Specifically, in implementation, a neural network model for speech recognition (such as Whisper, a general speech recognition model that is trained on a large dataset containing various audio and is also a multi-task model that can perform multi-language speech recognition, speech translation, and language recognition) can be used to process the original audio, extract speech features and convert them into text information, thereby obtaining the original text. This embodiment converts the speech content in the original audio into the original text, which is beneficial for the user to perform subsequent editing and proofreading.

[0037] After obtaining the original text, this embodiment can provide a text proofreading function to proofread and realign the original text. Then, based on the realigned original text, part of the text in the realigned original text is determined, and part of the text in the realigned original text is modified to the target content to obtain the target text. Specifically, this embodiment proofreads the original text based on the actual speech content of the original audio. The user can correct the incorrect words or sentences in the original text to make the original text more consistent with the actual speech content. After proofreading, this embodiment can input the proofread original text into the realignment module to realign the proofread original text with the actual content of the original audio to obtain the start and end times of each word in the original text. By realigning the original text, this embodiment can make the original text consistent with the content of the original audio, thereby improving the recognition accuracy. This process can be implemented based on a deep learning model for speech and text alignment, which can accurately obtain the correspondence between the timestamps in the original audio and the original text, thereby achieving realignment. After completing the realignment, the user selects part of the text from the aligned original text, such as one or consecutive words. As Figure 3 shown, the circular boxes represent the words in the original text, where the white part is the part that does not need to be modified, the green part represents the selected part of the text, that is, the content to be modified, and the yellow part represents the modified part. To reduce the computational burden of the long original audio, this embodiment can segment the original audio. T represents the start and end times of the segmented audio segments (such as Figure 3 T1 to T2 in it, and each audio segment may contain one or more words), and t represents the start and end times of each word. By segmenting the original audio and then analyzing the text and audio in each audio segment, this embodiment ensures the accurate matching of speech and text and improves the consistency between the audio and text content. Finally, this embodiment can modify part of the text in the realigned original text to the target content to obtain the target text. For example, if the realigned original text is "I love music" and part of the text is "love", then the part of the text "love" can be modified to the target content "always listen", and the target text is "always listen".

[0038] By introducing a noise reduction, proofreading, and realignment mechanism, this embodiment solves the deficiencies of traditional speech recognition technology in audio editing. Especially in a noisy environment, the system effectively removes background noise through the noise reduction module to ensure the clarity of the speech content, thereby improving the recognition accuracy. The user can proofread and modify the automatically recognized original text in real time, and then accurately align the modified original text with the original audio through the realignment step. This process effectively improves the matching degree between speech and text and enhances the accuracy and efficiency of speech editing.

[0039] Step S200: Delete the audio segment corresponding to the partial text in the original audio to obtain a reference audio.

[0040] Specifically, in this embodiment, based on the realigned original text, the start and end times of each word can be obtained. Then, the start and end times of each word are matched with the partial text in the realigned original text to obtain the start and end times of the partial text. When determining the start and end times, each word and its corresponding timestamp can be obtained. If the word is Chinese without spaces, it is based on characters; if the word is English with spaces, it is based on words, and the time is accurate to milliseconds (ms). Finally, in this embodiment, based on the start and end times of the partial text, the reference audio can be obtained. When determining the reference audio, in this embodiment, based on the start and end times of the partial text, the audio segment corresponding to the partial text can be determined. Then, the audio segment corresponding to the partial text in the original audio is deleted, and the remaining audio intervals are spliced to obtain the reference audio, which can be used for subsequent audio splicing. It should be noted that if there are multiple determined partial texts, there will also be multiple audio segments corresponding to the partial texts. When obtaining the reference audio, the audio segments corresponding to multiple partial texts in the original text are deleted in the order of timestamps, and a mark is made in the deletion order when each audio segment corresponding to a partial text is deleted. After all the audio segments corresponding to the partial texts are deleted, the remaining audio intervals can be spliced according to the marked order to obtain the reference audio.

[0041] For example, the realigned original text is "I love music", the start and end times of "I " in the original text "I love music" are t1 and t2, the start and end times of "love" are t3 and t4, and the start and end times of "music" are t5 and t6. If the partial text is "love", the audio segment corresponding to the partial text is the audio interval from t3 to t4. In this embodiment, the audio interval from t3 to t4 in the original audio can be deleted, and then the remaining audio intervals are spliced to obtain the reference audio. Therefore, the reference audio is the audio interval from t1 to t2 and from t5 to t6 corresponding to "I music".

[0042] Step S300: Generate the target audio corresponding to the target text based on the original text, the target text, and the reference audio.

[0043] After obtaining the target text and the reference audio, in this embodiment, part of the text in the original text can be replaced with the target text to obtain the edited complete text, and then the edited complete text and the reference audio are input into a preset speech editing model together to output the target audio corresponding to the target text. The speech editing model can be a text-to-speech tool (e.g., Voicecraft), or, adopting the Transformer architecture and combined with the token rearrangement process, it can achieve the ability to efficiently generate speech in the audio sequence. Based on this speech editing model, the target audio is generated through autoregressive inference, and it can reproduce the ambient sound, speaker timbre, and intonation characteristics of the reference audio. As Figure 4 shown, when the reference audio "I music" and the edited complete text "I always listen music" are input into the speech editing model together, the speech editing model autoregressively outputs the audio of "I music always listen" according to the reference audio, then deletes the audio of "I music", and finally obtains the audio of "always listen". Therefore, the target audio content corresponding to the output target text is "always listen".

[0044] Further, after obtaining the target audio, in this embodiment, the target audio can be spliced with the reference audio according to the timestamp. Then, smoothing processing is performed on the splicing boundary to obtain the edited audio. The speech editing model in this embodiment is generated frame by frame in an autoregressive manner and can automatically predict the stop. Therefore, the total length of the target audio is not considered during the audio splicing process.

[0045] In one implementation, when the target audio is spliced with the reference audio, in this embodiment, the resolution of all target audios can be upsampled by a preset value based on the super-resolution model of the neural network, such as sampled to 48 kHz. If the preset value is inconsistent with the resolution of the original audio (generally less than or equal to 48 kHz), the resolution of the target audio is adjusted to be the same as the resolution of the original audio. Then, according to the timestamp, the adjusted target audio is spliced into the middle of the reference audio, and smoothing processing is performed on the splicing boundary to obtain the edited audio. In another implementation, this embodiment can also upsample the reference audio to the preset value and then splice it with the target audio and perform smoothing processing on the splicing boundary to obtain the edited audio. Thus, it can be seen that this embodiment can ensure that the quality of the output audio is consistent with the input, the front and back are smoothly connected and natural, and meets the user's needs. At the same time, by introducing the audio upsampling technology, the sampling rate and sound quality of the audio are improved to ensure that the edited audio effect is as close as possible to the natural performance of the original audio.

[0046] In other implementation manners, this embodiment can also adjust the parameters in the edited audio as needed. For example, by adjusting the parameters to avoid slurring and unnatural repetitions, or by adjusting the variation range of the voice, or by adjusting the duration of the maximum audio segment when splicing audio once. By adjusting these parameters, the user can further optimize the generated target audio to make it more in line with the requirements. In other implementation manners, this embodiment can further adjust the generation parameters of the edited audio for more refined audio synthesis. The user can set different voice parameters (such as speech rate, text segmentation length, editing range, etc.) to optimize the audio effect, so as to meet scenarios with higher requirements for audio quality, such as dubbing, translation and other fields.

[0047] In other implementation manners, if there is a lot of background noise in the original audio, noise reduction processing can be performed on the original audio before speech recognition to separate the background noise part from the speech part, and only perform speech recognition on the speech part. When splicing the target audio and the reference audio, the background noise part of the original video can be obtained, and the background noise part of the original video is synthesized with the edited audio to restore the background noise, so that the edited audio retains the sound field environment of the original audio. Specifically, if the duration of the edited audio is less than the duration of the original audio, the background noise part of the original video is truncated, and then the truncated background noise part is spliced, and the boundary of the splicing is smoothed, and the spliced background noise part is synthesized with the edited audio, and the duration of the spliced background noise part is equal to the duration of the edited audio. If the duration of the edited audio is greater than the duration of the original audio, the background noise part of the original video is extended (such as repeating a certain section of the background noise) so that the duration of the extended background noise part is equal to the duration of the edited audio; finally, the extended background noise part is synthesized with the edited audio.

[0048] In other implementation manners, after the edited audio is generated, this embodiment can obtain the reference text and match the reference text with the edited audio to determine whether the generated edited audio meets the expectations. If the matching between the reference text and the edited audio fails, the original audio is re-edited. Specifically, this embodiment can re-modify the target text, and at this time, realignment is not required. For example, Figure 3 for the part of "re-modifying the target text", the target text is Figure 3 the yellow part after editing in Figure 3As shown in the part of "retaining the edited audio and modifying other words", at this time, the edited audio is used as the original audio, and steps such as recognition, alignment, and synthesis are repeated, and other words are modified. For example, it is re-determined that part of the text is from t7 to t9, and then based on the re-determined part of the text, the reference audio and the re-determined target text are re-determined. Finally, based on the re-determined reference audio and the re-determined target text, the edited audio is re-determined.

[0049] In summary, compared with the existing voice editing methods, the present invention does not rely on a locally installed professional voice editing system and has strong usability. The present invention provides an audio editing interaction process. Only by loading a voice editing model and uploading an audio can the audio editing effect be obtained in a timely manner, reducing the requirements for the user's device, avoiding the dependence on local GPUs and CPUs in professional software, enabling ordinary users to easily use it without being restricted by the device performance. In addition, the present invention solves the deficiencies in traditional audio editing by introducing mechanisms such as noise reduction, proofreading, and re-alignment. Especially in a noisy environment, the present invention can effectively remove background noise through a noise reduction module to ensure the clarity of the audio content, thereby improving the accuracy of speech recognition. Users can also proofread and modify the recognized original text in real time, and then accurately align the modified text with the original audio through the re-alignment step, effectively improving the matching degree between speech and text and enhancing the accuracy and efficiency of voice editing.

[0050] In addition, since the present invention only generates a reference audio for part of the text and then generates a target audio based on the reference audio and the target text, the computational burden is reduced, which is convenient for improving the voice editing efficiency.

[0051] Moreover, the present invention also introduces a segmentation mechanism for segmenting the original audio, solving the problems of lag and crash in the existing audio editing system when processing long audio. This mechanism adjusts the analyzed audio length to avoid overloading the system when processing overly long audio, improving the computational efficiency. In addition, in some implementation manners, the present invention can also support users to select a suitable editing area and limit the analyzed audio length, greatly improving the synthesis efficiency and the stability of the system.

[0052] Based on the above embodiments, the present invention also provides an audio editing system, which is used to implement the steps in the above method embodiments, such as Figure 5As shown in the figure, the system includes: an original audio processing module 10, a target text determination module 20, and a target audio generation module 30. The original audio processing module 10 can be used to implement noise reduction, recognition, text conversion, re-alignment of text and audio, etc. of the original audio. Specifically, the original audio processing module 10 is used to determine the original text corresponding to the original audio, and obtain a target text obtained by modifying part of the text in the original text. The target text determination module 20 is used to delete the audio segment corresponding to the part of the text in the original audio to obtain a reference audio. The target audio generation module 30 is used to generate a target audio corresponding to the target text based on the original text, the target text, and the reference audio.

[0053] The working principles of the modules in the audio editing system of this embodiment are the same as those of the steps in the above method embodiment, and will not be elaborated here.

[0054] Each module in the above audio editing system can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor in the terminal in hardware form or independent of it, or stored in the memory in the terminal in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0055] Based on the above embodiments, the present invention also provides a terminal, and the principle block diagram of the terminal can be as Figure 6 shown. The terminal may include one or more processors 100 ( Figure 6 only one is shown in the figure), a memory 101, and a computer program 102 stored in the memory 101 and executable on one or more processors 100. For example, an audio editing program. When one or more processors 100 execute the computer program 102, each step in the audio editing method embodiment can be implemented. Or, when one or more processors 100 execute the computer program 102, the functions of each module / unit in the audio editing system embodiment can be implemented, and no limitation is made here.

[0056] In one embodiment, the so-called processor 100 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.

[0057] In one embodiment, the memory 101 may be an internal storage unit of the electronic device, such as the hard disk or memory of the electronic device. The memory 101 may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 101 may also include both the internal storage unit and the external storage device of the electronic device. The memory 101 is used to store computer programs and other programs and data required by the terminal. The memory 101 may also be used to temporarily store data that has been output or is to be output.

[0058] Those skilled in the art can understand that Figure 6 the principle block diagram shown in

[0059] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, operational database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An audio editing method, characterized in that: The method comprises: Determine an original text corresponding to the original audio, and obtain a target text after modifying a portion of the original text; Deleting the audio segment corresponding to the part of the text in the original audio to obtain a reference audio; Based on the original text, the target text and the reference audio, a target audio corresponding to the target text is generated.

2. The audio editing method according to claim 1, characterized in that: The step of obtaining a target text after modifying a portion of the original text includes: Proofreading and realigning the original text; A portion of the text in the realigned original text is determined, and the portion of the text in the realigned original text is modified into target content to obtain the target text.

3. The audio editing method according to claim 2, characterized in that: The proofreading and realigning of the original text includes: Proofreading the original text based on the actual speech content of the original audio; The proofread original text is realigned with the actual content of the original audio to obtain the start and end time of each word in the original text.

4. The audio editing method according to claim 3, characterized in that: The deleting the audio segment corresponding to the part of the text in the original audio to obtain the reference audio includes: Matching the start and end time of each word with the partial text in the original audio to obtain the start and end time of the partial text; The reference audio is obtained based on the start and end time of the partial text.

5. The audio editing method according to claim 4, characterized in that: The step of obtaining the reference audio based on the start and end time of the partial text includes: Determining an audio segment corresponding to the partial text based on the start and end time of the partial text; The audio segment corresponding to the partial text in the original audio is deleted, and the remaining audio intervals are spliced ​​to obtain the reference audio.

6. The audio editing method according to claim 1, characterized in that: The step of generating a target audio corresponding to the target text based on the original text, the target text and the reference audio includes: After replacing part of the original text with the target text, an edited complete text is obtained; The edited complete text and the reference audio are input into a preset speech editing model, and the target audio corresponding to the target text is output.

7. The audio editing method according to claim 1, characterized in that: The method further comprises: splicing the target audio with the reference audio according to the timestamp; The spliced ​​boundaries are smoothed to obtain the edited audio.

8. The audio editing method according to claim 7, characterized in that: The step of splicing the target audio with the reference audio according to the timestamp includes: Upsampling the resolution of the target audio to a preset value; If the preset value is inconsistent with the resolution of the original audio, adjusting the resolution of the target audio to be consistent with the resolution of the original audio; The adjusted target audio is concatenated with the reference audio according to the timestamp.

9. The audio editing method according to claim 8, characterized in that: The method further comprises: The speech part and the background sound part in the original audio are obtained, and the background sound part is deleted.

10. The audio editing method according to claim 9, characterized in that: The method further comprises: The background sound portion of the original video is obtained, and the background sound portion of the original video is synthesized with the edited audio to restore the background sound.

11. The audio editing method according to claim 10, characterized in that: The background sound part of the original video is synthesized with the edited audio, comprising: If the duration of the edited audio is shorter than the duration of the original audio, the background sound portion of the original video is cut off; The truncated background sound parts are spliced ​​together, and the spliced ​​boundaries are smoothed. The spliced ​​background sound parts are synthesized with the edited audio, wherein the duration of the spliced ​​background sound parts is equal to the duration of the edited audio.

12. The audio editing method according to claim 10, characterized in that: The background sound part of the original video is synthesized with the edited audio, comprising: If the duration of the edited audio is longer than the duration of the original audio, the background sound portion of the original video is extended so that the duration of the extended background sound portion is equal to the duration of the edited audio; Synthesize the expanded background sound part with the edited audio.

13. The audio editing method according to claim 7, characterized in that: The method further comprises: Obtaining a reference text and matching the reference text with the edited audio; If the reference text fails to match the edited audio, the original audio is re-edited.

14. The audio editing method according to claim 13, characterized in that: Re-editing the original audio, including: The target text is re-edited, and a new edited audio is obtained based on the re-edited target text and the reference audio.

15. The audio editing method according to claim 14, characterized in that: Re-editing the original audio also includes: Based on the edited audio, re-determine a portion of the text in the edited audio, and based on the re-determined portion of the text, re-determine the reference audio and re-determine the target text; The target audio is re-determined based on the re-determined reference audio and the re-determined target text.

16. An audio editing system, characterized in that: The system is used to implement the steps of the audio editing method according to any one of claims 1 to 15, and the system comprises: An original audio processing module, used to determine the original text corresponding to the original audio, and obtain a target text after modifying part of the original text; A target text determination module, used for deleting the audio segment corresponding to the part of the text in the original audio to obtain a reference audio; A target audio generation module is used to generate a target audio corresponding to the target text based on the original text, the target text and the reference audio.

17. A terminal, characterized in that: The terminal includes a memory, a processor, and an audio editing program stored in the memory and executable on the processor. When the processor executes the audio editing program, the steps of the audio editing method according to any one of claims 1 to 15 are implemented.

18. A computer-readable storage medium, characterized in that: An audio editing program is stored on the computer-readable storage medium, and the audio editing program implements the steps of the audio editing method according to any one of claims 1 to 15 on the computer-readable storage medium.

Citation Information

Patent Citations

  • Voice data processing method and device

    CN109036422A

  • Sound and text realignment and information presentation method and device, electronic equipment and storage medium

    CN113761865A

  • Text-based voice editing method and system, electronic equipment and storage medium

    CN115966196A

  • Voice editing and optimizing method and device, equipment and storage medium

    CN117409762A

  • A speech editing and synthesis method and system based on autoregressive model

    CN119763542A

Cited By

  • Audio processing method and device, electronic equipment and storage medium

    CN121459818A