Editing method and apparatus, device, and storage medium

By recording target timbre reference audio and generating target audio, the problems of low efficiency and unnatural audio in media content editing are solved, more efficient and natural audio generation is achieved, and the user experience is improved.

WO2025195420A1PCT designated stage Publication Date: 2025-09-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/083482
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-22
Filing Date
2025-03-19
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

In the prior art, users are inefficient in adding dubbing to media content and the generated audio is unnatural, affecting the viewing experience.

Method used

By recording target timbre reference audio on the timbre configuration interface, obtaining input text and generating target audio, the input text is read aloud using the target timbre to generate an audio editing result or a video editing result.

Benefits of technology

It improves editing efficiency, generates more natural audio, and enhances the user viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025083482_25092025_PF_FP_ABST
    Figure CN2025083482_25092025_PF_FP_ABST
Patent Text Reader

Abstract

According to embodiments of the present invention, provided are an editing method and apparatus, a device, and a storage medium. The method comprises: in response to an operation of adding a target timbre on a timbre configuration interface, recording a reference audio formed by using the target timbre to read reference text aloud; in response to an operation of triggering audio generation on an editing interface, acquiring input text for audio generation; in response to a selection operation on the target timbre, generating target audio on the basis of the reference audio and the input text, the target audio comprising a speech formed by using the target timbre to read the input text aloud; and on the basis of the target audio, generating an audio editing result or a video editing result. In this way, a timbre in speech data can be used to generate at least part of a target audio. Therefore, a user can use a desired timbre in media content generation, thereby advantageously improving the editing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device and storage medium for editing

[0001] This application claims priority to the Chinese invention patent application entitled “Methods, devices, equipment and storage media for editing” and application number 202410339725.7, filed on March 22, 2024. The entire contents of that application are incorporated herein by reference. Technical Field

[0002] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, apparatuses, devices, and storage media for editing. Background Art

[0003] Currently, more and more applications are designed to provide various services to users. For example, users can publish, browse, and view media content within an application. Media content can include various types of content, such as videos, images, image collections, text, and audio. Users can also perform any appropriate operations on media content within the application, such as editing the published media content. Summary of the Invention

[0004] In a first aspect of the present disclosure, an editing method is provided. The method includes: in response to an operation of adding a target timbre on a timbre configuration interface, recording reference audio of a reference text read aloud using the target timbre; in response to an operation of triggering audio generation on an editing interface, obtaining input text for audio generation; and in response to an operation of selecting the target timbre, generating target audio based on the reference audio and the input text, the target audio including a voice reading aloud using the target timbre of the input text; and generating an audio editing result or a video editing result based on the target audio.

[0005] In a second aspect of the present disclosure, a device for editing is provided. The device includes: an audio recording module configured to, in response to an operation of adding a target timbre on a timbre configuration interface, record reference audio of a reference text read aloud in the target timbre; a text acquisition module configured to, in response to an operation of triggering audio generation on an editing interface, acquire input text for audio generation; an audio generation module configured to, in response to an operation of selecting a target timbre, generate target audio based on the reference audio and the input text, the target audio including a voice reading aloud the input text in the target timbre; and a result generation module configured to generate an audio editing result or a video editing result based on the target audio.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.

[0008] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the method according to the first aspect of the present disclosure is implemented.

[0009] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0011] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] 2A to 2R illustrate schematic diagrams of example interfaces according to some embodiments of the present disclosure;

[0013] 3A to 3G are schematic diagrams showing example interfaces according to other embodiments of the present disclosure;

[0014] FIG4 illustrates a flow chart of a process for editing according to some embodiments of the present disclosure;

[0015] FIG5 shows a block diagram of an apparatus for editing according to some embodiments of the present disclosure; and

[0016] FIG6 shows a block diagram of a device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION

[0017] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0018] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0019] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, for example, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0020] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0021] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0022] As used herein, the term "in response to" refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of executing a subsequent action executed in response to the event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is satisfied. For example, in some cases, a subsequent action may be executed immediately upon the occurrence of the event or the satisfaction of the condition; in other cases, the subsequent action may be executed some time after the occurrence of the event or the satisfaction of the condition.

[0023] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0024] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.

[0025] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.

[0026] As briefly mentioned above, users can also perform any appropriate type of operation on the media content in the application, such as editing the media content to be published. Traditionally, users can add dubbing to the media content to be published by speaking or selecting machine dubbing. However, when speaking, the user needs to recite or read the script, and usually needs to record repeatedly, which results in low dubbing efficiency. When machine dubbing, the user usually needs to select a timbre from a variety of predetermined timbres provided by the application, and the application can generate dubbing audio based on the selected timbre and script. However, the predetermined timbres provided by the application are often more mechanical, which will cause the generated audio to sound unnatural, which will affect the viewing experience of users viewing the media content.

[0027] The embodiments of the present disclosure propose an improved solution for editing. According to various embodiments of the present disclosure, if an operation of adding a target timbre is detected on the timbre configuration interface, a reference audio of a reference text read aloud in the target timbre is recorded. If an operation of generating audio is triggered on the editing interface, the input text for audio generation is obtained. In response to the operation of selecting the target timbre, a target audio is generated based on the reference audio and the input text, and the target audio includes a voice reading aloud the input text in the target timbre. Then, based on the target audio, an audio editing result or a video editing result is generated.

[0028] In this way, target audio can be generated based on reference audio with a target timbre and input text associated with the media content to be generated. This can facilitate users to use audio with a desired timbre to generate target audio during editing. For example, users can use a timbre that is similar to their own without having to tediously record audio themselves. This can advantageously improve the efficiency of editing (e.g., video editing).

[0029] Example embodiments of the present disclosure are described below with reference to the accompanying drawings.

[0030] Example scenario

[0031] 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In the example environment 100, an application 120 is installed in a terminal device 110. A user 140 can interact with the application 120 via the terminal device 110 and / or a device attached to the terminal device 110.

[0032] In some embodiments, the application 120 may be a content sharing application, a content editing application, a content creation application, etc. The application 120 can provide various services related to media content (also referred to as media content items, content items, media items, etc.) to the user 140, including browsing, commenting, forwarding, creating (e.g., shooting and / or editing), and publishing of media content.

[0033] In the environment 100 of FIG1 , if the application 120 is active, the terminal device 110 may present an interface 150 of the application 120. The interface 150 may include various interfaces provided by the application 120, such as a media content presentation interface, a media content creation interface, a media content publishing interface, etc. The application 120 may provide a media content editing function (for example, the application 120 may be an editing application) to support editing (e.g., editing) of media content in the application 120.

[0034] In some embodiments, the terminal device 110 communicates with the server 130 to enable the supply of services to the application 120. The terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.). The server 130 can be various types of computing systems / servers that can provide computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, and the like.

[0035] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.

[0036] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0037] Sample interface

[0038] Figures 2A to 2R show schematic diagrams of example interfaces 200A to 200R (also referred to as examples 200A to 200R) according to some embodiments of the present disclosure. It should be understood that the interfaces shown in the drawings are merely examples, and various interface designs may exist. The various graphical elements in the interface may have different arrangements and different visual representations, one or more of which may be omitted or replaced, and one or more other elements may also be present. The embodiments of the present disclosure are not limited in this respect.

[0039] The interfaces shown in Examples 200A to 200R may be presented at the terminal device 110. For ease of discussion, Examples 200A to 200R will be described with reference to the environment 100 of FIG.

[0040] In environment 100, terminal device 110 presents an interface 150 of application 120. In some cases, the interface 150 presented by terminal device 110 can be a specific editing interface of application 120. Terminal device 110 can present media content in the editing interface and receive user editing operations on the media content via the editing interface. The editing interface can also be referred to as a clipping interface, a creation interface, etc. In this article, each "media content" includes one or more types of content, such as video, image, animation, image set (for example, an image set with sound), audio, text, etc., and each "media content" includes at least audio type content (which can be simply referred to as audio). In some embodiments, the audio of the media content can include dubbing audio of the target text associated with the media content. The target text can, for example, include subtitles, narration, etc. of the media content. The target text can also be the text included in the media content, for example, the text superimposed on the image presented in the media content. It can be understood that the target text can be any appropriate text, and this disclosure is not limited to this.

[0041] As shown in Figure 2A, example 200A shows an example of an editing interface. In example 200A, media content and a plurality of operation controls for editing the media content can be presented. A plurality of operation controls can include an audio editing control (such as can include control 201) for editing the audio of the media content. Terminal equipment 110 can, for example, determine to receive an audio editing indication for the target text associated with the media content in response to receiving a triggering operation to control 201. The triggering operation can include but is not limited to a click operation, a double-click operation, a long press operation, a sliding operation, a drag operation, etc.

[0042] In some embodiments, terminal device 110 can be triggered in response to control 201 and present example 200B as shown in Figure 2B. In some embodiments, terminal device 110 can be triggered for the first time in response to control 201 within a predetermined time period (such as one day, one week, one month, etc.) and present example 200B. Alternatively or additionally, terminal device 110 can also present example 200B in response to control 201 being triggered at every turn. Example 200B includes prompt panel 210. Prompt information associated with audio editing can be presented in prompt panel 210. Terminal device 110 can present the audio editing interface in response to receiving the triggering operation of the control 211 in prompt panel 210.

[0043] Figures 2C to 2E show some examples of audio editing interfaces. In some embodiments, application 120 can provide audio editing services only when the user logs in to the application. In this case, if the user has not logged in to application 120, terminal device 110 can provide example 200C as shown in Figure 2C. Example 200C includes login control 223. Terminal device 110 can determine that the user has logged into the application in response to receiving a trigger operation for login control 223. After the user logs in to application 120, terminal device 110 can provide example 200D or example 200E as shown in Figure 2D or Figure 2E. If the user has not generated a timbre before, terminal device 110 can present example 200D. If the user has already generated a timbre before, terminal device 110 can present example 200E.

[0044] In some embodiments, the audio editing interface (e.g., Examples 200C to 200E) may include an audio configuration entry 220, which is an example of a timbre configuration interface. The terminal device 110 may receive recorded reference audio having a target timbre based on a timbre addition instruction received via the audio configuration entry 220. The target timbre may be used, for example, to simulate the timbre of a target subject (e.g., user 140). The reference audio may include, for example, pitch, timbre, frequency, and the like.

[0045] In some embodiments, the application 120 may also provide at least one timbre. If the timbre that the application 120 can provide is relatively diverse, in order to facilitate the user to obtain the timbre that he or she desires quickly and conveniently, the terminal device 110 may classify the multiple timbres and present multiple navigation tags in the audio configuration entry 220. Each navigation tag may, for example, correspond to multiple timbres of one type. For example, the multiple navigation tags may include "hot", "dialect", "boys' timbre", "girls' timbre" and so on. For example, the terminal device 110 may present the timbre of the type corresponding to the target navigation tag in the audio configuration entry 220 in response to the selection operation of the target navigation tag.

[0046] In some embodiments, the multiple navigation tags presented by the terminal device 110 may also have different levels. For example, the navigation tags "My" (e.g., tag 221), "Popular", "Dialect", "Male Voice", "Female Voice", etc. may be the first level, and the navigation tags "Favorites", "Like", "Clone Voice" (e.g., tag 224), etc. may be the second level. The timbre corresponding to the second-level navigation tag may, for example, be part of at least one timbre corresponding to the first-level navigation tag (e.g., the timbre corresponding to tag 224 is the timbre in the timbre corresponding to tag 221). It should be understood that this is merely exemplary and is not intended to limit the solutions disclosed herein.

[0047] In some embodiments, the terminal device 110 can present a timbre associated with a user (e.g., user 140) in response to receiving a trigger operation on the tag 221. The timbre associated with the user may include, for example, a timbre collected by the user, a timbre liked by the user, a timbre created by the user, and the like. The terminal device 110 can then present a timbre created by the user in response to receiving a trigger operation on the tag 224. The timbre created by the user can be a timbre created based on the user's own reference audio, which can be determined as a target timbre. That is, the user can select a timbre as the target timbre from the user-created timbres presented after the tag 224 is triggered.

[0048] The audio editing interface package may include a timbre adding control. For example, in example 200D, the audio configuration entry 220 may include a control 225. In example 200E, the audio configuration entry 220 may include a control 226. If the user desires to create a target timbre, the user may create the target timbre by triggering the timbre adding control. The terminal device 110 may, for example, determine that a triggering operation for timbre adding has been received in response to receiving a triggering operation for timbre adding control. The terminal device 110 may present an information introduction interface about timbre adding in response to receiving a triggering operation for timbre adding based on the audio configuration entry 220.

[0049] For example, in response to receiving a trigger operation for a timbre-adding control in the audio editing interface, terminal device 110 may present example 200F. Example 200F includes an information introduction interface 230. Information introduction interface 230 may include a control 232. In response to receiving a trigger operation for control 232, terminal device 110 may determine that a timbre-adding instruction has been received. In response to receiving the timbre-adding instruction, terminal device 110 may present a guidance interface.

[0050] In some embodiments, in order to ensure that the acquired data is data authorized by the user, the information introduction interface 230 may further include a control 231 and text associated with the control 231 (e.g., the text "I have read and agreed to "XXXXXX""). The terminal 110 may present the control 231 in the information introduction interface 230 in a first style (e.g., the style shown in FIG2F ) by default. For example, the terminal device 110 may switch to presenting the control 231 in the information introduction interface 230 in a second style (e.g., the style shown in FIG2G ) in response to receiving a trigger operation for the control 231. For example, the terminal device 110 may receive a trigger operation for the control 232 only when the control 231 is presented in the second style. For example, in the example 200F shown in FIG2F , since the control 231 is presented in the information introduction interface 230 in the first style, the terminal device 110 does not receive a trigger operation for the control 232, and the control 232 may be in an unselectable state, for example. For example, the terminal device 110 may present example 200G shown in FIG2G in response to receiving a trigger operation on the control 231 shown in FIG2F . In example 200G, since the control 231 is presented in the second style in the information introduction interface 230, the terminal device 110 may receive a trigger operation on the control 232, and the control 232 may be in a selectable state, for example.

[0051] In some embodiments, in addition to the interfaces shown in Figures 2C to 2G above, the terminal device 110 may also present interfaces as shown in Figures 2H to 2J. Figures 2H and 2I show other examples of audio editing interfaces. In some embodiments, the terminal device 110 may provide an editing interface as shown in Figure 2A only when the user has logged into the application 120. In this case, the terminal device 110 may present example 200B as shown in Figure 2B in response to receiving a triggering operation on the control 201 in Figure 2A. The terminal device 110 may then present an audio editing interface as shown in Figure 2H or Figure 2I in response to receiving a triggering operation on the control 211 in Figure 2B.

[0052] In some embodiments, example 200H and example 200G may also include an audio configuration entry 220, and multiple navigation tags may also be presented in the audio configuration entry 220. Multiple navigation tags may, for example, include a navigation tag "Popular" (e.g., tag 227) and a navigation tag "Personal Tone" (e.g., tag 228). The terminal device 110 may, for example, present at least one tone corresponding to tag 227 in the audio configuration entry 220 by default, that is, the terminal device 110 may present example 200H by default in response to receiving a trigger operation on the control 211 of Figure 2B. The terminal device 110 may, for example, switch to presenting example 200I as shown in Figure 2I in response to receiving a trigger operation on tag 228. The terminal device may, for example, provide a tone addition control (e.g., control 229) in example 200I.

[0053] In response to receiving a trigger operation on control 229, terminal device 110 may present example 200J as shown in FIG2J . Example 200J includes information introduction interface 240. Information introduction interface 240 may include control 242. In response to receiving a trigger operation on control 242, terminal device 110 may determine that a timbre addition instruction has been received. In response to receiving the timbre addition instruction, terminal device 110 may present a guidance interface.

[0054] In some embodiments, to ensure that the acquired data is user-authorized data, the information introduction interface 240 may further include a control 241 and text associated with the control 241 (e.g., the text "I have read and agree to XXXXXX"). For example, in response to the control 241 being triggered, the terminal device 110 may present the control 241 in a second style (e.g., the style shown in FIG. 2J ). When the control 241 is presented in the second style, the terminal device 110 may receive a triggering operation on the control 242.

[0055] In some embodiments, the guidance interface may include reference text for being read aloud in the voice recording. In some embodiments, the guidance interface may further include a recording control. The terminal device 110 may receive audio data recorded by a sound input component (e.g., a microphone) in response to triggering of the recording control in the guidance interface to serve as reference audio for the target object. In some embodiments, in addition to receiving the reference audio in real time, the terminal device 110 may also use pre-stored local audio data as the reference audio.

[0056] Examples 200K to 200P shown in Figures 2K to 2P illustrate multiple examples of a guide interface. As shown in Figure 2K, example 200K includes an area 250 for presenting a reference text and a recording control 251. The terminal device 110 can present example 200L in response to receiving a trigger operation on the recording control 251. In example 200L, the terminal device 110 can present the recording control 251 in a target style, and the target style can indicate that the terminal device 110 is currently receiving audio data recorded by the sound input component to be used as a reference audio of the target object (e.g., user 140). The audio data here can be, for example, audio data of the user 140 reading aloud the reference text. In example 200L, the terminal device 110 can, for example, stop receiving audio data in response to receiving a trigger operation again, and / or in response to the end of a previously received trigger operation. For example, if the terminal device 110 previously started receiving audio data in response to receiving a click operation on the recording control 251. The terminal device 110 may stop receiving audio data in response to receiving a click operation on the recording control 251 again. If the terminal device 110 previously started receiving audio data in response to receiving a long press operation on the recording control 251. For example, when receiving audio data, the terminal device 110 may present a prompt text (such as the text "Recording, please read aloud") in the guidance interface indicating that audio data is currently being received. The terminal device 110 may stop receiving audio data in response to the end of the long press operation.

[0057] In some embodiments, to ensure the quality of subsequently generated timbre, terminal device 110 can independently detect the quality of received audio data. For example, in response to detecting that the audio data does not meet a preset condition, terminal device 110 can present a prompt regarding the preset condition and a re-recording control. The preset conditions may include, but are not limited to, a first condition indicating that the duration of the audio data is within a predetermined duration range, a second condition indicating that the signal-to-noise ratio of the audio data is below a threshold, a third condition indicating that the difference between the text being read aloud and a reference text is less than a threshold difference, and so on. The predetermined duration range, threshold signal-to-noise ratio (i.e., the threshold corresponding to the signal-to-noise ratio), and threshold difference can all be predetermined by a user associated with application 120 (e.g., user 140 or backend staff of application 120) and / or determined independently by terminal device 110. For example, if the predetermined duration range is 50 to 70 seconds, terminal device 110 may determine that the duration of the audio data is not within the predetermined duration range in response to the audio data being 40 seconds, and thus determine that the audio data does not meet the first condition. The terminal device 110 may determine whether the audio data meets the second condition based on a comparison result of the signal-to-noise ratio and the threshold signal-to-noise ratio. For example, if the signal-to-noise ratio of the audio data is 60 and the threshold signal-to-noise ratio is 50, the terminal device 110 may determine that the signal-to-noise ratio reaches the threshold signal-to-noise ratio, and further determine that the audio data does not meet the second condition.

[0058] The terminal device 110 can, for example, use any appropriate method to detect the difference between the text read aloud in the audio data and the reference text. For example, the terminal device 110 can detect the difference between the text read aloud in the audio data and the reference text based on a pre-set rule or algorithm. The terminal device 110 can also use a trained model to detect the difference between the text read aloud in the audio data and the reference text. The model here can include a language model (LM), for example. This model can be deployed locally on the terminal device 110 or on other electronic devices (such as the server 130). The terminal device 110 can, for example, call a model deployed on other electronic devices via a communication connection with other electronic devices. Similarly, the terminal device 110 can determine a difference score based on the determined difference (for example, the smaller the difference, the higher the difference score) and a threshold difference score determined based on the threshold difference. The terminal device 110 can determine whether the audio data meets the third condition based on the comparison result between the difference score and the threshold difference clone. For example, if the difference score is 40 and the threshold difference score is 60, the terminal device 110 may determine that the difference between the text read aloud in the audio data and the reference text is greater than the threshold difference, and further determine that the audio data does not meet the third condition.

[0059] In some embodiments, if the terminal device 110 detects that the audio data does not meet the third condition, the prompt information about the preset condition presented by the terminal device 110 may include highlighting one or more characters in the reference text that are not read aloud, highlighting one or more characters in the reference text that are read aloud incorrectly, and presenting qualitative description text about the difference, etc. The qualitative description text here may indicate that the audio data does not meet the third condition.

[0060] For example, in example 200M shown in FIG. 2M , terminal device 110 may, in response to detecting that the audio data does not satisfy the third condition, highlight one or more characters that were not read aloud in the reference text presented in area 250 and / or highlight one or more characters that were incorrectly read aloud in the reference text. Terminal device 110 may also, in response to detecting that the audio data does not satisfy the third condition, present qualitative description text 252 in area 250.

[0061] The re-recording control and the recording control may be the same control or different controls. Exemplarily, as shown in FIG2M , the re-recording control may be, for example, a recording control 251. The terminal device 110 may, for example, present a prompt text (e.g., the text “Re-recording”) at the recording control 251 indicating that audio data can be received again. In response to triggering the recording control 251, the terminal device 110 may receive additional audio data captured by the sound input component to serve as reference audio.

[0062] In some embodiments, in response to the received audio data satisfying a preset condition, the terminal device 110 may determine the received audio data as reference audio having a target timbre. Exemplarily, the terminal device 110 may present example 200N as shown in FIG2N . Example 200N may include a window 260. Window 260 is used to present prompt text (e.g., text “XX%… in clone timbre generation”) indicating the progress of determining the reference audio based on the received audio data (i.e., the reference audio). Window 260 also includes a cancel control 261. The terminal device 110 may cancel the determination of the reference audio in response to receiving a trigger operation on the cancel control 261.

[0063] In some embodiments, to avoid accidental touch by the user, the terminal device 110 can also present example 200O as shown in Figure 2O in response to receiving a trigger operation for canceling the control 261. Example 200O includes window 270-1. Window 270-1 is used to present prompt text indicating whether to cancel the determination of the reference audio (for example, the text "Cloning is about to be completed. Cancellation will cause the recording material to be lost. Are you sure you want to cancel?"), a determination control 271, and a cancel control 272. The terminal device 110 can cancel the determination of the reference audio based on the audio data in response to receiving a trigger operation for the determination control 271. The terminal device 110 can return to presenting example 200N in response to receiving a trigger operation for canceling the control 272, and continue to determine the reference audio based on the audio data.

[0064] In some embodiments, the terminal device 110 may also present example 200P as shown in FIG2P in response to a failure to determine the reference audio based on the audio data. Example 200P includes window 270-2. Window 270-2 is used to present prompt text indicating that the determination of the reference audio based on the audio data failed (e.g., the text "Exception handling failed, please try again"). In some embodiments, the prompt text in window 270-2 may also indicate the reason for the failure. For example, if the determination of the reference audio based on the audio data fails due to a network problem, the text presented in window 270-1 may be "Network exception handling failed, please check the network." In some embodiments, different options for the next step may be provided for different reasons for failure. For example, in the example of FIG2P , a "Retry" control is provided in the interface. If this control is triggered, an attempt can be made to determine the reference audio based on the audio data using the previously recorded voice again. For another example, a "Rerecord" control may be provided in the interface. If this "Rerecord" control is triggered, the user can return to a guidance interface, such as the guidance interface shown in FIG2K .

[0065] In some embodiments, the terminal device 110 may determine that the reference audio is received in response to the reference audio being determined, and present a playback interface for the target timbre. Example 200Q shown in Figure 2Q shows an example of a playback interface. As shown in Figure 2Q, area 281 of example 200Q may include one or more audio clips. The one or more audio clips here each have a target timbre. In some embodiments, the received audio data may correspond to a first language, and at least one of the one or more audio clips corresponds to a second voice different from the first language. For example, the received audio data may correspond to Chinese, the one or more audio clips may include two audio clips, and one audio clip may correspond to Chinese (for example, the audio clip "Chinese Example Sentence" shown in the figure), and one audio clip may correspond to English (for example, the audio clip "English Example Sentence" shown in the figure).

[0066] Example 200Q may also include corresponding playback controls for one or more audio segments (e.g., the audio segment "Chinese Example Sentence" has a playback control 282, and the audio segment "English Example Sentence" has a playback control 283). Example 200Q may also include an identification configuration entry 284 for a target timbre. The identification configuration entry 284 may present an identification of a control to be presented at the audio configuration entry corresponding to the target timbre (e.g., the text "Timbre 01"). The terminal device 110 may receive user input via a triggering operation of a control 285 in the identification configuration entry 284, and determine the received user input as an identification.

[0067] Example 200Q may also include a confirmation control 286. The terminal device 110 may determine that a confirmation indication for the target timbre has been received in response to receiving a trigger operation on the confirmation control 286. The terminal device 110 may present a usage control for the target timbre at the audio configuration entry in response to receiving a confirmation indication. For example, if the terminal device 110 creates a target timbre in response to receiving a trigger operation on the control 225 in FIG. 2D , the terminal device 110 may present the example 200E shown in FIG. 2E in response to the target timbre being added. The usage control for the newly created target timbre (for example, the usage control "Timbre 01" shown in the figure) may be presented in the audio configuration entry 220 of Example 200E. The terminal device 110 may receive a designation of this target timbre via the usage control. For example, the terminal device 110 may receive a designation of this target timbre via the usage control "Timbre 01".

[0068] The specific steps for determining reference audio with a target timbre have been described above. The specific steps for generating the target audio will now be described in conjunction with the accompanying figures. The audio configuration entry may present corresponding usage controls for one or more timbres (also referred to as corresponding to one or more reference audio with the corresponding timbres). In response to a usage control for a particular timbre among the one or more timbres being triggered, the terminal device 110 may determine the reference audio corresponding to that timbre as the reference audio with the target timbre. For example, as shown in FIG2E , the audio configuration entry 220 may present a usage control for a single timbre (e.g., the usage control "Timbre 01" shown in the figure). Since only one usage control is included, the terminal device 110 may, in response to receiving a trigger operation for that usage control, directly determine the timbre corresponding to that usage control as the target timbre and determine the reference audio with the target timbre. If the audio configuration entry 220 may present usage controls for multiple timbres, the terminal device 110 may, in response to receiving a trigger operation for a particular usage control, determine the timbre corresponding to that usage control as the target timbre and determine the reference audio with the target timbre.

[0069] In some embodiments, the terminal device 110 can also play the audio of at least a portion of the target text read aloud by the timbre in response to receiving a preset operation on the control for using one or more timbres. For example, as shown in Figure 2E, the terminal device 110 can play the audio of at least a portion of the target text read aloud by the timbre 01 in response to receiving a preset operation on the control for using the timbre. That is, the user can audition the audio of at least a portion of the target text read aloud by the timbre by performing a preset operation on the control for using the timbre. Thus, based on the reading effects of different created timbres, the user can select the created timbre with the best effect as the target timbre.

[0070] In some embodiments, the terminal device 110 may also obtain input text for audio generation. For example, if an operation triggering audio generation is detected on the editing interface, such input text may be obtained. This input text can be any appropriate text obtained through any appropriate means. For example, it can be user-defined text. After the terminal device 110 determines reference audio with a target timbre and obtains the input text, it may generate target audio based on the reference audio and the input text. The target audio includes a voice reading the input text in the target timbre. This target audio is audio that speaks the input text in the target timbre. For example, as shown in Figures 2E and 2M, in response to receiving a trigger operation for the control "Timbre 01", the terminal device 110 may determine the timbre corresponding to the control "Timbre 01" as the target timbre and the corresponding reference audio as the reference audio with the target timbre. The audio configuration entry 220 may include a confirmation control 222. In response to receiving a trigger operation for the confirmation control 222, the terminal device 110 may determine to generate the target audio. The target audio may be generated locally by the terminal device 110 or generated by the server 130 providing service support for the application 120. The terminal device 110 may, for example, obtain the target audio from the server 130 via a communication connection with the server 130. During the target audio generation process, the terminal device 110 may present an example 200R as shown in FIG2R . In response to obtaining the target audio, the terminal device 110 may present an example 200A as shown in FIG2A . For example, the terminal device 110 may present an element 202 in the example 200A. The element 202 may indicate the target audio.

[0071] The terminal device 110 can generate an audio editing result or a video editing result based on the target audio, for example, generate media content. In some embodiments, the editing result generated by the terminal device 110 based on the target audio can include a preset mark. For example, the portion of the media content corresponding to the target audio can have a preset mark. The preset mark can indicate that the current audio data is generated using the target timbre. The terminal device 110 can determine the preset mark based on any appropriate method. For example, the terminal device 110 can determine the preset mark with the help of a model. For example, the preset mark can include a clear watermark and a dark watermark. Both the clear watermark and the dark watermark can include any appropriate content such as text and images.

[0072] In some embodiments, if the input text is updated, the terminal device 110 may also update the target audio based on the updated input text in response to the update of the input text. For example, if the previous input text is the text "XXXXXX" and the updated input text is the text "YYYYYY", the terminal device 110 may update the target audio based on the updated input text in response to the update of the input text. The updated target audio may be audio that reads the updated input text in the target timbre.

[0073] Some examples of editing are described above in conjunction with Figures 2A to 2R. The interfaces shown in Figures 2A to 2R can be implemented on one type of terminal device, such as a smartphone. Other examples of content editing are described below in conjunction with Figures 3A to 3G. The interfaces shown in Figures 3A to 3G can be implemented on another type of terminal device, such as a desktop computer, a portable computer, etc. However, it should be understood that this is merely an example and is not intended to limit the scope of this disclosure.

[0074] Figures 3A to 3G show schematic diagrams of example interfaces 300A to 300G (also referred to as examples 300A to 300G) according to some embodiments of the present disclosure. It should be understood that the interfaces shown in the drawings are merely examples, and various interface designs may exist. The various graphical elements in the interface may have different arrangements and different visual representations, one or more of which may be omitted or replaced, and one or more other elements may also exist. The embodiments of the present disclosure are not limited in this respect.

[0075] The interface shown in example 300A to example 300G can also be presented at terminal device 110. For ease of discussion, example 300A to example 300G will be described with reference to the environment 100 of Figure 1. Example 300A and example 300B show some instances of the audio editing interface. In some embodiments, as shown in Figures 3A and 3B, example 300A and example 300B can include a region 310 and a region 320 for presenting media content. Terminal device 110 can be triggered in response to the navigation tag "read aloud" (i.e., tag 321) in region 320, and regard region 320 as an audio configuration entry. The region 320 as the audio configuration entry can include multiple navigation tags. Multiple navigation tags can include a navigation tag "personal timbre" (i.e., tag 322). Terminal device 110 can present the timbre that the user has created in response to receiving a trigger operation to tag 322. After tag 322 is triggered, the user can select a timbre from the user-created timbres presented as a target timbre (i.e., the reference audio corresponding to the timbre is the reference audio with the target timbre). The audio editing interface may include controls for adding timbre. For example, example 300A includes control 323, and example 300B includes controls 324 and 325.

[0076] In some embodiments, the terminal device 110 may also present example 300C as shown in FIG3C in response to receiving a preset operation (e.g., a right-click operation) for the media content. Example 300C shows an example of an audio configuration entry. Example 300C may also include a timbre adding control (e.g., control 331).

[0077] For example, the terminal device 110 may present an information introduction interface about adding a timbre in response to receiving a trigger operation on a timbre adding control. The terminal device 110 may present a guidance interface in response to receiving a trigger operation on a predetermined control in the information introduction interface. Alternatively or additionally, in some embodiments, the terminal device 110 may also present a guidance interface directly in response to receiving a timbre adding instruction (e.g., a trigger operation on a timbre adding control).

[0078] Examples 300D and 300E illustrate some examples of guidance interfaces. As shown in FIG3D , Example 300D may include reference text and a recording control 341. In response to receiving a trigger operation on the recording control 341, the terminal device 110 may receive audio data recorded by the sound input component to serve as reference audio for the target object. In some embodiments, Example 300D may further include a selection entry 342 for the sound input component. The terminal device 110 may determine the target sound input component selected by the user via the selection entry 342 and receive audio data recorded by the target sound input component to serve as reference audio for the target object.

[0079] Furthermore, in response to the received audio data satisfying a preset condition, the terminal device 110 may determine the reference audio based on the received audio data. The terminal device 110 may present example 300E. Example 300E may include a window 350. Window 350 is used to present prompt text indicating the progress of determining the reference audio based on the received audio data (e.g., the text "Clone sound generation in XX%..."). Window 350 also includes a cancel control 351. In response to receiving a trigger operation on the cancel control 351, the terminal device 110 may cancel the determination of the reference audio.

[0080] In some embodiments, the terminal device 110 may present a playback interface for the target timbre in response to the reference audio being determined (e.g., Example 300F). As shown in Figure 3F, area 361 of Example 300F may include one or more audio clips. The one or more audio clips here each have a target timbre. In some embodiments, the received audio data may correspond to a first language, and at least one of the one or more audio clips corresponds to a second voice different from the first language. For example, the received audio data may correspond to Chinese, the one or more audio clips may include two audio clips, and one audio clip may correspond to Chinese (e.g., the audio clip "Chinese Example Sentence" shown in the figure), and one audio clip may correspond to English (e.g., the audio clip "English Example Sentence" shown in the figure).

[0081] Example 300F may also include corresponding playback controls for one or more audio segments (e.g., the audio segment "Chinese Example Sentence" has playback controls 362, and the audio segment "English Example Sentence" has playback controls 363). Example 300F may also include an identification configuration entry 364 for a target timbre. Identification configuration entry 364 may present an identification of a usage control corresponding to the target timbre to be presented at the audio configuration entry (e.g., the text "Timbre 01"). Terminal device 110 may determine the identification of the usage control via identification configuration entry 364.

[0082] Example 300F may also include a confirmation control 365 and a return control 366. In response to receiving a trigger operation on confirmation control 365, the terminal device 110 may determine that a confirmation indication for the target timbre has been received. In response to receiving the confirmation indication, the terminal device 110 may present usage controls for the target timbre (also referred to as usage controls for reference audio having the target timbre) at the audio configuration entry. For example, if the terminal device 110 determines that reference audio having the target timbre has been generated in response to receiving a trigger operation on control 323 in FIG. 3A , the terminal device 110 may present example 300B shown in FIG. 3B , in response to completing the determination of the reference audio (also referred to as the determination of the target timbre). Example 300B may present usage controls (e.g., the usage controls "Timbre 01" shown in the figure) for the newly created target timbre (also referred to as the reference audio having the target timbre). In response to triggering the usage controls (e.g., the usage controls "Timbre 01"), the terminal device 110 may generate the target audio based on the reference audio having the target timbre and the acquired input text.

[0083] Terminal device 110 can determine to return to present the guide interface in response to receiving a trigger operation to return control 366. In some embodiments, to avoid accidental touches, terminal device 110 can, for example, present example 300G as shown in Figure 3G in response to receiving a trigger operation to return control 366. Example 300G includes a determination control 371 and a cancellation control 372. Terminal device 110 can determine to return to present the guide interface in response to receiving a trigger operation to determine control 371. Terminal device 110 can continue to present the playback interface in response to receiving a trigger operation to cancel control 372. Thus, if the user is dissatisfied with the generated target timbre, they can return to regenerate the target timbre.

[0084] In summary, according to the embodiments of the present disclosure, target audio can be generated based on reference audio with a target timbre and input text associated with the media content to be generated. This can facilitate users to use audio with a desired timbre to generate target audio during editing. For example, users can use a timbre that is similar to their own without having to tediously record audio themselves. This can advantageously improve the efficiency of content editing (e.g., video editing).

[0085] Example Process

[0086] FIG4 shows a flow chart of a process 400 for editing according to some embodiments of the present disclosure. The process 400 may be implemented at the terminal device 110. The process 400 is described below with reference to FIG1.

[0087] In block 410 , in response to an operation of adding a target timbre on the timbre configuration interface, the terminal device 110 records a reference audio of a reference text read aloud in the target timbre.

[0088] In block 420 , the terminal device 110 obtains input text for audio generation in response to an operation of triggering audio generation on the editing interface.

[0089] In block 430 , the terminal device 110 generates a target audio based on the reference audio and the input text in response to a selection operation of the target timbre. The target audio includes a voice reading the input text in the target timbre.

[0090] In block 440 , the terminal device 110 generates an audio editing result or a video editing result based on the target audio.

[0091] In some embodiments, the operation of adding a target timbre on the timbre configuration interface includes triggering a timbre adding control in the timbre configuration interface.

[0092] In some embodiments, recording reference audio of a reference text read aloud in a target timbre includes: in response to an operation of adding a target timbre, presenting a guidance interface for guiding the user to record voice; and in response to triggering a recording control in the guidance interface, receiving audio data recorded by a sound input component to be used as reference audio.

[0093] In some embodiments, the guidance interface includes reference text for being read aloud in the voice recording.

[0094] In some embodiments, process 400 further includes: in response to detecting that the audio data does not meet the preset condition, presenting prompt information and a re-recording control about the preset condition; and in response to triggering the re-recording control, receiving additional audio data captured by the sound input component to be used as reference audio.

[0095] In some embodiments, the preset conditions include at least one of the following: a first condition indicating that the duration corresponding to the audio data is within a predetermined duration range, a second condition indicating that the signal-to-noise ratio of the audio data is lower than a threshold, or a third condition indicating that the difference between the text read aloud in the audio data and the reference text is less than a threshold difference, wherein the reference text is text presented for being read aloud in a voice recording.

[0096] In some embodiments, the audio data does not satisfy the third condition, and the prompt information presented about the preset condition includes at least one of the following: highlighting one or more words that are not read aloud in the reference text, highlighting one or more words that are read incorrectly in the reference text, and presenting qualitative descriptive text about the differences.

[0097] In some embodiments, generating target audio based on reference audio and input text includes: in response to the reference audio being recorded, presenting a playback interface for the target timbre, the playback interface including corresponding playback controls for one or more audio clips, each of the one or more audio clips having a target timbre; in response to receiving a confirmation indication for the target timbre via the playback interface, presenting usage controls for the target timbre in the timbre configuration interface; and in response to triggering the usage controls, generating the target audio based on the reference audio and input text.

[0098] In some embodiments, the playback interface further includes an identification configuration entry for the target timbre, and the presented usage control has an identification specified by the identification configuration entry.

[0099] In some embodiments, the reference audio corresponds to a first language and at least one of the one or more audio segments corresponds to a second language different from the first language.

[0100] In some embodiments, presenting a guidance interface for guiding a user to record a voice includes: presenting an information introduction interface about adding a tone in response to receiving a trigger operation for adding a tone based on a tone configuration interface; and presenting a guidance interface in response to detecting an operation of adding a target tone via the information introduction interface.

[0101] In some embodiments, the timbre configuration interface further presents corresponding usage controls for one or more added timbres.

[0102] In some embodiments, process 400 further includes, in response to a preset operation of a control for a first voice of the one or more added voices, playing audio of at least a portion of the input text spoken in the first voice.

[0103] In some embodiments, the portion of the audio editing result or the video editing result corresponding to the target audio has a preset flag, which indicates that the current audio data is generated using the target timbre.

[0104] In some embodiments, process 400 further includes, in response to an update of the input text, updating the target audio based on the updated input text.

[0105] Example devices and equipment

[0106] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above-described methods or processes. FIG5 shows a block diagram of an apparatus 500 for editing according to some embodiments of the present disclosure. Apparatus 500 may be implemented as or included in terminal device 110. Each module / component in apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.

[0107] As shown in Figure 5, the device 500 includes an audio recording module 510, which is configured to record a reference audio of a reference text read aloud in the target timbre in response to an operation of adding a target timbre on the timbre configuration interface. The device 500 also includes a text acquisition module 520, which is configured to acquire an input text for audio generation in response to an operation of triggering audio generation on the editing interface. The device 500 also includes an audio generation module 530, which is configured to generate a target audio based on the reference audio and the input text in response to an operation of selecting the target timbre, wherein the target audio includes a voice reading aloud the input text in the target timbre. The device 500 also includes a result generation module 540, which is configured to generate an audio editing result or a video editing result based on the target audio.

[0108] In some embodiments, the operation of adding a target timbre on the timbre configuration interface includes triggering a timbre adding control in the timbre configuration interface.

[0109] In some embodiments, the audio recording module 510 includes: a guidance interface presentation module, configured to present a guidance interface for guiding the user to record voice in response to receiving a tone addition indication; and a first audio data receiving module, configured to receive audio data recorded by the sound input component in response to triggering the recording control in the guidance interface for use as reference audio.

[0110] In some embodiments, the guidance interface includes reference text for being read aloud in the voice recording.

[0111] In some embodiments, the device 500 also includes: a re-recording module, configured to present prompt information about the preset conditions and re-recording controls in response to detecting that the audio data does not meet the preset conditions; and a second audio data receiving module, configured to receive additional audio data captured by the sound input component in response to triggering the re-recording control for use as reference audio.

[0112] In some embodiments, the preset conditions include at least one of the following: a first condition indicating that the duration corresponding to the audio data is within a predetermined duration range, a second condition indicating that the signal-to-noise ratio of the audio data is lower than a threshold, or a third condition indicating that the difference between the text read aloud in the audio data and the reference text is less than a threshold difference, wherein the reference text is text presented for being read aloud in a voice recording.

[0113] In some embodiments, the audio data does not satisfy the third condition, and the prompt information presented about the preset condition includes at least one of the following: highlighting one or more words that are not read aloud in the reference text, highlighting one or more words that are read incorrectly in the reference text, and presenting qualitative descriptive text about the differences.

[0114] In some embodiments, the audio generation module 530 includes: a playback interface presentation module, configured to present a playback interface for a target timbre in response to a reference audio being recorded, the playback interface including corresponding playback controls for one or more audio clips, each of the one or more audio clips having a target timbre; a usage control presentation module, configured to present usage controls for the target timbre in a timbre configuration interface in response to receiving a confirmation indication for the target timbre via the playback interface; and a target audio generation module, configured to generate target audio based on the reference audio and input text in response to triggering the usage controls.

[0115] In some embodiments, the playback interface further includes an identification configuration entry for the target timbre, and the presented usage control has an identification specified by the identification configuration entry.

[0116] In some embodiments, the reference audio corresponds to a first language and at least one of the one or more audio segments corresponds to a second language different from the first language.

[0117] In some embodiments, the guide interface presentation module includes: an information introduction interface presentation module, configured to present an information introduction interface about tone addition in response to receiving a trigger operation for tone addition based on the tone configuration interface; and a second guide interface presentation module, configured to present the guide interface in response to detecting the operation of adding the target tone via the information introduction interface.

[0118] In some embodiments, the timbre configuration interface further presents corresponding usage controls for one or more added timbres.

[0119] In some embodiments, the device 500 further includes an audio playback module configured to play audio of at least a portion of the input text read aloud using the first tone in response to a preset operation of a control using the first tone among the one or more added tones.

[0120] In some embodiments, a portion of the audio editing result or the video editing result corresponding to the target audio has a preset flag, and the preset flag indicates that the current audio data is generated using the target timbre.

[0121] In some embodiments, the apparatus 500 further includes: an updating module configured to update the target audio based on the updated input text in response to an update of the input text.

[0122] The units and / or modules included in the device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units and / or modules in the device 500 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0123] FIG6 shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 600 shown in FIG6 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 600 shown in FIG6 can be used to implement the terminal device 110 of FIG1 .

[0124] As shown in FIG6 , electronic device 600 is a general-purpose electronic device. Components of electronic device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit 610 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 600.

[0125] The electronic device 600 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 600.

[0126] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG6 , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 620 may include a computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0127] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 600 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 600 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0128] The input device 650 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 660 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 600 may also communicate with one or more external devices (not shown) through the communication unit 640 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 600, or with any device that allows the electronic device 600 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0129] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0130] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0131] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0132] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0133] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0134] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. An editing method, comprising: In response to an operation of adding a target timbre on the timbre configuration interface, recording a reference audio of a reference text read aloud with the target timbre; In response to an operation of triggering audio generation on the editing interface, obtaining input text for audio generation; as well as In response to a selection operation of the target timbre, generating a target audio based on the reference audio and the input text, the target audio including a voice reading the input text in the target timbre; as well as Based on the target audio, an audio editing result or a video editing result is generated.

2. The method according to claim 1, wherein the operation of adding the target timbre on the timbre configuration interface comprises triggering the timbre adding control in the timbre configuration interface.

3. The method according to claim 1, wherein recording a reference audio of the reference text read aloud in the target voice comprises: In response to the operation of adding the target timbre, presenting a guidance interface for guiding the user to record a voice; as well as In response to triggering of a recording control in the guide interface, audio data recorded by a sound input component is received to be used as the reference audio. The method according to claim 3 , wherein the guidance interface includes the reference text for being read aloud in the voice recording.

5. The method according to claim 3, further comprising: In response to detecting that the audio data does not meet a preset condition, presenting prompt information about the preset condition and a re-recording control; as well as In response to triggering of the re-recording control, additional audio data captured by the sound input component is received for use as the reference audio.

6. The method according to claim 5, wherein the preset condition comprises at least one of the following: The first condition indicates that the duration of the audio data is within a predetermined duration range. a second condition indicating that the signal-to-noise ratio of the audio data is below a threshold, or A third condition indicates that the text read aloud in the audio data differs from a reference text by less than a threshold difference, wherein the reference text is text presented for being read aloud in a speech recording.

7. The method according to claim 6, wherein the audio data does not satisfy the third condition, and presenting prompt information about the preset condition comprises at least one of the following: highlighting one or more characters in the reference text that are not read aloud, highlighting one or more characters in the reference text that are mispronounced, A qualitative description of the difference is presented.

8. The method of claim 1 , wherein generating the target audio based on the reference audio and the input text comprises: In response to the reference audio being recorded, presenting a playback interface for the target timbre, the playback interface including corresponding playback controls for one or more audio clips, the one or more audio clips each having the target timbre; In response to receiving a confirmation indication for the target timbre via the playback interface, presenting a usage control for the target timbre in the timbre configuration interface; as well as In response to triggering of the usage control, the target audio is generated based on the reference audio and the input text.

9. The method according to claim 8, wherein the playback interface further comprises an identification configuration entry for the target timbre, and the presented usage control has an identification specified by the identification configuration entry.

10. The method of claim 8, wherein the reference audio corresponds to a first language and at least one of the one or more audio segments corresponds to a second language different from the first language.

11. The method according to claim 3, wherein presenting a guidance interface for guiding the user to record a voice comprises: In response to receiving a trigger operation for adding a timbre based on the timbre configuration interface, presenting an information introduction interface about adding the timbre; as well as In response to detecting the operation of adding the target timbre via the information introduction interface, the guide interface is presented.

12. The method according to claim 1, wherein the timbre configuration interface further presents corresponding usage controls for one or more added timbres.

13. The method according to claim 12, further comprising: In response to a preset operation using a control for a first tone among the one or more added tones, audio of at least a portion of the input text read aloud using the first tone is played.

14. The method according to claim 1, wherein a portion of the audio editing result or the video editing result corresponding to the target audio has a preset flag, and the preset flag indicates that current audio data is generated using the target timbre.

15. The method according to claim 1, further comprising: In response to an update of the input text, the target audio is updated based on the updated input text.

16. An apparatus for editing, comprising: an audio recording module configured to record a reference audio of a reference text read aloud with the target timbre in response to an operation of adding the target timbre on the timbre configuration interface; A text acquisition module is configured to acquire input text for audio generation in response to an operation of triggering audio generation on the editing interface; an audio generation module configured to generate a target audio based on the reference audio and the input text in response to a selection operation of the target timbre, the target audio including a voice reading the input text in the target timbre; as well as The result generating module is configured to generate an audio editing result or a video editing result based on the target audio.

17. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 15 when executed by the at least one processing unit.

18. A computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the method according to any one of claims 1 to 15.

19. A computer program product comprising a computer program, wherein the computer program implements the method according to any one of claims 1 to 15 when executed by a processor.

Citation Information

Patent Citations

  • Editing method and device, equipment and storage medium

    CN120692427A

  • Video dubbing method and device, computer device and computer readable storage medium

    CN110933330A

  • Audio editing method, electronic equipment and storage medium

    CN114023301A

  • Voice playing method and device, equipment and storage medium

    CN114121028A

  • Audio editing method and device, equipment and storage medium

    CN114915836A