Method and apparatus for content generation, device, and storage medium
By presenting an audio editing panel on the terminal device and using a machine learning model to generate text and audio that match the content entities, the complexity and poor quality of existing music generator technologies are solved, enabling lightweight and personalized audio and video generation.
Patent Information
- Application Number
- PCT/CN2025/096143
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-30
- Filing Date
- 2025-05-20
- Publication Date
- 2025-12-04
AI Technical Summary
Existing music generator technologies are subject to significant limitations and have complex processes, resulting in poor audio conversion quality and an inability to effectively generate audio that matches the content entity.
The system responds to users' audio editing requests via the terminal device, presents an audio editing panel, and uses a machine learning model to generate text and audio that match the content entity, which are then overlaid on the content entity to form a video.
It provides a lightweight, fun, and personalized way to generate audio and video, enhancing the user experience by generating audio that fits the content without requiring user input.
Smart Images

Figure CN2025096143_04122025_PF_FP_ABST
Abstract
Description
Methods, apparatus, devices, and storage media for content generation
[0001] This application claims priority to Chinese Patent Application No. 202410693952.X, filed on May 30, 2024, entitled "Method, Apparatus, Device and Storage Medium for Content Generation", the entire contents of which are incorporated herein by reference. Technical Field
[0002] The exemplary embodiments disclosed herein generally relate to the field of computers, and more particularly to methods, apparatus, devices, and computer-readable storage media for content generation. Background Technology
[0003] More and more applications are now designed to provide users with various services. Many applications support user messaging. In the process of people exchanging information via the internet, various types of audio have become an important medium for social expression and information exchange. Summary of the Invention
[0004] In a first aspect of this disclosure, a method for content generation is provided. The method includes: in response to an audio editing request, presenting an audio editing panel, the audio editing panel including at least an audio generation control; in response to detecting a trigger on the audio generation control, obtaining first text and a first audio corresponding to the first text, at least one of the first text and the first audio being determined based on a content entity to be edited; and adding the first text and the first audio to the content entity to obtain a first video, wherein the first text is overlaid on the content entity, and the first audio is configured as at least a portion of the audio corresponding to the first video.
[0005] In a second aspect of this disclosure, an apparatus for content generation is provided. The apparatus includes: an editing panel presentation module configured to present an audio editing panel in response to an audio editing request, the audio editing panel including at least an audio generation control; a text and audio acquisition module configured to acquire first text and a first audio corresponding to the first text in response to detecting a trigger on the audio generation control, at least one of the first text and the first audio being determined based on a content entity to be edited; and a video acquisition module configured to add the first text and the first audio to the content entity to obtain a first video, wherein the first text is overlaid on the content entity, and the first audio is configured to be at least a portion of the audio corresponding to the first video.
[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.
[0008] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.
[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0012] Figures 2A to 2F show schematic diagrams of example interfaces for content generation according to some embodiments of the present disclosure;
[0013] Figure 3 shows a flowchart of a content generation process according to some embodiments of the present disclosure;
[0014] Figure 4 shows a block diagram of an apparatus for content generation according to some embodiments of the present disclosure; and
[0015] Figure 5 shows a block diagram of an apparatus capable of implementing several embodiments of the present disclosure. Detailed Implementation
[0016] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0017] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0018] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0019] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0020] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0021] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0022] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0023] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.
[0024] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0025] As used in this paper, the term "model" refers to a system that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processors to process inputs and provide corresponding outputs. In this paper, "model" may also be referred to as a "machine learning model," a "machine learning network," or simply a "network," and these terms are used interchangeably. A model can also include different types of processors or networks.
[0026] As used herein, a “unit,” “operation unit,” or “subunit” can consist of any suitable machine learning model or network. As used herein, a set of elements or similar expressions can include one or more such elements. For example, “a set of convolutional units” can include one or more convolutional units.
[0027] The term "work" in this disclosure refers to any type of media content or media work, which includes one or more types of content, including but not limited to audio files, video files, image files, text files, etc. Specifically, a work can be a short video, music, images, image collections, multimedia clips, audiovisual materials, etc. This disclosure is not limited in this respect.
[0028] As briefly described above, various types of audio have become an important medium for social expression and information exchange in the process of people interacting with information via the internet. Currently, it is possible to generate audio and music, or audio with a musical flavor, using music generators. However, music generators themselves are technically limited, and the process is long and complex, resulting in poor audio conversion quality.
[0029] Embodiments of this disclosure propose a scheme for content generation. According to various embodiments of this disclosure, if a user initiates an audio editing request, the user's terminal device displays an audio editing panel that includes at least an audio generation control. If the user clicks the audio generation control, the terminal device obtains first text and a first audio corresponding to the first text based on the user's trigger. At least one of the first text and the first audio is determined based on the content entity to be edited. The terminal device then adds the first text and the first audio to the content entity to obtain a first video, wherein the first text is overlaid on the content entity, and the first audio is configured as at least a portion of the audio corresponding to the first video.
[0030] Therefore, without requiring user input and based on an understanding of content entities, it is possible to easily generate videos composed of text and its corresponding audio. This provides users with a lightweight, fun, and personalized way to generate audio and video, enhancing the user experience.
[0031] Example Environment
[0032] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Environment 100 includes one or more users 110-1, 110-2, 110-3, ..., 110-N who can send and receive messages through their respective associated terminal devices 120-1, 120-2, 120-3, ..., 120-N. For ease of discussion, users 110-1, 110-2, 110-3, ..., 110-N may be collectively referred to as user 110 or individually, and terminal devices 120-1, 120-2, 120-3, ..., 120-N may be collectively referred to as terminal device 120 or individually. In some scenarios, user 110 can publish and comment on works on a target platform through associated terminal device 120. In some scenarios, user 110 is also referred to as the publisher of the work.
[0033] Terminal device 120 may have an application 125 that supports message interaction installed (i.e., terminal device 120-1 has application 125-1 installed, terminal device 120-2 has application 125-2 installed, terminal device 120-3 has application 125-3 installed, ..., terminal device 120-N has application 125-N installed). It should be noted that the application 125 installed on different terminal devices 120 can be the exact same application or different applications (e.g., different versions). Application 125 can be any suitable application with message sending and receiving functions, such as a dedicated chat application, a social application, a content sharing application, an office support application, etc.
[0034] In environment 100 of Figure 1, if application 125 is active, terminal device 120 can display the user interface of application 125. This user interface can include various interfaces provided by application 125, such as user interfaces supporting message interaction, user interfaces supporting content browsing, message sending and receiving interfaces, etc. Through different user interfaces, application 125 can provide different content to user 110. Through appropriate means, such as clicking or selecting any appropriate element in the user interface, application 125 can also provide user 110 with the option to select and switch the presentation mode of related content.
[0035] In some embodiments, different terminal devices 120 can also communicate with server 130 via network 132 to provide message interaction services. Server 130 can provide functions such as management, configuration and maintenance of application 125.
[0036] Terminal device 120 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 120 may also support any type of user-facing interface (such as "wearable" circuitry). Server 130 can be any type of computing system / server capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.
[0037] It should be understood that the structure and function of environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0038] The following description, with reference to the accompanying drawings, outlines some exemplary embodiments of this disclosure. It should be understood that the pages shown in the drawings are merely examples, and various page designs are possible in practice. The graphic elements on the page may have different arrangements and visual representations, one or more elements may be omitted or replaced, and one or more other elements may also be present. The embodiments of this disclosure are not limited in this respect.
[0039] In the following description, the example embodiments will be primarily from the perspective of terminal device 120. It should be understood that the actions described relative to terminal device 120 can be performed by an application on terminal device 120, and / or can be performed by the application in conjunction with its server (e.g., a server).
[0040] The present disclosure's scheme for content generation is described below with reference to Figures 2A to 2F. Figures 2A to 2F illustrate schematic diagrams of example interfaces 201 to 206 for content generation according to some embodiments of the present disclosure.
[0041] In some embodiments, if terminal device 120 receives an audio editing request, it presents an audio editing panel that includes at least an audio generation control. An audio editing request is, for example, a request initiated by a user when they wish to publish or edit a work. In some examples, in a scenario where user 110 begins editing or publishing a work, if terminal device 120 receives a trigger operation from user 110 on the "Select Music" control, it presents an audio editing panel that includes at least an audio generation control.
[0042] In other examples, when user 110 begins editing or publishing a work on the content entity to be edited, if terminal device 120 receives a trigger operation from user 110 on the "Select Music" control, it will present an audio editing panel that includes at least the content entity and an audio generation control. As shown in the example interface 201 of Figure 2A, after terminal device 120 receives an audio editing request (e.g., user 110 clicks the "Select Music" control 215), it presents an audio editing panel 216 that includes at least the content entity 212 to be edited and an audio generation control 211. In some examples, the content entity 212 to be edited can be used to indicate the content included in the video and / or image (e.g., video and / or image, image collection).
[0043] In some examples, the audio generation control 211 may be presented by the terminal device 120 under the “Recommended” list 213 included in the audio editing panel 216. Alternatively, the audio generation control 211 may also be presented by the terminal device 120 under other lists included in the audio editing panel 216 (e.g., presented under the “Favorites” list 214), as shown in Figure 2A is merely exemplary.
[0044] In some embodiments, the visual style of the audio generation control is randomly selected from multiple candidate visual styles when the terminal device 120 presents the audio editing panel once or multiple times. In some examples, the visual style of the audio generation control 211 presented by the terminal device 120 each time the audio editing panel is presented is randomly selected from multiple candidate visual styles. It is understood that the terminal device 120 randomly presents different animation designs of the audio generation control 211 each time the audio editing panel is presented or during loading.
[0045] In some embodiments, if the terminal device 120 detects a trigger operation on the audio generation control, it obtains the first text and the first audio corresponding to the first text. In some embodiments, the terminal device 120 determines the first text and the first audio based on the content entity to be edited. As shown in the example interfaces 201 to 203 of Figures 2A to 2C, if the terminal device 120 detects a trigger operation by the user 110 on the audio generation control 211, it determines the text 232 (e.g., text XXXXX) and the automatic dubbing A 231 corresponding to the text 232 based on the content entity to be edited 212, as shown in Figure 2C. In some examples, after the user 110 clicks the audio generation control 211, as shown in the interface 202 of Figure 2B, the terminal device 120 can display a prompt message 221 (e.g., "Generating...") in the area corresponding to the audio generation control 211 to inform the user 110 that the audio generation process is currently underway.
[0046] In some embodiments, upon detecting a user's triggering of the audio generation control 211, an audio generation request is sent to a server (e.g., server 130). This audio generation request may include the current content entity 212 or descriptive information about the content entity 212. The server generates first text based on the content entity by invoking a machine learning model. The machine learning model used to generate the text may be a generative model implemented based on a language model. The input to the machine learning model may be prompt words and the content entity 212 (or descriptive information about the content entity 212), and the output may be the first text. The first text may be text with a specific style suitable for being added to the content entity. After obtaining the first text, the server may also invoke a text-to-speech (TTS) model to generate first audio (speech) corresponding to the first text. In some embodiments, the invocation of the model may be executed locally on the terminal device. The terminal device 120 may invoke the machine learning model to generate the first text and invoke the TTS model to generate the first audio.
[0047] In some embodiments, after obtaining the first text and the first audio, the terminal device 120 adds the first text and the first audio to a content entity to obtain a first video. In some embodiments, the terminal device 120 overlays the first text onto the content entity, and the first audio is configured as at least a portion of the audio corresponding to the first video.
[0048] As shown in the example interfaces 203 to 204 of Figures 2C to 2D, after obtaining text 232 and the automatically generated voiceover A 231, the terminal device 120 adds text 232 and automatically generated voiceover A 231 to the content entity 212 to form a video. That is, the terminal device 120 presents text 232 on the content entity 212 and configures automatically generated voiceover A 231 as part of the audio corresponding to the video. In some examples, if the content entity 212 is a static image or a collection of images, adding automatically generated voiceover A 231 will result in a video. If the content entity 212 is a video, adding automatically generated voiceover A 231 will still result in a video.
[0049] In some embodiments, the terminal device 120 may obtain the first text by sampling a portion of the content entity and extracting first semantic information corresponding to the sampled portion using a semantic model. Then, the terminal device 120 generates the first text using a machine learning model based on the first semantic information and the prompt information. In some examples, when the content entity 212 is a video, the terminal device 120 samples several frames by frame extraction. When the content entity 212 is an image, the terminal device 120 samples a portion of the image region.
[0050] For example, when content entity 212 is a video, terminal device 120 extracts frames from the video to obtain partial frames. Then, terminal device 120 extracts semantic information from these partial frames using a semantic model. Next, based on the semantic information and prompts, terminal device 120 uses a machine learning model to understand the video and generates new, engaging, and user-expected text based on it. Understandably, by sampling partial content from content entity 212, terminal device 120 can avoid unauthorized use of information at remote locations and ensure data security.
[0051] In some embodiments, the first timbre type of the first audio is generated based on content entities using a machine learning model. In some examples, the machine learning model may run remotely, such as on a server, or locally, such as on a terminal device. Accordingly, the first audio may be generated by performing text-to-speech on the first text based on the first timbre type. In some examples, the timbre type corresponding to the audio of the read-aloud text 232 may also be generated based on content entities using a machine learning model.
[0052] In some embodiments, the terminal device 120 can determine a user-specified timbre type as the first timbre type. Before creating text and audio, candidate timbre types can be provided to the user, and the user's selection can be received. For example, the terminal device 120 can determine the timbre type A selected by the user as the timbre type for reading text 232, and generate an automatic voiceover A 231 with the type A timbre through a TTS function. In this process, text 232 is still automatically generated by a machine learning model, while the audio timbre can be selected by the user.
[0053] In some embodiments, a timbre type can be randomly selected from a timbre library as the first timbre type. For example, whenever text is generated by a machine learning model, a timbre of type B can be randomly selected from the timbre library and determined as the timbre type for reading text 232. Then, an automatic dubbing A 231 with the type B timbre can be generated using the TTS function. It should be understood that the timbres or target timbres available for user selection in this disclosure are timbres that are already in the timbre library and are licensed for use.
[0054] In summary, when a user cannot find audio that matches the content entity, the terminal device 120 can generate an audio that matches the content entity by using a machine learning model or by randomly selecting a timbre from a timbre library.
[0055] In some embodiments, the text style of the first text is determined by a machine learning model based on the content entity. Terminal device 120 uses the machine learning model to determine the text style of text 232 based on the content entity. For example, if the machine learning model determines that the content entity is a lively style, then terminal device 120 ultimately determines that the text style of text 232 is a lively style. If the machine learning model determines that the content entity is a somber style, then terminal device 120 ultimately determines that the text style of text 232 is a solemn style. In some embodiments, the first timbre type of the first audio is determined by a machine learning model based on the content entity. The timbre type may indicate a timbre that matches the content entity, and / or may be a timbre that matches the first text.
[0056] The preceding text described how the terminal device, in response to a user's triggering of the audio generation control, generates text and its corresponding audio, which is then applied to the video. The following text will continue to describe, with reference to Figures 2D to 2F, the adjustment of the text and its corresponding audio.
[0057] In some embodiments, the terminal device 120 may also present adjustment controls associated with the first video. The adjustment controls are for adjusting first text and first audio in the target content. If the terminal device 120 detects a trigger operation on the adjustment controls, it obtains second text and second audio to replace the first text and first audio. The terminal device 120 can determine the second text and second audio based on the content entity. Then, the terminal device 120 adds the second text and second audio to the content entity to obtain the second video. The terminal device 120 overlays the second text onto the content entity and configures the second audio as at least a portion of the audio corresponding to the second video.
[0058] As shown in the example interfaces 204 to 205 in Figures 2D to 2E, the terminal device 120 can also present adjustment controls 233 in association with the first video, allowing the user 110 to adjust the text 232 and the automatic voice-over A 231. If the terminal device 120 detects that the user 110 clicks the adjustment control 233, it obtains the text 252 (e.g., text YYYYY) determined again based on the content entity and the corresponding automatic voice-over B 251. Subsequently, the terminal device 120 replaces the text 232 in the first video with the text 252 and replaces the automatic voice-over A 231 in the first video with the automatic voice-over B 251 to form the second video. In some examples, the text 252 is overlaid on the content entity 212.
[0059] In some examples, if terminal device 120 detects a trigger on adjustment control 233, it can adjust the text in the first video to another text. If terminal device 120 detects a trigger on adjustment control 233, it can also randomly adjust the timbre of the audio in the first video to another timbre.
[0060] In some embodiments, if the terminal device 120 detects a trigger on the adjustment control, it presents a text style selection entry and / or a timbre selection entry. The terminal device 120 receives a selection of a second text style through the text style selection entry and / or a selection of a second timbre type through the timbre selection entry. Then, the terminal device 120 obtains a second text with the second text style and a second audio with the second timbre type to replace the first text and the first audio, respectively.
[0061] In some examples, if terminal device 120 detects a trigger on the adjustment controls, it presents a text style selection entry and / or a tone selection entry. If terminal device 120 detects that user 110 clicks the text style selection entry, it presents a list of text styles. If terminal device 120 detects that user 110 clicks the tone selection entry, it presents a list of tones. For example, if user 110 selects text style X from the text style list, terminal device 120 obtains text with text style X based on the text style entry. Terminal device 120 reads the text with text style X aloud in a first tone to form a second audio. Then, terminal device 120 adds the second audio and the text with style X to the content entity to generate a second video.
[0062] For example, if user 110 selects tone Y from the tone list, terminal device 120 obtains a second audio with tone Y based on the tone selection entry. Terminal device 120 reads aloud a first text (e.g., text 232) with tone Y to form the second audio. Then, terminal device 120 adds the second audio and text with style X to the content entity to generate a second video. For example, if user 110 selects tone YY from the tone list and text style XX from the text style list, terminal device 120 obtains a second audio based on the tone selection entry and / or text style selection entry. This second audio is formed by reading aloud text with text style XX (e.g., text 232) with tone YY. Then, terminal device 120 adds the second audio and text with style XX to the content entity to generate a second video.
[0063] In some embodiments, the terminal device 120 may also present an adjustment panel 261, which may include adjustment controls 233, text-to-speech control 263, and editing controls 264. In some examples, the terminal device 120 may present the adjustment panel 261 if it detects that the user 110 has triggered an action on a region associated with the text 232. Alternatively, the terminal device 120 may present the adjustment panel 261 if it detects that the user 110 has triggered an action on a settings control; this disclosure is not limited thereto.
[0064] In some embodiments, if the terminal device 120 detects a triggering of an editing operation on the first text, it presents a text editing box corresponding to the first text. The terminal device 120 receives the updated third text through the text editing box and overlays the third text onto the content entity. Accordingly, the terminal device 120 removes the first audio from the content entity.
[0065] As shown in the example interface 206 of Figure 2F, if the terminal device 120 detects that the user 110 has triggered an editing operation on the text 232, it presents the text editing box corresponding to the text 232. In some examples, the terminal device 120 presents the text editing box corresponding to the text 232 if it detects that the user 110 has clicked the editing control 264. Alternatively, the terminal device 120 may also present the text editing box corresponding to the text 232 if it detects that the user 110 has clicked the area 261 associated with the text 232. Then, the terminal device 120 receives the updated third text (e.g., the text AAAAA) through the text editing box and overlays the third text onto the content entity 212. Accordingly, the terminal device 120 removes the automatic voiceover A 231 from the content entity 212.
[0066] In some embodiments, if the terminal device 120 detects a text-to-speech request for third text, it obtains the third audio corresponding to the third text. Then, the terminal device 120 adds the third audio to the content entity to obtain the third video. As shown in the example interface 206 of Figure 2F, if the user 110 clicks the text-to-speech control 263, the terminal device 120 can convert the third text into the corresponding third audio. Then, the terminal device 120 adds the third audio and the third text to the content entity to obtain the third video.
[0067] In some examples, if user 110 clicks the text-to-speech control 263, terminal device 120 can use a machine learning model or randomly select a target timbre from a timbre library. Then, terminal device 120 converts the third text into the corresponding third audio with the target timbre.
[0068] In summary, without user input, an audio clip can be generated based on an understanding of the content entity. Correspondingly, if the user cannot find background music that fits the theme, the terminal device can generate matching audio based on the content entity, thus providing a lightweight, fun, and personalized audio experience.
[0069] Figure 3 shows a flowchart of a content generation process 300 according to some embodiments of the present disclosure. Process 300 can be implemented at a terminal device 120. Process 300 is described below with reference to Figure 1.
[0070] In box 310, terminal device 120 responds to an audio editing request by presenting an audio editing panel, which includes at least audio generation controls.
[0071] In box 320, in response to detecting a trigger on the audio generation control, the terminal device 120 obtains first text and a first audio corresponding to the first text, wherein at least one of the first text and the first audio is determined based on the content entity to be edited.
[0072] In box 330, terminal device 120 adds first text and first audio to a content entity to obtain a first video, wherein the first text is overlaid on the content entity and the first audio is configured as at least a portion of the audio corresponding to the first video.
[0073] In some embodiments, the first text and / or the first audio are generated based on content entities using a machine learning model.
[0074] In some embodiments, the first audio is generated by performing text-to-speech on the first text based on a first timbre type.
[0075] In some embodiments, the first timbre type is determined by at least one of the following: determining a user-specified timbre type as the first timbre type, or randomly selecting a first timbre type from a timbre library.
[0076] In some embodiments, the text style of the first text and / or the first timbre type of the first audio are determined based on the content entity using a machine learning model.
[0077] In some embodiments, process 300 further includes: presenting an adjustment control associated with a first video, the adjustment control indicating adjustment of first text and first audio in target content; in response to detecting a trigger on the adjustment control, obtaining second text and second audio to replace the first text and first audio respectively, at least one of the second text and second audio being determined based on a content entity; and adding the second text and second audio to the content entity to obtain a second video, wherein the second text is overlaid on the content entity and the second audio is configured as at least a portion of the audio corresponding to the second video.
[0078] In some embodiments, obtaining the second text and the second audio includes: in response to detecting a trigger on an adjustment control, presenting a text style selection entry and / or a timbre selection entry; receiving a selection of a second text style via the text style selection entry, and / or receiving a selection of a second timbre type via the timbre selection entry; and obtaining the second text having the second text style and the second audio having the second timbre type for replacing the first text and the first audio, respectively.
[0079] In some embodiments, process 300 further includes: in response to a triggering of an editing operation on the first text, presenting a text editing box corresponding to the first text; receiving updated third text via the text editing box; overlaying the third text onto the content entity; and removing the first audio from the content entity.
[0080] In some embodiments, process 300 further includes: in response to detecting a text-to-speech request for the third text, obtaining a third audio corresponding to the third text; and adding the third audio to a content entity to obtain a third video.
[0081] In some embodiments, the visual style of the audio generation control is randomly selected from multiple candidate visual styles during one or more presentations.
[0082] In some embodiments, obtaining the first text includes: sampling a portion of the content from the content entity; extracting first semantic information corresponding to the portion of the content using a semantic model based on the portion of the content; and generating the first text using a machine learning model based on the first semantic information and prompt information.
[0083] Figure 4 shows a schematic structural block diagram of an apparatus 400 for content generation according to certain embodiments of the present disclosure. The apparatus 400 may be implemented as or included in the terminal device 120. The various modules / components in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0084] As shown in the figure, the device 400 includes a page rendering module 410 configured to render an audio editing panel in response to an audio editing request. The audio editing panel includes at least an audio generation control. The device 400 also includes a text and audio acquisition module 420 configured to acquire first text and a corresponding first audio in response to detecting a trigger on the audio generation control. At least one of the first text and the first audio is determined based on the content entity to be edited. The device 400 also includes a video acquisition module 430 configured to add the first text and the first audio to the content entity to obtain a first video, wherein the first text is overlaid on the content entity, and the first audio is configured to be at least a portion of the audio corresponding to the first video.
[0085] In some embodiments, the first text and / or the first audio are generated based on content entities using a machine learning model.
[0086] In some embodiments, the first audio is generated by performing text-to-speech on the first text based on a first timbre type.
[0087] In some embodiments, the text and audio acquisition module 420 further includes a timbre type determination module, configured to determine the timbre type specified by the user as the first timbre type, or to randomly select the first timbre type from the timbre library.
[0088] In some embodiments, the text style of the first text and / or the first timbre type of the first audio are determined based on the content entity using a machine learning model.
[0089] In some embodiments, the video acquisition module 410 is further configured to present an adjustment control associated with a first video, the adjustment control indicating adjustment of first text and first audio in target content; in response to detecting a trigger on the adjustment control, acquire second text and second audio to replace the first text and first audio respectively, at least one of the second text and second audio being determined based on a content entity; and add the second text and second audio to the content entity to obtain a second video, wherein the second text is overlaid on the content entity and the second audio is configured as at least a portion of the audio corresponding to the second video.
[0090] In some embodiments, the text and audio acquisition module 420 is further configured to, in response to detecting a trigger on the adjustment control, present a text style selection entry and / or a timbre selection entry; receive a selection of a second text style via the text style selection entry, and / or receive a selection of a second timbre type via the timbre selection entry; and acquire a second text with a second text style and a second audio with a second timbre type for replacing the first text and the first audio, respectively.
[0091] In some embodiments, the text and audio acquisition module 420 is further configured to, in response to an editing operation on the first text, present a text editing box corresponding to the first text; receive updated third text via the text editing box; overlay the third text onto the content entity; and remove the first audio from the content entity.
[0092] In some embodiments, the video acquisition module 430 is further configured to, in response to detecting a text-to-speech request for the third text, acquire the third audio corresponding to the third text; and add the third audio to the content entity to obtain the third video.
[0093] In some embodiments, the visual style of the audio generation control is randomly selected from multiple candidate visual styles during one or more presentations.
[0094] In some embodiments, the text and audio acquisition module 420 further includes a text generation module configured to sample partial content from the content entity; extract first semantic information corresponding to the partial content using a semantic model based on the partial content; and generate first text using a machine learning model based on the first semantic information and prompt information.
[0095] Figure 5 shows a block diagram illustrating an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 500 shown in Figure 5 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 500 shown in Figure 5 can be used to implement the terminal device 120 of Figure 1.
[0096] As shown in Figure 5, the electronic device 500 is in the form of a general-purpose electronic device. Components of the electronic device 500 may include, but are not limited to, one or more processing units or processors 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processor 510 may be a physical or virtual processor and is capable of performing various processes according to the programs stored in the memory 520. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 500.
[0097] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.
[0098] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0099] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0100] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0101] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0102] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0103] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0104] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0106] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for content generation, comprising: in response to an audio editing request, presenting an audio editing panel, the audio editing panel comprising at least an audio generation control; in response to detecting a trigger of the audio generation control, obtaining a first text and a first audio corresponding to the first text, at least one of the first text and the first audio being determined based on a content entity to be edited; and adding the first text and the first audio into the content entity to obtain a first video, wherein the first text is overlaid presented on the content entity and the first audio is configured as at least a portion of audio corresponding to the first video.
2. The method of claim 1, wherein the first text and / or the first audio are generated based on the content entity with a machine learning model.
3. The method of any one of claims 1-2, wherein the first audio is generated based on a first voice type performing text-to-speech on the first text, and wherein the first voice type is determined by at least one of: determining a user-specified voice type as the first voice type, or randomly selecting the first voice type from a voice library.
4. The method of claim 2, wherein a text style of the first text and / or the first voice type of the first audio are determined based on the content entity with the machine learning model.
5. The method of any one of claims 1-4, further comprising: presenting an adjustment control in association with the first video, the adjustment control indicating for adjusting the first text and the first audio in the target content; in response to detecting a trigger of the adjustment control, obtaining a second text and a second audio for replacing the first text and the first audio respectively, at least one of the second text and the second audio being determined based on the content entity; and adding the second text and the second audio into the content entity to obtain a second video, wherein the second text is overlaid presented on the content entity and the second audio is configured as at least a portion of audio corresponding to the second video.
6. The method of claim 5, wherein obtaining a second text and a second audio comprises: in response to detecting a trigger of the adjustment control, presenting a text style selection portal and / or a voice selection portal; receiving a selection of a second text style via the text style selection portal and / or a selection of a second voice type via the voice selection portal; and obtaining a second text with the second text style and the second audio with the second voice type for replacing the first text and the first audio respectively.
7. The method of any one of claims 1-6, further comprising: in response to a trigger of an editing operation on the first text, presenting a text editing box corresponding to the first text; receiving an updated third text via the text editing box; overlaying presenting the third text on the content entity; and removing the first audio from the content entity. 8. The method of claim 7, further comprising: in response to detecting a text-to-speech request for the third text, obtaining third audio corresponding to the third text; and adding the third audio to the content entity to obtain a third video.
9. The method of any one of claims 1-8, wherein in one or more presentations, a visual style of the audio generation control is randomly selected from a plurality of candidate visual styles.
10. The method of any one of claims 1-9, wherein obtaining first text comprises: sampling a portion of content from the content entity; based on the portion of content, extracting first semantic information corresponding to the portion of content using a semantic model; and based on the first semantic information and prompt information, generating the first text using a machine learning model.
11. An apparatus for content generation, comprising: an edit panel presentation module configured to present an audio edit panel in response to an audio edit request, the audio edit panel comprising at least an audio generation control; a text and audio obtaining module configured to obtain first text and first audio corresponding to the first text in response to detecting a trigger of the audio generation control, at least one of the first text and the first audio being determined based on a content entity to be edited; and a video obtaining module configured to add the first text and the first audio into the content entity to obtain a first video, wherein the first text is superimposed presented on the content entity and the first audio is configured as at least a portion of audio corresponding to the first video.
12. An electronic device, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method of any one of claims 1-10.
13. A computer-readable storage medium having computer-executable instructions stored therein, the computer-executable instructions being executable by a processor to implement the method of any one of claims 1-10.
14. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method of any one of claims 1-10.
Citation Information
Patent Citations
Video dubbing method and device, computer device and computer readable storage medium
CN110933330A
Method for generating video through AI intelligent image-text
CN115988149A
Method and system for converting video into descriptive audio
CN116647730A
Video generation method and device based on neural network
CN116708951A
Method, Apparatus and System For Regenerating Voice Intonation In Automatically Dubbed Videos
US20160021334A1