Audio and video generation method and device, equipment, storage medium and program product

By acquiring the target speech of the target object in the video generation interface and performing speech synthesis processing, the timbre of the target object is cloned, which solves the problem of low audio and video generation efficiency in the video editing process and achieves more efficient audio and video generation.

CN122053876APending Publication Date: 2026-05-15XIAOHONGSHU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAOHONGSHU TECH CO LTD
Filing Date
2026-03-27
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

How can we improve the efficiency of audio and video generation during the video editing process in existing technologies?

Method used

The video generation interface is displayed, which includes an audio synthesis access point to obtain the target speech of the target object. The target speech has a target timbre and is used to perform speech synthesis processing on the target text to obtain synthesized speech with the target timbre. Then, the target audio and video are displayed. The target audio and video are obtained by merging the first video and the synthesized speech.

Benefits of technology

By cloning the timbre of the target object, the efficiency of audio and video generation during the video editing process is improved, avoiding the blockage caused by switching the video generation interface to an independent timbre cloning entry, thus increasing the efficiency of audio and video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122053876A_ABST
    Figure CN122053876A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an audio and video generation method and device, equipment, a storage medium and a program product, and the method comprises the steps: displaying a video generation interface which comprises an audio synthesis access entry; wherein the video generation interface is used for generating a first video based on at least one media material; obtaining target voice of the target object through the audio synthesis access entry; wherein the target voice has a target tone, and the target voice is used for performing voice synthesis processing on the target text to obtain a synthesized voice with the target tone; and displaying a target audio / video, wherein the target audio / video is obtained by combining the first video and the synthesized voice. According to the embodiment of the invention, the audio and video generation efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an audio and video generation method, apparatus, device, storage medium, and program product. Background Technology

[0002] With the rapid development of artificial intelligence and computer vision technologies, video content creation has been widely applied across various fields, from social media platforms and film and television production to advertising creativity. Video content creation has become an important form of information dissemination, entertainment, and commercial promotion. Among these methods, video editing is the most common way to create video content. Users can upload video editing materials, and then, based on these materials and existing video editing templates, generate and display multiple different videos for users to choose from. However, there is a need to add audio during the video editing process. Therefore, improving the efficiency of audio and video generation is a pressing technical problem that needs to be solved. Summary of the Invention

[0003] The technical problem to be solved by the embodiments of this application is to provide an audio and video generation method, apparatus, device, storage medium and program product that can improve the efficiency of audio and video generation.

[0004] On one hand, embodiments of this application provide an audio / video generation method, which includes: The video generation interface is displayed, and the video generation interface includes an audio synthesis access point; wherein, the video generation interface is used to generate a first video based on at least one media material; The target speech of the target object is obtained through the audio synthesis access portal; wherein, the target speech has a target timbre, and the target speech is used to perform speech synthesis processing on the target text to obtain synthesized speech with the target timbre; The target audio and video are displayed, which are obtained by merging the first video and the synthesized speech.

[0005] On the other hand, embodiments of this application provide an audio / video generation apparatus, which includes: A display unit is used to display a video generation interface, which includes an audio synthesis access point; wherein, the video generation interface is used to generate a first video based on at least one media material; The acquisition unit is used to acquire the target speech of the target object via the audio synthesis access portal; wherein the target speech has a target timbre, and the target speech is used to perform speech synthesis processing on the target text to obtain synthesized speech with the target timbre; The display unit is also used to display target audio and video, which is obtained by merging the first video and the synthesized speech.

[0006] In one implementation, the acquisition unit acquires the target speech of the target object via the audio synthesis access point, including: The target speech is acquired via the audio synthesis access point; or... At least one candidate speech of the target object is displayed via the audio synthesis access portal; in response to a selection operation of the at least one candidate speech, the target speech is determined; or... The target speech is determined based on the at least one media material via the audio synthesis access point.

[0007] In one embodiment, the acquisition unit acquires the target speech via the audio synthesis access point, including: The reference text and voice acquisition entry are displayed via the audio synthesis access point; The target speech output by the target object in response to the reference text is acquired via the speech acquisition port.

[0008] In one embodiment, the voice acquisition port includes a voice input control; the acquisition unit acquires the target voice output by the target object in response to the reference text via the voice acquisition port, including: In response to receiving and continuously receiving a voice input operation, the system acquires user-inputted voice during the duration of the voice input operation and displays a first voice operation area, wherein the voice input operation includes performing a first operation on the voice input control; in response to a voice input lock operation, the system continues to acquire user-inputted voice, wherein the voice input lock operation includes performing a second operation and stopping the first operation during the duration of the voice input operation; in response to a voice input pause operation, the system pauses receiving user-inputted voice and displays a second voice operation area instead of the first voice operation area; the second voice operation area is used to transmit the acquired target voice; or, In response to receiving and continuously receiving a voice input operation, the system acquires user-inputted voice during the duration of the voice input operation, the voice input operation including performing a first operation on the voice input control; in response to a voice input lock operation, the system continues to acquire user-inputted voice, the voice input lock operation including receiving a second operation and stopping the first operation during the duration of the voice input operation; in response to a voice transmission trigger event, the system transmits the acquired target voice; or, In response to receiving and continuously receiving a voice input operation, the system acquires the user's voice input during the duration of the voice input operation; in response to a voice transmission trigger event, the system transmits the acquired target voice.

[0009] In one embodiment, the display unit is further configured to display a voice transmission prompt message and a cancel transmission control after the acquisition unit acquires the target voice through the audio synthesis access point; the voice transmission prompt message is used to indicate that the target voice is being transmitted. The audio and video generation device further includes a processing unit, which is used to cancel the transmission of the target voice in response to a trigger operation of the cancel send control.

[0010] In one embodiment, the display unit is further configured to display speech synthesis prompt information and a cancel synthesis control after the acquisition unit acquires the target speech of the target object via the audio synthesis access entry; the speech synthesis prompt information is used to indicate that the target speech is being processed by speech synthesis. The audio and video generation device further includes a processing unit, which is used to cancel the speech synthesis processing of the target speech in response to a trigger operation of the cancel synthesis control.

[0011] In one embodiment, the display unit is further configured to display at least one synthesized speech having the target timbre; The audio and video generation device further includes a playback unit, which is used to play the target synthesized speech in response to a playback operation on the target synthesized speech, wherein the target synthesized speech is any one of the at least one synthesized speech; In response to the selection operation of the target synthesized speech, the display unit is triggered to display the target audio and video, which is obtained by merging the first video and the target synthesized speech.

[0012] On the other hand, embodiments of this application provide a computer device, which includes a memory, a communication interface, and a processor, wherein the memory, the communication interface, and the processor are interconnected; the memory stores a computer program, and the processor calls the computer program stored in the memory to implement the above-mentioned audio and video generation method.

[0013] On the other hand, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described audio and video generation method.

[0014] On the other hand, embodiments of this application provide a computer program product, which includes a computer program stored in a computer storage medium; the processor of a computer device reads the computer program from the computer storage medium and executes the computer program, causing the computer device to perform the above-described audio and video generation method.

[0015] In this embodiment, a video generation interface can be displayed. The video generation interface includes an audio synthesis access point. The video generation interface is used to generate a first video based on at least one media material. Target speech of a target object is obtained through the audio synthesis access point. The target speech has a target timbre and is used to perform speech synthesis processing on the target text to obtain synthesized speech with the target timbre. Then, the target audio and video are displayed. The target audio and video are obtained by merging the first video and the synthesized speech. Since the synthesized speech is obtained by cloning the timbre of the target object, this embodiment can provide more possibilities for video editing, allowing users to directly use cloned timbres during video editing. This avoids switching from the video generation interface to a separate timbre cloning entry point, which would interrupt video editing. Therefore, this embodiment can clone the timbre of the target object during video editing, improving the efficiency of audio and video generation. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the background art, the accompanying drawings used in the embodiments of this application or the background art will be described below.

[0017] Figure 1 This is a schematic diagram of the structure of an audio / video generation system provided in an embodiment of this application; Figure 2 This is a flowchart illustrating an audio / video generation method provided in an embodiment of this application; Figure 3A This is a schematic diagram of the video generation interface provided in an embodiment of this application; Figures 3B to 3D This is a schematic diagram of the material selection interface provided in an embodiment of this application; Figure 3E This is a schematic diagram illustrating the generation of prompt information provided in an embodiment of this application; Figure 3F This is a schematic diagram of the instruction input entry provided in an embodiment of this application; Figures 3G to 3I This is a schematic diagram of the first video browsing interface provided in an embodiment of this application; Figure 3J This is a schematic diagram of the video generation interface provided in an embodiment of this application; Figure 3K This is a schematic diagram of the voice input interface provided in an embodiment of this application; Figure 3L This is a schematic diagram of the video generation interface provided in an embodiment of this application; Figure 3M as well as Figure 3N This is a schematic diagram of the voice input interface provided in an embodiment of this application; Figure 3O This is a schematic diagram of the audio browsing interface provided in an embodiment of this application; Figures 3P to 3S This is a schematic diagram of the voice input interface provided in an embodiment of this application; Figure 3T This is a schematic diagram of the audio browsing interface provided in an embodiment of this application; Figure 3U as well as Figure 3V This is a schematic diagram of the content editing interface provided in an embodiment of this application; Figure 4 This is a flowchart illustrating another video generation method provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a video generation device provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0019] Furthermore, in the description of the embodiments of this application, the terms "first," "second," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0020] The descriptions involved in the embodiments of this application will be defined below: Published content: This refers to content information pre-published by users on the online platform. It can include text, images, audio, video, and other audio-visual materials, and can be displayed in formats such as notes, articles, and video files. The specific content information included in published content can be adjusted accordingly for different scenarios, and the content information in the published content can be determined by the publisher.

[0021] "In response to" indicates the state in which a corresponding event occurs or a condition is met. The timing of subsequent actions performed in response to this event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately when the event occurs or the condition is met; while in other cases, subsequent actions may be performed some time after the event occurs or the condition is met.

[0022] Triggered actions refer to actions or responses that are automatically executed when specific conditions or events occur. Triggered actions for controls mainly refer to events and behaviors triggered after a user interacts with UI controls (such as buttons, text boxes, drop-down menus, etc.). These triggered actions can include the following types: 1. Click event: The user clicks a button, link, or icon.

[0023] 2. Mouse Over / Mouse Enter: The mouse pointer moves over the control and hovers.

[0024] 3. Text Input Event (Input): The user enters or modifies text in the text box.

[0025] 4. Change Event: The user changes the options in the drop-down menu or selection box.

[0026] 5. Focus events (Focus / Blur): Controls gaining or losing input focus, etc.

[0027] In specific embodiments of this application, user-related data is involved, such as target voice, first video, target audio and video, etc. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of related data must comply with local laws, regulations, and standards. For example, this application can display prompt information before and during the collection of user-related data to inform the user that their related data is being collected. This ensures that the application only begins executing the steps for acquiring user-related data after receiving confirmation from the user regarding the prompt information; otherwise (i.e., without receiving confirmation from the user regarding the prompt information), the steps for acquiring user-related data end, meaning that user-related data is not acquired.

[0028] The audio and video generation method in this embodiment is applied to an audio and video generation device, which is located on a client. The client can run on a computer device, which has one or more processors, a memory, and one or more applications. The one or more applications are stored in the memory and configured to be executed by the processor to implement the video generation method. The computer device can be a smart terminal, such as a mobile phone, wearable device, in-vehicle device, tablet computer, network device, and smart computer.

[0029] This application provides an audio and video generation system, exemplarily, such as... Figure 1 As shown, Figure 1 This is a schematic diagram of an audio / video generation system provided in an embodiment of this application. The system may include at least one client 100 (each client 100 integrating a video generation device) and a server 200. Each client 100 runs a computer-readable storage medium corresponding to an audio / video generation method to execute the method. The server 200 communicates and interacts with at least one client 100.

[0030] Any client 100 can display a video generation interface, which includes an audio synthesis access point. The server 200 can generate a first video based on at least one media material determined by the client 100 through the video generation interface. The client 100 can obtain the target speech of the target object via the audio synthesis access point. The target speech has a target timbre, and the client 100 sends the target speech to the server 200. The server 200 can perform speech synthesis processing on the target text based on the target speech to obtain synthesized speech with the target timbre. The server 200 can also merge the first video and the synthesized speech to obtain the target audio-visual content. The server 200 sends the target audio-visual content to the client 100, and the client 100 displays the target audio-visual content.

[0031] The server 200 in the audio and video generation system can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0032] Any client 100 and server 200 can communicate directly or indirectly through wired or wireless communication, and this application does not impose any restrictions.

[0033] Understandable, Figure 1The servers or at least one client in the audio and video generation system shown do not constitute a limitation on the embodiments of the present invention. That is, the number of servers, the type of servers, the number of at least one client, the type of at least one client, or the number and type of devices contained in each server or client do not affect the overall implementation of the technical solution in the embodiments of the present invention, and can all be considered as equivalent substitutions or derivatives of the technical solutions claimed in the embodiments of the present invention.

[0034] Based on the above description, please refer to Figure 2 , Figure 2 This is a flowchart illustrating an audio / video generation method provided in an embodiment of this application. This audio / video generation method can be applied in a client application, such as... Figure 2 The audio and video generation method shown includes, but is not limited to, steps S201 to S203, wherein: S201, Display the video generation interface, which includes an audio synthesis access point. The video generation interface is used to generate a first video based on at least one media material.

[0035] The target object can log in to the client using the target account and then perform steps S201-S203 through the client. The audio synthesis access entry is used to clone the target object's target timbre to obtain the synthesized speech corresponding to the first video. The synthesized speech has the target timbre. Then, the first video and the synthesized speech are merged to obtain the target audio-visual product. Compared to manually recording the speech corresponding to each first video generated by the target object, this embodiment can improve the speech generation efficiency, thereby improving the audio-visual product generation efficiency.

[0036] by Figure 3A Taking the video generation interface shown as an example, when the client is editing a video, it can display the video generation interface 301. The video generation interface 301 may include an audio synthesis access point 301a. When the target object wants to generate the audio corresponding to the first video, it can trigger the audio-video synthesis access point 301a. The client can obtain the target object's target audio through the audio synthesis access point 301a. The target audio has a target timbre and is used to perform speech synthesis processing on the target text to obtain synthesized audio with the target timbre. The synthesized audio can be as shown in "Timbre Version 1" in the audio browsing interface 302. Optionally, if the first video has been successfully generated based on at least one media material during the process of generating synthesized audio, the first video can also be displayed in the audio browsing interface 302. Optionally, the audio browsing interface can be the first video browsing interface described below.

[0037] Optionally, the audio synthesis access point may include an audio synthesis access control, as shown in the audio synthesis access point in the video generation interface 301. Alternatively, the audio synthesis access point may be a voice command control entry point, whereby the target object can input a voice command. If the voice command and the voice command control entry point match, the client can obtain the target object's target voice through the audio synthesis access point. For example, assuming the voice command control entry point is configured as the "voice synthesis" voice command, when the target object inputs the "voice synthesis" voice command through the voice input device (e.g., microphone) of the computer device running the client, the client can determine that the voice command input by the target object matches the voice command control entry point, and thus obtain the target object's target voice through the audio synthesis access point. If the target object inputs the "open voice" voice command through the voice input device of the computer device running the client, the client can determine that the voice command input by the target object does not match the voice command control entry point, and thus refuse to obtain the target object's target voice. Optionally, to guide the target user to accurately input a voice command that matches the voice command control entry, the client can display voice command prompts when showing the video generation interface. These prompts guide the user to access the audio synthesis access entry using a voice command that matches the voice command control entry. Alternatively, the audio synthesis access entry can be a gesture command control entry. The target user can input a gesture. If the gesture input by the target user matches the gesture command control entry, the client can obtain the target user's target speech via the audio synthesis access entry. For example, assuming the gesture command control entry is configured as a "five-finger fist" gesture, when the target user inputs a "five-finger fist" gesture using the gesture input device (e.g., camera, infrared sensor) of the computer device running the client, the client can determine that the gesture input matches the gesture command control entry and thus obtain the target user's target speech via the audio synthesis access entry. If the target user inputs a "five-finger spread" gesture using the gesture input device of the computer device running the client, the client can determine that the gesture input does not match the gesture command control entry and thus refuse to obtain the target user's target speech. Further optionally, in order to guide the target user to accurately input a gesture that matches the gesture command control entry, the client can display gesture prompts when displaying the video generation interface. The gesture prompts are used to prompt the user to access the audio synthesis access entry by using a gesture that matches the gesture command control entry.

[0038] Optionally, the video generation interface 301 may also include at least one inference information, which is used to demonstrate the inference process of generating a first video based on at least one media material. The generation method of the first video may include: determining the inference process of generating the first video based on at least one media material, and performing video generation processing on at least one media material based on the inference process. The inference process includes at least one of the following inference information: highlight segments in at least one media material, a description of the content of at least one media material, the video theme of the first video, a description of the video content of the first video, and a summary of the video content of the first video. The inference information is determined based on at least one media material.

[0039] 1. At least one highlight segment in the media footage.

[0040] Highlight clips, also known as highlight video clips, are the most important and exciting segments in a video production. They attract the viewer's attention, leave a lasting impression, and even evoke emotional resonance. Highlight clips are typically generated by connecting at least one shot according to specific editing logic and techniques. Based on their placement within the final video (such as the first video), highlight clips can be further divided into opening highlight clips, ending highlight clips, and middle highlight clips. Opening highlight clips are usually located at the beginning of the final video, ending highlight clips at the end, and middle highlight clips in the middle. For example, using a series of rapidly changing shots at the beginning of a video to present an overview of the entire travel process constitutes an opening highlight clip. Similarly, using a series of shots with a concluding meaning (such as a silhouette or closing credits) at the end of a video constitutes an ending highlight clip.

[0041] Optionally, highlight recognition can be performed on at least one media clip to determine highlight segments of at least one media clip. Optionally, highlight recognition of at least one media clip to determine at least one highlight segment of at least one media clip can be performed by: sampling at least one media clip to obtain multiple sample frames, determining the attribute information of each sample frame, and determining the highlight segment corresponding to at least one media clip based on the attribute information of at least one media clip, the multiple sample frames, a target highlight template, and the highlight sample frames corresponding to multiple highlight shots. The attribute information of the sample frames includes a semantic feature vector, which indicates the image semantic information of the sample frame.

[0042] A highlight clip can include at least one of an image, a video cover, and text. An image refers to a single frame of an image; a video cover refers to a video clip, which can be clicked to play the corresponding video; and text refers to descriptive information about any highlight clip, describing its content.

[0043] Optionally, semantic understanding can be performed on the highlight segments to obtain segment descriptions. These segment descriptions are then input into a target large-scale language model, allowing the model to determine at least one of the following: the video theme, the video content description, or the video content summary of the first video. The target large-scale language model can be trained under supervised supervision.

[0044] Optionally, semantic understanding can be performed on the highlight clips to obtain a fragment description of the highlight clips, and frames can be extracted from at least one media material to obtain at least one image frame. Semantic understanding can be performed on the at least one image frame to obtain a text description of the at least one image frame. The fragment description of the highlight clips and the text description of the at least one image frame are input into the target large language model, so that the target large language model can be based on at least one of the following: the video theme of the first video, the video content description of the first video, and the video content summary of the first video.

[0045] 2. A description of the content of at least one media material.

[0046] Among them, the content description of at least one media material is used to describe the content of at least one media material.

[0047] Optionally, the content description of at least one media material may include at least one of the following: character information, scene information, and event information of the media material, wherein the character information, scene information, and event information are determined based on at least one media material. Specifically, the character information describes the characters in at least one media material; the character information describes the scenes in at least one media material; and the event information describes the events that occur in at least one media material. For example, the character information of the media material may be: "The core subject is a small white dog. In addition, there are several other dogs of different breeds, and the owner of the small white dog"; the scene information of the media material may be: "Including an outdoor grassy area, sunny, where the dog runs, plays, and interacts with other dogs; inside a car, the dog observes and looks around; indoors, the dog is sitting relaxed on a sofa"; and the event information of the media material may be: "The core is a scene of the owner taking the small white dog outdoors, and there are scenes of interaction and play with other pet dogs. Furthermore, inside the car and indoors, the dog has warm interactions with its owner, such as being cared for in the car, being leashed outdoors, and being accompanied indoors." III. The theme of the first video.

[0048] The video theme can refer to the content represented by the video, that is, what the video is specifically about. For example, if the video is specifically about travel and scenery, then the theme of the first generated video can be determined as travel and scenery. The video theme can be determined based on at least one media material.

[0049] Optionally, the video theme is determined based on at least one of character information, scene information, and event information. Character information is used to describe the characters in at least one media clip; character information is used to describe the scenes in at least one media clip; and event information is used to describe the events that occur in at least one media clip. The character information, scene information, and event information can be input into a target large language model to determine the video theme.

[0050] Optionally, the video theme can be determined based on at least one highlight segment of media material. Specifically, semantic understanding can be performed on the highlight segment to obtain a segment description (including at least one of character information, scene information, and event information), and the segment description of the highlight segment can be input into a target large language model to determine the video theme.

[0051] Optionally, the video theme can be determined based on at least one media clip and at least one highlight segment of the media clip. Specifically, semantic understanding can be performed on at least one media clip to obtain a text description of the media clip, semantic understanding can be performed on the highlight segment to obtain a segment description of the highlight segment, and the text description of at least one media clip and the segment description of the highlight segment can be input into a target large language model to determine the video theme.

[0052] IV. Description of the video content of the first video.

[0053] The video content description is used to describe the video content of the first video.

[0054] Optionally, the video content description includes at least one of the following: story title or theme, narrative order, and content description; the narrative order is used to describe the order of the video content in the first video.

[0055] The story title or theme could be something like "An Unforgettable Trip"; the narrative order describes the sequence or storyline of the video content in the first video. A storyline is a linear arrangement and narration of the story's direction, development, and ending. It encompasses a series of logically related events that drive the story forward until the final conclusion. It can be a single story; for example, the storyline could be "climbing a mountain," then "seeing flowers," and finally "going home." The content description describes the video content of the first video; for example, the content description could be "First, I climbed a mountain, saw various flowers in the mountain, and then returned home happily."

[0056] Optional, such as Figure 3A As shown, the video generation interface 301 may also include a media management entry 301b, which is used to manage the media materials used to generate the first video. For example, the client can update at least one media material used to generate the first video via the media management entry 301b, and then display the updated at least one inference information, which is used to demonstrate the inference process of generating the first video based on the updated at least one media material.

[0057] Optionally, semantic understanding can be performed on at least one media material to obtain a text description of at least one media material. The text description of at least one media material can be input into a target large language model, so that the target large language model can determine the video content description of the first video based on the input text description.

[0058] Optionally, semantic understanding can be performed on the highlight segments to obtain segment descriptions of the highlight segments. These segment descriptions can then be input into the target large language model, allowing the target large language model to determine the video content description of the first video based on the input text descriptions.

[0059] Alternatively, a text description of at least one media material and a fragment description of a highlight segment can be input into the target large language model, so that the target large language model can determine the video content description of the first video based on the input text description and the fragment description of the highlight segment.

[0060] Alternatively, the text description, the segment description of the highlight clip, and the video theme can be input into the target large language model, so that the target large language model can determine the video content description of the first video based on the input text description and video theme.

[0061] Optionally, there can be multiple first videos, and therefore multiple video content descriptions, with one video content description corresponding to one first video. For example, if there are four first videos, there will also be four video content descriptions, with one video content description corresponding to one first video. Each video content description includes at least one of the following: a corresponding story title or theme, narrative order, and content description. Optionally, a preset number of first videos can be pre-configured, allowing the client to generate a preset number of first videos based on at least one media asset.

[0062] V. Summary of the video content of the first video.

[0063] The video content summary is used to summarize the video content of the first generated video.

[0064] Optionally, the video content summary includes at least one of the following from the first video: video title, video script, video background music, video voice-over, and video style.

[0065] For example, the content of the video on First Video can be summarized as "An Unforgettable Trip," featuring everyday narrative text, relaxing music, and a minimalist, everyday style. The video title is "An Unforgettable Trip"; the text is described as "everyday narrative text," which can be understood as First Video's typical video text type; the music is described as "relaxing music," which can be understood as First Video's typical music type; the style is described as "minimalist, everyday style," which can be understood as First Video's typical video style type. The voiceover can be understood as First Video's typical voiceover type, such as "funny voice."

[0066] Optionally, there can be multiple first videos, and therefore multiple video content summaries, with one first video corresponding to one video content summary. For example, if there are four first videos, there will also be four video content summaries, with one video content summary corresponding to one first video. Each video content summary summarizes the video content of the corresponding first video. Optionally, a preset number of first videos can be pre-configured, allowing the client to generate a preset number of first videos based on at least one media asset.

[0067] Optionally, in response to a trigger action on the video update control, at least one candidate video style option can be displayed; and in response to a trigger action on the target video style option, a new first video can be displayed, where the target video style option is any one of the at least one candidate video style option. Figure 3AFor example, the client can respond to a trigger operation on the video update control 302a in the audio browsing interface 302 to display at least one candidate video style option, and the client can respond to a trigger operation on the target video style option to display a new first video.

[0068] Optionally, the audio browsing interface may also include a second editing control, which can respond to a trigger operation on the second editing control to display a video editing interface, which includes a first video, and can respond to an editing operation on the first video to obtain the edited first video.

[0069] The video editing interface may include at least one editing control, which can respond to a trigger operation on any editing control to perform corresponding editing operations on the first video, resulting in an edited first video. The editing controls include, but are not limited to, at least one of the following: special effects controls, text controls, clipping controls, filter controls, background controls, and audio selection controls. Specifically, special effects controls are used to add special effects to the first video, such as "star surround effect" or "meteor shower effect"; text controls are used to add text content to the first video, such as adding the text "fragments of life"; clipping controls are used to edit the first video; for example, if the first video is 20 seconds long, the target can click the clipping control to edit the first 10 seconds of the 20-second video into the edited first video; filter controls are used to add filter effects to the first video, such as "whitening" or "texture" filters; background controls are used to adjust the background color of the first video, such as changing the background color to "yellow" or "blue"; and audio selection controls are used to set the target audio for the first video. Figure 3A For example, the audio browsing interface 302 may also include a second editing control 302b, and the client may display a video editing interface in response to a trigger operation on the second editing control 302b.

[0070] Optionally, published content can also be generated based on the target audio / video. The published content includes the target audio / video, a title, and the main body of the content. The title and main body are derived from the theme of the target audio / video. Figure 3A For example, the audio browsing interface 302 may also include a third confirmation control 302c, namely "Next". The target object can click the third confirmation control 302c, so that the client can respond to the trigger operation of the third confirmation control 302c, that is, respond to the confirmation operation of the target audio and video, and display the publishing content editing interface. The publishing content editing interface includes the target audio and video, the title of the publishing content, and the body of the publishing content. The target object can edit the content to be published in the publishing content editing interface and publish the content to be published after the editing is completed.

[0071] In one optional implementation, the first video is displayed in the audio browsing interface, and the first video is displayed in full-screen mode in the audio browsing interface, that is, the display area of ​​the first video is equal to the display area of ​​the audio browsing interface. In this embodiment, by displaying the first video in full-screen mode, interference from other elements can be avoided, thereby improving the viewing experience of the first video and enhancing the user experience.

[0072] Optionally, the audio browsing interface may also include a video cover for the first video. The video cover for the first video may include any one of the following: the first frame, the last frame, any frame of the first video, or a configuration image corresponding to the first video. The configuration image corresponding to the first video may refer to any image that the target object can configure as the video cover for the first video.

[0073] In one alternative implementation, there are multiple first videos, and multiple first videos are displayed; in response to a selection operation on the multiple first videos, the first video selected by the selection operation is displayed.

[0074] When there are multiple first videos, the audio browsing interface can display multiple first videos. The target user can switch between them by swiping, using voice commands, or using a switching control. The selection operation for multiple first videos can include triggering an action on any one of the multiple first videos.

[0075] In one alternative implementation, if there is at least one first video, then the audio browsing interface can display the video cover of at least one first video, as well as the first video corresponding to the preset video cover in the at least one video cover.

[0076] The video cover of any first video may include any one of the following: the first frame, the last frame, any frame of the first video, and a configuration image corresponding to the first video. The configuration image corresponding to any first video may refer to any image that the target object can configure as the video cover for the first video. Preset video covers may include video covers at a preset display position or video covers with a preset cover style. Specifically, at least one video cover can be displayed in a sequence, for example, video cover 1 is displayed before video covers 2 and 3, and video cover 2 is displayed before video cover 3. If the preset display position is the first display position, then video cover 1 is the preset video cover; if the preset display position is the middle display position, then video cover 2 is the preset video cover; if the preset display position is the last display position, then video cover 3 is the preset video cover. At least one video cover can include at least one cover style. For example, video cover 1 is cartoon style, video cover 2 is artistic style, and video cover 3 is retro style. If the preset cover style is cartoon style, then video cover 1 is the preset video cover; if the preset cover style is artistic style, then video cover 2 is the preset video cover; if the preset cover style is retro style, then video cover 3 is the preset video cover. Optionally, the cover style of any video cover can match the video style of the first video corresponding to that video cover. For example, if the video style of the first video corresponding to video cover 1 is cartoon style, then video cover 1 can be cartoon style; similarly, if the video style of the first video corresponding to video cover 2 is artistic style, then video cover 2 can be artistic style.

[0077] In one alternative implementation, the first video corresponding to any one of the at least one video covers may also be displayed in response to a trigger operation on any one of the video covers.

[0078] In one alternative implementation, the first video and its video tag are displayed via an audio browsing interface.

[0079] The video tags may include at least one of the following: video theme, video title, text style, visual style, and dubbing type.

[0080] The video theme, video title, text style, visual style, and voice-over type of any first video are determined based on the content of at least one media asset. The video theme or title can refer to the content the video represents, i.e., what the video is specifically about. For example, if the content of at least one media asset is "summer trip and enjoying different scenery," then the video theme or title of the generated first video can be determined as "travel scenery." The text style can refer to the style of the text in the video, such as a minimalist style. The visual style refers to the visual style of the video, such as a simple style. The voice-over type refers to the type of music in the video, such as an upbeat style.

[0081] by Figure 3A For example, the audio browsing interface 302 also includes the video title of the first video, namely "Cute Pets", and the visual style, namely "Simple Style".

[0082] Optionally, the video tag for the first video can be highlighted or enhanced. Highlighting or enhancing the tag can include at least one of the following: bolding, italicizing, displaying with a different text color, displaying with a different font, or displaying with a different font size.

[0083] In one optional implementation, the client can display a media selection interface, which includes one or more media materials. The target object can select at least one media material, thereby allowing the client to acquire at least one media material. The media selection interface includes one or more media materials, which may include at least one of the following: video footage, image footage, or live photos. Figure 3B For example, Figure 3B This is a schematic diagram of the media selection interface provided in the embodiments of this application, wherein the media selection interface 303 can be displayed, and the media selection interface 303 includes multiple media media 303a.

[0084] Optionally, a content recommendation page can be displayed, which includes at least one published piece of content and a first material selection entry. The material selection interface can be displayed in response to a trigger operation on the first material selection entry.

[0085] by Figure 3C For example, Figure 3CThis is a schematic diagram of the material selection interface provided in the embodiments of this application. The client can display a content recommendation page 306, which includes at least one published content and a first material selection entry 306a. The target object can browse or interact with at least one published content on the content recommendation page 306, and the client can display the material selection interface 307 in response to the trigger operation on the first material selection entry 306a.

[0086] Optionally, a trending content page can be displayed. This page includes trending content and a primary material selection entry point. It can respond to triggering actions on the primary material selection entry point by displaying a material selection interface. Trending content can be determined based on the popularity of published content. The popularity of any published content can be determined based on its interaction volume, which is determined by the interaction actions taken by various users on that published content. These interaction actions could include liking, commenting, saving, and sharing.

[0087] Optionally, a message list page can be displayed, which includes conversation messages with other objects and a first material selection entry. The material selection interface can be displayed in response to a trigger operation on the first material selection entry.

[0088] Optionally, a personal page may be displayed, which includes object information of the target object for providing at least one media material, a first material selection entry and / or a second material selection entry, and may display a material selection interface in response to a trigger operation on the first material selection entry or the second material selection entry.

[0089] by Figure 3D For example, Figure 3D This is a schematic diagram of the material selection interface provided in the embodiments of this application. The client can display a personal page 308. The personal page 308 includes object information of the target object for providing at least one media material, a first material selection entry 308a, and a second material selection entry 308b. The client can display the material selection interface 309 in response to a trigger operation on the first material selection entry 308a or the second material selection entry 308b.

[0090] Optionally, one or more media materials can be one or more media materials stored in the computer device where the client is located. In other words, one or more media materials in the material selection interface are one or more media materials stored in the computer device.

[0091] Optionally, the media selection interface may also include at least one media type control, which can respond to a trigger operation for any media type and display at least one media material corresponding to that media type on the media selection interface.

[0092] In this embodiment, the target object can filter the media types of one or more media materials displayed on the media selection interface. Each media type control corresponds to a media type (video material, image material, live image material). The target object can click on any media type control, such as clicking the "video" control, so that the client can display one or more video materials on the media selection interface.

[0093] In this embodiment of the application, by displaying at least one media type control on the media selection interface, users can quickly filter media media of different media types, thereby helping users to more conveniently select the media media of the desired media type, improving the efficiency of users in selecting media media, and enhancing the user interaction experience.

[0094] Optionally, at least one media material can be acquired in response to a selection operation for at least one media material among one or more media materials, and in response to a confirmation operation for at least one media material.

[0095] Optionally, the selection operation for at least one media asset may include a trigger operation for at least one media asset.

[0096] Optionally, the associated positions of each media material may include selection buttons. The associated position of any media material may include the display position of any media material or the adjacent position of any media material, etc., and this application embodiment does not limit this. That is to say, the associated position of at least one media material may include a selection button corresponding to at least one media material. The target object can click the selection button of the associated position of at least one media material, thereby triggering the selection operation of at least one media material among one or more media materials. Optionally, the material selection interface may also include a second video generation control. After the target object selects at least one media material, it can click the second video generation control, and then the client can respond to the triggering operation of the second video generation control to obtain at least one media material.

[0097] by Figure 3BFor example, in the media selection interface 303, the associated position of media material 303b among multiple media materials can include a selection button 303c. The target object can click the selection button 303c to select media material 303b. The selection button 303c can include the selection order of the corresponding media materials. For example, if the selection button 303c includes "1", it means that the selection order of media material 303b corresponding to the selection button 303c is the first. Similarly, in the media selection interface 303, the associated position of media material 303d among multiple media materials can include a selection button 303e. The target object can click the selection button 303e to select media material 303d. The selection button 303e can include the selection order of the corresponding media materials. For example, if the selection button 303e includes "2", it means that the selection order of media material 303d corresponding to the selection button 303e is the second. The media selection interface 303 also includes a second video generation control 303f. The target object can click the second video generation control 303f, thereby allowing the client to obtain at least one media material.

[0098] Optionally, in response to the second video generation operation, a second video browsing interface is displayed. The second video browsing interface includes a video cover of at least one second video generated based on at least one media asset, and a first video generation control. The first video generation control is used to trigger the first video generation operation.

[0099] by Figure 3B For example, the client can respond to the trigger operation of the second video generation control 303f to obtain at least one media material, the server can generate at least one second video and at least one video cover of the second video based on the at least one media material, and the client can display the video cover 304a of at least one second video and the first video generation control 304b on the second video browsing interface 304 based on at least one media material.

[0100] The second video is generated differently from the first video; that is, the generation method of the second video does not include the reasoning process for determining the second video. The second video can be generated based on at least one media material and a preset video generation template.

[0101] Optionally, in response to the second video generation operation, at least one video cover of a second video, the second video corresponding to a preset video cover in the at least one video cover, and a first video generation control are displayed. Optionally, responding to the second video generation operation may include responding to a trigger operation on the second video generation control.

[0102] by Figure 3BFor example, the client can respond to at least one media material and to a trigger operation on the second video generation control 303f in the material selection interface 303, displaying at least one second video cover 304a, the second video 304c corresponding to the preset video cover in the at least one video cover, and the first video generation control 304b in the second video browsing interface 304.

[0103] Optionally, responding to the first video generation operation may include responding to a trigger operation on the first video generation control; or a preset voice operation; or a preset gesture operation, etc.

[0104] by Figure 3B For example, the second video browsing interface 304 includes a first video generation control 304b. The client can respond to the trigger operation of the first video generation control 304b and display the first video 305a on the first video browsing interface 305.

[0105] Optionally, the first video and its video tag can be displayed via the first video browsing interface.

[0106] The video tags may include at least one of the following: video theme, video title, text style, visual style, and dubbing type.

[0107] The video theme, video title, text style, visual style, and voice-over type of any first video are determined based on the content of at least one media asset. The video theme or title can refer to the content the video represents, i.e., what the video is specifically about. For example, if the content of at least one media asset is "summer trip and enjoying different scenery," then the video theme or title of the generated first video can be determined as "travel scenery." The text style can refer to the style of the text in the video, such as a minimalist style. The visual style refers to the visual style of the video, such as a simple style. The voice-over type refers to the type of music in the video, such as an upbeat style.

[0108] by Figure 3B For example, the first video browsing interface 305 also includes the video title 305b of the first video, namely "An Unforgettable Trip", and the visual style 305c, namely "Simple Style".

[0109] Optionally, the second video browsing interface may also include a first confirmation control. If the target object wants to generate content to be published based on any of the second videos, that is, if the target object is satisfied with any of the second videos, then the target object does not need to trigger the first video generation control, but can directly click the first confirmation control to... Figure 3BFor example, the second video browsing interface 304 may also include a first confirmation control 304d. The client can respond to the trigger operation of the first confirmation control 304d to display the content editing interface. The content editing interface includes any of the second videos. The target object can edit the content to be published in the content editing interface and publish the content to be published after editing. The content to be published includes any of the second videos.

[0110] Optionally, after at least one media material is selected, the selection button is in the selected state. The selection button in the selected state includes the selection order of the corresponding media material. If the target object wants to cancel the selection of any media material, the target object can click the selection button of the associated position of any media material again, and then cancel the selection of any media material. At this time, the selection button is in the unselected state.

[0111] Optionally, if any media material is a video, then the associated position of that media material can display its video duration. The associated position of any media material can include its display position or adjacent positions, etc., and this embodiment does not limit this. For example, if "0:13" is displayed adjacent to any media material, it indicates that the video duration of that media material is 13 seconds.

[0112] Optionally, in response to a selection operation on at least one media material among one or more media materials, a thumbnail media material of at least one media material may be displayed; in response to a trigger operation on a thumbnail media material, the media material corresponding to any thumbnail media material may be displayed.

[0113] The client can respond to a selection operation of at least one media material among one or more media materials by displaying a thumbnail of at least one media material on the material selection interface. The display order of the at least one thumbnail media material is the same as the selection order of the at least one media material. For example, if the selection order of media material 1 is the first and the selection order of media material 2 is the second, then the display order of the thumbnail media material of media material 1 is the first and the display order of the thumbnail media material of media material 2 is the second. The client can respond to a trigger operation of any thumbnail media material by displaying the media material corresponding to any thumbnail media material.

[0114] by Figure 3B For example, in response to a selection operation of at least one media material among multiple media materials, the client can display a thumbnail media material 303g of at least one media material on the material selection interface 303. The target object can click on the thumbnail media material 303g, so the client can respond to the trigger operation of the thumbnail media material 303g and display the media material corresponding to the thumbnail media material.

[0115] Optionally, the associated location of any media thumbnail may include a deselection control, which can deselect the media material corresponding to that media thumbnail in response to a triggering operation of the deselection control at the associated location of any media thumbnail. The associated location of any media thumbnail may include the display position of any media thumbnail or an adjacent position of any media thumbnail, etc.

[0116] Optionally, the media selection interface may also include a first editing control, which can respond to a trigger operation on the first editing control to display the media editing interface. The media editing interface includes at least one media material, and can respond to an editing operation on at least one media material to obtain at least one edited media material.

[0117] The media editing interface may include at least one editing control, which can respond to a trigger operation on any editing control to perform corresponding editing operations on at least one media material, resulting in at least one edited media material. The editing controls include, but are not limited to, at least one of the following: effects controls, text controls, clipping controls, filter controls, background controls, and audio selection controls. Specifically, effects controls are used to add effects to at least one media material, such as "star surround effect" or "meteor shower effect"; text controls are used to add text content to at least one media material, such as adding the text "fragmented life"; clipping controls are used to edit at least one media material, for example, if at least one media material is a 20-second video clip, the target object can click the clipping control to edit the first 10 seconds of the 20-second video clip into at least one edited media material; filter controls are used to add filter effects to at least one media material, such as "whitening" or "texture" filters; background controls are used to adjust the background color of at least one media material, such as adjusting the background color to "yellow" or "blue"; and audio selection controls are used to set the target audio for at least one media material.

[0118] by Figure 3B For example, the material selection interface 303 may also include a first editing control 303h, and the client can respond to the trigger operation of the first editing control 303h to display the material editing interface.

[0119] Optionally, the first editing control may include information on the quantity of at least one media material. For example, if the quantity of at least one media material is two, then the first editing control may display the quantity information "Next (2)" for at least one media material, which indicates that the quantity of at least one media material is two.

[0120] Optionally, the client can also receive the first video generation operation via the second video preview interface, and in response to the first video generation operation, display a generation prompt message to indicate that the first video is being generated; in response to the first video generation completion event, cancel the display of the generation prompt message and display an audio browsing interface, which includes the first video.

[0121] In this embodiment, since a certain amount of time is needed to determine the reasoning process and perform packaging and rendering when generating the first video, a generation prompt message can be displayed as a notification. Figure 3E For example, Figure 3E This is a schematic diagram of the generation prompt information provided in an embodiment of this application. The client can display a second video browsing interface 310, which includes a first video generation control 310a. In response to a trigger operation on the first video generation control 310a (i.e., a first video generation operation), the client can display generation prompt information 311a on the generation prompt interface 311. In response to a first video generation completion event, the client can cancel the display of the generation prompt information and display the first video 312a on the audio browsing interface 312. Optionally, the generation prompt information may include a determined highlight segment.

[0122] Optionally, the generation prompt interface may also include an exit control, which can respond to a triggering operation to exit the generation prompt interface. After the first video is generated, the personal page includes an entry point for displaying the first video. Figure 3E For example, the generation prompt interface 311 may also include an exit control 311b. The client can exit the generation prompt interface 311 in response to the trigger operation of the exit control 311b. At this time, the first video continues to be generated. After the first video is generated, the target can click to enter the personal page. The personal page displays the display entry of the first video, for example, the display entry is "Drafts". When the target clicks the drafts, the video cover of the first video can be displayed. The client responds to the trigger operation of the video cover of the first video and displays the first video. Optionally, the generation prompt interface may also include an exit prompt message. The exit prompt message is used to prompt the exit of the generation prompt interface and the personal page includes the display entry of the first video. For example, the exit prompt message can be "You can go to the drafts on the personal page to view the generation result".

[0123] In one optional implementation, the first video is displayed on a first video browsing interface, which further includes an instruction input field; the method further includes: receiving an update instruction via the instruction input field; and, in response to the update instruction, displaying the updated first video on the first video browsing interface, wherein the updated first video is generated based on the update instruction and at least one media material.

[0124] The updated first video is displayed on the first video browsing interface. Optionally, the first video browsing interface may include the updated first video and a command input field, meaning that users can continue to enter update commands on the first video browsing interface.

[0125] The update instruction is used to update at least one of the following inference information: at least one highlight clip in at least one media clip, at least one media clip content description, the video theme of the first video, the video content description of the first video, and the video content summary of the first video. After the update is completed, the updated inference process can be obtained. The updated inference process includes at least one of the following inference information: updated highlight clip, updated media clip content description, updated video theme, updated video content description, and updated video content summary.

[0126] After any inference information is updated, the dependent inference information of that inference information is updated based on the updated inference information. The dependent inference information of any inference information can be understood as the dependent inference information generated based on that inference information.

[0127] If an update instruction targets a highlight segment in at least one media clip, then the updated highlight segment can be determined based on the update instruction and at least one media clip. The dependent inference information for the highlight segment includes the video theme, video content description, and video content summary. The video theme, video content description, and video content summary are then updated based on the updated highlight segment, thereby determining the updated video theme, updated video content description, and updated video content summary.

[0128] If an update instruction specifies a content description for at least one media asset, then the updated content description can be determined based on the update instruction. The dependent inference information for the content description includes the video theme, video content description, and video content summary. The video theme, video content description, and video content summary are then updated based on the updated content description, thereby determining the updated video theme, updated video content description, and updated video content summary.

[0129] If the update instruction specifies a video theme for the first video, then the updated video theme can be determined based on the update instruction, highlight clips, and source material descriptions. The video theme's dependent inference information includes the video content description and video content summary. Therefore, the video content description and video content summary are updated based on the updated video theme, thus determining the updated video content description and updated video content summary.

[0130] If the update instruction specifies a video content description for the first video, then the updated video content description can be determined based on the update instruction, highlight clips, source content description, and video theme. The video content description's dependent inference information includes a video content summary; therefore, the video content summary is updated based on this updated video content description to determine the updated video content summary.

[0131] If the update instruction specifies a summary of the video content for the first video, then the updated summary of the video content can be determined based on the update instruction, highlight clips, source material descriptions, video theme, and video content description. The video content summary does not rely on inference information.

[0132] If the first video does not meet the target object's video generation requirements, the target object can input an update command. The update command can supplement or modify the first video, thereby displaying the updated first video. Specifically, the inference process can be updated according to the update command to obtain the updated inference process. Then, based on the generation method of the first video and the updated inference process, at least one media material is processed to generate a video, resulting in the updated first video.

[0133] The reasoning process can be updated using update instructions, allowing for the updating of any or multiple pieces of inference information. For example, the reasoning process of the first video might include a video content description, such as the story description "first, we climbed a mountain, saw various flowers in the mountain, and finally went home." The target user's update instruction might be "after seeing various flowers, we took many photos, and finally went home." Therefore, the reasoning process can be updated using these instructions to obtain the updated reasoning process. Then, based on the first video's generation method and the updated reasoning process, at least one media clip can be processed to generate the updated first video. Similarly, if the first video's reasoning process includes information from three highlight segments, and the target user's update instruction might be "want more highlights," the reasoning process can be updated using these instructions to obtain the updated reasoning process. Then, based on the first video's generation method and the updated reasoning process, at least one media clip can be processed to generate the updated first video.

[0134] In one alternative implementation, the command input entry point may include a command input control or a command input area. Figure 3F For example, Figure 3F This is a schematic diagram of the instruction input entry provided in the embodiments of this application, wherein the audio browsing interface 313 includes an instruction input control 313a, and the audio browsing interface 314 includes an instruction input area 314a.

[0135] Optionally, the audio browsing interface may also include input prompts for inputting update commands, which are determined based on the content of at least one media asset.

[0136] The input prompts can be displayed at any location on the audio browsing interface. The input prompts, determined based on the content of at least one media asset, can include: retrieving input prompts from an input prompt database that match the content of at least one media asset. The server constructs an input prompt database, which includes input prompts corresponding to different media content. For example, input prompts for travel content could be "Where did you go today? What did you do?", and input prompts for food content could be "What did you eat today?". For instance, if the content of at least one media asset is about holiday travel, then the content of at least one media asset can be determined to be travel content, and the input prompts matching the travel content can be retrieved from the input prompt database as "Where did you go today? What did you do?".

[0137] Optionally, input prompts can be displayed at an associated location of the command input field. The associated location of the command input field may include the display location of the command input field or a location adjacent to the command input field, etc., and this embodiment does not limit this.

[0138] Optionally, the audio browsing interface may also include preset prompts. These preset prompts are used to prompt for the input of update commands and are pre-set. Specifically, the server or client can pre-set preset prompts to be displayed on the audio browsing interface. For example, the preset prompt could be "Continue to supplement, making the final product even better."

[0139] Optionally, the command input entry includes a command input control or a command input area. Triggering operations on the command input entry include triggering operations on the command input control or triggering operations on the command input area. A triggering operation on the command input entry can be understood as an input operation on the command input entry to input an update command. The input operation on the command input entry may include, but is not limited to: directly pulling up the keyboard to edit input within the command input entry, pasting operations within the command input entry, importing operations within the command input entry, and voice input operations. Therefore, the embodiments of this application provide diverse input methods, which can improve the convenience of information input and enhance the efficiency of video generation.

[0140] Optionally, the command input area can be displayed in response to a trigger action on the command input control. Optionally, displaying the command input area may include adding a new command input area or replacing the command input control with the command input area.

[0141] Optionally, in response to an input operation on the command input field, an update command can be entered by pulling up the keyboard in the audio browsing interface in response to a trigger operation on the command input field to enter the update command.

[0142] by Figure 3F For example, the first video browsing interface 313 includes a command input control 313a, and the first video browsing interface 314 includes a command input area 314a. The target object can click the command input control 313a or the command input area 314a, thereby allowing the client to respond to a trigger operation on the command input control 313a or the command input area 314a by raising the keyboard 315a in the first video browsing interface 315. That is, when the target object clicks the command input control 313a or the command input area 314a, the keyboard changes from a hidden state to a raised state, and the target object can input an update command in the command input area 315b using the keyboard 315a. Optionally, the input update command can be displayed in the command input area 315b. Optionally, the associated position of the command input area 315b may also include an input confirmation control 315c. The client can respond to a trigger operation on the input confirmation control 315c to confirm the input update command and display the updated first video 316a in the audio browsing interface 316. The associated position of the instruction input area 315b may include the display position of the instruction input area 315b or the adjacent position of the instruction input area 315b, etc.

[0143] Optionally, an automatic save function is also provided in the command input area. When an update command is entered in the command input area, the update command entered in the command input area can be automatically saved, and a save prompt message can be displayed in the command input area. This save prompt message is used to indicate that the entered update command has been automatically saved. For example, a save prompt message "13:45 automatically saved" can be displayed in the command input area. Optionally, the update command entered in the command input area can be automatically saved according to a preset time interval, such as once per minute or once per second. This embodiment of the application does not limit this.

[0144] Optionally, when the number of text characters that the target object needs to be entered in the command input area is limited, character prompts can also be displayed in the command input area. These prompts include the number of characters currently entered and the maximum number of characters supported by the command input area. For example, if the command input area displays the character prompt "10002 / 10000", "10002" represents the number of characters currently entered, and "10000" represents the maximum number of characters supported by the command input area. When the number of characters currently entered exceeds the maximum number of characters, further character input is not possible. It should be understood that the embodiments of this application may limit the maximum number of characters supported by the command input area, or they may not limit the maximum number of characters supported by the command input area.

[0145] Optionally, in response to input operations in the command input area, the input prompt message is replaced with the entered update command. Optionally, the font and color of the input prompt message can be set as needed, such as setting the font size to 18 points and the color to gray. Optionally, when the update command is entered via the keyboard, the keyboard distance from the bottom is a preset value (e.g., 50 points).

[0146] Optionally, when the keyboard is hidden, the input prompt information displayed in the first video browsing interface is the first input prompt information; when the keyboard is raised, the input prompt information displayed in the first video browsing interface is the second input prompt information. The second input prompt information includes the first input prompt information; that is, the first input prompt information can be part or all of the second input prompt information. The second input prompt information is used to prompt for input update commands and is determined based on the content of at least one media material. For example, the second input prompt information could be "What time did you go?", "Did anything interesting happen?", or "Where did you go?".

[0147] Optionally, the import operation in the instruction input field may include: an import control in the first video browsing interface, which can import update instructions in the first video browsing interface in response to a trigger operation on the import control, and display the imported update instructions in the instruction input area.

[0148] Optionally, voice input can be performed directly on the first video browsing interface, and the input voice can be converted into text. In response to the voice input operation in the audio browsing interface, the voice input data is displayed in the command input area of ​​the audio browsing interface, and the voice input data is used as an update command.

[0149] In one alternative implementation, in response to a regeneration operation, a new first video is displayed on the first video browsing interface, the new first video being different from the first video.

[0150] Optionally, the regeneration operation may include, for example, a voice control operation. Alternatively, if the first video browsing interface includes a video update control, then the regeneration operation may include an operation triggered in response to the video update control.

[0151] Optionally, the first video browsing interface may also include a video update control. In response to a trigger operation on the video update control, a new first video may be displayed on the first video browsing interface. The new first video is different from the original first video.

[0152] If the target user is not satisfied with the generated first video, or wants to try generating a different first video, they can click the video update control in the first video browsing interface. The client will then respond to this trigger by displaying a new first video in the first video browsing interface. If there is only one first video, the new first video will be completely different from the first video; if there are multiple first videos, the new first video will not be completely identical to the first video. The new first video not being completely identical to the first video can include: the new first video being partially the same as the first video; or the new first video being completely different from the first video.

[0153] by Figure 3G For example, Figure 3G This is a schematic diagram of the first video browsing interface provided in the embodiments of this application. The first video browsing interface 317 includes a first video 317a and a video update control 317b. The client can respond to the trigger operation of the video update control 317b to display a new first video 318a in the audio browsing interface 318.

[0154] Optionally, the new first video is generated based on a new inference process and at least one media material. The new inference process includes at least one of the following: a new highlight segment in at least one media material, a new content description of at least one media material, a video theme of the new first video, a video content description of the new first video, and a video content summary of the new first video.

[0155] The reasoning process for the new first video differs from that of the first video. In other words, while both the new and first video reasoning processes are determined based on at least one media source, they are not identical. The new first video is generated based on this new reasoning process.

[0156] Optionally, in response to a trigger operation on the video update control, a new first video can be generated by processing the highlight segment based on the inference process; or, in response to a trigger operation on the video update control, the inference process can be updated to obtain a new inference process, and the highlight segment can be generated by processing the video based on the new inference process to obtain a new first video.

[0157] In this embodiment of the application, a video update control is displayed on the first video browsing interface. When the target object is not satisfied with the generated first video, or when the target object wants to try to generate another first video, the target object can click the video update control to generate a new first video, which can improve the interactive experience of the target object when generating videos.

[0158] In one alternative implementation, in response to a triggering operation on the video update control, a generation prompt message is displayed to indicate that a first video is being generated; in response to a completion event of the first video generation, the generation prompt message is de-displayed, and an audio browsing interface is displayed, which includes the first video.

[0159] by Figure 3H For example, Figure 3H This is a schematic diagram of the first video browsing interface provided in the embodiments of this application. The first video browsing interface 319 includes a first video 319a and a video update control 319b. The client can respond to the trigger operation of the video update control 319b and display the generation prompt information 320a on the generation prompt page 320. The client can also respond to the first video generation completion event, cancel the display of the generation prompt information, and display the first video 321a on the first video browsing interface 321.

[0160] In one alternative implementation, the video type of the new first video is matched with the type preference of the target object used to input the update instruction.

[0161] The video type can refer to the video's presentation style, such as abstract or concise. The target object's type preference can be determined based on its historical interaction information. Specifically, this historical interaction information can include the target object's historical video browsing and generation information. For example, if the target object's historical video browsing information indicates that most of the videos it viewed were concise, then its type preference can be determined to be concise. Similarly, if the target object's historical video generation information indicates that most of the videos it generated were concise, then its type preference can be determined to be concise. Likewise, if both the historical video browsing and generation information indicate that the target object's historical video browsing and generation information are concise, then its type preference can be determined to be concise. When generating a new first video, the video type of the new first video is also a concise type, depending on the type preference of the target object used to input the update instruction. For example, if the type preference is concise, the video type of the generated new first video is also concise.

[0162] In one alternative implementation, the video type of the new first video matches the video type specified by the target object used to input the update instruction. That is, when generating a new first video, the target object inputting the update instruction can specify the video type of the generated new first video, thereby satisfying the video generation requirements of the target object.

[0163] Optionally, in response to a triggering action on the video update control, at least one candidate video style option may be displayed, and in response to a triggering action on the target video style option, a new first video may be displayed, wherein the target video style option is any one of the at least one candidate video style option.

[0164] by Figure 3I For example, Figure 3IThis is a schematic diagram of a first video browsing interface provided in an embodiment of this application. In this interface, the client can respond to a trigger operation on the video update control 322a in the first video browsing interface 322, displaying at least one candidate video style option 323a on the video type selection page 323. The client can also respond to a trigger operation on a target video style option, displaying a new first video. Optionally, the video type selection page 323 may further include a second confirmation control 323b. The client can respond to both trigger operations on the target video style option and the second confirmation control 323b to display a new first video. Optionally, the video type selection page 323 may also include a video style option expansion control 323c. The client can respond to a trigger operation on the video style option expansion control 323c to display other candidate video style options, which are not entirely identical to the at least one candidate video style option.

[0165] Optionally, the video type selection page can be displayed above the audio browsing interface. The display area of ​​the video type selection page can be less than or equal to the display area of ​​the audio browsing interface. The video type selection page can be displayed as a pop-up or a floating layer, for example. Optionally, if the display area of ​​the video type selection page is less than the display area of ​​the audio browsing interface, the client can respond to a trigger operation on a display location other than the video type selection page in the audio browsing interface, cancel the generation of a new first video, and display the audio browsing interface. Optionally, the video type selection page can also include a deselection control. The client can respond to a trigger operation on the deselection control, cancel the generation of a new first video, and display the audio browsing interface.

[0166] Optionally, the first video browsing interface may also include a second editing control, which can respond to a trigger operation on the second editing control to display a video editing interface. The video editing interface includes a first video and can respond to an editing operation on the first video to obtain the edited first video.

[0167] The video editing interface may include at least one editing control, which can respond to a trigger operation on any editing control to perform corresponding editing operations on the first video, resulting in an edited first video. The editing controls include, but are not limited to, at least one of the following: special effects controls, text controls, clipping controls, filter controls, background controls, and audio selection controls. Specifically, special effects controls are used to add effects to the first video, such as "star surround effect" or "meteor shower effect"; text controls are used to add text content to the first video, such as adding the text "fragments of life"; clipping controls are used to edit the first video, for example, if the first video is 20 seconds long, the target can click the clipping control to edit the first 10 seconds of the 20-second video into the edited first video; filter controls are used to add filter effects to the first video, such as "whitening" or "texture" filters; background controls are used to adjust the background color of the first video, such as changing the background color to "yellow" or "blue"; and audio selection controls are used to set the target audio for the first video, for example, the audio selection control may include an audio synthesis access point.

[0168] by Figure 3I For example, the first video browsing interface 322 may also include a second editing control 322b, and the client may display the video editing interface in response to a trigger operation on the second editing control 322b.

[0169] Optionally, the first video can be played in response to a trigger operation on the first video. Further optionally, the first video can be stopped from playing in response to a trigger operation on the first video.

[0170] Optionally, the first video browsing interface may also include a video playback control, which can play the first video in response to a trigger operation on the video playback control.

[0171] Further optionally, in response to a triggering operation on the video playback control, a video pause control can be displayed, and in response to a triggering operation on the video pause control, playback of the first video can be stopped. Displaying the video pause control can include: replacing the video playback control with a video pause control, or adding a video pause control to the audio browsing interface.

[0172] Optionally, the first video browsing interface may also include a video progress bar for the first video. Responding to a trigger operation on the video progress bar, the first video can be played from the corresponding video time point. That is, the target object can drag the progress bar, allowing the client to start playing the first video at different video time points, improving the target object's interactive experience. Optionally, the associated position of the video progress bar may include the corresponding video time point and / or the video duration of the first video. For example, if the video time point of the first video corresponding to the video progress bar is "00:01", it indicates that the first second of the first video is currently playing, and if the video duration of the first video is "00:07", it indicates that the video duration of the first video is 7 seconds. The associated position of the video progress bar may include the display position of the video progress bar or adjacent positions of the video progress bar (such as left or right), etc., and this embodiment does not limit this.

[0173] Optionally, the first browsing interface may also include a first return control, which can respond to a trigger operation on the first return control, save the first video, and display the parent page of the audio browsing interface. The parent audio browsing interface can be a second video browsing interface. Figure 3I For example, the first video browsing interface 322 may also include a first return control 322c. The client can respond to the trigger operation of the first return control 322c, save the first video, and display the previous page of the first video browsing interface.

[0174] Optionally, in response to a trigger operation on the first return control, the options to clear content and save content can be displayed; in response to a trigger operation on the option to clear content, the first video can be deleted and the previous page of the audio browsing interface can be displayed; or, in response to a trigger operation on the option to save content, the first video can be saved and the previous page of the first video browsing interface can be displayed.

[0175] Optionally, the first video browsing interface can also provide an automatic save function. When the first video browsing interface includes the first video, the first video in the audio browsing interface can be automatically saved.

[0176] Optionally, after the first video is saved, the profile page includes an entry point to display the first video. Specifically, after the client saves the first video, the target audience can click to enter the profile page, which displays an entry point to display the first video, for example, displayed as "Drafts". When the target audience clicks on the Drafts folder, the video cover of the first video will be displayed, and the client will respond to the trigger operation on the video cover of the first video and display the first video.

[0177] Optionally, the first video browsing interface may further include synthesized speech information. In response to a trigger operation on the speech information, a speech selection interface may be displayed. The speech selection interface includes speech information for at least one candidate speech. In response to a trigger operation on the speech information of any one of the at least one candidate speech, the target speech may be replaced with that candidate speech. Optionally, the associated location of the synthesized speech information may include an audio shutdown control. In response to a trigger operation on the audio shutdown control, playback of the synthesized speech may be stopped. The associated location of the synthesized speech information may include its display position or an adjacent position (such as left or right).

[0178] S202, the target speech of the target object is obtained through the audio synthesis access point. The target speech has a target timbre. The target speech is used to perform speech synthesis processing on the target text to obtain synthesized speech with the target timbre.

[0179] The target text includes the text corresponding to the first video. The target text can be input by the target object or text generated by the server based on the video content of the first video that matches the first video.

[0180] In one alternative implementation, the target speech may be acquired in real time via an audio synthesis entry point, for example, by acquiring the target speech through an audio synthesis access entry point. For instance, the client may activate the microphone in response to a trigger operation on the audio synthesis entry point and acquire the target speech via the microphone.

[0181] Optionally, the client can display reference text and voice acquisition entry via an audio synthesis access point; and acquire the target voice output by the target object in response to the reference text via the voice acquisition entry point.

[0182] by Figure 3JTaking the video generation interface shown as an example, the client can respond to the trigger operation of the audio synthesis entry point by displaying authorization prompts and authorization controls. The authorization prompts are used to prompt the target object to authorize the output target speech; examples include "Requesting your voice authorization," and authorization controls include a "Clone my voice" button. The client can also respond to the trigger operation of the authorization controls by displaying reference text and a voice capture control (i.e., the voice capture entry point). Optionally, the client can also respond to the trigger operation of the authorization controls by displaying voice capture prompts, which prompt the target object to output the target speech based on the reference text; examples include "Read the example sentence, record your voice" or "Your recorded voice will be used for timbre cloning. Please read the example sentence with the desired tone and emotion," etc. Optionally, the client can also respond to the trigger operation of the authorization controls by displaying text update controls, which display updated reference text; examples include a "Change text" button. Optionally, the voice acquisition entry point may also include a first acquisition rule prompt message. This message prompts the target object that the output duration of the target voice must be greater than or equal to a preset duration threshold. For example, the first acquisition rule prompt message might say, "Press and hold to record for more than 5 seconds." Optionally, the client can respond to a trigger operation on the voice acquisition control by switching its display style from a first display style to a second display style. The first display style refers to the display style before the trigger operation, indicating that the device is not currently in voice acquisition mode. The second display style refers to the display style after the trigger operation, indicating that the device is currently in voice acquisition mode. After acquiring the target voice, the target text can be processed into speech synthesis to obtain synthesized speech with the target timbre, which is then displayed on the audio browsing interface.

[0183] Optionally, the client may respond to receiving and continuously receiving voice input operations, during the duration of the voice input operation, acquire the user's input voice, and display a first voice operation area, wherein the voice input operation includes performing a first operation on the voice input control; respond to a voice input lock operation, continue acquiring the user's input voice, wherein the voice input lock operation includes performing a second operation and stopping the first operation during the duration of the voice input operation; respond to a voice input pause operation, pause receiving the user's input voice and display a second voice operation area to replace the first voice operation area; the second voice operation area is used to send the acquired target voice.

[0184] by Figure 3K For example, Figure 3K This is a schematic diagram of a voice input interface provided in an embodiment of this application, wherein the voice input interface 324 includes a voice input control 324a.

[0185] The first voice operation area can be displayed in any area of ​​the voice input interface. The first operation can be a pressing operation on the voice input control. That is, when the user presses the voice input control, the client will continuously collect the user's voice input, and when the user stops pressing the voice input control, the client will stop collecting the user's voice input.

[0186] In one alternative implementation, the voice input interface may further include at least one of the following controls: an input lock control and a voice delete control. The input lock control is used to trigger continuous acquisition of user-inputted voice by performing a second operation and stopping the first operation, and the voice delete control is used to delete the acquired voice.

[0187] Optionally, an input lock control, a voice delete control, and a voice input control can be displayed in the first voice operation area. That is, the first voice operation area may include at least one of the following controls: an input lock control, a voice delete control, and a voice input control.

[0188] by Figure 3K For example, the client responds by receiving and continuously receiving voice input operations; that is, the client receives and continuously receives the user's pressing operation on the voice input control 324a in the voice input interface 324. It continuously collects the user's voice input and displays a first voice operation area 324b. The first voice operation area 324b may include the voice input control 324a, an input lock control 324c, and a voice deletion control 324d. The input lock control 324c is used to trigger continuous collection of user voice input by performing a second operation and stopping the first operation. The second operation is different from the first operation; for example, the first operation is a pressing operation, and the second operation is a sliding operation. The voice deletion control 324d is used to delete the collected voice.

[0189] Optionally, the voice input interface may also include a voice sending prompt message, which prompts the user to send the captured voice. The voice sending prompt message can instruct the user to stop the first operation, thereby allowing the client to send the captured voice to the server. For example, the voice sending prompt message could be "Release to complete," meaning that when the user stops pressing the voice input control, the client will send the voice captured during the pressing period to the server. Optionally, the voice sending prompt message can be displayed in the first voice operation area. Optionally, the voice sending prompt message can be displayed at an associated position of the voice input control. The associated position of the voice input control can refer to the display position of the voice input control, or an adjacent position, such as displaying it above the voice input control. In this embodiment, by displaying the voice sending prompt message, the user is prompted about the voice sending method, improving the user's interactive experience and enhancing user engagement.

[0190] Optionally, in response to a voice deletion operation, the captured voice can be deleted, which may include performing a third operation and stopping the first operation during the duration of the voice input operation.

[0191] Optionally, the voice input interface may also include a voice deletion control, in which case the third operation may include any of the following: 1. Move from the location of the first operation's contact point to the location of the voice deletion control.

[0192] The user can continuously press the voice input control, at which point the contact point is located. While continuously pressing, the user can move the contact point to the location of the voice deletion control. Furthermore, the user can stop the first operation after performing this third operation—that is, stop pressing after moving to the location of the voice deletion control—and the client will then delete the captured voice.

[0193] 2. Move a preset distance from the location of the contact point of the first operation.

[0194] The specific method for moving the preset distance from the contact point of the first operation can be as follows: move the preset distance from the contact point of the first operation towards the location of the voice deletion control. Specifically, the user can continuously press the voice input control, at which point the contact point is the location of the voice input control. While continuously pressing, the user can move the preset distance towards the location of the voice deletion control; for example, the preset distance is 2cm. If the voice deletion control is to the left of the voice input control, the user can move 2cm to the left while continuously pressing. Furthermore, the user can stop the first operation after performing the third operation, that is, stop pressing after moving the preset distance towards the location of the voice deletion control. Then, the client will delete the captured voice.

[0195] Optionally, in response to a third operation performed during the duration of the voice input operation, a deletion prompt message can be displayed to prompt the user to delete the captured voice. For example, when the user moves from the location of the first operation's contact point to the location of the voice deletion control, or moves a preset distance from the location of the first operation's contact point, the deletion prompt message "Release to Cancel" is displayed to prompt the user that the captured voice can be deleted after the first operation ends. Optionally, the deletion prompt message can be displayed in the first voice operation area. Optionally, the deletion prompt message can be displayed at an associated position of the voice deletion control. The associated position of the voice deletion control can refer to the display position of the voice deletion control, or an adjacent position of the voice deletion control, for example, it can be displayed above the voice deletion control. In this embodiment, by displaying a deletion prompt message, the user can be prompted about the voice deletion method, improving the user's interactive experience and enhancing user engagement.

[0196] In response to the voice input lock operation, the client will continue to collect the user's voice input. At this time, the user can continue to input voice without freeing their hands, reducing the operational burden when performing voice input and improving the interactive experience of voice input.

[0197] In one alternative implementation, the voice input interface further includes an input lock control; the second operation includes any of the following: 1. Move from the location of the contact point of the first operation to the location of the input locking control.

[0198] The user can continuously press the voice input control, at which point the contact point is located. While continuously pressing, the user can move the contact point to the location of the input lock control. Furthermore, the user can stop the first operation after performing the second operation—that is, stop pressing after moving to the input lock control. The client then continues to collect the user's voice input. In other words, the user can continue to input voice while keeping their hands free.

[0199] 2. Move a preset distance from the location of the contact point of the first operation.

[0200] The specific method for moving the preset distance from the contact point of the first operation can be as follows: move the preset distance from the contact point of the first operation towards the location of the input locking control. Specifically, the user can continuously press the voice input control, at which point the contact point is the location of the voice input control. While continuously pressing, the user can move the preset distance towards the location of the input locking control, for example, 2cm. If the input locking control is to the right of the voice input control, the user can move 2cm to the right while continuously pressing. Furthermore, the user can stop the first operation after performing the second operation, i.e., stop pressing after moving the preset distance towards the location of the input locking control. Then, the client continues to collect the user's voice input, meaning that the user can continue to input voice while freeing their hands.

[0201] Optionally, in response to a second operation performed during the duration of the voice input operation, a lock prompt message can be displayed to prompt the user to continue voice input. For example, when moving from the contact point of the first operation to the location of the input lock control, or moving a preset distance from the contact point of the first operation, the client displays the lock prompt message "Release Lock" to indicate to the user that after ending the first operation, they can continue voice input with their hands free. Optionally, the lock prompt message can be displayed in the first voice operation area. Optionally, the lock prompt message can be displayed at an associated position of the input lock control. The associated position of the input lock control can refer to the display position of the input lock control, or an adjacent position, such as above the input lock control. In this embodiment, by displaying the lock prompt message, the user is prompted about the voice input lock method, improving the user's interactive experience and enhancing user engagement.

[0202] Optionally, an input lock control can be displayed in the first voice operation area. Figure 3K For example, the voice input interface includes a first voice operation area 324b, the first voice operation area 324b includes an input locking control 324c, and the second operation may include moving from the position of the contact point of the first operation to the position of the input locking control 324c.

[0203] In one alternative implementation, in response to the voice input lock operation, the user's voice input continues to be acquired. The input lock control in the voice input interface can be replaced with an input pause control, and the voice input control in the voice input interface can be replaced with a voice send control for display. The input pause control is used to pause the reception of the user's voice input, and the voice send control is used to send the target voice.

[0204] Optionally, the voice input interface may include a first voice operation area, where an input lock control, a voice input control, and a voice delete control can be displayed. In response to a voice input lock operation, the interface can continue to collect user-inputted voice and update the first voice operation area. Specifically, updating the first voice operation area can involve replacing the input lock control with an input pause control and replacing the voice input control with a voice send control.

[0205] by Figure 3K For example, the voice input interface includes a first voice operation area 324b. In response to a voice input lock operation, the client can replace the input lock control 324c with an input pause control 324e, and the voice input control 324a with a voice send control 324f. The user can click the input pause control 324e, causing the client to pause receiving user-inputted voice; or, the user can click the voice send control 324f, causing the client to send the acquired target voice. The first voice operation area 324b may also include a voice delete control 324d. The user can click the voice delete control 324d, causing the client to delete the acquired target voice.

[0206] The voice input pause operation can be triggered by the user triggering the input pause control in the first voice operation area, or by the user performing a preset gesture operation.

[0207] Optionally, the first voice operation area may include an input pause control. In response to a voice input pause operation, pausing the reception of user input and displaying a second voice operation area to replace the first voice operation area can be achieved in the following manner: in response to a trigger operation on the input pause control, pausing the reception of user input and displaying a second voice operation area to replace the first voice operation area.

[0208] by Figure 3K For example, the first voice operation area 324b may include an input pause control 324e. The user can click the input pause control 324e, so that the client can respond to the trigger operation of the input pause control 324e, pause the reception of the user's voice input and display the second voice operation area 324f to replace the first voice operation area.

[0209] The second voice operation area is used to play, send, or edit the captured target voice. Specifically, editing the captured target voice can include deleting or truncating the captured target voice. For example, if the target voice is a 6-second voice recording, then the 6-second voice recording can be deleted, or the last 4 seconds of the 6-second voice recording can be truncated. Therefore, the first 2 seconds of the 6-second voice recording will be deleted.

[0210] In this embodiment, the acquired target speech can be edited. For example, if a 6-second segment of the acquired target speech is entirely noise, the user can delete it. Similarly, if the first 2 seconds of the acquired target speech are noise and the last 4 seconds are normal speech, the last 4 seconds can be truncated, thus deleting the noise from the first 2 seconds. Editing the acquired target speech enriches the user's interactive operations with the input speech, thereby improving the user's voice input experience.

[0211] In one optional implementation, the second voice operation area includes a voice deletion control and a voice playback control area. The voice deletion control is used to delete the target voice, and the voice playback control area includes at least one of a voice sending control, a voice playback control, and a playback progress control bar. The voice sending control is used to send the target voice, the playback progress control bar is used to adjust the playback start time of the target voice, and the voice playback control is used to start playing the target voice from the playback start time.

[0212] Optionally, in response to a trigger operation on the voice playback control, the target voice can be played starting from the playback start time, and a voice pause playback control can be displayed; alternatively, in response to a trigger operation on the voice pause playback control, the playback of the target voice can be paused. Displaying the voice pause playback control can control the switching of the control's state, for example, switching the continuous playback state corresponding to the voice playback control to the pause playback state corresponding to the voice pause playback control. The control state includes both continuous playback and pause playback states.

[0213] Optionally, the client can send target voice in response to a trigger action on the voice sending control. The target voice can refer to the user's voice input before editing or the user's voice input after editing. The client can send the target voice to the server in response to the user's trigger action on the voice sending control.

[0214] Optionally, the playback progress control bar can also include the duration of the target audio, for example, displaying the audio duration as "0:06", which means that the duration of the target audio is 6 seconds.

[0215] Optionally, in response to adjustments made to the playback progress control bar, the starting point for playback of the target audio can be determined; and the corresponding audio data can be played starting from the starting point for playback of the target audio.

[0216] The adjustment operation can be a sliding operation. The audio time point corresponding to the end point of the sliding operation is the starting point for playing the target audio. For example, if the audio time point corresponding to the end point of the sliding operation is "0:04", then the client will start playing the target audio from the 4th second.

[0217] Optionally, a voice-activated playback prompt can be displayed on the voice playback interface. This prompt instructs the user to slide the playback progress control bar to start playing the corresponding voice data from the end point of the slide. For example, the prompt could read, "Slide on the recording to play from any position." Optionally, the prompt can be displayed in the second voice operation area. Optionally, the prompt can be displayed in an associated position within the voice control area. This associated position could refer to the display position of the voice control area or an adjacent position, such as above it. In this embodiment, by adjusting the start time of the target voice playback using the playback progress control bar, the target voice can be played from any point in time, thereby improving the user's voice operation interaction experience and enhancing user engagement.

[0218] In one optional implementation, in response to receiving and continuously receiving a voice input operation, the system can acquire user-inputted voice during the duration of the voice input operation, the voice input operation including performing a first operation on the voice input control; in response to a voice input locking operation, the system can continue acquiring user-inputted voice, the voice input locking operation including receiving a second operation and stopping the first operation during the duration of the voice input operation; and in response to a voice transmission trigger event, the system can transmit the acquired target voice. The second operation can be found in the relevant descriptions in the above embodiments and will not be repeated here.

[0219] Optionally, the voice input interface may also include at least one of the following controls: an input lock control and a voice delete control. The input lock control is used to trigger continuous acquisition of user-inputted voice by performing a second operation and stopping the first operation, and the voice delete control is used to delete the acquired voice.

[0220] by Figure 3K For example, Figure 3KThis is a schematic diagram of a voice input interface provided in an embodiment of this application. The client responds to and continuously receives voice input operations; that is, the client receives and continuously receives the user's pressing operation on the voice input control 324a in the voice input interface 324. It continuously collects the user's input voice. The voice input interface 325 may also include an input locking control 324c and a voice deletion control 324d. The input locking control 324c is used to trigger continuous collection of user-input voice by performing a second operation and stopping the first operation. The second operation is different from the first operation; for example, the first operation is a pressing operation, and the second operation is a sliding operation. The voice deletion control 324d is used to delete the collected voice. As shown in the voice input interface 326, when the contact point of the first operation moves from the location of the voice input control to the location of the input locking control 324c, a voice sending control 324f will be displayed in the voice input interface 327. The client can respond to the triggering operation of the voice sending control 324f to determine the collected target voice. In response to the first operation, the client moves its contact point from the location of the voice input control to the location of the input lock control 324c. On the voice input interface 327, the input lock control is replaced with an input pause control 324e. In response to the triggering operation of the input pause control 324e, the client pauses receiving user-inputted voice, and on the voice input interface 328, a voice deletion control 324d and a voice playback control area 324g are displayed. The user can click the voice deletion control 324d, allowing the client to delete the target voice message. The voice playback control area 324g includes a voice send control 324f, a voice playback control 324h, and a playback progress control bar. The user can click the voice send control 324f, allowing the client to send the target voice message to the server. Users can trigger the playback progress control bar to adjust the start time of the target audio playback. For example, if the initial start time of the target audio is 0:00 and its duration is 6 seconds, the user can drag the playback progress control bar to adjust the initial start time and obtain the target playback start time. The target playback start time can range from 0:00 to 0:06. Users can click the audio playback control 324h, allowing the client to start playing the target audio from that start time.

[0221] Optionally, the voice input interface may also include a voice sending prompt message, which prompts the user to send the captured voice. The voice sending prompt message can instruct the user to stop the first operation, thereby allowing the client to send the captured voice to the server. For example, the voice sending prompt message could be "Release to complete," meaning that when the user stops pressing the voice input control, the client will send the voice captured during the pressing period to the server. Optionally, the voice sending prompt message can be displayed at an associated position of the voice input control. This associated position can refer to the display position of the voice input control, or an adjacent position, such as above the voice input control. In this embodiment, displaying the voice sending prompt message can prompt the user on how to send voice messages, improving the user's interactive experience and enhancing user engagement.

[0222] Optionally, in response to a voice deletion operation, the captured voice can be deleted. The voice deletion operation includes performing a third operation and stopping the first operation during the duration of the voice input operation. The third operation can be found in the relevant description in the above embodiments, and will not be repeated here.

[0223] The voice sending operation can be triggered by the user triggering the voice sending control or by the user performing a preset gesture operation.

[0224] Optionally, if the voice input interface includes a voice sending control, then in response to a voice sending operation, the captured target voice can be sent in the following way: in response to a trigger operation on the voice sending control, the captured target voice can be sent.

[0225] In one alternative implementation, in response to receiving and continuously receiving a voice input operation, the user-inputted voice is acquired during the duration of the voice input operation; in response to a voice transmission trigger event, the acquired target voice is transmitted.

[0226] For example, in response to receiving and continuously receiving voice input operations—that is, receiving and continuously receiving user presses on the voice input control in the voice input interface—the client will continuously capture the target voice input and display a voice sending prompt message on the voice input interface. This prompt message can instruct the target to stop pressing the voice input control; for example, the prompt message could be "Release to complete." When the client detects that the target has stopped pressing the voice input control, it can send the target voice captured during the pressing period to the server.

[0227] In one alternative implementation, at least one candidate speech of the target object can be displayed via an audio synthesis access point, and the target speech can be determined in response to a selection operation of the at least one candidate speech. The at least one candidate speech of the target object may be pre-acquired and stored.

[0228] by Figure 3L For example, Figure 3L This is a schematic diagram of the video generation interface provided in this application embodiment. The audio synthesis access point can be the audio synthesis access control in the video generation interface. In response to a trigger operation on the audio synthesis access control, the client can display a voice selection interface. The voice selection interface can include at least one candidate voice of the target object. The client can play the candidate voice in response to a trigger operation on any candidate voice. If the target object wants to clone the timbre of a candidate voice to generate synthesized voice corresponding to the first video, the target object can select that candidate voice. In response to the selection operation on at least one candidate voice, the client determines the candidate voice selected by the selection operation as the target voice.

[0229] In one optional implementation, the target speech can be determined based on at least one media material via an audio synthesis access point. For example, if at least one media material carries audio, then the audio carried by at least one media material can be used as the target speech. Optionally, if multiple media materials carry audio, then the audio carried by each media material can be concatenated to obtain the target speech, or an audio file can be randomly selected from the audio carried by multiple media materials as the target speech, or the audio file whose speech parameters meet preset conditions among the audio carried by multiple media materials can be used as the target speech. Speech parameters may include pitch, intensity, duration, etc., and the speech parameters meeting preset conditions may be, for example, the audio file with the longest playback duration, or the audio file with an intensity greater than a preset intensity threshold, or the audio file with a pitch greater than a preset pitch threshold, etc. As another example, the target speech may be the audio carried by the video being edited during video editing. The edited video may include, for example, a first video, then the target speech may be the audio in the first video, which may be generated based on at least one media material, or it may be the audio inherent in at least one media material.

[0230] In one optional implementation, after acquiring the target speech via the audio synthesis access point, a speech transmission prompt and a cancel transmission control can be displayed. The speech transmission prompt indicates that the target speech is being transmitted. The client can also cancel the transmission of the target speech in response to triggering the cancel transmission control.

[0231] by Figure 3M For example, Figure 3MThis is a schematic diagram of the voice input interface provided in this application embodiment. After the client collects the target voice through the audio synthesis access portal, the client can send the target voice to the server, so that the server can perform speech synthesis processing on the target text based on the target voice to obtain synthesized voice with the target timbre. During the process of the client sending the target voice, the client can display a voice sending prompt message and a cancel sending control, such as "Timbre uploading in progress". The client can also cancel sending the target voice in response to the triggering operation of the cancel sending control, such as a "Cancel" button.

[0232] In one optional implementation, after obtaining the target speech of the target object via the audio synthesis access point, a speech synthesis prompt message and a cancel synthesis control can be displayed. The speech synthesis prompt message is used to indicate that the target speech is being processed by speech synthesis. The client can also cancel the speech synthesis processing of the target speech in response to triggering the cancel synthesis control.

[0233] by Figure 3N For example, Figure 3N This is a schematic diagram of the voice input interface provided in this application embodiment. After the client collects the target voice through the audio synthesis access portal, the client can send the target voice to the server, so that the server can perform speech synthesis processing on the target text based on the target voice to obtain synthesized voice with the target timbre. During the server's generation of synthesized voice, the client can display speech synthesis prompts and a cancel send control, such as "Timbre cloning in progress". The client can also cancel the speech synthesis processing of the target voice in response to the triggering operation of the cancel send control, such as a "Cancel" button.

[0234] S203, Display target audio and video, which is obtained by merging the first video and the synthesized speech.

[0235] In one alternative implementation, the client may display at least one synthesized speech having the target timbre, and play the target synthesized speech in response to a playback operation on the target synthesized speech, wherein the target synthesized speech is any one of the at least one synthesized speech. The client may also trigger the display of the target audio-visual content in response to a selection operation on the target synthesized speech, the target audio-visual content being obtained by merging the first video and the target synthesized speech.

[0236] by Figure 3O For example, Figure 3OThis is a schematic diagram of an audio browsing interface provided in an embodiment of this application. The client can display at least one synthesized voice with the target timbre. Different synthesized voices have different language types, voice styles, emotional expressions, or modal particles. The language type is, for example, Chinese or English; the voice style is, for example, a deep voice, a sweet voice, or a child's voice; the emotional expression is, for example, cheerful or deep; and the modal particle is, for example, "ne" or "la". The target object can name any synthesized voice to obtain its voice identifier. The target object can also long-press any synthesized voice to play it. If the target object wants to merge the first video and the synthesized voice identified as "Timbre Version 1", the target object can select the synthesized voice identified as "Timbre Version 1". For example, after clicking the synthesized voice identified as "Timbre Version 1", the target object can click the "Save Timbre" button. The client can respond to the selection operation of the target synthesized voice by sending the target synthesized voice to the server so that the server can merge the first video and the target synthesized voice to obtain the target audio and video. After the server sends the target audio and video to the client, the client can display the target audio and video.

[0237] In one optional implementation, the client can display reference text and a voice acquisition entry via an audio synthesis access point. The voice acquisition entry may further include a first acquisition rule prompt, which prompts the target object that the output duration of the target voice must be greater than or equal to a preset duration threshold. For example, the first acquisition rule prompt may include... Figure 3P The voice input interface shows the option to "Press and hold to record for more than 5 seconds". If the output duration of the target voice is less than the preset duration threshold, the client can display a voice acquisition failure message. This message indicates the reason for the failure, for example... Figure 3P The voice input interface shows "Insufficient recording time, please record again".

[0238] In one optional implementation, the client or server can set a voice acquisition duration threshold. During the acquisition of target voice through the voice acquisition portal, if the interval between the acquisition duration and the voice acquisition duration threshold is less than a preset duration, the client can display a acquisition termination prompt. This prompt indicates that the acquisition of target voice will stop after the preset duration. For example, assuming the voice acquisition duration threshold is 30 seconds and the preset duration is 4 seconds, after the target inputs 26 seconds of voice, the client can display an acquisition termination prompt. The acquisition termination prompt might be something like this: Figure 3Q The voice input interface shown displays "Stop recording in 4 seconds." The client can also update the acquisition termination prompt to a countdown timer to indicate to the target that the recording of the target's voice is about to stop. The countdown timer information is as follows: Figure 3QThe “00:01” in the voice input interface indicates that there is 1 second left before the target voice is stopped.

[0239] In one alternative implementation, if the target object does not output the target speech in response to the reference text, the client can display speech output guidance information to guide the target object to output the target speech in response to the reference text. The speech output guidance information may include, for example... Figure 3R The voice input interface shown prompts "Please read aloud the example sentence and record it."

[0240] In one optional implementation, the client sends the target speech to the server, which then performs speech synthesis processing on the target text based on the target speech to obtain synthesized speech with the target timbre. If the server fails to generate the synthesized speech, it can send the result to the client, indicating that the synthesized speech generation was unsuccessful. The client can display a generation failure message corresponding to the result to prompt the target user to re-output the target speech. The generation failure message might include, for example... Figure 3S The voice input interface displays "Voice cloning failed, please try again." Optionally, the server may have failed to generate synthesized speech due to network issues, or the server's voice cloning algorithm may have failed to generate synthesized speech based on the target speech. Optionally, the client can further display the reason for the synthesized speech generation failure, such as network issues or algorithm problems.

[0241] In one alternative implementation, the client can display the speech identifier of any synthesized speech in response to a naming operation on that synthesized speech. For example... Figure 3T The audio browsing interface shown may include at least one synthesized speech and a voice identifier input area. The target user can click on any synthesized speech and enter its voice identifier in the input area. If the voice identifier does not conform to the naming rules, the client can display a naming failure message, prompting the target user to re-enter the voice identifier. The naming failure message may be, for example, "Naming unavailable".

[0242] In one alternative implementation, content to be published can also be generated based on the target audio and video. The content to be published includes a first video, a title of the published content, and the body of the published content. The title of the published content and the body of the published content are obtained based on the video theme.

[0243] Optionally, in response to a confirmation operation for the first video, a content editing interface can be displayed, which includes: the first video, the title of the content to be published, and the body of the content to be published; in response to a publishing operation for the first video, content to be published containing the first video can be generated and published.

[0244] Specifically, the title and body text of the published content can be determined based on the video's theme. For example, the video's theme can be input into the content generation model, which can then determine the title and body text of the published content that match the video's theme. The content generation model can be obtained through supervised or unsupervised training.

[0245] Optionally, the title of the published content can match the video title, and the body of the published content can match the video content description.

[0246] Optionally, matching the title of the published content with the title of the video can include having the same title for both, and matching the body of the published content with the description of the video content can include having the same body of the published content with the description of the video content.

[0247] by Figure 3U For example, Figure 3U This is a schematic diagram of the content editing interface provided in the embodiments of this application. The audio browsing interface 329 may also include a third confirmation control 329a, namely "Next". The target object can click the third confirmation control 329a, so that the client can respond to the trigger operation of the third confirmation control 329a, that is, respond to the confirmation operation of the first video, and display the content editing interface 330. The content editing interface 330 includes the target audio and video 330a, the content title 330f, and the content body 330g. The target object can edit the content to be published in the content editing interface 330 and publish the content to be published after the editing is completed.

[0248] Optionally, the content editing interface may also include a video playback control, which can respond to a trigger operation on the video playback control to play the first video. Optionally, the video playback control may be displayed in an associated position with the video cover of the first video. Figure 3U For example, the content editing interface 330 can also include a video playback control 330c, and the client can respond to the trigger operation of the video playback control 330c to play the first video.

[0249] Optionally, the content editing interface may also include a second return control, which can respond to a trigger operation on the second return control, save the content to be published, and display the parent page of the content editing interface. The parent page of the content editing interface can be an audio browsing interface. Figure 3U For example, the content editing interface 330 may also include a second return control 330e. The client can respond to the trigger operation of the second return control 330e, save the content to be published, and display the parent page of the content editing interface, namely the audio browsing interface 329.

[0250] Optionally, the content editing interface may also include a cover selection control, which can display at least one candidate video cover in response to a trigger operation on the cover selection control, and determine the first video cover for the first video in response to a trigger operation on the first video cover, with the content editing interface including the first video cover. Optionally, the cover selection control may be displayed in an associated position with the video cover of the first video.

[0251] by Figure 3U For example, the content editing interface 330 can also include a cover selection control 330b, which allows the target object to select and confirm the first video cover of the first video.

[0252] Optionally, the content editing interface may also include a publishing control. In response to a publishing operation on the content to be published, generating and publishing the content containing the first video can be done in the following way: Responding to a trigger operation on the publishing control, generating and publishing the content containing the first video. Figure 3U For example, the content editing interface 330 can also include a publishing control 330d. The client can respond to the trigger operation of the publishing control 330d, generate content to be published containing the first video, and publish the content to be published.

[0253] In this embodiment of the application, when a user publishes content including a first video, the target audio and video, the title of the published content, and the body of the published content can be displayed in the content editing interface. The title and body of the published content are automatically obtained based on the video theme, which can reduce the time required for users to edit and publish content and improve the efficiency of users' content publishing.

[0254] Optionally, after publishing content to be published, a content information block of the content to be published can be displayed on the personal page of the recipient of the content. Any recipient can click on the content information block to display the content details page of the content to be published, which includes the detailed content information of the content to be published. Optionally, after publishing content to be published, a content information block of the content to be published can be displayed on the content recommendation page of any recipient. Any recipient can click on the content information block to display the content details page of the content to be published, which includes the detailed content information of the content to be published. Further optional, if the recipient of the content to be published has set their content to be private, then only that recipient can view the content to be published. The publishing recipient and the target recipient are the same person.

[0255] Optionally, the content editing interface may include a data input area. In response to input operations performed on the data input area of ​​the content editing interface, the input data is displayed, and the content to be published is generated. Input operations on the data input area may include, but are not limited to: directly pulling up the keyboard to edit input in the data input area, pasting in the data input area, importing in the data input area, and voice input. Optionally, the input data includes at least the title and body of the content to be published; that is, the target can modify or supplement the title and body of the content in the content editing interface. Therefore, this application provides diverse input methods, which can improve the convenience of data input and the efficiency of publishing content.

[0256] Optionally, in response to an input operation in the data input area, the input data can be displayed in the following way: in response to a trigger operation in the data input area, the keyboard is pulled up in the information input interface to edit and input the input data, and the input data is displayed.

[0257] by Figure 3V For example, Figure 3V This is a schematic diagram of the content editing interface provided in this application embodiment. The content editing interface 331 includes a data input area 331a. The target object can click on the data input area 331a, so that the client can respond to the trigger operation on the data input area 331a and pull up the keyboard 332a in the content editing interface 332. That is, when the target object clicks on the data input area 332b, the keyboard changes from a hidden state to a pulled-up state. The target object can input data in the data input area 332b through the keyboard 332a. The client can respond to the input data in the data input area 332b and display the input data in the data input area 333a in the content editing interface 333.

[0258] Optionally, an automatic save function is also provided in the data input area. When input data exists in the data input area, the input data can be automatically saved, and a save prompt message can be displayed in the data input area. This save prompt message is used to indicate that the input data has been automatically saved. For example, a save prompt message "13:45 automatically saved" can be displayed in the data input area. Optionally, the input data in the data input area can be automatically saved according to a preset time interval, such as once per minute or once per second. This application embodiment does not limit this.

[0259] Optionally, when the number of text characters that can be entered in the data input area is limited, character prompts can also be displayed in the data input area. These prompts include the number of characters currently being entered and the maximum number of characters supported by the data input area. For example, if the data input area displays the character prompt "10002 / 10000", "10002" represents the number of characters currently being entered, and "10000" represents the maximum number of characters supported by the data input area. If the number of characters currently being entered exceeds the maximum number of characters, further character input is not possible. It should be understood that the embodiments of this application may limit the maximum number of characters supported by the data input area, or they may not limit the maximum number of characters supported by the data input area.

[0260] Optionally, the import operation in the data input area may include: an import control in the content editing interface, which can import input data in the content editing interface in response to a trigger operation on the import control, and display the imported input data in the data input area.

[0261] Optionally, voice input can be performed directly in the content editing interface, and the input voice can be converted into text. In response to the voice input operation in the content editing interface, the voice input data is displayed in the data input area of ​​the content editing interface, and the voice input data is used as input data.

[0262] Optionally, the content editing interface may also include a video cover for the target audio / video, which may include the first frame of the first video, the last frame of the first video, or any frame of the first video.

[0263] Optionally, in response to a trigger operation targeting the video cover of the target audio / video, at least one candidate video cover can be displayed; in response to a trigger operation targeting the first video cover, the first video cover of the target audio / video can be determined, and the content editing interface includes the first video cover. In other words, when a target user wants to change the video cover of the target audio / video, the target user can click on the video cover of the target audio / video, thus the client can display at least one candidate video cover, and the target user can then select the first video cover as the video cover of the target audio / video.

[0264] Optionally, the content editing interface may also include a cover selection control, which can display at least one candidate video cover in response to a trigger operation on the cover selection control, and determine the first video cover of the target audio / video in response to a trigger operation on the first video cover, with the content editing interface including the first video cover. Optionally, the cover selection control may be displayed in an associated position with the video cover of the target audio / video.

[0265] Optionally, the content editing interface may also include a video playback control, which can play the target audio / video in response to a trigger operation on the video playback control. Optionally, the video playback control may be displayed in an associated position on the video cover of the target audio / video.

[0266] Optionally, the content editing interface may also include a publishing control. In response to a publishing operation on the content to be published, generating and publishing the content to be published containing the first video can be done in the following way: in response to a trigger operation on the publishing control, generating and publishing the content to be published containing the first video.

[0267] Optionally, the content editing interface may also include a second return control, which can respond to a trigger operation on the second return control, save the content to be published, and display the parent page of the content editing interface. The parent page of the content editing interface can be an audio browsing interface.

[0268] Optionally, in response to a trigger operation on the second return control, the content clearing option and the save option can be displayed; in response to a trigger operation on the content clearing option, the content to be published can be deleted, and the previous page of the published content editing interface can be displayed; or, in response to a trigger operation on the save option, the content to be published can be saved, and the previous page of the published content editing interface can be displayed.

[0269] Optionally, the content editing interface can also provide an auto-save function, which can automatically save the input data when the content editing interface includes input data.

[0270] Optionally, after saving the content to be published, the personal page includes an entry point for displaying the content. Specifically, after the client saves the content to be published, the target audience can click to enter the personal page, which displays an entry point for the content to be published, for example, the entry point is "Drafts". When the target audience clicks on the drafts, the content information block of the content to be published will be displayed. The client responds to the trigger operation on the content information block of the content to be published and displays the content editing interface.

[0271] In this embodiment, a video generation interface is displayed, which includes an audio synthesis access point. The video generation interface is used to generate a first video based on at least one media material. The target speech of the target object is obtained through the audio synthesis access point. The target speech has a target timbre. The target speech is used to perform speech synthesis processing on the target text to obtain synthesized speech with the target timbre. Then, the target audio and video are displayed. The target audio and video are obtained by merging the first video and the synthesized speech, which can provide more possibilities for video editing and make it easier for users to directly use cloned timbres during the video editing process. This avoids switching from the video generation interface to a separate timbre cloning entry point, which would interrupt the video editing process. Therefore, this embodiment can clone the timbres of the target object during the video editing process, thereby improving the efficiency of audio and video generation.

[0272] Based on the above description, please refer to Figure 4 , Figure 4 This is a flowchart illustrating another video generation method provided in an embodiment of this application. This video generation method can be applied to both clients and servers, such as... Figure 4 The video generation method shown includes, but is not limited to, steps S401 to S407, wherein: S401. The client displays the video generation interface, which includes an audio synthesis access point.

[0273] The video generation interface may include a reasoning process for generating a first video based on at least one media material. The reasoning process includes at least one of the following reasoning information: highlight clips in at least one media material, content description of at least one media material, video theme of the first video, video content description of the first video, and video content summary of the first video.

[0274] Optionally, if at least one media asset includes video and image assets, the server can extract frames from the video asset to obtain extracted frame assets. Semantic understanding can then be performed on the extracted frame assets and image assets to obtain a content description for at least one media asset. Optionally, a large-scale graph-text model can be used to perform semantic understanding on the extracted frame assets and image assets to obtain a content description for at least one media asset. The large-scale graph-text model can be obtained through supervised training. Optionally, the video asset can be extracted at a preset frame extraction frequency.

[0275] Optionally, the content description of at least one media asset may include at least one of the following: role information of the media asset, scene information of the media asset, and event information of the media asset. The role information of the media asset is the "core subject" returned in the content description; the scene information of the media asset is the "main scene type" returned in the content description; and the event information of the media asset is the "key action" returned in the content description.

[0276] Optionally, the video theme of the first video can be determined based on the character information, scene information, and event information of the media material.

[0277] Optionally, if at least one media material includes video material, then highlight recognition can be performed on the video material to identify highlight segments of at least one media material. Semantic understanding can then be performed on these highlight segments to obtain segment descriptions. Alternatively, a large-scale image-text model can be used to perform semantic understanding on the highlight segments to obtain segment descriptions. This large-scale image-text model can be obtained through supervised training.

[0278] Optionally, when determining at least one highlight segment of a media material, scene filtering can be performed on the highlight segment of the at least one media material to filter out highlight segments containing redundant scenes, thereby improving the segment quality of the finally determined highlight segment.

[0279] Optionally, when performing semantic understanding on highlight segments, the semantic understanding of highlight segments can be combined with the content description of at least one media material, thereby obtaining a more accurate segment description of the highlight segments.

[0280] Optionally, the content description of at least one media asset is used to describe at least one media asset as a whole; the segment description of the highlight segment is used to describe the highlight segment. The content description includes the characters, scenes, and events in at least one media asset, as well as the interactions and contextual relationships between the characters, scenes, and events. The highlight segment includes the characters, scenes, and events in the highlight segment, as well as the interactions and contextual relationships between the characters, scenes, and events.

[0281] Optionally, highlight recognition of video footage to determine at least one highlight segment of the media footage can be performed as follows: The video footage is sampled to obtain multiple sample frames; the attribute information of each sample frame is determined; and based on the video footage, the attribute information of the multiple sample frames, the target highlight template, and the highlight sample frames corresponding to multiple highlight shots, at least one highlight segment of the media footage is determined. The attribute information of the sample frames includes a semantic feature vector, which indicates the image semantic information of the sample frame.

[0282] Optionally, based on the content description of at least one media material and the fragment description of the highlight fragment, the narrative of each highlight fragment and its corresponding fragment description can be rearranged to obtain the rearranged fragment descriptions of each highlight fragment and its corresponding fragment description, thereby giving the final generated first video a sense of story and improving the quality of the generated video.

[0283] Optionally, the target large language model may include a target reordering large language model. In this case, the content description of at least one media material, along with the segment description of highlight segments and corresponding reordering instructions, can be input into the target reordering large language model. The target reordering large language model can then output a narrative order, as well as the merging and / or sorting results of highlight segments based on the narrative order. Further, each highlight segment can be reordered according to its segment description to obtain the reordered highlight segments. Optionally, the content description of at least one media material, the segment description of highlight segments, and corresponding reordering instructions can be input into the target reordering large language model. The target reordering large language model can be obtained through supervised training. The segment order of the sorted highlight segments corresponds to the narrative order. Each sorted highlight segment has a corresponding segment number and segment description; the segment number indicates the segment order of the highlight segment.

[0284] Optionally, the target large language model may include a target text generation large language model. In this case, the segment description of the highlight clip, the narrative order output by the target rearrangement large language model, the merging and / or sorting results of the highlight clips based on the narrative order, and the text generation instructions can be input into the target text generation large language model, so that the target text generation large language model outputs the video text of the first video and the video tags of the first video.

[0285] The video tags can include at least one of the following: video title tag, video text tag, video music tag, video voice-over tag, and video style tag. Further, the video content summary of the first video can be determined based on its video tags. The video content summary includes at least one of the following: video title, video text, video music, video voice-over, and video style. Optionally, the video content summary of the first video can be determined by arranging and combining the video tags. Optionally, the video content summary of the first video can be determined by mapping the video tags to text. For example, the video content summary could be "《An Unforgettable Trip》, daily narrative text, paired with relaxing music, and a simple, everyday style," where the video title is "《An Unforgettable Trip》"; the video text is "daily narrative text"; the video music is "paired with relaxing music"; and the video style is "simple, everyday style packaging."

[0286] Optionally, the video content description can be determined by generating the video script of the first video output by the large language model based on the target script, and the merged and / or sorted results of the highlight segments output by the target reordering large language model. The video content description includes at least one of the following: story title or theme, narrative order, and content description.

[0287] Optionally, the story title or theme can be obtained from the video script of the first video output by the target text generation large language model.

[0288] Optionally, the narrative order can be obtained from the merged and / or sorted results of the highlight segments output by the target reordering large language model. Optionally, the merged and / or sorted results of the highlight segments include the segment number of each highlight segment, and the segment number of each highlight segment is removed when obtaining the narrative order.

[0289] Optionally, the content description can be obtained from the video script of the first video output by the large language model that generates the target text. Optionally, the text in the video script in a preset order can be used as the content description. For example, the first two sentences of the video script can be used as the content description.

[0290] Optionally, if there are multiple first videos, the video content description for each first video can be obtained from the corresponding video script and the merged and / or sorted results of highlight clips. Optionally, each first video may have a different video style.

[0291] Optionally, if the server receives an update instruction from the target object, it can rearrange the narrative of each highlight segment and its corresponding segment description based on the update instruction, the content description of at least one media material, and the segment description of the highlight segment. This can make the final generated first video have a sense of story and meet the video generation requirements of the object, thereby improving the quality of the generated video.

[0292] Optionally, the target large language model may include a target reordering large language model. In this case, the content description of at least one media material, update instructions, corresponding reordering instructions, and fragment description of highlight segments can be input into the target reordering large language model. Thus, the target reordering large language model can output the narrative order, as well as the merging and / or sorting results of highlight segments based on the narrative order.

[0293] S402: The client obtains the target speech of the target object through the audio synthesis access point. The target speech has the target timbre.

[0294] The specific implementation of step S402 in this embodiment can be found in the description of step S202 above, and will not be repeated here.

[0295] S403, The client sends the target voice to the server.

[0296] S404. The server performs speech synthesis processing on the target text based on the target speech to obtain synthesized speech with the target timbre.

[0297] The server can extract features from the target speech, obtaining the target object's voiceprint features, speech dynamics features, and emotional features. Then, using these features, it performs text-to-speech processing on the target text to obtain synthesized speech with the target timbre. Optionally, the server can use a neural network model to extract features from the target speech, and then use the extracted features to synthesize speech on the target text to obtain synthesized speech with the target timbre. Neural network models include generative adversarial networks, variational autoencoders, or autoregressive models.

[0298] S405 The server merges the first video and the synthesized speech to obtain the target audio and video.

[0299] S406. The server sends the target audio and video to the client.

[0300] S407, The client displays the target audio and video.

[0301] In this embodiment, the client displays a video generation interface, which includes an audio synthesis access point. The client obtains the target speech of the target object through the audio synthesis access point. The target speech has a target timbre. The client sends the target speech to the server. The server performs speech synthesis processing on the target text based on the target speech to obtain synthesized speech with the target timbre. The server merges the first video and the synthesized speech to obtain the target audio and video. The server sends the target audio and video to the client, and the client displays the target audio and video, which can improve the efficiency of audio and video generation.

[0302] This application also provides a computer storage medium storing program instructions, which, when executed, are used to implement the corresponding methods described in the above embodiments.

[0303] This application provides a computer program product, which includes a computer program stored in a computer storage medium. The processor of a computer device reads the computer program from the computer storage medium and executes the computer program, causing the computer device to perform the corresponding methods described in the above embodiments.

[0304] See also Figure 5 , Figure 5 This is a schematic diagram of the structure of a video generation device provided in an embodiment of this application.

[0305] In one implementation of the video generation apparatus of this application, the video generation apparatus includes the following structure.

[0306] Display unit 501 is used to display a video generation interface, the video generation interface including an audio synthesis access point; wherein, the video generation interface is used to generate a first video based on at least one media material; The acquisition unit 502 is used to acquire the target speech of the target object via the audio synthesis access port; wherein the target speech has a target timbre, and the target speech is used to perform speech synthesis processing on the target text to obtain synthesized speech with the target timbre; The display unit 501 is also used to display target audio and video, which is obtained by merging the first video and the synthesized speech.

[0307] In one embodiment, the acquisition unit 502 acquires the target speech of the target object via the audio synthesis access port, including: The target speech is acquired via the audio synthesis access point; or... At least one candidate speech of the target object is displayed via the audio synthesis access portal; in response to a selection operation of the at least one candidate speech, the target speech is determined; or... The target speech is determined based on the at least one media material via the audio synthesis access point.

[0308] In one embodiment, the acquisition unit 502 acquires the target speech via the audio synthesis access point, including: The reference text and voice acquisition entry are displayed via the audio synthesis access point; The target speech output by the target object in response to the reference text is acquired via the speech acquisition port.

[0309] In one embodiment, the voice acquisition port includes a voice input control; the acquisition unit 502 acquires the target voice output by the target object in response to the reference text via the voice acquisition port, including: In response to receiving and continuously receiving a voice input operation, the system acquires user-inputted voice during the duration of the voice input operation and displays a first voice operation area, wherein the voice input operation includes performing a first operation on the voice input control; in response to a voice input lock operation, the system continues to acquire user-inputted voice, wherein the voice input lock operation includes performing a second operation and stopping the first operation during the duration of the voice input operation; in response to a voice input pause operation, the system pauses receiving user-inputted voice and displays a second voice operation area instead of the first voice operation area; the second voice operation area is used to transmit the acquired target voice; or, In response to receiving and continuously receiving a voice input operation, the system acquires user-inputted voice during the duration of the voice input operation, the voice input operation including performing a first operation on the voice input control; in response to a voice input lock operation, the system continues to acquire user-inputted voice, the voice input lock operation including receiving a second operation and stopping the first operation during the duration of the voice input operation; in response to a voice transmission trigger event, the system transmits the acquired target voice; or, In response to receiving and continuously receiving a voice input operation, the system acquires the user's voice input during the duration of the voice input operation; in response to a voice transmission trigger event, the system transmits the acquired target voice.

[0310] In one embodiment, the display unit 501 is further configured to display a voice transmission prompt message and a cancel transmission control after the acquisition unit acquires the target voice through the audio synthesis access port; the voice transmission prompt message is used to indicate that the target voice is being transmitted. The audio and video generation device further includes a processing unit 503, which is used to cancel the transmission of the target voice in response to a trigger operation of the cancel send control.

[0311] In one embodiment, the display unit 501 is further configured to display speech synthesis prompt information and a cancel synthesis control after the acquisition unit 502 acquires the target speech of the target object via the audio synthesis access port; the speech synthesis prompt information is used to indicate that the target speech is being processed by speech synthesis. The audio and video generation device further includes a processing unit 503, which is used to cancel the speech synthesis processing of the target speech in response to the triggering operation of the cancel synthesis control.

[0312] In one embodiment, the display unit 501 is further configured to display at least one synthesized speech having the target timbre; The audio and video generation device further includes a playback unit 504, which is used to play the target synthesized speech in response to a playback operation on the target synthesized speech, wherein the target synthesized speech is any one of the at least one synthesized speech; In response to the selection operation of the target synthesized speech, the display unit 501 is triggered to display the target audio and video, which is obtained by merging the first video and the target synthesized speech.

[0313] In this embodiment, the display unit 501 displays a video generation interface, which includes an audio synthesis access point; wherein, the video generation interface is used to generate a first video based on at least one media material; the acquisition unit 502 acquires the target speech of the target object via the audio synthesis access point; wherein, the target speech has a target timbre, and the target speech is used to perform speech synthesis processing on the target text to obtain synthesized speech with the target timbre; the display unit 501 displays the target audio and video, which is obtained by merging the first video and the synthesized speech, thereby improving the efficiency of audio and video generation.

[0314] See also Figure 6 , Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device in this embodiment includes a power supply module and other structures, and includes a processor 601, a memory 602, and a communication interface 603. The processor 601, the memory 602, and the communication interface 603 can exchange data, and the processor 601 implements the corresponding video generation method.

[0315] The memory 602 may include volatile memory, such as random-access memory (RAM); the memory 602 may also include non-volatile memory, such as flash memory, solid-state drive (SSD), etc.; the memory 602 may also include a combination of the above types of memory.

[0316] Processor 601 may be a central processing unit (CPU). Processor 601 may also be a combination of a CPU and a GPU. In a computer device, multiple CPUs and GPUs may be included as needed for corresponding video generation. In one embodiment, memory 602 is used to store program instructions. Processor 601 can invoke program instructions to implement the various methods described above in the embodiments of this application.

[0317] The communication interface 603 may include a display screen, microphone, or speaker, etc.

[0318] In one possible implementation, the processor 601 of the computer device calls program instructions stored in memory 602 to display a video generation interface via communication interface 603, the video generation interface including an audio synthesis access point; wherein, the video generation interface is used to generate a first video based on at least one media material; to obtain target speech of a target object via the audio synthesis access point; wherein, the target speech has a target timbre, the target speech is used to perform speech synthesis processing on target text to obtain synthesized speech with the target timbre; and to display target audio and video, the target audio and video being obtained by merging the first video and the synthesized speech.

[0319] In one embodiment, the processor 601 obtains the target speech of the target object via the audio synthesis access port, including: The target speech is acquired via the audio synthesis access point; or... At least one candidate speech of the target object is displayed via the audio synthesis access portal; in response to a selection operation of the at least one candidate speech, the target speech is determined; or... The target speech is determined based on the at least one media material via the audio synthesis access point.

[0320] In one embodiment, the processor 601 acquires the target speech via the audio synthesis access point, including: The reference text and voice acquisition entry are displayed via the audio synthesis access point; The target speech output by the target object in response to the reference text is acquired via the speech acquisition port.

[0321] In one embodiment, the voice acquisition port includes a voice input control; the processor 601 acquires the target voice output by the target object in response to the reference text via the voice acquisition port, including: In response to receiving and continuously receiving a voice input operation, the system acquires user-inputted voice during the duration of the voice input operation and displays a first voice operation area, wherein the voice input operation includes performing a first operation on the voice input control; in response to a voice input lock operation, the system continues to acquire user-inputted voice, wherein the voice input lock operation includes performing a second operation and stopping the first operation during the duration of the voice input operation; in response to a voice input pause operation, the system pauses receiving user-inputted voice and displays a second voice operation area instead of the first voice operation area; the second voice operation area is used to transmit the acquired target voice; or, In response to receiving and continuously receiving a voice input operation, the system acquires user-inputted voice during the duration of the voice input operation, the voice input operation including performing a first operation on the voice input control; in response to a voice input lock operation, the system continues to acquire user-inputted voice, the voice input lock operation including receiving a second operation and stopping the first operation during the duration of the voice input operation; in response to a voice transmission trigger event, the system transmits the acquired target voice; or, In response to receiving and continuously receiving a voice input operation, the system acquires the user's voice input during the duration of the voice input operation; in response to a voice transmission trigger event, the system transmits the acquired target voice.

[0322] In one embodiment, the processor 601 is further configured to perform the following operations: After the target speech is acquired via the audio synthesis access point, a speech transmission prompt and a cancel send control are displayed; the speech transmission prompt is used to indicate that the target speech is being sent. In response to the triggering operation of the cancel send control, the sending of the target voice is canceled.

[0323] In one embodiment, the processor 601 is further configured to perform the following operations: After obtaining the target speech of the target object through the audio synthesis access portal, a speech synthesis prompt message and a cancel synthesis control are displayed; the speech synthesis prompt message is used to indicate that the target speech is being processed by speech synthesis. In response to the triggering operation of the cancel synthesis control, the speech synthesis processing of the target speech is canceled.

[0324] In one embodiment, the processor 601 is further configured to perform the following operations: Display at least one synthesized speech having the target timbre; In response to a playback operation on a target synthesized speech, the target synthesized speech is played, wherein the target synthesized speech is any one of the at least one synthesized speech; In response to the selection operation of the target synthesized speech, the display unit is triggered to display the target audio and video, which is obtained by merging the first video and the target synthesized speech.

[0325] In this embodiment, the processor 601 displays a video generation interface, which includes an audio synthesis access point. The video generation interface is used to generate a first video based on at least one media material. Target speech of a target object is obtained via the audio synthesis access point. The target speech has a target timbre and is used to perform speech synthesis processing on target text to obtain synthesized speech with the target timbre. Target audio and video are displayed, which are obtained by merging the first video and the synthesized speech, thereby improving the efficiency of audio and video generation.

[0326] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

[0327] The above-disclosed embodiments are merely some of the embodiments of this application, and should not be construed as limiting the scope of this application. Those skilled in the art can understand that all or part of the processes for implementing the above embodiments, and equivalent changes made in accordance with the claims of this application, still fall within the scope of this application.

Claims

1. A method for generating audio and video, characterized in that, The method includes: The video generation interface is displayed, and the video generation interface includes an audio synthesis access point; wherein, the video generation interface is used to generate a first video based on at least one media material; The target speech of the target object is obtained through the audio synthesis access portal; wherein, the target speech has a target timbre, and the target speech is used to perform speech synthesis processing on the target text to obtain synthesized speech with the target timbre; The target audio and video are displayed, which are obtained by merging the first video and the synthesized speech.

2. The method as described in claim 1, characterized in that, The process of obtaining the target speech of the target object via the audio synthesis access point includes: The target speech is acquired via the audio synthesis access point; or... At least one candidate speech of the target object is displayed via the audio synthesis access portal; in response to a selection operation of the at least one candidate speech, the target speech is determined; or... The target speech is determined based on the at least one media material via the audio synthesis access point.

3. The method as described in claim 2, characterized in that, The acquisition of the target speech via the audio synthesis access point includes: The reference text and voice acquisition entry are displayed via the audio synthesis access point; The target speech output by the target object in response to the reference text is acquired via the speech acquisition port.

4. The method as described in claim 3, characterized in that, The voice acquisition entry point includes a voice input control; the acquisition of the target voice output by the target object in response to the reference text via the voice acquisition entry point includes: In response to receiving and continuously receiving a voice input operation, the system acquires user-inputted voice during the duration of the voice input operation and displays a first voice operation area, wherein the voice input operation includes performing a first operation on the voice input control; in response to a voice input lock operation, the system continues to acquire user-inputted voice, wherein the voice input lock operation includes performing a second operation and stopping the first operation during the duration of the voice input operation; in response to a voice input pause operation, the system pauses receiving user-inputted voice and displays a second voice operation area instead of the first voice operation area; the second voice operation area is used to transmit the acquired target voice; or, In response to receiving and continuously receiving a voice input operation, the system acquires user-inputted voice during the duration of the voice input operation, the voice input operation including performing a first operation on the voice input control; in response to a voice input lock operation, the system continues to acquire user-inputted voice, the voice input lock operation including receiving a second operation and stopping the first operation during the duration of the voice input operation; in response to a voice transmission trigger event, the system transmits the acquired target voice; or, In response to receiving and continuously receiving a voice input operation, the system acquires the user's voice input during the duration of the voice input operation; in response to a voice transmission trigger event, the system transmits the acquired target voice.

5. The method as described in claim 2, characterized in that, After acquiring the target speech via the audio synthesis access point, the process further includes: The system displays a voice transmission prompt and a cancel send control; the voice transmission prompt is used to indicate that the target voice is being sent. In response to the triggering operation of the cancel send control, the sending of the target voice is canceled.

6. The method as described in claim 1, characterized in that, After obtaining the target speech of the target object via the audio synthesis access point, the process further includes: Displays a speech synthesis prompt and a cancel speech synthesis control; the speech synthesis prompt is used to indicate that the target speech is being processed by speech synthesis. In response to the triggering operation of the cancel synthesis control, the speech synthesis processing of the target speech is canceled.

7. The method as described in claim 1, characterized in that, The method further includes: Display at least one synthesized speech having the target timbre; In response to a playback operation on a target synthesized speech, the target synthesized speech is played, wherein the target synthesized speech is any one of the at least one synthesized speech; In response to the selection operation of the target synthesized speech, the target audio and video are triggered to be displayed, which are obtained by merging the first video and the target synthesized speech.

8. An audio / video generation device, characterized in that, The device includes: A display unit is used to display a video generation interface, which includes an audio synthesis access point; wherein, the video generation interface is used to generate a first video based on at least one media material; The acquisition unit is used to acquire the target speech of the target object via the audio synthesis access portal; wherein the target speech has a target timbre, and the target speech is used to perform speech synthesis processing on the target text to obtain synthesized speech with the target timbre; The display unit is also used to display target audio and video, which is obtained by merging the first video and the synthesized speech.

9. A computer device, characterized in that, The computer device includes a memory, a communication interface, and a processor, wherein the memory, the communication interface, and the processor are interconnected; the memory stores a computer program, and the processor calls the computer program stored in the memory to implement the audio and video generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the audio / video generation method as described in any one of claims 1 to 7.

11. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer storage medium; the processor of the computer device reads the computer program from the computer storage medium, and the processor executes the computer program, causing the computer device to perform the audio and video generation method as described in any one of claims 1 to 7.