Video generation method, video processing method, apparatus, device, and storage medium

By generating initial prompt text and object action-driven data using a large model, and combining this with existing materials to generate an initial video, and then processing and adjusting the video, this technology solves the problems of insufficient controllability, short duration, low quality, and lack of personalization in existing video generation, achieving efficient and personalized video generation and processing.

CN119110134BActive Publication Date: 2026-08-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2024-09-18
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing technologies suffer from several drawbacks when generating videos: insufficient controllability and editability, short video duration and low quality, lack of personalized features, high hardware resource consumption, and difficulty in generating continuous high-quality videos in complex scenarios.

Method used

Initial prompt text is generated by using a large model. Combined with object action-driven data and materials, an initial video is generated. Finally, video processing methods are used to adjust the material attributes to improve video quality and personalization.

Benefits of technology

It reduces the manpower and time costs of video generation, improves the controllability and quality of video generation, shortens the production cycle, and enhances the personalized features of videos and their performance in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119110134B_ABST
    Figure CN119110134B_ABST
Patent Text Reader

Abstract

The present disclosure provides a video generation method, relates to the technical field of artificial intelligence, and particularly relates to the technical fields of large models, video processing and virtual digital humans. A specific implementation scheme is as follows: according to an initial text input by a user, a plurality of initial prompt texts are determined, wherein the plurality of initial prompt texts include an initial content prompt text and an initial material prompt text; according to the initial content prompt text, a video content text and at least one initial object action driving data corresponding to the video content text are determined; and according to the at least one initial object action driving data and at least one initial material corresponding to the at least one initial material prompt text, an initial video is generated. The present disclosure also provides a video processing method and device, an electronic device and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of large-scale models, video processing, and virtual digital humans, and can be applied to various video production scenarios such as social media videos, marketing videos, educational training videos, news report videos, entertainment videos, e-commerce videos, and stylized animated videos. More specifically, this disclosure provides a video generation method, a video processing method, an apparatus, an electronic device, and a storage medium. Background Technology

[0002] With the development of artificial intelligence technology, the application scenarios of large-scale models are constantly expanding. Based on artificial intelligence technology, videos can be generated from user-input text. Summary of the Invention

[0003] This disclosure provides a video generation method, a video processing method, an apparatus, a device, and a storage medium.

[0004] According to one aspect of this disclosure, a video generation method is provided, the method comprising: determining a plurality of initial prompt texts based on initial text input by a user, wherein the plurality of initial prompt texts include initial content prompt texts and initial material prompt texts; determining video content texts and at least one initial object action-driven data corresponding to the video content texts based on the initial content prompt texts; and generating an initial video based on the at least one initial object action-driven data and at least one initial material corresponding to the at least one initial material prompt text.

[0005] According to another aspect of this disclosure, a video processing method is provided, the method comprising: determining at least one adjustment prompt text and at least one attribute adjustment information based on adjustment text corresponding to a video to be processed, wherein the video to be processed corresponds to at least one material to be adjusted; adjusting the attribute information of at least one material to be adjusted corresponding to at least one adjustment prompt text based on at least one attribute adjustment information corresponding to at least one adjustment prompt text, thereby obtaining at least one adjusted material; and obtaining a processed video based on at least one adjusted material.

[0006] According to another aspect of this disclosure, a video generation apparatus is provided, the apparatus comprising: a first determining module, configured to determine a plurality of initial prompt texts based on initial text input by a user, wherein the plurality of initial prompt texts include initial content prompt texts and initial material prompt texts; a second determining module, configured to determine video content texts and at least one initial object motion driving data corresponding to the video content texts based on the initial content prompt texts; and a generation module, configured to generate an initial video based on the at least one initial object motion driving data and at least one initial material corresponding to the at least one initial material prompt text.

[0007] According to another aspect of this disclosure, a video processing apparatus is provided, the apparatus comprising: a third determining module, configured to determine at least one adjustment prompt text and at least one attribute adjustment information based on adjustment text corresponding to a video to be processed, wherein the video to be processed corresponds to at least one material to be adjusted; an adjusting module, configured to adjust the attribute information of at least one material to be adjusted corresponding to at least one adjustment prompt text based on at least one attribute adjustment information corresponding to at least one adjustment prompt text, thereby obtaining at least one adjusted material; and an obtaining module, configured to obtain a processed video based on at least one adjusted material.

[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method provided according to this disclosure.

[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods provided according to this disclosure.

[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method provided according to this disclosure.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0013] Figure 1 This is an exemplary system architecture diagram of a video generation method, video processing method and apparatus applicable according to an embodiment of the present disclosure;

[0014] Figure 2 This is a flowchart of a video generation method according to an embodiment of the present disclosure;

[0015] Figure 3 This is a schematic diagram of script data according to one embodiment of the present disclosure;

[0016] Figure 4 This is a schematic diagram of multiple tasks corresponding to multiple prompt texts according to an embodiment of the present disclosure;

[0017] Figure 5 This is a schematic diagram of a video generation method according to an embodiment of the present disclosure;

[0018] Figure 6A This is a schematic diagram of a visual interface according to an embodiment of the present disclosure;

[0019] Figure 6B and Figure 6C This is a schematic diagram of different video frames of an initial video according to an embodiment of the present disclosure;

[0020] Figure 7 This is a flowchart of a video processing method according to another embodiment of the present disclosure;

[0021] Figure 8 This is a block diagram of a video generation apparatus according to an embodiment of the present disclosure;

[0022] Figure 9 This is a block diagram of a video processing apparatus according to another embodiment of the present disclosure; and

[0023] Figure 10 This is a block diagram of an electronic device to which video generation methods and / or video processing methods can be applied, according to an embodiment of the present disclosure. Detailed Implementation

[0024] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0025] To design and produce a beautiful and compliant video, one or more designers, directors, and actors with excellent aesthetic sense and rich experience need to spend a lot of time to complete it. This results in high labor costs, especially for professionals, in the video production process, as well as high time costs.

[0026] In some embodiments, a conversational large model can be used to generate videos based on natural language input from the user.

[0027] However, when using large models to generate videos, the controllability and editability of the videos are insufficient. When generating complex scenes, large models (such as Sora) may struggle to accurately simulate physical behavior, resulting in videos that lack detail. For example, interactions between objects may appear unnatural or fail to follow the laws of physics. Even if some large models (such as Runwaygen3) can generate high-quality videos, they still struggle to handle complex interactions between characters and objects, making it difficult for the generated videos to meet user needs.

[0028] Furthermore, when using artificial intelligence (AI) technology to generate videos, the resulting videos tend to be short. As the length of the generated videos increases, so do the errors and inconsistencies. For example, generating a video that is several minutes long using AI technology may require a long generation time, and the resulting video quality is also low, making AI-based video generation technology unsuitable for generating long videos.

[0029] Furthermore, in the process of generating videos using artificial intelligence technology, general data can be used for automated video generation, resulting in videos that lack personalized features and have weak recognizability, making it difficult to meet users' personalized needs.

[0030] Furthermore, the maturity and stability of AI-based video generation technologies are relatively low, exhibiting significant performance variations across different scenarios, particularly struggling to generate continuous, high-quality videos in complex environments. While AI-based video generation technologies have brought numerous conveniences and possibilities to video creation, their controllability, editability, compliance, and personalization levels still need improvement, and the significant hardware resource consumption necessitates further optimization.

[0031] Therefore, in order to generate videos efficiently, this disclosure provides a video generation method and a video processing method, and the system architecture of applying this method will be described below.

[0032] Figure 1 This is a schematic diagram of an exemplary system architecture for applying a video generation method, video processing method, and apparatus according to an embodiment of this disclosure. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0033] like Figure 1As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0034] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0035] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0036] It should be noted that the video generation method and video processing method provided in this disclosure embodiment can generally be executed by server 105. Correspondingly, the video generation apparatus and video processing apparatus provided in this disclosure embodiment can generally be located in server 105. The video generation method and video processing method provided in this disclosure embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the video generation apparatus and video processing apparatus provided in this disclosure embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.

[0037] As can be understood, the system architecture of this disclosure has been described above, and the method of this disclosure will be described below.

[0038] Figure 2 This is a flowchart of a video generation method according to an embodiment of the present disclosure.

[0039] like Figure 2 As shown, the method 200 may include operations S210 to S230.

[0040] In operation S210, multiple initial prompt texts are determined based on the initial text entered by the user.

[0041] In this embodiment of the disclosure, the initial prompt text can be determined using various methods. For example, the initial text can be segmented into words. Based on the result of word segmentation, the initial prompt text is determined. If the initial text is "A host is introducing the news," the initial prompt text can be "host" and "news."

[0042] In this embodiment of the disclosure, the multiple initial prompt texts include initial content prompt text and initial material prompt text. For example, the initial prompt text "News" can be the initial content prompt text. The initial prompt text "Host" can be the initial material prompt text.

[0043] In operation S220, based on the initial content prompt text, the video content text and at least one initial object action driving data corresponding to the video content text are determined.

[0044] In this embodiment of the disclosure, the initial object action driving data includes driving data corresponding to the object's lip movements. For example, based on the initial prompt text "news," one or more news articles from a certain day can be obtained as the video content text. Based on text-to-speech (TTS) technology, one or more lip movement data corresponding to the video content text can be obtained. Based on this one or more lip movement data, one or more initial object action driving data can be determined.

[0045] In operation S230, an initial video is generated based on at least one initial object action-driven data and at least one initial material corresponding to at least one initial material prompt text.

[0046] For example, the initial material corresponding to the initial prompt text "Host" can be a virtual avatar. This virtual avatar includes a head and lips. Using the initial object motion-driven data, the virtual avatar's lips can be driven to perform one or more lip movements to achieve one or more of the aforementioned lip movements, resulting in one or more video frames. Based on these video frames, the initial video can be obtained.

[0047] This disclosure utilizes initial prompt text to achieve video generation, reducing the labor and time costs required for video generation. Generating video content text using initial prompt text and determining action-driven data ensures greater consistency and coordination between the video text content and the actions displayed by objects in the video, thereby improving video quality.

[0048] Furthermore, through the embodiments of this disclosure, when the scene in the video is a three-dimensional scene, production costs can be significantly reduced. With the necessary materials for video generation prepared, videos with three-dimensional scenes can be quickly generated through text descriptions. For low- to medium-requirement projects requiring short delivery cycles, such as virtual human live streams and promotional videos, the production and delivery cycle can be significantly shortened to within a few days. In addition, 3D scene designers can focus their efforts on material optimization and overall scene design, minimizing repetitive work and improving overall video quality and generation efficiency.

[0049] As can be understood, the method of this disclosure has been described above, and the initial prompt text of this disclosure will be described below.

[0050] In some embodiments of operation S210 described above, determining multiple initial prompt texts based on the initial text input by the user includes: determining initial script data based on the initial text and the user's attribute information; and determining multiple initial prompt texts based on the initial script data.

[0051] In this embodiment of the disclosure, the user's attribute information includes information such as the user's industry and the actual application scenario. For example, when a user uses an artificial intelligence product for the first time, they can enter the text "Help me generate a video using this digital human" in the input box of the product's visual interface. Then, after the user authorizes the request, the user's authorization attribute information can be obtained.

[0052] In this embodiment, multiple initial prompt texts can be determined using a large model based on the initial text and user attribute information. The large model can be fine-tuned using multiple sample texts and multiple preset prompt texts. The initial prompt texts are determined from multiple preset prompt texts using the large model. The large model can be a Large Language Model (LLM). The preset prompt texts can be standardized prompt texts. Sample texts can be historical texts input by a user into the large model, historical texts input by multiple users with similar attributes, or texts with high similarity generated based on historical texts input by users; this disclosure does not impose any limitations on these. The large model can be a conversational large model such as Wenxin Yiyan. Through this embodiment, using a conversational large model, short natural language text prompts can be used to quickly generate videos based on already produced materials in a material production platform. The large model is fine-tuned using preset prompt texts, which can correspond to the identification text of the materials, so that the fine-tuned large model can quickly determine the material corresponding to the initial prompt text or the adjusted prompt text from multiple materials.

[0053] In this embodiment of the disclosure, the initial script data includes at least one of the following: initial script outline data, initial script storyboard data, and initial script reference frame data. The following will be combined with... Figure 3 Please provide an explanation.

[0054] Figure 3 This is a schematic diagram of script data according to one embodiment of the present disclosure.

[0055] like Figure 3 As shown, user30 can input the natural language text "A presenter is introducing the news" into the large model llm30. The large model llm30 can decompose the natural language text into multiple decomposition results. These multiple decomposition results can include "presenter," "news," etc. Based on these multiple decomposition results, the script structure is determined from the script structure library stru30 using the script framework fw30 corresponding to the user's attribute information. This script structure can include, for example, at least one of an outline, storyboard, and reference screen. Accordingly, the script data determined using the script structure can include at least one of the following: script outline data sc30, script storyboard description data b30, and script reference screen data p30.

[0056] The script outline data sc30 is equivalent to a script outline provided by the screenwriter. The script outline data sc30 can include the script outline text "A presenter is introducing the news".

[0057] The storyboard description data b30 corresponds to the storyboard description text "The camera moves from a wide shot to a close-up, then to a medium shot." This storyboard description text can correspond to common camera movement patterns used in news scenes. The storyboard description data can also correspond to at least one storyboard data point. For example... Figure 3 As shown, at least one storyboard data set may include storyboard data bd31, storyboard data bd32, storyboard data bd33, and storyboard data bd34, etc. Storyboard data bd31 may include at least one of scene description data scene31, lens pointer data lens31, action data action31, audio data audio31, lighting data light31, and duration data dur31. Scene description data scene31 may indicate scene type, surrounding layout, etc. Lens pointer data lens31 may indicate lens type and lens movement. Action data action31 may indicate head and body movements of objects. Audio data audio31 may indicate sound effects and background music. Lighting data light31 may indicate character lighting and scene lighting. Duration data dur31 may indicate the duration of a storyboard shot.

[0058] The script reference screen data p30 includes video size data, focal length data, depth of field data, and motion rhythm data. Motion rhythm data 31 indicates the speed at which the object performs the aforementioned limb movements. The script reference screen data p30 corresponds to the script reference screen description text such as "16:9 composition, overall smooth rhythm, etc."

[0059] It is understandable that, if the aforementioned natural language text is used as the initial text, the aforementioned script data can be used as the initial script data. The script outline data sc30, script storyboard description data b30, and script reference screen data p31 can be used as the initial script outline data, initial script storyboard description data, and initial script reference screen data, respectively. The initial script storyboard description data corresponds to at least one initial storyboard data, which includes at least one of the following: initial scene description data, initial camera direction data, initial action data, initial audio data, initial lighting data, and initial duration data. The initial script reference screen data includes at least one of the following: initial video size data, initial focal length data, and initial depth of field data.

[0060] Through the embodiments of this disclosure, a highly professional script (script outline, reference screen, storyboard, etc.) that matches user attributes can be generated using concise natural language text, so that video can be quickly generated based on the script.

[0061] As can be understood, the initial script data disclosed above has been explained; the following will combine... Figure 4 Descriptions for multiple initial prompt text lines.

[0062] Figure 4 This is a schematic diagram of multiple tasks corresponding to multiple prompt texts according to an embodiment of the present disclosure.

[0063] like Figure 4 As shown, user40 can provide natural language text40. The natural language text40 could be "A presenter is introducing the news".

[0064] In this embodiment of the disclosure, multiple initial prompt texts can be determined based on the initial script data. For example, based on the aforementioned script outline text, script storyboard description text, and script reference screen description text, script text st40 can be obtained by processing using a large model. Script text st40 could be something like, "A female anchor in a blue business suit stands in a sci-fi style studio, eloquently introducing the day's international news. The camera moves from a wide shot to a close-up and then to a medium shot. The total duration is 30 seconds, with a 16:9 aspect ratio and a generally slow pace." It can be understood that the aforementioned script text can be obtained by adding descriptive text related to scenes and objects to the script outline text. Based on this script text, multiple prompt texts can be determined. These multiple prompt texts can correspond to multiple tasks. These multiple tasks may include content generation task m401, motion generation task m402, object determination task m403, scene determination task m404, shot determination task m405, lighting determination task m406, and compositing output task m407.

[0065] In this embodiment of the disclosure, the multiple prompt texts may include content prompt text. The content prompt text may be "international current affairs news of the day". This content prompt text may serve as the initial content prompt text.

[0066] In this embodiment, a large model can be used to generate initial content text based on initial content prompt text. For example, to perform the content generation task m401, the large model can search for international current affairs news based on the date and generate content text tn40. Content text tn40 could be: "This afternoon, the lunar probe return capsule held an opening ceremony, marking a successful conclusion to the lunar exploration mission's engineering phase. According to a report yesterday by the broadcasting company of country A, the head of the health bureau of a certain country announced that there are currently no major public health events. The president of country B signed a decree approving unit C's decision regarding the use of drones; this decree has been published on the website of the presidential palace of country B." This content text can serve as the initial content text.

[0067] For example, after executing the action generation task m402, limb motion driving data and head motion driving data of the object can be obtained. The limb motion driving data can correspond to the object's overall movement, movements above the waist, foot movements, and hand movements. After executing the object determination task m403, for example, a digital human can be determined as the object. The object's limbs and head can be driven by the aforementioned limb motion driving data and head motion driving data, respectively, to perform relevant actions. After executing the scene determination task m404, the scene in which the object is located can be determined. To execute the shot determination task m405, shot description information can be determined based on the object's position. Shot description information is related to the camera's movement. The object's position includes the object's head, eyes, mouth, waist, left hand, right hand, left foot, right foot, left leg, and right leg.

[0068] The lighting determination task m406 can include one or more of the following tasks: fill light task m4061, lighting style determination task m4062, color temperature determination task m4063, and hue determination task m4064. The execution result of fill light task m4061 can correspond to various fill light methods such as front lighting, backlighting, side lighting, side-front lighting, side-backlighting, top lighting, and edge lighting. The execution result of lighting style determination task m4062 can correspond to multiple lighting styles, including flat lighting, horror lighting, 1920s lighting, and cyberpunk lighting. The execution result of color temperature determination task m4063 is related to warm white temperature, neutral white temperature, and cool white temperature. The execution result of hue determination task m4064 corresponds to red and blue hues.

[0069] Video output task m407 can include one or more of the following tasks: audio synthesis task m4071, color grading task m4072, subtitle synthesis task m4073, size configuration task m4074, and rendering output task m4075. Based on audio synthesis task m4071, background sounds, ambient sounds, and sound effects can be synthesized into an audio file. Based on color grading task m4072, the overall color tone of the video can be adjusted. Based on subtitle synthesis task m4073, one or more of the following can be performed: subtitle encapsulation, font settings, and subtitle effects encapsulation. Size configuration task m4074 corresponds to multiple video sizes, including 16:9 and 4:3. Rendering output task m4075 corresponds to one or more video frames and also to the video format. The video frame file format can be Joint Photographic Experts Group (JPEG) format, Portable Network Image (PNG) format, etc.

[0070] As you can understand, the above text has explained several initial prompts. The following text will further explain how to generate the video.

[0071] Figure 5 This is a schematic diagram of a video generation method according to an embodiment of the present disclosure.

[0072] like Figure 5 As shown, the user can input natural language text as initial text text50. Based on this initial text text50, initial script data can be determined. The initial script data may include initial script text st50. Based on the initial script text st50, multiple initial prompt texts can be determined. These multiple initial prompt texts may include initial content prompt text. It can be understood that the above explanation regarding natural language text text40 and script text st40 also applies to initial text text50 and initial script text st50. The following will explain some methods for determining object action-driven data.

[0073] In some embodiments of this disclosure, in some implementations of the above-described operation S220, determining the video content text and at least one initial object action-driven data corresponding to the video content text based on the initial content prompt text includes: generating the initial content text based on the initial content prompt text; determining the initial audio content data, time information, and at least one initial object action-driven data corresponding to the initial content text. For example, the above-described text-to-speech technology can be used to determine the initial audio content data audio501 corresponding to the initial content text tn50. The initial audio content data audio501 can correspond to time information. This time information is related to the duration of the initial audio content. Some methods for determining the object action-driven data will be described below. It is understood that the above description of the initial content text tn40 also applies to the initial content text tn50, and will not be repeated here.

[0074] In some embodiments, determining the initial audio content data, timing information, and at least one initial object motion-driven data corresponding to the initial content text includes: determining initial head motion-driven data based on the initial audio content data. For example, the initial head motion-driven data hdrive50 may correspond to the lip movements of the initial audio content data.

[0075] In some embodiments, determining the initial audio content data, timing information, and at least one initial object motion-driven data corresponding to the initial content text includes: determining the initial limb motion-driven data based on the initial content text.

[0076] In this embodiment of the disclosure, determining the initial limb movement driving data based on the initial content text includes: determining at least one initial content subtext of the initial content text based on its text structure; and determining at least one first initial limb movement driving subdata corresponding to the at least one initial content subtext. For example, the text structure can indicate the segmentation information of the text. The at least one content subtext of the content text tn50 may include: "This afternoon, the lunar exploration project's return capsule held an opening ceremony, marking a successful conclusion to the lunar exploration mission's engineering implementation phase.", "According to a report yesterday by the broadcasting company of country A, the head of the health bureau of a certain country announced that there are currently no major public health events.", and "The president of country B signed a decree approving the decision of unit C regarding the use of drones, and the decree has been published on the website of the presidential palace of country B." Different content subtexts may correspond to different actions. Thus, at least one action material can be determined from multiple action materials as at least one first action material corresponding to the content subtext. The driving data corresponding to the first action material can be used as the first initial limb movement driving subdata body510.

[0077] In this embodiment of the disclosure, determining the initial limb movement driving data based on the initial content text includes: determining at least one second initial limb movement driving sub-data corresponding to the initial content text based on time information. For example, based on the time information of the initial audio content data audio501 mentioned above, at least one second action material corresponding to the entire initial content text can be determined from multiple action materials. The second action material can increase the richness of the object's actions. The driving data corresponding to at least one second action material can be used as the second initial limb movement driving sub-data body520.

[0078] In this embodiment of the disclosure, determining the initial limb movement driving data based on the initial content text includes: determining at least one initial limb movement driving data based on at least one first initial limb movement driving sub-data and at least one second initial limb movement driving sub-data. For example... Figure 5 As shown, by fusing at least one first initial limb movement driving sub-data body510 and at least one second initial limb movement driving sub-data body520, at least one initial limb movement driving data bdrive50 can be determined.

[0079] In some embodiments, determining the initial audio content data, timing information, and at least one initial object motion-driven data corresponding to the initial content text includes: determining at least one initial object motion-driven data based on the initial head motion-driven data and the initial limb motion-driven data. For example... Figure 5 As shown, the initial object motion driving data can be determined based on the initial head motion driving data hdrive50 and at least one initial limb motion driving data bdrive50.

[0080] As we have explained above, some methods for determining the initial object's action-driven data will now be explained below, along with some methods for generating video.

[0081] In some embodiments, at least one initial material cue text includes initial style material cue text. In some embodiments of the above-described operation S220, generating an initial video based on at least one initial object motion-driven data and at least one initial material corresponding to at least one initial material cue text includes: determining initial object material based on initial lighting material corresponding to the initial style material cue text and at least one initial object motion-driven data. For example... Figure 5 As shown, the multiple initial prompt texts determined based on the initial script text st50 can include initial style material prompt text. Based on the initial style material prompt text, initial lighting material sm501 can be determined. Based on initial lighting material sm501, the virtual image corresponding to the initial prompt text "host," and at least one initial object motion-driven data, initial object material sm502 can be determined. Based on this initial object material, an initial video can be generated.

[0082] In this embodiment of the disclosure, the multiple initial prompt texts also include at least one of initial scene material prompt text and initial shot description prompt text. For example, the initial scene material prompt text can be "studio". The multiple scene materials can include indoor scene materials and outdoor scene materials. Outdoor scene materials can correspond to scenes such as mountains, grasslands, city rooftops, and treehouses in forests, and can also correspond to different weather materials. Different weather materials include sunny outdoor scene materials, rainy outdoor scene materials, and snowy outdoor scene materials. Indoor scene materials can include office materials, studio materials, and classroom materials, etc. The initial scene material sm503 corresponding to the initial scene material prompt text "studio" can be the aforementioned studio material. The initial shot description prompt text can correspond to the shot description information shot50. The shot description information can indicate the type of shot movement, focal length, and angle, etc.

[0083] In this embodiment of the disclosure, generating an initial video based on initial object material includes: obtaining multiple first initial video frames based on the initial object material, initial lighting material, and initial scene material corresponding to the initial scene material prompt text. For example, multiple first initial video frames are obtained based on the initial object material, initial lighting material, initial scene material, and initial shot description information corresponding to the initial shot description prompt text. Figure 5As shown, the initial object material sm502 can correspond to multiple different actions. Based on the initial object material sm502, the initial scene material sm503, and the shot description information shot50, multiple composite results r50 corresponding to the multiple actions are determined. Based on the multiple composite results r50 and the initial lighting material sm501, multiple first initial video frames f51 can be obtained. Based on the multiple first initial video frames f51, an initial video can be generated. It can be understood that the initial video can be directly determined based on multiple first initial video frames. Alternatively, one or more of the multiple first initial video frames can be selected to generate the initial video, as will be explained below.

[0084] In this embodiment of the disclosure, generating an initial video based on a plurality of first initial video frames includes: determining initial video editing information based on at least one of initial audio content data corresponding to initial content text, initial background sound effect material corresponding to initial style material prompt text, initial editing style material corresponding to initial style material prompt text, and initial transition style material corresponding to initial style material prompt text. For example... Figure 5 As shown, based on the initial style material prompt text, the initial background sound effect material sm504, the initial editing style material sm505, and the initial transition style material sm506 can be identified, and the initial video clip information clip50 can be determined. The initial video clip information can include one or more editing conditions. Editing conditions can be that the actions in different frames are the same. One of the two video frames that meet the editing conditions can be deleted. It is understood that editing conditions can also be other conditions, and this disclosure does not limit them.

[0085] In this embodiment of the disclosure, generating an initial video based on a plurality of first initial video frames includes: obtaining at least one second initial video frame based on initial video clip information and the plurality of first initial video frames. For example, at least one second video frame f52 can be obtained based on initial video clip information clip50 and the plurality of first initial video frames f51. Next, the initial video is generated based on the at least one second initial video frame.

[0086] In this embodiment of the disclosure, generating an initial video based on at least one second initial video frame includes: encapsulating at least one second initial video frame based on initial video encapsulation material corresponding to the initial style material prompt text to generate the initial video. For example... Figure 5 As shown, the initial style material prompt text can correspond to the initial video container material sm507. The initial video container material sm507 can indicate the format of the video container. Based on the initial video container material package50, at least one second initial video frame f52 is containerized to obtain the initial video v50.

[0087] As can be understood, the method of this disclosure has been explained above, and the post-processing method of the initial video of this disclosure will be explained below.

[0088] In some embodiments, sequence labels can be determined for multiple video frames in an initial video to adjust the order of the video frames. For example, content subtext can correspond to multiple video frames. The order of video content corresponding to different content subtexts in the video can be changed according to the order of the video frames.

[0089] In some embodiments, one or more video frames in the initial video that meet the deletion criteria can be deleted. The deletion criteria may include at least one of the following: the presence of a motion error, repetitive motion, or stillness.

[0090] In some embodiments, the playback rate of the initial video can be adjusted.

[0091] In some embodiments, an external application programming interface (API) can be called to render a tile image. The tile image can be an icon or other image displayed in the video. An external audio generation API can also be called to adjust the audio or sound effects of the initial video. If subtitles exist in the initial video, their font format can also be adjusted.

[0092] As you can understand, the post-processing methods for the video have been explained above. The following will explain the initial video.

[0093] In some embodiments, generating the initial video includes: displaying the initial video on a visual interface. The following will combine... Figures 6A to 6C Please provide an explanation.

[0094] Figure 6A This is a schematic diagram of a visual interface according to an embodiment of the present disclosure.

[0095] like Figure 6A As shown, the visual interface i60 may include an input box ib60. The user uses the input box ib60 to enter the initial text "A presenter is introducing the news." Next, an initial video can be generated based on this initial text and displayed on the visual interface i60.

[0096] Figure 6B and Figure 6C This is a schematic diagram of different video frames of an initial video according to an embodiment of the present disclosure.

[0097] like Figure 6B As shown, at the first moment, the visual interface i60 displays the initial video frame f621. (As...) Figure 6C As shown, at the second moment, the visual interface i60 displays video frame f622 of the initial video. The objects' actions differ between video frames f621 and f622.

[0098] As you can understand, the above has described the video generation method disclosed herein, and the following will further explain the application scenarios of this method.

[0099] In some embodiments, the methods described above can be used to create videos for social media. For example, using the methods described above and the conversational big model, users can quickly obtain interesting and engaging videos for display on social media platforms and conveniently share them daily.

[0100] In some embodiments, the methods described above can be used to create corporate advertising and marketing videos. For example, using the methods and conversational models described above, relevant personnel within a company can easily create promotional videos that showcase the features of the company's products and its brand philosophy, thereby improving advertising effectiveness and attracting more potential customers.

[0101] In some embodiments, the above methods can be used to create educational and training videos. For example, using the above methods and the dialogic model, teachers can transform teaching content into vivid and engaging videos, enhancing students' learning experience and enabling them to learn knowledge more intuitively.

[0102] In some embodiments, the methods described above can be used to produce news report videos. For example, using the methods described above and a conversational big data model, journalists, directors, and others can quickly produce news report videos, increasing the speed and impact of news dissemination and allowing viewers to understand news events more intuitively.

[0103] In some embodiments, the above methods can be used to create entertainment and leisure videos. For example, using the above methods and a conversational big model, users can create home movies, travelogues, etc., to record their lives in video format during entertainment and leisure activities.

[0104] In some embodiments, the above methods can be used to create e-commerce operation videos. For example, in the e-commerce industry, using the above methods and conversational models, livestreamers, merchants, and others can create product display videos to attract more customers to buy goods and improve efficiency.

[0105] In some embodiments, the methods described above can be used to create stylized animated videos. For example, using the methods described above and dialogic large models, animation creators can quickly create animated clips, significantly improving animation production efficiency and bringing a rich visual experience to the audience.

[0106] Through this embodiment, the video production process is entirely user-centric, following a standard commercial video production workflow. At each key stage, a large-scale model is used to input text prompts, assisting users in quickly achieving the desired effects. In this embodiment, the large-scale model can be a simple Chinese text input tool, supporting full Chinese input and demonstrating a more accurate understanding of Chinese semantics.

[0107] As can be understood, the application scenarios of the video generation method disclosed herein have been explained above. The video processing method disclosed herein will be explained below.

[0108] Figure 7 This is a flowchart of a video processing method according to another embodiment of the present disclosure.

[0109] like Figure 7 As shown, the method 700 may include operations S740 to S760.

[0110] In operation S740, at least one adjustment prompt text and at least one attribute adjustment information are determined based on the adjustment text corresponding to the video to be processed.

[0111] In this embodiment of the disclosure, the video to be processed can be an already generated video. For example, the video to be processed can be the initial video v50 described above.

[0112] In this embodiment of the disclosure, the video to be processed corresponds to at least one material to be adjusted. For example, the video to be processed may be generated using at least one material. Any material in the image to be processed can be used as a material to be adjusted. The aforementioned initial scene material can be used as a material to be adjusted. The attributes of the material may include size, position, color, content, etc.

[0113] In this embodiment of the disclosure, the adjustment text can be input by the user for the generated video. For example, the adjustment text could be "Please change the scene to a blue color scheme for me".

[0114] In this embodiment of the disclosure, various methods can be used to determine the adjustment prompt text and attribute adjustment information. For example, the adjustment text can be segmented into words. Based on the result of word segmentation, the adjustment prompt text and attribute adjustment information are determined. Taking the first noun as the adjustment prompt text and the verb and related words as the attribute adjustment information as an example, "scene" can be used as an adjustment prompt text, and the attribute adjustment information can be determined based on "change to blue".

[0115] In operation S750, based on at least one attribute adjustment information corresponding to at least one adjustment prompt text, the attribute information of at least one material to be adjusted corresponding to at least one adjustment prompt text is adjusted to obtain at least one adjusted material.

[0116] In this embodiment of the disclosure, the adjustment prompt text can correspond to one or more attribute adjustment information. The adjustment prompt text can also correspond to one or more materials to be adjusted. For example, the adjustment prompt text "scene" can correspond to a scene material and also to attribute adjustment information determined based on "change to blue tones". The color of the scene material can be adjusted to blue to obtain the adjusted scene material.

[0117] When operating the S760, a processed video is obtained based on at least one adjusted source material.

[0118] In this embodiment of the disclosure, a processed video can be generated based on the adjusted and unadjusted materials. Alternatively, the adjusted materials can be used to replace the materials corresponding to the adjustment prompt text to obtain the processed video.

[0119] This disclosure utilizes adjustment prompts and attribute adjustment information to modify source material, enabling targeted video modifications and reducing the labor and time costs associated with video generation. Furthermore, the ability to modify videos using user-inputted adjustment text lowers the barrier to targeted modification, allowing users without artistic expertise to create polished videos and effectively improving user experience.

[0120] In some embodiments, the video to be processed is obtained from an initial video, and the attribute information of at least one material to be adjusted corresponding to at least one adjustment prompt text is adjusted: in response to hitting the material to be adjusted corresponding to the adjustment prompt text from at least one material to be adjusted, the identification text of the material to be adjusted corresponding to the adjustment prompt text is displayed on the visual interface.

[0121] In some embodiments, determining at least one adjustment prompt text and at least one attribute adjustment information based on the adjustment text corresponding to the video to be processed includes: determining at least one adjustment prompt text and at least one attribute adjustment information using a large model based on the adjustment text. For example, the adjustment prompt text "scene" and the attribute adjustment information "change to blue tones" can be determined using the aforementioned large model llm30. Through embodiments of this disclosure, by utilizing a dialogic large model, the style and tone of a video can be quickly and in real-time changed based on the input text prompts.

[0122] In some embodiments, the video to be processed is obtained from an initial video, adjusting the attribute information of at least one material to be adjusted corresponding to at least one adjustment prompt text: in response to finding the material to be adjusted corresponding to the adjustment prompt text from at least one material to be adjusted, the identification text of the material to be adjusted corresponding to the adjustment prompt text is displayed on the visual interface. For example, the material to be processed used when generating the initial video v50 can be used as the material to be adjusted. From a plurality of materials to be adjusted, the material to be adjusted corresponding to the adjustment prompt text "scene" can be determined. This material to be adjusted can be the aforementioned initial scene material. The text "Successfully located scene" can be displayed on the visual interface to display the identification text "scene" of the material to be adjusted. Through the embodiments of this disclosure, displaying the identification text of the material corresponding to the adjustment prompt text on the visual interface allows the user to obtain standard and standardized text describing the material in the video, which helps to improve the efficiency of video adjustment and further improve the user experience.

[0123] In this embodiment of the disclosure, adjusting the attribute information of at least one material to be adjusted corresponding to at least one adjustment prompt text, based on at least one attribute adjustment information corresponding to at least one adjustment prompt text, includes: adjusting the attribute information of the material to be adjusted corresponding to the prompt text based on the attribute adjustment information corresponding to the prompt text. For example, based on the attribute adjustment information "change to blue" corresponding to the adjustment prompt text "scene", the scene material can be adjusted to a blue color scheme to obtain a blue scene, which is then used as the adjusted material.

[0124] In some embodiments, obtaining a processed video based on at least one adjusted source material may include displaying the processed video on a visual interface. For example, the processed video may be displayed on a visual interface.

[0125] It is understood that the method of this disclosure has been described above, and the apparatus of this disclosure will be described below.

[0126] Figure 8 This is a block diagram of a video generation apparatus according to an embodiment of the present disclosure.

[0127] like Figure 8 As shown, the device 800 may include a first determining module 810, a second determining module 820, and a generating module 830.

[0128] The first determining module 810 is used to determine multiple initial prompt texts based on the initial text input by the user. These multiple initial prompt texts include initial content prompt texts and initial material prompt texts.

[0129] The second determining module 820 is used to determine the video content text and at least one initial object action driving data corresponding to the video content text based on the initial content prompt text.

[0130] The generation module 830 is used to generate an initial video based on at least one initial object action-driven data and at least one initial material corresponding to at least one initial material cue text.

[0131] In some embodiments, the first determining module includes: a first determining submodule, configured to determine initial script data based on initial text and user attribute information; and a second determining submodule, configured to determine multiple initial prompt texts based on the initial script data.

[0132] In some embodiments, the initial script data includes at least one of initial script outline data, initial script storyboard description data, and initial script reference frame data. The initial script storyboard description data corresponds to at least one initial storyboard data, which includes at least one of initial scene description data, initial camera pointing data, initial motion data, initial audio data, initial lighting data, and initial duration data. The initial script reference frame data includes at least one of initial video size data, initial focal length data, and initial depth of field data.

[0133] In some embodiments, the second determining module includes: a first generating submodule, configured to generate initial content text based on the initial content prompt text; and a third determining submodule, configured to determine initial audio content data, timing information, and at least one initial object action driving data corresponding to the initial content text.

[0134] In some embodiments, the third determining submodule includes: a first determining unit, configured to determine initial head motion driving data based on initial audio content data; a second determining unit, configured to determine initial limb motion driving data based on initial content text; and a third determining unit, configured to determine at least one initial object motion driving data based on the initial head motion driving data and the initial limb motion driving data.

[0135] In some embodiments, the second determining unit includes: a first determining subunit, configured to determine at least one initial content subtext of the initial content text based on the text structure of the initial content text; a second determining subunit, configured to determine at least one first initial limb movement driving subdata corresponding to the at least one initial content subtext; a third determining subunit, configured to determine at least one second initial limb movement driving subdata corresponding to the initial content text based on time information; and a fourth determining subunit, configured to determine at least one initial limb movement driving data based on at least one first initial limb movement driving data and at least one second initial limb movement driving data.

[0136] In some embodiments, at least one initial material cue text includes initial style material cue text. The generation module includes: a fourth determining submodule, configured to determine initial object material based on initial lighting material corresponding to the initial style material cue text and at least one initial object motion-driven data; and a second generation submodule, configured to generate an initial video based on the initial object material.

[0137] In some embodiments, at least one initial material prompt text further includes an initial scene material prompt text. The second generation submodule includes: a first obtaining unit, configured to obtain multiple first initial video frames based on initial object material, initial lighting material, and initial scene material corresponding to the initial scene material prompt text; and a first generation unit, configured to generate an initial video based on the multiple first initial video frames.

[0138] In some embodiments, the plurality of initial cue texts further include initial shot description cue text. The first obtaining unit is further configured to: obtain a plurality of first initial video frames based on initial object material, initial lighting material, initial scene material, and initial shot description information corresponding to the initial shot description cue text.

[0139] In some embodiments, at least one initial material prompt text further includes initial style material prompt text. The first generation unit includes: a fifth determining subunit, configured to determine initial video clip information based on at least one of initial audio content data corresponding to the initial content text, initial background sound effect material corresponding to the initial style material prompt text, initial editing style material corresponding to the initial style material prompt text, and initial transition style material corresponding to the initial style material prompt text; a first obtaining subunit, configured to obtain at least one second initial video frame based on the initial video clip information and multiple first initial video frames; and a generation subunit, configured to generate an initial video based on the at least one second initial video frame.

[0140] In some embodiments, the generation subunit is further configured to: encapsulate at least one second initial video frame based on the initial video encapsulation material corresponding to the initial style material cue text, so as to generate an initial video.

[0141] In some embodiments, the first determining module is further configured to: determine multiple initial prompt texts based on the initial text and the user's attribute information using a large model. The large model is obtained by fine-tuning multiple sample texts and multiple preset prompt texts. The initial prompt texts are determined from the multiple preset prompt texts using the large model.

[0142] In some embodiments, the generation module includes: a first display submodule for displaying the initial video on a visual interface.

[0143] Figure 9This is a block diagram of a video processing apparatus according to another embodiment of the present disclosure.

[0144] like Figure 9 As shown, the device 900 may include a third determining module 940, an adjusting module 950, and an obtaining module 960.

[0145] The third determining module 940 is used to determine at least one adjustment prompt text and at least one attribute adjustment information based on the adjustment text corresponding to the video to be processed. The video to be processed corresponds to at least one material to be adjusted.

[0146] The adjustment module 950 is used to adjust the attribute information of at least one material to be adjusted corresponding to at least one adjustment prompt text based on at least one attribute adjustment information corresponding to at least one adjustment prompt text, so as to obtain at least one adjusted material.

[0147] Module 960 is used to obtain the processed video based on at least one adjusted source material.

[0148] In some embodiments, the video to be processed is obtained from an initial video, and the adjustment module includes: a second display submodule, configured to display the identification text of the video to be adjusted corresponding to the adjustment prompt text on a visual interface in response to hitting the video to be adjusted from at least one video to be adjusted.

[0149] In some embodiments, the obtaining module includes a third display submodule for displaying the processed video on a visual interface.

[0150] In some embodiments, the third determining module is further configured to: determine at least one adjustment prompt text and at least one attribute adjustment information based on the adjustment text using a large model.

[0151] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0152] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0153] Figure 10A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0154] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.

[0155] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0156] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as at least one of video generation methods and video processing methods. For example, in some embodiments, at least one of the video generation methods and video processing methods may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by computing unit 1001, it can perform one or more steps of at least one of the video generation method and video processing method described above. Alternatively, in other embodiments, computing unit 1001 can be configured to perform at least one of the video generation method and video processing method by any other suitable means (e.g., by means of firmware).

[0157] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0158] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0159] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM) or flash memory, optical fiber, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0160] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) monitor or a liquid crystal display (LCD)); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0161] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0162] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

[0163] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0164] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A video generation method, comprising: Based on the initial text input by the user and the user's attribute information, the initial script data is determined using a large model. The initial script data includes initial script outline data, initial script storyboard description data, and initial script reference screen data. Based on the initial script data, multiple initial prompt texts are determined, wherein the multiple initial prompt texts include initial content prompt texts and initial material prompt texts; Based on the initial content prompt text, determine the video content text and at least one initial object action driving data corresponding to the video content text; From a plurality of initial materials, determine at least one initial material corresponding to at least one initial material prompt text, wherein the at least one initial material prompt text includes an initial scene material prompt text, an initial style material prompt text, and an initial shot description prompt text; An initial video is generated based on at least one of the initial object action-driven data and at least one initial material corresponding to at least one of the initial material prompt texts.

2. The method according to claim 1, wherein, The initial script storyboard description data corresponds to at least one initial storyboard data, and the initial storyboard data includes at least one of the following: initial scene description data, initial camera indication data, initial action data, initial audio data, initial lighting data, and initial duration data. The initial script reference image data includes at least one of the following: initial video size data, initial focal length data, and initial depth of field data.

3. The method of claim 1, wherein, The step of determining the video content text and at least one initial object action-driven data corresponding to the video content text based on the initial content prompt text includes: Generate initial content text based on the initial content prompt text; Determine the initial audio content data, time information, and at least one of the initial object action-driven data corresponding to the initial content text.

4. The method of claim 3, wherein, The determination of the initial audio content data, time information, and at least one of the initial object action-driven data corresponding to the initial content text includes: Based on the initial audio content data, determine the initial head movement driving data; Based on the initial content text, determine the initial limb movement driving data; Based on the initial head motion driving data and the initial limb motion driving data, at least one initial object motion driving data is determined.

5. The method of claim 4, wherein, The step of determining the initial limb movement driving data based on the initial content text includes: Based on the text structure of the initial content text, at least one initial content subtext of the initial content text is determined; Determine at least one first initial limb movement driving sub-data corresponding to at least one of the initial content sub-texts; Based on the time information, at least one second initial limb movement driver sub-data corresponding to the initial content text is determined; At least one initial limb movement driving data is determined based on at least one first initial limb movement driving data and at least one second initial limb movement driving data.

6. The method according to claim 1, wherein, The step of generating the initial video based on at least one of the initial object action-driven data and at least one initial material corresponding to at least one of the initial material prompt texts includes: The initial object material is determined based on the initial lighting material corresponding to the initial style material prompt text and at least one of the initial object motion driving data; The initial video is generated based on the initial object material.

7. The method according to claim 6, wherein, The process of generating the initial video based on the initial object material includes: Based on the initial object material, the initial lighting material, and the initial scene material corresponding to the prompt text of the initial scene material, multiple first initial video frames are obtained; The initial video is generated based on a plurality of the first initial video frames.

8. The method according to claim 7, wherein, The step of obtaining multiple first initial video frames based on the initial object material, the initial lighting material, and the initial scene material corresponding to the initial scene material prompt text includes: Multiple first initial video frames are obtained based on the initial object material, the initial lighting material, the initial scene material, and the initial shot description information corresponding to the initial shot description prompt text.

9. The method according to claim 7, wherein, The step of generating the initial video based on a plurality of the first initial video frames includes: The initial video editing information is determined based on at least one of the following: the initial audio content data corresponding to the initial content text, the initial background sound effect material corresponding to the initial style material prompt text, the initial editing style material corresponding to the initial style material prompt text, and the initial transition style material corresponding to the initial style material prompt text. Based on the initial video clip information and multiple first initial video frames, at least one second initial video frame is obtained; The initial video is generated based on at least one of the second initial video frames.

10. The method of claim 9, wherein, Generating the initial video based on at least one second initial video frame includes: Based on the initial video encapsulation material corresponding to the initial style material prompt text, at least one second initial video frame is encapsulated to generate the initial video.

11. The method according to claim 1, wherein, The large model was obtained by fine-tuning multiple sample texts and multiple preset prompt texts. The initial prompt text is determined from multiple preset prompt texts using the large model.

12. The method of claim 1, wherein, The generation of the initial video includes: The initial video is displayed in the visual interface.

13. A video processing method, comprising: Based on the adjustment text corresponding to the video to be processed, at least one adjustment prompt text and at least one attribute adjustment information are determined, wherein the video to be processed corresponds to at least one material to be adjusted; Based on at least one attribute adjustment information corresponding to at least one of the adjustment prompt texts, adjust the attribute information of at least one material to be adjusted corresponding to at least one of the adjustment prompt texts to obtain at least one adjusted material; Based on at least one of the adjusted materials, a processed video is obtained; The video to be processed is obtained from an initial video, which is generated by the method according to any one of claims 1 to 12.

14. The method of claim 13, wherein, The adjustment of at least one of the attribute information of the material to be adjusted, corresponding to at least one of the adjustment prompt texts, includes: In response to finding the material to be adjusted corresponding to the adjustment prompt text from at least one of the materials to be adjusted, the identification text of the material to be adjusted corresponding to the adjustment prompt text is displayed on the visual interface.

15. The method of claim 13, wherein, The process of obtaining the processed video based on at least one of the adjusted materials includes: The processed video is displayed on a visual interface.

16. The method of claim 13, wherein, The step of determining at least one adjustment prompt text and at least one attribute adjustment information based on the adjustment text corresponding to the video to be processed includes: Based on the adjusted text, at least one adjustment prompt text and at least one attribute adjustment information are determined using a large model.

17. A video generation apparatus, comprising: The first determining module is used to determine initial script data based on the initial text input by the user and the user's attribute information using a large model. The initial script data includes initial script outline data, initial script storyboard description data, and initial script reference screen data. Based on the initial script data, multiple initial prompt texts are determined, wherein the multiple initial prompt texts include initial content prompt texts and initial material prompt texts. The second determining module is used to determine video content text and at least one initial object action driving data corresponding to the video content text based on the initial content prompt text; and to determine at least one initial material corresponding to at least one initial material prompt text from multiple initial materials, wherein the at least one initial material prompt text includes initial scene material prompt text, initial style material prompt text and initial shot description prompt text; A generation module is used to generate an initial video based on at least one of the initial object action-driven data and at least one initial material corresponding to at least one of the initial material prompt texts.

18. A video processing apparatus, comprising: The third determining module is used to determine at least one adjustment prompt text and at least one attribute adjustment information based on the adjustment text corresponding to the video to be processed, wherein the video to be processed corresponds to at least one material to be adjusted; An adjustment module is used to adjust the attribute information of at least one material to be adjusted corresponding to at least one adjustment prompt text according to at least one attribute adjustment information corresponding to at least one adjustment prompt text, so as to obtain at least one adjusted material; An acquisition module is used to obtain a processed video based on at least one of the adjusted materials; The video to be processed is obtained from an initial video, which is generated by the video generation apparatus according to claim 17.

19. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 16.

20. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 16.

21. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 16.