Method, device, equipment and product for generating video

By using a multimodal model to mix multiple effects during video generation, the problem of operational complexity and visual inconsistencies caused by selecting a single effect in existing technologies is solved, achieving efficient and unified video creation and an enhanced visual experience.

CN121012971APending Publication Date: 2025-11-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511149142.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

In existing technologies, users can only select a single effect in a video, resulting in numerous operation steps and complex video editing. Furthermore, multiple effect videos differ in visual style and motion continuity, affecting the visual experience and creative efficiency.

Method used

A method and apparatus are provided to simultaneously apply multiple special effects during video generation using a multimodal model. Users can select multiple special effects in a special effects selection component and generate a mixed special effects video using special effects description text. The multimodal model generates the video based on a reference image and special effects description text.

Benefits of technology

It reduces user operation steps, improves video creation efficiency, and generates videos with more consistent visual style and motion continuity, enhancing visual experience and immersion, and meeting more personalized creative needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121012971A_ABST
    Figure CN121012971A_ABST
Patent Text Reader

Abstract

The invention relates to a method, a device, equipment and a product for generating a video. The method comprises the following steps: displaying a special effect selection component in response to a detected interactive operation of a user for a mixed special effect component, and displaying a candidate special effect list in the special effect selection component; the method further includes determining a special effect description text in response to receiving a first special effect and a second special effect selected by the user from the list of candidate special effects. Further, the method includes acquiring a video generated by the multi-modal model based on a reference image provided by the user and the special effect description text, in which the first special effect and the second special effect are at least partially applied to the reference image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of computers, and more specifically to methods, apparatus, devices, and computer program products for generating video. Background Technology

[0002] With the development of machine learning technology, machine learning-based video generation and special effects processing technologies are increasingly being used in video creation and personalized content generation. Traditional video special effects production typically relies on professional video editing software and manual design processes. This process requires users to possess high levels of professional skills and a long production cycle. Furthermore, significant manual intervention is often required in stages such as material segmentation, effects overlay, and motion generation, leading to high costs and low efficiency in video production. In recent years, utilizing generative multimodal models to automate the processing of input images or videos to generate special effects videos has gradually become a new trend. Summary of the Invention

[0003] In embodiments of this disclosure, a method, apparatus, electronic device, and computer program product for generating video are provided.

[0004] In a first aspect of this disclosure, a method for generating a video is provided. The method includes displaying an effects selection component, showing a list of candidate effects, in response to detecting an interaction by a user with a hybrid effects component. The method further includes determining effect description text in response to receiving a first effect and a second effect selected by the user from the list of candidate effects. Furthermore, the method includes acquiring a video generated by a multimodal model based on a reference image provided by the user and the effect description text, in which the first and second effects are at least partially applied to the reference image.

[0005] In a second aspect of this disclosure, an apparatus for generating video is provided. The apparatus includes an effects selection component display module configured to display an effects selection component, displaying a list of candidate effects, in response to detecting a user interaction with a hybrid effects component. The apparatus also includes an effects description text determination module configured to determine effects description text in response to receiving a first and a second effect selected by the user from the list of candidate effects. Furthermore, the apparatus includes a video acquisition module configured to acquire a video generated by a multimodal model based on a reference image and effects description text provided by the user, in which the first and second effects are at least partially applied to the reference image.

[0006] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes a processor. The electronic device also includes a memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform the method according to a first aspect of this disclosure.

[0007] In a fourth aspect of this disclosure, a computer program product is provided, comprising computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to a first aspect of this disclosure.

[0008] In a fifth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions, which are executed by a processor to implement the method according to the first aspect.

[0009] The summary section is intended to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 A schematic diagram of an example environment in which several embodiments of the present disclosure may be implemented is shown;

[0012] Figure 2 A flowchart of a method for generating video according to some embodiments of the present disclosure is shown;

[0013] Figures 3A to 3K A schematic diagram illustrating an example of generating a video by mixing two special effects according to some embodiments of the present disclosure;

[0014] Figure 4 A schematic diagram illustrating an example of blending multiple effects by selecting multiple components in an effects selection component, according to some embodiments of the present disclosure;

[0015] Figure 5 A schematic diagram illustrating an example of saving multiple preset effects according to some embodiments of the present disclosure is shown;

[0016] Figure 6 A schematic diagram of an example system for generating video by mixing multiple special effects according to some embodiments of the present disclosure is shown;

[0017] Figure 7 Block diagrams of apparatus for generating video according to some embodiments of the present disclosure are shown; and

[0018] Figure 8 A schematic block diagram of an electronic device according to some embodiments of the present disclosure is shown.

[0019] In all the accompanying figures, the same or similar reference numerals denote the same or similar elements. Detailed Implementation

[0020] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0021] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0022] For example, upon receiving a user's proactive request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0023] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0024] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0026] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects unless explicitly stated. Other explicit and implicit definitions may also be included below.

[0027] In recent years, the use of generative multimodal models to automate the processing of input images or videos to generate special effects videos has gradually become a new trend. This technology can acquire reference images provided by the user (e.g., photos of people) and use image recognition, object detection, and segmentation techniques to identify the main objects and their features within the images. Then, the technology can apply special effects to the reference images based on the user-selected video effects template, such as changing the shape of the person, generating motion, and creating scene transitions, thereby automatically synthesizing the target video. For example, a user can upload a photo of a person and select the "person turns into a bee and flies away" effect template. The system can automatically generate a video based on the uploaded photo and the selected effect template, in which the person in the photo gradually transforms into a bee and flies away from the frame.

[0028] However, in some related technologies, users can only select one effect from the effects library and then apply that single effect to the input reference image. When users want to present two or more effects simultaneously in the same video, these technologies require generating multiple videos with a single effect separately, and then compositing these videos into the target video through subsequent video editing. This not only increases the number of steps but also requires additional video editing skills, reducing the user's creative efficiency. Furthermore, since multiple effect videos are generated separately and then composited, the various video segments may differ in visual style, lighting conditions, and the continuity of the main actions, resulting in abrupt scene transitions and a lack of smooth transitions in the final video, affecting the visual experience. In addition, being able to select only one effect template means that users cannot flexibly combine effects during the generation process. For example, users cannot achieve a continuous change of "a person turning into a bee while simultaneously making a crying expression" within the same generation task, limiting the flexibility of video creation.

[0029] Therefore, embodiments of this disclosure provide a scheme for user-generated video. In this scheme, in response to detecting a user's interactive operation on a hybrid effects component, a processing device can display an effects selection component, which displays a list of candidate effects. Then, in response to receiving a first effect and a second effect selected by the user from the list of candidate effects, the processing device can determine effect description text. The processing device can then acquire a video generated by a multimodal model based on a reference image and effect description text provided by the user, in which the first and second effects are at least partially applied to the reference image.

[0030] In this way, users can select multiple effect templates simultaneously in a single video generation task, and the system can blend these effects during the video generation process. Compared to related technologies that require generating multiple single-effect videos separately and then compositing them in post-production, this solution reduces user steps and improves video creation efficiency. Furthermore, the generated videos exhibit greater consistency in visual style, lighting, and subject movement, enhancing the visual appeal and immersion. In addition, users can freely combine different effect templates according to their creative needs, achieving diverse presentation methods such as "applying the first effect to the subject first and then transitioning to the second effect" or "simultaneously overlaying two effects on the subject," thus meeting more personalized creative requirements.

[0031] Figure 1 A schematic diagram of an example environment 100 in which various embodiments of this disclosure may be implemented is shown. For example... Figure 1 As shown, environment 100 includes processing device 102. Processing device 102 can be any device with processing capabilities. For example, processing device 102 can include smartphones, tablets, laptops, desktop computers, workstations, or smart wearable devices. In environment 100, processing device 102 can run application 104. Users can interact with application 104 through processing device 102. In application 104, users can create videos using materials such as images, videos, or live photos. For example, application 104 can be a content publishing platform, video creation platform, video editing platform, or social platform.

[0032] In environment 100, a user can initiate a request to create a video. For example, application 104 can provide components for creating new videos (e.g., a "Create Effects Video" button) on the user interface. Clicking this component triggers the request to create a video. Upon receiving the request, a video creation page 106 can be displayed in application 104. On video creation page 106, the user can upload a reference image 108, which application 104 will use to generate the video. For example, application 104 can request authorization from the user to access the local photo album of processing device 102 or the user's cloud photo album. After obtaining authorization, application 104 can retrieve multiple images from the local or cloud photo album and present them to the user for selection. The user can then select reference image 108 from these images. The selected reference image 108 or a thumbnail of reference image 108 can be displayed on video creation page 106.

[0033] In environment 100, a blending effects component 110 may be displayed on the video creation page 106. This component prompts the user to generate a video by blending multiple effects. In response to detecting a user interaction with the blending effects component 110, application 104 may display an effects selection component 112, which shows a candidate effects list 114. The user can use the effects selection component 112 to select multiple effects to blend from the candidate effects list 114. The effects selection component 112 may be, for example, a pop-up, a panel, or a page. In the example shown in environment 100, the candidate effects list 114 includes effects 116-1, 116-2, and 116-3 (collectively referred to as effects 116). Effect 116 may be a video or animation, which may include a specific storyline (e.g., a person transforming into a bee and flying out of the frame from the right). When effect 116 is applied to reference image 108 to generate video, it can blend objects in reference image 108 with the storyline in effect 116. Effect 116 can also have a specific painting style (e.g., the painting style of a specific painter, ink painting style, or abstract style, etc.), so that the generated video can have the painting style of effect 116. In representing a specific painting style, effect 116 can also be an image.

[0034] In environment 100, a user can select multiple effects to blend from a candidate effects list 114. For example, the user can select effects 116-1 and 116-2 for blending. In response to receiving effects 116-1 and 116-2 selected by the user from the candidate effects list 114, application 104 can determine effect description text 118. For example, the storyline of effect 116-1 may include flowers growing on a person's face, and the storyline of effect 116-2 may include a person transforming into a futuristic warrior. Effect description text 118 can be used as a cue word for a multimodal model to enable the multimodal model to generate video based on a reference image. In some embodiments, effect description text 118 may describe at least a portion of effect 116-1 and at least a portion of effect 116-2. For example, effect description text 118 may be "Transform into a futuristic warrior, the background changes from realistic to sci-fi, and flowers grow on the face during the transformation." In some embodiments, the effects description text 118 may be an instruction to mix effects 116-1 and effects 116-2 in a multimodal model. For example, the effects description text 118 may be “to generate a video by overlaying the provided multiple effects together”.

[0035] In environment 100, application 104 can provide reference image 108 and effects description text 118 to multimodal model 120 to generate video 122. In video 122, effects 116-1 and effects 116-2 are at least partially applied to reference image 108. For example, in video 122, a character in reference image 108 can gradually transform into a futuristic warrior, and during the transformation, flowers gradually grow on the character's face. Multimodal model 120 refers to a generative model capable of simultaneously processing information from different modalities (e.g., text, images, audio, video, etc.). Multimodal model 120 can simultaneously receive one or more of image information, text information, or video information as input, and encode, align, and fuse features from different modalities through a deep learning network to generate a video that meets the constraints of the input.

[0036] In some embodiments, application 104 may provide multimodal model 120 with reference image 108 and effects description text 118 (e.g., "Transform into a future warrior, the background changes from realistic to sci-fi, and flowers grow on the face during the transformation"). Multimodal model 120 may generate video 122 based on reference image 108 and effects description text 118. In some embodiments, application 104 may provide multimodal model 120 with reference image 108, effects description text 118 (e.g., "Generate a video by overlaying multiple provided effects"), effects 116-1, and effects 116-2 to generate video 122. Multimodal model 120 may be deployed on processing device 102 or may be a service provided by a third party. After multimodal model 120 generates video 122, application 104 may obtain video 122 from multimodal model 120 for displaying video 122 on a user interface or allowing the user to export video 122 to storage on processing device 102.

[0037] It should be understood that, for the sake of brevity, Figure 1 Only a specific number of effects 116 are shown, but this is not intended to limit the number of effects 116 in the candidate effects list 114. In other implementations, the candidate effects list 114 may include any number of effects 116. Furthermore, the user can select any number of effects 116 from the candidate effects list 114 to use for generating video 122.

[0038] In this way, users can select multiple effects 116 simultaneously in a single video generation task, and the system can blend these effects 116 during the video generation process, reducing user steps and improving video creation efficiency. Furthermore, the generated video 122 exhibits greater consistency in visual style, lighting, and subject movement, enhancing its visual appeal and immersive experience. In addition, users can freely combine different effects 116 according to their creative needs, achieving diverse presentation methods such as "applying the first effect to the subject first and then transitioning to the second effect" or "simultaneously overlaying two effects on the subject," thus satisfying more personalized creative requirements.

[0039] The following will combine Figures 2 to 8 The process according to embodiments of this disclosure is described in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and not intended to limit the scope of this disclosure. It should be understood that the embodiments described below may also include additional actions not shown and / or actions shown may be omitted, and the scope of this disclosure is not limited in this respect.

[0040] Figure 2A flowchart illustrating a method 200 for generating video according to some embodiments of the present disclosure is shown. Method 200 can be performed by a processing device. For example, method 200 can be performed by… Figure 1 The processing device 102 performs the operation. The processing device may include, for example, a smartphone, tablet computer, laptop computer, desktop computer, workstation, or smart wearable device.

[0041] In box 202, in response to detecting a user interaction with the hybrid effects component, the processing device can display an effects selection component, which shows a list of candidate effects. For example, in... Figure 1 In the environment 100 shown, in response to detecting a user interaction with the blending effects component 110, application 104 can display an effects selection component 112, which displays a candidate effects list 114. The user can use the effects selection component 112 to select multiple effects to blend from the candidate effects list 114. The candidate effects list 114 includes effects 116-1, 116-2, and 116-3. Effect 116 can be a video or animation, which may include a specific storyline. When effect 116 is applied to reference image 108 to generate a video, it can blend objects in reference image 108 with the storyline in effect 116. Effect 116 can also have a specific painting style, allowing the generated video to have the painting style of effect 116. In representing a specific painting style, effect 116 can also be an image.

[0042] In box 204, in response to receiving the first and second effects selected by the user from the list of candidate effects, the processing device can determine the effect description text. For example, in... Figure 1 In the illustrated environment 100, a user can select effects 116-1 and 116-2 for blending. In response to receiving the selection of effects 116-1 and 116-2 by the user from a candidate effects list 114, application 104 can determine effect description text 118. Effect description text 118 can be used as a cue word for a multimodal model to generate video based on a reference image. In some embodiments, effect description text 118 can describe at least a portion of effect 116-1 and at least a portion of effect 116-2. In some embodiments, effect description text 118 can be an instruction for the multimodal model to blend effects 116-1 and 116-2.

[0043] In box 206, the processing device can acquire video, which is generated by a multimodal model based on a user-provided reference image and effect description text, in which a first and a second effect are at least partially applied to the reference image. For example, in... Figure 1In the illustrated environment 100, application 104 can provide reference image 108 and effects description text 118 to multimodal model 120 to generate video 122. In video 122, effects 116-1 and effects 116-2 are applied at least partially to reference image 108. In some embodiments, application 104 can provide reference image 108 and effects description text 118 to multimodal model 120. Multimodal model 120 can generate video 122 based on reference image 108 and effects description text 118. In some embodiments, application 104 can provide reference image 108, effects description text 118, effects 116-1, and effects 116-2 to multimodal model 120. Multimodal model 120 can generate video 122 based on reference image 108, effects description text 118, effects 116-1, and effects 116-2. The multimodal model 120 can be deployed on the processing device 102 or it can be a service provided by a third party. After the multimodal model 120 generates video 122, the application 104 can obtain video 122 from the multimodal model 120 for displaying video 122 on the user interface or allowing the user to export video 122 to the storage device of the processing device 102.

[0044] In this way, users can select multiple effects simultaneously in a single video generation task, and the system can blend these effects during the video generation process, reducing user steps and improving video creation efficiency. Furthermore, the generated videos exhibit greater consistency in visual style, lighting, and main action transitions, enhancing visual appeal and immersion. In addition, users can freely combine different effects according to their creative needs, satisfying more personalized creative requirements.

[0045] Figures 3A to 3K A schematic diagram of example 300, which generates a video by mixing two special effects according to some embodiments of the present disclosure, is shown. Figure 3A A schematic diagram illustrating an example of the initial state of a video creation page according to some embodiments of the present disclosure is shown. Figure 3B A schematic diagram illustrating an example of selected reference images according to some embodiments of the present disclosure is shown. Figure 3C A schematic diagram illustrating an example of displaying a selected reference image on a video creation page according to some embodiments of the present disclosure. Figure 3D A schematic diagram showing examples of preview effects according to some embodiments of this disclosure is provided. Figure 3E A schematic diagram illustrating an example of the initial state of an effects selection component according to some embodiments of the present disclosure is shown. Figure 3F A schematic diagram is shown illustrating an example of an effect being selected in an effects selection component according to some embodiments of the present disclosure. Figure 3GA schematic diagram is shown illustrating an example of two effects being selected in an effects selection component according to some embodiments of the present disclosure. Figure 3H The diagram illustrates an example of a video creation page after a user has completed selecting effects, according to some embodiments of the present disclosure. Figure 3I A schematic diagram illustrating an example of a process for generating video according to some embodiments of the present disclosure is shown. Figure 3J A schematic diagram illustrating an example of a video creation page during the video generation process, according to some embodiments of the present disclosure. Figure 3K A schematic diagram illustrating an example of viewing a generated video according to some embodiments of this disclosure is shown.

[0046] In some embodiments, the processing device may display a text input box for receiving a description of the effect. In response to receiving a first effect and a second effect selected by the user from the list of candidate effects, the processing device may enter the effect description text into the text input box. By displaying the effect description text, the user can understand how the first and second effects are mixed, thus enabling the user to have reasonable expectations for the generated video content and improving the user experience. Furthermore, by entering the effect description text into the text input box, the user has the opportunity to modify the effect description text to better suit their needs in the generated video.

[0047] like Figure 3A As shown, Example 300 includes a video creation page 302. The video creation page 302 displays a reference image component 304, a read-only text input box 306, a target effect component 308, and a disabled generation component 310. Furthermore, functions associated with preset effects can be provided in area 312 of the video creation page 302. For example... Figure 3A As shown, area 312 may include a clear component 314, a blending effects component 316, and multiple preset effects (e.g., such as...). Figure 3A Special effects 1 to 6 are shown.

[0048] On the video creation page 302, users can select reference images from their album to generate videos by clicking the reference image component 304. Users can switch the text input box 306 from read-only to edit mode by clicking on or near it. Then, users can enter effect description text in the text input box 306 to describe the desired video or effects to be included in the video. When the user has not selected a reference image or the text input box 306 is empty, the generation component 310 is disabled. After the user has selected a reference image and entered effect description text in the text input box 306, the generation component 310 can be made available. At this time, users can trigger the generation of a video based on the reference image and effect description text by clicking the generation component 310.

[0049] In addition to manually entering the effect description text, users can also automatically fill the text input box 306 with the corresponding preset effect description text by clicking on the preset effects in area 312. For example, when the user clicks on effect 2, the application can obtain the preset effect description text for effect 2 (e.g., "Transform into a future warrior, the background changes from realistic to sci-fi") and automatically fill the preset effect description text into text input box 306. After the preset effect description text is filled into text input box 306, the user can further modify the text in text input box 306 to add the desired video content based on the preset effect. The selected preset effect can be displayed in the target effect component 308. For example, the target effect component 308 can display the thumbnail, logo, and other information of the selected preset effect.

[0050] When a user wants to blend multiple effects in a video, they can click on the blending effects component 316 to display an effects selection component, which shows a list of candidate effects. The user can then select multiple preset effects to blend. For example, when the user selects effects 1 and effects 2 through the blending effects component 316, the application can determine the corresponding effect description text, which is automatically filled into the text input box 306. The corresponding effect description text could be, for example, "Transform into a futuristic warrior, the background changes from realistic to sci-fi, and flowers grow on the face during the transformation" or "Generate a video by overlaying multiple provided effects together." Furthermore, the application can display thumbnails of effects 1 and effects 2 in the target effects component 308 to indicate to the user that they have selected to blend effects 1 and 2 to generate a video. The user can then click on the generation component 516 to trigger the generation of a video based on the reference image and the effect description text. In addition, users can also clear the selected preset effects or the blend of preset effects by clicking the clear component 314.

[0051] After detecting that the user clicked on the reference image component 304, the application can display as follows: Figure 3B The image selection component 320 is shown. (As shown...) Figure 3B As shown, the image selection component 320 can display at least one image that the user has authorized the application to access. The user can then select one image from the at least one image and import it into the application as a reference image. For example, the user can select image 322 by clicking on image 322. The user can then import image 322 into the application and designate it as the reference image by clicking on the add component 324.

[0052] like Figure 3C As shown, after acquiring image 322 as a reference image, the application can fill the reference image component 304 with a thumbnail of image 322 so that the user knows that image 322 has been uploaded as a reference image. In example 300, the user can select multiple effects by clicking the blending effects component 316 to apply multiple effects to the reference image by overlaying them.

[0053] In some embodiments, the processing device may display a first slot and a second slot in the effect selection component. In response to detecting a user's selection of a first effect, the processing device may fill the first effect into the first slot. In response to detecting a user's selection of a second effect, the processing device may fill the second effect into the second slot. In this way, the selected effects can be visually displayed on the user interface, making it easier for users to clearly understand which effects will be mixed, thereby improving the efficiency and user experience of effect selection.

[0054] In some embodiments, in response to a first effect being filled into a first slot, the processing device can display a clear component for the first slot in the effect selection component. In response to detecting a user interaction with the clear component, the processing device can remove the first effect from the first slot. This provides users with an intuitive and efficient way to remove effects, allowing them to flexibly adjust effect selections and improving the convenience and interactive experience of effect selection.

[0055] In some embodiments, the candidate effects in the candidate effects list include videos, and the candidate effects are displayed at a first size. The processing device may display a scaling component for the candidate effects in the effects selection component. In response to detecting a user interaction with the scaling component, the processing device may play the candidate effects at a second size, which is larger than the first size. In this way, users can more clearly view the visual effects and details of the video, thereby improving the accuracy of effects selection and the user experience.

[0056] like Figure 3DAs shown, after detecting a user click on the blending effects component 316, the application can display the effects selection component 326. The effects selection component 326 displays slots 328 and 330. Furthermore, the effects selection component 326 can also display a list of candidate effects. The list of candidate effects can include multiple effects. For example... Figure 3D As shown, multiple effect components can be displayed in the effect selection component 326. Each effect component can display a thumbnail of the corresponding effect, the playback duration of the effect, and a corresponding zoom component. For example, effect component 332 displays the cover of effect 4, the playback duration of effect 4, and the corresponding zoom component 334. After detecting that the user clicks on the zoom component 334, the application can display as follows: Figure 3E The effect preview component 336 is shown.

[0057] like Figure 3E As shown, the effects preview component 336 may include a video playback component 338. Users can use the video playback component 338 to play the effects and view their details. The video playback component 338 can be larger than the effects component 332, allowing users to clearly view the details of the effects. Users can close the effects preview component 336 by clicking the close component 340.

[0058] return Figure 3D Users can click on candidate effects to fill slot 328. For example, when the application detects that the user has selected effect 1, it can fill effect 1 into slot 328. Figure 3F As shown, slot 328 can display the cover of effect 1. After slot 328 is filled, slot 330 can switch from below slot 328 to above slot 328 to indicate to the user that the next selected effect will be filled into slot 330. Furthermore, after slot 328 is filled with an effect, a clear component 342 corresponding to slot 328 can be displayed above or near slot 328. When the user clicks the clear component 342, the application can remove the filled effect from slot 328. Then, the effect selection component 326 can return to... Figure 3D The state shown.

[0059] If slot 328 is already filled with an effect, when the application detects that the user clicks on another effect in the candidate effect list, it can fill that effect into slot 330. For example, Figure 3F As shown, slot 328 is already filled with effect 1. When the application detects that the user clicks on effect 2, it can fill effect 2 into slot 330. Figure 3GAs shown, the cover of Effect 2 can be displayed in slot 330. Additionally, a clear component corresponding to slot 330 can be displayed above or near slot 330; this clear component is used to remove the effects already filled in slot 330. When both slots 328 and 330 are filled with effects, the state of the completion component 344 can switch from disabled to enabled. Users can select mixed effects by clicking completion 344. When the application detects that a user has clicked completion component 344, it can close effect selection component 326 and return to the video creation page 302, as shown. Figure 3H As shown.

[0060] In some embodiments, in response to receiving a first effect and a second effect selected by the user from a list of candidate effects, the processing device may display a target effect component to indicate to the user that the first and second effects have been selected for video generation. This provides the user with clear confirmation feedback, helping them to promptly verify whether the selected effects meet expectations, thereby reducing the possibility of misselection and improving the accuracy of operations before video generation.

[0061] In some embodiments, in response to receiving a first effect and a second effect selected by the user from a list of candidate effects, the processing device may display a preset mixed effect component. This component indicates to the user that a mixed effect exists and is obtained by combining the first and second effects. In response to detecting a user interaction with the preset mixed effect component, the processing device may enter effect description text into a text input box. This helps users intuitively understand the combination relationship of effects, improving the predictability and controllability of the generated result. Furthermore, automatically entering effect description text into the text input box when the user triggers the preset mixed effect component reduces manual input steps, improving the efficiency and ease of interaction in video generation.

[0062] In some embodiments, the processing device can display the cover of a first effect and the cover of a second effect within preset effect components. This allows for an intuitive visual presentation of the elements comprising the effect combination, facilitating quick identification and review of the selected effects by the user, thereby enhancing the overall interactive experience.

[0063] In some embodiments, the processing device can receive user modifications to the effect description text via a text input box. The processing device can then generate a video using a multimodal model based on a reference image and the modified effect description text. This approach provides users with flexible customization capabilities for the effect content, making the generated video more aligned with personalized creative needs and enhancing the controllability and diversity of the generated video.

[0064] like Figure 3HAs shown, after receiving two effects selected by the user, the application can determine the effect description text associated with the selected effects and automatically fill the effect description text into the text input box 306. For example, in... Figure 3H In the example shown, the effect description text associated with the two selected effects could be "Transform into a futuristic warrior, the background changes from realistic to sci-fi, and flowers grow on your face during the transformation." Furthermore, the target effect component 308 on the video creation page 302 can display the cover or thumbnail of the two selected effects to indicate to the user the combination of effects to be used in the generated video. In addition, after receiving the two effects selected by the user, the application can add a preset blend effect component 346 in area 312. The preset blend effect component 346 corresponds to a preset blend effect, which consists of the two selected effects. Multiple covers of the multiple effects included in the preset blend effect can be displayed in the preset blend effect component 346. The preset blend effect can be stored in association with the user's account. Thus, when the user next enters the video creation page 302, the application can load and display the preset blend effect, allowing the user to quickly select the previously set blend effect by clicking the preset blend effect component 346, improving the efficiency of video creation. In addition, users can modify the effect description text through text input box 306 to better match their creative vision. For example, a user can change "flowers grow from the face" to "flowers grow from the mouth." This way, in the generated video, the flowers will grow from the mouth of the person in the reference image, instead of from other parts of the face.

[0065] After detecting a user click on the generation component 310, the application can utilize a multimodal model and generate a video based on reference images and effect description text. During video generation, the application can display, for example... Figure 3I The video generation page 348 is shown. The progress of the video generation process can be displayed on video generation page 348. Furthermore, the video generation process can run in the background, and users can return to the previous screen by clicking component 350 before the video generation process is complete. Figure 3J The video creation page shown is 302.

[0066] like Figure 3JAs shown, since the video generation process is not yet complete, a progress indicator pop-up window 352 can be displayed on the video creation page 302. The progress indicator pop-up window 352 displays the progress of the video generation process to remind the user that the video generation process is not yet finished. In addition, the progress indicator pop-up window 352 can also display thumbnails of reference images, the cover of the selected effect, or a combination thereof. When the user clicks the close button on the progress indicator pop-up window 352, a pop-up window can prompt the user to view the generated video on the history task page. Then, the user can create a new blending effect by clicking the blending effect component 316. In some embodiments, the video creation page 302 can display only a specific number of preset blending effect components 346 (e.g., only one preset blending effect component). If the user has already set a preset blending effect, when the user sets a preset blending effect again, the application can prompt the user via a pop-up window that the new preset blending effect will be filled into the preset blending effect component 346 and overwrite the original preset blending effect component 346. Since there are multiple ways to combine special effects, and users can easily set new preset mixed effects through the mixed effect component 316, limiting the number of preset mixed effect components 346 can prevent a large number of idle preset mixed effects from appearing on the page and push individual preset effects out of the page's display area.

[0067] After the video generation process is complete, the application can display something like this: Figure 3K The video preview page 354 is shown. Video preview page 354 displays a video playback component 356, a music component 358, a music selection component 360, and an export component 362. Users can view the generated video through the video playback component 356. When the user selects an effect with associated music, the generated video can be configured with the music associated with the selected effect. The music component 358 displays the name of the configured music. This helps users quickly locate music that matches the selected effect, improving video creation efficiency. Furthermore, users can click the music selection component 360 to reselect music for the video. When the application detects that the user clicks the export component 362, it can export the generated video to the user's device or cloud storage, or publish the generated video to a content sharing platform or social media platform.

[0068] In some embodiments, the processing device may display a third slot in the effect selection component. In response to detecting a user's selection of a third effect, the processing device may fill the third effect into the third slot. In response to receiving a first effect, a second effect, and a third effect selected by the user from the candidate effect list, the processing device may determine effect description text, which describes at least a portion of the first effect, at least a portion of the second effect, and at least a portion of the third effect. In this way, more effects can be mixed, resulting in richer combinations of effects.

[0069] Figure 4 A schematic diagram of example 400, illustrating how multiple effects are blended by selecting multiple components in an effects selection component according to some embodiments of the present disclosure, is shown. Figure 4 As shown, when a user clicks on a blending effects component on the video creation page (e.g., in...), Figure 3C After clicking the blending effects component 316 on the video creation page 302, the application can display the effects selection component 402. The effects selection component 402 displays slots 328, 330, and 408. Additionally, the effects selection component 402 can also display a list of candidate effects. The candidate effects list can include multiple effects. Users can click on a candidate effect to fill it into slot 404. For example, when the application detects that the user has selected effect 1, it can fill effect 1 into slot 404. After slot 404 is filled, slot 406 can switch from below slot 404 to above slot 404 to indicate to the user that the next selected effect will be filled into slot 406. For example, when the application detects that the user has selected effect 2, it can fill effect 2 into slot 406. After slot 406 is filled, slot 408 can switch from below to above slot 406 to indicate to the user that the next selected effect will be filled into slot 408. For example, when the application detects that the user has selected effect 3, it can fill effect 3 into slot 408. When slots 404, 406, and 408 are all filled with effects, the user can complete the blending effect settings by clicking the completion component 412. The application can then determine the effect description text based on effect 1, effect 2, and effect 3, which may include at least a portion of the content of effect 1, at least a portion of the content of effect 2, and at least a portion of the content of effect 3.

[0070] In some embodiments, the aforementioned blending effect is a first blending effect, and the preset blending effect component is a first preset blending effect component. In response to receiving a fourth and a fifth effect selected by the user from a list of candidate effects, the processing device can display a second preset blending effect component. This second preset blending effect component is used to indicate to the user that a second blending effect exists and that the second blending effect is obtained by blending the fourth and fifth effects. In response to detecting a user interaction with the second preset blending effect component, the processing device can enter a second prompt word in a text input box. The second prompt word describes at least a portion of the fourth effect and at least a portion of the fifth effect. In this way, the application can store more preset blending effects, preventing previously set preset blending effects from being overwritten by new preset blending effects.

[0071] Figure 5 A schematic diagram of an example 500 showing the storage of multiple preset effects according to some embodiments of the present disclosure is shown. Figure 5 As shown, a preset blending effect component 506 is displayed in area 504 of the video creation page 502, and the preset blending effect component 506 is bound to a blending effect. When the user creates a new blending effect using the blending effect component 518, the application can add a new preset blending effect component 508 in area 504 and bind the preset blending effect component 508 to the new blending effect. In this way, the user can enter the corresponding effect description text in the text input box 510 by clicking on either the preset blending effect component 506 or the preset blending effect component 508. Furthermore, the target effect component 512 can display multiple effects corresponding to the currently selected preset blending effect component.

[0072] In some embodiments, in response to receiving a first effect and a second effect selected by the user from a list of candidate effects, the processing device can acquire a first cue word describing the first effect and a second cue word describing the second effect. The processing device can then acquire effect description text generated based on the first and second cue words. In this way, accurate text content describing the mixed effects can be automatically generated, reducing the burden of manual writing by the user, improving the accuracy and consistency of effect descriptions, and thus enhancing the efficiency and quality of video generation.

[0073] In some embodiments, the processing device can generate a mixed cue word based on a first cue word and a second cue word. The mixed cue word includes the first cue word, the second cue word, and instructions for mixing the first and second cue words. The processing device can then obtain a third cue word generated by a language model based on the mixed cue word as the effect description text. Compared to directly concatenating the first and second cue words to generate the effect description text, semantic analysis and fusion of the first and second cue words result in a more natural effect description text in terms of syntactic structure, semantic logic, and temporal relationships, reducing the problems of awkward sentences, unclear logic, or ambiguous meaning caused by direct concatenation. Furthermore, by adapting the content of the two cue words, this method can reflect the transition or superposition relationship between the two effects in the description, thereby more accurately guiding the multimodal model to generate a mixed effect video that meets expectations, rather than simply presenting the two effects independently.

[0074] In some embodiments, the first special effect includes a first storyline, the second special effect includes a second storyline, and the video includes at least a portion of the first storyline and at least a portion of the second storyline.

[0075] Figure 6 A schematic diagram of an example system 600 for generating video by mixing multiple special effects according to some embodiments of the present disclosure is shown. Figure 6 As shown, system 600 can acquire special effects 602 and 604 selected by the user. Then, system 600 can acquire the prompt word 606 corresponding to special effect 602 and the prompt word 608 corresponding to special effect 604. For example, prompt word 606 could be "flowers grow on the face," and prompt word 608 could be "transform into a futuristic warrior, the background changes from realistic to sci-fi." Then, system 600 can generate a mixed prompt word 610 based on prompt words 606 and 608. Mixed prompt word 610 can include prompt words 606 and 608, and also includes instructions for mixing prompt words 606 and 608. For example, the instruction could be "output a new prompt word based on the storyline of the two provided prompt words. The new prompt word should integrate the storylines of the two prompt words and conform to normal logic. The new storyline can be appropriately polished."

[0076] After generating the hybrid cue word 610, the system 600 can utilize a large language model (LLM) 612 and generate a new cue word 614 based on the hybrid cue word 610. The cue word 614 may include at least a portion of the storyline described by cue word 606 and at least a portion of the storyline described by cue word 608. For example, the cue word 614 could be "transform into a future warrior, the background changes from realistic to sci-fi, and flowers grow on the face during the transformation." Then, the system 600 can utilize a multimodal model 618 (e.g., Figure 1 The multimodal model 120 in the image generates video 602 based on cue words 614 and reference images 616. Thus, video 602 may include at least a portion of the storyline in special effects 602 and at least a portion of the storyline in special effects 604.

[0077] Figure 7 A block diagram of an apparatus 700 for generating video according to some embodiments of the present disclosure is shown. Figure 7 As shown, the device 700 includes an effects selection component display module 702, configured to display an effects selection component in response to detecting a user's interactive operation on a hybrid effects component, the effects selection component displaying a list of candidate effects. The device 700 also includes an effects description text determination module 704, configured to determine effects description text in response to receiving a first and a second effect selected by the user from the list of candidate effects. Furthermore, the device 700 includes a video acquisition module 706, configured to acquire a video generated by a multimodal model based on a reference image and effects description text provided by the user, in which the first and second effects are at least partially applied to the reference image.

[0078] It is understood that the device 700 disclosed herein can achieve at least one of the many advantages that the methods or processes described above can achieve. For example, a user can select multiple effect templates simultaneously in a single video generation task, and the device 700 can blend multiple effects during the video generation process. Compared to related technologies that require generating multiple single-effect videos separately and then compositing them in post-production, the device 700 can reduce the user's operational steps and improve the efficiency of video creation. In addition, the generated videos are more unified in terms of visual style, lighting consistency, and subject action transitions, thereby enhancing the visual appeal and immersion of the videos. Furthermore, users can freely combine different effect templates according to their creative needs, achieving diverse presentation methods such as "applying the first effect to the subject first and then transitioning to the second effect" or "simultaneously overlaying two effects on the subject," thereby meeting more personalized creative needs.

[0079] Figure 8 A block diagram of a device 800 capable of implementing various embodiments of the present disclosure is shown. (See diagram for example.) Figure 8As shown, device 800 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 801, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 802 or loaded from storage unit 808 into random access memory (RAM) 803. The RAM 803 can also store various programs and data required for the operation of device 800. The CPU / GPU 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804. Although not shown in... Figure 8 As shown, device 800 may also include a coprocessor.

[0080] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0081] The various methods or processes described above can be executed by CPU / GPU 801. For example, in some embodiments, the methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by CPU / GPU 801, one or more steps or actions in the methods or processes described above can be performed.

[0082] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0083] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0084] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or external storage device. The network may include copper cables, fiber optic cables, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0085] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​and conventional procedural programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0086] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0087] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0088] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0089] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for generating video, comprising: In response to detecting a user's interaction with the hybrid effects component, an effects selection component is displayed, which shows a list of candidate effects; In response to receiving a first and a second effect selected by the user from the list of candidate effects, a description text for the effects is determined; as well as The video is acquired, which is generated by a multimodal model based on a reference image provided by the user and the effect description text, in which the first effect and the second effect are at least partially applied to the reference image.

2. The method according to claim 1, further comprising: Display a text input box, which is used to receive a description of the special effect; as well as In response to receiving the first and second effects selected by the user from the list of candidate effects, the effect description text is entered into the text input box.

3. The method according to claim 1, further comprising: The first slot and the second slot are displayed in the special effects selection component; In response to detecting a user's selection of the first effect, the first effect is filled into the first slot; and In response to detecting the user's selection of the second effect, the second effect is filled into the second slot.

4. The method according to claim 3, further comprising: In response to the first effect being filled into the first slot, a clear component for the first slot is displayed in the effect selection component; as well as In response to detecting a user interaction with the clear component, the first effect is removed from the first slot.

5. The method according to claim 3, further comprising: The third slot is displayed in the effects selection component; In response to detecting the user's selection of a third effect, the third effect is filled into the third slot; as well as In response to receiving the first effect, the second effect, and the third effect selected by the user from the list of candidate effects, the effect description text is determined, the effect description text describing at least a portion of the first effect, at least a portion of the second effect, and at least a portion of the third effect.

6. The method of claim 1, wherein the candidate effects in the candidate effects list include videos, the candidate effects are displayed at a first size, and the method further comprises: A scaling component for the candidate effects is displayed in the effects selection component; as well as In response to detecting a user interaction with the scaling component, the candidate effect is played at a second size, which is larger than the first size.

7. The method according to claim 1, further comprising: In response to receiving the first effect and the second effect selected by the user from the candidate effect list, a target effect component is displayed, which is used to prompt the user that the first effect and the second effect have been selected for generating the video.

8. The method according to claim 2, further comprising: In response to receiving the first effect and the second effect selected by the user from the candidate effect list, a preset mixed effect component is displayed, which is used to prompt the user that a mixed effect exists and that the mixed effect is obtained by mixing the first effect and the second effect; as well as In response to detecting a user interaction with the preset hybrid effects component, the effect description text is entered into the text input box.

9. The method according to claim 8, further comprising: The cover of the first effect and the cover of the second effect are displayed in the preset mixed effects component.

10. The method of claim 8, wherein the blending effect is a first blending effect, and the method further comprises: In response to receiving a fourth and a fifth effect selected by the user from the candidate effect list, the first hybrid effect is replaced in the preset hybrid effect component with a second hybrid effect, which is obtained by mixing the fourth and the fifth effect.

11. The method of claim 8, wherein the blending effect is a first blending effect, the preset blending effect component is a first preset blending effect component, and the method further comprises: In response to receiving the fourth and fifth effects selected by the user from the candidate effect list, a second preset mixed effect component is displayed. The second preset mixed effect component is used to prompt the user that a second mixed effect exists and that the second mixed effect is obtained by mixing the fourth and fifth effects. as well as In response to detecting a user interaction with the second preset hybrid effect component, a second prompt word is entered into the text input box, the second prompt word describing at least a portion of the fourth effect and at least a portion of the fifth effect.

12. The method of claim 2, wherein generating the video using the multimodal model based on the reference image and the special effects description text comprises: The text input box is used to receive modifications made by the user to the description text of the special effect; as well as The video is generated using the multimodal model based on the reference image and the modified effects description text.

13. The method according to claim 1, further comprising: In response to receiving the first effect and the second effect selected by the user from the candidate effect list, a first prompt word for describing the first effect and a second prompt word for describing the second effect are obtained; as well as Obtain the effect description text generated based on the first prompt word and the second prompt word.

14. The method of claim 13, wherein obtaining the effect description text generated based on the first prompt word and the second prompt word comprises: A mixed prompt word is generated based on the first prompt word and the second prompt word, wherein the mixed prompt word includes the first prompt word, the second prompt word, and an instruction for mixing the first prompt word and the second prompt word; as well as Obtain a third prompt word generated by the language model based on the hybrid prompt word as the effect description text.

15. The method of claim 1, wherein the first special effect comprises a first storyline, the second special effect comprises a second storyline, and the video comprises at least a portion of the first storyline and at least a portion of the second storyline.

16. An apparatus for generating video, comprising: The special effects selection component display module is configured to display a special effects selection component in response to detecting a user's interactive operation on the hybrid special effects component, wherein the special effects selection component displays a list of candidate special effects; The special effects description text determination module is configured to determine special effects description text in response to receiving a first special effect and a second special effect selected by the user from the candidate special effects list; as well as A video acquisition module is configured to acquire the video, which is generated by a multimodal model based on a reference image provided by a user and the effect description text, in which the first effect and the second effect are at least partially applied to the reference image.

17. An electronic device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1 to 15.

18. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of claims 1 to 15.