Video generation method and apparatus, electronic device, and medium

CN122802726APending Publication Date: 2026-09-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510336348.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

当前大多视频工作者在进行视频创作时,在确定了视频的主题之后,都需要手动选取相应的素材并进行剪辑和文案编辑,受限于人工能力的制约,无法快速地生成视频作品,降低了制作视频作品的效率,且不容易保持视频作品风格与作者风格统一

Benefits of technology

[0107] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objectives and other advantages of this disclosure may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802726A_ABST
    Figure CN122802726A_ABST
Patent Text Reader

Abstract

This disclosure provides a video generation method, apparatus, electronic device, and medium. The video generation method includes: obtaining a reference video storage space address; obtaining multiple reference videos from the reference video storage space address; extracting common style text from the multiple reference videos using a multimodal model; obtaining an image set, wherein the images in the image set have image tags; determining multiple matching images in the image set based on the matching of the common style text and the image tags; and generating a target video based on the multiple matching images using the multimodal model. Embodiments of this disclosure can improve the efficiency of video generation while meeting the conditions of a desired video style. Embodiments of this disclosure can be applied to various scenarios such as short video production and work summary video production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of video processing, and in particular to a video generation method, apparatus, electronic device, and medium. Background Technology

[0002] With the widespread use of short video platforms in people's lives, more and more video creators are entering the field of video production. Currently, most video creators, after determining the theme of their video, need to manually select relevant materials and edit and transcribe the script. Limited by manual capabilities, this prevents the rapid generation of video works, reducing efficiency and making it difficult to maintain consistency between the video's style and the creator's personal style. Although the market provides some templates for video creators to use in creating videos and related scripts, this method still relies on video creators manually selecting and filtering materials, failing to quickly generate the desired video works and effectively improve production efficiency. Summary of the Invention

[0003] This disclosure provides a video generation method, apparatus, electronic device, and medium that can improve the efficiency of video generation while meeting the conditions of a desired video style.

[0004] According to one aspect of this disclosure, a video generation method is provided, comprising:

[0005] Obtain the storage space address of the reference video;

[0006] Multiple reference videos are obtained from the reference video storage space address;

[0007] Using a multimodal model, common style text is extracted from the multiple reference videos;

[0008] Obtain an image set, wherein the images in the image set have image tags;

[0009] Based on the matching of the common style text and the image tags, multiple matching images are determined in the image set;

[0010] Using the multimodal model, a target video is generated based on the multiple matched images.

[0011] According to one aspect of this disclosure, a video generation apparatus is provided, the video generation apparatus comprising:

[0012] The first acquisition unit is used to acquire the storage space address of the reference video.

[0013] The second acquisition unit is used to acquire multiple reference videos from the reference video storage space address;

[0014] The extraction unit is used to extract common style text from the multiple reference videos using a multimodal model;

[0015] The third acquisition unit is used to acquire an image set, wherein the images in the image set have image tags;

[0016] The determining unit is used to determine multiple matching images in the image set based on the matching of the common style text and the image tags;

[0017] The generation unit is used to generate a target video based on the multiple matching images using the multimodal model.

[0018] Optionally, the determining unit is specifically used for:

[0019] Display the aforementioned common style text;

[0020] In response to positive and negative markers on the displayed common style text, positive and negative marker text is identified in the common style text;

[0021] Images whose image tags match the negative text are filtered out to obtain a set of filtered images;

[0022] In the filtered image set, images whose image tags match the positive label text are identified, thus obtaining the multiple matching images.

[0023] Optionally, the determining unit is specifically used for:

[0024] Extract the first seed keyword from the negatively labeled text;

[0025] The first seed keyword is expanded using a thesaurus to obtain the first expanded keyword;

[0026] From the image set, images whose image tags match the first expanded keyword are filtered out to obtain the filtered image set.

[0027] Optionally, the determining unit is specifically used for:

[0028] Extract the second seed keyword from the positively labeled text;

[0029] The second seed keyword is expanded using a thesaurus to obtain the second expanded keyword;

[0030] In the filtered image set, images whose image tags match the second expanded keywords are identified to obtain the plurality of matching images.

[0031] Optionally, the generation unit is specifically used for:

[0032] Using the multimodal model, the target video is generated based on the multiple matching images and according to the positive and negative labeled text.

[0033] Optionally, the generation unit is specifically used for:

[0034] The multiple matching images are divided into categories according to their generation time.

[0035] For each category, the matching images under that category are divided into subcategories according to the subject.

[0036] The multiple matching images are displayed according to the class, wherein, under each class, the subclasses into which the matching images under the class are divided are displayed;

[0037] In response to the selection of the displayed subclass, the target video is generated based on the matching image under the subclass.

[0038] Optionally, the generation unit is specifically used for:

[0039] In response to the selection of the displayed subclass, the matching image under the subclass is obtained;

[0040] Using the multimodal model, common video features and common text features of the multiple reference videos are extracted from the multiple reference videos;

[0041] Using the multimodal model, the target video is generated based on the matching images under the subclass, according to the common video features and the common text features.

[0042] Optionally, the generation unit is specifically used for:

[0043] Using the multimodal model, based on the matched images under the subclass, and according to the common text features, video scripts and video narration are generated;

[0044] Display the video text and, in response to editing operations on the video text, generate an edited video text;

[0045] Display the video narration, and generate an edited video narration in response to an editing operation on the video narration;

[0046] Display the matching image under the subclass, and generate the edited matching image in response to the editing operation of the matching image;

[0047] Using the multimodal model, based on the edited matched image and the edited video narration, the target video is generated according to the common video features, and the target video is displayed in association with the edited video text.

[0048] Optionally, the generation unit is specifically used for:

[0049] In response to the editing operation on the anchor object in the anchor image of the matching image, a first edit is performed on the anchor object in the anchor image;

[0050] Identify the anchor object in other matching images besides the anchor image in the matching image;

[0051] For the anchor objects in the other matching images, perform the same second edit as the first edit to generate the edited matching image.

[0052] Optionally, the generation unit is specifically used for:

[0053] Display image range input control;

[0054] The image range input control receives the input image range.

[0055] Within the range of input images, identify the anchor object in other matching images besides the anchor image.

[0056] Optionally, the generation unit is specifically used for:

[0057] The multimodal model is used to perform image text annotation on the matching images within the input image range to obtain multiple annotation information corresponding to each matching image;

[0058] The confidence level of each annotation in the matching image is determined based on the matching images within the image range and the multiple annotation information corresponding to the matching images.

[0059] The anchor object confidence of the anchor object in the matching image is determined based on the confidence of each annotation information in the matching image.

[0060] If the confidence level of the anchor object in the matched image is greater than a first confidence threshold, the anchor object in the matched image is identified.

[0061] Optionally, the first editing is an elimination process, and the generation unit is specifically used for:

[0062] The anchor objects in the other matching images are masked to obtain the masked areas;

[0063] The image content within the masked area is removed to obtain the edited matching image.

[0064] Optionally, the extraction unit is specifically used for:

[0065] Based on the first sub-model, the second sub-model, and the third sub-model, a base model is generated, wherein the first sub-model is used for mutual generation between texts, the second sub-model is used for mutual generation between text and images, and the third sub-model is used for mutual generation between text and video.

[0066] Obtain a video sample set, and generate a data sample set based on the video sample set;

[0067] Based on the data sample set, the basic model is trained and optimized to obtain the multimodal model.

[0068] Optionally, the first sub-model includes a first encoder and a first decoder, the second sub-model includes a second encoder and a second decoder, and the third sub-model includes a third encoder and a third decoder;

[0069] The extraction unit is specifically used for:

[0070] The outputs of the first encoder, the second encoder, and the third encoder are connected to the input of the central encoder, and the output of the central encoder is connected to the inputs of the first decoder, the second decoder, and the third decoder, respectively, to form the basic model.

[0071] Optionally, the extraction unit is specifically used for:

[0072] Visual features, audio features, and text features are extracted from video samples in the video sample set.

[0073] Obtain the class target label of the video sample;

[0074] Based on the visual features, audio features, text features, and class tags of the video samples, data samples corresponding to the video samples are generated, thereby obtaining the data sample set based on the video sample set.

[0075] Optionally, the extraction unit is specifically used for:

[0076] Audio information and subtitle text information are extracted from the video samples in the video sample set;

[0077] Image information is extracted from the video sample, and image text is annotated on the image information to obtain annotated text information;

[0078] Based on the subtitle text information and the annotation text information, they are integrated into text information;

[0079] The visual features are extracted from the video samples, the audio features are extracted from the audio information, and the text features are extracted from the text information.

[0080] Optionally, the extraction unit is specifically used for:

[0081] Extract metadata from the video samples;

[0082] Generate a content description for the video sample;

[0083] Input the metadata and content description into the pre-classification model to obtain the initial class target label;

[0084] The initial class label is adjusted to obtain the class label.

[0085] Optionally, the extraction unit is specifically used for:

[0086] The data sample set is divided into a first sample subset and a second sample subset;

[0087] Based on the first sample subset, the base model is trained to obtain the trained model;

[0088] Based on the second sample subset, the trained model is optimized to obtain the multimodal model.

[0089] Optionally, the data samples in the first sample subset include a first visual feature, a first audio feature, a first text feature, and a first type of target tag;

[0090] The extraction unit is specifically used for:

[0091] The first visual feature, the first audio feature, and the first text feature are input into the base model to obtain the predicted category;

[0092] Based on the predicted category and the first type of target tag, calculate the first loss function;

[0093] Based on the first loss function, the base model is trained to obtain the trained model.

[0094] Optionally, the data samples in the second sample subset include second visual features, second audio features, second text features, and a second type of target tag;

[0095] The extraction unit is specifically used for:

[0096] Obtain a task instance, which includes an instance image and an instance text description;

[0097] Using the trained model, an instance video is generated based on the instance image and the instance text description;

[0098] Identify the video category of the example video;

[0099] Extract the second visual feature, the second audio feature, and the second text feature from the example image and the example text description;

[0100] From the data samples of the second sample subset, obtain the second type of target tag corresponding to the second visual feature, the second audio feature, and the second text feature;

[0101] Based on the video category and the second target tag, calculate the second loss function;

[0102] Based on the second loss function, the trained model is tuned to obtain the multimodal model.

[0103] According to one aspect of this disclosure, an electronic device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the video generation method as described above.

[0104] According to one aspect of this disclosure, a computer-readable storage medium is provided, the storage medium storing a computer program that, when executed by a processor, implements the video generation method described above.

[0105] According to one aspect of this disclosure, a computer program product is provided, comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the video generation method as described above.

[0106] This disclosure embodiment can analyze multiple reference videos uploaded by a video creator using a multimodal large model to extract common style text from the reference videos. Through this common style text, the style commonly adopted by the multiple reference videos that the video creator wants to reference can be clearly identified. Based on this, the multimodal large model can automatically filter images in an image set by combining the common style text to obtain images that match the style of the reference videos the video creator wants to reference, i.e., matching images. This process eliminates the need for the video creator to manually select and filter materials, improving the efficiency of video generation. Furthermore, this disclosure embodiment can also utilize the multimodal model to generate a target video based on multiple matching images. The entire target video generation process does not require the video creator to manually edit the filtered images and the accompanying text, further improving the efficiency of video production.

[0107] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objectives and other advantages of this disclosure may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0108] The accompanying drawings are provided to further understand the technical solutions of this disclosure and constitute a part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.

[0109] Figure 1 This is a system architecture diagram of a video generation method applied according to an embodiment of the present disclosure;

[0110] Figure 2A These are schematic diagrams of a picture set according to embodiments of the present disclosure;

[0111] Figures 2B-2C This is a schematic diagram illustrating an application scenario of the video generation method according to embodiments of the present disclosure in short video production;

[0112] Figures 3A-3B This is a schematic diagram illustrating an application scenario of the video generation method according to embodiments of the present disclosure in the production of work summary videos;

[0113] Figure 4 This is a main flowchart of a video generation method according to an embodiment of the present disclosure;

[0114] Figure 5 This is a schematic diagram of the structure of a multimodal model according to an embodiment of the present disclosure;

[0115] Figure 6 This is a schematic diagram of the structure of a third sub-model according to an embodiment of the present disclosure;

[0116] Figure 7 This is a schematic diagram of the structure of an image-text pair learning model according to an embodiment of the present disclosure;

[0117] Figure 8 This is a schematic diagram of subtitle text information according to an embodiment of the present disclosure;

[0118] Figure 9 This is a schematic diagram illustrating the process of adding text annotations to image information according to an embodiment of the present disclosure.

[0119] Figure 10 This is a schematic diagram illustrating the expansion of the first seed keyword using a thesaurus according to an embodiment of the present disclosure;

[0120] Figure 11 This is a schematic diagram illustrating the expansion of a second seed keyword using a thesaurus according to an embodiment of the present disclosure;

[0121] Figure 12A This is a schematic diagram illustrating how multiple matching images are categorized according to their generation time, based on an embodiment of this disclosure.

[0122] Figure 12B This is a schematic diagram illustrating how matching images under a class are divided into subclasses according to the subject, based on an embodiment of this disclosure.

[0123] Figures 13A-13E This is a schematic diagram illustrating the generation of a target video based on matching images under subclasses according to common video features and common text features, according to an embodiment of this disclosure.

[0124] Figures 14A-14B This is a schematic diagram illustrating the first editing of an anchor object in an anchor image in response to an editing operation on an anchor object in an anchor image in a matching image, according to an embodiment of the present disclosure;

[0125] Figures 14C-14F This is a schematic diagram of generating an edited matching image by performing the same second editing as the first editing on anchor objects in other matching images according to embodiments of this disclosure;

[0126] Figure 15 This is a block diagram of a video generation apparatus according to an embodiment of the present disclosure;

[0127] Figure 16 This is a terminal structure diagram of a video generation method according to an embodiment of the present disclosure;

[0128] Figure 17 This is a server structure diagram of a video generation method according to an embodiment of the present disclosure. Detailed Implementation

[0129] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this disclosure.

[0130] Before providing a further detailed description of the embodiments of this disclosure, the terms and concepts used in these embodiments are explained, and they are subject to the following interpretations:

[0131] Diffusion models are generative models that generate high-resolution images by simulating the diffusion process of matter in a medium. Diffusion models are inspired by non-equilibrium thermodynamics. They define a Markov chain of diffusion steps to slowly add random noise to the data and then learn to reverse the diffusion process, constructing the desired data sample from the noise. Unlike VAEs or flow-based models, diffusion models are learned with a fixed procedure, and the latent variables have high dimensionality (the same as the original data). Diffusion models include forward diffusion and backward diffusion.

[0132] Forward diffusion: Noise is gradually added to the original image until it becomes completely noisy. The original image can be any image. Adding noise to the original image is done in noisy steps. In each noisy step, the image before and after adding noise is recorded for use in back diffusion.

[0133] Backdiffusion: Starting from a completely noisy state, noise is gradually removed to restore a clear image. This process is achieved through a series of denoising steps, which correspond to the respective noise-adding steps. The noise-adding and denoising steps are collectively called diffusion steps. For example, in forward diffusion, the noise-adding steps are 1, 2, ..., T, while in backdiffusion, the denoising steps are T, T-1, ..., 1. Denoising step T corresponds to noise-adding step T, denoising step T-1 corresponds to noise-adding step T-1...

[0134] Contrastive learning: The basic principle of contrastive learning is to train the model by constructing pairs of positive and negative samples. Positive samples are those with similar features, while negative samples are those with different features. By designing a contrastive loss function, the model strives to maximize the similarity between positive samples and minimize the similarity between negative samples.

[0135] With the widespread use of short video platforms in people's lives, more and more video creators are entering the field of video production. Currently, most video creators, after determining the theme of their video, need to manually select relevant materials and edit and transcribe the script. Limited by manual capabilities, this prevents the rapid generation of video works, reducing efficiency and making it difficult to maintain consistency between the video's style and the creator's personal style. Although the market provides some templates for video creators to use in creating videos and related scripts, this method still relies on video creators manually selecting and filtering materials, failing to quickly generate the desired video works and effectively improve production efficiency.

[0136] Based on this, embodiments of this disclosure provide a video generation method, apparatus, electronic device, and medium. The video generation method provided by embodiments of this disclosure can automatically filter images from an image set by combining common style text from multiple reference videos to obtain matching images that conform to the style of the reference videos desired by the video creator. This process eliminates the need for the video creator to manually select and filter materials, improving the efficiency of video generation. Furthermore, embodiments of this disclosure can also utilize a multimodal model to generate a target video based on multiple matching images. The entire target video generation process does not require the video creator to manually edit the filtered images and the accompanying text, supporting rapid generation of target videos and improving the efficiency of video production.

[0137] Figure 1 This is a system architecture diagram of the video generation method according to embodiments of the present disclosure. It includes a terminal 140, an Internet 130, a gateway 120, a server 110, etc.

[0138] Terminal 140 can take various forms, including desktop computers, laptops, PDAs (personal digital assistants), mobile phones, in-vehicle terminals, home theater terminals, and dedicated terminals. Furthermore, it can be a single device or a collection of multiple devices. Terminal 140 can communicate with the Internet 130 via wired or wireless means to exchange data. It is understood that the video generation method of this disclosure embodiment can be deployed on terminal 140 as an application.

[0139] Server 110 refers to a computer system capable of providing certain services to terminal 140. Compared to ordinary terminal 140, server 110 has higher requirements in terms of stability, security, and performance. Server 110 can be a single high-performance computer in a network platform, a cluster of multiple high-performance computers, a portion of a single high-performance computer (e.g., a virtual machine), or a combination of portions of multiple high-performance computers (e.g., virtual machines). In a distributed system, multiple servers 110 are interconnected as computing and storage units to transmit data and coordinate task processing using shared communication lines.

[0140] Gateway 120, also known as an internetwork connector or protocol converter, is a computer system or device that acts as a translator, enabling network interconnection at the transport layer. It bridges the gap between two systems using different communication protocols, data formats, languages, or even completely different architectures. Gateways can also provide filtering and security functions. Messages sent from terminal 140 to server 110 are forwarded to the corresponding server 110 via gateway 120. Messages sent from server 110 to terminal 140 are also forwarded to the corresponding terminal 140 via gateway 120.

[0141] This disclosure can be applied to short video production. A short video is short-duration video content that uses video as a medium and is disseminated through an internet platform. Short videos are typically short, usually within a few minutes. It is understood that the video generation method of this disclosure can be deployed as an application plugin in the user's terminal photo album. (Refer to...) Figure 2A , Figure 2A This is a screenshot of the interface of a photo album application installed on a user's terminal. Users can access the application by clicking the "Tell me your thoughts..." dialog box. Figure 2B The interface shown. Furthermore, users can... Figure 2B Enter the account link of the video creator you want to reference in the dialog box at the bottom (e.g., Figure 2B Users can upload links like "https: / / www.zhanghao1.com" and "https: / / www.zhanghao2.com", or directly upload short videos they want to reference, allowing the application to analyze them. The application summarizes the overall style of the videos from multiple perspectives, including visual presentation, content theme, emotional resonance, and multimedia content, thus indicating areas for learning. Furthermore, users can mark content they want to learn with dotted lines (e.g.,...). Figure 2B The description of "growth-oriented bloggers, cute and sweet, fresh and charming, with a refined lifestyle, skilled in both academics and daily life, documenting their academic progress and daily growth, etc." uses solid lines to mark content that needs to be avoided (e.g.,...). Figure 2B (e.g., "food recommendations"). This process allows the application to clearly identify which specific aspects of content the user wants to refer to. After marking the content they want to learn and the content they want to avoid, users can enter their needs, such as... Figure 2C As shown, this allows the application to filter and organize all the pictures in the user's album based on what the user wants to learn, what the user wants to avoid, and the user's needs, thus obtaining the desired materials. Finally, after the application exports the file based on the user's needs, the user can click... Figure 2C The "Go to View" control in the app allows users to navigate to a dedicated folder organized by the app and view the images filtered by the app.

[0142] This disclosure can also be applied to the production of work summary videos. A work summary video is a summary presented in video format. It can combine various elements such as images, sound, and text to review, analyze, and summarize work progress over a period of time. As mentioned above, users can access the video summary by clicking the "Tell me your thoughts..." dialog box in the interface of the photo album application installed on their terminal. Figure 3A The interface shown. Furthermore, users can... Figure 3A In the dialog box at the bottom, enter the links to the work summary template videos you want to reference (e.g., "https: / / www.muban1.com" and "https: / / www.muban2.com"), or directly upload the work summary template videos you want to reference so that the application can analyze them. Because of the special nature of work summary template videos, the application will summarize the overall style of multiple work summary template videos from multiple perspectives, such as video structure, video focus, and the way the work content is presented, thus indicating directions for users to learn from. Furthermore, users can mark the content they want to learn using dotted lines (e.g., ...). Figure 3A The summary includes sections such as "Introduction: Briefly introduce the time frame and background of the summary. Work Achievements: List in detail the main tasks completed and the results achieved during this period. Lessons Learned: Analyze the problems and challenges encountered during the work process and summarize the lessons learned. Future Planning: Propose future work plans and improvement measures." A solid line is used to mark content that you want to avoid (e.g., ...). Figure 2B The "Goal Setting: Clearly list the work goals for this stage in the introduction. Results Comparison: In the results section, compare the actual results with the goals and analyze the reasons." approach allows the application to clearly identify which aspects the user wants to refer to. After marking the content they want to learn and the content they want to avoid, users can enter their needs, such as... Figure 3BAs shown, this allows the application to filter and organize all work-related images in the user's album based on what the user wants to learn, what the user wants to avoid, and the user's needs, thus obtaining the desired materials. Then, after the application exports the file based on the user's needs, the user can click... Figure 3B Below the chat box that generates a custom folder, there's a "Go to View" control that takes you to the interface of the custom folder organized by the application, allowing you to view the images filtered by the application. Finally, after confirming the filtered images and ensuring they meet your expectations, you can enter your requirements for generating a work summary video. The application will then generate the video based on the filtered images and the learning points and avoidance points marked by the user. After the application exports the work summary video, the user can click... Figure 3B The "Go to View" control below the chat box that generates the work summary video allows you to jump to the interface of the work summary video generated by the application and view it.

[0143] It should be understood that the above description only illustrates some application scenarios of this disclosure. The business scenarios to which this disclosure can be applied may include, but are not limited to, the specific embodiments described above.

[0144] General Description of Embodiments in this Disclosure

[0145] In related technologies, most video creators, after determining the theme of their video, need to manually select relevant materials and edit the script. Limited by human capabilities, this hinders the rapid generation of video works, reducing efficiency and making it difficult to maintain consistency between the video's style and the creator's. While some templates are available for video creators to use to create videos and related scripts, this method still relies on manual material selection and filtering, failing to quickly generate the desired video and effectively improve production efficiency.

[0146] Some embodiments of this disclosure provide a video generation method, apparatus, electronic device, and medium. It can automatically filter images from an image set based on the common style adopted by multiple reference videos to obtain matching images that conform to the style of the reference videos desired by the video creator. This eliminates the need for the video creator to manually select and filter materials, improving the efficiency of video generation. Furthermore, it can utilize a multimodal model to generate the target video. The entire target video generation process does not require the video creator to manually edit the filtered images or add text, supporting rapid target video generation and further improving video production efficiency.

[0147] According to one embodiment of this disclosure, a video generation method is provided. This method can be applied to scenarios such as short video production and work summary video production.

[0148] like Figure 4 As shown, a video generation method according to an embodiment of this disclosure includes:

[0149] Step 410: Obtain the reference video storage space address;

[0150] Step 420: Obtain multiple reference videos from the reference video storage space address;

[0151] Step 430: Using a multimodal model, extract common style text from multiple reference videos;

[0152] Step 440: Obtain the image set, wherein the images in the image set have image tags;

[0153] Step 450: Based on the matching of common style text and image tags, identify multiple matching images in the image set;

[0154] Step 460: Using a multimodal model, generate the target video based on multiple matching images.

[0155] The following is a brief description of steps 410-460 above.

[0156] In step 410, the address of the reference video storage space is obtained.

[0157] As is understandable, a reference video refers to a video that the user wants to refer to. The reference video's storage address can be the storage location of the video on the user's terminal, an internet link to the video, or an internet link to the account of the video's creator. Furthermore, the reference video's storage address is generally provided by the user through typing.

[0158] In step 420, multiple reference videos are obtained from the reference video storage space address.

[0159] According to embodiments of this disclosure, if the reference video storage address is the storage address of the video the user wants to reference on the user's terminal, the reference video can be obtained by reading the storage address of the video on the user's terminal. If the reference video storage address is an internet link to the video the user wants to reference, the reference video can be obtained by redirecting to the internet link of the video. If the reference video storage address is an internet link to the account of the video creator the user wants to reference, multiple reference videos created by that video creator can be obtained by redirecting to the internet link of the video creator's account.

[0160] For example, refer to Figure 2B If "https: / / www.zhanghao1.com" is the internet link to the first video creator's account that the user wants to reference, and "https: / / www.zhanghao2.com" is the internet link to the second video creator's account that the user wants to reference, then after obtaining these two links, the application can redirect to both "https: / / www.zhanghao1.com" and "https: / / www.zhanghao2.com" to access multiple reference videos created by the first video creator and multiple reference videos created by the second video creator that the user wants to reference.

[0161] In step 430, a multimodal model is used to extract common style text from multiple reference videos.

[0162] It is understandable that a multimodal model is a type of artificial intelligence model, capable of processing and fusing various types of input data, such as text, images, audio, and video, to achieve more comprehensive and accurate task understanding and execution. In step 430, the multimodal model can analyze multiple input reference videos to obtain common style text among them. Common style text refers to descriptions of the shared style used across multiple different reference videos. Through this common style text, the style of the reference video the user wants to refer to can be clearly identified.

[0163] For example, if the theme of reference video 1 entered by the user is beauty, sharing daily life and campus life, and the theme of reference video 2 entered by the user is food, sharing daily life and campus life, then it can be concluded that the common style text of reference video 1 and reference video 2 is sharing daily life and campus life.

[0164] In step 440, an image set is obtained, wherein the images in the image set have image tags.

[0165] It is understood that an image set refers to a collection of multiple images stored on a user's terminal. In this embodiment of the disclosure, an image set can also be understood as a collection of all images stored in the user's terminal's photo album. Moreover, each image in the image set has an image tag, which is a tag used to describe the characteristic attributes of the image. An image can have multiple image tags. For example, if image A shows beef noodles, image B shows a university gate, and image C shows a waterfall, then the image tag for image A is "food," the image tag for image B is "university life," and the image tag for image C is "travel and scenery."

[0166] Furthermore, image labels can be obtained using a multimodal model. All images in the image set can be used as input to the multimodal model, allowing it to analyze the feature attributes of each image and output a label for each. Alternatively, image labels can be obtained using a trained label extraction model. When determining image labels, all images in the image set are used as input to the label extraction model, which extracts and analyzes objects in each image, thus outputting a label for each image. The label extraction model is an image processing model used to determine the image labels for each image. It analyzes each object in the image to comprehensively determine the corresponding image label. For example, if the label extraction model extracts objects such as chopsticks, chili oil, a bowl, and beef noodles from image A, then the image label for image A can be determined to be "food." If the label extraction model extracts objects such as sunset, mountains, trees, and rivers from image B, then the image label for image B can be determined to be "tourism" and "scenery." If the label extraction model extracts objects such as a face, lipstick, and eyeshadow palette from image C, then it can be determined that the image label corresponding to image C is "beauty".

[0167] In step 450, based on the matching of common style text and image tags, multiple matching images are identified in the image set.

[0168] Understandably, a matching image refers to an image in a collection that matches the common style text. Matching images are determined by matching the common style text with the image tags. If an image in the collection has tags that match the common style text, it means that the image's characteristics match the style the user wants to refer to, and the image can be identified as a matching image. If the image tags do not match the common style text, it means that the image's characteristics do not match the style the user wants to refer to, and the image can be filtered out. Through this process, images matching the common style text can be automatically filtered, eliminating the need for video creators to manually select and filter materials, thus improving the efficiency of video generation.

[0169] For example, if the common style text is "travel," the image set includes image A tagged with "food," image B tagged with "daily life," image C tagged with "travel and scenery," image D tagged with "work and study," and image E tagged with "travel and food." Based on the matching of the common style text and image tags, images C and E can be identified as matching images.

[0170] In step 460, a target video is generated based on multiple matching images using a multimodal model.

[0171] Understandably, in step 460, the multimodal model can analyze and edit multiple matching input images to obtain the target video. The target video refers to a video generated with reference to the style of a reference video, and it includes background music, subtitles, and text. This content can be learned by the multimodal model from multiple tourism-related videos. This process eliminates the need for video producers to manually edit the selected images and text, thus improving video production efficiency.

[0172] For example, if multiple matching images include multiple images tagged with travel, food, and scenery, then in the target video generated by the multimodal model based on the multiple matching images, the subtitles corresponding to each matching image can be introductions about attractions and food, the background music of the target video can be light and beautiful music, and the text of the target video can be travel insights and experiences.

[0173] According to embodiments of this disclosure, after generating the target video, and with the user's permission and authorization obtained, the application can automatically publish the target video on the user's various session platforms without requiring the user to manually publish it on each session platform.

[0174] The embodiments of steps 410 to 460 described above can analyze multiple reference videos uploaded by the video creator using a multimodal large model to extract common style text from the multiple reference videos. Through the common style text, the style commonly adopted by the multiple reference videos that the video creator wants to reference can be clearly identified. Based on this, the multimodal large model can automatically filter images in the image set in conjunction with the common style text to obtain images that match the style of the reference videos that the video creator wants to reference, i.e., matching images. This process eliminates the need for the video creator to manually select and filter materials, improving the efficiency of video generation. Furthermore, embodiments of this disclosure can also utilize a multimodal model to generate a target video based on multiple matching images. The entire target video generation process does not require the video creator to manually edit the filtered images and the accompanying text, further improving the efficiency of video production.

[0175] The above is a general description of steps 410 to 460. Since steps 410, 420, and 440 have been described sufficiently above, the specific implementation processes of steps 430, 450, and 460 will be described in detail below.

[0176] Detailed description of step 430

[0177] In step 430, a multimodal model is used to extract common style text from multiple reference videos.

[0178] In one embodiment, the multimodal model is generated in the following manner:

[0179] Step 510: Generate a base model based on the first sub-model, the second sub-model, and the third sub-model. The first sub-model is used for mutual generation between texts, the second sub-model is used for mutual generation between text and images, and the third sub-model is used for mutual generation between text and video.

[0180] Step 520: Obtain the video sample set and generate a data sample set based on the video sample set;

[0181] Step 530: Based on the data sample set, train and optimize the basic model to obtain a multimodal model.

[0182] Steps 510 to 530 are described in detail below.

[0183] In step 510, a base model is generated based on the first sub-model, the second sub-model, and the third sub-model. The first sub-model is used for mutual generation between texts, the second sub-model is used for mutual generation between text and images, and the third sub-model is used for mutual generation between text and video.

[0184] Understandably, the base model refers to a model that has undergone large-scale pre-training and can analyze general text, images, and videos. Furthermore, by training and optimizing the base model, a multimodal model can be obtained. Specifically, the first sub-model in the base model refers to a deep learning model used for processing and analyzing text data, and this first sub-model can be used for text-to-text generation. For example, the input to the first sub-model can be an article, and the output can be the sentiment expressed in that article. The input to the first sub-model can also be a question, and the output can be the answer to that question. The second sub-model in the base model refers to a deep learning model used for processing and analyzing visual data (such as images and videos), and this second sub-model can be used for text-to-image generation. For example, the input to the second sub-model can be a photo containing a face, and the output can be the identity corresponding to that face. The input to the second sub-model can also be a specific description, and the output can be an image matching the description. The third sub-model in the base model refers to a deep learning model used for processing and analyzing video, and this third sub-model can be used for text-to-video generation. For example, the input to the third sub-model can be surveillance video, and the output can be a specific person in the surveillance video and their behavior. Alternatively, the input to the third sub-model can be a specific description, and the output can be a video that matches the description.

[0185] For example, refer to Figure 5 Here, we will explain the structure of the base model by examining how it processes the input data: After the data (which can be text, images, or other formats) and the target request are input into the base model, the gating network determines which sub-models will participate in the current task based on the input data and the target request. (If the target request is to analyze the viewpoint in an article and the input data is text, then the third sub-model will participate. If the target request is to identify objects in an image and the input data is an image, then the second sub-model will participate. If the target request is to analyze the viewpoint expressed in a video and the video's visual expression, then the first, second, and third sub-models will participate.) Weights are then assigned to each sub-model participating in the current task. Each sub-model then analyzes the input data and outputs its results. The final output of the base model is obtained by weighting the weights assigned by the gating network to the sub-models participating in the current task and combining the outputs of the sub-models participating in the current task.

[0186] For example, refer to Figure 6Here, the structure of the third sub-model is explained by examining its processing of the input video data: After the video data is input, the encoder encodes each frame of the video data, forming an intermediate representation of the latent space corresponding to the video data. Then, the video compression network performs dimensionality reduction on this intermediate representation, resulting in dimensionality-reduced video data. Next, each frame of the dimensionality-reduced video data is divided into small blocks (each block records the temporal and spatial information of the video frame at that location), and noise is continuously added to these blocks until each block is converted into a completely noisy image. Then, using the diffusion model and the instructions from the image-text pair learning model, the completely noisy images are progressively denoised to output a clear representation of the video frame in the latent space. The decoder then processes this clear representation to reconstruct the clear video frame image. Through this process, the third sub-model clearly defines how to process the video.

[0187] Furthermore, this is combined with Figure 7 The process of the image-text pair learning model processing the input data is explained below: The text processing model receives the labels "tea set, ceramic cup, wooden tea tray," "stationery, metal cup, stone tea tray," and the text description "A hand is pouring tea into a ceramic cup, the ceramic cup is..." as input, and extracts features from these texts. The image processing model receives the image of "A hand is pouring tea into a ceramic cup" and extracts features from the image. Then, a binary cross-entropy loss function is used to compare the difference between the probability distribution predicted by the text processing model and the probability distribution of the true labels, and a logistic regression loss function is used to compare the difference between the probability predicted by the image processing model and the actual labels. The text processing model and the image processing model are optimized through their respective loss functions to learn to better extract features from the input data and perform classification. Finally, by combining the features of the text and the image, the image-text pair learning model can learn the association between image and text pairs (e.g., understanding that the ceramic cup in the input image and the "ceramic cup" in the text description belong to the same category). Understandably, image-text pair learning models can be applied to scenarios such as image captioning generation, visual question answering, image-based content search, and augmented reality applications.

[0188] The specific method for "generating a base model based on the first sub-model, the second sub-model, and the third sub-model" will be described in detail below.

[0189] In step 520, a video sample set is obtained, and a data sample set is generated based on the video sample set.

[0190] As is understandable, a video sample set refers to a large collection of video content, covering various types such as travel, food, lifestyle, and education. Furthermore, video sample sets can be collected from various session platforms using web scraping techniques or application programming interfaces (APIs).

[0191] According to embodiments of this disclosure, the video sample sets are often sourced from diverse sources, with varying video quality and styles, and containing a large amount of irrelevant content such as watermarks and advertisements. Therefore, these video sample sets cannot be directly used for training the base model and require processing to generate a data sample set suitable for training the base model. Here, the data sample set refers to a collection of data that has had irrelevant content removed and can be directly used for training the base model.

[0192] The specific method for “obtaining a video sample set and generating a data sample set based on the video sample set” will be described in detail below.

[0193] In step 530, the base model is trained and optimized based on the data sample set to obtain a multimodal model.

[0194] According to embodiments of this disclosure, training refers to the process of optimizing a base model using a data sample set. By training the base model, its performance on certain specific tasks can be improved, such as classifying input videos or generating videos based on input text and images. Tuning refers to the process of further optimizing the loss function of the base model based on the data sample set after training it. By tuning the base model, the accuracy of the data output by the base model can be further improved. It is understood that after training and tuning the base model, the current base model can be determined as a multimodal model.

[0195] The specific method for "training and optimizing the basic model based on the data sample set to obtain a multimodal model" will be described in detail below.

[0196] The embodiments of steps 510 to 530 described above can generate a base model based on the first sub-model, the second sub-model, and the third sub-model, enabling the base model to process and understand multiple types of data, thereby improving the accuracy of prediction and classification tasks. After generating the base model, embodiments of this disclosure will further train and optimize the base model based on a data sample set to obtain a multimodal model. This process can improve the performance of the base model on certain specific tasks and further enhance the accuracy of the data output by the base model.

[0197] In one embodiment, the first sub-model includes a first encoder and a first decoder, the second sub-model includes a second encoder and a second decoder, and the third sub-model includes a third encoder and a third decoder;

[0198] Step 510 includes:

[0199] Step 610: Connect the outputs of the first encoder, the second encoder, and the third encoder to the input of the central encoder, and connect the output of the central encoder to the inputs of the first decoder, the second decoder, and the third decoder, respectively, to form the basic model.

[0200] Step 610 is described in detail below:

[0201] Understandably, the first encoder is built into the first sub-model and is used to extract text information from the data input to the base model. The second encoder is built into the second sub-model and is used to extract image information from the data input to the base model. The third encoder is built into the third sub-model and is used to extract video information from the data input to the base model. For example, if the data input to the base model is a video introducing Huangshan Mountain, accompanied by background music of a guzheng (Chinese zither), then the first encoder can extract subtitles introducing Huangshan Mountain from the input data, the second encoder can extract each frame of the video from the input data, and the third encoder can extract background music and the time point of each frame in the video from the input data. By setting up the first, second, and third encoders, information from various domains can be quickly captured from the input data.

[0202] Furthermore, the input of the central encoder is connected to the output of the first encoder, the output of the second encoder, and the output of the third encoder. The central encoder is used to integrate and fuse the outputs of the first, second, and third encoders. The output of the central encoder is connected to the inputs of the first, second, and third decoders, respectively, forming the basic model. Specifically, the first decoder converts the output of the central encoder into text when the desired output format is text. The second decoder converts the output of the central encoder into an image when the desired output format is image. The third decoder converts the output of the central encoder into video when the desired output format is video.

[0203] Understandably, the above settings can convert the output of the central encoder into the required output format, ensuring that the base model can produce high-quality output in various fields.

[0204] The embodiment of step 610 described above, by embedding a first encoder in the first sub-model, a second encoder in the second sub-model, and a third encoder in the third sub-model, can quickly capture information from various domains from the input data. Furthermore, by connecting the output of the central encoder to the inputs of the first decoder, the second decoder, and the third decoder, respectively, the output of the central encoder can be converted into the desired output format by each decoder, ensuring that the basic model can produce high-quality output in various domains.

[0205] In one embodiment, step 520 includes:

[0206] Step 710: Extract visual features, audio features, and text features from the video samples in the video sample set;

[0207] Step 720: Obtain the class target label of the video sample;

[0208] Step 730: Based on the visual features, audio features, text features, and class tags of the video samples, generate data samples corresponding to the video samples, thereby obtaining a data sample set based on the video sample set.

[0209] Steps 710 to 730 are described in detail below:

[0210] In step 710, visual features, audio features, and text features are extracted from the video samples in the video sample set.

[0211] As is understandable, a video sample set includes multiple video samples. A video sample refers to a segment of video content, which can be a complete video clip or a short segment extracted from a long video. By analyzing and extracting from video samples, visual features, audio features, and text features can be obtained. Visual features refer to key information used to represent visual content, describing objects, scenes, actions, etc., in video frames. Audio features refer to the characteristics used to describe and quantify the audio data in the video samples, representing different aspects of the audio signal, such as timbre, pitch, loudness, and rhythm. Text features refer to elements or attributes extracted from video samples that characterize the text content. Text features can be used in text analysis, natural language processing tasks, and other fields to help algorithms better understand and process text.

[0212] In one embodiment, before extracting visual, audio, and text features from the video samples in the video sample set, it is necessary to filter out low-quality video samples that have too low a resolution or cannot provide useful information (e.g., severe shaking). Furthermore, after filtering out low-quality video samples, the remaining video samples need to be preprocessed to remove irrelevant content such as watermarks and edited advertisements, thereby improving the overall quality of the video samples.

[0213] The specific methods for "extracting visual features, audio features, and text features from video samples in the video sample set" will be described in detail below.

[0214] In step 720, the class target label of the video sample is obtained.

[0215] Understandably, category tags refer to labels used to indicate the type of a video sample. For example, if video sample A introduces the Huashan Scenic Area, then the category tag for video sample A would be "tourism" and "scenery." If video sample B introduces how to make tiramisu, then the category tag for video sample B would be "cooking" and "pastry." If video sample C introduces one's university life, then the category tag for video sample C would be "daily life" and "study."

[0216] The specific method for "obtaining the class tags of video samples" will be described in detail below.

[0217] In step 730, data samples corresponding to the video samples are generated based on the visual features, audio features, text features, and class tags of the video samples, thereby obtaining a data sample set based on the video sample set.

[0218] Understandably, the data samples corresponding to video samples refer to data that has had irrelevant content removed from the video samples and can be directly used for training the base model. The data samples corresponding to video samples are composed of the visual features, audio features, text features, and class tags of the video samples. This setup allows the base model to learn the visual, audio, and text features corresponding to each class tag, thus facilitating the subsequent generation of videos with each class tag by the base model and ensuring the accuracy of the base model's output.

[0219] Furthermore, since the video sample set includes multiple video samples of different types, and each video sample has a corresponding data sample, a data sample set can be obtained based on multiple video samples in the video sample set. This processing helps the base model learn the visual, audio, and textual features in different types of video samples (e.g., travel, daily life, food, etc.), facilitating the subsequent combination of information from different modalities to generate videos that meet user expectations and improving the generalization ability of the base model.

[0220] The embodiments of steps 710 to 730 described above can generate data samples corresponding to video samples based on the visual features, audio features, text features, and class tags of the video samples. This setup allows the base model to learn the visual, audio, and text features corresponding to each class tag, facilitating the subsequent generation of videos with each class tag and ensuring the accuracy of the base model's output. Furthermore, the data sample set obtained from the video sample set in this disclosure embodiment helps the base model learn the visual, audio, and text features in different types of video samples, facilitating the subsequent combination of information from different modalities to generate videos that meet user expectations, thus improving the generalization ability of the base model.

[0221] In one embodiment, step 710 includes:

[0222] Step 810: Extract audio information and subtitle text information from the video samples in the video sample set;

[0223] Step 820: Extract image information from the video sample and annotate the image information with image text to obtain annotated text information;

[0224] Step 830: Based on the subtitle text information and the annotation text information, integrate them into text information;

[0225] Step 840: Extract visual features from video samples, extract audio features from audio information, and extract text features from text information.

[0226] Steps 810 to 840 are described in detail below:

[0227] In step 810, audio information and subtitle text information are extracted from the video samples in the video sample set.

[0228] Understandably, audio information refers to the background sounds in the video sample. For example, if the video sample depicts a comedian performing crosstalk, and the background sounds are canned laughter and applause, then the audio information extracted from this video sample will be the canned laughter and applause. If the video sample depicts a scene from an urban romance drama, and the background sounds are pop songs, then the audio information extracted from this video sample will be the pop songs.

[0229] Subtitle text information refers to the subtitles in the video sample. Subtitles help the base model better understand the content of the video sample. For example, refer to... Figure 8If the caption at the bottom of the video sample reads "During my drive, I found the roads here to be very empty", then the subtitle text information extracted from this video sample will be "During my drive, I found the roads here to be very empty".

[0230] In step 820, image information is extracted from the video sample, and the image information is annotated with image text to obtain annotated text information.

[0231] As can be understood, image information refers to the frames within each video sample. By breaking down the video sample frame by frame, multiple images arranged chronologically can be obtained. After annotating these images with text, the corresponding annotated text information can be obtained. Image text annotation refers to identifying specific regions within the image information and labeling them based on their feature attributes. The annotated text information refers to the feature attributes corresponding to each specific region in the image information. Furthermore, the annotated text information is not necessarily a specific object; it can also be descriptive terms such as style.

[0232] For example, refer to Figure 9 The image information extracted from the video sample is a photo taken in the living room. After annotating the image information with text, the following elements can be identified: living room, sofa, dog, sitting, lamp, blanket, picture frame, fashion, door, book, stool, etc. Based on this, it can be determined that the annotated text information includes the following elements: living room, sofa, dog, sitting, lamp, blanket, picture frame, fashion, door, book, stool, etc.

[0233] In step 830, the subtitle text information and annotation text information are integrated into text information.

[0234] It is understandable that both subtitle text information and annotation text information belong to text. By integrating subtitle text information and annotation text information into text information, it can help the model to understand the content of video samples more comprehensively and output a more accurate structure.

[0235] For example, if the subtitle text in a video sample includes the sentence, "This time we went on a road trip to Dunhuang and experienced the magnificent desert scenery. On this ancient and mysterious land, we not only felt the weight of history but also experienced the wonders of nature," and the labeled text includes elements such as camels, desert, crowds, majestic mountains, murals, sculptures, and springs, then the sentences "This time we went on a road trip to Dunhuang and experienced the magnificent desert scenery. On this ancient and mysterious land, we not only felt the weight of history but also experienced the wonders of nature" and "camels, desert, crowds, majestic mountains, murals, sculptures, and springs" can be integrated into a single text message.

[0236] In step 840, visual features are extracted from video samples, audio features are extracted from audio information, and text features are extracted from text information.

[0237] According to embodiments of this disclosure, to facilitate subsequent training of the base model on video samples, feature extraction processing is required on the content of video samples, audio information, and text information to convert them into an expression form that the base model can understand. Based on this, firstly, color histograms (typically used to describe the overall color distribution of an image), scale-invariant feature transforms (typically used to detect and describe local features), and deep learning features (typically used to capture high-level semantic information in images) corresponding to each frame of the video sample can be extracted. Then, color histograms, scale-invariant feature transforms, and deep learning features can be identified as visual features, facilitating the base model's understanding of the visual style adopted in the video sample. Simultaneously, Mel-frequency cepstral coefficients (typically used for speech recognition and audio classification) and spectrograms (typically representing the frequency distribution of an audio signal at different time points) can be extracted from the audio sample. Then, Mel-frequency cepstral coefficients and spectrograms can be identified as audio features, facilitating the base model's understanding of the audio style adopted in the video sample. Term frequency (TF) and inverse document frequency (IVF) can be extracted from text information. TF and IVF are typically used to count the frequency of each word in a text and are typically used to assess the importance of a word to a document in a document set or corpus. Then, TF and IVF can be identified as text features, making it easier for the base model to understand the style adopted by the text in the video samples.

[0238] The embodiments of steps 810 to 840 described above can integrate the subtitle text information and annotation text information of the video samples into text information. This processing helps the model to more comprehensively understand the content of the video samples, resulting in a more accurate output structure. Furthermore, embodiments of this disclosure can also extract visual features from the video samples, audio features from the audio information, and text features from the text information, converting the video samples, audio information, and text information into an expression form that the basic model can understand, facilitating subsequent training of the basic model on the video samples.

[0239] In one embodiment, step 720 includes:

[0240] Step 910: Extract metadata from the video sample;

[0241] Step 920: Generate a content description for the video sample;

[0242] Step 930: Input the metadata and content description into the pre-classification model to obtain the initial class target labels;

[0243] Step 940: Receive the adjustment to the initial class target label to obtain the class target label.

[0244] Steps 910 to 940 are described in detail below:

[0245] In step 910, metadata is extracted from the video sample.

[0246] It is understandable that metadata extracted from video samples refers to descriptive information related to the video samples, and metadata can provide a detailed description of the content of the video samples. Furthermore, metadata includes technical metadata (such as resolution, codec, frame rate, etc.), additive metadata (the user who created the video sample, shooting time, recording location, etc.), and structural metadata (the structural relationships between individual video frames in the video sample). Specialized libraries can be used in programming languages ​​to read the metadata from video samples.

[0247] In step 920, a content description of the video sample is generated.

[0248] It is understandable that "content description" refers to a description of the specific content within a video sample. For example, if the specific content of video sample A is an introduction to the various routes within the Jiuzhaigou scenic area, then the content description of video sample A is a travel guide to the Jiuzhaigou scenic area. If the specific content of video sample B is an introduction to how to cook cola chicken wings, then the content description of video sample B is a tutorial on cooking cola chicken wings.

[0249] Furthermore, the content description of the video samples can be obtained from a pre-trained content description model. This content description model is a deep learning model used to generate descriptions of the specific content within the video samples. By inputting the video samples into the content description model, the model will output the corresponding content description. This process eliminates the need for manual review of each video sample, significantly improving the efficiency of determining the content description for each sample.

[0250] In step 930, metadata and content descriptions are input into the pre-classification model to obtain initial class target labels.

[0251] As we can understand, a pre-classification model is a model that allows users to determine the type of a video sample. By inputting the metadata and content description of the video sample into the pre-classification model, the model outputs an initial class label for the video sample. This initial class label represents the preliminary classification result of the video sample. Through the processing of the pre-classification model, manual browsing and summarization of video samples are eliminated, thus improving the efficiency of determining the initial class label for each video sample.

[0252] According to embodiments of this disclosure, a pre-classification model can be trained using a semi-supervised learning method. The process of training a pre-classification model using a semi-supervised learning method is as follows: First, by manually browsing video samples, the category of each video sample is determined based on its content, subject, and style (for example, if a video sample describes a trip to Huangshan, its category can be determined as tourism, scenery, and personal experience). Then, the metadata and content description corresponding to each video sample, as well as the category of the video sample, are recorded, and this recorded data is used as the training set. Afterward, this training set is used to train a preliminary classification model, which is the pre-classification model. By using a semi-supervised learning method to train the pre-classification model, a large amount of unlabeled metadata and content description (i.e., the metadata and content description corresponding to video samples whose categories have not yet been determined) can be utilized, improving data utilization. Moreover, by combining unlabeled metadata and content description with labeled metadata and content description (i.e., the metadata and content description corresponding to video samples whose categories have been determined), the pre-classification model can learn more generalized feature representations, reducing the risk of overfitting.

[0253] In step 940, the adjustment of the initial class target label is received to obtain the class target label.

[0254] Understandably, due to limitations in the training set size of the pre-classification model, the initial class labels obtained through the pre-classification model may not accurately reflect the type of the video sample. Therefore, manual adjustment of erroneous initial class labels is necessary to obtain class labels that correctly reflect the type of the video sample. Here, the class labels can also be understood as the types of video samples obtained after adjusting the initial class labels.

[0255] Furthermore, after adjusting the initial class labels of the video samples to obtain the class labels, the metadata and content descriptions corresponding to the video samples, along with the class labels of the video samples, can be used as a training set and re-input into the pre-classification model to further optimize the model. This process improves the accuracy of the pre-classification model in determining the initial class labels corresponding to video samples, thereby increasing the efficiency of classifying video samples.

[0256] The embodiments described in steps 910 to 940 above can input the metadata and content description of the video samples into the pre-classification model to obtain the initial class tags corresponding to the video samples. Through the processing of the pre-classification model, manual browsing and summarization of the video samples is eliminated, improving the efficiency of determining the initial class tags corresponding to the video samples. Furthermore, it can also receive manual adjustments to the initial class tags to obtain class tags, and use the metadata and content description corresponding to the video samples, as well as the class tags of the video samples, as a training set, and re-input them into the pre-classification model to further optimize the pre-classification model and improve the accuracy of the pre-classification model in determining the initial class tags corresponding to the video samples.

[0257] In one embodiment, step 530 includes:

[0258] Step 1010: Divide the data sample set into a first sample subset and a second sample subset;

[0259] Step 1020: Based on the first sample subset, train the base model to obtain the trained model;

[0260] Step 1030: Based on the second sample subset, optimize the trained model to obtain a multimodal model.

[0261] Steps 1010 to 1030 are described in detail below:

[0262] In step 1010, the data sample set is divided into a first sample subset and a second sample subset.

[0263] Understandably, the first sample subset refers to the dataset used to train the base model. This first sample subset also serves as the input data during the base model's learning process. The base model uses this first sample subset to learn the relationships between visual features, audio features, text features, and class tags, and to adjust its own weights and parameters. The second sample subset refers to the dataset used to evaluate the performance of the trained model during the training process and to adjust the parameters of the trained model. Furthermore, the first and second sample subsets are independent.

[0264] Furthermore, in the process of dividing the data sample set into a first subset and a second subset, the ratio of the first subset to the second subset can be determined based on the size of the data sample set and the task requirements of the base model. Generally, the size of the first subset can be four times or more than the size of the second subset. By dividing the data sample set in this way, overfitting can be effectively avoided, while ensuring the generalization ability of the multi-model approach in practical applications.

[0265] In one embodiment, the data sample set can be further divided into a first sample subset, a second sample subset, and a third sample subset. The third sample subset is a dataset used to evaluate the final performance of the multimodal model, and it is typically used after the multimodal model has been trained and tuned to evaluate its performance on unseen data. Furthermore, when dividing the data sample set into the first, second, and third sample subsets, their proportions can be determined based on the size of the data sample set and the task requirements of the base model. Generally, the first sample subset can be set to 60%-80% of the data sample set, the second sample subset to 10%-20%, and the third sample subset to 10%-20%.

[0266] In step 1020, the base model is trained based on the first sample subset to obtain the trained model.

[0267] As is understandable, a post-trained model refers to the model obtained after training the base model using a first subset of samples. Compared to the base model, the post-trained model can obtain more accurate class tags based on the visual, audio, and textual features of the input.

[0268] According to embodiments of this disclosure, during the training of the base model based on a first sample subset, the first sample subset is first input into the base model to enable it to calculate predicted values. Then, the loss value between the predicted values ​​and the known true labels is calculated to complete forward propagation. Following this, backpropagation is performed. First, the gradient of the loss function with respect to the model parameters of the base model is calculated, and then the model parameters of the base model are updated to reduce the loss value. By repeating the above forward and backpropagation steps, the trained model can be obtained.

[0269] The specific method for "training the base model based on the first sample subset to obtain the trained model" will be described in detail below.

[0270] In step 1030, the trained model is tuned based on the second sample subset to obtain a multimodal model.

[0271] It is understandable that in step 1030, the multimodal model refers to the final model obtained after fine-tuning the trained model using a second subset of samples. Compared to the trained model, the multimodal model can also obtain more accurate class tags based on the input visual features, audio features, and text features.

[0272] According to embodiments of this disclosure, the process of fine-tuning the trained model based on a second sample subset is performed periodically. That is, whenever a forward and backward propagation is performed on the base model based on a first sample subset, or after multiple forward and backward propagations, the parameters of the trained model need to be adjusted using a second sample subset to prevent overfitting.

[0273] The specific method for "optimizing the trained model based on the second sample subset to obtain a multimodal model" will be described in detail below.

[0274] The embodiments of steps 1010 to 1030 described above can train the base model based on a first sample subset to obtain a trained model that can generate more accurate class tags based on the input visual features, audio features, and text features. The trained model can also be periodically tuned based on a second sample subset to obtain the final multimodal model. This process effectively prevents overfitting of the multimodal model.

[0275] In one embodiment, the data samples in the first sample subset include a first visual feature, a first audio feature, a first text feature, and a first type of target tag;

[0276] Step 1020 includes:

[0277] Step 1110: Input the first visual feature, the first audio feature, and the first text feature into the base model to obtain the predicted category;

[0278] Step 1120: Calculate the first loss function based on the predicted category and the first type of target tag;

[0279] Step 1130: Based on the first loss function, train the base model to obtain the trained model.

[0280] Steps 1110 to 1130 are described in detail below:

[0281] In step 1110, the first visual feature, the first audio feature, and the first text feature are input into the base model to obtain the predicted category.

[0282] It is understandable that the data sample set includes visual features, audio features, text features, and class tags. After dividing the data sample set into a first sample subset and a second sample subset, in order to distinguish the data in the first sample subset and the data in the second sample subset, the visual features in the first sample subset can be identified as the first visual features, the audio features in the first sample subset can be identified as the first audio features, and the text features in the first sample subset can be identified as the first text features.

[0283] According to embodiments of this disclosure, during the training of the base model based on a first subset of samples, the first visual features, first audio features, and first text features corresponding to the video samples are first input into the base model, so that the base model predicts the category tags of the video samples based on these data, thereby obtaining the predicted category. The predicted category is the category tag of the video sample predicted by the base model.

[0284] In step 1120, a first loss function is calculated based on the predicted category and the first type of target label.

[0285] It is understandable that the first type of target tag refers to the target tag in the first sample subset. Simultaneously, the first type of target tag is also the target tag corresponding to the first visual feature, the first audio feature, and the first text feature in the first sample subset. The first loss function is a function used to measure the difference between the predicted category and the first type of target tag. Based on the predicted category and the first type of target tag, the first loss function corresponding to the base model can be determined. Furthermore, the first loss function can be used to guide the training process of the base model. By minimizing the first loss function, the accuracy of the base model in determining the predicted category and generating the target video can be improved.

[0286] In step 1130, the base model is trained based on the first loss function to obtain the trained model.

[0287] Understandably, during the training of the base model, the forward propagation process includes steps 1110 and 1120, and the backpropagation process includes calculating the gradient of the first loss function with respect to the model parameters (which may be the weights of the base model), and updating the model parameters of the base model based on the gradient. Then, by iteratively executing the above forward and backpropagation processes to minimize the first loss function, the iteration stops when the first loss function is less than a preset first loss function threshold, or when the number of iterations reaches the first threshold, and the base model at this point is determined as the trained model.

[0288] The embodiments of steps 1110 to 1130 described above can input the first visual feature, the first audio feature, and the first text feature into the base model to obtain the predicted category, and calculate the first loss function based on the predicted category and the first type of target tag. By determining the first loss function, the training process of the base model can be guided. By minimizing the first loss function, the accuracy of the base model in determining the predicted category and generating the target video can be improved. Furthermore, the embodiments of this disclosure will also train the base model based on the first loss function to minimize the first loss function, obtaining the trained model. Through this processing, the base model can learn the relationship between the first visual feature, the first audio feature, the first text feature, and the first type of target tag from the first sample subset, effectively improving the performance of the base model.

[0289] In one embodiment, the data samples in the second sample subset include second visual features, second audio features, second text features, and a second type of target tag;

[0290] Step 1030 includes:

[0291] Step 1210: Obtain the task instance. The task instance includes an instance image and an instance text description.

[0292] Step 1220: Using the trained model, generate instance videos based on instance images and instance text descriptions;

[0293] Step 1230: Identify the video category of the example video;

[0294] Step 1240: Extract the second visual features, the second audio features, and the second text features from the instance image and instance text description;

[0295] Step 1250: Obtain the second type of target tag corresponding to the second visual feature, the second audio feature, and the second text feature from the data samples of the second sample subset;

[0296] Step 1260: Calculate the second loss function based on video category and second target tag;

[0297] Step 1270: Based on the second loss function, optimize the trained model to obtain a multimodal model.

[0298] Steps 1210 to 1270 are described in detail below:

[0299] In step 1210, a task instance is obtained, which includes an instance image and an instance text description.

[0300] According to embodiments of this disclosure, a task instance refers to a task that simultaneously processes and analyzes information from textual, visual, and auditory modalities. By determining task instances, the data understanding and processing capabilities of the trained model can be enhanced. Furthermore, the task instance can be determined based on user needs. For example, if the user's need is to enhance the trained model's ability to generate text and images, then the task instance can be text and its corresponding image. If the user's need is to enhance the trained model's text generation capability, then the task instance can be set to text.

[0301] In one embodiment, the user's requirement is to enhance the post-trained model's ability to generate text, images, and videos. Task instances include instance images and instance text descriptions. Instance images are images used as input to the post-trained model to assist it in understanding visual information. Instance text descriptions are descriptions of the instance images, along with related animations and sound effects. These descriptions are input along with the instance images to the post-trained model to further assist it in understanding the instance images and outputting videos that conform to both the instance images and the instance text descriptions. This approach enhances the data understanding and processing capabilities of the post-trained model.

[0302] In step 1220, the trained model is used to generate an instance video based on the instance image and instance text description.

[0303] It is understandable that by using multimodal modalities to generate instance videos based on instance images and instance text descriptions, deep learning technology can be used to fuse instance images and instance text descriptions to generate dynamic video content that conforms to instance images and instance text descriptions, i.e., instance videos.

[0304] For example, if an instance image shows a egret standing on a riverbank, and the instance text description is "The sound of flowing water is gentle; an egret stands on the riverbank, occasionally looking down to see if there are any small fish in the water," after inputting the instance image and instance text description into the trained model, the trained model will generate a video that matches the instance image and instance text description, namely, an egret standing on the riverbank, occasionally looking down to see if there are any small fish in the water, with the background sound being the sound of flowing water.

[0305] In step 1230, the video category of the instance video is identified.

[0306] It is understandable that the video category refers to the type corresponding to the instance video. For example, as mentioned above, if the content of instance video A is a white egret standing on the riverbank, occasionally looking down at the river to see if there are small fish, and the background sound is the sound of flowing water, then the video category of instance video A is Animals, Nature, Ecology.

[0307] Furthermore, a pre-classification model can be used to identify the video category of an example video. The example video is taken as input and fed into the pre-classification model, which then analyzes and processes the video and outputs the corresponding video category.

[0308] In step 1240, second visual features, second audio features, and second text features are extracted from the instance image and instance text description.

[0309] It is understandable that, since the instance image describes visual features, and the instance text description describes the instance image, as well as the associated animations and sound effects, second visual features, second audio features, and second text features can be extracted from the instance image and instance text description.

[0310] For example, if the instance image is a sunset at the beach, the instance text description is "The clouds in the sky are dyed orange by the setting sun, reflecting on the endless beach. Seagulls circle on the horizon, their cries echoing across the empty beach." Based on this, the image of a sunset at the beach can be extracted from the instance image, i.e., the second visual feature. Second text features such as clouds, sunset, beach, seagulls, and sand can be extracted from the instance text description, and the cries of seagulls can be extracted from the instance text description, identified as the second audio feature.

[0311] In step 1250, the second type of target tag corresponding to the second visual feature, the second audio feature, and the second text feature is obtained from the data samples of the second sample subset.

[0312] It is understandable that the data sample set includes visual features, audio features, text features, and class tags. After dividing the data sample set into a first sample subset and a second sample subset, in order to distinguish the data in the first sample subset and the data in the second sample subset, the visual features in the second sample subset can be identified as the second visual features, the audio features in the second sample subset can be identified as the second audio features, and the text features in the second sample subset can be identified as the second text features.

[0313] Furthermore, the second type of target tag refers to the target tag in the second sample subset. At the same time, the second type of target tag is also the target tag corresponding to the second visual feature, the second audio feature, and the second text feature in the second sample subset.

[0314] In step 1260, a second loss function is calculated based on the video category and the second target tag.

[0315] Understandably, the second loss function is used to measure the difference between video categories and the second target tag. Based on the video categories and the second target tag, the corresponding second loss function for the trained model can be determined. Furthermore, the second loss function can be used to fine-tune the trained model; by minimizing the second loss function, the accuracy of the trained model in generating target videos can be improved.

[0316] In step 1270, the trained model is tuned based on the second loss function to obtain a multimodal model.

[0317] Understandably, in the process of fine-tuning the trained model, the forward propagation process includes using the trained model to generate instance videos based on instance images and instance text descriptions, identifying the video categories of the instance videos, and calculating the second loss function based on the video categories and the second target tag. The backpropagation process includes calculating the gradient of the second loss function with respect to the model parameters (which can be the weights of the trained model) of the trained model, and updating the model parameters of the trained model according to the gradient. Then, by iteratively executing the above forward and backpropagation processes to minimize the second loss function, the iteration stops when the second loss function is less than a preset second loss function threshold, or when the number of iterations reaches the second threshold, thus completing the fine-tuning process of the trained model. The trained model at this point is then identified as a multimodal model.

[0318] The embodiments of steps 1210 to 1220 described above can utilize the trained model to generate instance videos based on instance images and instance text descriptions, identify the video category of the instance videos, and calculate a second loss function based on the video category and a second target tag. By determining the second loss function, the optimization process of the trained model can be guided. By minimizing the second loss function, the accuracy of the target videos generated by the trained model can be improved. Furthermore, embodiments of this disclosure will also optimize the trained model based on the second loss function to minimize it, obtaining a multimodal model. Through this processing, the trained model can fully integrate text information, image information, and audio information, thereby generating more accurate and higher-quality target videos.

[0319] Detailed description of step 450

[0320] In step 450, based on the matching of common style text and image tags, multiple matching images are identified in the image set.

[0321] In one embodiment, step 450 includes:

[0322] Step 1310: Display common style text;

[0323] Step 1320: In response to the positive and negative markers of the displayed common style text, identify the positive and negative marker text in the common style text;

[0324] Step 1330: Filter out the images whose image tags match the negative text to obtain the filtered image set;

[0325] Step 1340: In the filtered image set, identify images whose image labels match the positive tag text to obtain multiple matching images.

[0326] Steps 1310 to 1340 are described in detail below.

[0327] In step 1310, common style text is displayed.

[0328] It is understandable that, such as Figure 2B As shown, after the multimodal model extracts common style text from multiple reference videos input by the user, it will display the common style text on the interface. Through the common style text, users can intuitively understand the style adopted by the reference video they want to refer to.

[0329] In step 1320, in response to positive and negative markers on the displayed common style text, positive marker text and negative marker text are identified in the common style text.

[0330] Understandably, users can not only view common style text through the terminal interface, but also annotate desired learning points and unwanted avoidance points within the common style text using dashed and solid lines respectively. In other words, positive annotations on the displayed common style text represent the annotation processing performed by the user for desired learning points, and the positively labeled text refers to the text of the desired learning points within the common style text. Similarly, negative annotations on the displayed common style text represent the annotation processing performed by the user for unwanted avoidance points, and the negatively labeled text refers to the text of the unwanted avoidance points within the common style text.

[0331] For example, refer to Figure 2B Based on the user's dotted-line annotations (positive markers) of the displayed common style text, the text identified as positive can be categorized as: "Growth-oriented blogger, cute and sweet, fresh and elegant, with a refined lifestyle, skilled in academics and life, recording the process of entering higher education and daily growth, etc."; "Sharing campus life and learning methods, from internship experience and career planning to travel diaries," and "Growth records: They record their growth trajectory, including academic achievements, the development of personal interests, and the establishment of interpersonal relationships, providing fans with a growth blueprint to refer to." Based on the user's solid-line annotations (negative markers) of the displayed common style text, the text identified as negative can be categorized as "food recommendations, etc."

[0332] In step 1330, images whose image labels match the negative text are filtered out, resulting in a filtered image set.

[0333] Understandably, matching image tags with negative text is difficult because image tags are typically in the form of fields, while negative text is usually in the form of long sentences. Therefore, fields can be extracted from the negative text, and images whose image tags match the fields in the negative text can be filtered out. This can also be understood as filtering out images from the image set that use a style the user does not expect. The filtered image set refers to the collection of images remaining after removing images with the undesirable style.

[0334] For example, if the negative label text is "food recommendations, etc.", the field "food" can be extracted from the negative label text, and all images in the image set labeled "food" can be filtered out. The set of images remaining in the image set is then determined as the filtered image set.

[0335] The specific method for "filtering out images whose image tags match the negative text to obtain the filtered image set" will be described in detail below.

[0336] In step 1340, in the filtered image set, images whose image labels match the positive label text are identified, resulting in multiple matching images.

[0337] Understandably, as mentioned above, matching image tags with positive tag text is difficult because image tags are generally in the form of fields, while positive tag text is usually in the form of long sentences. Therefore, fields can be extracted from the positive tag text, and images whose image tags match the fields in the positive tag text can be selected from the filtered image set. This can also be understood as selecting images from the filtered image set that use the style the user expects. These matching images are the images in the filtered image set that use the style the user expects.

[0338] For example, if the positive tag text is "The current account has posted a large number of videos featuring scenery from various countries", the field "scenery" can be extracted from the positive tag text, and images with the tag "scenery" can be selected from the filtered image set as matching images.

[0339] The specific method for "identifying images whose image labels match the positive text in the filtered image set to obtain multiple matching images" will be described in detail below.

[0340] The embodiments of steps 1310 to 1340 described above can filter out images in the image set whose image tags match negative markers and identify images whose image tags match positive marker text to obtain multiple matching images. This processing allows for quick and automatic selection of materials based on the user's desired style, eliminating the need for manual selection and filtering by the user, thus improving the efficiency of video generation.

[0341] In one embodiment, step 1330 includes:

[0342] Step 1410: Extract the first seed keyword from the negatively labeled text;

[0343] Step 1420: Expand the first seed keyword using the thesaurus to obtain the first expanded keyword;

[0344] Step 1430: Filter out images from the image set whose image tags match the first expanded keyword to obtain the filtered image set.

[0345] Steps 1410 to 1430 are described in detail below.

[0346] In step 1410, the first seed keyword is extracted from the negatively labeled text.

[0347] Understandably, the first seed keyword refers to a meaningful field extracted from the negatively labeled text. For example, if the negatively labeled text is "Most of the content in their accounts is related to games," then the first seed keyword extracted from the negatively labeled text is "games." If the negatively labeled text is "They record their beauty tips," then the first seed keyword extracted from the negatively labeled text is "beauty."

[0348] In step 1420, the first seed keyword is expanded using a thesaurus to obtain the first expanded keyword.

[0349] Understandably, a thesaurus is a collection or database containing a set of synonyms and related words. Thesauruses are typically used to provide equivalents or similar expressions for words. Expanding the first seed keyword using a thesaurus means identifying synonyms or near-synonyms of the first seed keyword. The first seed keyword, along with its synonyms or near-synonyms, constitutes the first expanded keyword. Expanding the first seed keyword using a thesaurus helps enhance its diversity and expressive power, improving the accuracy of subsequent image filtering within the image set.

[0350] For example, refer to Figure 10 If the negatively labeled text is "food recommendations, etc.", after extracting the first seed keyword "food" from the negatively labeled text, a thesaurus can be used to expand "food" to obtain the first expanded keywords "food", "delicacies", "delicious", "rare delicacies", and "fine dishes".

[0351] In step 1430, images whose image tags match the first expanded keyword are filtered out from the image set to obtain the filtered image set.

[0352] It is understandable that images whose image tags match the first expanded keywords are images in the image set that use a style that the user does not expect. The embodiments of this disclosure obtain a filtered image set by filtering out images in the image set whose image tags match the first expanded keywords. This automatically filters out images that do not meet the user's expectations without relying on the user, thus improving the efficiency of generating the target video.

[0353] The embodiments of steps 1410 to 1430 described above can extract a first seed keyword from the negatively labeled text and expand the first seed keyword using a thesaurus to obtain a first expanded keyword. This process helps enhance the diversity and expressiveness of the first seed keyword, improving the accuracy of subsequent image filtering in the image set. Furthermore, the embodiments of this disclosure can automatically filter out images that do not meet user expectations (i.e., images whose image tags match the first expanded keyword) without relying on the user, improving the efficiency of generating the target video.

[0354] In one embodiment, step 1340 includes:

[0355] Step 1510: Extract the second seed keyword from the positively marked text;

[0356] Step 1520: Expand the second seed keyword using the thesaurus to obtain the second expanded keyword;

[0357] Step 1530: In the filtered image set, identify images whose image tags match the second expanded keywords to obtain multiple matching images.

[0358] Steps 1510 to 1530 are described in detail below.

[0359] In step 1510, the second seed keyword is extracted from the positively labeled text.

[0360] Understandably, the second seed keyword refers to a meaningful field extracted from the positive tag text. For example, if the positive tag text is "They record their own growth trajectory," then the second seed keyword extracted from the positive tag text would be "growth trajectory." If the positive tag text is "Usually using soft colors and a fresh layout," then the second seed keywords extracted from the positive tag text would include "soft" and "fresh."

[0361] In step 1520, the second seed keyword is expanded using a thesaurus to obtain the second expanded keyword.

[0362] Understandably, expanding the second seed keyword using a thesaurus refers to identifying synonyms or near-synonyms of the second seed keyword through the thesaurus. The second seed keyword, along with its synonyms or near-synonyms, constitutes the expanded second seed keyword. Expanding the second seed keyword using a thesaurus helps enhance its diversity and expressiveness, improving the accuracy of subsequent image selection from the filtered image set.

[0363] For example, refer to Figure 11 If the positive tag text is: "Campus life, learning method sharing...", after extracting the second seed keywords "campus life" and "learning method" from the positive tag text, for the second seed keyword "campus life", we can use a thesaurus to expand "campus life" to obtain the second expanded keywords "campus life", "school life", "campus daily life", "campus time", and "campus experience".

[0364] In step 1530, in the filtered image set, images whose image tags match the second expanded keywords are identified, resulting in multiple matching images.

[0365] It is understood that the images whose image tags match the second expanded keywords are images from the filtered image set that use the style desired by the user. The embodiments of this disclosure, by identifying images whose image tags match the second expanded keywords, obtain multiple matching images. This allows for automatic filtering of images that meet the user's expectations without relying on the user, thus improving the efficiency of generating the target video.

[0366] The embodiments of steps 1510 to 1530 described above can extract second seed keywords from positively labeled text and expand the second seed keywords using a thesaurus to obtain second expanded keywords. This process helps enhance the diversity and expressiveness of the second seed keywords, improving the accuracy of subsequent image filtering in the filtered image set. Furthermore, the embodiments of this disclosure can automatically filter images that meet user expectations (i.e., images whose image tags match the second expanded keywords) without relying on the user, improving the efficiency of generating the target video.

[0367] Detailed description of step 460

[0368] In step 460, a target video is generated based on multiple matching images using a multimodal model.

[0369] In one embodiment, step 460 includes:

[0370] Step 1610: Using a multimodal model, based on multiple matching images, generate the target video according to the positive and negative labeled text.

[0371] Step 1610 will be described in detail below.

[0372] Understandably, when generating the target video, the multimodal model matches corresponding visual, audio, and text elements to multiple matching images based on the description of the positive tag text. For example, if the positive tag text is "usually use soft colors and a clean layout," the multimodal model can add soft or clean filters to multiple matching images and determine the background music to be gentle and soothing light music, so that the final generated target video matches the expectations of the positive tag text.

[0373] Accordingly, when generating the target video, the multimodal model will, based on the description of the negatively labeled text, try to avoid using visual, audio, and text elements related to the negatively labeled text in the target video. For example, if the negatively labeled text is "They usually use dark-toned photos with eerie background music," the multimodal model will avoid adding dark filters to multiple matching images and avoid using eerie background music when processing multiple matching images, in order to avoid these elements that do not meet the user's expectations.

[0374] The embodiment of step 1610 described above can, when generating the target video, match corresponding visual elements, audio elements, and text elements for multiple matching images based on the description of the positively labeled text, and, based on the description of the negatively labeled text, try to avoid using visual elements, audio elements, and text elements related to the negatively labeled text in the target video. Through this processing, the target video can meet the user's expectations, ensuring the quality of the target video.

[0375] In one embodiment, step 460 includes:

[0376] Step 1710: Divide the multiple matching images into categories according to their generation time;

[0377] Step 1720: For each class, divide the matching images under the class into subclasses according to the subject.

[0378] Step 1730: Display multiple matching images according to categories, where, under each category, display the subcategories into which the matching images under that category are divided;

[0379] Step 1740: In response to the selection of the displayed subclass, generate the target video based on the matching images under the subclass.

[0380] Steps 1710 to 1740 are described in detail below.

[0381] In step 1710, the multiple matching images are divided into categories according to the generation time of the matching images.

[0382] It's understandable that multiple matching images can be categorized according to the time period to which they were generated. Here, the category of a matching image refers to the time period it corresponds to.

[0383] For example, if a user's request is to organize all photos of themselves and their partner over the past five years of marriage and generate a target video in chronological order, then after identifying multiple matching images in the image set based on matching common style text and image tags, the matching images can be categorized according to their shooting time into classes such as "first year of marriage," "second year of marriage," "third year of marriage," "fourth year of marriage," and "fifth year of marriage."

[0384] For example, refer to Figure 12A If the user's request is to publish content about their university years in a timeline, starting from the date they received their university admission notice, then after identifying multiple matching images in the image set based on matching common style text and image tags, these images can be categorized according to their shooting time, such as "Before University," "First Year (First Semester)," "Winter Break," and "First Year (Second Semester)."

[0385] In one embodiment, multiple matching images can also be categorized according to their image tags, where the category corresponding to a matching image refers to the image tag associated with that matching image. For example, if a user's request is to edit photos from their daily life and generate a video combining elements of food, travel, and beauty, then after identifying multiple matching images in the image set based on matching common style text and image tags, these images can be categorized into "food," "travel," and "beauty" categories according to their corresponding image tags.

[0386] In step 1720, for each class, the matching images under the class are divided into subclasses according to the subject.

[0387] Understandably, multiple matching images under a class can be divided into subclasses based on the subject and description within the matching image. Here, the subclass corresponding to a matching image refers to the description of the subject within that matching image.

[0388] For example, refer to Figure 12BAs mentioned above, if the user's request is to publish content about their university years chronologically, starting from the time they receive their university acceptance letter, after categorizing multiple matching images into "Before University," "First Year (First Semester)," "Winter Break," and "First Year (Second Semester)," we can see that for the "Before University" category, the matching images can be further divided into subcategories based on their subject, such as "Checking Notices," "Eating and Drinking," "Back-to-School Preparations," "Choosing Notebooks," and "Studying Hard."

[0389] In step 1730, multiple matching images are displayed according to categories, wherein, under each category, the subcategories into which the matching images under that category are displayed.

[0390] It is understandable that, such as Figure 12A As shown, after the application categorizes multiple matching images according to their generation time, it displays the matching images to the user according to their categories, allowing the user to view images taken in each time period and quickly check the matching images. For example... Figure 12B As shown, under each category, the subcategories of the matched images under that category will also be displayed, so that users can quickly understand the content of the matched images.

[0391] In step 1740, in response to the selection of the displayed subclass, a target video is generated based on the matching images under the subclass.

[0392] Understandably, if the number of matching images is too large, resulting in overly cluttered content, users can select the displayed subclasses to streamline the target video. This allows the multimodal model to generate the target video based on the matching images within the user-selected subclasses. This reduces redundant content in the target video, making it more aligned with user expectations and improving its quality.

[0393] For example, refer to Figure 12B As mentioned above, for the "Before Going to University" category, the matching images under this category can be divided into subcategories based on their subject, such as "Checking Notices," "Eating and Drinking," "Back-to-School Preparation," "Choosing a Notebook," and "Studying Hard." If a user finds the subcategories within the "Before Going to University" category too complex and selects the "Back-to-School Preparation" subcategory, the multimodal model will generate the target video based on the matching images under this subcategory.

[0394] The specific method for "generating the target video based on the matching image under the subclass in response to the selection of the displayed subclass" will be described in detail below.

[0395] The embodiments of steps 1710 to 1740 described above can classify multiple matching images according to their generation time. For each class, the matching images within that class are further divided into subclasses based on their main content. This processing allows users to quickly understand the content of the matching images. Furthermore, embodiments of this disclosure can also generate a target video based on the matching images within the subclasses in response to the user's selection of the displayed subclasses. This reduces redundant content in the target video, making it more in line with user expectations and improving its quality.

[0396] In one embodiment, step 1740 includes:

[0397] Step 1810: In response to the selection of the displayed subclass, obtain the matching image under the subclass;

[0398] Step 1820: Using a multimodal model, extract common video features and common text features from multiple reference videos;

[0399] Step 1830: Using a multimodal model, based on the matching images under the subclass, generate the target video according to common video features and common text features.

[0400] Steps 1810 to 1830 are described in detail below.

[0401] In step 1810, in response to the selection of the displayed subclass, the matching image under the subclass is obtained.

[0402] Understandably, after a user selects a subclass for display, the multimodal model will respond to the user's selection by retrieving the matching image for that subclass. Here, the subclass corresponding to the matching image refers to the description of the subject within that matching image; all matching images within the subclass are those with the same subject description.

[0403] In step 1820, a multimodal model is used to extract common video features and common text features from multiple reference videos.

[0404] It is understandable that common video features refer to the same video features used across multiple different reference videos. For example, if reference video A introduces a self-driving tour of western Sichuan and reference video B introduces the natural scenery of Tibet, it can be determined that the common video features of reference videos A and B are the use of bright and clear colors, the capture of magnificent and beautiful scenery, and the use of light and upbeat instrumental music as background music in both reference videos A and B.

[0405] Understandably, common textual features refer to the same textual features used across multiple different reference videos. For example, if reference video C introduces Zhu Ziqing's essay "Hurry," and reference video D recites Dai Wangshu's poem "Rainy Alley," it can be determined that the common textual feature of reference videos C and D is the use of poetic language, possessing a high degree of literary quality.

[0406] In step 1830, using a multimodal model, a target video is generated based on the matching images under the subclass, according to common video features and common text features.

[0407] Understandably, when generating the target video, the multimodal model fuses matching images with common video features and common text features. This can be understood as the multimodal model matching the user's desired visual, audio, and text elements to the matching images within a subclass based on common video and text features. This process ensures the final generated target video conforms to the user's desired style, thus improving the video's quality.

[0408] For example, if a user selects the subcategory "Sunset at the Seaside," and the common video features are bright, clear colors and upbeat instrumental music as background music, and the common text features are poetic language, then the multimodal model, based on the matching images under the "Sunset at the Seaside" subcategory, will generate a target video according to the common video and text features. The target video's subtitles and text will use poetic language to describe the sunset scene at the seaside (e.g., the clouds are dyed into a magnificent sunset glow, from deep purple to crimson, then to orange-yellow, layer upon layer, ever-changing. The sun gradually sinks into the sea, its rounded outline gradually disappearing in the waves, finally leaving only a magnificent glow illuminating half the world). Furthermore, the background music of the target video is upbeat instrumental music, and the filter for the matching images in the target video will be processed to be bright and clear.

[0409] The specific method of "using a multimodal model to generate target videos based on matching images under subclasses and according to common video features and common text features" will be described in detail below.

[0410] The embodiments of steps 1810 to 1830 described above can utilize a multimodal model to extract common video features and common text features from multiple reference videos, and process the matching images under subcategories according to these common video and text features to generate the target video. This processing allows for matching the user's desired visual, audio, and text elements to the matching images under subcategories, ensuring that the final generated target video conforms to the user's desired style and improving the quality of the target video.

[0411] In one embodiment, step 1830 includes:

[0412] Step 1910: Using a multimodal model, based on the matching images under the subclass, generate video scripts and video narration according to common text features;

[0413] Step 1920: Display the video script and generate the edited video script in response to editing operations on the video script;

[0414] Step 1930: Display video narration and, in response to editing operations on the video narration, generate the edited video narration;

[0415] Step 1940: Display the matching images under the subclass, and in response to the editing operation on the matching images, generate the edited matching images;

[0416] Step 1950: Using a multimodal model, based on the edited matching images and edited video narration, generate a target video according to common video features, and display the target video in association with the edited video text.

[0417] Steps 1910 to 1950 are described in detail below.

[0418] In step 1910, using a multimodal model, based on the matching images under the subclass, video scripts and video narration are generated according to common text features.

[0419] Understandably, video scripts refer to the textual descriptions written for a target video, which can enhance the video's appeal and expressiveness. Video narration refers to the spoken explanations or commentary added to the target video; it can be used to explain, describe, or comment on the content. During the generation of the target video, to ensure it better matches the user's desired style, the styles of the video scripts and narration need to conform to common textual characteristics.

[0420] For example, refer to Figures 13A to 13D If a user selects the "European Tourism" subcategory, the common text features are "good at describing scenery" and "using beautiful language." Based on this, using a multimodal model, and matching images under the "European Tourism" subcategory, the video script generated according to the common text features is as follows: Figure 13A As shown, this includes phrases like "Only after visiting Europe do you understand where their relaxed lifestyle comes from...". Based on matching images under the subcategory "European Tourism," video narration generated according to common text includes phrases like "I went to Europe once, and I finally understand where their relaxed lifestyle comes from...". Figure 13B , Figure 13C and Figure 13D As shown, it can be obtained that in Figure 13BThe video accompanying the matched image features the Eiffel Tower, with the caption "The Eiffel Tower is very tall." Figure 13C The video accompanying the matched image is set in Venice, and the accompanying narration reads, "We are in the canals of Venice." Figure 13D The video, which matches the image of a car speeding along an empty road, has a caption that reads "The road is very empty."

[0421] In step 1920, the video script is displayed, and in response to editing operations on the video script, the edited video script is generated.

[0422] Understandably, if a user is dissatisfied with the video text corresponding to the matched image generated by the multimodal model based on common text features, they can click on the following: Figure 13A The "Edit" control in the interface allows users to modify the video script, essentially editing it. The modified script can then be set as the final edited video script. Furthermore, after setting the final edited script, it needs to replace the original video script.

[0423] For example, refer to Figure 13A Regarding the video script, "Only after visiting Europe will you understand where their relaxed lifestyle comes from...", if users don't like the phrase "slow living is an art" in the video article, they can click... Figure 13A The "Edit" control in the interface was used to remove the phrase "Slow living is an art" from the video script. As a result, the edited video script no longer contains the phrase "Slow living is an art" compared to the original script.

[0424] In step 1930, video narration is displayed, and in response to editing operations on the video narration, edited video narration is generated.

[0425] Understandably, if a user is dissatisfied with the video narration corresponding to the matched image generated by the multimodal model based on common text features, they can click on the following: Figure 13B The "Edit" control at the top of the interface allows you to modify the video narration corresponding to the matched image, essentially editing the video narration. You can then set the modified narration as the edited version. Furthermore, after setting the edited narration, you need to replace the original narration with it.

[0426] For example, refer to Figures 13A to 13D Users can click Figure 13A To access the video, please refer to the following: Figures 13B to 13D The video narration editing interface has a "Edit" control at the top of each video narration editing interface. (This is for...) Figure 13CThe narration in the video "We are in the waterways of Venice" can be modified if the user believes it needs to be changed. Figure 13C Access the video narration editing interface via the "Edit" control at the top of the screen. If the user changes the video narration to "We stroll through the canals of Venice," then "We stroll through the canals of Venice" can be selected as the edited video narration.

[0427] In step 1940, the matching images under the subclass are displayed, and in response to the editing operation on the matching images, the edited matching images are generated.

[0428] Understandably, if a user is not satisfied with the matched images under a subcategory, they can click on an image such as... Figure 13E The "All" control allows users to view all matching images under a subcategory. Users can click on unsatisfactory matching images to modify them, effectively editing the images. The modified image is then designated as the edited matching image. Furthermore, after designating the edited matching image, it is used to replace the original matching image.

[0429] For example, refer to Figure 13C and Figure 13E If the user thinks Figure 13C The matching image in the image is not bright enough; you can improve it by clicking... Figure 13E The "All" control allows you to view all matching images under the "European Tourism" subcategory, and click on them as shown in the image. Figure 13C Matching images to Figure 13C Perform editing operations, Figure 13C The color tone has been adjusted to be brighter. The image with the adjusted color tone can then be selected as the edited matching image.

[0430] The specific method for "displaying the matching images under the subclass and generating the edited matching images in response to editing operations on the matching images" will be described in detail below.

[0431] In step 1950, using a multimodal model, based on the edited matching images and the edited video narration, a target video is generated according to common video features, and the target video is displayed in association with the edited video text.

[0432] Understandably, when generating the target video, the multimodal model fuses the edited matching image and the edited video narration according to common video features to generate the target video. This can also be understood as using the edited video narration as the narration for the edited matching image, and matching the user's desired visual, audio, and text elements to the edited matching image based on common video features. This process allows users to flexibly modify the content of the target video, facilitating personalized settings.

[0433] Furthermore, in this embodiment, after generating the target video, the target video is displayed in association with the edited video script, that is, the target video and the edited video script are displayed as a related whole. Based on this, the attractiveness and expressiveness of the target video can be enhanced through the edited video script.

[0434] The embodiments of steps 1910 to 1950 described above can support editing the video text, video narration, and matching images separately to generate edited video text, edited video narration, and edited matching images. This processing allows users to flexibly modify the content of the target video, facilitating personalized settings. Furthermore, embodiments of this disclosure can also utilize a multimodal model to generate a target video based on the edited matching images and edited video narration, according to common video features, and associate the target video with the edited video text for display, thereby enhancing the attractiveness and expressiveness of the target video through the edited video text.

[0435] In one embodiment, step 1940 includes:

[0436] Step 2010: In response to the editing operation on the anchor object in the anchor image of the matching image, perform the first editing on the anchor object in the anchor image;

[0437] Step 2020: Identify anchor objects in other matching images besides the anchor image in the matching images;

[0438] Step 2030: Perform the same second edit as the first edit on the anchor objects in other matching images to generate the edited matching images.

[0439] Steps 2010 to 2030 are described in detail below.

[0440] In step 2010, in response to an editing operation on the anchor object in the anchor image of the matching image, a first edit is performed on the anchor object in the anchor image.

[0441] As we can understand, an anchor image is a matching image that serves as the baseline for editing. In other words, if you edit the anchor object in an anchor image, modifications to other matching images containing anchor objects will be reflected in the changes made to the anchor image. The anchor object can be a specific item in the anchor image, or it can be a filter, lighting, or color element used to adjust the visual effects of the image.

[0442] Furthermore, editing operations on anchor objects in an anchor image can be user-defined operations such as deleting, zooming in or out, or modifying. Moreover, the first edit performed on an anchor object in an anchor image is similar to a user-defined edit, the only difference being that the user-defined edit is performed by the user, while the first edit is performed by the multimodal model.

[0443] For example, if matching image 1, matching image 2, matching image 3, and matching image 4 are identified in the image set, and all four matching images contain a cup, and matching image 1 is designated as the anchor image, then if the cup in matching image 1 is designated as the anchor object, the user can eliminate the cup in matching image 1 by clicking on the controls on the screen. Simultaneously, matching images 2, 3, and 4 will also simultaneously eliminate the cup. Similarly, if the filter in matching image 1 is designated as the anchor object, the user can adjust the filter in matching image 1 to a clean / fresh style filter by clicking on the controls on the screen. Simultaneously, matching images 2, 3, and 4 will also simultaneously adjust their filters to a clean / fresh style filter.

[0444] In step 2020, anchor objects in other matching images besides the anchor image are identified.

[0445] According to embodiments of this disclosure, after identifying the anchor object in the anchor image, the anchor object in the other matching images can be identified by extracting features from other matching images besides the anchor image and then comparing these features with the anchor object in the anchor image. Furthermore, a convolutional neural network can be used to extract features from the other matching images.

[0446] The specific method for "identifying anchor objects in other matching images besides the anchor image in the matching image" will be described in detail below.

[0447] In step 2030, the same second edit as the first edit is performed on the anchor objects in the other matching images to generate the edited matching images.

[0448] Understandably, the editing operations performed in the second edit are the same as those performed in the first edit. The only difference between the second and first edits is the subject being edited. The first edit modifies the anchor object in the anchor image, while the second edit modifies the anchor objects in other matching images. The edited matching image refers to the image obtained after performing the second edit on the anchor objects in the other matching images. By performing the same second edit on the anchor objects in the other matching images as the first edit, visual consistency can be ensured, greatly improving the video quality of the target video.

[0449] For example, refer to Figures 14A to 14F ,exist Figure 14A It's an anchor image. Figure 14A If the charging port is an anchor object, and the user edits the terminal screen to... Figure 14A Eliminating the charging port in the middle will result in the following: Figure 14B The modified anchor image is shown below. If the user... Figure 14A When eliminating the charging port in the image, the "Apply to all objects" control was checked, and the "Confirm" control was clicked. After obtaining the modified anchor image, the multimodal model will recognize it. Figure 14C and Figure 14E The two images show the charging ports, and each is paired with... Figure 14C and Figure 14E The charging port in the middle performs a cancellation operation to obtain Figure 14D and Figure 14F These two images were edited and matched.

[0450] The specific method for “performing a second edit on anchor objects in other matching images, the same as the first edit, to generate an edited matching image” will be described in detail below.

[0451] The embodiments of steps 2010 to 2030 described above can identify anchor objects in other matching images besides the anchor image in the matching images, and perform the same second editing as the first editing on the anchor objects in the other matching images to obtain the edited matching images. Through this process, the visual effects of other matching images can be modified on a large scale by referring to the visual effect of the anchor image without relying on manual editing, so that the visual effects of the other matching images are consistent with the visual effects of the anchor image. This not only ensures visual consistency and greatly improves the video quality of the target video, but also increases the efficiency of producing the target video.

[0452] In one embodiment, step 2020 includes:

[0453] Step 2110: Display the image range input control;

[0454] Step 2120: In the image range input control, receive the input image range;

[0455] Step 2130: Within the range of input images, identify anchor objects in other matching images besides the anchor image.

[0456] Steps 2110 to 2130 are described in detail below.

[0457] In step 2110, the image range input control is displayed.

[0458] It is understandable that the image range refers to the selection range of matching images other than the anchor image. The image range can be matching images other than the anchor image within a certain time period, matching images other than the anchor image within a folder divided according to image descriptions, or matching images other than the anchor image from multiple video frames in a video.

[0459] Furthermore, the image range input control refers to a control provided to the user for inputting an image range. Typically, the image range input control appears after the user performs the first edit on the anchor object in the anchor image, allowing the user to determine other matching images besides the anchor image by inputting the image range. This facilitates subsequent large-scale modification of the visual effects of other matching images by referencing the visual effect of the anchor image.

[0460] In step 2120, the input image range is received in the image range input control.

[0461] Understandably, after clicking the image range input control on the interface, the user can type in the image range, allowing the multimodal model to receive the user-inputted image range. Furthermore, the user can also click on... Figure 12A The class, or such Figure 12B A subclass of [subclass name], used to select a range of input images.

[0462] In step 2130, within the range of input images, anchor objects in other matching images besides the anchor image are identified.

[0463] Understandably, when a multimodal model receives an image from the user, it will first extract multiple matching images within the range of the user's input image and identify anchor objects in these matching images, so as to perform a second editing on the anchor objects in other matching images outside the anchor images.

[0464] For example, if the user enters an image range of 2020.01.01-2021.01.01 in the image range input control, the multimodal model will extract all matching images whose generation time is between 2020.01.01 and 2021.01.01, and identify the anchor objects in the matching images. If the user enters an image range of the subclass "European Tourism" in the image range input space, the multimodal model will extract all matching images under the subclass "European Tourism" to identify the anchor objects in these matching images.

[0465] The specific method for "identifying anchor objects in other matching images outside the anchor image within the range of input images" will be described in detail below.

[0466] The embodiments of steps 2110 to 2130 described above can identify anchor objects in other matching images besides the anchor image within the range of images input by the user, so as to facilitate the subsequent second editing of anchor objects in other matching images besides the anchor image, without requiring the user to manually select them in the matching images, thus improving the efficiency of generating the target video.

[0467] In one embodiment, step 2130 includes:

[0468] Step 2210: Use a multimodal model to perform image text annotation on the matching images within the input image range, and obtain multiple annotation information corresponding to each matching image;

[0469] Step 2220: Determine the confidence level of each annotation in the matching image based on the matching images within the image range and the multiple annotation information corresponding to the matching images;

[0470] Step 2230: Determine the anchor object confidence of the anchor object in the matching image based on the confidence of each annotation information in the matching image;

[0471] Step 2240: If the confidence score of the anchor object in the matching image is greater than the first confidence score threshold, identify the anchor object in the matching image.

[0472] Steps 2210 to 2240 are described in detail below.

[0473] In step 2210, a multimodal model is used to perform image text annotation on the matching images within the input image range, resulting in multiple annotation information corresponding to each matching image.

[0474] It is understandable that image text annotation for matching images refers to extracting features and main objects from the matching images. The annotation information refers to the features and main objects extracted from the matching images. For example, if the input image range includes matching image A, which depicts a sunset and mountains, and uses a sharp filter, then the main objects can be extracted from matching image A as the sunset and mountains, and the feature is the sharp filter. Therefore, the sunset, mountains, and sharp filter can be identified as the annotation information corresponding to the matching image.

[0475] In step 2220, the confidence level of each annotation in the matching image is determined based on the matching images within the image range and the multiple annotation information corresponding to the matching images.

[0476] Understandably, the confidence level of annotation information is used to indicate the degree of certainty the multimodal model has about that annotation information. A higher confidence level indicates that the annotation information is closer to the real subject or feature, and therefore more reliable; conversely, a lower confidence level indicates that the annotation information is less reliable. Based on this, by determining the confidence level of each annotation in the matching images, we can determine the approximate subject and features in each matching image. Matching images with similar subject and features will also have similar levels of confidence for each annotation and its corresponding confidence level.

[0477] For example, if the matching images within the image range include matching image A, matching image B, and matching image C. Matching image A has annotations for a cat, a dog, and a lawn, with a confidence level of 90% for the cat, 79% for the dog, and 80% for the lawn. Matching image B has annotations for coffee, a coffee table, and a cat, with a confidence level of 90% for the coffee, 60% for the coffee table, and 85% for the cat. Matching image C has annotations for a cat, a dog, and a lawn, with a confidence level of 80% for the cat, 89% for the dog, and 95% for the lawn. It can be seen that the annotations and corresponding confidence levels in matching images A and C are closer. Therefore, compared to matching images A and B, or B and C, the subject objects and features in matching images A and C are closer.

[0478] In step 2230, the anchor object confidence of the anchor object in the matching image is determined based on the confidence of each annotation information in the matching image.

[0479] Understandably, anchor object confidence refers to the degree of certainty the multimodal model has about an anchor object. A higher anchor object confidence indicates that the anchor object in the matching image is closer to the anchor object in the anchor image, allowing the anchor object to be extracted and further processed. Conversely, a lower anchor object confidence indicates that the anchor object in the matching image is less similar to the anchor object in the anchor image, suggesting a significant difference between the anchor object identified by the multimodal model and the anchor object in the anchor image.

[0480] In step 2240, if the confidence of the anchor object in the matching image is greater than the first confidence threshold, the anchor object in the matching image is identified.

[0481] Understandably, the first confidence threshold refers to the critical value at which the confidence level of the labeled information is determined to be unreliable. If the confidence level of the labeled information is less than or equal to the first confidence threshold, it indicates that the labeled information annotated by the multimodal model is unreliable. Conversely, if the confidence level of the labeled information is greater than the first confidence threshold, it indicates that the labeled information annotated by the multimodal model is reliable. Based on this, the reliability of the anchor objects identified by the multimodal model from the matched images can be determined by matching the confidence level of the anchor objects in the matched images with the first confidence threshold. If the confidence level of the anchor objects in the matched images is greater than the first confidence threshold, then the anchor objects identified by the multimodal model from the matched images are considered reliable.

[0482] For example, with a first confidence threshold of 90% and the anchor object being a laptop, if the confidence level for the laptop in matching image A is 80%, the confidence level for the laptop in matching image B is 92%, and the confidence level for the laptop in matching image C is 95%, it can be determined that the anchor objects identified by the multimodal model from matching images B and C are reliable, and a second edit can be performed on the anchor objects in matching images B and C.

[0483] The embodiments of steps 2210 to 2240 described above can determine the anchor object confidence of the anchor object in the matching image based on the confidence of each annotation information in the matching image, and identify the anchor object in the matching image when the anchor object confidence of the matching image is greater than a first threshold. This process eliminates the need for users to manually select the objects to be modified in the matching image, improving the efficiency of generating the target video.

[0484] In one embodiment, the first editing is an elimination process, and step 2030 includes:

[0485] Step 2310: Mask the anchor objects in other matching images to obtain the masked areas;

[0486] Step 2320: Remove the image content within the masked area to obtain the edited matching image.

[0487] Steps 2310 to 2320 are described in detail below.

[0488] In step 2310, the anchor objects in other matching images are masked to obtain the masked area.

[0489] Understandingly, masking anchor objects involves using a two-dimensional array of the same size as the anchor object to segment the anchor objects in the matching image. Each element in the two-dimensional array represents the visibility of the corresponding pixel of the anchor object in the matching image (e.g., 1 for visible, 0 for hidden), and the masking area refers to this two-dimensional array of the same size as the anchor object.

[0490] By masking anchor objects in the matching image, the multimodal model can be guided to focus on the anchor objects in the matching image, making it easier to remove the anchor objects in the subsequent process.

[0491] In step 2320, the image content within the masked area is removed to obtain the edited matching image.

[0492] Understandably, after determining the masking region in the matching image, the multimodal model can control the masking effect by modifying the values ​​of the two-dimensional array of the masking region, thereby eliminating the image content within the masking region and removing anchor objects in the matching image to obtain the edited matching image.

[0493] The embodiments of steps 2310 to 2320 described above, by masking the anchor objects in other matching images and eliminating the images within the masked area, can not only flexibly eliminate anchor objects, but also automate the processing of a large number of matching images, so that the visual effects of the matching images and the anchor images are consistent, thereby improving the generation efficiency of the target video.

[0494] Description of apparatus and devices according to embodiments of this disclosure

[0495] It is understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated in this embodiment, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0496] It should be noted that in various specific embodiments of this application, when processing data related to object characteristics, such as object attribute information or sets of attribute information, is required, the object's permission or consent will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require obtaining object attribute information, separate permission or consent from the object will be obtained through pop-ups or redirection to a confirmation page. Only after obtaining the object's separate permission or consent will the necessary object-related data for the proper functioning of these embodiments be acquired.

[0497] Figure 15 This is a schematic diagram of the structure of a video generation apparatus 1500 provided in an embodiment of the present disclosure. The video generation apparatus 1500 includes:

[0498] The first acquisition unit 1510 is used to acquire the storage space address of the reference video.

[0499] The second acquisition unit 1520 is used to acquire multiple reference videos from the reference video storage space address;

[0500] Extraction unit 1530 is used to extract common style text from multiple reference videos using a multimodal model;

[0501] The third acquisition unit 1540 is used to acquire an image set, wherein the images in the image set have image tags;

[0502] Unit 1550 is used to determine multiple matching images in the image set based on the matching of common style text and image tags;

[0503] The generation unit 1560 is used to generate a target video based on multiple matching images using a multimodal model.

[0504] Optionally, the determining unit 1550 is specifically used for:

[0505] Display common style text;

[0506] In response to positive and negative markers in the displayed common style text, identify positively marked text and negatively marked text in the common style text;

[0507] Images whose image tags match the negative text are filtered out, resulting in a set of filtered images.

[0508] In the filtered image set, images whose image labels match the positive tag text are identified, resulting in multiple matching images.

[0509] Optionally, the determining unit 1550 is specifically used for:

[0510] Extract the first seed keyword from the negatively labeled text;

[0511] The first seed keyword is expanded using a thesaurus to obtain the first expanded keyword;

[0512] From the image set, images whose image tags match the first expanded keyword are filtered out, resulting in the filtered image set.

[0513] Optionally, the determining unit 1550 is specifically used for:

[0514] Extract the second seed keyword from the positively labeled text;

[0515] The second seed keyword is expanded using a thesaurus to obtain the second expanded keyword;

[0516] In the filtered image set, images whose image tags match the second expanded keywords are identified, resulting in multiple matching images.

[0517] Optionally, the generating unit 1560 is specifically used for:

[0518] Using a multimodal model, a target video is generated based on multiple matching images and according to positive and negative labeled text.

[0519] Optionally, the generating unit 1560 is specifically used for:

[0520] Based on the generation time of the matching images, multiple matching images are divided into categories;

[0521] For each category, the matching images under that category are further divided into subcategories based on their subject.

[0522] Display multiple matching images according to categories, where within each category, display the subcategories into which the matching images within that category are divided;

[0523] In response to the selection of the displayed subclass, the target video is generated based on the matching images under the subclass.

[0524] Optionally, the generating unit 1560 is specifically used for:

[0525] In response to the selection of a subclass to display, retrieve the matching image under that subclass;

[0526] Using a multimodal model, common video features and common text features of multiple reference videos are extracted from multiple reference videos;

[0527] Using a multimodal model, target videos are generated based on matching images under subclasses, according to common video features and common text features.

[0528] Optionally, the generating unit 1560 is specifically used for:

[0529] Using a multimodal model, based on matching images under subclasses, video scripts and video narration are generated according to common text features;

[0530] Displays the video script and generates the edited video script in response to editing operations on the video script;

[0531] Displays video narration and generates edited video narration in response to editing operations on the video narration;

[0532] Displays the matching images under the subclass, and generates the edited matching images in response to editing operations on the matching images;

[0533] Using a multimodal model, based on the edited matching images and edited video narration, a target video is generated according to common video features, and the target video is displayed in association with the edited video text.

[0534] Optionally, the generating unit 1560 is specifically used for:

[0535] In response to an edit operation on the anchor object in the anchor image of the matching image, perform the first edit on the anchor object in the anchor image;

[0536] Identify anchor objects in other matching images besides the anchor image in the matching image;

[0537] For anchor objects in other matching images, perform the same second edit as the first edit to generate the edited matching images.

[0538] Optionally, the generating unit 1560 is specifically used for:

[0539] Display image range input control;

[0540] The image range input control receives the input image range.

[0541] Within the range of input images, identify anchor objects in other matching images besides the anchor image.

[0542] Optionally, the generating unit 1560 is specifically used for:

[0543] A multimodal model is used to perform image text annotation on matching images within the range of the input image, resulting in multiple annotation information corresponding to each matching image;

[0544] The confidence level of each annotation in the matching image is determined based on the matching images within the image range and the multiple annotation information corresponding to the matching images.

[0545] The anchor object confidence of the anchor object in the matching image is determined based on the confidence of each annotation information in the matching image.

[0546] If the confidence score of the anchor object in the matching image is greater than the first confidence threshold, the anchor object in the matching image is identified.

[0547] Optionally, the first editing is an elimination process, and the generation unit 1560 is specifically used for:

[0548] Mask the anchor objects in other matching images to obtain the masked areas;

[0549] The image content within the masked area is removed to obtain the edited matching image.

[0550] Optionally, the extraction unit 1530 is specifically used for:

[0551] Based on the first sub-model, the second sub-model, and the third sub-model, a base model is generated. The first sub-model is used for mutual generation between texts, the second sub-model is used for mutual generation between text and images, and the third sub-model is used for mutual generation between text and video.

[0552] Obtain a video sample set and generate a data sample set based on the video sample set;

[0553] Based on the data sample set, the basic model is trained and optimized to obtain a multimodal model.

[0554] Optionally, the first sub-model includes a first encoder and a first decoder, the second sub-model includes a second encoder and a second decoder, and the third sub-model includes a third encoder and a third decoder;

[0555] Extraction unit 1530 is specifically used for:

[0556] The outputs of the first encoder, the second encoder, and the third encoder are connected to the input of the central encoder, and the output of the central encoder is connected to the inputs of the first decoder, the second decoder, and the third decoder, respectively, to form the basic model.

[0557] Optionally, the extraction unit 1530 is specifically used for:

[0558] Visual features, audio features, and text features are extracted from video samples in the video sample set.

[0559] Obtain the class and target tags of the video samples;

[0560] Based on the visual features, audio features, text features, and class tags of video samples, data samples corresponding to the video samples are generated, thereby obtaining a data sample set based on the video sample set.

[0561] Optionally, the extraction unit 1530 is specifically used for:

[0562] Extract audio information and subtitle text information from video samples in the video sample set;

[0563] Image information is extracted from video samples, and image text is annotated to obtain annotated text information;

[0564] Based on the subtitle text information and the annotation text information, they are integrated into text information;

[0565] Visual features are extracted from video samples, audio features are extracted from audio information, and text features are extracted from text information.

[0566] Optionally, the extraction unit 1530 is specifically used for:

[0567] Extract metadata from video samples;

[0568] Content description of the generated video samples;

[0569] Input the metadata and content description into the pre-classification model to obtain the initial class target labels;

[0570] Receive adjustments to the initial class object tag to obtain the class object tag.

[0571] Optionally, the extraction unit 1530 is specifically used for:

[0572] The data sample set is divided into a first sample subset and a second sample subset;

[0573] Based on the first sample subset, the base model is trained to obtain the trained model;

[0574] Based on the second sample subset, the trained model is optimized to obtain a multimodal model.

[0575] Optionally, the data samples in the first sample subset include first visual features, first audio features, first text features, and a first type of target tag;

[0576] Extraction unit 1530 is specifically used for:

[0577] Input the first visual feature, the first audio feature, and the first text feature into the base model to obtain the predicted category;

[0578] Calculate the first loss function based on the predicted category and the first type of target tag;

[0579] The base model is trained based on the first loss function to obtain the trained model.

[0580] Optionally, the data samples in the second sample subset include second visual features, second audio features, second text features, and second type of target tags;

[0581] Extraction unit 1530 is specifically used for:

[0582] Retrieve a task instance, which includes an instance image and an instance text description;

[0583] Using the trained model, instance videos are generated based on instance images and instance text descriptions;

[0584] Identify the video category of the example video;

[0585] Extract second visual features, second audio features, and second text features from the example images and example text descriptions;

[0586] From the data samples of the second sample subset, obtain the second type of target tag corresponding to the second visual feature, the second audio feature, and the second text feature;

[0587] Calculate the second loss function based on video category and second target tag;

[0588] Based on the second loss function, the trained model is tuned to obtain a multimodal model.

[0589] Reference Figure 16 , Figure 16 To implement the structural block diagram of the terminal portion of the video generation method according to the embodiments of this disclosure, the terminal includes: a radio frequency (RF) circuit 1610, a memory 1615, an input unit 1630, a display unit 1640, a sensor 1650, an audio circuit 1660, a wireless fidelity (WiFi) module 1670, a processor 1680, and a power supply 1690, etc. Those skilled in the art will understand that... Figure 16 The terminal structure shown does not constitute a limitation on mobile phones or computers and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0590] The RF circuit 1610 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 1680; in addition, it transmits uplink data to the base station.

[0591] The memory 1615 can be used to store software programs and modules. The processor 1680 executes various functional applications and data processing of the content terminal by running the software programs and modules stored in the memory 1615.

[0592] The input unit 1630 can be used to receive input numeric or character information, and to generate key signal inputs related to the settings and function control of the content terminal. Specifically, the input unit 1630 may include a touch panel 1631 and other input devices 1632.

[0593] The display unit 1640 can be used to display input or provided information, as well as various menus of the content terminal. The display unit 1640 may include a display panel 1641.

[0594] Audio circuitry 1660, speaker 1661, and microphone 1662 provide an audio interface.

[0595] In this embodiment, the processor 1680 included in the terminal can execute the video generation method of the previous embodiment.

[0596] The terminals disclosed in this embodiment include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, and aircraft. The embodiments of this invention can be applied to various scenarios, including but not limited to short video production and work summary video production.

[0597] Figure 17 This is a partial structural block diagram of a server for implementing the video generation method of this disclosure. The server can vary significantly due to different configurations or performance, and may include one or more central processing units (CPUs) 1722 (e.g., one or more processors) and memory 1732, and one or more storage media 1730 (e.g., one or more mass storage devices) for storing application programs 1742 or data 1744. The memory 1732 and storage media 1730 may be temporary or persistent storage. The program stored in the storage media 1730 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server. Furthermore, the CPU 1722 may be configured to communicate with the storage media 1730 and execute the series of instruction operations in the storage media 1730 on the server.

[0598] The server may also include one or more power supplies 1726, one or more wired or wireless network interfaces 1750, one or more input / output interfaces 1758, and / or one or more operating systems 1741, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0599] The central processing unit 1722 in the server can be used to execute the video generation method of the present disclosure embodiments.

[0600] This disclosure also provides a computer-readable storage medium for storing program code for executing the video generation methods of the foregoing embodiments.

[0601] This disclosure also provides a computer program product comprising a computer program. A processor of a computer device reads and executes the computer program, causing the computer device to perform the video generation method described above.

[0602] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar terms and are not necessarily used to describe a particular order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “including,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.

[0603] It should be understood that in this disclosure, "at least one item" refers to one or more items, and "more than one item" refers to two or more items. "And / or" is used to describe the relationship between related content, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related content are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0604] It should be understood that in the description of the embodiments disclosed herein, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.

[0605] In the embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0606] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0607] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0608] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server 130, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.

[0609] It should also be understood that the various implementation methods provided in this disclosure can be combined arbitrarily to achieve different technical effects.

[0610] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.

Claims

1. A video generation method, characterized in that, include: Obtain the storage address of the reference video; Multiple reference videos are obtained from the reference video storage space address; Using a multimodal model, common style text is extracted from the multiple reference videos; Obtain an image set, wherein the images in the image set have image tags; Based on the matching of the common style text and the image tags, multiple matching images are determined in the image set; Using the multimodal model, a target video is generated based on the multiple matched images.

2. The video generation method according to claim 1, characterized in that, The process of matching the common style text with the image tags to determine multiple matching images in the image set includes: Display the aforementioned common style text; In response to positive and negative markers on the displayed common style text, positive and negative marker text is identified in the common style text; Images whose image tags match the negative text are filtered out to obtain a set of filtered images; In the filtered image set, images whose image tags match the positive label text are identified, thus obtaining the multiple matching images.

3. The video generation method according to claim 2, characterized in that, The step of filtering out the images whose image tags match the negative text to obtain a filtered image set includes: Extract the first seed keyword from the negatively labeled text; The first seed keyword is expanded using a thesaurus to obtain the first expanded keyword; From the image set, images whose image tags match the first expanded keyword are filtered out to obtain the filtered image set.

4. The video generation method according to claim 2, characterized in that, The step of identifying images whose image tags match the positive label text in the filtered image set, and obtaining the plurality of matching images, includes: Extract the second seed keyword from the positively labeled text; The second seed keyword is expanded using a thesaurus to obtain the second expanded keyword; In the filtered image set, images whose image tags match the second expanded keywords are identified to obtain the plurality of matching images.

5. The video generation method according to claim 1, characterized in that, The step of generating a target video based on the multiple matched images using the multimodal model includes: The multiple matching images are divided into categories according to their generation time. For each category, the matching images under that category are divided into subcategories according to the subject. The multiple matching images are displayed according to the class, wherein, under each class, the subclasses into which the matching images under the class are divided are displayed; In response to the selection of the displayed subclass, the target video is generated based on the matching image under the subclass.

6. The video generation method according to claim 5, characterized in that, In response to a selection of the displayed subclass, generating the target video based on the matching image under the subclass includes: In response to the selection of the displayed subclass, the matching image under the subclass is obtained; Using the multimodal model, common video features and common text features of the multiple reference videos are extracted from the multiple reference videos; Using the multimodal model, the target video is generated based on the matching images under the subclass, according to the common video features and the common text features.

7. The video generation method according to claim 6, characterized in that, The step of generating the target video using the multimodal model, based on the matching images under the subclass, according to the common video features and the common text features, includes: Using the multimodal model, based on the matched images under the subclass, and according to the common text features, video scripts and video narration are generated; Display the video text and, in response to editing operations on the video text, generate an edited video text; Display the video narration, and generate an edited video narration in response to an editing operation on the video narration; Display the matching image under the subclass, and generate the edited matching image in response to the editing operation of the matching image; Using the multimodal model, based on the edited matched image and the edited video narration, the target video is generated according to the common video features, and the target video is displayed in association with the edited video text.

8. The video generation method according to claim 7, characterized in that, The step of generating an edited matching image in response to an editing operation on the matching image includes: In response to the editing operation on the anchor object in the anchor image of the matching image, a first edit is performed on the anchor object in the anchor image; Identify the anchor object in other matching images besides the anchor image in the matching image; For the anchor objects in the other matching images, perform the same second edit as the first edit to generate the edited matching image.

9. The video generation method according to claim 8, characterized in that, The step of identifying the anchor object in other matching images besides the anchor image in the matching images includes: Display image range input control; The image range input control receives the input image range. Within the range of the input images, identify the anchor object in other matching images besides the anchor image.

10. The video generation method according to claim 1, characterized in that, The multimodal model is generated in the following way: Based on the first sub-model, the second sub-model, and the third sub-model, a base model is generated, wherein the first sub-model is used for mutual generation between texts, the second sub-model is used for mutual generation between text and images, and the third sub-model is used for mutual generation between text and video. Obtain a video sample set, and generate a data sample set based on the video sample set; Based on the data sample set, the basic model is trained and optimized to obtain the multimodal model.

11. The video generation method according to claim 10, characterized in that, The first sub-model includes a first encoder and a first decoder, the second sub-model includes a second encoder and a second decoder, and the third sub-model includes a third encoder and a third decoder; The generation of the base model based on the first sub-model, the second sub-model, and the third sub-model includes: The outputs of the first encoder, the second encoder, and the third encoder are connected to the input of the central encoder, and the output of the central encoder is connected to the inputs of the first decoder, the second decoder, and the third decoder, respectively, to form the basic model.

12. The video generation method according to claim 10, characterized in that, The step of acquiring a video sample set and generating a data sample set based on the video sample set includes: Visual features, audio features, and text features are extracted from video samples in the video sample set. Obtain the class target label of the video sample; Based on the visual features, audio features, text features, and class tags of the video samples, data samples corresponding to the video samples are generated, thereby obtaining the data sample set based on the video sample set.

13. The video generation method according to claim 12, characterized in that, The step of extracting visual features, audio features, and text features from video samples in the video sample set includes: Audio information and subtitle text information are extracted from the video samples in the video sample set; Image information is extracted from the video sample, and image text is annotated on the image information to obtain annotated text information; Based on the subtitle text information and the annotation text information, they are integrated into text information; The visual features are extracted from the video samples, the audio features are extracted from the audio information, and the text features are extracted from the text information.

14. The video generation method according to claim 12, characterized in that, The process of obtaining the class tag of the video sample includes: Extract metadata from the video samples; Generate a content description for the video sample; Input the metadata and content description into the pre-classification model to obtain the initial class target label; The initial class label is adjusted to obtain the class label.

15. The video generation method according to claim 10, characterized in that, The process of training and optimizing the base model based on the data sample set to obtain the multimodal model includes: The data sample set is divided into a first sample subset and a second sample subset; Based on the first sample subset, the base model is trained to obtain the trained model; Based on the second sample subset, the trained model is optimized to obtain the multimodal model.

16. The video generation method according to claim 15, characterized in that, The data samples in the second sample subset include second visual features, second audio features, second text features, and second type of target tags; The step of optimizing the trained model based on the second sample subset to obtain the multimodal model includes: Obtain a task instance, which includes an instance image and an instance text description; Using the trained model, an instance video is generated based on the instance image and the instance text description; Identify the video category of the example video; Extract the second visual feature, the second audio feature, and the second text feature from the example image and the example text description; From the data samples of the second sample subset, obtain the second type of target tag corresponding to the second visual feature, the second audio feature, and the second text feature; Based on the video category and the second target tag, calculate the second loss function; Based on the second loss function, the trained model is tuned to obtain the multimodal model.

17. A video generation apparatus, characterized in that, The video generation device includes: The first acquisition unit is used to acquire the storage space address of the reference video. The second acquisition unit is used to acquire multiple reference videos from the reference video storage space address; The extraction unit is used to extract common style text from the multiple reference videos using a multimodal model; The third acquisition unit is used to acquire an image set, wherein the images in the image set have image tags; The determining unit is used to determine multiple matching images in the image set based on the matching of the common style text and the image tags; The generation unit is used to generate a target video based on the multiple matching images using the multimodal model.

18. An electronic device comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the video generation method according to any one of claims 1 to 16.

19. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the video generation method according to any one of claims 1 to 16.

20. A computer program product comprising a computer program, characterized in that, The computer program is read and executed by the processor of the computer device, causing the computer device to perform the video generation method according to any one of claims 1 to 16.