Target theme video generation method and device, intelligent agent and electronic equipment

By generating promotional copy using a large model and combining it with object images, and selecting relevant materials from a video material library for editing and compositing, the problem of material shortage in live video streaming, advertising production, and e-commerce sales is solved, video quality is improved, and computing resource consumption is reduced.

CN121509782APending Publication Date: 2026-02-10BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511494785.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In scenarios such as live video streaming, advertising production, and e-commerce sales, existing technologies lack materials for specific target audiences when generating videos on target themes, resulting in poor marketing effectiveness. Furthermore, generative AIGC video production processes consume high computational resources and take a long time to generate.

Method used

The promotional text is generated by creating a large model, the insertion position of the materials is marked, relevant materials are selected from the video material library by combining the object image, and then arranged and synthesized to generate a target theme video.

Benefits of technology

It improves the logic and richness of video content, reduces computing resource consumption and generation time, and ensures that the visual materials of the target audience are consistent with the main text of the promotional material.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509782A_ABST
    Figure CN121509782A_ABST
Patent Text Reader

Abstract

The invention provides a target theme video generation method and device, an intelligent agent and electronic equipment, relates to the technical field of artificial intelligence, in particular to the technical fields of deep learning, large models, computer vision and the like, and can be applied to scenes such as video live broadcast, advertisement production, e-commerce sales and the like. The specific implementation scheme is as follows: obtaining object information of a target propaganda object, wherein the object information comprises an object name and an object image; generating a propaganda copywriting based on the object name through the large model; annotating the insertion position of the propaganda material in the propaganda copywriting to obtain propaganda position annotation information; selecting a target video material irrelevant to the target propaganda object from a video material library according to the propaganda copywriting and the propaganda position labeling information; according to the propaganda copywriting, the propaganda position labeling information and the object image, obtaining an object video material related to the target propaganda object; and according to the propaganda copywriting, the propaganda position labeling information, the target video material and the object video material, generating a target theme video of the target propaganda object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of natural language processing, deep learning, large models, and computer vision. Specifically, it relates to a method, apparatus, intelligent agent, and electronic device for generating target-themed videos, which can be applied to scenarios such as live video streaming, advertising production, and e-commerce sales. Background Technology

[0002] With the rapid development of internet technology, especially mobile internet and high-speed communication technology, video content has become a core medium for information dissemination, commercial promotion, and user interaction. In many scenarios such as live video streaming, advertising production, and e-commerce sales, high-quality, highly engaging targeted videos are key elements for attracting user attention, enhancing brand image, and ultimately achieving commercial conversion.

[0003] However, in the current technical practice of generating such promotional videos, especially in scenarios requiring rapid, mass production, there are still many technical bottlenecks and challenges. Summary of the Invention

[0004] This disclosure provides a method, apparatus, intelligent agent, and electronic device for generating target-themed videos.

[0005] According to a first aspect of this disclosure, a method for generating a target-themed video is provided, comprising: Obtain object information of the target advertising object, the object information including object name and object image; The large model generates promotional copy for the target promotional object based on the object name; The insertion positions of promotional materials in the aforementioned promotional text are marked to obtain promotional position marking information; Based on the promotional copy and the promotional location marking information, select target video materials from the video material library that are unrelated to the target promotional object; Based on the promotional copy, the promotional location marking information, and the object image, obtain object video materials related to the target promotional object; The target theme video of the target promotional object is obtained by arranging and synthesizing the promotional copy, the promotional location marking information, the target video material, and the object video material.

[0006] According to a second aspect of this disclosure, a target-themed video generation apparatus is provided, comprising: The acquisition module is used to acquire object information of the target advertising object, the object information including object name and object image; The copywriting management module is used to generate promotional copy for the target promotional object based on the object name using a large model; The copy management module is also used to mark the insertion positions of promotional materials in the promotional copy to obtain promotional position marking information; The retrieval module is used to select target video materials unrelated to the target promotional object from the video material library based on the promotional copy and the promotional location marking information; The object material enhancement module is used to obtain object video materials related to the target promotional object based on the promotional copy, the promotional location annotation information, and the object image; The editing and compositing module is used to edit and compose the promotional copy, the promotional location marking information, the target video material, and the object video material to obtain the target theme video of the target promotional object.

[0007] According to a third aspect of this disclosure, an intelligent agent is provided, comprising: The input module is used to receive input information; The processing module is used to determine the target task based on the input information received by the input module, determine the large model based on the target task, and execute the method described in the first aspect above by calling the large model to obtain output information; The output module is used to output the output information obtained by the processing module.

[0008] According to a fourth aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect above.

[0009] According to a fifth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first aspect above.

[0010] According to a sixth aspect of this disclosure, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in the first aspect above.

[0011] According to the technical solution disclosed herein, the problem of poor marketing effectiveness caused by the lack of specific promotional materials (such as products) in the traditional AIGC video production process of acquisition and editing in specific scenarios such as live video streaming, advertising production, and e-commerce sales can be solved. This allows the object visual materials in the generated target-themed video to be consistent with the main text related to the target promotional object, thereby improving the logic and richness of the video content. Furthermore, by improving upon the traditional acquisition and editing AIGC video production process, the problem of high computational resource costs (such as the memory usage, computational load, and communication overhead of the model during the inference stage) and long generation time in the generative AIGC video production process in specific scenarios such as live video streaming, advertising production, and e-commerce sales can be solved.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 A flowchart of a method for generating target-themed videos provided in this embodiment of the disclosure; Figure 2 A flowchart of a method for generating target-themed videos provided in this embodiment of the disclosure; Figure 3 A flowchart of a method for generating target-themed videos provided in this embodiment of the disclosure; Figure 4 A flowchart of a method for generating target-themed videos provided in this embodiment of the disclosure; Figure 5 A block diagram of a target-themed video generation apparatus provided in the embodiments of this disclosure; Figure 6 A block diagram of a target-themed video generation apparatus provided in the embodiments of this disclosure; Figure 7 A flowchart illustrating the target-topic video method provided in this embodiment of the disclosure; Figure 8 Example diagrams illustrating the alignment of materials using an arrangement strategy provided in this embodiment of the disclosure; Figure 9 A block diagram of an intelligent agent provided in an embodiment of this disclosure; Figure 10 This is a block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0014] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0015] The embodiments disclosed herein relate to the technical fields of natural language processing, large models (large language models), deep learning, and computer vision.

[0016] Artificial Intelligence (AI) is a new technological science that studies, develops, and applies theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence.

[0017] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods that enable effective communication between humans and computers using natural language. It is a discipline that uses computer technology to analyze, understand, and process natural language, treating computers as a powerful tool for language research. With computer support, it conducts quantitative research on language information and provides language descriptions that can be used by both humans and computers.

[0018] Large Language Models (LLMs) are deep learning models trained on large amounts of text data that can generate natural language text or understand the meaning of language text. LLMs can handle various natural language tasks, such as text classification, question answering, and dialogue, and are an important pathway to artificial intelligence.

[0019] Computer vision uses various imaging systems to replace visual organs as the means of input sensing, allowing computers to process and interpret information instead of the brain. The ultimate research goal of computer vision is to enable computers to observe and understand the world through vision, just like humans, and to have the ability to autonomously adapt to their environment.

[0020] An intelligent agent is a proxy capable of perceiving its environment and taking actions to achieve specific goals. It can be software, hardware, or a system, possessing autonomy, adaptability, and interactivity. By perceiving changes in the environment (e.g., through sensors or data input), an intelligent agent makes judgments and decisions based on its learned knowledge and algorithms, and then executes actions to influence the environment or achieve predetermined goals.

[0021] It should be noted that the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0022] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0023] It is worth noting that in the embodiments disclosed herein, certain software, components, models, and other existing solutions in the industry may be mentioned. These should be considered as exemplary, and their purpose is only to illustrate the feasibility of implementing the technical solutions disclosed herein. However, this does not mean that the applicant has used or necessarily used such solutions.

[0024] In related technologies, long-form video generation methods are mainly divided into two types: acquisition-editing and generation-based. The acquisition-editing paradigm primarily includes two core capabilities: material acquisition and material arrangement. Material acquisition relies on building a large-scale video material library in the early stages, retrieving scenes corresponding to the video script through semantic-level relevance retrieval, aligning video materials with text materials such as subtitles using arrangement strategies, and then producing the video through a rendering engine. The mainstream generation-based paradigm generates storyboard scenes corresponding to the script through a text-to-image model, then generates video materials through an image-to-video model. The finished video is composed of video material clips spliced ​​together according to an arrangement strategy.

[0025] Unlike typical videos for general entertainment or general knowledge scenarios, targeted promotional videos not only need to present engaging content but also effectively showcase the advertised target audience and conduct appropriate marketing within the video. However, current technological practices for generating such promotional videos, especially in scenarios requiring rapid, mass production, still face numerous technical bottlenecks and challenges.

[0026] Based on this, this disclosure provides a method, apparatus, intelligent agent, and electronic device for generating target-themed videos. It aims to address the problem in related technologies where the editorial AIGC (Artificial Intelligence Generated Content) video production process often lacks specific promotional material (such as products) in specific scenarios like live video streaming, advertising production, and e-commerce sales, leading to poor marketing results. This method ensures that the visual material of the target subject in the generated video is consistent with the text related to the target promotional object, improving the logical coherence and richness of the video content. By improving upon the traditional editorial AIGC video production process, it solves the problems of high computational resource costs (such as memory usage, computational load, and communication overhead during the inference phase of the model) and long generation time in specific scenarios like live video streaming, advertising production, and e-commerce sales.

[0027] The following description, with reference to the accompanying drawings, outlines a method, apparatus, intelligent agent, and electronic device for generating target-themed videos according to embodiments of this disclosure.

[0028] It should be noted that the execution subject of the target theme video generation method in this embodiment of the disclosure can be a target theme video generation device. The device can be implemented by software and / or hardware. The device can be configured in an electronic device, which may include, but is not limited to, a terminal, a server, etc.

[0029] It is worth noting that the target topic video generation method of this disclosure embodiment can be implemented by an autonomous agent based on a large language model, which can generate target topic videos based on a large model.

[0030] Figure 1 A flowchart illustrating a method for generating a target-themed video according to an embodiment of this disclosure. Figure 1 As shown, the method for generating the target theme video may include, but is not limited to, the following steps.

[0031] In step 101, the object information of the target advertising object is obtained.

[0032] In some embodiments, the target of promotion can be understood as the object to be promoted. For example, taking the application of this disclosure to an e-commerce sales scenario, the target of promotion can be a product being promoted (sold), such as a daily necessities product, but not limited to this, such as a cultural and tourism product.

[0033] In some embodiments, the object information may include, but is not limited to, object name and object image. The object name can be understood as the name of the target promotional object. Optionally, the object image may be a cover image (or detail image) of the target promotional object. For example, taking the application of this disclosure to an e-commerce sales scenario as an example, the target promotional object may be a product being promoted (sold), and the object image may be a product cover image.

[0034] In embodiments of this disclosure, the object information of the target promotional object can be obtained through user input. For example, an interactive interface can be provided for the user, through which the user can select or input object information of the target promotional object, thereby enabling the electronic device to obtain object information such as the object name and object image of the target promotional object through user input or selection.

[0035] In step 102, promotional copy for the target promotional object is generated based on the object name using a large model.

[0036] In the embodiments of this disclosure, a large model can be invoked to generate promotional copy for the target promotional object based on its object name. For example, the large model can be used to write promotional copy based on the object name and a prompt, resulting in the promotional copy output by the large model. The prompt can instruct the large model to understand the object based on its name and to write corresponding promotional copy based on that understanding.

[0037] In some embodiments, the aforementioned prompts may be pre-set or provided by the user. In some embodiments, the aforementioned prompts may include, but are not limited to, copywriting requirements, which may include, but are not limited to, the proportion of marketing copy in promotional copy (e.g., the proportion of marketing copy in promotional copy is 0.3, and the copywriting part is required to follow a strong narrative and weak marketing writing style), selecting plots based on the theme of the target promotional object, etc. For example, in terms of plot selection, the copywriting may be based on the theme (IP) of the target promotional object. For example, for daily necessities, "Theme A XX Mug", the IP / theme is "Theme A", and the promotional copy will be written based on the story of "Theme A"; or for books, "XXXX" books, the IP / theme is "financial education", and the promotional copy will be written based on financial knowledge.

[0038] In step 103, the insertion positions of promotional materials in the promotional copy are marked to obtain promotional position marking information.

[0039] In some embodiments, a large model can be invoked to mark the insertion positions of promotional materials in the promotional copy, thereby obtaining promotional position marking information, which can be used to find the insertion positions of target marketing materials in the subsequent video arrangement strategy.

[0040] In step 104, target video materials unrelated to the target audience are selected from the video material library based on the promotional copy and promotional location labeling information.

[0041] In some embodiments, promotional copy unrelated to the target audience can be identified from the promotional copy based on the promotional location annotation information. For example, the marketing copy in the promotional copy is determined based on the promotional location annotation information; that is, the copy corresponding to the insertion position of the promotional material in the promotional copy is the marketing copy. All other copy in the promotional copy is considered unrelated to the target audience. Video materials can be retrieved from a preset video material library based on this promotional copy to obtain the corresponding target video materials. Since the promotional copy used for video retrieval is unrelated to the target audience, the video content in the retrieved target video materials is also unrelated to the target audience. Instead, it consists of video content with a storyline based on the theme of the target audience, following a strong narrative and light marketing approach. This can attract users to consume the video content and promote conversion. The purpose of following a strong narrative and light marketing writing style in the copy is to keep the proportion of marketing copy in the entire promotional copy very low, thus keeping the cost of generating target materials in the subsequent target video material generation stage controllable.

[0042] In step 105, based on the promotional copy, promotional location marking information, and object image, obtain object video materials related to the target promotional object.

[0043] It is understandable that, for object-specific materials, video material libraries like the one mentioned above typically lack video materials with granularity down to the object type (such as product model). In such cases, this disclosure can process promotional text, promotional location annotation information, and object images to generate object-specific video materials related to the target promotional object.

[0044] In some embodiments, promotional copy related to the target target can be determined from the promotional copy based on the promotional location annotation information. Marketing copy within the promotional copy is then determined based on the promotional location annotation information; that is, the copy corresponding to the insertion position of the promotional material in the promotional copy is the marketing copy, which is the promotional copy related to the target target. Image-to-video processing is then performed on the promotional copy and the object image of the target target to obtain a generated video. This generated video serves as object video material (also called object marketing material) related to the target target.

[0045] In step 106, the promotional copy, promotional location marking information, target video materials, and object video materials are arranged and synthesized to obtain the target theme video of the target promotional object.

[0046] In some embodiments, an arrangement strategy can be used to align promotional text, target video footage, and object video footage based on promotional location annotation information; the aligned promotional text, target video footage, and object video footage can then be rendered and composited to generate a target-themed video of the target promotional object.

[0047] For example, the promotional copy can be split into time slots at the sentence level; for each time slot, the duration of the time slot is determined based on the TTS (Text To Speech) duration of the sentence in that time slot, such as using the TTS duration of the sentence in that time slot as the duration of that time slot; based on the promotional position labeling information, the video material to be filled in the time slot is determined from the target video material and the object video material, and the video material is filled into the time slot, until the material filling of all time slots is completed.

[0048] For example, based on the information marked on the advertising location, it can be determined whether the time slot should be filled with ordinary materials (such as target video materials) or object materials (such as object video materials). If the time slot should be filled with ordinary materials, high-resolution and relevant resources can be selected from the retrieved target video materials to fill the time slot. If the duration of the material is longer than the duration of the time slot, the material is truncated. Otherwise, more materials are added to fill the time slot. After filling each time slot in this way, the rendering engine automatically produces the complete video content, which is the target theme video of the target advertising object.

[0049] By implementing the embodiments of this disclosure, the problem of poor marketing results caused by the lack of specific promotional materials (such as products) in the traditional AIGC video production process of acquisition and editing in specific scenarios such as live video streaming, advertising production, and e-commerce sales can be solved. This ensures that the object visual materials in the generated target-themed video are consistent with the main text related to the target promotional object, thereby improving the logicality and richness of the video content. Improvements to the traditional acquisition and editing AIGC video production process can also solve the problems of high computational resource costs (such as memory usage, computational load, and communication overhead during the inference stage of the model) and long generation time in the generative AIGC video production process in specific scenarios such as live video streaming, advertising production, and e-commerce sales.

[0050] Figure 2 A flowchart illustrating a method for generating a target-themed video according to an embodiment of this disclosure. Figure 2 As shown, the method for generating the target theme video may include, but is not limited to, the following steps.

[0051] In step 201, the object information of the target advertising object is obtained.

[0052] In some embodiments, the object information described above may include, but is not limited to, object name and object image.

[0053] Optionally, step 201 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0054] In step 202, object understanding information is obtained by performing object understanding based on object name through the large model.

[0055] In some embodiments, the object understanding information may include, but is not limited to, the subject (IP) information of the target advertising object, and may also include the object type.

[0056] In the embodiments of this disclosure, a large model can be invoked to perform object understanding based on the object name of the aforementioned target promotional object. For example, the large model can perform object understanding based on the object name and prompt of the aforementioned target promotional object. For instance, object type and object theme information can be parsed from the object name. Taking a product as an example, the product name can be parsed to extract understanding information such as product type, product theme (IP), the pain points the product addresses, and advantages compared to competitors.

[0057] In step 203, promotional copy for the target audience is generated using a large model based on object understanding information.

[0058] For example, promotional copy can be written based on object-oriented understanding information and prompts using a large model to generate promotional copy for the target audience. These prompts can be pre-set or user-provided. In some embodiments, the prompts may include copywriting requirements, such as the proportion of marketing copy in the promotional copy (e.g., marketing copy accounts for 0.3% of the promotional copy, and the copywriting should follow a strong narrative with less marketing emphasis), and selecting a plot based on the theme of the target audience. For instance, the plot selection can be based on the theme (IP) of the target audience. For example, for everyday items like "Theme A XX Mug," the IP / theme is "Theme A," and the promotional copy would be written based on a story related to "Theme A"; similarly, for books like "XXXX," the IP / theme is "financial education," and the promotional copy would be written based on financial knowledge.

[0059] In step 204, the insertion positions of promotional materials in the promotional copy are marked to obtain promotional position marking information.

[0060] Optionally, step 204 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0061] In step 205, semantic extraction processing is performed on the promotional copy of the target promotional object, and the semantic information of the promotional copy is supplemented according to the context to obtain the target semantic information of the target promotional object.

[0062] It should be noted that the generated promotional copy usually cannot be directly used to retrieve materials because each sentence in the copy is written within a context. This can lead to a situation where a sentence lacks a direct subject because it has been mentioned in the preceding text; or, if the context describes an ancient story, but one sentence contains an ambiguous description of a scene, similar queries will result in the inability to retrieve the most relevant scenes. Therefore, in the embodiments of this disclosure, based on the generated promotional copy, corresponding retrieval copy can be generated according to the semantic information of the promotional copy. For example, semantic extraction processing can be performed on the promotional copy of the target promotional object. For instance, a large model can be used to perform semantic extraction processing on the promotional copy of the target promotional object, and the large model can determine whether the promotional copy contains ambiguous semantic expressions (e.g., road vs. ancient road vs. modern highway). The semantic information of the promotional copy can be supplemented according to the context, thereby obtaining the target semantic information of the target promotional object, ensuring that the scenes in the subsequently generated target-themed video remain consistent within the storyline.

[0063] It is worth noting that the large model used for semantic extraction of promotional copy here may be the same model as the large model used to generate promotional copy, or it may be a different model.

[0064] It should be noted that in some embodiments, steps 204 and 205 above can be executed simultaneously or in reverse order.

[0065] In step 206, a corresponding search copy is generated based on the target semantic information of the target advertising object.

[0066] In the embodiments of this disclosure, the promotional copy can be supplemented or rewritten according to the target semantic information of the target promotional object to obtain a corresponding search copy that is more suitable for retrieval.

[0067] In one possible implementation, the target semantic information of the target advertising object can be analyzed, the complex target semantic information can be decomposed into simple query intents, and corresponding search terms can be generated based on the query intents. These search terms can then constitute the search copy.

[0068] In one possible implementation, the target semantic information of the target audience can be transformed into one or more questions, and search terms can be extracted from the questions. These search terms can then constitute the search copy.

[0069] It's worth noting that directly using the original promotional text to search for visual (or video) footage may result in inconsistencies between frames. For example, if the promotional text tells an ancient story and mentions concepts like "roads," a direct search might return modern highways. However, considering the context, the text actually refers to ancient roads; the original text simply omitted certain semantic elements. Using the search term avoids the jarring effect of switching between ancient and modern scenes. Similarly, if the target audience is a product featuring a specific IP, and the promotional text is "Foxes are so cute," a direct search might return various images of wild foxes. However, considering the context, the fox refers to the specific IP's fox image, not wild foxes or foxes from other IPs. Searching the term helps the video footage more focused on images related to the semantics of the original text.

[0070] In step 207, target video materials unrelated to the target advertising object are selected from the video material library based on the search copy and advertising location label information.

[0071] In some embodiments, a first search copy unrelated to the target advertising object can be determined from the search copy based on the advertising position annotation information. For example, marketing copy in the search copy can be determined based on the advertising position annotation information; that is, the copy corresponding to the insertion position of the advertising material in the search copy is the marketing copy. All other copy in the search copy can be considered the first search copy unrelated to the target advertising object. Video materials can then be retrieved from the video material library based on the first search copy to obtain the corresponding target video materials.

[0072] In step 208, based on the search copy, advertising location label information, and object image, object video materials related to the target advertising object are obtained.

[0073] In some embodiments, a first promotional copy related to the target promotional object can be determined from the search text based on the promotional location annotation information. That is, the copy corresponding to the insertion position of the promotional material in the search text is the marketing copy, and this marketing copy is the first promotional copy related to the target promotional object. Material supplementation can be performed based on the first promotional copy and the object image of the target promotional object to obtain object video material related to the target promotional object.

[0074] In some embodiments, the optional implementation of supplementing materials based on the first promotional copy and the object image of the target promotional object to obtain object video materials related to the target promotional object may include: cleaning the object image to obtain a static image of the target promotional object, wherein the static image of the target promotional object does not contain text descriptions; and using a graph-generated video model to supplement the static image of the target promotional object with life scenes based on the first promotional copy to obtain object video materials related to the target promotional object.

[0075] For example, the image of the target promotional object can be used as the original material. A multimodal image generation model can be used to clean the image, retaining the target promotional object itself while removing textual descriptions. This ensures that garbled text will not appear in the video footage when the object is subsequently generated. The cleaned image (i.e., the static image of the target promotional object) can then be supplemented with real-life scenarios based on the first promotional text using an image-generated video model, thereby obtaining relevant video footage of the target promotional object.

[0076] It's worth noting that during the process of supplementing the static image of the target promotional object with a life-scenario-based element using the image-generated video model based on the initial promotional copy, the generation effort is kept as conservative as possible to minimize the illusion of authenticity. For example, the white-background cover image of a health-preserving kettle from brand XX is cleaned to obtain a static image of the kettle itself. The image-generated video model then supplements this static image with a life-scenario-based element, resulting in product video material for the kettle. For instance, this product video material could depict "a health-preserving kettle from brand XX filled with hot water and steaming, and then a user pouring a cup of hot water from the kettle." This method not only supplements the missing object video material but also ensures that the subject of the material is consistent with the target promotional object. Furthermore, because the proportion of the marketing copy in the entire promotional text is kept very low, the cost of generating the required object material is also controllable.

[0077] In step 209, the promotional copy, promotional location marking information, target video materials, and object video materials are arranged and synthesized to obtain the target theme video of the target promotional object.

[0078] Optionally, step 209 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0079] By implementing the embodiments of this disclosure, the problem of ambiguous text descriptions in traditional AIGC video production processes leading to the inability to retrieve the most relevant images and causing image skipping between previous and subsequent recalls can be solved, thereby improving the relevance of recalled images. It also addresses the issue of poor marketing effectiveness in specific scenarios such as live video streaming, advertising production, and e-commerce sales due to a lack of specific promotional material (e.g., products). This ensures that the object image material in the generated target-themed video is consistent with the text related to the target promotional object, improving the logical coherence and richness of the video content. Furthermore, improvements to the traditional AIGC video production process address the problems of high computational resource costs (such as memory usage, computational load, and communication overhead during the inference phase) and long generation times in generative AIGC video production processes in specific scenarios such as live video streaming, advertising production, and e-commerce sales.

[0080] Figure 3 A flowchart illustrating a method for generating a target-themed video according to an embodiment of this disclosure. Figure 3 As shown, the method for generating the target theme video may include, but is not limited to, the following steps.

[0081] In step 301, the object information of the target advertising object is obtained.

[0082] In some embodiments, the object information described above may include, but is not limited to, object name and object image.

[0083] Optionally, step 301 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0084] In step 302, promotional copy for the target promotional object is generated based on the object name using a large model.

[0085] Optionally, step 302 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0086] In step 303, the insertion positions of promotional materials in the promotional copy are marked to obtain promotional position marking information.

[0087] Optionally, step 303 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0088] In step 304, the promotional text is segmented into sentences, and the sentences obtained from the segmentation of the promotional text are used as corresponding subtitles.

[0089] In the embodiments of this disclosure, the promotional text can be segmented into sentences, and the sentences obtained from the segmentation of the promotional text can be used as corresponding subtitles. This makes it easier to output the subtitles and video footage together, improves the reading experience, and also enhances the aesthetic effect.

[0090] It should be noted that in some embodiments, steps 303 and 304 above can be executed simultaneously or in reverse order.

[0091] In step 305, the promotional text is converted into speech using TTS technology.

[0092] Optionally, the promotional copy can be analyzed to generate a TTS style, and then the promotional copy for the target audience can be converted into corresponding speech based on this TTS style using TTS technology.

[0093] It should be noted that in some embodiments, steps 304 and 305 above can be executed simultaneously or in reverse order.

[0094] In step 306, background music is selected through the promotional copy.

[0095] Optionally, the promotional copy can be analyzed to generate a background music style, and a background music search can be performed based on this style to obtain the corresponding background music.

[0096] It should be noted that in some embodiments, steps 305 and 306 above can be executed simultaneously or in reverse order.

[0097] In the embodiments of this disclosure, the corresponding subtitles, audio, and background music are obtained through promotional documents, which facilitates the subsequent arrangement and synthesis of subtitles, audio, background music, target video materials, object video materials, and text when generating target-themed videos, thereby further enhancing the richness of video content.

[0098] In step 307, target video materials unrelated to the target audience are selected from the video material library based on the promotional copy and promotional location labeling information.

[0099] Optionally, step 307 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0100] In step 308, based on the promotional copy, promotional location marking information, and object image, video materials related to the target promotional object are obtained.

[0101] Optionally, step 308 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0102] In step 309, the advertising location information, advertising copy, target video material, object video material, subtitles, audio and background music are arranged and synthesized to obtain the target theme video of the target advertising object.

[0103] In some embodiments, the aforementioned promotional text, target video material, object video material, subtitles, audio, and background music are aligned based on the promotional location annotation information using an arrangement strategy; the aligned promotional text, target video material, object video material, subtitles, audio, and background music are then rendered and synthesized to generate a target-themed video of the target promotional object.

[0104] For example, the promotional copy can be split into time slots at the sentence level. For each time slot, the duration of the time slot is determined based on the TTS (Text To Speech) duration of the sentence in that time slot, such as using the TTS duration of the sentence in that time slot as the duration of that time slot. Based on the promotional position labeling information, the material to be filled in the time slot is determined from ordinary materials (including target video materials, subtitles unrelated to the target promotional object, promotional copy, and audio) and object materials (including object video materials, subtitles related to the target promotional object, promotional copy, and audio), and the material is filled into the time slot until the material filling of all time slots is completed.

[0105] For example, based on the information marked on the advertising location, it can be determined whether the time slot should be filled with ordinary materials (such as target video materials) or object materials (such as object video materials). If the time slot should be filled with ordinary materials, high-resolution and relevant resources can be selected from the retrieved target video materials to fill the time slot. If the duration of the material is longer than the duration of the time slot, the material is truncated. Otherwise, more materials are added to fill the time slot. After filling each time slot in this way, the rendering engine automatically produces the complete video content, which is the target theme video of the target advertising object.

[0106] In the above embodiments, corresponding subtitles, audio, and background music can also be obtained from promotional documents. Based on the promotional location marking information, promotional copy, target video material, object video material, subtitles, audio, and background music, they can be arranged and synthesized to further enhance the content richness of the target theme video.

[0107] Figure 4 A flowchart illustrating a method for generating a target-themed video according to an embodiment of this disclosure. Figure 4 As shown, the method for generating the target theme video may include, but is not limited to, the following steps.

[0108] In step 401, the object information of the target advertising object is obtained.

[0109] In some embodiments, the object information described above may include, but is not limited to, object name and object image.

[0110] Optionally, step 401 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0111] In step 402, object understanding information is obtained by performing object understanding based on object name through the large model.

[0112] Optionally, step 402 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0113] In 403, promotional copy for the target audience is generated through a large model based on object understanding information.

[0114] Optionally, step 403 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0115] In step 404, the insertion positions of promotional materials in the promotional copy are marked to obtain promotional position marking information.

[0116] Optionally, step 404 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0117] In step 405, the promotional text is segmented into sentences, and the sentences obtained from the segmentation of the promotional text are used as corresponding subtitles.

[0118] Optionally, step 405 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0119] It should be noted that in some embodiments, steps 404 and 405 can be executed simultaneously or in reverse order.

[0120] In step 406, the promotional text is converted into speech using TTS technology.

[0121] Optionally, step 406 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0122] It should be noted that in some embodiments, steps 405 and 406 can be executed simultaneously or in reverse order.

[0123] In step 407, background music is selected through the promotional copy.

[0124] Optionally, step 407 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0125] It should be noted that in some embodiments, steps 406 and 407 above can be executed simultaneously or in reverse order.

[0126] In step 408, semantic extraction processing is performed on the promotional copy of the target promotional object, and the semantic information of the promotional copy is supplemented according to the context to obtain the target semantic information of the target promotional object.

[0127] Optionally, step 408 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0128] It should be noted that in some embodiments, steps 407 and 408 can be executed simultaneously or in reverse order.

[0129] In step 409, corresponding search copy is generated based on the target semantic information of the target advertising object.

[0130] Optionally, step 409 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0131] In step 410, target video materials unrelated to the target advertising object are selected from the video material library based on the search copy and advertising location label information.

[0132] Optionally, step 410 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0133] In step 411, based on the search copy, advertising location label information, and object image, object video materials related to the target advertising object are obtained.

[0134] Optionally, step 411 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0135] In step 412, the advertising location information, advertising copy, target video material, object video material, subtitles, audio and background music are arranged and synthesized to obtain the target theme video of the target advertising object.

[0136] Optionally, step 412 can be implemented using any of the implementation methods in the various embodiments of this disclosure. This disclosure does not limit this implementation and will not elaborate further.

[0137] In the above embodiments, corresponding subtitles, audio, and background music can also be obtained from promotional documents. Based on the promotional location marking information, promotional copy, target video material, object video material, subtitles, audio, and background music, they can be arranged and synthesized to further enhance the content richness of the target theme video.

[0138] Figure 5 A block diagram of a target-themed video generation apparatus provided in an embodiment of this disclosure. Figure 5 As shown, the target-themed video generation device may include: an acquisition module 501, a text management module 502, a retrieval module 503, an object material enhancement module 504, and an arrangement and compositing module 505.

[0139] The acquisition module 501 is used to acquire object information of the target advertising object, including object name and object image.

[0140] The copywriting management module 502 is used to generate promotional copy for the target promotional object based on the object name through a large model.

[0141] The copywriting management module 502 is also used to mark the insertion positions of promotional materials in promotional copywriting and obtain promotional position marking information.

[0142] The retrieval module 503 is used to select target video materials that are unrelated to the target audience from the video material library based on the promotional copy and promotional location labeling information.

[0143] The object material enhancement module 504 is used to obtain object video materials related to the target promotional object based on the promotional copy, promotional location annotation information, and object image.

[0144] The arrangement and compositing module 505 is used to arrange and composite the promotional copy, promotional location marking information, target video material, and object video material to obtain the target theme video of the target promotional object.

[0145] In some embodiments, the copywriting management module 502 is used to: perform object understanding based on object name through a large model to obtain object understanding information; and generate promotional copy for the target promotional object through the large model based on the object understanding information.

[0146] In some embodiments, the object understanding information includes at least the theme information of the target advertising object. The copywriting management module 502 is used to: write promotional copy based on the object understanding information and prompts using a large model, and generate promotional copy for the target advertising object, wherein the prompts include copywriting requirements, which include the proportion of marketing copy in the promotional copy and the selection of plots based on the theme of the target advertising object.

[0147] In some embodiments, such as Figure 6 As shown, the target-themed video generation device may further include a screen consistency management module 606. The screen consistency management module 606 is used to: perform semantic extraction processing on the promotional text of the target promotional object, and supplement the semantic information of the promotional text according to the context to obtain the target semantic information of the target promotional object; and generate corresponding search text based on the target semantic information of the target promotional object. Wherein, Figure 6 601-605 and Figure 5 The 501-505 series have the same function and structure.

[0148] In some embodiments, the retrieval module is used to: determine a first retrieval text that is unrelated to the target advertising object from the retrieval text based on the advertising location annotation information; and perform a video material retrieval in the video material library according to the first retrieval text to obtain the corresponding target video material.

[0149] In some embodiments, the object material enhancement module is used to: determine a first promotional text related to the target promotional object from the search text based on the promotional location annotation information; and supplement the material according to the first promotional text and the object image to obtain object video material related to the target promotional object.

[0150] In one possible implementation, the object material enhancement module is used to: clean the object image to obtain a static image of the target promotional object, which does not contain any text descriptions; and use a graph-generated video model to supplement the static image of the target promotional object with life scenes based on the first promotional copy to obtain object video materials related to the target promotional object.

[0151] In some embodiments, the orchestration and compositing module is used to: align promotional text, target video materials, and object video materials based on promotional location annotation information using an orchestration strategy; and render and composite the aligned promotional text, target video materials, and object video materials to generate a target-themed video of the target promotional object. For example, an optional implementation of aligning promotional text, target video materials, and object video materials based on promotional location annotation information using an orchestration strategy includes: splitting time slots at the sentence level of the promotional text; determining the duration of each time slot based on the text-to-speech (TTS) duration of the sentence in the time slot; determining the video materials to be filled in the time slots from the target video materials and object video materials based on the promotional location annotation information, and filling the time slots with the video materials, until all time slots have been filled.

[0152] In some embodiments, the copywriting management module is further configured to: segment the promotional copy into sentences, using the segments as corresponding subtitles; convert the promotional copy into speech using text-to-speech (TTS) technology; and select background music based on the promotional copy. In embodiments of this disclosure, the arrangement and synthesis module is configured to: align the promotional copy, target video material, object video material, subtitles, speech, and background music based on the promotional location annotation information using an arrangement strategy; and render and synthesize the aligned promotional copy, target video material, object video material, subtitles, speech, and background music to generate a target-themed video of the target promotional object.

[0153] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0154] The following will combine Figure 7 and Figure 8 Taking the application of this disclosure to e-commerce sales (or e-commerce live-streaming sales) scenarios as an example, the technical solution of this disclosure is described to facilitate understanding by those skilled in the art. The target promotional object can be the product being promoted in the e-commerce sales (or e-commerce live-streaming sales) scenario, and the target theme video can be a live-streaming video related to the product being promoted.

[0155] For example, such as Figure 7As shown, the system can obtain product information (such as product name, product cover image, etc.) selected by the user (or author) for sale (or product promotion). The copywriting management module first calls the large model to understand the product, parsing the product type, product theme (IP), the pain point the product solves, and its advantages compared to competitors through the product name. Then, based on the product understanding, the large model generates promotional copy. In order to attract users to consume video content rather than simply advertising, the copywriting is required to follow a strong narrative and low-marketing writing style. In terms of plot selection, it will try to be based on the product's IP. For example, for daily necessities, "Theme A XX Mug", the IP / theme is "Theme A", and the copywriting will be based on the story of "Theme A"; for books, "XXXX" books, the IP / theme is "financial education", and the copywriting will be based on financial knowledge. At the same time, a sentence of copywriting is usually around 30 words. Outputting it directly with the video would cause reading difficulties and be unsightly. Therefore, the copywriting needs to be divided into sentences and corresponding subtitles are output. Finally, it's worth mentioning that, to facilitate finding the insertion points for product marketing materials in subsequent video editing strategies, you can use a large model to mark the marketing positions in the copy.

[0156] The generated text often cannot be directly used to retrieve materials because each sentence in the text is written within context. This can lead to a situation where a sentence lacks a direct subject because it has been mentioned in the preceding text; or, if the context describes an ancient story, but one sentence contains an ambiguous description of the scene, similar queries will result in the inability to retrieve the most relevant scene. Therefore, based on the generated text, a scene consistency management module is needed to generate corresponding search text based on the semantic information of the text. The search text must satisfy the semantics of the original text, and a large model should be used to determine if the text contains ambiguous semantic expressions (e.g., road vs. ancient road vs. modern highway). The semantic information of the original text should be supplemented based on the context to ensure that the scenes in the video remain as consistent as possible within the storyline. Furthermore, due to the special nature of product-selling scenarios, buyers often choose to purchase based on the IP attribute of a product. The generation of product-selling videos will emphasize that the IP appearing in the scenes is consistent. For example, in a product-selling video for "XXX game disc," the video footage will consist of footage from that game, rather than interspersed with other similar game footage for marketing purposes.

[0157] For product materials, as mentioned earlier, material libraries typically lack video materials specific to product models, and finding similar competitor products could harm the client's interests. In such cases, material enhancement modules are used to supplement the materials. This disclosure uses the product cover image as the original material. First, a multimodal image generation model is used to clean the image, preserving the product itself while removing textual descriptions from the cover image. This step ensures that garbled text will not appear in the video footage when generating subsequent product video materials. The cleaned image is then used to supplement the static product image with a life-like scene using an image-to-video model, but the generation is kept as conservative as possible to reduce any illusions. For example, a white-background cover image of a health-preserving kettle from brand XX will be transformed through the material enhancement module into a kettle filled with hot water and steaming, with a user holding the kettle and pouring a cup of hot water. This method not only supplements the missing product video materials but also ensures that the subject matter and the product are consistent. Furthermore, because the proportion of marketing copy in the entire text is kept very low, the cost of generating product materials is also controllable.

[0158] After obtaining the retrieved video footage, product video footage generated through the image-based video model, subtitles, sales copy, audio, and background music, the video footage (including the retrieved video footage and product video footage), audio, background music (BGM), copy, and subtitles can be aligned using an arrangement strategy. The arrangement strategy includes two core functions: aligning the timelines of various materials (subtitles, visuals, audio, BGM) and sorting all retrieved video footage for each time slot. For example... Figure 8 As shown, a long video can be viewed as consisting of n video segments. To ensure that the video doesn't switch too frequently or loop due to insufficient footage, segments need to be broken down at the sentence level of the script. For each segment, the editing strategy first determines the segment length based on its time-to-segment (TTS) duration, and then determines whether to fill the segment with product footage or regular footage based on the product location information. If the segment is filled with regular footage, high-resolution and relevant resources are prioritized from the recalled footage to fill the segment. If the footage length exceeds the segment length, it is truncated; otherwise, more footage is added. This process is repeated until each segment is filled, after which the rendering engine automatically produces the complete video content.

[0159] This disclosure also provides an intelligent agent. Figure 9 A block diagram of an intelligent agent provided in an embodiment of this disclosure. (See diagram below.) Figure 9As shown, the intelligent agent may include an input module 901, a processing module 902, and an output module 903. The input module 901 receives input information; the processing module 902 determines a target task based on the input information received by the input module 901, determines a large model based on the target task, and executes the target-themed video generation method described in the above method embodiment by calling the large model to obtain output information; the output module 903 outputs the output information obtained by the processing module 902. For example, the target task may be a video generation task. The input information may include object information of the target promotional object. The output information may include a target-themed video of the target promotional object.

[0160] According to embodiments of this disclosure, this disclosure also provides an electronic device and a readable storage medium.

[0161] like Figure 10 The diagram shown is a block diagram of an electronic device according to an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0162] like Figure 10 As shown, the electronic device includes one or more processors 1001, a memory 1002, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise as required. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 10 Take processor 1001 as an example.

[0163] The memory 1002 is the non-transitory computer-readable storage medium provided in this disclosure. The memory stores instructions executable by at least one processor to cause the at least one processor to perform the target-themed video generation method provided in this disclosure. The non-transitory computer-readable storage medium of this disclosure stores computer instructions for causing a computer to perform the target-themed video generation method provided in this disclosure.

[0164] Memory 1002, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the target theme video generation method in this disclosure embodiment (e.g., appendix). Figure 5 The module shown includes the acquisition module 501, the document management module 502, the retrieval module 503, the object material enhancement module 504, and the arrangement and compositing module 505. Figure 6 The module shown includes the acquisition module 601, the text management module 602, the retrieval module 603, the object material enhancement module 604, the arrangement and compositing module 605, and the image consistency management module 606. The processor 1001 executes various functional applications and data processing of the server by running non-transient software programs, instructions, and modules stored in the memory 1002, thereby realizing the target theme video generation method in the above method embodiment.

[0165] The memory 1002 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 1002 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1002 may optionally include memory remotely located relative to the processor 1001, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0166] The electronic device may also include an input device 1003 and an output device 1004. The processor 1001, memory 1002, input device 1003, and output device 1004 can be connected via a bus or other means. Figure 10 Taking the example of a connection between China and Israel via a bus.

[0167] Input device 1003 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 1004 may include display device, auxiliary lighting device (e.g., LED), and haptic feedback device (e.g., vibration motor). The display device may include, but is not limited to, liquid crystal display (LCD), light-emitting diode (LED) display, and plasma display. In some embodiments, the display device may be a touch screen.

[0168] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, application-specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0169] These computational programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0170] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0171] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0172] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0173] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0174] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for generating a target-themed video, comprising: Obtain object information of the target advertising object, the object information including object name and object image; The large model generates promotional copy for the target promotional object based on the object name; The insertion positions of promotional materials in the aforementioned promotional text are marked to obtain promotional position marking information; Based on the promotional copy and the promotional location marking information, select target video materials from the video material library that are unrelated to the target promotional object; Based on the promotional copy, the promotional location marking information, and the object image, obtain object video materials related to the target promotional object; The target theme video of the target promotional object is obtained by arranging and synthesizing the promotional copy, the promotional location marking information, the target video material, and the object video material.

2. The method according to claim 1, wherein, The process of generating promotional copy for the target promotional object based on the object name using a large model includes: Object understanding information is obtained by performing object understanding based on the object name using a large model; Based on the object understanding information, the promotional copy for the target promotional object is generated through the large model.

3. The method according to claim 2, wherein, The object understanding information includes at least the theme information of the target advertising object; The process of generating promotional copy for the target audience based on the object understanding information through the large model includes: The large model uses the object's understanding information and prompts to write promotional copy, generating promotional copy for the target target. The prompts include copywriting requirements, which include the proportion of marketing copy in the promotional copy and selecting a plot based on the theme of the target target.

4. The method according to claim 1, further comprising: Semantic extraction processing is performed on the promotional copy of the target promotional object, and the semantic information of the promotional copy is supplemented according to the context to obtain the target semantic information of the target promotional object; Generate corresponding search copy based on the target semantic information of the target advertising object.

5. The method according to claim 4, wherein, The step of selecting target video materials unrelated to the target promotional object from the video material library based on the promotional copy and the promotional location marking information includes: Based on the advertising location labeling information, a first search text that is unrelated to the target advertising object is determined from the search text; Based on the first search query, a video material search is performed in the video material library to obtain the corresponding target video material.

6. The method according to claim 4, wherein, The step of obtaining object video materials related to the target promotional object based on the promotional copy, the promotional location marking information, and the object image includes: Based on the advertising location marking information, a first advertising copy related to the target advertising object is determined from the search copy; Based on the first promotional text and the object image, supplementary materials are obtained to obtain object video materials related to the target promotional object.

7. The method according to claim 6, wherein, The step of supplementing materials based on the first promotional copy and the object image to obtain object video materials related to the target promotional object includes: The object image is cleaned to obtain a static image of the target promotional object, wherein the object image contains no text descriptions. The image-generated video model supplements the static image of the target promotional object with life scenes based on the first promotional copy, thereby obtaining object video materials related to the target promotional object.

8. The method according to any one of claims 1-7, wherein, The step of arranging and synthesizing the promotional copy, the promotional location marking information, the target video material, and the object video material to obtain the target theme video of the target promotional object includes: Based on the promotional location annotation information, the arrangement strategy aligns the promotional copy, the target video material, and the object video material, including: splitting the promotional copy into time slots at the sentence level; for each time slot, determining the duration of the time slot based on the text-to-speech (TTS) duration of the sentence in the time slot; and based on the promotional location annotation information, determining the video material to be filled in the time slot from the target video material and the object video material, and filling the time slot with the video material, until all the material filling for the split time slots is completed. The aligned promotional text, the target video material, and the object video material are rendered and combined to generate a target-themed video of the target promotional object.

9. The method according to any one of claims 1-7, further comprising: The promotional text is segmented into sentences, and the sentences obtained from the segmentation of the promotional text are used as corresponding subtitles; The promotional copy was converted into speech using text-to-speech (TTS) technology. Choose background music from the promotional text.

10. The method according to claim 9, wherein, The step of arranging and synthesizing the promotional copy, the promotional location marking information, the target video material, and the object video material to obtain the target theme video of the target promotional object includes: Based on the advertising location labeling information, the editing strategy aligns the advertising copy, the target video material, the object video material, the subtitles, the audio, and the background music. The aligned promotional text, target video material, object video material, subtitles, audio, and background music are rendered and synthesized to generate a target-themed video of the target promotional object.

11. A target-themed video generation apparatus, comprising: The acquisition module is used to acquire object information of the target advertising object, the object information including object name and object image; The copywriting management module is used to generate promotional copy for the target promotional object based on the object name using a large model; The copy management module is also used to mark the insertion positions of promotional materials in the promotional copy to obtain promotional position marking information; The retrieval module is used to select target video materials unrelated to the target promotional object from the video material library based on the promotional copy and the promotional location marking information; The object material enhancement module is used to obtain object video materials related to the target promotional object based on the promotional copy, the promotional location annotation information, and the object image; The editing and compositing module is used to edit and compose the promotional copy, the promotional location marking information, the target video material, and the object video material to obtain the target theme video of the target promotional object.

12. An intelligent agent, comprising: The input module is used to receive input information; The processing module is configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and execute the method of any one of claims 1-10 by calling the large model to obtain output information; The output module is used to output the output information obtained by the processing module.

13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-10.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.

15. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-10.