Seamless brand integration method in text-to-video generation and terminal equipment
By building a brand knowledge base and a seamless brand integration approach through multi-agent collaboration, the problems of high computational resource investment and monetization difficulties in text-to-video generation services are solved, achieving natural integration of brand elements and commercial sustainability, and improving user experience and brand exposure.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE CHINESE UNIV OF HONG KONG (SHENZHEN)
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-12
AI Technical Summary
Existing text-to-video generation services suffer from huge computational resource investment and a lack of effective monetization methods, making it difficult for service providers to recoup their investment and achieve profitability. At the same time, traditional advertising methods damage the user experience and affect the sustainability of the service.
We build a brand knowledge base and generate a seamless brand integration method through multi-agent collaboration and iterative optimization, including brand selection, strategy generation, prompt word rewriting, and experience learning. This enables the automatic embedding and natural integration of brand elements, and utilizes adapters and multi-dimensional evaluation to ensure the generation effect.
It achieves a natural integration of brand elements in video generation, enhances user experience, provides a sustainable monetization path, reduces deployment costs, increases brand exposure, and establishes a win-win ecosystem for all three parties.
Smart Images

Figure CN122023618A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a seamless brand integration method and terminal device for text-to-video generation, belonging to the field of video generation technology. Background Technology
[0002] The rapid development of text-to-video (T2V) technology has opened up unprecedented opportunities for automated content creation, especially in the advertising production field. However, existing T2V services face severe commercialization challenges. On the one hand, model training and inference require huge investments in computing resources, including the procurement and maintenance of large-scale GPU clusters, the collection and processing of massive amounts of training data, and the ongoing costs of model optimization and iteration. On the other hand, the lack of effective monetization methods makes it difficult for service providers to recoup their investment and achieve profitability. Traditional intrusive advertising, such as interstitial ads and pop-up ads, severely damages the user experience, leading to user churn and ultimately undermining the sustainability of the service. Summary of the Invention
[0003] This invention provides a seamless brand integration method and terminal device for text-to-video generation, which can solve the problem that existing T2V services require huge computing resources for model training and inference, as well as lack effective monetization methods, making it difficult for service providers to recover their investment and achieve profitability.
[0004] On one hand, the present invention provides a seamless brand integration method in text-to-video generation, the method comprising:
[0005] S1. Construct a brand knowledge base containing multiple brand profiles, receive text prompts, and query the brand knowledge base for the target brand profile corresponding to the text prompts.
[0006] S2. Based on the text prompts and the target brand profile, generate a brand integration strategy, and optimize the text prompts according to the brand integration strategy to obtain optimized prompts;
[0007] S3. Evaluate the optimized prompt words from multiple dimensions and determine the target prompt words based on the evaluation results;
[0008] S4. Using the target prompts and the target brand profile, generate the target video.
[0009] Optionally, the construction of a brand knowledge base comprising multiple brand profiles in S1 specifically includes:
[0010] Based on the brand's initial profile, generate multiple test prompts, and based on the initial profile and each test prompt, generate a corresponding test video;
[0011] If the percentage of test videos that meet the preset requirements among all test videos of the brand is greater than or equal to the preset percentage, then the initial file of the brand will be stored in the brand knowledge base as the brand file of the brand.
[0012] If the percentage of test videos that meet the preset requirements among all test videos of the brand is less than the preset percentage, then the adapter of the brand is determined according to the initial file, and the initial file and adapter of the brand are stored as the brand file of the brand in the brand knowledge base.
[0013] Optionally, the adapter for the brand is determined based on the initial profile, specifically including:
[0014] Multiple comprehensive prompts are generated based on the initial profile; each comprehensive prompt includes a brand name and a trigger marker.
[0015] Based on the reference image in the initial file and the multiple comprehensive prompt words, multiple training videos are generated accordingly, and the multiple comprehensive prompt words and the corresponding multiple training videos are combined into a training dataset.
[0016] The text-to-video model is trained using the training dataset to obtain the adapter for the brand.
[0017] Optionally, based on the reference image in the initial file and the multiple synthesized cue words, multiple training videos are generated, specifically including:
[0018] The reference image in the initial file and the input text of each comprehensive prompt word are fed into the image model to generate an initial frame corresponding to each comprehensive prompt word;
[0019] The initial frame corresponding to each of the comprehensive prompt words is input into the image-to-video model to generate a training video corresponding to each of the comprehensive prompt words.
[0020] Optionally, S3 specifically includes:
[0021] The optimized prompts were evaluated across four dimensions: semantic fidelity, brand clarity, integration naturalness, and generation effect, yielding the evaluation results.
[0022] If the scores of all four dimensions in the evaluation results are greater than or equal to the preset threshold, then the optimized prompt word will be used as the target prompt word.
[0023] If the evaluation results contain modification suggestions, then the optimized prompt words are modified according to the modification suggestions to obtain the target prompt words.
[0024] Optionally, the optimized prompt words are modified according to the modification suggestions to obtain the target prompt words, specifically including:
[0025] If the modification suggestion is a re-optimization suggestion, then the optimized prompt word is used as the text prompt word, and the text prompt word is optimized according to the brand integration strategy and the re-optimization suggestion to obtain the optimized prompt word. Then S3 is executed again until the target prompt word is obtained.
[0026] If the proposed modification is to regenerate the strategy, then the optimized prompt word is used as the text prompt word, and S2 and S3 are re-executed until the target prompt word is obtained.
[0027] Optionally, S4 specifically includes:
[0028] The target prompt and the target brand profile are input into the video model to generate the target video.
[0029] Optionally, prior to S4, the method further includes:
[0030] If the target brand profile contains an adapter for the brand, then the adapter is loaded into the text-to-video model.
[0031] Optionally, after S4, the method further includes:
[0032] Collect user feedback on the target video;
[0033] Structured experience is extracted from the feedback information and stored in the brand knowledge base.
[0034] On the other hand, the present invention provides a terminal device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the seamless brand integration method in text-to-video generation as described above.
[0035] The beneficial effects that this invention can produce include:
[0036] This invention provides a seamless brand integration method for text-to-video generation, which for the first time systematically solves the triple challenge of automatically embedding brand elements during the generation of user-requested videos while maintaining semantic fidelity, ensuring brand recognizability, and achieving natural integration. This method opens up new avenues for the commercial and sustainable development of text-to-video services.
[0037] This invention provides a seamless brand integration method for text-to-video generation, employing a two-stage collaborative multi-agent framework. In the offline stage, a brand knowledge base is constructed through prior knowledge exploration and selective model adaptation, providing a foundation for subsequent integration. In the online stage, five specialized agents (brand selection, strategy generation, prompt word rewriting, evaluation, and experience learning) work collaboratively, achieving high-quality brand integration through iterative optimization. This invention is plug-and-play, integrating with various mainstream text-to-video generation models (such as Veo, Sora, and Kling) without requiring large-scale modifications to the underlying model, offering strong applicability and low deployment costs.
[0038] This invention provides a seamless brand integration method for text-to-video generation, aiming to establish a win-win ecosystem for all three parties. For brands, they gain organic exposure, with brand elements naturally appearing in user-generated content, resulting in higher acceptance and memorability compared to traditional advertising. For service providers, this invention enables the establishment of a sustainable revenue stream, charging brands for brand integration services to cover high computational costs and achieve profitability. For end users, they can continue to receive high-quality video generation services without explicit advertising interruptions, ensuring a fully guaranteed user experience. This model, while protecting the interests of all parties, lays the foundation for the long-term development of T2V technology. Attached Figure Description
[0039] Figure 1 A flowchart illustrating a seamless brand integration method for text-to-video generation provided in an embodiment of the present invention. Detailed Implementation
[0040] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0041] The core principle of this invention is to achieve semantic-aware integration of brands through multi-agent collaboration and knowledge base sharing. This principle is based on the organic combination of three key design ideas. First, addressing the differences in knowledge states of different brands in the pre-trained model, this invention automatically identifies brand types through a prior knowledge detection mechanism. For known brands, it directly utilizes the model's prior knowledge to achieve integration through prompt word optimization. For new brands, it creates brand-specific adapters through synthetic data generation and LoRA fine-tuning, injecting necessary brand knowledge. Second, brand integration requires a trade-off between multiple dimensions such as semantic fidelity, brand visibility, and integration naturalness. A single model cannot handle all constraints simultaneously. Therefore, this invention designs five specialized agents (brand selection, strategy generation, prompt word rewriting, evaluation, and experience learning) to collaborate, achieving efficient collaboration among agents through a dual memory mechanism (long-term brand knowledge base and short-term working context). Furthermore, this invention employs an iterative optimization strategy, whereby the evaluation agent performs multi-dimensional evaluations of each generated prompt word and decides whether to accept, revise, or replan based on the results. Through a hierarchical iterative mechanism, it ensures that the final output simultaneously meets the triple constraints, avoiding the problem of neglecting one aspect while focusing on another in simple methods. At the same time, the experience-learning agent transforms each integration result into reusable knowledge, enabling continuous improvement of the system.
[0042] Specifically, embodiments of the present invention provide a seamless brand integration method in text-to-video generation, such as... Figure 1 As shown, the method includes:
[0043] S1. Construct a brand knowledge base containing multiple brand profiles, receive text prompts, and query the target brand profile corresponding to the text prompts from the brand knowledge base.
[0044] The aforementioned construction of a brand knowledge base containing multiple brand profiles specifically includes:
[0045] Based on the brand's initial profile, generate multiple test prompts, and based on the initial profile and each test prompt, generate a corresponding test video;
[0046] If the percentage of test videos that meet the preset requirements is greater than or equal to the preset percentage among all the test videos of a brand, then the brand's initial profile will be stored in the brand knowledge base as the brand profile.
[0047] If the percentage of test videos that meet the preset requirements is less than the preset percentage among all the brand's test videos, then the brand's adapter is determined based on the initial profile, and the brand's initial profile and adapter are stored as the brand's profile in the brand knowledge base.
[0048] The input at this stage is the brand's initial profile, in which... This includes the brand name, brand category, set of reference images, and brand description text. The system first performs prior knowledge exploration of the brand, and then determines whether model-level adaptation is needed based on the exploration results.
[0049] Specifically, the system uses a prompt generator to create diverse test prompts that explicitly mention the brand name in different contextual scenarios. The T2V model generates corresponding videos based on these prompts, and a brand quality evaluator checks whether brand elements are accurately presented and whether visual features are correct. If more than 70% of the test videos successfully generate identifiable brand features, the system marks the brand as having sufficient prior knowledge and directly registers it in the brand knowledge base. If the success rate is below 70%, the brand is deemed to lack prior knowledge and needs to proceed to the next step of model adaptation.
[0050] The adapters whose brand is determined based on the initial files specifically include:
[0051] (1) Generate multiple comprehensive prompts based on the initial file; each comprehensive prompt contains a brand name and a trigger tag.
[0052] (2) Based on the reference image and multiple comprehensive prompt words in the initial file, generate multiple training videos and combine the multiple comprehensive prompt words and the corresponding multiple training videos into a training dataset.
[0053] The process involves generating multiple training videos based on reference images and multiple comprehensive prompts in the initial file. Specifically, this includes: first, inputting the reference images and each comprehensive prompt from the initial file into the image model to generate an initial frame corresponding to each comprehensive prompt; then, inputting the initial frame corresponding to each comprehensive prompt into the image-to-video model to generate a training video corresponding to each comprehensive prompt.
[0054] For brands lacking prior knowledge, the synthetic data generator first constructs a training dataset. The generator creates M comprehensive cue words containing the brand name and specific trigger markers. Then, using a reference image R provided by the brand and a text-to-image model, it generates corresponding initial frames. Next, an image-to-video model expands these initial frames into complete video sequences, i.e., training videos. This process generates a training dataset containing M training samples.
[0055] (3) Use the training dataset to train the text-to-video model to obtain the brand's adapter.
[0056] The system uses LoRA technology to fine-tune the T2V model and generate a brand-specific adapter AB. This adapter AB enables the model to accurately generate the brand's visual features when the cue words contain trigger markers.
[0057] The system stores all relevant brand information in the brand knowledge base, including the brand knowledge type (known brand or new brand), adapter weights (if applicable), reference visual patterns, and an initial experience pool.
[0058] After building the brand knowledge base, the brand selection agent queries the brand knowledge base based on the text prompts entered by the user, retrieves all available brand information, and analyzes the text prompts. The system identifies scene characteristics (including scene type, main objects, actions, and activities), assesses the semantic compatibility of each brand with the scene, and outputs the selected brand and its complete profile.
[0059] S2. Based on the text prompts and target brand profile, generate a brand integration strategy, and optimize the text prompts according to the brand integration strategy to obtain optimized prompts.
[0060] The strategy generation agent deeply analyzes the contextual characteristics of the user-input text prompts to identify potential brand integration points. It then queries the brand knowledge base's experience pool to retrieve historical cases similar to the current scenario, analyzing which strategies are effective in similar situations. Based on the scenario analysis and historical experience, the strategy generation agent generates specific brand integration strategies. This includes the way brand elements are presented, the timing of their presentation, and the extent of their presentation, ensuring that brand elements appear as an organic part of the scene.
[0061] Prompt word rewriting intelligent agent receiving brand integration strategy and text prompts The optimization and rewriting process follows four core principles: semantic preservation ensures the preservation of key subjects, actions, and style preferences in the text prompts; natural integration integrates brand elements as natural scene components; logical consistency maintains the temporal order and causal relationship coherence of the scene; and style consistency ensures compliance with the input specifications of the T2V model, ultimately generating optimized prompts.
[0062] S3. Conduct multi-dimensional evaluation of the optimized prompts and determine the target prompts based on the evaluation results.
[0063] Specifically, the process involves evaluating the optimized prompts across four dimensions: semantic fidelity, brand clarity, integration naturalness, and generation effect. If the scores for all four dimensions are greater than or equal to a preset threshold, the optimized prompt is selected as the target prompt. If the evaluation results include modification suggestions, the optimized prompt is modified according to these suggestions to obtain the target prompt.
[0064] Specifically, the process of modifying and optimizing prompts based on modification suggestions to obtain target prompts includes: if the modification suggestion is a re-optimization suggestion, the optimized prompts are used as text prompts, and the text prompts are optimized based on the brand integration strategy and the re-optimization suggestion to obtain optimized prompts. Then, S3 is executed again until the target prompts are obtained; if the modification suggestion is a strategy regeneration suggestion, the optimized prompts are used as text prompts, and S2 and S3 are executed again until the target prompts are obtained.
[0065] Evaluate the agent's performance on the generated optimization prompts. A comprehensive evaluation was conducted, including semantic fidelity (and...). The evaluation process assesses four dimensions: consistency, brand clarity (recognition and visibility), integration naturalness (the organic integration of scenes), and generated effect (expected video quality). If all dimensions score at the preset threshold, the evaluation agent accepts the optimized suggestion and proceeds to the video generation stage. If there are local issues but the overall strategy is correct, a revision suggestion (i.e., a re-optimization suggestion) is given. The optimized suggestion is then modified based on the brand integration strategy and the re-optimization suggestion, and the process returns to execute S3. If the integration strategy itself has problems and requires replanning (i.e., regenerating strategy suggestions), the process returns to execute S2 and S3.
[0066] S4. Generate target videos using target cue words and target brand profiles.
[0067] Specifically, the target prompts and target brand profiles are input into the video model to generate the target video.
[0068] In practical applications, the T2V model can be used to generate target videos based on the final accepted target prompts.
[0069] Furthermore, prior to S4, the method further includes:
[0070] If the target brand profile contains a brand adapter, then the adapter is loaded into the text-to-video model.
[0071] If the selected brand requires an adapter, the system conditionally loads the corresponding brand's adapter during the generation process, enabling the model to accurately generate the brand's visual features.
[0072] Furthermore, after S4, the method further includes:
[0073] Collect user feedback on the target video;
[0074] Extract structured lessons from feedback information and store them in the brand knowledge base.
[0075] After the target video is generated, the experience-learning agent collects user feedback (explicit or implicit) and abstracts this feedback into structured experience. For positive feedback, it extracts success patterns, records scene features, integrates strategies, and performance metrics. For negative feedback, it records the reasons for failure and strategies to avoid them. The abstracted structured experience is written back into the brand knowledge base's experience pool, achieving closed-loop learning.
[0076] The system completes brand integration tasks through the collaboration of five agents. Each agent achieves efficient collaboration through a dual-memory mechanism: all agents share access to a long-term brand knowledge base to obtain brand profiles, adapter weights, and historical experience, while simultaneously using short-term working contexts to pass intermediate results and status information, maintaining process continuity. The system employs an iterative improvement loop, repeatedly adjusting prompts or strategies based on evaluation feedback until quality requirements are met. Furthermore, through an experience feedback loop, each integration result is transformed into reusable knowledge, continuously improving overall performance.
[0077] The pseudocode of the algorithm flow of this invention is shown in Table 1 below.
[0078] Table 1. Pseudocode for Brand Integration Methods
[0079] Another embodiment of the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steps of the seamless brand integration method in text-to-video generation as described above.
[0080] The present invention produces the following significant beneficial effects:
[0081] (1) This invention has multiple advantages at the system design level. First, it is plug-and-play. This invention does not require modification of the underlying T2V model architecture, and brand integration is achieved only through prompt word optimization and optional adapter loading. This means that service providers can integrate this technology without modifying their existing systems, reducing deployment risks and costs. This invention can be seamlessly integrated with mainstream commercial models (such as Veo3, Sora2, Kling2.1) and open-source models, and has broad compatibility. The brand knowledge base and adapters can be managed and updated independently, supporting flexible brand registration, deregistration, and update operations. Continuous learning capability is another major architectural advantage of this invention. Through an experience-based learning mechanism, system performance continuously improves with the number of uses. The accumulation of successful experiences makes the system more efficient and accurate in handling similar scenarios, and can quickly identify the best integration strategy. This self-improvement capability is not available in static methods, providing a guarantee for the long-term operational value of the system.
[0082] (2) This invention establishes a sustainable business model for text-to-video generation services. This model creates new revenue streams for T2V service providers, enabling them to charge brands for integrated brand services, thereby covering high computing and operational costs and ultimately achieving profitability. For brands, they gain organic exposure in high-quality user-generated content; this integrated advertising is more effective than traditional intrusive advertising, with higher user acceptance and stronger brand recall. For end users, they can continue to enjoy high-quality video generation services for free or at a low price, as the service costs are borne by the brands, and the user experience is not significantly negatively affected.
[0083] In summary, this invention provides a promising T2V commercialization strategy, which is of great significance for ensuring the sustainable development of the content creation industry. This invention not only solves technical challenges but also establishes a commercial ecosystem, providing plug-and-play solutions for application scenarios such as advertising production, content creation, and brand marketing.
[0084] The above descriptions are merely a few embodiments of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any modifications or alterations made by those skilled in the art without departing from the scope of the technical solution of the present invention using the disclosed technical content are equivalent to equivalent implementation cases and fall within the scope of the technical solution.
Claims
1. A method for seamless brand integration in text-to-video generation, characterized in that, The method includes: S1. Construct a brand knowledge base containing multiple brand profiles, receive text prompts, and query the brand knowledge base for the target brand profile corresponding to the text prompts. S2. Based on the text prompts and the target brand profile, generate a brand integration strategy, and optimize the text prompts according to the brand integration strategy to obtain optimized prompts; S3. Evaluate the optimized prompt words from multiple dimensions and determine the target prompt words based on the evaluation results; S4. Using the target prompts and the target brand profile, generate the target video.
2. The method according to claim 1, characterized in that, The construction of a brand knowledge base comprising multiple brand profiles in S1 specifically includes: Based on the brand's initial profile, generate multiple test prompts, and based on the initial profile and each test prompt, generate a corresponding test video; If the percentage of test videos that meet the preset requirements among all test videos of the brand is greater than or equal to the preset percentage, then the initial file of the brand will be stored in the brand knowledge base as the brand file of the brand. If the percentage of test videos that meet the preset requirements among all test videos of the brand is less than the preset percentage, then the adapter of the brand is determined according to the initial file, and the initial file and adapter of the brand are stored as the brand file of the brand in the brand knowledge base.
3. The method according to claim 2, characterized in that, Determining the adapter for the brand based on the initial file specifically includes: Multiple comprehensive prompts are generated based on the initial profile; each comprehensive prompt includes a brand name and a trigger marker. Based on the reference image in the initial file and the multiple comprehensive prompt words, multiple training videos are generated accordingly, and the multiple comprehensive prompt words and the corresponding multiple training videos are combined into a training dataset. The text-to-video model is trained using the training dataset to obtain the adapter for the brand.
4. The method according to claim 3, characterized in that, Based on the reference image in the initial file and the multiple synthesized prompt words, multiple training videos are generated, specifically including: The reference image in the initial file and the input text of each comprehensive prompt word are fed into the image model to generate an initial frame corresponding to each comprehensive prompt word; The initial frame corresponding to each of the comprehensive prompt words is input into the image-to-video model to generate a training video corresponding to each of the comprehensive prompt words.
5. The method according to claim 1, characterized in that, S3 specifically includes: The optimized prompts were evaluated across four dimensions: semantic fidelity, brand clarity, integration naturalness, and generation effect, yielding the evaluation results. If the scores of all four dimensions in the evaluation results are greater than or equal to the preset threshold, then the optimized prompt word will be used as the target prompt word. If the evaluation results contain modification suggestions, then the optimized prompt words are modified according to the modification suggestions to obtain the target prompt words.
6. The method according to claim 5, characterized in that, Modify the optimized prompt words according to the modification suggestions to obtain the target prompt words, specifically including: If the modification suggestion is a re-optimization suggestion, then the optimized prompt word is used as the text prompt word, and the text prompt word is optimized according to the brand integration strategy and the re-optimization suggestion to obtain the optimized prompt word. Then S3 is executed again until the target prompt word is obtained. If the proposed modification is to regenerate the strategy, then the optimized prompt word is used as the text prompt word, and S2 and S3 are re-executed until the target prompt word is obtained.
7. The method according to claim 1, characterized in that, Specifically, S4 is: The target prompt and the target brand profile are input into the video model to generate the target video.
8. The method according to claim 7, characterized in that, Prior to S4, the method further includes: If the target brand profile contains an adapter for the brand, then the adapter is loaded into the text-to-video model.
9. The method according to claim 1, characterized in that, Following S4, the method further includes: Collect user feedback on the target video; Structured experience is extracted from the feedback information and stored in the brand knowledge base.
10. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the seamless brand integration method in text-to-video generation as described in any one of claims 1 to 9.