Artificial intelligence-based advertisement material automatic fission method and system
By monitoring advertising platform data in real time, extracting multimodal features, and selecting dynamic strategies, the system solves the problems of inefficiency and homogenization in advertising creative fission, achieving efficient and intelligent automatic fission of advertising materials and improving creative quality and campaign performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- FEIYU (GUANGZHOU) INTERACTIVE MEDIA CO LTD
- Filing Date
- 2025-09-30
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies are inefficient in advertising creative fission, rely on human experience, suffer from severe content homogenization, and have unstable results, failing to effectively inherit the successful elements of hit creative materials.
By monitoring the performance data of the advertising platform in real time, utilizing multimodal feature extraction and a central scheduling engine, a fission strategy is dynamically selected to generate and test derivative creative materials, forming a closed-loop optimization system.
It has enabled the automation, scaling, and intelligent proliferation of advertising creative ideas, improved the quality and diversity of derived creative materials, extended the life cycle of creative assets, and increased the return on investment.
Smart Images

Figure CN121258610B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital advertising, and in particular to a method and system for automatic fission of advertising creatives based on artificial intelligence. Background Technology
[0002] In the digital advertising field, to maximize the ROI of advertising campaigns, advertisers and operating platforms are constantly seeking ways to extend the lifecycle of high-performing ad creatives (often referred to as "hit creatives") and expand their reach. A common strategy is to "fission" proven hit creatives, that is, to generate new derivative creatives in batches by modifying or reorganizing their constituent elements, in order to replicate or continue their excellent campaign performance.
[0003] Currently, achieving ad creative viral marketing primarily relies on the following technical means: First, the traditional manual viral marketing method. This method is usually led by experienced operations personnel or creative designers who manually analyze the success factors of viral creative materials, such as visually impactful openings, captivating narrative rhythms, or high-conversion copywriting. They then manually re-edit the videos, replace storyboards, modify background music, or change copywriting to create new versions. However, this method heavily depends on human experience, making it inefficient, costly, and difficult to scale up, failing to quickly respond to rapidly changing market demands. Second, there is the automated replacement method based on preset templates. This method provides fixed video or image templates, allowing users to batch replace key elements such as product images, price tags, or slogans. While this method improves the efficiency of material production to some extent, its limitations are also significant. Because the template structure is fixed, it cannot dynamically learn and inherit the unique, market-proven narrative structure or visual language of non-template-based viral creative materials. This results in highly homogenized derivative materials, easily causing user fatigue and a rapid decline in click-through and conversion rates. Finally, with the development of artificial intelligence (AI) technology, some tools have emerged that utilize AI for content generation. These tools can provide assistance in specific stages, such as automatically generating ad copy or performing simple video editing. However, existing AI tools often have limited functionality and lack a holistic and in-depth understanding of viral creative materials. They typically cannot comprehensively analyze multimodal information such as video, audio, and text, and struggle to accurately identify and inherit the core combination logic behind the success of viral creative materials. Consequently, the quality of the generated derivative materials varies greatly, resulting in volatile campaign performance and an inability to guarantee consistently high conversion rates.
[0004] In summary, existing technologies for advertising creative viralization generally suffer from several drawbacks, including difficulty in balancing efficiency and creative quality, insufficient automation and content diversity, and a lack of mechanisms for inheriting the success factors of viral creative materials. Therefore, the industry urgently needs a technological solution that can automate, scale, and intelligently generate advertising creative viralization to effectively address these shortcomings. Summary of the Invention
[0005] The purpose of this invention is to provide an automatic fission method and system for advertising creatives based on artificial intelligence, which aims to solve the technical problems in the existing technology such as reliance on human experience, low efficiency, serious content homogenization, and unstable fission effect.
[0006] In a first aspect, embodiments of the present invention propose an automatic fission method for advertising creatives based on artificial intelligence, the method comprising: The application programming interface (API) monitors and aggregates creative performance data at the creative material level of one or more advertising platforms in real time. When at least one performance indicator in the performance data meets a data-driven trigger condition that can be dynamically optimized based on historical campaign data and business objectives, the corresponding best-selling creative material is automatically identified. Multimodal feature extraction is performed on popular creative materials to generate a structured feature vector representing the visual, textual, and audio information corresponding to the popular creative materials; A central scheduling engine dynamically combines and selects at least one fission strategy from a strategy library containing multiple fission strategies, based on diagnostic analysis of structured feature vectors. Based on the selected combination of fission strategies, the corresponding generative artificial intelligence model or application interface is invoked to transform the viral creative material in order to generate at least one new derivative creative material. New derivative creative materials are deployed to the advertising platform for testing, and performance data of the new derivative creative materials is collected. The performance data is then fed back to the central scheduling engine to update the subsequent viral strategy selection logic.
[0007] Preferably, the performance metrics include at least one of click-through rate, conversion rate, return on advertising investment, and cost per conversion.
[0008] Preferably, the step of extracting multimodal features from viral creative materials includes: using a convolutional neural network to extract the main visual framework and color distribution features of the viral creative materials; using automatic speech recognition to extract spoken text and using optical character recognition to extract subtitles within video frames, so as to jointly constitute text features; and using a beat detection algorithm to analyze the waveform of background music to extract audio rhythm features.
[0009] Preferably, in the step of generating at least one new derivative creative material, when the selected fission strategy includes a dynamic narrative reconstruction strategy, it includes: The video content of viral creative materials is automatically cut into multiple shots using a lens boundary detection algorithm; A narrative graph is constructed by using multiple storyboards as nodes and the temporal transitions between storyboards as edges. The narrative graph is processed using a graph neural network trained to predict new node sequences that maximize the expected delivery effect, in order to generate new node sequences and reassemble the storyboards based on the new node sequences.
[0010] Preferably, in the step of generating at least one new derivative creative material, when the selected fission strategy includes a storyboard replacement strategy, it includes: Feature vectors are extracted from target scenes in popular creative materials. In a pre-set material library, the dot product of the feature vectors of candidate scenes and target scenes is calculated and normalized according to the magnitude of each vector to obtain the cosine similarity that quantifies the similarity of content and style between the two in a multi-dimensional feature space. A set of candidate scenes with similarity higher than the preset similarity threshold is then selected. Furthermore, a pre-trained multimodal model is used to perform cross-modal semantic consistency verification between each candidate segment in the candidate segment set and the original context of the target segment, and the candidate segment with the highest verification score is selected for replacement.
[0011] Preferably, in the step of generating at least one new derivative creative material, when the selected fission strategy includes a pre-post expansion strategy, the following steps are included: Automatically detect climactic moments in viral creative materials using facial expression recognition APIs or audio loudness analysis; Place the climax scene at the starting point of new derivative creative material; It also utilizes a large language model to generate background text that logically forms a flashback relationship with the climax, using the climax segment as input. Based on the background text, it calls a text-to-video model to generate or matches corresponding video footage from a media library and splices it before the climax segment.
[0012] Preferably, in the step of generating at least one new derivative creative material, when the selected fission strategy includes a character replacement strategy, the following steps are included: Using voiceprint separation, the original spoken audio can be extracted from viral creative materials; It also calls the digital human generation application interface, takes the original spoken audio and the preset target digital human image as input, and generates a new video storyboard. The new video storyboard contains the target digital human, and uses a lip-sync model to ensure that the lip movements of the target digital human match the original spoken audio.
[0013] Preferably, the step of deploying new derivative creative materials to the advertising platform for testing is modeled as a multi-arm machine model, wherein each derivative creative material is treated as an independent arm; Its performance indicators during the testing period serve as reward signals; And through a strategy designed to balance exploration and utilization, with the optimization goal of maximizing cumulative rewards or minimizing cumulative regrets, the testing budget for each arm is dynamically allocated to quickly determine the optimal derivative creative material within a preset testing period.
[0014] Preferably, before the step of deploying the new derivative creative materials to the advertising platform for testing, the method further includes: New derivative creative materials are submitted to a compliance filter layer for automated review. The compliance filter layer compares them with a copyright database and an advertising regulations knowledge base, and automatically blocks derivative creative materials that do not meet the preset compliance standards.
[0015] Secondly, embodiments of the present invention propose an automatic advertising creative splitting system based on artificial intelligence, comprising: The data monitoring and triggering module is used to monitor and aggregate creative material granular performance data from one or more advertising platforms in real time, and automatically identify the corresponding best-selling creative material when at least one performance indicator in the performance data meets a data-driven triggering condition that can be dynamically optimized based on historical campaign data and business objectives. The multimodal feature extraction module is used to extract multimodal features from popular creative materials to generate a structured feature vector containing visual, textual, and audio information. The central scheduling and strategy selection module is used to dynamically combine and select at least one fission strategy from a strategy library containing multiple fission strategies based on the diagnostic analysis of structured feature vectors. The strategy execution and generation module is used to transform popular creative materials by calling corresponding generative artificial intelligence models or application programming interfaces according to the selected fission strategy combination, so as to generate at least one new derivative creative material. The strategy execution and generation module includes: a graph neural network processor for executing dynamic narrative reconstruction strategy; a cross-modal verification unit for calling a pre-trained multimodal model to ensure semantic coherence when executing the storyboard replacement strategy; and a digital human synthesis unit for calling a lip-sync model to ensure audio-visual synchronization when executing the character replacement strategy. It also includes a closed-loop testing and optimization module, which is used to deploy new derivative creative materials to the advertising platform for testing, collect performance data of the new derivative creative materials, and feed the performance data back to the central scheduling and strategy selection module to update the subsequent fission strategy selection logic. Beneficial effects
[0016] By monitoring the performance data of advertising platforms in real time and setting dynamically optimized data-driven trigger conditions, this invention can automatically and accurately identify creative materials with the potential to become viral hits. It achieves automated screening for the viral spread of high-quality creatives, overcoming the shortcomings of traditional manual viral spread methods that rely on subjective experience, are inefficient, and difficult to scale. Specifically, it introduces a step of multimodal feature extraction for viral creative materials, enabling in-depth analysis of the inherent logic behind the success of the materials from multiple dimensions such as visual, textual, and audio perspectives, ensuring that subsequent viral spreads are not simply element replacements. Based on this, the central scheduling engine can dynamically combine and select the most suitable viral spread strategy from the strategy library according to the diagnostic analysis results. This solves the problems of homogenized derivative content and inability to inherit the core narrative structure of viral hits caused by existing viral spread methods based on fixed templates or single AI tools, significantly improving the quality and diversity of derivative creative materials. By testing the newly generated derivative creative materials and feeding their performance data back to the central scheduling engine, the logic for selecting subsequent viral spread strategies can be continuously updated and optimized. This data-driven adaptive learning mechanism ensures that the entire fission system can continuously evolve, effectively balancing content innovation and performance stability, thereby systematically improving the success rate of new material delivery, extending the life cycle of creative assets, and ultimately bringing higher returns on investment to advertisers. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the embodiments will be briefly described below with reference to the accompanying drawings. It should be noted that the accompanying drawings are merely examples and not intended to limit the present invention; the same reference numerals in the drawings denote the same components. In the accompanying drawings: Figure 1 This is a flowchart of an AI-based automatic advertising material splitting method in an embodiment of the present invention; Figure 2 This is a schematic diagram of the functional modules of the AI-based automatic advertising material splitting system in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of the computer device proposed in the embodiments of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0019] For the first aspect, see [link / reference] Figure 1 As shown in the illustration, this invention proposes an automatic advertising creative creation method based on artificial intelligence. This method can be applied to an AI-based automatic advertising creative creation system, which can be executed by, but is not limited to, a server cluster with computing and storage resources, and interacts with the application programming interface (API) of one or more advertising platforms via a network. The method provided by this invention aims to solve the problems of existing technologies, such as reliance on human experience, low efficiency, severe content homogenization, and unstable creation effects, in advertising creative creation. By constructing a data-driven, strategy-rich, and closed-loop optimized automated process, it achieves efficient, large-scale, and intelligent re-creation of popular advertising creative materials. The core process of the method proposed in this embodiment includes automatic identification of popular creative materials, multimodal deep analysis, dynamic strategy combination creation based on diagnosis, generation of derivative materials, and final automated delivery testing and closed-loop optimization. The following detailed description of the method is provided in conjunction with specific implementation methods: Step S1: Monitor and aggregate creative material granularity performance data from one or more advertising platforms in real time through the application programming interface. When at least one performance indicator in the performance data meets a data-driven trigger condition that can be dynamically optimized based on historical campaign data and business objectives, the corresponding best-selling creative material is automatically identified.
[0020] Specifically, this step is the starting point of the entire automated viral marketing process. Its purpose is to accurately and promptly capture high-performing ad creatives in the market as "seeds" for subsequent viral marketing. For example, by establishing stable data connections with major advertising platforms (such as Facebook Ads Manager API, Google Ads API, TikTok for Business API, etc.), real-time or near real-time data retrieval is performed (e.g., a polling cycle that can be configured every 15 minutes). The retrieved data dimension is at the creative level, which means that the system can obtain detailed performance data for each individual ad video or image creative, rather than macro-level data for ad campaigns or ad sets.
[0021] Preferably, the performance metrics include at least one of click-through rate, conversion rate, return on advertising investment, and cost per conversion.
[0022] In this embodiment, the key performance indicators (KPIs) monitored and aggregated by the system are multi-dimensional, comprehensively reflecting the attractiveness, conversion rate, and cost-effectiveness of the content. These indicators specifically include, but are not limited to: Click-through rate (CTR), calculated as (number of clicks / number of impressions), directly reflects the visual appeal of the creative to the target audience and the strength of the content hook. Conversion rate (CVR), calculated as (number of conversions / number of clicks), measures the ability of the creative to guide users to complete a specific goal action (such as downloading, registering, or purchasing). Return on ad spend (ROAS), calculated as (revenue generated by advertising / advertising expenditure), is a core indicator for measuring the profitability of an advertising campaign. Cost per acquisition (CPA), calculated as (total advertising expenditure / number of conversions), reflects the cost of acquiring a valid conversion.
[0023] Secondly, a data-driven trigger condition is set up to determine whether a creative meets the "viral" standard. This trigger condition is not a fixed, unchanging value, but a floating threshold that can be dynamically optimized based on historical campaign data and current business objectives. For example, the system can define creatives with CTR or CVR ranking in the top 5% based on the performance data distribution of all creatives in the same category and channel over the past 30 days as viral; or, when the ROAS of a new creative exceeds a preset benchmark value (such as 2.0) within 24 hours of its launch, viral identification is triggered. This dynamic threshold setting allows the viral identification standard to adapt to changes in the market environment and the differences in objectives of different advertising campaigns.
[0024] Once one or more performance indicators of a creative material meet the trigger conditions, the system automatically marks it as a "viral creative material" and stores its ID, metadata, performance data, and the material file itself, such as the video file URL, into a dedicated "viral creative material pool." At the same time, it is sent to the subsequent fission processing queue to await further in-depth analysis.
[0025] Step S2 involves extracting multimodal features from the viral creative material to generate a structured feature vector representing the visual, textual, and audio information corresponding to the viral creative material.
[0026] Specifically, after identifying viral creative materials, a deep, "dissecting" analysis is needed to transform their unstructured multimedia content into machine-understandable, structured feature information, providing a basis for subsequent strategy selection.
[0027] Preferably, the step of multimodal feature extraction for viral creative materials includes: extracting the main visual framework and color distribution features of the viral creative materials using a convolutional neural network; extracting the spoken text using automatic speech recognition and extracting in-video subtitles using optical character recognition, to jointly constitute text features; and analyzing the waveform of background music using a beat detection algorithm to extract audio rhythm features. In this embodiment, multimodal feature extraction covers the following three main aspects: For visual feature extraction, the system first decodes the video footage into a sequence of image frames. For main visual framework and object detection, a pre-trained Convolutional Neural Network (CNN), such as ResNet or EfficientNet, is used to process keyframes and extract deep feature vectors that characterize the scene, composition, and core elements. Simultaneously, an object detection model (such as the "You Only Look Once" series, YOLO) is employed to identify and label key objects within the frames (such as people, specific products, brand logos, etc.) and their locations. Color distribution features: The system analyzes the overall hue, saturation, and brightness distribution of the video, as well as the color composition of keyframes. These features are closely related to user emotional responses and brand tone.
[0028] Text feature extraction: Text is key to conveying information and guiding conversion. For videos containing audio, the system uses an Automatic Speech Recognition (ASR) engine to convert the spoken content into a timestamped text script. For intra-frame subtitle extraction, the system uses Optical Character Recognition (OCR) technology to scan and extract embedded subtitles or titles from the video frame by frame. For semantic analysis, the extracted spoken text and subtitles are merged, and a Natural Language Processing (NLP) model is used for sentiment analysis (determining whether the text is positive, negative, or neutral), keyword extraction (such as high-value words like "limited-time offer" and "free trial"), and topic modeling.
[0029] Audio feature extraction: Audio is a crucial element in creating atmosphere and controlling rhythm. Background music analysis: Using beat detection algorithms, the waveform of background music is analyzed to extract features such as beats per minute (BPM) and rhythm intensity. These features are essential for the pacing of video editing. Mood and style classification: Audio classification models are used to determine the mood (e.g., upbeat, soothing, tense) and style (e.g., pop, classical, electronic) of the background music.
[0030] Through the above processing, the original viral creative material is transformed into a high-dimensional structured feature vector. This vector comprehensively describes all the key attributes of the material in terms of visual composition, information delivery, and emotional rhythm, forming a "genetic map".
[0031] Step S3 involves a central scheduling engine dynamically combining and selecting at least one fission strategy from a strategy library containing multiple fission strategies, based on diagnostic analysis of structured feature vectors.
[0032] Specifically, this step is the brain of the entire method. The core task of the Central Scheduling Engine is to act as an experienced AI creative director. It receives the multimodal feature vectors generated in the previous step, performs intelligent diagnosis on them, and then decides on the optimal fission plan.
[0033] First, the central scheduling engine performs diagnostic analysis on the feature vectors. For example, it analyzes the video's viewing curve data (such as the second-level completion rate) and, combined with the feature vectors, finds that the video's completion rate is as high as 95% in the first 3 seconds, indicating that its "hook" is very successful; however, there is a significant peak in viewership drop-off in the middle of the video (10-15 seconds). Through feature analysis, it is located that this period is when a real person appears on camera to explain product parameters, and the visuals are relatively monotonous.
[0034] Then, based on the diagnostic results, the engine dynamically combines and selects strategies from a pre-set strategy library. This strategy library integrates a variety of proven fission strategies for different optimization objectives. In this embodiment, the strategy library includes at least core strategies such as dynamic narrative reconstruction, storyboard replacement, pre-post expansion, and character replacement.
[0035] Continuing with the example above, based on the diagnostic conclusions, the engine might make the following decisions: Retain the successful "hook" at the beginning, therefore choosing the "retain the hit beginning" strategy. Optimize the middle section where there is significant drop-off, therefore combining the "character replacement" strategy (replacing the live narrator with a more novel digital human) and the "scene replacement" strategy (replacing the monotonous narration with more visually impactful product close-ups or usage scenarios). Do not change the overall script logic of the video, therefore not choosing the "re-editing" or "pre-post expansion" strategies.
[0036] Ultimately, the engine outputs a fission task instruction, which specifies the combination of strategies to be executed on the original viral creative material (e.g., {retain the viral opening, replace characters, replace storyboards}) and the specific parameters of each strategy. This diagnostic-based dynamic strategy combination mechanism ensures that each fission is a targeted optimization aimed at making up for the shortcomings of the original work or amplifying its advantages, rather than a blind, random modification.
[0037] Step S4: Based on the selected fission strategy combination, call the corresponding generative artificial intelligence model or application interface to transform the viral creative material to generate at least one new derivative creative material.
[0038] Specifically, this step is the execution phase of the viral marketing strategy. Based on the task instructions output by the central scheduling engine, the system invokes the corresponding AIGC (Artificial Intelligence Generated Content) module to perform the actual transformation of the viral creative materials. The following will elaborate on the specific implementation methods of several preferred viral marketing strategies: Preferably, when the selected fission strategy includes a dynamic narrative reconstruction strategy, the step includes: using a shot boundary detection algorithm to automatically cut the video content of the viral creative material into multiple shots; using the multiple shots as nodes and the temporal transition relationships between shots as edges to construct a narrative graph; and using a graph neural network trained to predict a new node sequence that can maximize the expected delivery effect to process the narrative graph to generate a new node sequence, and recombining the shots according to the new node sequence.
[0039] First, advanced shot boundary detection algorithms, such as HSV color histogram differencing or deep learning-based methods, are employed to precisely segment the continuous video stream into a series of independent shots. Then, each shot is abstracted as a "node" in a graph, with each node carrying its multimodal feature vector. The original temporal relationships and transitions between shots are abstracted as "edges" in the graph. In this way, the entire video's narrative structure is mathematically represented as a directed graph, or "narrative graph." Finally, the system utilizes a specially trained Generative Neural Network (GNN). This GNN is trained to learn the narrative patterns of viral videos and predict which node sequences (i.e., shot arrangements) will maximize the expected delivery effect (such as CTR or CVR). By reordering nodes, adding, deleting, or modifying edge relationships on the narrative graph, the GNN can generate one or more entirely new node sequences that still follow efficient narrative logic. The system then reassembles and combines these new sequences to generate structurally refreshed derivative creative materials.
[0040] Preferably, when the selected fission strategy includes a scene replacement strategy, the step includes: extracting feature vectors from the target scene in the viral creative material, and filtering a set of candidate scenes by calculating cosine similarity in a pre-set material library; and using a pre-trained multimodal model to perform cross-modal semantic consistency verification, and selecting the candidate scene with the highest verification score for replacement.
[0041] Feature vector matching and similarity matching: For a target scene that needs to be replaced, the system extracts its multimodal feature vector. Then, in a pre-set material library containing a large number of backup scenes, the system calculates the cosine similarity between the target scene and the feature vectors of all scenes in the library, using the formula similarity=(A·B) / (||A|| ||B||), and selects a set of candidate scenes with similarity scores higher than a preset threshold (e.g., ≥0.7).
[0042] Cross-modal semantic consistency verification is the core innovation of this strategy. To avoid visually similar but semantically incompatible replacement scenes (i.e., "semantic gaps"), the system utilizes a powerful pre-trained multimodal model (such as CLIP). This model understands the relationship between image content and text description. The system inputs the keyframes of each candidate scene, along with the text description of its original context (e.g., the corresponding spoken text extracted by ASR), into the CLIP model for cross-modal semantic consistency scoring. Ultimately, only the candidate scene with the highest score, i.e., the one that best fits the context semantically, is selected for replacement. For example, for a scene of "a woman drinking coffee in a coffee shop," the system will prioritize replacing it with "a man drinking juice in a park" rather than "a woman working in an office," because the former is more semantically consistent with "leisure and enjoyment."
[0043] Preferably, when the selected viral marketing strategy includes a pre-post expansion strategy, the step includes: automatically detecting the climax segment in the viral creative material; placing it at the beginning of the new material; and using a large language model to generate flashback text and corresponding visuals for splicing.
[0044] The system uses multi-dimensional analysis to locate the climax of a video. Visually, it can call facial expression recognition APIs (such as Face++) to detect frames with strong emotions such as "surprise" and "excitement"; audio-wise, it can detect sudden increases in volume or screams. The detected climax segments (usually 1-3 seconds) are directly edited and placed at the very beginning of new derivative creative materials as a powerful "hook" to attract the user's attention.
[0045] Large Language Models (LLMs) generate flashbacks, using the content of the climax as context. They then send a carefully crafted prompt to the LLM, such as GPT-4, requesting it to generate a logically connected flashback. For example, for the climax of "the phone fell from a great height and survived intact," the LLM might generate: "Can you believe it? We threw this new phone from the tenth floor…."
[0046] Based on the introductory text generated by LLM, a text-to-video (CTV) model (such as Sora or Runway) can be called to directly generate matching video footage, or semantically matching shots can be retrieved from the media library and seamlessly spliced with the preceding climax to form a complete, suspenseful, and engaging new opening.
[0047] Preferably, when the selected fission strategy includes a character replacement strategy, the step includes: separating the original spoken audio; and calling the digital human generation interface, inputting the audio and the target image, to generate a new video storyboard driven by a lip-sync model.
[0048] Utilizing advanced voiceprint separation technology, clean, noise-free original audio is precisely extracted from the complex audio track of the original video. Third-party digital human generation APIs (such as D-ID and Synthesia) or a self-developed digital human model are then called. The extracted audio and a pre-defined target digital human avatar (which can be a virtual avatar or a digital clone generated from a real-life photograph) are used as input. A crucial step in the generation process is using a lip-sync model, such as Wav2Lip, to ensure that the digital human's lip movements precisely match the input audio in every frame. This is key to enhancing realism and audience trust. Finally, a new video storyboard featuring the target digital human delivering the audio is generated and used to replace the corresponding portion of the original video.
[0049] Step S5: Deploy the new derivative creative materials to the advertising platform for testing, collect performance data of the new derivative creative materials, and feed the performance data back to the central scheduling engine to update the subsequent fission strategy selection logic.
[0050] Specifically, this step builds a data-driven, automated closed-loop optimization system to ensure that the fission behavior can continuously learn and evolve.
[0051] First, before the official launch, the system will conduct pre-launch screening and compliance review.
[0052] Preferably, a compliance filtering layer is included before the step of testing new derivative creative materials on the advertising platform.
[0053] The generated derivative materials are first submitted to a compliance filtering layer. This layer automatically compares them with a copyright database (checking for potential copyright infringement risks in background music and visual materials) and an advertising regulations knowledge base (checking for prohibited words and compliance with advertising laws in the copy), automatically blocking materials that do not meet preset compliance standards. In addition, the materials are fed into a pre-trained performance prediction model. Only materials with predicted performance metrics (such as predicted CTR) above a certain threshold are allowed to proceed to the next stage of human testing, significantly saving on testing budgets.
[0054] Secondly, the system employs an efficient A / B testing method to verify the selected derivative materials.
[0055] Preferably, the step of deploying new derivative creative materials to an advertising platform for testing is modeled as a multi-armed bandit (MAB) model. In this model: Each derivative creative asset to be tested is considered an independent "arm." Its key performance indicators (KPIs) during the testing period (such as CVR) are used as "reward" signals. The system dynamically allocates the testing budget to each "arm" using a strategy designed to balance "exploration" (trying new, uncertain-effect assets) and "exploration" (allocating more budget to assets known for their good performance) (such as the UCB algorithm or Thompson sampling). The goal of this method is to quickly and accurately determine which derivative creative asset is the best-performing "winner" within a pre-defined testing period (e.g., 24 hours) at minimal cost. Finally, the test results are used for closed-loop feedback and optimization. After the A / B test, the system collects performance data for all test assets and performs statistical significance tests (such as t-tests, p-value < 0.05).
[0056] If a winner emerges—that is, a derivative creative significantly outperforms the original viral hit—the system will automatically elevate it to the new "viral hit" benchmark and allocate more budget for its large-scale deployment. Simultaneously, this success story (i.e., "diagnosis results -> strategy combination -> successful results") will be recorded to strengthen the weight of the corresponding decision rules in the central scheduling engine. If there is no clear winner or all fail, the system will analyze the common characteristics of the failed creatives (e.g., all creatives that replaced the digital human performed worse) and feed this "lesson learned" back to the central scheduling engine to adjust or reduce the priority of selecting this viral strategy in the future. Through this complete closed-loop process of "identification-analysis-viral-testing-feedback," the method of this invention can form a continuously self-learning and self-optimizing intelligent system, continuously improving the efficiency and effectiveness of advertising creative production.
[0057] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0058] For the second aspect, please refer to... Figure 2 This invention also proposes an AI-based automatic creative creative generation system, adapted to the AI-based automatic creative creative generation method described in the preceding embodiments. This system is typically deployed on servers with high-performance computing capabilities (such as Graphics Processing Unit (GPU) clusters) and massive storage capacity, aiming to provide the digital advertising industry with an end-to-end, automated creative creative optimization solution. The system is described in detail below with reference to specific implementation methods. The system includes a data monitoring and triggering module, a multimodal feature extraction module, a central scheduling and strategy selection module, a strategy execution and generation module, and a closed-loop testing and optimization module. These modules collaborate through an internal data bus and API calls, forming an intelligent creative generation and optimization workflow.
[0059] The data monitoring and triggering module is used to monitor and aggregate creative material performance data at the creative material level from one or more advertising platforms in real time. When at least one performance indicator in the performance data meets a data-driven triggering condition that can be dynamically optimized based on historical campaign data and business objectives, the module automatically identifies the corresponding best-selling creative material.
[0060] Specifically, this module acts as the system's "data antennae," integrating multiple API clients to connect with mainstream advertising platforms such as Facebook Ads, Google Ads, and TikTok for Business. Through periodic (e.g., every 15 minutes) API polling, the module proactively retrieves performance data for each individual creative, such as click-through rate (CTR), conversion rate (CVR), and return on investment (ROAS). Internally, the module maintains a dynamic threshold calculation unit that continuously updates the triggering conditions for determining "viral" results based on historical data (e.g., the average performance of similar creatives over the past 30 days) and current business goals (e.g., the ROAS target set for a marketing campaign). When a creative's performance metrics meet these dynamic conditions, the trigger is activated. The module then packages the creative's ID, metadata, performance data, and source file address, temporarily storing it in the system's "viral creative pool" and issuing processing instructions to the multimodal feature extraction module.
[0061] The multimodal feature extraction module is used to extract multimodal features from viral creative materials to generate a structured feature vector containing visual, textual, and audio information. Specifically, this module is the system's "analysis engine," responsible for transforming unstructured video or image materials into machine-readable deep features. This module consists of multiple parallel processing pipelines: The visual processing pipeline incorporates a feature extractor based on a convolutional neural network (CNN, such as ResNet) to analyze scene, composition, and color distribution in video keyframes; it also integrates an object detection model (such as YOLO) to identify core elements in the image, such as people and products. The text processing pipeline integrates an Automatic Speech Recognition (ASR) engine and an Optical Character Recognition (OCR) engine to extract spoken text and embedded subtitles from the video, respectively. The extracted text data is then fed into a Natural Language Processing (NLP) unit for sentiment analysis and keyword extraction. The audio processing pipeline includes a beat detection algorithm unit to analyze the rhythmic features (such as BPM) of background music; and an audio classification model to determine the mood and style of the music. The outputs of all processing pipelines are ultimately integrated into a unified high-dimensional structured feature vector, which comprehensively depicts the "success genes" of viral creative materials and is then passed to the central scheduling and strategy selection module.
[0062] The central scheduling and strategy selection module is used to dynamically combine and select at least one fission strategy from a strategy library containing multiple fission strategies, based on the diagnostic analysis of structured feature vectors.
[0063] Specifically, this module is the system's decision-making brain. It receives structured feature vectors from the multimodal feature extraction module and combines them with the raw performance data of the content (e.g., second-level completion rate curves) to perform intelligent diagnosis. For example, the diagnosis unit might find that the retention rate is extremely high in the first 3 seconds of the content, but a certain explanatory segment in the middle leads to a serious drop-off of users. Based on this diagnosis, the strategy selection unit will query the internal strategy library—which predefines various viral strategies such as dynamic narrative reconstruction, scene replacement, pre-post expansion, and character replacement. The selection unit dynamically combines one or more viral strategies that best suit the current diagnosis result through a set of rule-based or machine learning model-based decision logic (e.g., combining the "retaining the viral beginning" strategy with the "character replacement" strategy), and sends the task instruction containing the strategy combination and related parameters to the strategy execution and generation module.
[0064] The strategy execution and generation module is used to transform popular creative materials by calling the corresponding generative artificial intelligence model or application interface according to the selected combination of fission strategies, so as to generate at least one new derivative creative material.
[0065] Specifically, this module serves as the system's creative factory, responsible for translating decisions into actual derivative creative ideas. It is a highly integrated AIGC (Artificial Intelligence Generated Content) platform, comprising multiple dedicated processing units to execute different fission strategies: The Graph Neural Network Processor (GNN) is activated when it receives an instruction to execute the "Dynamic Narrative Reconstruction Strategy." It first calls a shot boundary detection algorithm to segment the video into scenes, then constructs a narrative graph from the scene sequence, and uses a built-in Graph Neural Network (GNN) model to optimize and reorganize the graph structure, ultimately generating a video with a completely new narrative rhythm. The Cross-Modal Validation Unit plays a crucial role when executing the "Scene Replacement Strategy." After matching visually similar candidate scenes from the media library using cosine similarity, this unit calls a pre-trained multimodal model (such as CLIP) to perform semantic consistency checks on the candidate scenes and their original context text descriptions, ensuring that the replaced content is logically and semantically coherent, thus avoiding "semantic gaps." The Digital Human Synthesis Unit is invoked when executing the "Character Replacement Strategy." It first uses voiceprint separation technology to extract the original spoken audio, then sends the audio and the target digital human image to the Digital Human Generation API. During the synthesis process, this unit will specifically call a lip-sync model (such as Wav2Lip) to ensure that the lip movements of the generated digital human video are accurately matched with the audio, so as to guarantee the realism and credibility of the final product.
[0066] The closed-loop testing and optimization module is used to deploy new derivative creative materials to the advertising platform for testing, collect performance data of the new derivative creative materials, and feed the performance data back to the central scheduling and strategy selection module to update the subsequent viral strategy selection logic.
[0067] Specifically, this module serves as the system's "quality control and evolution center," responsible for verifying the viral marketing effect and driving system iteration. Its workflow includes: pre-launch screening, where generated derivative materials first undergo an automated review process connected to a copyright database and advertising regulations knowledge base. Materials passing this review are then fed into an effectiveness prediction model for initial screening. Automated A / B testing involves the module automatically creating A / B tests for materials passing the initial screening via the advertising platform's API. Budget allocation for the tests is intelligently managed by a multi-armed machine model, identifying winning materials with maximum efficiency and minimal cost. Feedback and optimization follow, with the module collecting detailed performance data from all test materials after the tests. This data is used to decide whether to scale up the winning materials, and more importantly, to provide feedback on the complete "diagnosis-strategy-result" learning case to the central scheduling and strategy selection module. This module utilizes these new data samples, through reinforcement learning or online learning, to continuously update its internal decision-making model, making future viral marketing strategy selection increasingly accurate and efficient.
[0068] Through the close collaboration of the above modules, the system proposed in this embodiment of the invention constructs a fully automated advertising creative fission ecosystem capable of self-perception, self-decision-making, self-execution, and self-optimization, thereby systematically solving the various problems mentioned in the background art.
[0069] Please see Figure 3 The third aspect of this application provides an AI-based automatic advertising material splitting device, including a memory and a processor connected in series. The memory stores a computer program, and the processor reads the computer program and executes the AI-based automatic advertising material splitting method as described in the first aspect of the application. Specifically, the memory may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, first-in-first-out (FIFO) memory, and / or last-in-first-out (FILO) memory, etc.; the processor may be, but is not limited to, microprocessors of the STM32F105 series, ARM (Advanced RISC Machines), x86 architecture processors, or processors with integrated NPU (neural-network processing units). The working process, working details, and technical effects of the device provided in the third aspect of this application can be found in the first aspect of the application, and will not be repeated here.
[0070] This fourth aspect of the embodiment provides a computer-readable storage medium storing instructions containing the AI-based automatic advertising material splitting method of the first aspect of the embodiment. Specifically, the computer-readable storage medium stores instructions that, when executed on a computer, perform the AI-based automatic advertising material splitting method of the first aspect. The computer-readable storage medium refers to a data storage carrier, which may include, but is not limited to, floppy disks, optical disks, hard disks, flash memory, USB flash drives, and / or Memory Sticks. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The working process, details, and technical effects of the computer-readable storage medium provided in this fourth aspect of the embodiment can be found in the first aspect of the embodiment, and will not be repeated here.
[0071] The fifth aspect of this embodiment provides a computer program product containing instructions that, when executed on a computer, cause the computer to perform the AI-based automatic fission method for advertising materials as described in the first aspect of this embodiment. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0072] The embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0073] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. Based on this understanding, the technical solution of the present invention can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0074] Finally, it should be noted that the above are merely preferred embodiments of the invention and are not intended to limit the scope of protection of the invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the invention should be included within the scope of protection of the invention.
Claims
1. A method for automatic fission of advertising creatives based on artificial intelligence, characterized in that, The method includes: The application programming interface (API) monitors and aggregates creative material performance data at the creative material level from one or more advertising platforms in real time. When at least one performance indicator in the performance data meets a data-driven trigger condition that can be dynamically optimized based on historical campaign data and business objectives, the corresponding best-selling creative material is automatically identified. Multimodal feature extraction is performed on the viral creative material to generate a structured feature vector representing the visual, textual, and audio information corresponding to the viral creative material. The multimodal feature extraction steps include: extracting the main visual framework and color distribution features of the viral creative material using a convolutional neural network; extracting the spoken text using automatic speech recognition and extracting in-video subtitles using optical character recognition to jointly constitute text features; and analyzing the waveform of the background music using a beat detection algorithm to extract audio rhythm features. A central scheduling engine dynamically combines and selects at least one fission strategy from a strategy library containing multiple fission strategies, based on diagnostic analysis of the structured feature vectors. Based on the selected fission strategy combination, the corresponding generative artificial intelligence model or application programming interface is invoked to transform the viral creative material to generate at least one new derivative creative material; wherein, in the step of generating at least one new derivative creative material, when the selected fission strategy includes a scene replacement strategy, the steps include: extracting feature vectors from the target scene in the viral creative material, and in a pre-set material library, calculating the dot product of the feature vectors of the candidate scene and the target scene and normalizing them according to the magnitude of their respective vectors to obtain a cosine similarity that quantifies the degree of content and style similarity between the two in a multi-dimensional feature space, and selecting a set of candidate scenes with similarity higher than a preset similarity threshold; and using a pre-trained multimodal model to perform cross-modal semantic consistency verification between each candidate scene in the candidate scene set and the original context of the target scene, and selecting the candidate scene with the highest verification score for replacement; The new derivative creative materials are deployed to the advertising platform for testing, and the performance data of the new derivative creative materials is collected. The performance data is then fed back to the central scheduling engine to update the subsequent fission strategy selection logic.
2. The method according to claim 1, characterized in that, The performance metrics include at least one of click-through rate, conversion rate, return on investment (ROI), and cost per conversion.
3. The method according to claim 1, characterized in that, In the step of generating at least one new derivative creative material, when the selected fission strategy includes a dynamic narrative reconstruction strategy, it includes: The video content of the viral creative material is automatically cut into multiple shots using a lens boundary detection algorithm; A narrative graph is constructed by using the multiple storyboards as nodes and the temporal transitions between storyboards as edges; The narrative graph is processed using a graph neural network trained to predict a new sequence of nodes that maximizes the expected delivery effect, to generate the new sequence of nodes, and the storyboards are reassembled based on the new sequence of nodes.
4. The method according to claim 1, characterized in that, In the step of generating at least one new derivative creative material, when the selected fission strategy includes a pre-post expansion strategy, the following steps are included: The climactic moments in the viral creative materials are automatically detected using facial expression recognition APIs or audio loudness analysis. Place the aforementioned climax scene at the starting point of the new derivative creative material; In addition, using a large language model, taking the climax segment as input, a prelude text that logically forms a reverse chronological relationship with the climax segment is generated, and a text-to-video model is invoked to generate or match corresponding video footage from a material library based on the prelude text, which is then spliced before the climax segment.
5. The method according to claim 1, characterized in that, In the step of generating at least one new derivative creative material, when the selected fission strategy includes a character replacement strategy, the following steps are included: The original spoken audio was extracted from the viral creative material using voiceprint separation. The system also calls the digital human generation application interface, takes the original spoken audio and the preset target digital human image as input, and generates a new video storyboard, wherein the new video storyboard includes the target digital human, and uses a lip-sync model to ensure that the lip movements of the target digital human match the original spoken audio.
6. The method according to claim 1, characterized in that, The step of testing the new derivative creative materials on an advertising platform is modeled as a multi-arm machine model, where each derivative creative material is treated as an independent arm. Its performance indicators during the testing period serve as reward signals; And through a strategy designed to balance exploration and utilization, with the optimization goal of maximizing cumulative rewards or minimizing cumulative regrets, the testing budget for each arm is dynamically allocated to quickly determine the optimal derivative creative material within a preset testing period.
7. The method according to claim 6, characterized in that, Before the step of deploying the new derivative creative materials to the advertising platform for testing, the following steps are also included: The new derivative creative materials are submitted to a compliance filter layer for automated review. The compliance filter layer compares them with a copyright database and an advertising regulations knowledge base, and automatically blocks derivative creative materials that do not meet the preset compliance standards.
8. An AI-based automatic advertising creative generation system, characterized in that, include: The data monitoring and triggering module is used to monitor and aggregate creative material granular performance data from one or more advertising platforms in real time, and automatically identify the corresponding best-selling creative material when at least one performance indicator in the performance data meets a data-driven triggering condition that can be dynamically optimized based on historical delivery data and business objectives. A multimodal feature extraction module is used to extract multimodal features from the viral creative material to generate a structured feature vector containing visual, textual, and audio information. The multimodal feature extraction module includes: a unit for extracting the main visual framework and color distribution features of the viral creative material using a convolutional neural network; a unit for extracting spoken text using automatic speech recognition and extracting in-video subtitles using optical character recognition to jointly constitute text features; and a unit for analyzing the waveform of background music using a beat detection algorithm to extract audio rhythm features. The central scheduling and strategy selection module is used to dynamically combine and select at least one fission strategy from a strategy library containing multiple fission strategies based on the diagnostic analysis of the structured feature vectors. The strategy execution and generation module is used to transform the viral creative material by calling the corresponding generative artificial intelligence model or application interface according to the selected fission strategy combination, so as to generate at least one new derivative creative material. The strategy execution and generation module includes a cross-modal verification unit. This cross-modal verification unit is used to extract feature vectors from the target scene in the viral creative material when executing the scene replacement strategy. In a pre-set material library, it calculates the dot product of the feature vectors of the candidate scene and the target scene and normalizes them according to the magnitude of their respective vectors to obtain a cosine similarity that quantifies the similarity of their content and style in a multi-dimensional feature space. It then selects a set of candidate scenes with similarity higher than a preset similarity threshold. Furthermore, it calls a pre-trained multimodal model to perform cross-modal semantic consistency verification between the original context of each candidate scene in the candidate scene set and the target scene, and selects the candidate scene with the highest verification score for replacement. The system also includes a closed-loop testing and optimization module, which is used to deploy the new derivative creative materials to the advertising platform for testing, collect the performance data of the new derivative creative materials, and feed the performance data back to the central scheduling and strategy selection module to update the subsequent fission strategy selection logic.
Citation Information
Patent Citations
Advertisement material production method and device, storage medium and computer equipment
CN117676048A
Video generation method and device based on AIGC and storage medium
CN120238693A
Intelligent advertisement material automatic optimization method, system and device and readable storage medium
CN120563172A