A marketing video generation method based on user portrait

By acquiring user behavior sequences and session context feature vectors, and using hidden Markov models and policy effectiveness evaluation models to generate marketing videos, this technology solves the problem of marketing videos being out of touch with user needs in existing technologies, achieving high relevance and credibility of content, and possessing self-learning optimization capabilities.

CN122317366APending Publication Date: 2026-06-30TONGLING UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TONGLING UNIV
Filing Date
2026-03-13
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing marketing video generation technologies fail to capture the dynamic psychological changes of users during the purchase decision-making process, resulting in a disconnect between the pushed content and users' real needs, and a lack of effective guarantees for the credibility and verifiability of the content.

Method used

By acquiring user behavior sequences and session context feature vectors, a hidden Markov model is used to identify the user's psychological decision-making stage. A strategy effectiveness evaluation model is constructed to quantify the effectiveness of marketing intervention strategies, a structured narrative instruction tree is generated, and a marketing video is synthesized through a conditional diffusion video generation model. A dual-channel feedback mechanism is integrated for self-calibration.

Benefits of technology

It achieves a close alignment between the content of marketing videos and users' psychological needs, enhances the relevance and persuasiveness of the content, ensures the credibility and transparency of the video content, and has self-learning capabilities to continuously optimize marketing effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122317366A_ABST
    Figure CN122317366A_ABST
Patent Text Reader

Abstract

This invention relates to the field of digital marketing technology, specifically disclosing a marketing video generation method based on user profiles. The method includes: acquiring the target user's behavioral sequence and conversation context feature vector over the past seven days; identifying the user's current psychological decision-making stage based on the behavioral sequence using a pre-trained Hidden Markov Model; quantifying the effectiveness of multiple marketing intervention strategies based on the psychological decision-making stage and the context features using a strategy effectiveness evaluation model; generating an intervention priority ranking vector and corresponding predicted conversion rates; and establishing a dual-channel feedback and self-calibration method to not only collect user behavioral feedback but also actively collect reference data from AI search engines, forming a comprehensive effect evaluation closed loop. By comparing actual conversion results with predicted conversion rates, the parameters of the strategy effectiveness evaluation model are continuously optimized, exhibiting long-term self-learning and evolution capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital marketing technology, specifically to a method for generating marketing videos based on user profiles. Background Technology

[0002] With the rapid development of the internet and e-commerce, personalized marketing has become a key means to improve user conversion rates. Currently, the industry generally uses recommendation systems based on user profiles to push advertising content.

[0003] Existing marketing video generation technologies rely solely on static user tags for coarse-grained content matching, failing to capture the dynamic psychological changes of users during the purchase decision-making process. This results in content being disconnected from users' actual needs. Furthermore, existing marketing video content generation is typically template-based, lacking quantitative evaluation of the combined effects of different marketing strategies, making it difficult to achieve precise strategy optimization. In addition, in the era of artificial intelligence, content not only needs to be user-oriented but also needs to meet the crawling and understanding requirements of AI search engines. However, existing technologies generally lack effective methods to ensure the credibility and verifiability of content.

[0004] To address the aforementioned shortcomings, a marketing video generation method based on user profiles is proposed. Summary of the Invention

[0005] In order to overcome the shortcomings of the prior art, the present invention provides a marketing video generation method based on user profiles to solve the problems mentioned in the background.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a marketing video generation method based on user profiles, comprising the following steps:

[0007] Step 1: Obtain the target user's behavior sequence and session context feature vector over the past seven days;

[0008] Step 2: Based on the behavioral sequence, identify the user's current psychological decision-making stage using a pre-trained Hidden Markov Model;

[0009] Step 3: Based on the psychological decision-making stage and the context feature vector, quantify the effectiveness of multiple marketing intervention strategies through a strategy effectiveness evaluation model, and generate an intervention priority ranking vector and the corresponding predicted conversion rate;

[0010] Step 4: Input the psychological decision-making stage and the intervention priority ranking vector into the intelligent script generation module to generate a structured narrative instruction tree. The nodes of the narrative instruction tree include content type, emotional tag and semantic instruction.

[0011] Step 5: Based on the narrative instruction tree, call the conditional diffusion video generation model to synthesize the final marketing video, wherein the conditional diffusion video generation model injects the semantic instructions, sentiment tags, and product attribute embedding vectors obtained by encoding product selling points into the nodes during the denoising process.

[0012] Step 6: Collect user behavior feedback on the final marketing video through front-end tracking points, and collect reference data of the video from the AI ​​search engine through the API interface to form a dual-channel feedback.

[0013] Step 7: Based on the comparison between the actual conversion results in the dual-channel feedback and the predicted conversion rate, update the parameters of the strategy effect evaluation model to achieve model self-calibration.

[0014] The beneficial effects of this invention are:

[0015] This invention uses a hidden Markov model to deeply analyze users' recent behavioral sequences, accurately identifying the user's current dynamic psychological decision-making stage. Compared with traditional user profiles based on static tags, the method of this invention enables the generation of marketing videos to closely match the user's real-time psychological needs in the decision-making process, significantly improving the relevance and persuasiveness of the content.

[0016] This invention, through the constructed strategy effectiveness evaluation model, can not only independently evaluate the effectiveness of different marketing intervention strategies such as price sensitivity, social proof, and limited-time offers, but also quantify their combined synergistic effect. Based on this evaluation result, the generated intervention priority ranking vector provides strategic guidance for subsequent video script generation, ensuring that the presentation weight and combination logic of each marketing element in the video content are optimal, effectively solving the problems of blind strategy application and poor results in existing technologies.

[0017] This invention uses a structured narrative instruction tree as a bridge connecting strategy evaluation and video synthesis, and drives a conditional diffusion video generation model to create multimodal content. It not only injects product attribute information, but also follows the emotional tags and semantic instructions determined by the strategy, ensuring a high degree of consistency and emotional resonance in the generated video at the visual, text, and voice levels. At the same time, by synchronously generating structured content credentials, it enhances the credibility and transparency of marketing content, enabling it to effectively reach users and be accurately understood and referenced by AI search engines.

[0018] This invention establishes a dual-channel feedback and self-calibration method, which not only collects user behavior feedback but also actively gathers reference data from AI search engines, forming a more comprehensive closed loop for effect evaluation. It compares actual conversion results with predicted conversion rates, continuously optimizes the parameters of the strategy effect evaluation model, and possesses long-term self-learning and evolution capabilities to constantly adapt to changes in market and user preferences, ensuring continuous improvement in marketing effectiveness. Attached Figure Description

[0019] The invention will now be further described with reference to the accompanying drawings.

[0020] Figure 1 A flowchart illustrating a marketing video generation method based on user profiles, provided as an embodiment of the present invention;

[0021] Figure 2 A flowchart illustrating the execution steps of the joint intervention effect evaluation module provided in this embodiment of the invention;

[0022] Figure 3 A flowchart illustrating the final marketing video synthesis process provided in this embodiment of the invention;

[0023] Figure 4 This is a closed-loop feedback architecture diagram of a marketing video generation method based on user profiles provided in an embodiment of the present invention. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] Please see Figures 1-4 As shown, this invention is a marketing video generation method based on user profiles, specifically including the following steps:

[0026] Step 1: Collect and construct user behavior and context feature vectors, including: obtaining all interaction records of the target user in the past seven days from the internal log database, cleaning and standardizing them to obtain an ordered behavior sequence; at the same time, collecting session-related context information and encoding it into a 16-dimensional context feature vector for confounding variable control in subsequent strategy effect evaluation modeling.

[0027] Furthermore, this includes: obtaining all interaction records generated by the target user in the past seven days from the user behavior log database. The user behavior log database is an internally deployed distributed event recording system that collects and stores user behavior events in real time through front-end tracking tools. The primary key is the user ID, and each interaction record includes a timestamp, behavior type, and product identifier, etc.

[0028] The acquired interaction records, after being cleaned and standardized, form an ordered sequence of behaviors, denoted as... ,in, This represents the timestamp of the i-th action event; Indicates the behavior type of the i-th behavior event; The identifier represents the product identifier associated with the i-th behavioral event; N is the length of the behavioral sequence.

[0029] It should be noted that the cleaning process for all interaction records mentioned above includes removing invalid events, standardizing behavior naming, and filling in time breakpoints; the standardization process includes time offset normalization and sequence reordering.

[0030] It needs to be further explained that, The value is an element from a predefined set of behavior categories, such as browsing, adding to cart, or exiting. Behavior types include, but are not limited to: accessing a product details page, adding a product to the cart, exiting the shopping process, and viewing other users' reviews. The product identifier is an internally unique string or integer used to associate with a specific product, facilitating personalized analysis and content generation based on product attributes. If a behavior event does not involve a specific product, then... Set to an empty value or the default placeholder.

[0031] Simultaneously, contextual information related to the target user's session is collected, including: terminal device model, network type, whether it is in a promotional period, city tier, and current time period. This contextual information is automatically collected by the front-end SDK and stored in the log database along with corresponding behavioral events; the contextual information is encoded into a 16-dimensional real vector. Let be denoted as the context feature vector.

[0032] It should be noted that the specific encoding method for the aforementioned contextual information is as follows: 5-dimensional device model, 3-dimensional network type, 1-dimensional promotion status, 1-dimensional city level, 3-dimensional time period, and 3-dimensional reserved dimension. After encoding, all feature vectors are normalized to the [0, 1] interval to ensure consistency in the dimensions of different features. This contextual feature vector serves as a confounding variable in the strategy effectiveness evaluation model, helping to distinguish whether changes in user behavior are driven by internal psychology or by the external environment. For example, an increase in users adding items to their cart during a promotion may be due to increased price sensitivity or the influence of the promotional atmosphere. The contextual feature vector effectively controls such confounding effects, thereby improving the accuracy of strategy effectiveness evaluation inferences.

[0033] Step 2: Identify the user's current psychological decision-making stage; analyze the user's behavioral sequence using a pre-trained Hidden Markov Model; the Hidden Markov Model treats the user's explicit behavior as observable symbols and the internal psychological state as hidden variables; infer the user's most likely current psychological state through the Viterbi decoding algorithm and encode it into a fixed-dimensional feature vector of the current decision-making environment.

[0034] Furthermore, this includes the following: constructing a user decision-making stage identification module to characterize the evolution of a user's psychological state during the purchase process, and defining a discrete set of psychological states using the user decision-making stage identification module. These correspond to four decision-making stages: interest stimulation, functional doubt, price hesitation, and trust.

[0035] Hidden Markov Models (HMMs) are used as state inference tools, where the explicit behavior of the target user is regarded as observable symbols, and a set of observation symbols is defined. The target user's internal psychological state is a hidden variable; the state transition probability matrix A and the observation emission probability matrix B of the hidden Markov model are both obtained offline by training the entire historical user data using the Baum-Welch algorithm.

[0036] Specifically, the state space of the Hidden Markov Model is 4, corresponding to four psychological states; the six behavior types in the observation symbol set are browsing product details page, adding products to shopping cart, exiting the shopping process, viewing other users' reviews, clicking on promotional information, and viewing product prices. These behaviors are all standard event types collected by front-end tracking points, and each behavior type corresponds to a unique event ID, which is used for subsequent state inference and strategy effect evaluation modeling.

[0037] It should be noted that the state transition probability matrix A contains elements This indicates the target user's psychological state. Shift to psychological state The probability of, where and The mental states represent the current state and the next state, respectively, and their values ​​are taken from the set of mental states. The observed emission probability matrix B contains the following elements. Indicates psychological state Observed behavior The probability of, where For the k-th user behavior type, the value is taken from the observation symbol set. .

[0038] Furthermore, the Viterbi decoding algorithm is invoked to process the behavior sequence. By performing optimal state path inference, we can obtain the psychological state that the target user is currently in with the highest probability. and its most recent state transition path .

[0039] It should be noted that the Viterbi decoding algorithm is used for behavioral sequences. The time complexity of performing the optimal state path search is O(n log n). Where N is the length of the action sequence, =4, Viterbi decoding algorithm outputs the current mental state. and its most recent state transition path This is used for subsequent decision-making and intervention calculations.

[0040] Current psychological state The input is fed into a pre-trained state encoder, which outputs a fixed-dimensional feature vector of the current decision-making environment. The state encoder is a single-layer fully connected neural network. The weights of the fully connected neural network are obtained through random initialization and gradient descent optimization during the training phase to ensure that different states are mapped to different regions in the semantic space. The feature vector of the current decision environment serves as one of the key inputs to the adaptive narrative generation engine, representing the semantic embedding of the target user's current decision context.

[0041] Step 3: Evaluate the effectiveness of the multi-dimensional intervention strategy and predict the conversion rate; Based on the identified psychological state, define three interventionizable processing variables, including: price sensitivity, social proof demand intensity, and perceived intensity of limited-time offers, and quantify their values ​​through user historical behavior; Then, introduce a strategy effect evaluation module based on Bayesian network to calculate the independent and joint intervention effects of each processing variable in the current state, and generate a three-dimensional intervention priority ranking vector according to effectiveness, while outputting the predicted conversion rate.

[0042] Specifically, this includes the following: the current mental state based on the output of the Viterbi decoding algorithm in step 2. Define price sensitivity Intensity of social proof requirements Perceived strength of limited-time offers Its value is quantified based on the user's historical behavioral characteristics and its effectiveness is only evaluated under the current psychological state; price sensitivity This measure reflects users' level of attention to product pricing, calculated by weighting behavioral metrics such as browsing time and add-to-cart rate across different price ranges; social proof demand intensity. This reflects the degree to which users rely on others' reviews or sales data, and is calculated by weighting behavioral indicators such as the number of reviews viewed and clicks on best-selling lists; perceived strength of limited-time offers. It is used to reflect the degree of user responsiveness to the timeliness of promotions, and is calculated by weighting indicators such as the time spent on the countdown page and the refresh frequency.

[0043] It should be noted that the aforementioned user historical behavior characteristics refer to a structured data set formed by various operations performed by users on the platform over a past period. After extraction, encoding, and standardization, this data is used to characterize users' consumption habits, interests, decision-making patterns, and other intrinsic attributes. These user historical behavior characteristics are stored in a user behavior log database; specifically, price sensitivity... The strength of social proof demand is calculated by weighting users' browsing time, add-to-cart frequency, and unsubscribe rate for products in different price ranges over the past 30 days. The perceived strength of the limited-time offer is calculated based on a combination of the number of times reviews were viewed and the frequency of clicking the "10,000 people have already purchased" tag. The evaluation criteria are determined based on the duration of page dwell and the frequency of page refreshes during the promotional period; all features are standardized before being input into the strategy effectiveness evaluation model to ensure consistency in cross-user comparisons.

[0044] Furthermore, to overcome the limitations of traditional univariate factor analysis, a joint intervention effect evaluation module was set up. This module is based on a Bayesian network-based strategy effect evaluation inference model, used to evaluate the independent and joint effects of multiple intervention variables under a specific psychological state. Specifically, when the current psychological state is input... After considering user historical behavioral characteristics, the joint intervention effectiveness evaluation module performs the following steps:

[0045] Step a1: Use a Bayesian network to learn the dependencies between variables from the full set of user data;

[0046] Step a2: Calculate the policy interaction coefficients among the treatment variables. This is to reflect whether there is a positive or negative synergistic effect;

[0047] Step a3: For each processing variable When calculating the intervention effect, it should be noted that when calculating the treatment variable... When considering the intervention effect, its own influence is taken into account, and it is also included in other treatment variables. The interaction terms; where the formula for calculating the intervention effect is as follows:

[0048] Where Y is a binary outcome variable, where 1 indicates a completed purchase and 0 indicates a non-complete purchase, and Z is a set of confounding variables, including age, historical category preference, city tier, etc. To handle variables and The policy interaction coefficients between them are learned through a Bayesian network from the full set of user data. Specifically... This indicates the synergistic effect between the intensity of positive social demand and the perceived intensity of limited-time offers. If its value is greater than 0, it indicates that the combined intervention of the two is more effective than the sum of their individual effects.

[0049] Step a4: Normalize the effects of each intervention and sort them from highest to lowest effectiveness to generate a three-dimensional intervention priority ranking vector. .

[0050] Step a5: Ranking vector based on intervention priority Construct a transformation prediction graph and estimate the conditional probability using a Bayesian network: Y1 represents the user conversion time, such as click or purchase. This indicates that the intervention strategy is manually set; the specific method for obtaining this conditional probability is: sorting the intervention priority vector. As input, it is substituted into the pre-trained policy effectiveness evaluation model, and the output is the predicted conversion rate. .

[0051] It should be noted that the joint intervention effect evaluation module is built based on the Bayesian network structure learning algorithm. The input is the user behavior feature vector and psychological state label, and the output is the independent effect of each treatment variable and its policy interaction coefficient with other variables. The joint intervention effect evaluation module is deployed on the server in software form. All parameters are trained offline on millions of user sessions and are updated regularly to adapt to market changes.

[0052] For each processing variable Calculate the effect of each individual treatment variable on the current mental state. Intervention effect Simultaneously, the combined effect of individual treatment variables with other variables is calculated; specifically, in terms of psychological state... Below, the intensity of demand for social proof alone The conversion rate was 12%, and the perceived strength of the limited-time offer when used alone was... The conversion rate was 14%, but using both methods simultaneously increased the conversion rate to 28%. and The combined intervention effect was significantly higher than the sum of the individual effects of the two, indicating a positive synergistic effect.

[0053] Furthermore, after the above calculations are completed, the system normalizes the intervention effect values ​​of the three treatment variables using a min-max normalization method, converting them into weights within the range of [0, 1]. The specific formula is as follows: ;in, and These represent the maximum and minimum values ​​of the intervention effects of the three treatment variables, respectively. Based on the calculated intervention effectiveness weights of each treatment variable, a three-dimensional intervention priority ranking vector p is generated. Specifically, if the weights of the treatment variables are arranged from highest to lowest effectiveness as p2, p3, and p1, then the three-dimensional intervention priority ranking vector is... Where p1 represents the intervention effectiveness weight of price sensitivity, p2 represents the intervention effectiveness weight of social proof demand intensity, and p3 represents the intervention effectiveness weight of limited-time offer perception intensity. The higher the value, the more effective the intervention factor is in the current state. The intervention priority ranking vector, as the strategy control signal of the adaptive narrative generation engine, directly determines the presentation weight and combination logic of each marketing element in the video.

[0054] It should be noted that the pre-trained strategy effectiveness evaluation model is a multivariate intervention response model built on a Bayesian network, used to quantify the impact of different marketing strategies on user conversion behavior. The model ranks vectors by intervention priority. As input, the output is the predicted conversion rate. This refers to the probability that a user will perform a conversion behavior such as clicking, adding to cart, or making a purchase under a given intervention strategy. The strategy effectiveness evaluation model is pre-trained offline using historical advertising data, which includes the strategy configuration used for each advertisement and its actual conversion results. By constructing a prediction graph of the transformation between variables, for example , and based on and Interaction edges between them are used to capture synergistic effects. Model parameters are learned through maximum likelihood estimation or variational inference methods and optimized using the mean squared error loss function. After training, the strategy effect evaluation model is deployed in the inference stage to quickly respond to strategy combinations generated by different narrative instruction trees and provide conversion effect predictions, thereby supporting subsequent video generation optimization and strategy iteration.

[0055] Step 4: Generate a structured narrative instruction tree; invoke a Transformer-based intelligent script generation module, receiving the current decision-making environment feature vector and intervention priority ranking vector as input; the intelligent script generation module recursively generates a hierarchical narrative instruction tree, where each node precisely specifies the type, sentiment tag, and specific semantic instruction of the video content unit. Further, this includes: invoking the intelligent script generation module, which consists of a Transformer-based sequence-to-tree decoder; the sequence-to-tree decoder receives the current decision-making environment feature vector. and intervention priority ranking vector Furthermore, a hierarchical narrative instruction tree is generated layer by layer through a recursive expansion strategy. The root node of the narrative instruction tree represents the overall narrative type, while the child nodes correspond to specific video content units.

[0056] Specifically, the narrative instruction tree is a rooted, ordered, attributed multi-branch tree, defined as follows: Root node: represents the overall narrative type, with no parent node; Internal nodes: have multiple child nodes, each corresponding to a video content unit; Leaf nodes: no longer split, representing the smallest executable content fragment; Each child node contains the following attribute fields: Content type, indicating whether the unit should use text, still image, dynamic video, or audio; Emotion tag, selected from a preset emotion dictionary, such as trust, urgency, and pleasure; Semantic instruction, an operation description in natural language form, such as showing over 100,000 users to purchase this product or inserting a 30-second countdown animation.

[0057] Furthermore, the sequence-to-tree decoder is built on the Tree-Transformer architecture and includes: a parent node attention module for aggregating parent node information; a sibling node masking mechanism to ensure that child nodes are generated in order; and an output layer containing three classification heads that predict: node type; content attribute triples; and the number of child nodes, respectively.

[0058] Furthermore, the recursive expansion strategy is implemented through the following steps: Initialization: Initialize the feature vector of the current decision environment. and intervention priority ranking vector The concatenated vector serves as the initial context vector. Root node generation: The decoder first predicts the narrative type of the root node and outputs its hidden state. Child node recursively generated: for the current node The model predicts the number of its child nodes. ,like If it is a leaf node, then mark it as such and stop expanding; if Then generate sequentially Each child node has 10 child nodes, and the generation of each child node is based on the hidden state of the parent node. With global context Fusion representation; Termination condition: Maximum depth limit

[0059] The system is set to 3 layers; or the predicted number of child nodes is 0; or the cumulative number of generated nodes reaches the upper limit.

[0060] It should be noted that the narrative instruction tree serves to provide a precise and executable control blueprint for multimodal video synthesis, ensuring that the generated content strictly adheres to the intervention strategy; specifically, for high-priority... and A composite child node is generated, whose content type is dynamic image plus text, the emotion tag is trust plus urgency, and the semantic instruction is to overlay real-time sales bullet comments and red countdown box; this composite child node will be passed to the video compositing module to drive the synchronous generation of images, subtitles and sound effects.

[0061] It should also be noted that the sequence-to-tree decoder described above differs from traditional sequence generation models, such as the standard Transformer decoder. At each step, the sequence-to-tree decoder not only predicts the next word, but also predicts the number of child nodes and the branch type of the current node, thereby constructing a tree-shaped output structure that conforms to semantic logic. Its training process uses a tree-shaped cross-entropy loss function and performs end-to-end optimization on a large-scale labeled dataset.

[0062] Specifically, when input At this time, the root node type is a composite guided type. The first-level child nodes have two: Node A: content type = dynamic video, emotion = trust, instruction = play short videos with positive user reviews; Node B: content type = text + image, emotion = urgency, instruction = "display the red prompt box 'Only 3 items left'". Node A has no child nodes, while Node B has one child node C: content type = countdown animation, emotion = urgency, instruction = overlay a 60-second countdown. This ultimately forms a multi-branch tree with a depth of 3, which is used by the video compositing module to render node by node.

[0063] Step 5: Synthesize a multimodal personalized marketing video; Initiate the video generation process based on the narrative instruction tree: Allocate a total duration of 15 seconds according to the intervention priority weight associated with each node; Call the conditional diffusion video generation model, injecting the content type, sentiment tag, and product attribute embedding vectors of the nodes during the denoising process to generate a 720p video clip; Ensure the consistency of video style by adding a style smoothing loss term between adjacent frames; Constrain the semantic alignment between the narration text and the visuals through the CLIP model, and ensure the emotional consistency between TTS speech and text through a sentiment classifier; Finally, output a standard MP4 format video of 1920×1080, 30fps, and 15 seconds. Specifically, this includes the following: Synthesizing the final marketing video based on the conditional diffusion video generation model using the narrative instruction tree, with the following specific implementation steps:

[0064] Step b1: Traverse all child nodes in the tree. For the h-th child node, obtain its associated intervention priority weight. That is, the intervention priority ranking vector The video duration is allocated proportionally to the components of the corresponding processing variables; the total duration is fixed at fifteen seconds, and the duration allocated to each node is calculated according to the following formula: M represents the total number of child nodes. The intervention priority weight is the value corresponding to the h-th child node. =The sum of the intervention priority weights of all child nodes, and 15 is the preset total video duration.

[0065] It should be noted that the calculation formulas for the allocated time at each node are used to achieve dynamic allocation of video time based on intervention priority. This is significant because it allocates limited video time resources proportionally according to the effectiveness weight of the marketing elements associated with each sub-node in the current user's psychological state, ensuring that high-priority content receives a more substantial presentation opportunity. Specifically, the total video time is fixed at 15 seconds as a resource pool, and the appropriate time is calculated for each sub-node h in the narrative instruction tree. Among them, molecules This represents the priority weight of the intervention factor corresponding to this node; a higher value indicates that the factor is more effective. The denominator represents the priority weight of the intervention factor at this node. It is the sum of the weights of all child nodes, used for normalization to ensure that the total duration of each node equals 15 seconds; in this way, the system implements a strategy-driven, resource-optimized content presentation mechanism, for example, when social proof... When a segment is deemed the most effective, its corresponding video clip will receive approximately 6.3 seconds of display time, significantly higher than other lower-priority elements, thereby maximizing conversion rates.

[0066] Specifically, the mapping relationship between the child node and its associated processing variable is defined by preset rules, as follows: when the semantic instruction of the child node contains keywords such as sales volume, user reviews, and 10,000 people having purchased, the corresponding processing variable, social proof demand strength, is determined. and take When a child node contains keywords such as countdown, limited time, or last chance, determine the perceived strength of the limited-time offer in its corresponding processing variable. and take When a child node involves information such as price comparison, discount level, original price / current price, the price sensitivity of its corresponding processing variable is determined. and take The above mapping relationships are stored in a preset configuration table and are dynamically loaded at runtime.

[0067] Step b2: Construct a conditional diffusion video generation module. This module uses an improved VideoCrafter-v2 diffusion model as its core generator, injecting the following conditional signals into each layer of the denoising network: the content type and sentiment tag of the current node; and the product attribute embedding vector. .

[0068] Specifically, the VideoCrafter-v2 diffusion model is improved by adding two conditional input channels to the original model; setting up dual-path cross-attention modules in each layer of the denoising network to process local node information in the narrative instruction tree and product attribute embedding vectors respectively; finally, adding a time step awareness module to improve inter-frame coherence; the improved diffusion model is deployed on a GPU cluster, supports batch inference, and outputs short videos with a resolution of 720p and a frame rate of 30fps.

[0069] In each layer of the denoising network, the following two types of conditional information are injected through a cross-attention mechanism: 1. The content type and sentiment label of the current node. Content type: such as dynamic images and text, speech and subtitles, etc., to guide the generator to select appropriate visual modalities; sentiment label: such as trust, urgency, joy, etc., to control the emotional expression of the screen tone, font style, animation rhythm, etc. The above information is encoded into a discrete category vector, and converted into a continuous feature representation through the embedding layer, and then fed into the cross-attention module to affect the pixel distribution in the denoising process. 2. Product attribute embedding vector The product attribute embedding vectors are obtained by encoding the product selling point text using a Chinese BERT model. This Chinese BERT model is pre-trained on Chinese corpora and can effectively capture the semantics of Chinese phrases. The specific process is as follows:

[0070] Step c1: Extract structured descriptions from the product database, such as: contains hyaluronic acid, suitable for sensitive skin, alcohol-free, lightweight and non-sticky; Step c2: Concatenate the above phrases into a sentence, such as: This product contains hyaluronic acid, is suitable for sensitive skin, has no added alcohol, and has a lightweight and non-sticky texture; Step c3: Encode this sentence using a Chinese BERT model: Input: A sequence of words obtained after segmenting the product selling points text, with a special marker [CLS] added to the beginning of the sequence; Output: The hidden state vector corresponding to the [CLS] marker in the last layer of the model. This vector is the final product attribute embedding vector. This is used to guide content association and style control in the subsequent video generation process; step c4, embedding vectors for product attributes. Perform L2 normalization to ensure that its numerical range is controllable and avoid gradient explosion.

[0071] Among them, the embedding vector of product attributes The L2 normalization formula is: ,in, , Represents the embedding vector of product attributes The x-th component, x = 1, 2, ..., 768; each It is a real number representing the value of the corresponding dimension in the hidden state output by the Chinese BERT model; specifically, if =[0.5, 0.8, 1.2, ...], then =0.5, =0.8, and so on.

[0072] During each denoising layer, perform the following operations: acquire the current noisy image. and conditional signals ,in, This is a content-based embedding vector, obtained from the content type of the child nodes through the embedding layer; The sentiment tag embedding vector is obtained by encoding the semantic tag. Embed vectors for product attributes; As a query and key, As values, calculate attention weights Weighted aggregation yields a new feature map. The result is used to update the denoised image; where V represents the value vector from the noisy image. The characteristics are represented.

[0073] Step b3: To avoid abrupt style changes between different narrative stages, add a style smoothing term for adjacent frames to the training loss function: ,in This represents the t-th frame of the video. This is a pre-trained visual style encoder, specifically based on Gram matrix features extracted from a VGG network.

[0074] Specifically, visual style encoder This is a style feature extraction module built based on a pre-trained VGG-19 network. The specific operations are as follows: For each frame of video... Input to the 10th layer of VGG-19 and extract the activation map output from the 10th layer. Where H and W are the height and width of the 10th layer space, and C is the number of channels, calculate its Gram matrix. Where a, b=1, ..., C, represents the correlation between the a-th and b-th channels, reflecting the texture style. This represents the activation value of the a-th channel at the s-th spatial location; s = 1, 2, ..., H×W are the spatial location indices, arranged in row-major order; flatten the Gram matrix into a vector. That is, to flatten a C×C matrix into a length of [length missing] by row-major order. Vectors; define style smoothing loss: , That is to refer to This loss is jointly optimized with the main task during training, with weight coefficients set to... This is used to suppress style abrupt changes between adjacent frames; specifically, if frame t has a blue background and white text, and frame t+1 suddenly changes to a red background and yellow text, then... and The differences are significant. The output is relatively high; adjust the generation strategy to maintain color consistency.

[0075] Step b4: Text-visual semantic alignment loss, which minimizes the embedding distance between the narration and the image using the CLIP model; Voice-text sentiment consistency loss, which forces the sentiment scores of the TTS voice and the original text to be close using a sentiment classifier.

[0076] Specifically, the text-visual semantic alignment loss is achieved by encoding the image frames using the CLIP-ViT-B / 32 model. Encoding as visual embeddings Narration text Encoded as text embedding Among them, the narration text The text content of each child node in the narrative instruction tree is used as input text to generate speech for the TTS module. This text is the text input encoded by the CLIP model, used to achieve cross-modal semantic alignment; cosine similarity is calculated. Construct the contrastive loss function: ,in =0.07 represents the temperature parameter, Q represents the number of negative samples, and u represents the quantity index. The speech-text sentiment consistency loss is achieved by predicting the sentiment score of the original text using a pre-trained sentiment classifier, specifically a BERT-based sentiment analysis model. Emotional score of TTS voice Define MSE loss. The two losses together constitute the total alignment loss: Specifically, if the original text is "This product is mild and non-irritating," with the emotional label "trust" and a score of 0.8, then the TTS voice should generate a voice with a steady tone and gentle intonation, and its emotional score should also be close to 0.8.

[0077] Step b5: Output a standard MP4 video file with a resolution of 1920×1080, a frame rate of 30 frames per second, and a duration of 15 seconds. This video file is used to push to the user's terminal and serves as the object of subsequent feedback collection. After the user watches the video file in the mini-program, data such as the duration of their stay, click behavior, and number of likes are recorded to evaluate the advertising effect.

[0078] Step 6: Generate content credentials and deploy a dual-channel feedback mechanism; Simultaneously with video generation, the system automatically generates a structured content credential, including an EEAT statement, a list of fact anchors, and a semantic summary vector, which is stored and distributed along with the video; The system also deploys dual-channel feedback collection: first, it records user behavior through front-end event tracking; second, it obtains the frequency of citations of the video by the AI ​​search engine through the API interface.

[0079] Furthermore, step 6 includes the following: Simultaneously with video generation, the content trust service module is invoked to generate a structured content certificate. This content certificate is organized in machine-readable JSON format, including an EEAT statement, a list of fact anchors, and a semantic summary vector. The EEAT statement records evidence of the content's professionalism, experientiality, authority, and credibility. The list of fact anchors lists key factual claims in the video and their verifiable source links. The semantic summary vector is a 512-dimensional real-valued vector, compressed and represented by a pre-trained semantic encoder of the core video content.

[0080] The aforementioned content credentials file is stored along with the video on the content delivery network and embedded in the front-end webpage using structured data markup conforming to the Schema.org standard, allowing generative search engines to crawl it automatically. Simultaneously, a dual-channel feedback collection mechanism is deployed: the behavioral channel records the video completion rate, whether the user clicked the action button, and whether a purchase was ultimately completed through front-end tracking; the latter is recorded as the actual conversion result. AI Citation Channel: Through data interfaces established with mainstream generative search engines, the frequency of citations of the video in AI question-answering scenarios is obtained regularly.

[0081] Step 7: Achieve long-term self-calibration and evolution of the model; compare the actual observed user conversion results with the conversion rate predicted by the strategy effect evaluation module in Step 3, and construct a regression loss function; this regression loss function is used to update the parameters of the strategy effect evaluation model on a daily basis, enabling continuous learning and optimization, and constantly improving the accuracy and conversion effect of marketing videos.

[0082] Furthermore, step 7 includes the following: converting the actual results... Conversion rate predicted in step 3 By comparing the results, a regression loss function is constructed: Among them, model parameters Includes the state transition probability matrix and policy interaction coefficients of the Hidden Markov Model. The system employs a daily, timed Adam optimization algorithm to achieve long-term self-calibration and evolutionary capabilities. Specifically, the Adam optimization algorithm is used to perform a gradient descent update with a learning rate of 0.001 to obtain the optimized parameters. This update operation is performed once a day to ensure that the strategy effectiveness evaluation model closely approximates real user behavior patterns.

[0083] Specifically, the optimized parameters It includes an updated state transition probability matrix for the Hidden Markov Model (HMM). The state transition probability matrix defines the probability that a user will transition from one psychological state to another. It is continuously revised by training with real conversion data every day, so that it can more accurately reflect the evolution of the actual decision-making path of the target user group in the current market environment. Specifically, if the data shows that a large number of users directly enter the trust establishment, state non-state, or functional doubt stages after seeing content with high social proof, the HMM will correspondingly increase the transition probability from interest arousal to trust establishment.

[0084] Optimized parameters The system also includes an updated strategy interaction coefficient, which quantifies the strength of synergistic or antagonistic effects between different marketing intervention strategies. By comparing predicted conversion rates with actual conversion results, it can learn which strategy combinations are truly effective in specific psychological states. Specifically, if the strategy effectiveness evaluation inference model finds that when users are in a price-hesitant state, the combined effect of limited-time offers and social proof far exceeds the sum of their individual effects, then the corresponding strategy interaction coefficient... This will be adjusted upwards, making it more likely that this highly synergistic combination will be recommended in future strategy evaluations.

[0085] By optimizing the state transition probability matrix and policy interaction coefficients of the Hidden Markov Model as described above, we can dynamically adjust its judgment logic on users' psychological states and the evaluation criteria for the effectiveness of marketing strategies, ensuring that the generated marketing videos always maintain a high degree of consistency with users' real behavioral patterns and psychological needs.

[0086] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0087] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0088] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0089] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0090] Finally, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for generating a marketing video based on a user profile, the method comprising: receiving a user profile; receiving a video; and generating a marketing video based on the user profile and the video. Includes the following steps: Step 1: Obtain the target user's behavior sequence and session context feature vector; Step 2: Based on the behavioral sequence, identify the user's current psychological decision-making stage using a pre-trained Hidden Markov Model; Step 3: Based on the psychological decision-making stage and the context feature vector, quantify the effectiveness of multiple marketing intervention strategies through a strategy effectiveness evaluation model, and generate an intervention priority ranking vector and the corresponding predicted conversion rate; Step 4: Input the psychological decision-making stage and the intervention priority ranking vector into the intelligent script generation module to generate a structured narrative instruction tree. The nodes of the narrative instruction tree include content type, emotional tag and semantic instruction. Step 5: Based on the narrative instruction tree, call the conditional diffusion video generation model to synthesize the final marketing video. In the denoising process, the conditional diffusion video generation model injects semantic instructions, sentiment tags, and product attribute embedding vectors obtained from product selling point text encoding of the nodes. Step 6: Collect behavioral feedback from target users on the final marketing video, and collect reference data of the final marketing video from the AI ​​search engine through the API interface to form a dual-channel feedback. Step 7: Based on the comparison between the actual conversion results in the dual-channel feedback and the predicted conversion rate, update the parameters of the strategy effect evaluation model. 2.The marketing video generation method based on user portrait according to claim 1, characterized in that, Step 1 specifically includes: The system retrieves all interaction records generated by the target user within seven days from the user behavior log database, collects and stores user behavior events in real time using front-end tracking tools, and then cleans and standardizes all the obtained interaction records to form an ordered sequence of behaviors.

3. The marketing video generation method based on user profiles according to claim 1, characterized in that, Step 2 specifically includes: A user decision-making stage identification module is constructed, defining a discrete set of psychological states, including interest arousal, functional doubt, price hesitation, and trust establishment. A Hidden Markov Model (HMM) is used as the state inference tool, with the target user's explicit behavior treated as observable symbols, defining a set of observable symbols, and the target user's internal psychological state as hidden variables. The Viterbi decoding algorithm is used to infer the optimal state path from the behavioral sequence, obtaining the target user's current psychological state with the highest probability and its most recent state transition path. The current psychological state is input into a pre-trained state encoder, which outputs a fixed-dimensional feature vector of the current decision-making environment.

4. The marketing video generation method based on user profiles according to claim 1, characterized in that, Step 3 specifically includes: based on the current psychological state output by the Viterbi decoding algorithm in step 2, defining three interventionizable processing variables, including price sensitivity, intensity of social proof need, and intensity of perceived limited-time offer; and evaluating the independent and joint effects of multiple intervention variables under a specific psychological state through a joint intervention effect evaluation module.

5. The marketing video generation method based on user profiles according to claim 4, characterized in that, After the current psychological state and the user's historical behavioral characteristics are input, the joint intervention effect evaluation module performs the following steps: Step a1: Use a Bayesian network to learn the dependencies between variables from the full set of user data; Step a2: Calculate the policy interaction coefficients between each processing variable; Step a3: Calculate the intervention effect for each treatment variable; Step a4: Normalize the effects of each intervention and sort them from high to low effectiveness to generate a three-dimensional intervention priority ranking vector; Step a5: Construct a conversion prediction map based on the intervention priority ranking vector, and estimate the conditional probability through a Bayesian network. The specific method for obtaining the conditional probability is as follows: take the intervention priority ranking vector as input, substitute it into the pre-trained strategy effect evaluation model, and output the predicted conversion rate.

6. The marketing video generation method based on user profiles according to claim 1, characterized in that, Step 4 specifically includes: calling the intelligent script generation module, which consists of a sequence-to-tree decoder based on the Transformer architecture; the sequence-to-tree decoder receives the current decision environment feature vector and the intervention priority sorting vector, and generates a narrative instruction tree with a hierarchical structure layer by layer through a recursive expansion strategy. The root node of the narrative instruction tree represents the overall narrative type, and the child nodes correspond to specific video content units.

7. The marketing video generation method based on user profiles according to claim 1, characterized in that, Step 5 specifically includes: synthesizing the final marketing video based on the conditional diffusion video generation model invoked according to the narrative instruction tree. The specific implementation steps are as follows: Step b1: Traverse all child nodes in the narrative instruction tree. For the h-th child node, obtain its associated intervention priority weight. That is, the intervention priority ranking vector The corresponding components of the processing variables are allocated video duration proportionally; the total duration is fixed at fifteen seconds, and the duration allocated to each node is calculated according to the formula: M represents the total number of child nodes. The intervention priority weight is the value corresponding to the h-th child node. The sum of the intervention priority weights of all child nodes, and 15 is the preset total video duration; Step b2: Construct a conditional diffusion video generation module. This module uses an improved VideoCrafter-v2 diffusion model as its core generator, injecting the following conditional signals into each layer of the denoising network: the content type and sentiment tag of the current node; and the product attribute embedding vector. ; Step b3: Add a style smoothing term for adjacent frames to the training loss function: ,in This represents the t-th frame of the video. For a pre-trained visual style encoder; Step b4: Text-visual semantic alignment loss and speech-text sentiment consistency loss; Step b5: Output the final marketing video.

8. The marketing video generation method based on user profiles according to claim 1, characterized in that, Step 6 specifically includes: simultaneously generating the final marketing video, invoking the content trust service module to generate a structured content certificate, which includes an EEAT statement, a list of fact anchors, and a semantic summary vector.

9. The marketing video generation method based on user profiles according to claim 1, characterized in that, Step 7 specifically includes: comparing the actual conversion results with the conversion rate predicted in step 3, constructing a regression loss function; and updating it through a daily scheduled Adam optimization algorithm.