A method for constructing a computational advertising creative intelligent agent and related equipment

CN121724693BActive Publication Date: 2026-08-11SOUTH CHINA UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]本申请实施例的主要目的在于提出一种基于场景语义驱动与数字孪生交互的计算广告创意智能体构造方法、系统、电子设备、存储介质及程序产品,以解决现有广告创意生成技术依赖人工、效率低下、个性化不足、缺乏对顾客交互场景深度感知的问题

Benefits of technology

[0019] The embodiments of this application include at least the following beneficial effects: This application provides a method, system, electronic device, storage medium, and program product for constructing a computational advertising creative intelligent agent. This solution, through scene semantic driving and digital twin interaction, combined with RPA automation and generative AI technology, achieves intelligent, automated, and personalized advertising creative generation, significantly improving the efficiency of advertising creative generation and customer interaction experience. Compared with traditional methods that rely on human experience or simple programmatic creatives, this application can more accurately understand customer intent, scene semantics, and interaction needs, thereby generating advertising content that better meets customer needs and improving advertising effectiveness and customer satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121724693B_ABST
    Figure CN121724693B_ABST
Patent Text Reader

Abstract

This application provides a method and related equipment for constructing a computational advertising creative intelligent agent, belonging to the interdisciplinary field of artificial intelligence and computational advertising. The method includes: real-time perception of customer behavior, environment, and intent through multimodal data to construct a dynamically updated scene semantic digital twin; based on the difference between the digital twin feedback and the advertiser's expectations, using RPA technology to decompose the creative generation task into sub-tasks; employing a learning-decision-modeling-prediction (LDMP) model combined with retrieval-enhanced generation (RAG) technology to retrieve materials from a creative knowledge base, and performing effect prediction and iterative optimization of the generated new creative in a digital twin simulation environment until a personalized advertising creative that meets expectations is output. This application, through the deep integration of digital twin, RPA, and generative AI, achieves a high degree of intelligence, automation, and personalization in advertising creative generation, significantly improving advertising effectiveness and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the interdisciplinary field of artificial intelligence and computational advertising, and in particular to a method for constructing a computational advertising creative intelligent agent and related equipment. Background Technology

[0002] With the widespread adoption of the internet and mobile devices, the advertising industry is undergoing a profound transformation from traditional one-way communication to intelligent, programmatic advertising. Traditional advertising creative production processes rely heavily on human experience, with fixed steps and lengthy cycles, making it difficult to meet today's market demands for agile, personalized, and iterative advertising content.

[0003] In recent years, the development of programmatic advertising and generative artificial intelligence (such as Large Language Modeling, LLM) technologies has provided new possibilities for the automation of advertising creative. However, existing programmatic creative optimization solutions mostly focus on simple element combinations (such as A / B testing of images and copy), lacking an understanding of customers' deep intentions, behavioral states, and environmental contexts in specific scenarios. This results in a gap between the generated advertising creatives and customers' real needs, making it difficult to achieve accurate personalized matching, thus limiting further improvements in advertising conversion efficiency and user experience. Summary of the Invention

[0004] The main objective of this application is to propose a method, system, electronic device, storage medium, and program product for constructing a computational advertising creative intelligent agent based on scene semantics-driven and digital twin interaction, in order to solve the problems of existing advertising creative generation technology relying on manual labor, low efficiency, lack of personalization, and lack of deep perception of customer interaction scenarios.

[0005] To achieve the above objectives, one aspect of this application proposes a method for constructing a computational advertising creative agent, the method comprising: Scene semantic perception steps: Real-time collection and analysis of customer behavior characteristics, environmental information and interaction intentions through multimodal data, to construct and dynamically update a digital twin reflecting the customer's real-time state and scene context; Task decomposition steps: Based on the difference between the actual customer feedback reflected by the digital twin and the advertiser's expected feedback, the advertising creative generation task is decomposed into a series of executable sub-tasks using Robotic Process Automation (RPA) technology; Simulation optimization steps: Based on the Learning, Decision, Modeling and Prediction LDMP model, according to the new advertising creative design principles corresponding to the sub-tasks, new advertising creatives are retrieved, generated and simulated in the advertising creative knowledge base until the customer feedback simulated by the digital twin meets the advertiser's expectations; Creative generation steps: Output the advertising creatives that have been simulated and optimized by the LDMP model.

[0006] In some embodiments, the idea generation step includes: The optimized advertising creative is presented to customers, and customer feedback data is collected. The feedback data is then used to update the digital twin and optimize the LDMP model, forming a closed-loop optimization process.

[0007] In some embodiments, the construction of the digital twin includes the construction of a semantic layer, a geometric layer, a physical layer, and a behavioral layer, specifically achieved through multi-source sensor data fusion, 3D modeling and driving, and LDMP behavioral simulation.

[0008] In some embodiments, the scene semantic awareness step, which involves constructing and dynamically updating the digital twin, specifically includes: Multimodal data is collected through visual sensors, voice sensors, environmental sensors, and intent recognition modules; Based on the multimodal data, a semantic layer of the digital twin is constructed through semantic encoding, and structured semantic tags are output. Based on RGB-D camera point cloud data, a 3D geometric model of the customer is generated using the Poisson surface reconstruction algorithm. The geometric model is then driven by facial motion unit (AU) and skeletal joint data to construct the geometric and physical layers of the digital twin. The structured semantic tags and customer behavior data are learned and used to make decisions based on the LDMP model, thus constructing the behavior layer of the digital twin.

[0009] In some embodiments, the task decomposition step utilizes RPA technology to decompose the advertising creative generation task into a series of executable sub-tasks, specifically including: Based on the hierarchical task network (HTN) planner, the advertising creative generation task is decomposed into objectives according to the three elements of "who to appeal to, what to appeal to, and how to appeal". Based on the decomposed objectives, a corresponding sequence of atomic tasks is generated, wherein the atomic tasks include at least one of keyword extraction, script generation, and multi-version rendering.

[0010] In some embodiments, the LDMP model in the simulation optimization step includes: Learning layer L: Utilizes graph neural network (GNN) to learn from historical interaction data and real-time semantic labels to construct a customer behavior evolution map; Decision layer D: Based on deep reinforcement learning or multi-armed gambling machine MAB model, select the optimal creative strategy from the preset advertising strategy action space; Modeling layer M: Predicts customer responses to advertising creatives using conditional generative adversarial networks (cGANs) or behavior trees instantiated in a virtual engine; Prediction layer P: Through a multi-task prediction model, it simultaneously outputs the predicted click-through rate (CTR) and conversion rate (CVR) of the ad creative.

[0011] In some embodiments, the simulation optimization step further includes a knowledge enhancement step: By utilizing the retrieval enhancement generation (RAG) technology, creative cases and materials related to the current task are retrieved from the advertising creative knowledge base composed of graph database and vector database through a hybrid sorting method that combines keyword retrieval and semantic retrieval, and then injected into the LDMP model for optimization. Among them, the conditional generative adversarial network cGAN is used to predict customers' reactions to advertising creatives.

[0012] In some embodiments, the use of retrieval-enhanced RAG (Retrieval-Based Generative Algorithm) technology to retrieve creative cases and materials relevant to the current task from an advertising creative knowledge base composed of graph databases and vector libraries, through a hybrid sorting method combining keyword retrieval and semantic retrieval, includes: A hybrid ranking strategy is used to retrieve relevant advertising materials from a vector database. The final score of the hybrid ranking strategy is obtained by weighted sum of keyword retrieval score and semantic retrieval score, wherein keyword retrieval is based on the TF-IDF algorithm and semantic retrieval is based on vector similarity calculation.

[0013] In some embodiments, the final score calculation formula for the hybrid ranking strategy is as follows:

[0014] in, These are the weighting coefficients. For keyword retrieval scores, This is used to score semantic retrieval.

[0015] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.

[0016] To achieve the above objectives, another aspect of this application proposes a computational advertising creative agent construction system, comprising: The multimodal data acquisition module is used to collect customer behavioral characteristics, environmental information, and interaction intentions. The scene semantic-driven digital twin module is used to build and dynamically update a digital twin that reflects the customer's state and scene context based on multimodal data; The RPA task decomposition module is used to break down the advertising creative generation task into machine-executable subtasks. The LDMP simulation optimization module is used to conduct simulation tests and effect predictions on new advertising creatives based on the LDMP model and advertising creative knowledge base. The advertising creative generation intelligent agent interactive interface module is used to integrate the various modules and receive input from advertisers or customers.

[0017] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described above.

[0018] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer program product, including a computer program that, when executed by a processor, implements the method described above.

[0019] The embodiments of this application include at least the following beneficial effects: This application provides a method, system, electronic device, storage medium, and program product for constructing a computational advertising creative intelligent agent. This solution, through scene semantic driving and digital twin interaction, combined with RPA automation and generative AI technology, achieves intelligent, automated, and personalized advertising creative generation, significantly improving the efficiency of advertising creative generation and customer interaction experience. Compared with traditional methods that rely on human experience or simple programmatic creatives, this application can more accurately understand customer intent, scene semantics, and interaction needs, thereby generating advertising content that better meets customer needs and improving advertising effectiveness and customer satisfaction. Attached Figure Description

[0020] Figure 1 This is an architecture diagram of a computational advertising creative intelligent agent construction system according to an embodiment of this application; Figure 2 This is a flowchart of a computational advertising creative intelligent agent construction system according to an embodiment of this application; Figure 3 This is a flowchart illustrating the generative computational advertising creative intelligence agent in an embodiment of this application; Detailed Implementation To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0022] Before providing a detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained first. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.

[0023] 1) RPA technology: Robotic Process Automation is a business process automation technology based on software robots and artificial intelligence (AI).

[0024] 2) LDMP Model: LDMP interactive digital twin model based on deep learning for learning, decision-making, modeling, and prediction.

[0025] 3) Poisson Surface Reconstruction Algorithm: In computer graphics and computer vision, Poisson surface reconstruction is an algorithm used to reconstruct a surface model from a set of sampled points. This method is particularly suitable for generating high-quality 3D models from scattered point cloud data. It is based on the Poisson equation, a partial differential equation used to solve how to smoothly interpolate a continuous surface from a given set of sampled points.

[0026] 4) Multi-Armed Bandit (MAB): This is a classic problem in reinforcement learning and a framework for solving the "exploitation dilemma." The MAB problem is a mathematical model that formally describes how to maximize long-term gains by balancing "exploration" and "exploitation" in an uncertain environment. Its application in online advertising includes selecting one ad creative from multiple options to maximize click-through rate or conversion rate.

[0027] The traditional advertising creative production process (mainly before 2010) is characterized by fixed steps, clear hierarchy, and long cycle. Each link is highly dependent on offline collaboration and human experience. Once it enters the subsequent links, it is difficult to make reverse adjustments. Overall, it is more inclined to the communication logic of "one-way brand output".

[0028] After 2020, with the popularization of interactive media such as short videos and social media, and the maturity of AI and big data technologies, the advertising creative production process has shifted to "customer-centric", showing agile, data-driven, and iterative characteristics. Creativity is no longer a "one-time output" but a "dynamic optimization process".

[0029] However, existing programmatic advertising creative optimization focuses on element combinations, lacking an understanding of customers' real interaction scenarios and deep intentions. The creation model cannot keep up with the development of other intelligent links in terms of speed, scale and standards, making it difficult to gain a foothold in the digital advertising chain. The rapid development of large language models (LLM) and generative artificial intelligence in recent years has brought opportunities to reconstruct the advertising creation model.

[0030] In view of this, the embodiments of this application aim to provide a method, electronic device, storage medium and program product for constructing a computational advertising creative intelligent agent based on scene semantic driving and digital twin interaction, to solve the problems of traditional advertising creative generation methods that rely on human experience, are inefficient, lack personalization and lack customer interaction perception, and adopt methods such as multimodal interaction, digital twin modeling, RPA (Robotic Process Automation), generative AI and deep learning integration to improve the intelligence, personalization and programmatic nature of advertising creative generation.

[0031] The computational advertising creative intelligent agent construction method provided in this application relates to technical fields such as data mining, generative AI, digital twins, RPA, and advertising creative generation and optimization. The computational advertising creative intelligent agent construction method provided in this application can be applied to terminals, servers, or software running on terminals or servers. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited to these; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing the computational advertising creative intelligent agent construction method, but is not limited to the above forms.

[0032] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0033] Please see Figure 1 , Figure 1 This is an optional architecture diagram of a computational advertising creative intelligent agent construction system provided in this application embodiment. The system mainly includes five core modules: a multimodal data acquisition module, a digital twin module, an RPA task decomposition module, an LDMP simulation optimization module, and an advertising creative generation intelligent agent interactive interface module. Multimodal data acquisition module: Composed of a series of sensors, used to collect behavioral characteristics (such as gaze focus, gesture operation), environmental information (such as light intensity, location coordinates, timestamp) and interactive intentions (such as voice commands, text input) of customers in the physical scene of viewing advertisements.

[0034] In some embodiments, the multimodal data acquisition module includes: Visual sensors are used to capture customers' 3D skeletal movements and facial expressions; A voice sensor, combined with a speech recognition model, converts speech into text; Environmental sensors are used to collect information on light intensity, temperature, humidity, and location. The intent recognition module is used to integrate visual, voice, and text inputs and parse the customer's task intent through a multimodal Transformer.

[0035] Digital Twin Module: Receives data from the module and constructs a dynamic digital twin containing geometric, physical, semantic, and behavioral layers through semantic parsing, 3D modeling, and physical simulation.

[0036] Specifically, this digital twin module is used to interpret the multimodal data collected by the real-time multimodal data acquisition module and transform it into the corresponding states and behaviors of the digital twin. These states and behaviors reflect the customer's real-time feedback on the advertising creative.

[0037] RPA Task Decomposition Module: Receives semantic information output from the digital twin module and uses the HTN planner to decompose complex creative correction tasks into atomic sub-tasks such as "keyword extraction" and "script generation".

[0038] Specifically, the RPA task decomposition module is used to implement a mapping mechanism that translates the semantic interpretation from the digital twin module to subtasks. These tasks reflect the need to correct the discrepancy between the customer feedback expected by the advertising creative and the actual feedback reflected in the digital twin; that is, the task objective is to revise existing advertising creatives until the current customer feedback is adjusted to the feedback expected by the advertiser. This constructs different task intents for each subtask, which correspond to their respective subtask agents. These agents execute the decomposed subtask steps and, based on the analysis results of the task intents, provide new design principles for the advertising creative.

[0039] LDMP Simulation Optimization Module: Integrates an LLM-based advertising creative knowledge base and LDMP model. Based on the task instructions of the RPA module, it performs knowledge retrieval, strategy decision-making, virtual simulation, and effect prediction to generate and optimize advertising creatives.

[0040] Specifically, the LDMP simulation optimization module is used to select and design new advertising creatives in the knowledge base based on the new design principles of advertising creatives proposed by the RPA task decomposition module, until the customer feedback from the twin simulation meets expectations.

[0041] The intelligent agent interactive interface module for generating advertising creatives serves as the central control and interaction entry point, integrating and coordinating the operation of various modules while allowing advertisers or customers to input instructions or feedback, forming an optimized closed loop.

[0042] Specifically, the ad creative generation intelligent agent interactive interface module integrates the above modules and can accept direct input from advertisers and customers when needed. This input will influence LDMP optimization, enabling it to converge more quickly to new ad creatives that meet expectations. For example, it can call upon content from existing ad creative knowledge bases based on instructions and integrate it into the LDMP optimization process.

[0043] Please see Figure 2 , Figure 2 This is an optional flowchart of a method for constructing a computational advertising creative agent provided in an embodiment of this application. The method includes the following steps: S1. Scene semantic perception step: Real-time collection and analysis of customer behavior characteristics, environmental information and interaction intent through multimodal data.

[0044] For example, customers interact with the system through voice, gestures, text, etc. The multimodal data acquisition module captures these interactions and environmental data (such as lighting and location) in real time, and performs data fusion and intent recognition through technologies such as multimodal Transformer and Kalman filtering.

[0045] S2. Build and dynamically update a digital twin that reflects the customer's real-time status and context.

[0046] In some embodiments, a dynamic digital twin is constructed using the data collected in step S1. Specifically, this includes the following steps: S21. Use Poisson surface reconstruction and Unreal Engine Metahuman technology to build the geometric model.

[0047] S22. Use ST-GCN (Spatiotemporal Graph Convolutional Network) to semantically encode behavioral data and generate structured semantic labels (such as {"intent": "purchase_decision", "environment": "high_illumination"}).

[0048] S23. Initialize the behavior layer of the digital twin based on the LDMP model framework, so that it can simulate the decision-making process of real customers.

[0049] S3. Task decomposition steps: Based on the difference between the actual customer feedback reflected by the digital twin and the advertiser's expected feedback, the advertising creative generation task is decomposed into a series of executable sub-tasks using RPA technology.

[0050] In some embodiments, the "real feedback" reflected by the digital twin is compared with the advertiser's "expected feedback." If a discrepancy exists, RPA task decomposition is triggered. The HTN planner breaks down the creative revision task. For example, if the goal is to increase purchase intent, it may be broken down into atomic tasks such as ["extract the core selling points of the product", "generate promotional scripts", "render 3 video variations"].

[0051] S4. Simulation Optimization Steps: Based on the Learning, Decision, Modeling and Prediction LDMP model, according to the new advertising creative design principles corresponding to the sub-tasks, new advertising creatives are retrieved, generated and simulated in the advertising creative knowledge base until the customer feedback simulated by the digital twin meets the advertiser's expectations.

[0052] Step S4 is the core optimization loop of the system. In some embodiments, the LDMP module performs the following sub-steps: Learning layer (L): Utilizes Gated GNN to integrate creative cases retrieved from RAG and real-time user profiles to update the user-product association matrix.

[0053] Decision layer (D): The UCB algorithm is used to explore and utilize multiple strategies (emotion, function, promotion) to select the current optimal creative strategy.

[0054] Modeling layer (M): Instantiates digital twins and advertising scenarios in the Unity engine, and simulates possible customer reactions to new creative ideas (such as staying or skipping) through Conditional Behavior Tree (CBT).

[0055] Prediction layer (P): Using a model combining DeepFM and CrossNet, the CTR and CVR of the simulated idea are predicted and confidence level is calibrated.

[0056] S5. Creative Generation Steps: Output the advertising creatives that have been simulated and optimized by the LDMP model, collect customer feedback data, and use the feedback data to update the digital twin and optimize the LDMP model, forming a closed-loop optimization process.

[0057] In some embodiments, based on the results of LDMP optimization, the system invokes generative AI (such as GPT-4 for copywriting and Stable Diffusion for images) to generate new advertising creatives. These creatives are then input back into the system (returning to step S1), forming a closed-loop optimization until the simulation feedback from the digital twin reaches the advertiser's desired metrics.

[0058] Below, in conjunction with Figure 3 The present application provides a detailed description and explanation of the embodiments and specific implementation methods.

[0059] This embodiment provides a generative computational advertising creative agent that integrates RPA and is based on scene semantics-driven and digital twin interaction, enabling dynamic perception and modeling of customer behavior, environmental state, and task intent, including: An interactive model is built to capture customers' behavioral characteristics, environmental information, and interaction intentions in specific scenarios in real time. Essentially, it's a dynamically updated digital twin that reflects the state and changes of customers as the audience in real-world scenarios. Based on this, the twin integrates multimodal interaction capabilities, fusing and processing visual, voice, and text input from customers to achieve a "insight-production-presentation-execution" operational logic, thereby generating computational advertising creatives that effectively engage with customers.

[0060] To this end, RPA technology is first integrated to break down complex advertising creatives into executable sub-tasks of "who to appeal to, what to appeal to, and how to appeal," forming automated instructions such as "extracting keywords provided by advertisers, using a short video AI script generator, and intelligently generating multiple scripts that meet the creative requirements within seconds." Through seven dimensions of indicators, including number of followers, likes, comments, favorites, reposts, duration, and keywords, the effectiveness of advertising communication is evaluated and the process is optimized, thereby improving the effectiveness and accuracy of creative production in terms of brand and advertising effects.

[0061] Secondly, a scene-semantic driven digital twin model is constructed and combined with a human-machine collaboration mode to achieve a complete closed loop from task perception, task decomposition, task execution to task feedback, thereby enabling the intelligent agent to make autonomous decisions and execute in complex scenarios. This digital twin model consists of two parts: 1) LLM-based advertising creative knowledge base: This knowledge base utilizes the powerful knowledge integration and reasoning capabilities of LLM to systematically store and manage massive amounts of advertising materials, marketing strategies, customer profiles, success stories, etc., and provides intelligent agents with real-time and accurate creative inspiration and material support through retrieval-enhanced generation (RAG) technology.

[0062] 2) Deep Learning-Based LDMP Interactive Digital Twin Model: This invention proposes a Learning, Decision-making, Modeling, and Predicting (LDMP) model to simulate customer reactions and interpret scene semantics. This model interacts with customers through multimodal methods (such as voice, gestures, and text) to enhance the naturalness and flexibility of human-computer interaction. Simultaneously, it draws upon industry-leading click-through rate (CTR) / conversion rate (CVR) prediction model architecture, continuously optimizing the digital twin model and RPA task execution strategy through a feedback mechanism. It analyzes massive amounts of campaign data to predict the effects of different creative variations generated by the agent, thereby guiding the agent to select and optimize the final advertising plan presented to the customer.

[0063] (1) Constructing a scene semantic perception layer (dynamic digital twin generation) It mainly consists of two parts: data acquisition and digital twin generation. These will be described in detail below.

[0064] 1.1) Real-time acquisition and semantic parsing of multimodal data Deploy sensor networks to capture in real time the behavioral characteristics (such as gaze focus and gesture operation), environmental information (such as light intensity, location coordinates, and timestamps) and interactive intentions (such as voice commands and text input) of customers in the physical scene of viewing advertisements, providing input for building dynamic digital twins.

[0065] The sensors include: a) Visual sensors – using RGB-D cameras (such as Azure Kinect) to capture customers’ 3D skeletal movements, facial micro-expressions (AU units), and object interaction trajectories.

[0066] Pose estimation uses the OpenPose model

[0067] in For the confidence level of the key points, Scoring is given for limb connectivity.

[0068] b) Voice sensor – a microphone array for directional sound pickup, combined with an end-to-end speech recognition model (Wav2Vec2.0) to convert speech into text:

[0069] c) Environmental Sensors – IoT devices collect temperature and humidity (DHT22), illumination (BH1750), and location (UWB positioning module) data to generate an environmental state vector. .

[0070] d) Intent Recognition Module – Integrates visual, speech, and text input, and parses task intent using a multimodal Transformer:

[0071] Output intent category labels (such as "product inquiry", "price comparison", "purchase decision"), and trigger task decomposition when the confidence level is >0.85.

[0072] If necessary, Kalman filtering can be considered for multi-source data fusion.

[0073] in For Kalman gain, For the observed values, This is the observation matrix.

[0074] The core principle of Kalman filtering in multimodal data fusion is to achieve optimal estimation of the system state under noise interference by combining a system dynamic model with multi-sensor observation data through a recursive algorithm. Essentially, it is a closed-loop process of "prediction-correction," aiming to minimize the covariance of the estimation error.

[0075] 1.2) Constructing a scene-semantic driven digital twin By leveraging data collected through sensor networks, a dynamic twin reflecting the customer's real-time state (physiological, psychological, and behavioral) and environmental context can be constructed, supporting the personalized generation of advertising creatives.

[0076] a) First, Metahuman is built using Unreal Engine 5.6. Based on the collected customer body data, 3D modeling and texture mapping are completed to form the geometric layer of the digital twin, and the appearance of the physical entity is reproduced with high fidelity.

[0077] a.1) Multi-source data → Semantic label mapping Input customer behavior data, including: skeletal joint coordinates (OpenPose output), environmental parameters (light intensity, temperature and humidity, geographical location), and interaction intent (NLP parsing results of speech / text); then perform semantic encoding, and use a spatiotemporal graph convolutional network (ST-GCN) to fuse spatiotemporal features:

[0078] in It is an adjacency matrix. These are learnable weights.

[0079] Finally, the output is structured semantic tags, for example: { "user_behavior": "ad_watch(confidence=0.91)", "environment": "high_illumination(>800lux)", "intent": "purchase_decision(urgency=high)" } b) Secondly, the physical characteristics of the human body are simulated based on the Unity physics engine, and the physical layer of the digital twin is constructed by using rigid body dynamics and simulation of mechanical properties. b.1) Parameterized generation of dynamic twins First, geometric modeling is implemented. Based on RGB-D camera point cloud data, a 3D model of the customer is generated using the Poisson surface reconstruction algorithm.

[0080] in Let be the normal vector field of the point cloud.

[0081] Then, behavior-driven processing is completed. Facial expressions are processed by the FACS system through the mapping of 52-dimensional BlendShape parameters, and virtual skeletons are driven by BVH motion data streams. ), The output is generated in real time by the facial motion unit detector. Pose optimization is achieved using inverse kinematics (IK). , This is the joint position function.

[0082] b.2) Environmental semantic fusion and real-time rendering The first step is the implementation of environment mapping, the core of which is the lighting model, which requires physically based rendering (PBR) materials to respond to HDR environment maps:

[0083] in It is the bidirectional reflection distribution function (BRDF).

[0084] Secondly, there is dynamic updating (sensor data → Unreal Engine parameter binding) and virtual-real synchronization (using a publish / subscribe model, transmitting data streams via Apache Kafka).

[0085] c) Finally, based on the LDMP model, a dynamic semantic-driven intelligent response is realized to construct the behavior layer of the digital twin.

[0086] The closed loop of the LDMP model is as follows: L – Learning Layer The input consists of historical interaction data and real-time semantic labels. The method involves using a graph neural network (GNN) to construct a graph of customer behavior evolution.

[0087] The output is the probability distribution of behavioral patterns (such as the transition probability of "browsing → purchasing").

[0088] D – Decision-making level Selecting the optimal advertising strategy based on deep reinforcement learning:

[0089] The reward function .

[0090] M – Modeling layer Constructing a conditional generative adversarial network (cGAN) to predict customer responses:

[0091] Enter ad creative Output the estimated interaction rate.

[0092] P – Predicting layer Integrated multi-task learning model for simultaneous CTR / CVR output:

[0093] Among them, the weight-shared hidden layer Improve efficiency.

[0094] (2) RPA task decomposition and LDMP simulation optimization This section accepts the scene semantics from section (1), and then follows the core logic chain of "RPA task decomposition → LDMP simulation optimization" to finally realize the intelligent generation of advertising ideas.

[0095] 2.1) RPA task decomposition, the steps of which include: 2.1.1) Implement a semantic-to-subtask mapping mechanism First, input the structured tags output from the scene semantic layer in step one (e.g., {intent: "urgent_purchase", product: "smartwatch"}). Then, based on the Hierarchical Task Network (HTN) planner, the ad creative is decomposed into atomic operations: # HTN Planning Pseudocode def htn_planner(semantic_input): # First level: Goal decomposition ("Three elements of the appeal") goals = { "WHO": segm_audience(semantic_input.user_profile), #To whom should the complaint be directed? "WHAT": extract_keywords(semantic_input.intent), # What is the request? "HOW": select_creative_format(semantic_input.env) # How to make the request } # Second Layer: Atomic Task Generation subtasks = [] if goals["HOW"] == "video": subtasks.extend([ {"task": "keyword_extraction", "params": {"source": "brand_materials"}}, {"task": "script_generation", "params": {"LLM": "GPT-4-turbo"}}, {"task": "multivariant_rendering", "params": {"count": 5}} ]) return subtasks# Outputs the machine-executable instruction chain 2.1.2) Complete knowledge-enhanced RAG retrieval. The core is to implement the knowledge base architecture, which mainly includes: Storage layer: Neo4j graph database stores entity relationships (product-audience-creative case studies). Index layer: Embedded advertising creatives in the ChromaDB vector library (CLIP-ViT feature extraction) The search formulas used include: Keyword search (exact match)

[0096] Semantic retrieval (vector similarity)

[0097] Mixed sorting (α = 0.3) The final output includes Top-3 related creative case studies and materials (in JSON format).

[0098] 2.2) Implement LDMP four-stage simulation optimization 2.2.1) Dynamic Knowledge Injection (Learning Layer L) Input: RAG search results + real-time customer profile Knowledge Fusion: Integrating Heterogeneous Data Using Gated Graph Neural Network (GGNN)

[0099] Output an enhanced customer-product relationship matrix .

[0100] 2.2.2) Multi-strategy game (decision level D) Action space A = {Emotional appeal, functional demonstration, promotional stimulus} Decision Model – Balance Exploration and Utilization Using Multi-Armed Gambling Machines (MABs)

[0101] Where T is the total number of attempts, and n a The number of times the action can be selected.

[0102] Finally, the optimal creative strategy is output. .

[0103] 2.2.3) Digital Twin Simulation (Modeling Layer M) First, complete the virtual scene construction and instantiate it in the Unity engine: a) Customer Metahuman b) Ad display interface (size / position adjusted according to contextual semantics) Next, interactive response simulation was conducted, using Conditional Behavior Tree (CBT) to model the customer decision-making process: If you see an ad then if match interests then Duration of stay = Weibull (λ=2.5, k=1.2) else Slide to skip() end end 2.2.4) Quantitative evaluation of effects (prediction layer P) Construct the following multi-task prediction model:

[0104] Among them, the DeepFM component captures low-order feature combinations. CrossNet Components: Explicitly Modeling High-Order Feature Interactions Confidence calibration uses

[0105] Among the parameters ( b) Optimization is performed using order-preserving regression.

[0106] In summary, the generative computational advertising creative agent proposed in this embodiment, which integrates RPA and is based on scene semantics-driven and digital twin interaction, can effectively improve the intelligence, automation, and personalization of advertising creative generation, thereby enhancing advertising effectiveness and customer satisfaction. Experimental results based on real datasets show that the framework designed in this embodiment achieves a 94.00% pass rate on the expert-built dataset and a 100.00% pass rate on the LLM-generated dataset, indicating that the framework proposed in this embodiment achieves good results in the network intent recognition task.

[0107] Compared with the prior art, the present invention has the following significant advantages: 1) Deep Scene Understanding: Through multimodal perception and digital twin technology, a virtual avatar that dynamically reflects the customer's state and intention is constructed, realizing a deep and real-time semantic understanding of the advertising scenario, which surpasses the traditional tag-based user profile.

[0108] 2) Closed-loop dynamic optimization: It innovatively combines RPA task automation with LDMP model simulation optimization to form a complete closed loop of "perception-decision-simulation-correction", which enables advertising creatives to be continuously iterated and optimized based on user feedback, realizing true "human-in-the-loop" intelligence.

[0109] 3) High degree of personalization and automation: By utilizing generative AI and RAG technologies, it can quickly and automatically generate a large number of personalized creative variations that conform to the semantics of the scene, and conduct pre-simulation testing through digital twins, which greatly improves the efficiency and accuracy of creative production.

[0110] 4) Synergistic effect of technology integration: This invention is not a simple stacking of multiple technologies, but through systematic architectural design, it enables cutting-edge technologies such as digital twins, RPA, generative AI, and deep learning to generate a powerful synergistic effect, which jointly solves the core pain points in the field of advertising creativity and has non-obviousness.

[0111] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0112] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0113] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0114] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0115] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0116] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0117] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented in the embodiments of this program product are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments. The executable computer program code or "code" used to perform the various embodiments can be written in high-level programming languages ​​such as C, C++, Python, Smalltalk, Java, JavaScript, Visual Basic, Structured Query Language (e.g., Transact-SQL), Perl, or in various other programming languages.

[0118] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0119] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0120] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0121] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0122] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0123] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0124] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0125] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0126] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0127] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0128] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for constructing a computational advertising creative intelligent agent, characterized in that, The method includes the following steps: Scene semantic perception steps: Real-time collection and analysis of customer behavior characteristics, environmental information and interaction intentions through multimodal data, to construct and dynamically update a digital twin reflecting the customer's real-time state and scene context; Task decomposition steps: Based on the difference between the actual customer feedback reflected by the digital twin and the advertiser's expected feedback, the advertising creative generation task is decomposed into a series of executable sub-tasks using RPA technology; Simulation optimization steps: Based on the Learning, Decision, Modeling and Prediction LDMP model, according to the new advertising creative design principles corresponding to the sub-tasks, new advertising creatives are retrieved, generated and simulated in the advertising creative knowledge base until the customer feedback simulated by the digital twin meets the advertiser's expectations; Creative generation steps: Output the advertising creatives that have been simulated and optimized using the LDMP model; Building and dynamically updating digital twins specifically includes: Multimodal data is collected through visual sensors, voice sensors, environmental sensors, and intent recognition modules; Based on the multimodal data, a semantic layer of the digital twin is constructed through semantic encoding, and structured semantic tags are output. Based on RGB-D camera point cloud data, a 3D geometric model of the customer is generated using the Poisson surface reconstruction algorithm. The geometric model is then driven by facial motion unit (AU) and skeletal joint data to construct the geometric and physical layers of the digital twin. The structured semantic tags and customer behavior data are learned and used to make decisions based on the LDMP model, thus constructing the behavior layer of the digital twin.

2. The method according to claim 1, characterized in that, In the task decomposition step, RPA technology is used to break down the advertising creative generation task into a series of executable sub-tasks, specifically including: Based on the hierarchical task network (HTN) planner, the advertising creative generation task is decomposed into objectives according to the three elements of "who to appeal to, what to appeal to, and how to appeal"; Based on the decomposed objectives, a corresponding sequence of atomic tasks is generated, wherein the atomic tasks include at least one of keyword extraction, script generation, and multi-version rendering.

3. The method according to claim 1, characterized in that, The LDMP model in the simulation optimization step includes: Learning layer L: Utilizes graph neural network (GNN) to learn from historical interaction data and real-time semantic labels to construct a customer behavior evolution map; Decision layer D: Based on deep reinforcement learning or multi-armed gambling machine MAB model, select the optimal creative strategy from the preset advertising strategy action space; Modeling layer M: Predicts customer responses to advertising creatives using conditional generative adversarial networks (cGANs) or behavior trees instantiated in a virtual engine; Prediction layer P: Through a multi-task prediction model, it simultaneously outputs the predicted click-through rate (CTR) and conversion rate (CVR) of the ad creative.

4. The method according to claim 1, characterized in that, The simulation optimization step also includes a knowledge enhancement step: By utilizing the retrieval enhancement generation (RAG) technology, creative cases and materials related to the current task are retrieved from the advertising creative knowledge base composed of graph database and vector database through a hybrid sorting method that combines keyword retrieval and semantic retrieval, and then injected into the LDMP model for optimization. Among them, the conditional generative adversarial network cGAN is used to predict customers' reactions to advertising creatives.

5. The method according to claim 4, characterized in that, The aforementioned retrieval-enhanced RAG (Research-Based Aggregation) technology retrieves creative cases and materials relevant to the current task from an advertising creative knowledge base comprised of graph databases and vector databases. This retrieval is achieved through a hybrid sorting method combining keyword retrieval and semantic retrieval, including: A hybrid ranking strategy is used to retrieve relevant advertising materials from a vector database. The final score of the hybrid ranking strategy is obtained by weighted sum of keyword retrieval score and semantic retrieval score, wherein keyword retrieval is based on the TF-IDF algorithm and semantic retrieval is based on vector similarity calculation.

6. The method according to claim 5, characterized in that, The final score calculation formula for the hybrid sorting strategy is as follows: in, These are the weighting coefficients. For keyword retrieval scores, This is used to score semantic retrieval.

7. A computational advertising creative intelligent agent construction system, characterized in that, include: The multimodal data acquisition module is used to collect customer behavioral characteristics, environmental information, and interaction intentions. The scene semantic-driven digital twin module is used to build and dynamically update a digital twin that reflects the customer's state and scene context based on multimodal data; The RPA task decomposition module is used to break down the advertising creative generation task into machine-executable subtasks. The LDMP simulation optimization module is used to conduct simulation tests and effect predictions on new advertising creatives based on the LDMP model and advertising creative knowledge base. The advertising creative generation intelligent agent interactive interface module is used to integrate the various modules and receive input from advertisers or customers; The scene semantic-driven digital twin module constructs and dynamically updates the digital twin, specifically including: Multimodal data is collected through visual sensors, voice sensors, environmental sensors, and intent recognition modules; Based on the multimodal data, a semantic layer of the digital twin is constructed through semantic encoding, and structured semantic tags are output. Based on RGB-D camera point cloud data, a 3D geometric model of the customer is generated using the Poisson surface reconstruction algorithm. The geometric model is then driven by facial motion unit (AU) and skeletal joint data to construct the geometric and physical layers of the digital twin. The structured semantic tags and customer behavior data are learned and used to make decisions based on the LDMP model, thus constructing the behavior layer of the digital twin.

8. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Intelligent delivery parameter optimization method and system based on data delivery feedback effect

    CN120031612A