Interactive advertisement generation method and device, equipment and storage medium

By using multimodal analysis and user behavior simulation, interactive ads are generated and adjusted, solving the problems of homogenization and prediction lag, and achieving efficient and diversified interactive ad generation.

CN121998707APending Publication Date: 2026-05-08广州三七极耀网络科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
广州三七极耀网络科技有限公司
Filing Date
2025-12-31
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing interactive advertising technologies suffer from severe homogenization and a lack of synergy with user interaction, and the prediction of results is delayed, making it difficult to adjust advertising content in a timely manner.

Method used

By parsing text data in a multimodal manner to generate text tags, 3D model features, and interaction logic, semantic alignment and association binding are performed. Combined with device-adaptive rendering and user behavior simulation, interactive advertisements are generated and adjusted.

Benefits of technology

It has enriched the diversity of interactive ads, improved generation efficiency, enabled timely adjustments to ad performance, and enhanced the user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998707A_ABST
    Figure CN121998707A_ABST
Patent Text Reader

Abstract

The invention provides an interactive advertisement generation method and device, equipment and a storage medium, and the method comprises the steps: obtaining inputted text data, and carrying out the multi-mode analysis of the text data, and obtaining a text label, a three-dimensional model feature and interaction logic; performing semantic alignment and association binding processing on the text tag, the three-dimensional model feature and the interaction logic to obtain rendering data; performing equipment adaptation rendering on the rendering data to generate advertisement content, and performing user behavior simulation prediction on the advertisement content to obtain a prediction result; and adjusting the advertisement content based on the prediction result to generate an interactive advertisement. The interactive advertisement generated by the scheme is richer and more diversified, the prediction effect of the interactive advertisement can be obtained and adjusted in time, and the generation efficiency and the actual use effect of the interactive advertisement are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to an interactive advertising generation method, apparatus, device, and storage medium. Background Technology

[0002] Interactive advertising is a new type of advertising that delivers information through two-way interaction. Unlike traditional static advertising, interactive ads allow users to interact with the content, thus enhancing the user experience. For example, during the display of an interactive ad, users can try it out, interact with the ad content, and thus disseminate information in a more vivid and engaging way.

[0003] In related technologies, the generation of interactive ads primarily relies on template-based generation tools and monomodal AI generation technology. However, ads generated in this way suffer from severe homogenization and lack synergy with user interaction. Furthermore, because traditional A / B testing requires several days to obtain feedback, the predictive effects of interactive ads exhibit a certain lag after generation, hindering timely content release and adjustments. Summary of the Invention

[0004] This application provides an interactive advertising generation method, apparatus, device, and storage medium, solving the problems of severe homogenization of generated advertising content, lack of synergy in user interaction, and a certain lag in prediction results in related technologies. By performing multimodal analysis on text data to generate interactive advertisements, its diversity is enriched. By utilizing user behavior simulation mechanisms to obtain prediction results and make corresponding adjustments, the predicted effect of interactive advertisements can be obtained and adjusted in a timely manner, improving the generation efficiency and actual use effect of interactive advertisements.

[0005] Firstly, this application provides a method for generating interactive advertisements, including: The system acquires the entered text data and performs multimodal parsing on the text data to obtain text labels, 3D model features, and interaction logic. The text labels, the 3D model features, and the interaction logic are semantically aligned and associated to obtain rendering data. The rendered data is used to generate advertising content through device-adaptive rendering, and the advertising content is used to simulate and predict user behavior to obtain prediction results. Based on the prediction results, the advertising content is adjusted to generate interactive ads.

[0006] Optionally, the step of performing multimodal parsing on the text data to obtain text labels, 3D model features, and interaction logic includes: The text data is parsed using a set language model to obtain structured policy instructions; Based on the structured strategy instructions and the set model library, corresponding text labels, 3D model features, and interaction logic are generated.

[0007] Optionally, the step of generating corresponding text labels, 3D model features, and interaction logic based on the structured strategy instructions and the set model library includes: Determine the text tags corresponding to the structured strategy instructions; The visual features, sound features, and interactive features associated with the text tags are determined by setting up a visual element library, a sound effect library, and a dynamic effect library; The three-dimensional model features corresponding to the structured strategy instruction are determined based on the visual features, the sound effect features, and the set three-dimensional model library. The interaction logic corresponding to the structured strategy instruction is determined based on the interaction features and the 3D model library.

[0008] Optionally, the step of semantically aligning and associating the text labels, the 3D model features, and the interaction logic to obtain rendering data includes: Define a virtual timeline, which includes virtual key points and interaction states; Align the text labels, the 3D model features, and the interaction logic with the virtual key points and interaction states to obtain aligned text labels, 3D model features, and interaction logic; Semantic hotspots are determined in the aligned 3D model features, and the aligned text labels and interaction logic are bound to the semantic hotspots to obtain rendering data.

[0009] Optionally, the step of performing device-adaptive rendering on the rendered data to generate advertising content includes: Obtain the built-in device performance graph; Multiple progressively tiered advertising content are generated based on the device performance graph and the rendering data.

[0010] Optionally, the step of performing user behavior simulation prediction on the advertising content to obtain the prediction result includes: An advertising interaction prototype is generated based on the advertising content. The advertising interaction prototype is subjected to a simulated operation of preset user behavior, and the churn node corresponding to the simulated operation is recorded to obtain a prediction result containing at least one churn node.

[0011] Optionally, adjusting the advertising content based on the prediction result to generate interactive ads includes: In the advertising content, determine one or more adjustable elements corresponding to the lost node and the associated element adjustment values; Based on the element adjustment values, the corresponding adjustable elements in the advertisement content are adjusted to generate an interactive advertisement.

[0012] Secondly, this application provides an interactive advertising generation device, comprising: The acquisition module is used to acquire the entered text data; The parsing module is used to perform multimodal parsing on the text data to obtain text labels, 3D model features, and interaction logic; The information processing module is used to perform semantic alignment and association binding on the text labels, the 3D model features, and the interaction logic to obtain rendering data; The content generation module is used to perform device-adaptive rendering on the rendering data to generate advertising content, perform user behavior simulation and prediction on the advertising content to obtain prediction results, and adjust the advertising content based on the prediction results to generate interactive advertisements.

[0013] Thirdly, this application also provides an interactive advertising generation device, the device comprising: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the interactive advertising generation method as described in any of the preceding first aspects.

[0014] Fourthly, this application also provides a storage medium for storing computer-executable instructions, which, when executed by a computer processor, are used to perform the interactive advertising generation method as described in any of the first aspects above.

[0015] In the solution provided in this application embodiment, the input text data is obtained, and multimodal parsing is performed on the text data to obtain text tags, 3D model features, and interaction logic. Then, semantic alignment and association binding processing is performed on the text tags, 3D model features, and interaction logic to obtain rendering data. Device-adaptive rendering is performed on the rendering data to generate advertising content, and user behavior simulation prediction is performed on the advertising content to obtain prediction results. Finally, the advertising content is adjusted based on the prediction results to generate interactive advertisements. By performing multimodal parsing on text data to generate interactive advertisements, its diversity is enriched. By using the user behavior simulation mechanism to obtain prediction results and make corresponding adjustments, the prediction effect of interactive advertisements can be obtained in a timely manner and adjustments can be made, thereby improving the generation efficiency and actual use effect of interactive advertisements. Attached Figure Description

[0016] Figure 1This is a flowchart of an interactive advertisement generation method provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a method for multimodal parsing of text data provided in an embodiment of this application; Figure 3 This is a flowchart illustrating a method for generating corresponding text labels, 3D model features, and interaction logic based on structured strategy instructions, as provided in an embodiment of this application. Figure 4 This is a flowchart illustrating a method for obtaining rendered data through semantic alignment and association binding, as provided in an embodiment of this application. Figure 5 This is a flowchart illustrating a method for obtaining prediction results by simulating user behavior in advertising content, as provided in an embodiment of this application. Figure 6 This is a block diagram of the module structure of an interactive advertising generation device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an interactive advertising generation device provided in an embodiment of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. It should also be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as being processed sequentially, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. A process can be terminated when its operation is completed, but it may also have additional steps not included in the drawings. A process can correspond to a method, function, procedure, subroutine, subroutine, etc.

[0018] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0019] The interactive ad generation method provided in this application can automatically generate and adjust interactive ads. Unlike traditional methods that use template-based generation tools (such as Canva) and single-modal AI generation technologies (such as DALL-E for image generation), this method avoids the problem of severe ad homogenization. It addresses the current situation where 70% of advertisers report ad schemes generated by existing tools have a similarity exceeding 60%. This method significantly improves the efficiency of effect prediction and adjustment compared to existing technologies. Traditional A / B testing requires 3-7 days of real-world deployment to obtain effective data and suffers from cross-modal fragmentation.

[0020] The interactive advertising generation method provided in this application embodiment can be executed by a device with computing power, such as a server, laptop, or desktop computer.

[0021] Figure 1 This is a flowchart of an interactive advertisement generation method provided in an embodiment of this application. Figure 1 As shown, the interactive ad generation method includes: Step S101: Obtain the entered text data, and perform multimodal parsing on the text data to obtain text labels, 3D model features, and interaction logic.

[0022] The text data can be raw text information input by the user, such as product descriptions, advertising copy, or creative requirements, for example, "summer promotion of athletic shoes." Multimodal parsing refers to using natural language processing and artificial intelligence technologies to extract structured information from text data.

[0023] Among them, text tags are keywords parsed from text. For example, for "summer promotion of sports shoes", the corresponding parsed text tags could be "breathable" and "lightweight". 3D model features can be 3D model attributes inferred from the specific description of the text data, such as shape, material, color, and size. For example, the 3D model feature parsed for the text data "summer promotion of sports shoes" is "mesh structure". Interaction logic is user interaction rules identified or inferred from the text data, such as "click to jump" and "drag to rotate". For example, the interaction logic parsed for the text data "summer promotion of sports shoes" is "360° rotation".

[0024] In one embodiment, the multimodal parsing of text data to obtain text tags, 3D model features, and interaction logic can be implemented as follows: A pre-trained natural language understanding model is used to analyze the text, and named entity recognition and semantic role labeling are used to extract key text tags. Simultaneously, a 3D generative model is used to transform descriptions of object shapes and appearances in the text data into parameterized 3D model features. Then, combined with a rule engine and intent recognition algorithm, the user's desired interaction method is parsed from the text data, ultimately outputting structured text tags, 3D model features, and interaction logic. For example, taking the text data "summer promotion of athletic shoes" as an example, it is ensured that when the generated AR shoe model rotates, the corresponding features of the advertising text are automatically highlighted.

[0025] Step S102: Perform semantic alignment and association binding on text labels, 3D model features, and interaction logic to obtain rendering data.

[0026] Among them, semantic alignment processing can ensure that data from different sources are consistent and coordinated in meaning, while association binding processing enables the splicing of different data elements, namely structured text labels, 3D model features and interaction logic, into a unified and operable dataset.

[0027] In one embodiment, the way to obtain rendering data by semantically aligning and associating text labels, 3D model features, and interaction logic can be as follows: the text labels are mapped to 3D model features through a semantic matching algorithm, for example, associating the "red" label with the model's color attribute; at the same time, the interaction logic is bound to specific 3D model parts or scene elements to ensure that user operations can trigger the correct response; finally, a complete rendering data package containing geometric information, material maps, animation sequences, and interactive event triggers is generated.

[0028] Step S103: Perform device-adaptive rendering on the rendered data to generate advertising content, and perform user behavior simulation prediction on the advertising content to obtain prediction results.

[0029] Among them, device-adaptive rendering refers to optimizing rendering output based on the performance and display characteristics of different terminal devices; user behavior simulation prediction refers to predicting user behavior responses by simulating user interaction with advertisements and obtaining corresponding prediction results, which can be quantitative evaluation data derived from simulation prediction.

[0030] In one embodiment, the method of generating ad content by performing device-adaptive rendering on rendered data and obtaining prediction results by simulating user behavior on the ad content can be as follows: A rendering engine can adaptively adjust rendering parameters based on the target device's screen resolution, GPU performance, and platform characteristics to generate adapted ad content. Correspondingly, historical interaction data and machine learning models can be used to simulate the clicks, browsing duration, and other behaviors of different user groups on the ad in a virtual environment. By analyzing the simulated data, the interactive effect and user engagement of the ad can be predicted, and the prediction results of key indicators can be output. For example, taking the aforementioned original text data "Summer Promotion of Sports Shoes" as an example, after generating an interactive ad and simulating user trial behavior, a 62% probability is obtained that the user will click the "Share to Social Platform" button, thus obtaining the corresponding prediction result.

[0031] Another example is the interactive advertising solution generated for the input text data "Immersive experience of new lipstick products" which is "AR virtual makeup combined with color matching mini-game". Through user behavior simulation and prediction, it is found that "23% of users will complete the entire path from trying on colors to scoring in the game and then receiving coupons".

[0032] Step S104: Adjust the advertising content based on the prediction results to generate interactive ads.

[0033] Adjusting ad content based on prediction results can optimize ad content and interaction design to generate the final interactive ad delivery. Optionally, adjusting ad content based on prediction results to generate interactive ads can involve: analyzing weaknesses in the prediction results, such as low click-through rate areas or low interaction duration modules, and automatically adjusting the visual effects of the 3D model, the triggering conditions of the interaction logic, or the layout of ad elements; iteratively optimizing through an A / B testing framework to finally generate an interactive ad version that has been data-validated, adapted to multiple devices, and has high expected interactive effects. For example, one optimization method is: when the tracking latency of a visual element in an interactive ad running on an Android device is detected to be greater than 200ms, the visual element is automatically downgraded to a 2D overlay scheme.

[0034] As described above, the process involves acquiring input text data, performing multimodal parsing to obtain text tags, 3D model features, and interaction logic, then semantically aligning and binding these elements to obtain rendering data. This rendering data is then used for device-adaptive rendering to generate advertising content. User behavior simulation is then used to predict the advertising content and generate a prediction result. Finally, the advertising content is adjusted based on the prediction result to generate an interactive advertisement. By performing multimodal parsing on text data to generate interactive advertisements, the diversity of these advertisements is enriched. Furthermore, by utilizing user behavior simulation mechanisms to obtain prediction results and making corresponding adjustments, the predictive effect of interactive advertisements can be obtained and adjusted in a timely manner, improving the generation efficiency and actual usage effect of interactive advertisements.

[0035] Figure 2 This is a flowchart illustrating a method for multimodal parsing of text data provided in an embodiment of this application, such as... Figure 2 As shown, it includes: Step S201: Obtain the input text data, perform multimodal parsing on the text data through the set language model to obtain structured strategy instructions, and generate corresponding text labels, 3D model features and interaction logic based on the structured strategy instructions and the set model library.

[0036] This large language model is generated based on a massive amount of literature on advertising, marketing psychology, and product design. It can reverse engineer the commercial and emotional intentions corresponding to text data. For example, if the input is "summer promotion of sports shoes", the corresponding output structured instructions are "{"core appeal": "establish the perception of coolness and lightweight technology", "target emotion": "inspire a sense of eagerness and vitality", "key information hierarchy": ["visual impact (mesh)" > "physical cues (breathability)" > "technological credibility (lightweight materials)"]}.

[0037] Upon receiving the structured strategy instruction, the system generates corresponding text labels, 3D model features, and interaction logic based on the instruction and the established model library. Optionally, such as... Figure 3 As shown, Figure 3 This is a flowchart illustrating a method for generating corresponding text labels, 3D model features, and interaction logic based on structured strategy instructions, as provided in an embodiment of this application. Figure 3 As shown, it includes: Step S2011: Determine the text tags corresponding to the structured strategy instructions.

[0038] In determining the text labels corresponding to structured policy instructions, a trained large language model can be used to perform deep semantic deconstruction of the structured policy instructions. By analyzing the emotional tone and cognitive goals in the intent, a set of text labels with weights and emotional connotations can be inferred and generated. For example, for the intent of the text data "to create a feeling of coolness", the model will output multi-dimensional text labels such as "transparent", "icy blue", and "burden-free".

[0039] Step S2012: Determine the visual features, sound features, and interactive features associated with the text label through the set visual element library, sound effect library, and dynamic effect library.

[0040] In one embodiment, a visual element library, a sound effect library, and a dynamic effect library are pre-defined. For a determined text tag, keyword matching can be used to determine the visual features, sound effect features, and interactive features corresponding to the text tag in the visual element library, sound effect library, and dynamic effect library, respectively.

[0041] Step S2013: Determine the 3D model features corresponding to the structured strategy instructions based on visual features, sound effect features, and the set 3D model library.

[0042] In one embodiment, when determining the 3D model features, the 3D model features corresponding to the structured strategy instruction are determined based on the aforementioned determined visual features, sound effect features, and the established 3D model library. Optionally, a semantic-based coarse screening is first performed in the 3D model library based on visual features (such as "mesh texture" and "cool blue") to obtain a basic shoe model. Subsequently, the basic model is fine-tuned in real time through the established differentiable rendering and procedural generation modules. For example, based on the specific parameters of "mesh texture," a procedural perforated mesh is automatically generated in a specific area of ​​the shoe upper (such as the side waist); and based on the color psychology parameters of "cool blue," a blue material with subtle warm and cool variations is mixed on the material sphere. The final output 3D model features can be, for example, a composite of the basic model and parametric local features.

[0043] Step S2014: Determine the interaction logic corresponding to the structured strategy instruction based on the interaction features and the 3D model library.

[0044] In one embodiment, the interaction logic is determined based on the aforementioned determined interaction features and the established 3D model library. The 3D model library records 3D models of multiple different objects in different scenes. Each 3D model corresponds to one or more interactive structures. Taking a determined shoe model as an example, the interactive structure can be a partial mesh upper or the entire shoe body. After obtaining the interaction features and the interactive structures of the 3D models in the 3D model library, the interaction features are matched with the interactive structures. Specifically, multiple interaction features to be matched can be predefined for each interactive structure. The current interaction feature can be matched with multiple unmatched interaction features. If a match is successful, the current interaction feature is associated with the interactive structure of the 3D model. Then, the discrete associated features and interactive structures can be compiled into a behavior script compatible with the rendering engine using a set interaction logic compiler. This behavior script defines the trigger conditions, action sequences, and ending states, thus constituting the complete interaction logic. For example, the trigger condition could be "a finger hovers over the mesh area," and the action sequence could be "increasing the transparency of the mesh material from 0.3 to 0.8 within 0.2 seconds."

[0045] Step S202: Perform semantic alignment and association binding on text labels, 3D model features, and interaction logic to obtain rendering data.

[0046] Step S203: Perform device-adaptive rendering on the rendered data to generate advertising content, and perform user behavior simulation prediction on the advertising content to obtain prediction results.

[0047] Step S204: Adjust the advertising content based on the prediction results to generate interactive ads.

[0048] As described above, by acquiring the input text data, multimodal parsing of the text data using a set language model yields structured strategy instructions. Based on these structured strategy instructions and the set model library, corresponding text tags, 3D model features, and interaction logic are generated. Then, semantic alignment and association binding are performed on the text tags, 3D model features, and interaction logic to obtain rendering data. Device-adaptive rendering is performed on the rendering data to generate advertising content. User behavior simulation and prediction are then performed on the advertising content to obtain prediction results. Finally, the advertising content is adjusted based on the prediction results to generate interactive advertisements. By performing multimodal parsing on text data to generate interactive advertisements, its diversity is enriched. By using a user behavior simulation mechanism to obtain prediction results and make corresponding adjustments, the prediction effect of interactive advertisements can be obtained and adjusted in a timely manner, improving the generation efficiency and actual usage effect of interactive advertisements.

[0049] Figure 4This is a flowchart illustrating a method for obtaining rendered data through semantic alignment and association binding processing, as provided in an embodiment of this application. Figure 4 As shown, it includes: Step S401: Obtain the entered text data, and perform multimodal parsing on the text data to obtain text labels, 3D model features, and interaction logic.

[0050] Step S402: Define a virtual timeline, align text labels, 3D model features, and interaction logic with virtual key points and interaction states to obtain aligned text labels, 3D model features, and interaction logic. Determine semantic hot zones in the aligned 3D model features, and bind the aligned text labels and interaction logic to the semantic hot zones to obtain rendering data.

[0051] The virtual timeline can be understood as a defined "timeline script." For interactive displays of different 3D models, multiple key time points, or virtual keypoints, are pre-marked on this timeline. For example, in a 30-second advertisement showcasing a shoe model, key time points are pre-defined in the "timeline script" corresponding to that shoe model. Examples could be: the shoe appears at second 0, it starts rotating automatically at second 5, and the camera zooms in for a close-up of the shoe's upper at second 10. Then, text labels, 3D model features, and interaction logic are "anchored" to the most relevant key points on this virtual timeline. The anchoring process involves calculating the correlation between the corresponding text labels, 3D model features, and interaction logic and the semantic description of each key time point. The key time point with the highest correlation value is then identified as the key point anchored by the corresponding text label, 3D model feature, and interaction logic. For example, text labels (such as "breathable technology"), 3D model features (such as the mesh structure of the shoe upper) and interactive logic (such as "highlight after clicking") are "anchored" to the 10th second, the 5th second and the 10th second of the virtual timeline, respectively.

[0052] Furthermore, in the aligned data, semantic hotspots of the 3D model features are determined, and the aligned text labels and interaction logic are bound to these semantic hotspots to obtain the rendered data. Optionally, semantic hotspots can be determined by manual annotation or by designating the model's preset structure as interactive hotspot areas. For example, an invisible click box can be marked on the mesh material structure of a shoe, and this box can be named "breathable mesh hotspot." After obtaining the semantic hotspots, the aligned text labels and interaction logic are bound to the determined semantic hotspots in the form of programming scripts. The binding condition can be the precise state of the current virtual timeline. Each text label and interaction logic is identified with its corresponding usage state. During binding, each precise state of the virtual timeline is traversed. When the precise state matches the usage state, the corresponding text label and / or interaction logic are bound to that precise state. For example, two precise states can be predefined for the shoe model: state A (global display) and state B (close-up of the shoe upper). In state A, when the advertisement is in the "global display" state, if the user clicks on the hot area of ​​the shoe upper, a prompt box will pop up displaying "breathable mesh". When the advertisement is in the "close-up of the shoe upper" state, if the user clicks on the same hot area of ​​the shoe upper, it will be "highlighted".

[0053] Step S403: Perform device-adaptive rendering on the rendered data to generate advertising content, and perform user behavior simulation prediction on the advertising content to obtain prediction results.

[0054] Step S404: Adjust the advertising content based on the prediction results to generate interactive ads.

[0055] As described above, by acquiring the input text data, performing multimodal parsing to obtain text tags, 3D model features, and interaction logic, and then performing semantic alignment and association binding on the text tags, 3D model features, and interaction logic to obtain rendering data, performing device-adaptive rendering on the rendering data to generate advertising content, and performing user behavior simulation prediction on the advertising content to obtain prediction results, and finally adjusting the advertising content based on the prediction results to generate interactive advertisements, the multimodal parsing of text data to generate interactive advertisements enriches their diversity. By using the user behavior simulation mechanism to obtain prediction results and making corresponding adjustments, the prediction effect of interactive advertisements can be obtained in a timely manner and adjustments can be made, improving the generation efficiency and actual usage effect of interactive advertisements.

[0056] In one embodiment, device-adaptive rendering of rendering data to generate advertising content includes: obtaining a built-in device performance graph, and generating multiple progressively tiered advertising content based on the device performance graph and the rendering data. The device performance graph is a built-in and continuously updated graph-style data, including typical GPU computing power, memory, and network values ​​for various mobile phones, tablets, and VR headsets. Corresponding rendering processing methods can be pre-set for different device performance levels to obtain corresponding advertising content. Taking high-end and mid-range mobile devices as examples, for high-end mobile devices, a more detailed polygonal shoe model is used, configured with a real-time ray-traced reflection display mode, and implemented using a complex global particle system; for mid-range mobile devices, the more detailed polygonal shoe model is adjusted to a less detailed optimized version, and the real-time ray-traced reflection is replaced with a pre-baked cube map; the complex global particle system is simplified to a 2D animation in screen space. Simultaneously, an interface is reserved in the rendering data to allow users to automatically select a "smoothness priority" or "image quality priority" mode during loading or runtime.

[0057] Figure 5 This is a flowchart illustrating a method for obtaining prediction results by simulating user behavior in advertising content, as provided in an embodiment of this application. Figure 5 As shown, it includes: Step S501: Obtain the entered text data, and perform multimodal parsing on the text data to obtain text labels, 3D model features, and interaction logic.

[0058] Step S502: Perform semantic alignment and association binding on text labels, 3D model features, and interaction logic to obtain rendering data.

[0059] Step S503: Perform device-adaptive rendering on the rendering data to generate advertising content, generate an advertising interaction prototype based on the advertising content, simulate preset user behaviors on the advertising interaction prototype, and record the churn nodes of the corresponding simulated operations to obtain a prediction result containing at least one churn node.

[0060] In one embodiment, a low-fidelity, fast-loading ad interaction prototype is pre-generated before final rendering. By integrating graph neural networks and reinforcement learning path prediction algorithms, the interaction process of various typical users (such as fast browsers, detail explorers, and sharing users) with the ad interaction prototype is simulated. The drop-off nodes of the corresponding simulated operations are recorded to obtain prediction results containing at least one drop-off node (such as directly exiting a certain information page), which are used to adjust the final ad content.

[0061] Step S504: Adjust the advertising content based on the prediction results to generate interactive ads.

[0062] In one embodiment, after obtaining a prediction result containing at least one churn node, the process of adjusting the ad content can be as follows: In the ad content, determine one or more adjustable elements corresponding to the churn node and their associated adjustment values; based on these adjustment values, adjust the corresponding adjustable elements in the ad content to generate an interactive ad. Specifically, during ad generation, a parameterized content element library is predefined. Each adjustable element in the library (such as button position, color code, text string, animation duration, and model display angle) is assigned a unique identifier and a set of adjustable parameters. Simultaneously, a mapping table is established to logically associate each adjustable element with one or more key experience objectives (such as "increasing click-through rate," "reducing bounce rate," and "extending dwell time"). When user behavior simulation prediction or actual monitoring data identifies a "churn node" (i.e., a point where users significantly interrupt their experience), the attribution analysis model can be used to calculate one or more experience dimension defects that are most likely to cause churn (e.g., "information overload," "operational confusion," "lack of incentives"). Then, based on the aforementioned mapping table, a set of candidate adjustable elements directly related to these defect dimensions can be retrieved in reverse. For each identified candidate adjustable element, a pre-defined optimization strategy function library can be invoked. This function library applies preset optimization heuristic rules based on the identified defect type (e.g., to address "operational confusion," the rule is "increase the size of the main action button and enhance visual contrast"; to address "lack of incentives," the rule is "add progress prompts or instant reward feedback at key points"). This calculates one or more specific, quantified element adjustment values ​​for each adjustable element (e.g., adjust the width of button A from 80 pixels to 120 pixels, and increase the background color contrast value from 4.5:1 to 7:1). Accordingly, after obtaining the element adjustment value, the rendering template or generation engine of the ad content is called to generate an adjustable element with the corresponding element adjustment value, so as to integrate it into the ad content.

[0063] As described above, by acquiring the input text data, performing multimodal parsing to obtain text tags, 3D model features, and interaction logic, and then performing semantic alignment and association binding on the text tags, 3D model features, and interaction logic to obtain rendering data, performing device-adaptive rendering on the rendering data to generate advertising content, generating an advertising interaction prototype based on the advertising content, simulating preset user behaviors on the advertising interaction prototype, and recording the churn nodes of the corresponding simulated operations to obtain a prediction result containing at least one churn node, and finally adjusting the advertising content based on the prediction result to generate interactive ads, the multimodal parsing of text data to generate interactive ads enriches their diversity. By using a user behavior simulation mechanism to obtain prediction results and make corresponding adjustments, the prediction effect of interactive ads can be obtained and adjusted in a timely manner, improving the generation efficiency and actual usage effect of interactive ads.

[0064] Figure 6 This is a block diagram of the module structure of an interactive advertisement generation device provided in an embodiment of this application. The device is used to execute an interactive advertisement generation method provided in the above embodiment, and has corresponding functional modules and beneficial effects for executing the method. Figure 6 As shown, the device specifically includes: The acquisition module 101 is used to acquire the entered text data; The parsing module 102 is used to perform multimodal parsing on the text data to obtain text labels, three-dimensional model features, and interaction logic; Information processing module 103 is used to perform semantic alignment and association binding processing on the text labels, the three-dimensional model features and the interaction logic to obtain rendering data; The content generation module 104 is used to perform device-adaptive rendering on the rendering data to generate advertising content, perform user behavior simulation prediction on the advertising content to obtain prediction results, and adjust the advertising content based on the prediction results to generate interactive advertisements.

[0065] As can be seen from the above scheme, the input text data is obtained, multimodal parsing is performed on the text data to obtain text tags, 3D model features, and interaction logic, and then semantic alignment and association binding processing is performed on the text tags, 3D model features, and interaction logic to obtain rendering data. The rendering data is then subjected to device-adaptive rendering to generate advertising content, and user behavior simulation prediction is performed on the advertising content to obtain prediction results. Finally, the advertising content is adjusted based on the prediction results to generate interactive advertisements. By performing multimodal parsing on text data to generate interactive advertisements, its diversity is enriched. By using the user behavior simulation mechanism to obtain prediction results and make corresponding adjustments, the prediction effect of interactive advertisements can be obtained in a timely manner and adjustments can be made, thereby improving the generation efficiency and actual use effect of interactive advertisements.

[0066] In one possible embodiment, the parsing module 102 is specifically used for: The text data is parsed using a set language model to obtain structured policy instructions; Based on the structured strategy instructions and the set model library, corresponding text labels, 3D model features, and interaction logic are generated.

[0067] In one possible embodiment, the parsing module 102 is specifically used for: Determine the text tags corresponding to the structured strategy instructions; The visual features, sound features, and interactive features associated with the text tags are determined by setting up a visual element library, a sound effect library, and a dynamic effect library; The three-dimensional model features corresponding to the structured strategy instruction are determined based on the visual features, the sound effect features, and the set three-dimensional model library. The interaction logic corresponding to the structured strategy instruction is determined based on the interaction features and the 3D model library.

[0068] In one possible embodiment, the information processing module 103 is specifically used for: Define a virtual timeline, which includes virtual key points and interaction states; Align the text labels, the 3D model features, and the interaction logic with the virtual key points and interaction states to obtain aligned text labels, 3D model features, and interaction logic; Semantic hotspots are determined in the aligned 3D model features, and the aligned text labels and interaction logic are bound to the semantic hotspots to obtain rendering data.

[0069] In one possible embodiment, the content generation module 104 is specifically used for: Obtain the built-in device performance graph; Multiple progressively tiered advertising content are generated based on the device performance graph and the rendering data.

[0070] In one possible embodiment, the content generation module 104 is specifically used for: An advertising interaction prototype is generated based on the advertising content. The advertising interaction prototype is subjected to a simulated operation of preset user behavior, and the churn node corresponding to the simulated operation is recorded to obtain a prediction result containing at least one churn node.

[0071] In one possible embodiment, the content generation module 104 is specifically used for: In the advertising content, determine one or more adjustable elements corresponding to the lost node and the associated element adjustment values; Based on the element adjustment values, the corresponding adjustable elements in the advertisement content are adjusted to generate an interactive advertisement.

[0072] Figure 7 This is a schematic diagram of the structure of an interactive advertising generation device provided in an embodiment of this application, such as... Figure 7 As shown, the device includes a processor 201, a memory 202, an input device 203, and an output device 204; the number of processors 201 in the device can be one or more. Figure 7 Taking a processor 201 as an example; the processor 201, memory 202, input device 203, and output device 204 in the device can be connected via a bus or other means. Figure 7 Taking a bus connection as an example, the memory 202, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions or modules corresponding to an interactive advertising generation method in this embodiment. The processor 201 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 202, thereby implementing the aforementioned interactive advertising generation method. The input device 203 can be used to receive input digital or character information and generate key signal inputs related to user settings and function control of the device. The output device 204 may include a display screen or other display device.

[0073] This application also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform an interactive advertisement generation method, the method comprising: The system acquires the entered text data and performs multimodal parsing on the text data to obtain text labels, 3D model features, and interaction logic. The text labels, the 3D model features, and the interaction logic are semantically aligned and associated to obtain rendering data. The rendered data is used to generate advertising content through device-adaptive rendering, and the advertising content is used to simulate and predict user behavior to obtain prediction results. Based on the prediction results, the advertising content is adjusted to generate interactive ads.

[0074] It is worth noting that in the above embodiment of the interactive advertising generation method system, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this application.

[0075] Note that the above are merely preferred embodiments and the technical principles applied in this application. Those skilled in the art will understand that the embodiments of this application are not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the embodiments of this application. Therefore, although the embodiments of this application have been described in detail through the above embodiments, the embodiments of this application are not limited to the above embodiments. More other equivalent embodiments may be included without departing from the concept of the embodiments of this application, and the scope of the embodiments of this application is determined by the scope of the appended claims.

Claims

1. A method for generating interactive advertisements, characterized in that, include: The system acquires the entered text data and performs multimodal parsing on the text data to obtain text labels, 3D model features, and interaction logic. The text labels, the 3D model features, and the interaction logic are semantically aligned and associated to obtain rendering data. The rendered data is used to generate advertising content through device-adaptive rendering, and the advertising content is used to simulate and predict user behavior to obtain prediction results. Based on the prediction results, the advertising content is adjusted to generate interactive ads.

2. The interactive advertisement generation method according to claim 1, characterized in that, The process of performing multimodal parsing on the text data to obtain text labels, 3D model features, and interaction logic includes: The text data is parsed using a set language model to obtain structured policy instructions; Based on the structured strategy instructions and the set model library, corresponding text labels, 3D model features, and interaction logic are generated.

3. The interactive advertisement generation method according to claim 2, characterized in that, The generation of corresponding text labels, 3D model features, and interaction logic based on the structured strategy instructions and the set model library includes: Determine the text tags corresponding to the structured strategy instructions; The visual features, sound features, and interactive features associated with the text tags are determined by setting up a visual element library, a sound effect library, and a dynamic effect library; The three-dimensional model features corresponding to the structured strategy instruction are determined based on the visual features, the sound effect features, and the set three-dimensional model library. The interaction logic corresponding to the structured strategy instruction is determined based on the interaction features and the 3D model library.

4. The interactive advertisement generation method according to claim 1, characterized in that, The process of semantic alignment and association binding of the text labels, the 3D model features, and the interaction logic to obtain rendering data includes: Define a virtual timeline, which includes virtual key points and interaction states; Align the text labels, the 3D model features, and the interaction logic with the virtual key points and interaction states to obtain aligned text labels, 3D model features, and interaction logic; Semantic hotspots are determined in the aligned 3D model features, and the aligned text labels and interaction logic are bound to the semantic hotspots to obtain rendering data.

5. The interactive advertisement generation method according to any one of claims 1-4, characterized in that, The step of generating advertising content by performing device-adaptive rendering on the rendered data includes: Obtain the built-in device performance graph; Multiple progressively tiered advertising content are generated based on the device performance graph and the rendering data.

6. The interactive advertisement generation method according to any one of claims 1-4, characterized in that, The step of performing user behavior simulation prediction on the advertising content to obtain the prediction result includes: An advertising interaction prototype is generated based on the advertising content. The advertising interaction prototype is subjected to a simulated operation of preset user behavior, and the churn node corresponding to the simulated operation is recorded to obtain a prediction result containing at least one churn node.

7. The interactive advertisement generation method according to claim 6, characterized in that, The step of adjusting the advertising content based on the prediction results to generate interactive ads includes: In the advertising content, determine one or more adjustable elements corresponding to the lost node and the associated element adjustment values; Based on the element adjustment values, the corresponding adjustable elements in the advertisement content are adjusted to generate an interactive advertisement.

8. An interactive advertising generation device, characterized in that, include: The acquisition module is used to acquire the entered text data; The parsing module is used to perform multimodal parsing on the text data to obtain text labels, 3D model features, and interaction logic; The information processing module is used to perform semantic alignment and association binding on the text labels, the 3D model features, and the interaction logic to obtain rendering data; The content generation module is used to perform device-adaptive rendering on the rendering data to generate advertising content, perform user behavior simulation and prediction on the advertising content to obtain prediction results, and adjust the advertising content based on the prediction results to generate interactive advertisements.

9. An interactive advertising generation device, the device comprising: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the interactive advertising generation method as described in any one of claims 1-7.

10. A storage medium storing computer-executable instructions, which, when executed by a computer processor, are used to perform the interactive advertising generation method as described in any one of claims 1-7.