Cross-border e-commerce personalized advertisement generation system based on generative adversarial network

Through the cross-border e-commerce personalized advertising generation system based on generative adversarial networks, user information is collected and optimized in real time, which solves the problems of long production cycle and high cost in cross-border e-commerce advertising generation, and realizes efficient and personalized advertising generation and improves click-through rate.

CN120612137APending Publication Date: 2025-09-09LINGYA INTELLIGENT TECHNOLOGY (SUZHOU) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510514624.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies for cross-border e-commerce advertising generation have problems such as long production cycles, high labor costs, serious content homogeneity, and inability to achieve dynamic real-time optimization, resulting in slow advertising iteration and poor results.

Method used

A cross-border e-commerce personalized advertising generation system based on generative adversarial networks is adopted. Through the multimodal generation module, multi-dimensional discrimination module and dynamic optimization module, user information is collected in real time, advertising content is generated and optimized, and advertising features are generated using the Transformer's NLP module and StyleGAN3 architecture. Dynamic optimization is performed by combining cross-modal fusion and real-time behavioral data feedback.

Benefits of technology

Significantly improve generation efficiency, reduce production costs, enhance ad personalization and click-through rates, support multiple ad formats, and achieve rapid iteration and personalized ad generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612137A_ABST
    Figure CN120612137A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of advertisement generation, in particular to a cross-border e-commerce personalized advertisement generation system based on a generative adversarial network. The method comprises the following steps: S1, acquiring first information input by a user in real time through an acquisition module; s2, outputting a first result based on the first information through a multi-modal generation module; s3, evaluating the first result for at least one time through a multi-dimensional judgment module so as to optimize the first result and generate a second result; s4, dynamically optimizing the second result through a dynamic optimization module, and generating a third result; according to the cross-border e-commerce personalized advertisement generation system based on the generative adversarial network, cross-border e-commerce personalized advertisements are generated, so that the investigation, communication, planning and design time of users can be shortened, the generation efficiency is greatly improved, and the manufacturing cost is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of advertisement generation, and in particular to a cross-border e-commerce personalized advertisement generation system based on a generative adversarial network. Background Art

[0002] The current mainstream approach uses a linear workflow: market research takes 3-5 days, copywriting planning takes 2+ days, visual design takes 3+ days, and A / B testing takes 3+ days, for a total of 10-23 days. A single creative requires at least one planner, designer, and optimizer each. Traditional A / B testing, especially during ad iteration, requires a complete execution cycle, making dynamic, real-time optimization impossible. This results in lengthy design cycles and high labor costs.

[0003] Therefore, it is necessary to provide a new technical solution to overcome the above-mentioned defects. Summary of the Invention

[0004] The purpose of the present invention is to provide a cross-border e-commerce personalized advertising generation system based on a generative adversarial network that can effectively solve the above-mentioned technical problems.

[0005] In order to achieve the purpose of the present invention, the following technical solutions are adopted:

[0006] S1. Collecting first information input by a user in real time through a collection module;

[0007] S2. Outputting a first result based on the first information through a multimodal generation module;

[0008] S3. Evaluate the first result at least once using a multi-dimensional discrimination module to optimize the first result and generate a second result;

[0009] S4. Dynamically optimize the second result through a dynamic optimization module to generate a third result.

[0010] Furthermore, the multimodal generation module includes:

[0011] A first advertisement feature generating unit, which generates a first advertisement feature based on a Transformer NLP module;

[0012] A second advertisement feature generation unit, based on the CV module of the StyleGAN3 architecture, generates a second advertisement feature;

[0013] The cross-modal fusion unit uses a cross-modal fusion network to optimize the matching of the generated first advertisement feature and the second advertisement feature, and outputs a first result.

[0014] Furthermore, the multi-dimensional discrimination module includes:

[0015] User profile matching evaluation unit, which evaluates the compatibility between generated content and the target customer profile based on at least 200 user characteristics;

[0016] The business value prediction unit predicts click-through rate and conversion rate using trained click-through rate and conversion rate prediction models;

[0017] The aesthetic quality assessment unit uses the Noise-Contrastive Estimation method to evaluate the aesthetic score of graphic content.

[0018] Furthermore, step S4 includes:

[0019] Build a Kafka+Flink streaming computing architecture to continuously collect real-time user behavior data and feed the real-time behavior data back to the multimodal generation module.

[0020] Furthermore, the first information includes: product features, target customer group portraits, and style parameters;

[0021] The product features include: SPU information, product description, specifications, price, category, and image feature vector;

[0022] The style parameters include: style parameters set by the user in the system and style parameters automatically recommended by the system.

[0023] Furthermore, the method for constructing the target customer group portrait includes: collecting basic information from the target group characteristics entered by the user to generate a preliminary user portrait; and matching relevant characteristic data of the same type of target customer groups in the user database to improve the preliminary user portrait.

[0024] Furthermore, the user database is constructed based on the target customer group and includes:

[0025] Clean the original cross-border user dataset and extract cross-border e-commerce-specific features such as multilingual preferences, cross-border logistics selection records, and tariff sensitivity.

[0026] Convert the user's local purchase period to the target market time zone;

[0027] Feature engineering technology is used to transform the cleaned attribute information into a multi-dimensional feature vector to quantitatively capture users' sensitivity to international payment methods and shipping time.

[0028] Build a user relationship graph and define the criteria for determining adjacency relationships in cross-border e-commerce scenarios;

[0029] A multilingual semantic alignment model is used to map heterogeneous language tags into a unified feature space.

[0030] Furthermore, the first advertisement feature is text; and the second advertisement feature is a graphic.

[0031] Furthermore, the second advertisement feature generation unit uses a pre-trained visual feature extraction network to extract high-dimensional image feature vectors from product images.

[0032] Furthermore, the cross-modal fusion unit adopts an attention mechanism to achieve information alignment and fusion between the first advertising feature and the second advertising feature.

[0033] Compared with the existing technology, the present invention has the following beneficial effects: the cross-border e-commerce personalized advertising generation system based on generative adversarial networks of the present invention generates cross-border e-commerce personalized advertisements through the cross-border e-commerce personalized advertising generation system based on generative adversarial networks of this application, which can reduce the user's research, communication, planning, and design time, greatly improve the generation efficiency, and greatly reduce the production cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0035] Figure 1 This is a flow chart of the cross-border e-commerce personalized advertising generation system based on generative adversarial networks of the present invention;

[0036] Figure 2 This is a flow chart of the multimodal generation module of the cross-border e-commerce personalized advertising generation system based on generative adversarial networks of the present invention;

[0037] Figure 3 This is a flowchart of the multi-dimensional discrimination module of the cross-border e-commerce personalized advertising generation system based on the generative adversarial network of the present invention. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments.

[0039] It should be understood that, although the various steps in the flow chart of each embodiment of the present invention are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence according to the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in order, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0040] like Figures 1 to 3 As shown, the cross-border e-commerce personalized advertising generation system based on the generative adversarial network of the present invention includes:

[0041] S1. The collection module collects first information input by the user at the user terminal in real time; the user terminal includes: a mobile phone, a tablet, a computer, etc.

[0042] S2. Outputting a first result based on the user's first information through a multimodal generation module, where the multimodal generation module is in communication with the acquisition module.

[0043] S3. Evaluate the first result at least once through the multi-dimensional discrimination module to optimize the first result and generate a second result; the multi-dimensional discrimination module is communicatively connected to the multimodal generation module.

[0044] S4. Dynamically optimize the second result through a dynamic optimization module to generate a third result.

[0045] The current industry largely relies on manual creative production. A single ad production requires collaborative efforts across multiple stages, including market research, copywriting, and visual design. This takes an average of 7-15 days and costs as much as $500-5,000, severely limiting the speed of ad iteration. For example, the manual creative production process requires 3-5 days of market research, 2 days of copywriting, 3 days of visual design, and 3 days of A / B testing before a campaign can be launched. While some existing automated tools attempt to improve efficiency, they are limited by template-based generation systems and single-modal content production techniques, enabling only basic parameter adjustments like color replacement. This leads to significant homogeneity in ad content, with image-text matching rates below 60% and style options limited to five basic types, making it difficult to meet the needs of over 200 market segments. This ultimately results in an industry-wide average click-through rate (CTR) consistently below 1.2% and a conversion rate below 0.5%. These tools, for example, rely on product databases, pre-set template libraries (approximately 500 templates), parametric color or text replacement, batch generation, or the use of single-modal generation architectures like CNNs or RNNs, which only support single-content generation, such as text or images.

[0046] These technical shortcomings create a vicious cycle: manual intervention drives up production costs, forcing companies to adopt inefficient templates. The rigidity of these templates exacerbates content homogeneity, leading to lower conversion rates for low-quality ads, which in turn incentivizes more frequent manual optimization efforts. Existing solutions face technical limitations in efficiency, personalization, and real-time performance, severely hindering the precise delivery and conversion effectiveness of cross-border e-commerce advertising.

[0047] Generating cross-border e-commerce personalized advertisements through the cross-border e-commerce personalized advertisement generation system based on generative adversarial networks in this application can reduce production costs by 92%, and the single cost is less than US$40; the generation efficiency is increased by 300 times, and the average generation time is less than 15 seconds; the supported advertising formats are expanded to 8 types, such as: pictures and texts, short videos, 3D displays, etc.; the click-through rate of the target customer group is increased by 210%, which is 60 times higher than the self-optimization click-through rate of the traditional system model.

[0048] In step S1, the first information is the user's design requirements, which include: product features, target customer group portrait, and style parameters;

[0049] Product features include: SPU information, product description, specifications, price, category, image feature vector, etc.

[0050] The method for constructing a target customer group portrait includes: collecting basic information from the target group characteristics entered by the user, such as age, gender, location, user age, consumption preferences, behavioral history, device characteristics, etc., to generate a preliminary user portrait; at the same time, matching the relevant characteristic data of the same type of target customer groups in the user database to improve and refine the preliminary user portrait.

[0051] Style parameters include: style parameters set by the user in the system and style parameters automatically recommended by the system.

[0052] The user database is built based on the target customer group. The specific process is as follows:

[0053] Clean the original cross-border user dataset to extract cross-border e-commerce-specific features such as multilingual preferences, cross-border logistics selection records, and tariff sensitivity.

[0054] Convert the user's local purchase period to the target market time zone to ensure consistent time zone information;

[0055] Feature engineering techniques are used to transform the cleaned attribute information into a multidimensional feature vector. The vector contains indicators such as currency preference coding and logistics service weighting to quantitatively capture users' sensitivity to international payment methods and shipping time.

[0056] Build a user relationship graph and define the criteria for determining adjacency relationships in cross-border e-commerce scenarios, including: participating in cross-border shopping communities together; purchasing the same overseas niche brand products; and interacting with product reviews in languages ​​other than the platform's official language. A user adjacency matrix is ​​constructed based on these adjacency relationships and processed using an improved graph attention network. The graph attention network introduces bias weight parameters based on the user's country of residence and dynamically adjusts the attention weights between users from different countries based on historical consumption data. For example, in the beauty and cosmetics category, the interaction weights between Southeast Asian users and European and American users will be automatically differentiated.

[0057] In response to the unique cross-platform social behavior of cross-border e-commerce users, the system obtains the public label data of social platforms entered by users, desensitizes the data, and then uses a multilingual semantic alignment model to map heterogeneous language labels to a unified feature space, and uses sparse matrix decomposition technology under the federated learning framework to achieve feature dimensionality reduction.

[0058] In step S2, the multimodal generation module includes: a first advertising feature generation unit, which generates a first advertising feature based on the Transformer NLP module, and the first advertising feature is personalized advertising copy or description for cross-border e-commerce. The personalized advertising copy or description supports 12 mainstream national languages; unsupervised pre-training is performed on a large-scale multilingual corpus, an autoregressive prediction task is adopted, the objective function is cross-entropy loss, and supervised fine-tuning is performed using an annotated product copy dataset to optimize the accuracy and fluency of personalized advertising copy or description generation. A dynamic beam width adjustment strategy is adopted, and Beam Search with a beam width of K=8 is used in the initial generation stage to ensure core semantic accuracy. When the generated length exceeds 15 tokens, Top-p sampling (p=0.9) is switched to ensure diversity. Finally, the best 20-50 outputs are selected from the multiple generated candidate copies through a diversity sorting algorithm to ensure language diversity and semantic coherence. The model parameters are adjusted using the Adam or AdamW optimizer, and the learning rate decay strategy is combined to ensure training stability and efficient convergence.

[0059] The second advertising feature generation unit, based on the CV module of the StyleGAN3 architecture, constructs a generation network including a generator according to the style preferences in the user portrait to generate the second advertising feature; this solution solves the texture stickiness problem, is suitable for dynamic content generation, and supports high-quality image generation; the second advertising feature generation unit uses a pre-trained visual feature extraction network (such as CNN) to extract high-dimensional image feature vectors from product images, which is suitable for image generation tasks. The adversarial mechanism between the discriminator and the generator can improve the generation quality. Pre-training avoids training from scratch, accelerates convergence and reduces data dependence; in this embodiment, layer-by-layer convolutional upsampling can also be used to generate high-resolution images and embed a style modulation mechanism; the style control module maps the style information to a multi-dimensional latent space and injects style features into each convolutional layer through a style transfer network. The multi-dimensionality includes 128 dimensions, 512 dimensions, etc. The use of 512 dimensions improves style diversity, makes the transition smoother, and requires higher computing power;

[0060] Cross-modal fusion unit: Utilize the cross-modal fusion network to optimize the matching between the generated first advertising feature and the second advertising feature, and realize the information alignment and fusion between the first advertising feature and the second advertising feature through the attention mechanism, so that the content of the generated first advertising feature and the second advertising feature matches in style and semantics, and generates a first result; alignment and fusion include: using the CLIP model to calculate the image and text matching degree and automatically correct it, for example, replacing conflicting elements, adjusting color matching, etc.; the cross-modal fusion unit preferably adopts a multi-head attention mechanism to interactively encode image and text features, so that the two complement each other semantically, thereby improving the consistency of the overall content.

[0061] In step S3, the multi-dimensional discrimination module optimizes the first result and generates a second result including:

[0062] The second result is a candidate result output after adversarial optimization based on the first result;

[0063] User Profile Matching Assessment Unit: This unit assesses the compatibility of generated content with the target customer profile based on at least 200 user characteristics, such as age, gender, location, consumption preferences, behavioral history, and device information.

[0064] Business Value Prediction Unit: This unit uses trained click-through rate and conversion rate prediction models to evaluate the effectiveness of generated content in commercial applications. A dataset is generated based on the target customer's user profile and the first click-through rate. This dataset is then input into the trained click-through rate and conversion rate prediction models to obtain a predicted conversion rate.

[0065] Filtering unit: Automatically screens and detects taboo patterns, such as the prohibition of cow elements in the Indian market, alcohol logos in the Middle East, and politically sensitive symbols in Europe and the United States. The filtering unit is incrementally trained once a day and fully updated once a week to ensure the latest data.

[0066] Aesthetic Quality Assessment Unit: Utilizes the Noise-Contrast Estimation method to construct a two-dimensional evaluation system to assess the aesthetics of graphic and text content, ensuring that the content is visually and textually appealing.

[0067] The aesthetic quality assessment unit is assessed using the following methods:

[0068] By calculating the color distribution entropy (preferably 0.78-1.2) based on the principle of contrast and complementarity, and combining the brand's main color coverage (≥65%) and the logo significance coefficient (F value > 0.4), a dynamic balance model of visual coordination and brand consistency was established.

[0069] YOLOv7 was used to segment the product area (accounting ≥40%), the promotion layer (25±5%), and the background layer. Eye movement experiment data was combined to construct a gaze path heat map and optimize the spatial weight distribution of core elements.

[0070] The CTR-CVR correlation features extracted from the high-conversion sample library (such as the layout grid golden ratio of 0.618 and the use of positive emotional fonts > 82%) are linked with the commercial value prediction unit to establish a multimodal feature vector:

[0071] The marginal contribution of each feature to the conversion rate is quantified through the gradient boosting tree, driving the generative model to automatically optimize key parameters such as the brand element implantation intensity (15% to 25%) and the promotion information hierarchy (Z-index 3-5).

[0072] The 200+ dimensional features output by the user portrait matching evaluation unit are cross-modally fused with the aesthetic feature vector through the Attention mechanism to achieve personalized aesthetic adaptation: for example: for the youth group: inject high-saturation colors (HSV-V ≥ 80%) and dynamic visual elements (frame change rate 0.5-2Hz); for business users: adopt a modular grid layout (8 / 12 column system) and increase the proportion of professional fonts (serif > 70%).

[0073] In step S4, the second result is dynamically optimized to generate a third result, including: multi-dimensional aesthetic evaluation of candidate advertising plans, automatic generation of an optimization report containing color fine-tuning suggestions, layout reconstruction plans and element replacement strategies, supporting advertising designers to achieve Pareto optimality of aesthetic expression and business goals while maintaining brand genes; specifically, it includes: building a Kafka+Flink streaming computing architecture, using a data pipeline with a delay of less than 500ms, continuously collecting real-time user behavior data, and feeding the extracted real-time behavior data back to the multimodal generation module; real-time behavior data includes: user feedback, click popularity, stay time, etc. ; Build a streaming computing architecture of Kafka+Flink, use a data pipeline with a delay of less than 500ms, feed back real-time behavior data such as user clicks and interactions to the multimodal generation module, and feed back the extracted real-time behavior data to the multimodal generation module; the incremental learning framework model update cycle is less than 2 hours, which can quickly respond to changes in the market and user behavior, realize online learning and strategy adjustment, enable the system to adapt to new data in a short time, and improve the quality and matching of generated content by continuously updating the generation strategy; among them, the third result is the result generated after optimization based on the second result, specifically in the form of pictures, texts, short videos, 3D displays or a combination thereof.

[0074] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0075] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0076] It should be understood that those skilled in the art can make improvements or changes based on the above description, and all such improvements and changes should fall within the scope of protection of the appended claims of the present invention.

Claims

1. A cross-border e-commerce personalized advertising generation system based on generative adversarial networks, characterized by: S1. Collecting first information input by a user in real time through a collection module; S2. Outputting a first result based on the first information through a multimodal generation module; S3. Evaluate the first result at least once using a multi-dimensional discrimination module to optimize the first result and generate a second result; S4. Dynamically optimize the second result through a dynamic optimization module to generate a third result.

2. The cross-border e-commerce personalized advertising generation system based on a generative adversarial network according to claim 1, characterized in that: The multimodal generation module includes: A first advertisement feature generating unit, which generates a first advertisement feature based on a Transformer NLP module; A second advertisement feature generation unit, based on the CV module of the StyleGAN3 architecture, generates a second advertisement feature; The cross-modal fusion unit uses a cross-modal fusion network to optimize the matching of the generated first advertisement feature and the second advertisement feature, and outputs a first result.

3. The cross-border e-commerce personalized advertising generation system based on a generative adversarial network according to claim 1, characterized in that: The multi-dimensional discrimination module includes: User profile matching evaluation unit, which evaluates the compatibility between generated content and the target customer profile based on at least 200 user characteristics; The business value prediction unit predicts click-through rate and conversion rate using trained click-through rate and conversion rate prediction models; The aesthetic quality assessment unit uses the Noise-Contrastive Estimation method to evaluate the aesthetic score of graphic content.

4. The cross-border e-commerce personalized advertising generation system based on a generative adversarial network according to claim 1, characterized in that: Step S4 includes: Build a Kafka+Flink streaming computing architecture to continuously collect real-time user behavior data and feed the real-time behavior data back to the multimodal generation module.

5. The cross-border e-commerce personalized advertising generation system based on a generative adversarial network according to claim 1, characterized in that: The first information includes: product features, target customer group portraits, and style parameters; The product features include: SPU information, product description, specifications, price, category, and image feature vector; The style parameters include: style parameters set by the user in the system and style parameters automatically recommended by the system.

6. The cross-border e-commerce personalized advertising generation system based on a generative adversarial network according to claim 5, characterized in that: The method for constructing the target customer group portrait includes: collecting basic information from the target group characteristics entered by the user to generate a preliminary user portrait; and matching relevant characteristic data of the same type of target customer groups in the user database to improve the preliminary user portrait.

7. The cross-border e-commerce personalized advertising generation system based on a generative adversarial network according to claim 6, characterized in that: The user database is constructed based on the target customer group and includes: Clean the original cross-border user dataset and extract cross-border e-commerce-specific features such as multilingual preferences, cross-border logistics selection records, and tariff sensitivity. Convert the user's local purchase period to the target market time zone; Feature engineering technology is used to transform the cleaned attribute information into a multi-dimensional feature vector to quantitatively capture users' sensitivity to international payment methods and shipping time. Build a user relationship graph and define the criteria for determining adjacency relationships in cross-border e-commerce scenarios; A multilingual semantic alignment model is used to map heterogeneous language tags into a unified feature space.

8. The cross-border e-commerce personalized advertising generation system based on a generative adversarial network according to claim 2, characterized in that: The first advertisement feature is text; the second advertisement feature is a graphic.

9. The cross-border e-commerce personalized advertising generation system based on a generative adversarial network according to claim 2, characterized in that: The second advertisement feature generation unit uses a pre-trained visual feature extraction network to extract a high-dimensional image feature vector from the product image.

10. The cross-border e-commerce personalized advertising generation system based on generative adversarial networks according to claim 2, characterized in that: The cross-modal fusion unit uses an attention mechanism to achieve information alignment and fusion between the first advertising feature and the second advertising feature.

Citation Information

Cited By

  • Advertisement content generation method and device, electronic equipment and readable storage medium

    CN121352881A

  • Advertisement push picture generation method and system based on style migration

    CN121353454A