Product copywriting generation method based on conditional variational autoencoder and retrieval enhancement

Through the combination of conditional variational automatic encoder and graph neural network, the diversity and personalization problems of the general large language model when generating product copywriting in specific fields are solved, and the product copy generation with clear style and diverse styles is achieved, which improves the generation efficiency and copy quality.

CN120297243BActive Publication Date: 2025-08-19NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510787349.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-08-19
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

When generating product copywriting in a specific field, the existing general language model is difficult to meet the description of multiple product attributes at the same time, and the generated copywriting lacks diversity and personalization, and cannot adapt to the needs of complex semantic expressions.

Method used

Using a method based on conditional variational autoencoder and graph neural network, candidate inference paths are generated through multi-hop graph inference, combined with conditional variational autoencoder and ON-LSTM gated hierarchical structure, diverse product copywriting is generated, and through iterative optimization of evaluation modules and thinking tree, copywriting that meets user needs is finally generated.

Benefits of technology

It realizes clear and diverse product copywriting generation, improves the expression ability and generation efficiency of the model, and ensures the relevance and diversity of copywriting and user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297243B_ABST
    Figure CN120297243B_ABST
Patent Text Reader

Abstract

The present invention discloses a product copy generation method based on conditional variational autoencoder and retrieval enhancement. The method designs a reasoning module, a generation module, an evaluation module, a memory module, and a prompt engineering module. Based on a graph neural network, the method adopts path pruning and multi-hop reasoning for user input content. In the reasoning module, path reasoning is used to prune low-scoring paths, improve the efficiency of reasoning, and reason about implicit relationships. The generation module introduces a layer-by-layer variational autoencoder, and through conditional variables, the generated candidate product copy can meet specific conditional requirements. The evaluation module comprehensively scores the candidate product copy, constructs a memory module to store the dialogue between the user and the generation module, and the prompt engineering module performs preset rounds of iteration based on a thinking tree structure to obtain the final product copy. The method designed by the present invention can enhance the personalization and authenticity of the copy, and realize the generation of clear and diverse advertising copy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence large models, and in particular to a method for generating product copy based on conditional variational autoencoders and retrieval enhancement. Background Art

[0002] In the fields of artificial intelligence and natural language processing (NLP), with the rise of deep learning and large language models (LLMs) such as GPT, BERT, and T5, copywriting generation has become a key technology for content creation, advertising marketing, and personalized recommendations. Traditional copywriting generation methods primarily rely on template-based and rule-driven strategies, such as text synthesis based on predefined sentence structures. While these methods can generate structured text in specific scenarios, their flexibility and diversity are limited, making them difficult to adapt to complex semantic expressions and personalized needs.

[0003] However, despite LLM's excellent performance in open-domain text generation, it still faces challenges in generating high-quality copy for specific domains, such as product advertisements. General-purpose large models lack vertical coverage (such as specific product categories and industry terminology), making it difficult to simultaneously describe multiple product attributes. Furthermore, the emphasis on each attribute varies, requiring specific emphasis based on user needs. While ensuring relevance, the generated copy is diverse and scalable. Although fine-tuned general-purpose large models have the ability to generate text for specific domains, the rapid pace of product updates and iterations means that some combinations are rare or untrained, necessitating the introduction of external knowledge bases to supplement these. Summary of the Invention

[0004] The purpose of the present invention is to provide a product copy generation method based on conditional variational autoencoder and retrieval enhancement, which can realize the generation of product copy with clear style and diversity, and enhance the product copy generation capability in vertical fields.

[0005] To achieve the above functions, the present invention designs a product copy generation method based on conditional variational autoencoder and retrieval enhancement, and executes steps S1 to S6 to complete the generation of product copy:

[0006] Step S1: extract keywords from the product copy requirement description entered by the user;

[0007] Step S2: Using keywords as input, a graph neural network model-based reasoning module is constructed. Multi-hop graph reasoning is used to generate candidate reasoning paths. The candidate reasoning paths undergo path embedding, path pruning, and score calculation. Several candidate reasoning paths are selected as output paths based on the scores, and implicit keywords are extracted and output.

[0008] Step S3: Taking keywords and hidden keywords as input, a generation module is constructed, consisting of an embedding layer, an encoder, a reparameterization module, a decoder, and an output layer. The generation module integrates the conditional variational autoencoder structure and the ON-LSTM gated hierarchical structure to output several candidate product copywriting;

[0009] Step S4: Building an evaluation module based on the large language model, comprehensively scoring the candidate product copywriting, and selecting a product copywriting from the candidate product copywriting based on the comprehensive score as the product copywriting output by the evaluation module;

[0010] Step S5: Construct a memory module to store the user input content and the generated product copy;

[0011] Step S6: Based on the thinking tree structure with a preset number of layers, a prompt engineering module is constructed, and the memory module, reasoning module, generation module and evaluation module are called to perform preset rounds of iteration to obtain the final product copy.

[0012] Beneficial effects: Compared with the prior art, the advantages of the present invention include:

[0013] 1. The reasoning module proposed in this paper learns node representations in a graph structure through an information propagation mechanism. It captures the relationships between entities through adjacency matrices and relationship features, and enhances the model's expressiveness through multi-layer reasoning, enabling graph neural networks to effectively perform multi-hop reasoning. Path pruning reduces model training time while maintaining existing retrieval accuracy.

[0014] 2. The generation module proposed in this invention passes its input through the encoder, into the latent space, and then into the decoder of the ON-LSTM gating hierarchy. In this process, a latent vector is learned and input into different ON-LSTM gating hierarchies with different weights. The more important attributes are located at higher levels of the ON-LSTM gating hierarchy and will be retained longer. The latent vector also contains the newly added information, achieving controllable generation while diversifying the generated text.

[0015] 3. The evaluation module based on the thinking tree proposed in this invention uses the thinking tree strategy to realize iterative copy generation, making copy generation more diversified. The Top-K strategy of the voting mechanism retains K optimal paths, combines the greedy algorithm with random sampling, balances local optimality and global optimality, and avoids excessive restrictions on copy generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 Flowchart of a method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to an embodiment of the present invention;

[0017] Figure 2is a flow chart of a reasoning module provided according to an embodiment of the present invention;

[0018] Figure 3 is a schematic diagram of a generation module provided according to an embodiment of the present invention;

[0019] Figure 4 Schematic diagram of an evaluation module and a thinking tree structure provided according to an embodiment of the present invention. DETAILED DESCRIPTION

[0020] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.

[0021] The method for generating product copy based on conditional variational autoencoder and retrieval enhancement provided by the embodiment of the present invention is as follows: Figure 1 , execute steps S1 to S6 to complete the generation of product copy:

[0022] Step S1: extract keywords from the product description entered by the user using an existing large language model;

[0023] In an embodiment, the system receives a product copy requirement description input by a user, for example: "I want to generate a product advertisement copy for a white top, made of denim, with embroidery, holes, and a simple style." Using an existing pre-trained large language model (such as a Transformer structure), the system performs entity recognition and intent extraction on the input content, identifying key entity items such as: "top, denim, white, simple, embroidery, holes." Subsequently, the entity items are aligned with standard entities in the knowledge graph to construct an initial triple structure, such as: ("top", "material", "denim") ("top", "color", "white") ("top", "style", "simple") ("top", "style", "embroidery") ("jacket", "style", "holes").

[0024] Step S2: Using keywords as input, a graph neural network model-based reasoning module is constructed. Multi-hop graph reasoning is used to generate candidate reasoning paths. The candidate reasoning paths undergo path embedding, path pruning, and score calculation. Several candidate reasoning paths are selected as output paths based on the scores, and implicit keywords are extracted and output.

[0025] Reference Figure 2 The reasoning module is used to obtain the product description information entered by the user and obtain relevant attributes through database retrieval and multi-hop graph reasoning. The reasoning module is based on the GNN PEP-GNN (Path Encoding and Pruning) model, which combines the instruction-driven reasoning mechanism with graph neural networks to achieve efficient knowledge reasoning.

[0026] The specific steps of step S2 are as follows:

[0027] Step S2.1: Load a pre-trained graph neural network model, including graph convolution kernels, attention weight matrices, and historical path memory parameters. The graph neural network model is trained on a large-scale corpus of product copywriting and domain knowledge graphs to enhance its domain adaptability and reasoning capabilities.

[0028] Assume that the knowledge graph is ,in is a collection of entities, is a set of edges, each edge represents a semantic relationship between two entities;

[0029] Setting the inference depth parameter ,implement Round multi-hop graph reasoning, in the knowledge graph i In layer propagation, entity nodes Representation By its neighbor nodes The information is obtained through weighted aggregation, where Representation node The set of neighbor nodes of node The set of all connected nodes; the formula for weighted information aggregation is as follows:

[0030] ;

[0031] in, For entity nodes the expression; Represents the nodes in the adjacency matrix i and nodes j connectivity relationship; Represents a slave node i To Node j relational embedding; is the neighbor aggregation weight matrix; is the relationship weight matrix; is the activation function;

[0032] The propagation of each hop is similar to the diffusion of a signal in a knowledge graph, thus constructing a multi-hop entity relationship link.

[0033] In this embodiment, the inference depth parameter is set , performing three rounds of information propagation in the knowledge graph. In each hop, based on the current entity representation and the knowledge graph structure, the neighboring node information is aggregated and the attention score is calculated. The reasoning process is as follows:

[0034] Jump 1: Starting with "top", aggregate the first-level attributes such as "denim", "white", and "simple";

[0035] Jump 2: Jump from "white" to its semantically related words, such as "clean", "elegant", etc.

[0036] Jump 3: Expand from "embroidery" to abstract descriptions such as "exquisite", "feminine", and "personalized design".

[0037] Each hop propagation updates the node status and records the path propagation trajectory and its score, and finally generates a set of candidate reasoning paths.

[0038] Step S2.2: To support path-level reasoning and selection, it is necessary to construct node representations and vectorize the entire path. Start, go through multiple intermediate entities until the target entity Path ;

[0039] To effectively model the reasoning path, a path embedding is constructed for each path. The path embedding vector is used to represent the semantic chain formed from the starting entity, through multiple intermediate entities and relationships, to the target node. The path is modeled using the Concat splicing method:

[0040] ;

[0041] in, is the path embedding vector obtained by modeling the reasoning path, The first i The embedding representation of an entity, Represents a splicing operation. This approach maintains the semantic order of entities in the path and can provide finer-grained control signals for subsequent pruning and generation;

[0042] Use Concat to concatenate modeled paths Expressed as:

[0043] ;

[0044] in, 、 is an intermediate entity, express and relationship, express and relationship, Lis the total number of entities; this method preserves semantics and order, which helps conditional information serve as input in subsequent pruning and scoring, and accurately guides the style and details of product copywriting.

[0045] Step S2.3: In order to effectively filter out the path with the highest relevance to the input query in the high-dimensional path space, for each path , perform path pruning, including static pruning and semantic pruning;

[0046] Static pruning for each path Calculating path scores , to measure its information content and stability:

[0047] ;

[0048] in, For the i entities, L is the total number of entities; Represents the L2 norm, which is extremely sensitive to the embedding dimension and can effectively filter out paths with severe semantic drift;

[0049] Path scoring can identify semantic drift or noise paths and set the threshold , only keep Path;

[0050] Semantic pruning for each path and the user input query vector Calculate cosine similarity :

[0051] ;

[0052] Keep only path to ensure the semantic relevance and accuracy of the reasoning results.

[0053] Step S2.4: Sort the paths according to their scores, starting with the path with the highest score, and select several paths in the following format: , extract tail entities, which contain implicit keywords;

[0054] In the embodiment, the above process is expressed as:

[0055] Reasoning path: top → material → denim → style → holes → style → personality;

[0056] Extract implicit keywords: "personality display", "denim", "ripped design".

[0057] Step S3: Taking keywords and hidden keywords as input, a generation module is constructed, consisting of an embedding layer, an encoder, a reparameterization module, a decoder, and an output layer. The generation module integrates the Conditional Variational Autoencoder (CVAE) structure and the ON-LSTM gated hierarchical structure to output several candidate product copywriting.

[0058] Reference Figure 3 , the specific steps of step S3 are as follows:

[0059] Step S3.1: The embedding layer converts word IDs into dense vectors using a word embedding function; adjusts the dimensions of the dense vectors, which include sequence length, batch size, and hidden layer dimensions; and feeds the embedded vectors into the encoder.

[0060] The input to the generation module consists of two parts: Keywords (from the product knowledge graph): top, denim, white, simple, embroidered, jacket, ripped; and implicit keywords (from the three-hop inference path of the graph neural network): clean, elegant, feminine, and personalized design. These keywords are concatenated with the implicit keywords and used as the input to the generation module. This input is then fed into the tokenizer for processing. The tokenizer uses a pretrained dictionary to segment the input text into semantic units and encodes them into vector tensors, which serve as the encoder input.

[0061] Step S3.2: The encoder adopts the conditional variational autoencoder structure, and the formula is as follows:

[0062] ;

[0063] ;

[0064] in, is the standard deviation of the potential distribution, which is calculated using the sub-neural network layer logvar_layer in the conditional variational autoencoder structure to control the degree of discreteness of the potential distribution. is the mean of the potential distribution, which is calculated using the sub-neural network layer mu_layer in the conditional variational autoencoder structure to determine the center position in the latent space. is hidden state, is the weight matrix used to transform the hidden state of the encoder Mapping to the mean of the underlying distribution , is the bias vector;

[0065] Step S3.3: The reparameterization module samples the latent vector from the Gaussian distribution :

[0066] ;

[0067] in, represents noise sampled from a standard normal distribution, represents element-wise multiplication, is the standard deviation of the underlying distribution, is the mean of the underlying distribution, represents a normal distribution;

[0068] The low-dimensional latent vector thus sampled , depends on the input keyword condition information.

[0069] Step S3.4: The decoder converts the latent vector Expand to the target sequence length and pass the latent vector Transmitted to different layers of the ON-LSTM gating hierarchy to generate hidden states and generate token distributions;

[0070] The decoder is based on the ON-LSTM model, which adopts a hierarchical structure and uses potential vectors The design of this structure enables the generative model to flexibly capture semantic information at different levels and effectively separate details from themes.

[0071] The core design idea of the decoder is to transform the latent vector Injected into different layers of the ON-LSTM model to better model the long-term and short-term information of the copy. The decoder is composed of multiple layers of ON-LSTM stacked together with the latent vector The ON-LSTM model itself dynamically models the allocation of long-term and short-term memory through a cumulative softmax mechanism (cumax). High-level layers are responsible for topic semantics. These layers (shallow layers near the input) primarily receive subspace features in the latent vector related to topic and style, controlling overall semantic consistency. Lower layers are responsible for local details. These layers (deeper layers near the output) receive subspace features in the latent vector related to fine-grained details (such as attributes and modifiers), flexibly adjusting the generation of textual details. Therefore, the layered injection of latent vectors is not simply a monolithic addition to all layers; rather, it is selective, proportional, and modular.

[0072] In order to make the latent vector It affects the copywriting process without destroying the sentence structure and main information. It adopts a layered injection mechanism to inject latent vectors into the text. By gradually injecting the latent variables into the ON-LSTM model at different levels, this method can effectively control the influence of latent variables on different levels, allowing the model to flexibly generate text that meets the control conditions while ensuring semantic consistency.

[0073] In order to make the latent vector To control the text generation process, the gating mechanism has been improved. The gating mechanism in the ON-LSTM model consists of a master forget gate, a master input gate, and candidate memory. Through these gating mechanisms, the model can flexibly forget and update information at different levels, ensuring the controllability of the generation process.

[0074] Assume that the ON-LSTM model contains L layers, and each layer index is In order to achieve hierarchical control of potential information, the potential vector corresponding to each layer is defined Transformation:

[0075] ;

[0076] Where, For the The multi-layer perceptron is a specialized learning layer that is responsible for transforming the original latent vector Extract the subspace information related to the current level .

[0077] In the gating calculation of each layer, For example, the latent variable They are injected into the main forget gate, main input gate and conventional gate respectively, and the formula is as follows:

[0078] The Master Forget Gate uses the cumax activation function to generate a progressive forgetting weight from 0 to 1 by accumulating softmax to control the historical cell state. The retention ratio is as follows:

[0079] ;

[0080] Where, Represents the main forget gate, 、 、 Indicates that the main forget gate corresponds to the current input , upper hidden state and the latent vector The weight matrix of .

[0081] The Master Input Gate forms a complementary relationship with the Forget Gate, and the formula is as follows:

[0082] ;

[0083] Where, represents the main input gate, 、 、 Indicates that the main input gates correspond to the current input , upper hidden state and the latent vector The weight matrix of

[0084] In addition, in order to fuse the main gate and the regular gate, a reconstruction gate mechanism is introduced.

[0085] Conventional gates are the input gate and forget gate in standard LSTM, and their definitions are as follows:

[0086] Input Gate:

[0087] ;

[0088] Where, represents the input gate, 、 Indicates that the input gate corresponds to the current input , upper hidden state The weight matrix, is the learnable offset.

[0089] Forget Gate:

[0090] ;

[0091] Where, represents the forget gate, 、 Indicates that the forget gate corresponds to the current input , upper hidden state The weight matrix, is the learnable offset.

[0092] Output gate:

[0093] ;

[0094] Where, represents the output gate, 、 Indicates that the output gate corresponds to the current input , upper hidden state The weight matrix, is the learnable offset.

[0095] Reconstructing the forget gate Defined as:

[0096] ;

[0097] Reconstructing the input gate Defined as:

[0098] ;

[0099] in Represents element-wise multiplication.

[0100] Candidate memory update: At each time step, the candidate memory integrates the current input , upper hidden state and the latent vector , the tanh activation function constrains the value to the range [-1, 1] to ensure stable gradient flow, and introduces a learnable offset when calculating linear transformations Candidate memory The update formula is as follows:

[0101] ;

[0102] Where, 、 、 Indicates that the candidate memories correspond to the current input , upper hidden state and the latent vector The weight matrix, is the learnable offset.

[0103] Finally, the cell state and hidden state The update is done by the following formula:

[0104] ;

[0105] ;

[0106] in, Is the output gate, which determines the current cell state Through this mechanism, the model can effectively control the long-term dependencies and short-term details of information when generating text.

[0107] This hierarchical latent variable injection mechanism enables high-level neurons to stably maintain the main theme of the text, and low-level neurons to flexibly fine-tune details, achieving diversity and consistency in copywriting style and content.

[0108] The decoder generates layers according to the following process:

[0109] 1. According to the latent vector Generate potential vectors at each layer through different MLP modules ;

[0110] 2. In each layer of ON-LSTM, the potential vector Inject each gate calculation separately;

[0111] 3. Pass the hidden state and cell state down layer by layer until the final text sequence is output.

[0112] Through layered injection, the model achieves the organic coordination of semantics of different granularities in the generation process, enhancing the naturalness and controllability of the generated copy.

[0113] In this embodiment, the generation module generates three candidate advertising copy for the input keyword set: "top, denim, white, simple, embroidery, jacket, ripped", as shown below:

[0114] Copy 1: This stylish white denim jacket, with its simple and elegant white color, highlights your personality and charm. Paired with ripped jeans, it's a truly stylish and alluring jacket. The delicate embroidery accentuates a woman's elegance. The denim fabric is comfortable and skin-friendly, making it a comfortable and relaxed fit.

[0115] Copywriting 2: This white embroidered denim jacket combines fashion and comfort. The ripped holes enhance the street feel, and the simple design highlights the high-end feel, creating a unique fashion charm.

[0116] Copywriting 3: This white top is full of personality, with a design that combines ripped holes with embroidery. The denim material is soft and durable, combining practicality and beauty, making it a versatile item for spring and autumn.

[0117] Step S3.5: The output layer uses a fully connected layer to map the hidden state to the vocabulary dimension, generate multiple different candidate product copywriting, and calculate the total loss function.

[0118] The training sample set for the generation module consists of two parts: 1. Product copy text data, used to train the model's copywriting capabilities; 2. Product description keyword data, used as input for generation. Each training sample consists of content (keywords) and summary (complete copy), and training samples are sourced from both public and private business datasets. During the training phase, the variational lower bound (ELBO) is optimized, and the total loss function includes reconstruction loss and KL divergence loss.

[0119] Reconstruction Loss is used to measure the difference between the model-generated sequence and the original input. The mean squared error (MSE) is used as the measure of reconstruction error. The formula is as follows:

[0120] ;

[0121] in, represents the reconstruction loss, represents the true value at the t-th time step, represents the predicted value generated by the model at the t-th time step, is the total length of the sequence.

[0122] KL divergence loss (Kullback-Leibler Divergence Loss) is used to measure the difference between the posterior distribution of the potential vector and the standard normal distribution. It is calculated as follows:

[0123] ;

[0124] in, is the KL divergence loss, is the standard deviation of the underlying distribution, is the mean of the underlying distribution, is the latent vector Dimensions;

[0125] The total loss function (Total Loss) is to balance the reconstruction quality and the latent space regularization effect, as shown in the following formula:

[0126] ;

[0127] in, represents the total loss function, is the regulating factor.

[0128] Step S4: Building an evaluation module based on the large language model, comprehensively scoring the candidate product copywriting, and selecting a product copywriting from the candidate product copywriting based on the comprehensive score as the product copywriting output by the evaluation module;

[0129] Reference Figure 4The evaluation module uses a large language model to evaluate candidate product copywriting. Large language models are based on existing third-party large model APIs, including but not limited to iFlytek Spark, Dark Side of the Moon KIMI, Douyin Doubao, Tongyi Qianwen, and Tencent Hunyuan. Candidate product copywriting is evaluated using two evaluation criteria based on quality and relevance. The evaluation criteria include a scoring mechanism and a voting mechanism. The scoring mechanism requires the large language model to score candidate product copywriting with a score range of 0 to 1, retaining two decimal places. The voting mechanism requires the large language model to cast a unique vote on the candidate product copywriting. To facilitate the extraction of scoring and voting results, standard prompts are preset, requiring the model to only return formatted content (for example, "Rating: [0.86, 0.74, 0.81]; Votes: 1") to facilitate subsequent data string processing. Then, a simple counting variable is used to count the number of votes. In order to unify the dimensions, the number of votes for the candidate product copy is normalized (for example, if the total number of votes is 5 and a certain copy has 3 votes, its normalized vote score is 0.6). The comprehensive score of the score and the number of votes is calculated by weighted calculation, and the candidate product copy is selected based on the comprehensive score as the product copy output by the evaluation module.

[0130] If the evaluation module is in the last round of iteration, the candidate product copy with the highest comprehensive score is selected as the product copy output by the evaluation module. If it is not the last round of iteration, the greedy algorithm Top-K is used to retain the top K (for example, K=2) candidate product copies with the highest comprehensive scores as the product copy output by the evaluation module and enter the next round of iteration.

[0131] Step S5: Construct a memory module to store the user input content and the generated product copy;

[0132] The memory module is used to store and retrieve historical conversation content using a MySQL database. It converts user input and generated product text into a structured triplet of product-keyword-text and stores it in the MySQL database. The database supports manual entry by users to accommodate rapid access to new knowledge.

[0133] Step S6: Based on a tree of thought (TOT) structure with a preset number of layers, a prompt engineering module is constructed, and the memory module, reasoning module, generation module, and evaluation module are called to perform a preset number of iterations to obtain the final product copy.

[0134] Reference Figure 4 ,The thinking tree structure includes root node, middle node, leaf node, and number of thinking tree layers;

[0135] If this is the first iteration, the root node is the keywords and implicit keywords, such as "top, denim, white, simple, embroidery, jacket, ripped" and "clean, elegant, feminine, personalized design." If this is a subsequent iteration, the root node is the product copy generated in the previous iteration, as well as the preset keywords related to the focus of this iteration.

[0136] The intermediate node is the list of candidate product copywriting in the current iteration;

[0137] The leaf node is the final product copy output;

[0138] The number of layers in the thinking tree is the number of iteration rounds, which is preset by the user.

[0139] In the prompt engineering module, relevant keywords are preset for each round of iteration, focusing on the product entity (such as product type and brand), user group, product characteristics and attributes (such as color and style), and main style and scene.

[0140] In the embodiment, a 4-layer thinking tree structure is set up and 4 rounds of iteration are performed:

[0141] Layer 1: Generate multiple copywriting styles based on the keyword "tops + jackets" related to the product category;

[0142] The second layer: Focusing on the user group-related keywords "female + personality", focusing on potential user groups and providing precise copywriting;

[0143] Level 3: Describe the product based on its characteristics (color, style, material, etc.) such as "denim + white + holes";

[0144] Level 4: Final polish and sublimation of the copy around the style and scene of "simplicity + embroidery + elegance";

[0145] In the prompt engineering module, the steps for building and pruning the thinking tree structure are as follows:

[0146] Step S6.1: If this is the first iteration, enter the keywords and the implicit keywords output by the inference module. If this is a subsequent iteration, call the product copy generated in the previous iteration in the memory module and enter the preset keywords related to the focus of this iteration;

[0147] Step S6.2: Call the generation module to generate a list of candidate product copywriting in the current iteration, with corresponding labels {1, 2, ..., N}, where N is the number of candidate product copywriting;

[0148] Step S6.3: Call the evaluation module to obtain the comprehensive score of each candidate product copy;

[0149] Step S6.4: To maintain the diversity of generated product copy, use the greedy Top-K algorithm to select the K candidate product copies with the highest comprehensive scores to enter the next iteration;

[0150] Step S6.5: After completing the preset round of iterations, the product copy with the highest comprehensive score is selected from the candidate product copies and output as the final product copy.

[0151] In this example, the final product copy output is: "A fashionable white denim jacket, simple and elegant white, highlights individual charm. Paired with ripped jeans, it is full of individual charm. The embroidery design is exquisite and delicate, highlighting the elegant temperament of women. The denim fabric is comfortable and skin-friendly, making it comfortable and comfortable to wear."

[0152] The following is an application embodiment of the present invention:

[0153] The inference model was trained 15,000 times. In this example, the performance of the PEP-GNN model was evaluated using four datasets: HotpotQA, 2WikiMultihopQA, WebQuestions, and TriviaQA. The F1 metric used was the harmonic mean of precision and recall. Since each dataset has different characteristics, the experimental results use the F1 values of HotpotQA, 2WikiMultihopQA, and PopQA. The HotpotQA dataset used in this example is an English multi-hop question-answering dataset containing approximately 112,000 manually constructed questions. The answers to questions rely on the combined reasoning of information from two or more documents. The 2WikiMultihopQA dataset is an English multi-hop question-answering dataset built on Wikipedia, containing approximately 20,000 questions. The PopQA dataset is an open-domain question-answering dataset primarily used to evaluate the knowledge retention and recall capabilities of large-scale language models. This dataset contains approximately 7,500 English question-and-answer pairs, covering high-frequency, popular facts. Sourced from Wikipedia, query logs, and Q&A communities, the questions are diverse in their formulation and possess good generalization capability. Each question is trained on a cross-document graph neural network using information from two Wikipedia entity pages. The training parameters include a Top-K of 3, a path pruning threshold of 0.3, 3 multi-hop inferences, a batch size of 8, a dropout of 0.3, 1000 test samples, and 15,000 training epochs. See Table 1 for a comparison of the accuracy of the text generated by each model:

[0154] Table 1. Comparison of copywriting accuracy

[0155]

[0156] As shown in Table 1, the PEP-GNN reasoning module proposed in this embodiment has a significant performance improvement compared to other RAG framework reasoning models, according to comparative experimental results. The present invention improves both F1 and EM accuracy at different test levels. The reasoning accuracy of the reasoning module provides better keyword support for the subsequent generation module.

[0157] The baseline model used was Chatglm3-6b, whose performance, as reflected by its indicators, was midway between the three selected large models. Based on this, the CVPH model was introduced. After training the generation module 20,000 times, the number of thought tree layers was 4, the number of generated texts was 5, the number of Top-K paths was 3, and the multi-hop reasoning depth was 3. In this example, two datasets were used: AdvertiseGen, which contains 98,375 pieces of WeChat clothing advertising copy, and TaoBaoGen, which contains 1,048,576 pieces of Taobao product advertising copy. This example uses a conditional variational autoencoder and an ON-LSTM gated hierarchical structure as examples, and uses the currently popular large language models Qwen, LLaMA, and Chatglm as control groups to compare and verify the technical effectiveness of the text generation method provided in this example. The four evaluation metrics used are Self-Bleu, Rogue-1, Rouge-2, and Rouge-L, which are used to assess the diversity of text generation. Lower scores indicate higher diversity in generated text. Rogue is used to evaluate the similarity between generated text and reference text. Rogue-1 measures the weight of words, Rogue-2 measures the weight of bigrams, and Rogue-L measures the overlap of the longest common subsequence. A higher score indicates a higher similarity between the generated text and the reference. The specific training parameters are as follows: The Transformer-based encoder and decoder both contain 6 encoder layers and 6 decoder layers. The sequence length is 2048, the hidden layer size is 4096, the number of attention heads is 32, the latent vector size is 256, the UD network block size is 32, the dropout weight is 0.1, and the number of ON-LSTM layers is 6. All models underwent 10 cycles of LoRa fine-tuning; refer to Table 2 for a comparison of the model-generated text indicator scores:

[0158] Table 2. Comparison of model-generated text index scores

[0159]

[0160] In Table 2, CAVE represents a conditional variational autoencoder, Onlstm represents an ON-LSTM gated hierarchical structure, and tot represents a mind tree. It is not difficult to see from Table 2 that the text generation method of the improved variational autoencoder proposed in this embodiment can generate more diverse advertising text while improving the controllable generation of conditional text. According to the results of comparative experiments and ablation experiments, the text generation method of the improved variational autoencoder based on a large language model provided in this embodiment has a Self-Bleu index reduction of at least 40% compared to the first three large models, and the reuse rate of generated text is reduced, while the Rogue-1 index has an improvement of at least 8.56%, the Rogue-2 index has an improvement of at least 25.22%, and the Rogue-L has also been slightly improved, indicating that while ensuring diversity, the keyword relevance of the text remains at a certain level. After using the hierarchical progression structure, the model of the present invention has a 5% improvement in long sequence performance, verifying that the CVAE model and the ON-LSTM gated hierarchical structure have a positive enhancement effect on copy generation. The generated text has a higher correlation with the reference words and is more diverse, making it a reliable and practical copy generation solution.

[0161] In summary, the product copywriting generation method based on conditional variational autoencoder and retrieval enhancement designed by the present invention meets the relevance and diversity of copywriting generation, thus having practical application scenarios.

[0162] The embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the scope of knowledge possessed by ordinary technicians in this field without departing from the spirit of the present invention.

Claims

1. A product copywriting generation method based on conditional variational autoencoder and retrieval enhancement, characterized in that: Execute steps S1 to S6 to complete the generation of product copy: Step S1: extract keywords from the product copy requirement description entered by the user; Step S2: Using keywords as input, a graph neural network model-based reasoning module is constructed. Multi-hop graph reasoning is used to generate candidate reasoning paths. The candidate reasoning paths undergo path embedding, path pruning, and score calculation. Several candidate reasoning paths are selected as output paths based on the scores, and implicit keywords are extracted and output. Step S3: Taking keywords and hidden keywords as input, a generation module is constructed, consisting of an embedding layer, an encoder, a reparameterization module, a decoder, and an output layer. The generation module integrates the conditional variational autoencoder structure and the ON-LSTM gated hierarchical structure to output several candidate product copywriting; Step S4: Building an evaluation module based on the large language model, comprehensively scoring the candidate product copywriting, and selecting a product copywriting from the candidate product copywriting based on the comprehensive score as the product copywriting output by the evaluation module; Step S5: Construct a memory module to store the user input content and the generated product copy; Step S6: Based on the thinking tree structure with a preset number of layers, a prompt engineering module is constructed, and the memory module, reasoning module, generation module and evaluation module are called to perform preset rounds of iteration to obtain the final product copy.

2. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 1, characterized in that: The specific steps of step S2 are as follows: Step S2.1: Assume that the knowledge graph is ,in is a collection of entities, For edge sets; set the inference depth parameter ,implement Round multi-hop graph reasoning, in the knowledge graph i In layer propagation, entity nodes Representation By its neighbor nodes The information is obtained through weighted aggregation, where Representation node The formula for weighted aggregation of information is as follows: ; in, For entity nodes the expression; Represents the nodes in the adjacency matrix i and nodes j connectivity relationship; Represents a slave node i To Node j relational embedding; is the neighbor aggregation weight matrix; is the relationship weight matrix; is the activation function; Step S2.2: Given a starting entity in the knowledge graph Start, go through multiple intermediate entities until the target entity Path , and adopt splicing modeling: ; in, 、 is an intermediate entity, express and relationship, express and relationship, L is the total number of entities; Step S2.3: For each path , perform path pruning, including static pruning and semantic pruning; Static pruning for each path Calculating path scores : ; in, For the i entities, L is the total number of entities; represents the L2 norm; Setting thresholds , only keep Path; Semantic pruning for each path and the user input query vector Calculate cosine similarity : ; Keep only Path; Step S2.4: Sort the paths according to their scores, starting with the path with the highest score, and select several paths in the following format: , extract tail entities, which contain implicit keywords.

3. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 1, characterized in that: The specific steps of step S3 are as follows: Step S3.1: The embedding layer converts word IDs into dense vectors using a word embedding function; adjusts the dimensions of the dense vectors, which include sequence length, batch size, and hidden layer dimensions; and feeds the embedded vectors into the encoder. Step S3.2: The encoder adopts the conditional variational autoencoder structure, and the formula is as follows: ; ; in, is the standard deviation of the underlying distribution, is the mean of the underlying distribution, is hidden state, is the weight matrix, is the bias vector; Step S3.3: The reparameterization module samples the latent vector from the Gaussian distribution : ; in, represents noise sampled from a standard normal distribution, represents element-wise multiplication, is the standard deviation of the underlying distribution, is the mean of the underlying distribution, represents a normal distribution; Step S3.4: The decoder converts the latent vector Expand to the target sequence length and pass the latent vector Transmitted to different layers of the ON-LSTM gating hierarchy to generate hidden states and generate token distributions; Step S3.5: The output layer uses a fully connected layer to map the hidden state to the vocabulary dimension, generate multiple different candidate product copywriting, and calculate the total loss function.

4. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 3, characterized in that: The total loss function in step S3.5 includes reconstruction loss and KL divergence loss, and the weighted sum of the two is calculated: ; in, represents the total loss function, represents the reconstruction loss, represents the KL divergence loss, is the regulating factor.

5. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 1, characterized in that: The evaluation module in step S4 calls the large language model to evaluate the candidate product copywriting, using two evaluation criteria: a scoring mechanism and a voting mechanism. The scoring mechanism requires the large language model to score the candidate product copywriting, with a score range between 0 and 1; the voting mechanism requires the large language model to cast a unique vote on the candidate product copywriting, calculate the weighted comprehensive score of the score and the number of votes, and select from the candidate product copywriting based on the comprehensive score as the product copywriting output by the evaluation module.

6. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 5, characterized in that: In step S4, if the evaluation module is in the last round of iteration, the candidate product copy with the highest comprehensive score is selected as the product copy output by the evaluation module. If it is not the last round of iteration, the greedy algorithm Top-K is used to retain the top K candidate product copies with the highest comprehensive scores as the product copy output by the evaluation module and enter the next round of iteration.

7. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 1, characterized in that: The memory module in step S5 uses a MySQL database to convert the user input and the generated product copy into a structured triple form of product-keyword-copy and store it in the MySQL database.

8. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 1, characterized in that: The prompt engineering module in step S6, wherein the thinking tree structure includes a root node, an intermediate node, a leaf node, and a number of thinking tree layers; If this is the first iteration, the root node is the keyword and implicit keyword. If this is a subsequent iteration, the root node is the product copy generated in the previous iteration and the preset keywords related to the focus of this iteration. The intermediate node is the list of candidate product copywriting in the current iteration; The leaf node is the final product copy output; The number of layers in the thinking tree is the number of iteration rounds, which is preset by the user.

9. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 8, characterized in that: In the prompt engineering module in step S6, for each iteration, key words related to the focus are preset, including the product subject, user group, product characteristics and attributes, and main style and scene.

10. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 8, characterized in that: In the prompt engineering module in step S6, the steps of constructing and pruning the thinking tree structure are as follows: Step S6.1: If this is the first iteration, enter the keywords and the implicit keywords output by the inference module. If this is a subsequent iteration, call the product copy generated in the previous iteration in the memory module and enter the preset keywords related to the focus of this iteration; Step S6.2: Call the generation module to generate a list of candidate product copywriting in the current iteration, with corresponding labels {1, 2, ..., N}, where N is the number of candidate product copywriting; Step S6.3: Call the evaluation module to obtain the comprehensive score of each candidate product copy; Step S6.4: Use the greedy Top-K algorithm to select the K candidate product copywritings with the highest comprehensive scores to enter the next round of iteration; Step S6.5: After completing the preset round of iterations, the product copy with the highest comprehensive score is selected from the candidate product copies and output as the final product copy.

Citation Information

Patent Citations

  • Inference method and device for multi-hop question and answer of knowledge graph and storage medium

    CN118503366A

  • Iterative design method and system for cultural and creative products based on user emotion feedback

    CN119295185A