Commodity copywriting generation method based on conditional variation automatic encoder and retrieval enhancement

Through the combination of conditional variational automatic encoder and graph neural network, the diversity and flexibility of the general large language model in generating product copywriting in specific fields is solved, and the product copy generation with clear styles is achieved, which improves the generation ability and relevance of copywriting.

CN120297243AActive Publication Date: 2025-07-11NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202510787349.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-11
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

When generating product copywriting in a specific field, the existing general language model is difficult to meet the description of multiple product attributes at the same time, and the generated copywriting lacks diversity and flexibility, and cannot adapt to complex semantic expressions and personalized needs.

Method used

The inference module based on conditional variational autoencoder and graph neural network is adopted, combining multi-hop graph inference and path pruning, and the generation module integrates the conditional variational autoencoder and ON-LSTM gated hierarchical structure, and generates diverse product copy through iterative optimization of thinking tree.

Benefits of technology

It realizes clear and diverse product copy generation, improves the generation ability and diversity of copywriting in specific fields, and maintains the relevance and consistency of generated copywriting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297243A_ABST
    Figure CN120297243A_ABST
Patent Text Reader

Abstract

The invention discloses a commodity copywriting generation method based on a conditional variation automatic encoder and retrieval enhancement, which comprises the following steps: designing a reasoning module, a generation module, an evaluation module, a memory module and a prompt engineering module, aiming at user input content, based on a graph neural network, adopting path pruning and multi-hop reasoning, and generating a commodity copywriting in the reasoning module; pruning of a low-score path is achieved through path reasoning, the reasoning efficiency is improved, and an implicit relation is reasoned; the generation module introduces a layer-by-layer variational auto-encoder and enables the generated candidate commodity copywriting to meet specific condition requirements through conditional variables, the evaluation module performs comprehensive scoring on the candidate commodity copywriting, the memory module is constructed to store conversation between a user and the generation module, and the prompt engineering module performs iteration of a preset round based on a thinking tree structure. According to the method designed by the invention, the personalization and reality sense of the copywriting can be improved, and the generation of diversified advertisement copywriting with clear styles is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large artificial intelligence models, and particularly to a method for generating product copywriting based on conditional variational autoencoders and retrieval enhancement. Background Art

[0002] In the fields of artificial intelligence and natural language processing (NLP), with the rise of deep learning and large language models (LLMs) (such as GPT, BERT, T5, etc.), copywriting generation has become a key technology for content creation, advertising and marketing, and personalized recommendations. Traditional copywriting generation methods mainly rely on template filling and rule-driven strategies, such as text synthesis based on predefined sentence structures. Although these methods can generate structured text in specific scenarios, their flexibility and diversity are limited, and it is difficult to adapt to complex semantic expressions and personalized needs.

[0003] However, although LLMs perform excellently in open-domain text generation, they still face challenges in generating high-quality copywriting in specific fields (such as product advertisements, etc.). General large models lack understanding of vertical domains (such as specific product categories, industry terms), and it is difficult to simultaneously describe multiple product attributes, and the emphasis on each attribute is not completely consistent. It is necessary to have different emphases according to the user's needs, and while ensuring that the copywriting is relevant to the user, the generated copywriting should have diversity and extensibility. After being fine-tuned, general large models have the ability to generate text for specific fields, but the product update and iteration speed is fast, and some combinations are relatively few or not trained. At this time, it is necessary to introduce an external knowledge base to supplement it. Summary of the Invention

[0004] The object of the present invention is to provide a method for generating product copywriting based on conditional variational autoencoders and retrieval enhancement, which can realize the generation of product copywriting with clear style and diversity, and enhance the ability to generate product copywriting in vertical domains.

[0005] To achieve the above functions, the present invention designs a method for generating product copywriting based on conditional variational autoencoders and retrieval enhancement, and performs steps S1 - S6 to complete the generation of product copywriting: Step S1: Extract keywords for the description of the product copywriting requirements input by the user; Step S2: Using the keywords as input, construct an inference module based on a graph neural network model, and use multi-hop graph reasoning to generate candidate inference paths. The candidate inference paths go through path embedding, path pruning, and score calculation, and several candidate inference paths are selected as output paths according to the scores, and the implicit keywords are extracted and output; Step S3: Using the keywords and implicit keywords as inputs, construct a generation module composed of an embedding layer, an encoder, a reparameterization module, a decoder, and an output layer. The generation module integrates the conditional variational autoencoder structure and the ON-LSTM gated hierarchical structure to output several candidate product copywriting texts; Step S4: Based on the large language model, construct an evaluation module to comprehensively score the candidate product copywriting texts, and select from the candidate product copywriting texts according to the comprehensive score as the product copywriting text output by the evaluation module; Step S5: Construct a memory module to store the user input content and the generated product copywriting texts; Step S6: Based on the thinking tree structure with a preset number of layers, construct a prompt engineering module, call the memory module, the reasoning module, the generation module, and the evaluation module, and perform iterative operations for a preset number of rounds to obtain the final product copywriting text.

[0006] Beneficial effects: Compared with the prior art, the advantages of the present invention include: 1. The reasoning module proposed by the present invention learns the representation of nodes in the graph structure through the information propagation mechanism. It captures the relationships between entities through the adjacency matrix and relationship features, and enhances the expression ability of the model through multi-layer reasoning, enabling the graph neural network to effectively perform multi-hop reasoning. The path pruning method reduces the training time of the model while maintaining the existing retrieval accuracy; 2. For the generation module proposed by the present invention, its input passes through the encoder, is transmitted to the latent space, and then to the decoder of the ON-LSTM gated hierarchical structure. During the process, a latent vector is learned and input to different ON-LSTM gated hierarchies through different weights. The more emphasized attributes are located higher in the ON-LSTM gated hierarchy and will be retained for a longer time. Moreover, the latent vector contains new information, realizing controllable generation while making the copywriting generation diversified; 3. For the evaluation module based on the thinking tree proposed by the present invention, it uses the thinking tree strategy to realize iterative copywriting generation, making the copywriting generation more diversified. The Top-K strategy of the voting mechanism retains the K optimal paths, combines the greedy algorithm with random sampling, balances the local optimum and the global optimum, and avoids the over-restriction of copywriting generation. Description of the Drawings

[0007] Figure 1 is a flowchart of the product copywriting generation method based on the conditional variational autoencoder and retrieval enhancement according to the embodiment of the present invention; Figure 2 is a flowchart of the reasoning module according to the embodiment of the present invention; Figure 3 is a schematic diagram of the generation module according to the embodiment of the present invention; Figure 4It is a schematic diagram of an evaluation module and a thought tree structure provided according to an embodiment of the present invention. Detailed implementation manners

[0008] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.

[0009] The method for generating product copy based on conditional variational autoencoder and retrieval enhancement provided by the embodiment of the present invention, referring to Figure 1 , execute step S1-step S6 to complete the generation of product copy: Step S1: For the description of the product copy requirement input by the user, use an existing large language model to extract keywords; In the embodiment, the system receives the description of the product copy requirement input by the user, for example: "I want to generate a product advertisement copy about a white top, the fabric is denim, the pattern is embroidery, there are holes, and the style is simple". Use an existing pre-trained large language model (such as the Transformer structure) to perform entity recognition and intent extraction on the input content, and identify key entity items: "top, denim, white, simple, embroidery, hole", etc. Subsequently, align the entity items with the standard entities in the knowledge graph to construct an initial triple structure, such as: ("top", "material", "denim") ("top", "color", "white") ("top", "style", "simple") ("top", "style", "embroidery") ("coat", "style", "hole").

[0010] Step S2: Use the keyword as the input to construct an inference module based on the graph neural network model, generate candidate inference paths using multi-hop graph inference, the candidate inference paths go through path embedding, path pruning and score calculation, and select several candidate inference paths as the output paths according to the scores, and extract and output implicit keywords; Referring to Figure 2 , the inference module is used to obtain the description information of the product copy requirement input by the user, and obtain relevant attributes through database retrieval and multi-hop graph inference. The inference module is based on the PEP-GNN (Path Encoding and Pruning) model of GNN, which combines the instruction-driven inference mechanism and the graph neural network to achieve efficient knowledge inference.

[0011] The specific steps of step S2 are as follows: Step S2.1: Load a pre-trained graph neural network model, including structures such as graph convolutional kernels, attention weight matrices, and historical path memory parameters. The graph neural network model is trained on a large-scale product copy and domain knowledge graph joint corpus to enhance its domain adaptability and inference ability.

[0012] Suppose the knowledge graph is , where is the set of entities, is the set of edges, and each edge represents the semantic relationship between two entities; Set the inference depth parameter , and execute rounds of multi-hop graph reasoning. In the propagation of the i th layer in the knowledge graph, the representation of the entity node is obtained by weighted aggregation of the information of its neighbor nodes , where represents the set of neighbor nodes of node , that is, the set of all nodes connected to node ; the formula for information weighted aggregation is as follows: ; Among them, is the representation of the entity node ; represents the connectivity relationship between node i and node j in the adjacency matrix; represents the relationship embedding from node i to node j ; is the neighbor aggregation weight matrix; is the relationship weight matrix; is the activation function; The propagation of each hop is similar to the diffusion of signals in the knowledge graph, thus constructing a multi-hop entity relationship link.

[0013] In the embodiment, set the inference depth parameter , and perform three rounds of information propagation in the knowledge graph. In each hop, according to the current entity representation and the knowledge graph structure, aggregate the information of adjacent nodes and calculate the attention score. The inference process is as follows: The 1st hop: Starting from the "top", aggregate the first-level attributes such as "denim", "white", "simple", etc.; The 2nd hop: Jump from "white" to its semantically related words, such as "clean", "elegant", etc.; The 3rd hop: Expand from "embroidery" to abstract descriptions such as "exquisite", "feminine", "personalized design", etc.

[0014] Each hop of propagation updates the node state and records the path propagation trajectory and its score, and finally generates a set of candidate inference paths.

[0015] Step S2.2: To support path-level inference selection, node representations need to be constructed and the entire path needs to be vectorized for modeling. Given a path in the knowledge graph that starts from the starting entity and passes through multiple intermediate entities until the target entity ; ; To effectively model the inference path, a path embedding is constructed for each path. The path embedding vector is used to represent the semantic chain formed by starting from the starting entity, passing through multiple intermediate entities and relationships, and reaching the target node. The Concat splicing method is used to model the path: ; where is the path embedding vector obtained by modeling the inference path, is the embedding representation of the i th entity in the path, and represents the splicing operation. This method preserves the semantic order of the entities and can provide a finer-grained control signal for subsequent pruning and generation; The path after being modeled using Concat splicing is represented as: ; where , are intermediate entities, represents the relationship between and , represents the relationship between and , and L is the total number of entities; this method preserves semantics and order and helps the conditional information to be used as an input basis in subsequent pruning and scoring, accurately guiding the style and details of the product copywriting.

[0016] Step S2.3: To effectively screen out the path with the highest relevance to the input query in the high-dimensional path space, for each path , path pruning is performed, including static pruning and semantic pruning; Static pruning calculates the path score for each path to measure its information content and stability: ; where is the i th entity, and L is the total number of entities; represents the L2 norm, which is extremely sensitive to embedding dimensions and can effectively filter out paths with severe semantic drift; The path score can identify semantic drift or noisy paths and set a threshold , and only retain the paths; Semantic pruning calculates the cosine similarity for each path with the user input query vector : ; ; Only retain the paths to ensure the semantic relevance and accuracy of the inference results.

[0017] Step S2.4: Sort each path according to the path score, and starting from the path with the highest score, select several paths in sequence, in the form of: , extract the tail entity, which contains implicit keywords; In the embodiment, the above process is expressed as: Inference path: top → material → denim → style → hole → style → personality publicity; Extract implicit keywords: "personality publicity", "denim", "hole design".

[0018] Step S3: Using the keywords and implicit keywords as inputs, construct a generation module composed of an embedding layer, an encoder, a reparameterization module, a decoder, and an output layer. The generation module integrates the structure of a Conditional Variational Autoencoder (CVAE) and an ON-LSTM gated hierarchical structure to output several candidate product copywriting; Referring to Figure 3 , the specific steps of step S3 are as follows: Step S3.1: The embedding layer uses a word embedding function to convert the word Id into a dense vector; adjust the dimension of the dense vector, and the dimension of the dense vector includes the sequence length, batch size, and hidden layer dimension; send the embedded vector to the encoder; The input content of the generation module includes the following two parts: Keywords (from the product knowledge graph): top, denim, white, simple, embroidery, coat, hole. Implicit keywords (from the three-hop inference path of the graph neural network): clean, elegant, feminine, personalized design. Concatenate and combine the above keywords and implicit keywords as the input content of the generation module, and send it to the tokenizer for processing. The tokenizer cuts the input text into several semantic units based on the pre-trained dictionary and encodes it into a vector tensor form as the input of the encoder.

[0019] Step S3.2: The encoder adopts the structure of a conditional variational autoencoder, and the formula is as follows: ; ; wherein, is the standard deviation of the latent distribution, calculated using the sub-neural network layer logvar_layer in the conditional variational autoencoder structure, controlling the dispersion degree of the latent distribution, is the mean of the latent distribution, calculated using the sub-neural network layer mu_layer in the conditional variational autoencoder structure, determining the central position in the latent space, is the hidden state, is the weight matrix, used to map the hidden state of the encoder to the mean of the latent distribution, is the bias vector; Step S3.3: The reparameterization module samples a latent vector from the Gaussian distribution: ; wherein, represents the noise sampled from the standard normal distribution, represents element-wise multiplication, is the standard deviation of the latent distribution, is the mean of the latent distribution, represents the normal distribution; The low-dimensional latent vector sampled therefrom depends on the input keyword conditional information.

[0020] Step S3.4: The decoder extends the latent vector to the target sequence length and transmits it to different layers of the ON-LSTM gated hierarchical structure through the latent vector to generate hidden states and generate a token distribution; The decoder part is based on the ON-LSTM model, adopting a hierarchical progression structure, and guiding the copywriting generation process through the latent vector . The design of this structure enables the generation model to flexibly capture semantic information at different levels and effectively separate details from the theme.

[0021] The core design idea of the decoder is to inject the latent vector into different levels of the ON-LSTM model in order to better model the long-term and short-term information of the copywriting. The decoder is composed of multiple layers of ON-LSTM stacked together, combined with the latent vector Guide the hierarchical structure division of each layer of neurons. The ON-LSTM model itself dynamically models long-term and short-term memory allocation through the cumulative softmax mechanism (cumax). The high layers are responsible for topic semantics. The high layers (the shallow layers closer to the input) mainly receive the subspace features related to topics and styles in the latent vector, controlling the overall semantic consistency; the low layers are responsible for local details. The low layers (the deep layers closer to the output) receive the subspace features related to fine-grained details (such as attributes and modifiers) in the latent vector, flexibly adjusting the generation of text details. Therefore, the hierarchical injection of the latent vector is not simply adding the whole to all layers, but injecting selectively, proportionally, and modularly.

[0022] To make the latent vector have an impact on the copywriting generation process without destroying the sentence structure and backbone information, a hierarchical injection mechanism is adopted to inject the latent vector into the ON-LSTM model step by step according to different levels. This method can effectively control the influence degree of latent variables on different levels, enabling the model to flexibly generate text that meets the control conditions while ensuring semantic consistency.

[0023] To enable the latent vector to have a control effect on the text generation process, the gating mechanism is improved. The gating mechanism in the ON-LSTM model consists of the Master Forget Gate, the Master Input Gate, and the Candidate Memory. Through these gating mechanisms, the model can flexibly forget and update information at different levels, ensuring the controllability of the generation process.

[0024] Suppose the ON-LSTM model contains L layers, and the index of each layer is . To achieve hierarchical control of latent information, define the corresponding latent vector transformation for each layer: ; In the formula, is a multi-layer perceptron specifically learned for the th layer, responsible for extracting the subspace information related to the current level from the original latent vector .

[0025] In the gating calculation of each layer, taking the th layer as an example, the latent variable is respectively injected into the Master Forget Gate, the Master Input Gate, and the regular gating, and the formulas are as follows: Master Forget Gate, through the cumax activation function, whose principle is to generate a progressive forgetting weight from 0 to 1 through cumulative softmax, controls the retention ratio of the historical cell state as follows: ; In the formula, represents the Master Forget Gate, , , represent the weight matrices of the Master Forget Gate corresponding to the current input , the upper hidden state and the latent vector respectively.

[0026] Master Input Gate, which forms a complementary relationship with the Forget Gate, is as follows: ; In the formula, represents the Master Input Gate, , , represent the weight matrices of the Master Input Gate corresponding to the current input , the upper hidden state and the latent vector respectively; In addition, in order to fuse the master gate and the conventional gate, a reconstruction gate mechanism is introduced.

[0027] Conventional gates are the input gate and the forget gate in the standard LSTM, and the related definitions are as follows: Input gate: ; In the formula, represents the input gate, , represent the weight matrices of the input gate corresponding to the current input , the upper hidden state respectively, is the learnable offset.

[0028] Forget gate: ; In the formula, represents the forget gate, , represent the weight matrices of the forget gate corresponding to the current input , the upper hidden state respectively, is the learnable offset.

[0029] Output gate: ; In the formula, represents the output gate, , represent the weight matrices of the output gate corresponding to the current input , the upper hidden state respectively, and is the learnable offset.

[0030] The reconstructed forget gate is defined as: ; The reconstructed input gate is defined as: ; where represents element-wise multiplication.

[0031] Candidate memory update: At each time step, the candidate memory integrates the current input , the upper hidden state and the latent vector , and constrains the value within the range of [-1, 1] through the tanh activation function to ensure stable gradient flow. A learnable offset is introduced during the calculation of the linear transformation. The update formula for the candidate memory is as follows: ; In the formula, , , represent the weight matrices of the candidate memory corresponding to the current input , the upper hidden state and the latent vector respectively, and is the learnable offset.

[0032] Finally, the update of the cell state and the hidden state is completed through the following formula: ; ; where is the output gate, which determines the influence of the current cell state on the output. Through this mechanism, the model can effectively control the long-term dependencies and short-term details of information when generating text.

[0033] This hierarchical latent variable injection mechanism enables high-level neurons to stably maintain the main theme of the text, and low-level neurons to flexibly fine-tune the details, achieving diversity and consistency in the copywriting style and content.

[0034] The decoder generates layers according to the following process: 1. According to the latent vector Generate potential vectors of each layer through different MLP modules ; 2. In each layer of ON-LSTM, the potential vector Inject each gated calculation separately; 3. Pass the hidden state and cell state downward layer by layer until the final text sequence is output.

[0035] Through layered injection, the model achieves organic coordination of semantics of different granularities in the generation process, enhancing the naturalness and controllability of the generated text.

[0036] In the embodiment, the generation module generates three candidate advertisements for the input keyword set: "top, denim, white, simple, embroidery, jacket, hole", as shown below: Copy 1: Fashionable white denim jacket, simple and elegant white, highlights individual charm, paired with ripped jeans, full of individual charm. The embroidery design is exquisite and delicate, highlighting the elegant temperament of women. Denim fabric is comfortable and skin-friendly, and it is comfortable to wear.

[0037] Copywriting 2: This white embroidered denim jacket combines fashion and comfort. The ripped holes enhance the street feel, and the simple design highlights the high-end feel, creating a different kind of fashion charm.

[0038] Copywriting 3: The white top is full of personality, with a design that combines ripped holes with embroidery. The denim material is soft and durable, combining practicality and beauty, making it a versatile item for spring and autumn.

[0039] Step S3.5: The output layer uses a fully connected layer to map the hidden state to the vocabulary dimension, generate multiple different candidate product copywriting, and calculate the total loss function.

[0040] The training sample set of the generation module consists of two parts: 1. Product copy text data, which is used to train the model's copywriting expression ability; 2. Product description keyword data, which is used as the input condition for generation; each training sample consists of content (keywords) and summary (complete copy), and the training sample sources include public data sets and private business data sets. In the training stage, the variational lower bound (Evidence Lower Bound, ELBO) is used as the optimization target, and the total loss function includes reconstruction loss and KL divergence loss; The reconstruction loss is used to measure the difference between the sequence generated by the model and the original input. The mean squared error (MSE) is adopted as the metric for the reconstruction error, and the formula is as follows: ; where, represents the reconstruction loss, represents the true value at the t-th time step, represents the predicted value generated by the model at the t-th time step, is the total length of the sequence.

[0041] The Kullback-Leibler divergence loss is used to measure the difference between the posterior distribution of the latent vector and the standard normal distribution, and the specific calculation is as follows: ; where, is the KL divergence loss, is the standard deviation of the latent distribution, is the mean of the latent distribution, is the latent vector 's dimension; The total loss function is to balance the reconstruction quality and the regularization effect of the latent space, as shown in the following formula: ; where, represents the total loss function, is the adjustment factor.

[0042] Step S4: Build an evaluation module based on the large language model, comprehensively score the candidate product copywriting, and select from the candidate product copywriting according to the comprehensive score as the product copywriting output by the evaluation module; Refer to Figure 4, the evaluation module calls a large language model to evaluate the candidate product copywriting. The large language model is an existing third-party large model API, including but not limited to iFlytek Spark, Darkside Kimi, ByteDance Doubao, Tongyi Qianwen, and Tencent Hunyuan. According to quality and relevance, two evaluation criteria are adopted to evaluate the candidate product copywriting. The evaluation criteria include a scoring mechanism and a voting mechanism. The scoring mechanism requires the large language model to score the candidate product copywriting, and the score range is between 0 and 1, with two decimal places reserved; the voting mechanism requires the large language model to cast a single vote on the candidate product copywriting. To facilitate the extraction of scoring and voting results, a preset standard prompt is used, and the model is required to only return formatted content (for example: "Score: [0.86, 0.74, 0.81]; Vote: 1"), which is convenient for subsequent processing of data strings. Then, a simple counting variable is used to count the number of votes. To unify the dimension, the number of votes for the candidate product copywriting is normalized (for example, if the total number of votes is 5 and a certain copywriting gets 3 votes, then its normalized number of votes is 0.6). The comprehensive score of the score and the number of votes is calculated according to the weighted calculation, and the candidate product copywriting is selected according to the comprehensive score as the product copywriting output by the evaluation module.

[0043] If the evaluation module is the last round of iteration, the candidate product copywriting with the highest comprehensive score is selected as the product copywriting output by the evaluation module. If it is not the last round of iteration, the greedy algorithm Top-K is adopted, and the top K (such as K = 2) candidate product copywriting with the highest comprehensive score is retained as the product copywriting output by the evaluation module and enters the next round of iteration.

[0044] Step S5: Build a memory module to store the user input content and the generated product copywriting; The memory module is used to store and retrieve historical conversation content using a MySQL database. The user input and the generated product copywriting are converted into a structured triple form of product-keyword-copywriting and stored in the MySQL database. The database supports manual entry by users to adapt to the rapid access of new knowledge.

[0045] Step S6: Based on the tree of thought (tot) structure with a preset number of layers, build a prompt engineering module, call the memory module, reasoning module, generation module, and evaluation module, and perform iterations for a preset number of rounds to obtain the final product copywriting.

[0046] Refer to Figure 4 , the tree of thought structure includes a root node, intermediate nodes, leaf nodes, and the number of layers of the tree of thought; If it is the first iteration, the root nodes are the keywords and implicit keywords, such as: "top, denim, white, simple, embroidered, coat, ripped" and "clean, elegant, feminine, personalized design"; if it is a subsequent iteration, the root nodes are the product copy generated in the previous iteration and the keywords related to the focus of this iteration preset. The intermediate nodes are the list of candidate product copies in the current iteration. The leaf nodes are the final output of the product copy. The number of layers of the thinking tree is the number of iterations, which is preset by the user.

[0047] In the prompt engineering module, for each iteration, keywords related to the focus are preset. The focuses include the product entity (such as product category, brand), user group, product features and attributes (such as color, style), main style and scenario.

[0048] In the embodiment, a 4-layer thinking tree structure is set up and 4 iterations are carried out: The first layer: Generate multiple style copy candidates around the keyword "top + coat" related to the product category. The second layer: Focus on the potential user group and provide accurate copy around the keyword "female + personality" related to the user group. The third layer: Describe the product features (color, style, material, etc.) "denim + white + ripped". The fourth layer: Finally polish and enhance the copy around the style and scenario "simple + embroidered + elegant". In the prompt engineering module, the steps for constructing and pruning the thinking tree structure are as follows: Step S6.1: If it is the first iteration, input the keywords and the implicit keywords output by the reasoning module. If it is a subsequent iteration, call the product copy generated in the previous iteration in the memory module and the keywords related to the focus of this iteration preset. Step S6.2: Call the generation module to generate the list of candidate product copies in the current iteration, corresponding to the labels {1, 2,..., N}, where N is the number of candidate product copies. Step S6.3: Call the evaluation module to obtain the comprehensive score of each candidate product copy. Step S6.4: In order to retain the diversity of product copy generation, use the greedy algorithm Top-K to select the K candidate product copies with the highest comprehensive scores to enter the next iteration. Step S6.5: After completing the preset number of iterations, select the product copy with the highest comprehensive score from the candidate product copies as the final output of the product copy.

[0049] In the embodiment, the final product copywriting output is: "A fashionable white denim jacket. The simple and elegant white color showcases personal charm. Paired with ripped jeans, it is full of personal charm. The embroidery design is delicate and exquisite, highlighting the elegant temperament of women. The denim fabric is comfortable and skin-friendly, making you feel comfortable when wearing it." The following is an application embodiment of the present invention: The inference model is trained 15,000 times. In this embodiment, four datasets, namely HotpotQA, 2wikimultihopqa, WebQuestions, and TriviaQA, are selected to evaluate the performance of the PEP-GNN model. The evaluation metric F1 used is the harmonic mean of precision and recall. Since each dataset has different directions, the F1 values of HotpotQA, 2wikimultihopqa, and PopQA are selected for the experimental results. The HotpotQA dataset selected in this embodiment is an English multi-hop question-answering dataset, which contains approximately 112,000 artificially constructed question-answering data in total. The questions require joint reasoning based on information from two or more documents to obtain the answers. The 2WikiMultihopQA dataset is an English multi-hop question-answering dataset constructed on Wikipedia, with a total of approximately 20,000 questions. The PopQA dataset is an open-domain question-answering dataset, mainly used to evaluate the knowledge memory and recall ability of large-scale language models. This dataset contains approximately 7,500 English question-answer pairs in total, covering high-frequency common facts, sourced from Wikipedia, query logs, and question-answering communities. The question expressions are diverse and have good evaluation value for generalization ability. For each question, cross-document inference graph neural network training parameter settings are required by combining the information of two Wikipedia entity pages. Top-K is 3, the path pruning threshold is 0.3, the number of multi-hop inferences is 3, batch_size is 8, dropout is 0.3, the number of test samples is 1000, and the number of training times is 15,000. The comparison table of the accuracy of the copywriting text generated by each model is shown in Table 1: Table 1. Comparison table of the accuracy of copywriting text

[0050] It is not difficult to see from Table 1 that for the PEP-GNN inference module proposed in this embodiment, according to the comparative experimental results, this model has a good performance improvement compared with the inference models of other RAG frameworks. The present invention has improved the F1 and EM accuracies at different test levels, and the inference accuracy of the inference module provides better keyword support for the subsequent generation module.

[0051] The benchmark model selected is Chatglm3-6b, and its performance reflected by the indicators is in the middle of the three large models selected. On this basis, the CVPH model is introduced, and the generation module is trained 20,000 times. The number of layers of the thought tree is 4, the number of generated texts is 5, the Top-K path is 3, and the multi-hop reasoning depth is 3. In this embodiment, two datasets are selected. AdvertiseGen is the copywriting of WeChat business clothing advertisements, containing 98,375 pieces of data, and TaoBaoGen is the copywriting of Taobao commodity advertisements, containing 1,048,576 pieces of data. In this embodiment, the conditional variational autoencoder and the ON-LSTM gated hierarchical structure are taken as examples, and the current popular large language models Qwen, LLaMA, and Chatglm are used as the control group to compare and verify the technical effect of the copywriting generation method provided in this embodiment. The four evaluation indicators used are Self-Bleu, Rogue-1, Rouge-2, and Rouge-L: used to evaluate the diversity of text generation. The lower the score, the higher the diversity of the generated text. Rogue is used to evaluate the similarity between the generated text and the reference text. Rogue-1: measures the weight of words, Rogue-2: measures the weight of bigrams, Rogue-L: measures the coincidence degree of the longest common subsequence. The higher the score, the higher the similarity between the generated text and the reference. The specific training parameters are as follows: Both the encoder and decoder based on Transformer contain 6 encoder layers and 6 decoder layers. The sequence length is 2048, the hidden layer size is 4096, the number of attention heads is 32, the latent vector size is 256, the UD network block size is 32, the Dropout weight is 0.1, and the number of ON-LSTM layers is 6. All models are fine-tuned by Lora for 10 epochs; The comparison of the model-generated text index scores refers to Table 2: Table 2. Comparison Table of Model-Generated Text Index Scores

[0052] In Table 2, CAVE represents the conditional variational autoencoder, Onlstm represents the ON-LSTM gated hierarchical structure, and tot represents the thought tree. It is not difficult to see from Table 2 that the proposed improved variational autoencoder copywriting generation method in this embodiment can generate more diverse advertising copy while improving the controllable generation of conditional text. According to the results of the comparative experiment and ablation experiment, for the copywriting generation method based on the improved variational autoencoder of the large language model provided in this embodiment, compared with the first three large models, the Self-Bleu index of the model of the present invention is reduced by at least 40%, and the reuse rate of the generated copy is reduced, while the Rogue-1 index is increased by at least 8.56%, the Rogue-2 index is increased by at least 25.22%, and the Rogue-L also has a small increase, indicating that while ensuring diversity, the keyword relevance of the copy still remains at a certain level. After the model of the present invention uses the hierarchical progression structure, the long sequence performance is improved by 5%, verifying that the CVAE model and the ON-LSTM gated hierarchical structure have a positive enhancement effect on copywriting generation, and the generated text has higher relevance and diversity with the reference words, which is a reliable and practical copywriting generation scheme.

[0053] In summary, the commodity copywriting generation method based on the conditional variational autoencoder and retrieval enhancement designed by the present invention meets the relevance and diversity of copywriting generation, thus having an actual application scenario.

[0054] The embodiments of the present invention have been described in detail above with reference to the drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. A method for generating product copy based on conditional variational autoencoders and retrieval enhancement, characterized in that Execute steps S1 - S6 to complete the generation of product copywriting: Step S1: Extract keywords from the description of the product copywriting requirements input by the user. Step S2: Using the keywords as input, construct an inference module based on the graph neural network model. Employ multi-hop graph reasoning to generate candidate inference paths. The candidate inference paths undergo path embedding, path pruning, and score calculation. Select several candidate inference paths as the output paths according to the scores, and extract and output implicit keywords. Step S3: Using the keywords and implicit keywords as input, construct a generation module composed of an embedding layer, an encoder, a reparameterization module, a decoder, and an output layer. The generation module integrates the conditional variational autoencoder structure and the ON-LSTM gated hierarchical structure, and outputs several candidate product copywritings. Step S4: Construct an evaluation module based on the large language model, comprehensively score the candidate product copywritings, and select from the candidate product copywritings according to the comprehensive score as the product copywriting output by the evaluation module. Step S5: Construct a memory module to store the user input and the generated product copywriting. Step S6: Based on the thinking tree structure with a preset number of layers, construct a prompt engineering module, call the memory module, inference module, generation module, and evaluation module, and perform iterations for a preset number of rounds to obtain the final product copywriting.

2. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 1, wherein The specific steps of Step S2 are as follows: Step S2.1: Assume the knowledge graph is , where is the entity set, is the edge set; set the inference depth parameter , and perform rounds of multi-hop graph reasoning. In the i -th layer propagation in the knowledge graph, the representation of the entity node is obtained by weighted aggregation of the information of its neighbor nodes , where represents the neighbor node set of the node . The formula for information weighted aggregation is as follows: ; Among them, is the representation of the entity node ; represents the connectivity relationship between nodes i and j in the adjacency matrix; represents the relation embedding from node i to j ; is the neighbor aggregation weight matrix; is the relation weight matrix; is the activation function; Step S2.2: Given a path in the knowledge graph that starts from the starting entity and passes through multiple intermediate entities until the target entity , and adopt splicing modeling: ​ ; Among them, and are intermediate entities, indicating and 's relationship, indicating and 's relationship, L is the total number of entities; Step S2.3: For each path , perform path pruning, including static pruning and semantic pruning; Static pruning is performed for each path Calculate the path score : ; Among them, is the i th entity, L is the total number of entities; represents the L2 norm; Set a threshold value , and only retain 's path; Semantic pruning is performed on each path and the user input query vector to calculate the cosine similarity : ; Only retain 's path; Step S2.4: Sort each path according to the path score, and starting from the path with the highest score, select several paths in sequence, in the form of: , extract the tail entity, which contains implicit keywords.

3. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 1, wherein The specific steps of Step S3 are as follows: Step S3.1: The embedding layer uses the word embedding function to convert the word Id into a dense vector; adjust the dimension of the dense vector, and the dimension of the dense vector includes sequence length, batch size, and hidden layer dimension; send the embedded vector to the encoder. Step S3.2: The encoder adopts the conditional variational autoencoder structure, and the formula is as follows: ; ; wherein, is the standard deviation of the latent distribution, is the mean of the latent distribution, is the hidden state, is the weight matrix, is the bias vector; Step S3.3: The reparameterization module samples a latent vector from a Gaussian distribution : ; wherein, denotes noise sampled from a standard normal distribution, denotes element-wise multiplication, is the standard deviation of the latent distribution, is the mean of the latent distribution, denotes a normal distribution; Step S3.4: The decoder extends the latent vector to the target sequence length and transmits it through the latent vector to different layers of the ON-LSTM gating hierarchy to generate hidden states and generate a token distribution; Step S3.5: The output layer uses a fully connected layer to map the hidden state to the vocabulary dimension, generates multiple different candidate product copywritings, and calculates the total loss function.

4. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 3, characterized in that The total loss function in Step S3.5 includes the reconstruction loss and the KL divergence loss, and calculates the weighted sum of the two: ; Among them, represents the total loss function, represents the reconstruction loss, represents the KL divergence loss, is the adjustment factor.

5. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 1, wherein In Step S4, the evaluation module calls the large language model to evaluate the candidate product copywritings, adopting two evaluation criteria, the scoring mechanism and the voting mechanism. The scoring mechanism requires the large language model to score the candidate product copywritings, and the score range is between 0 and 1; the voting mechanism requires the large language model to vote uniquely on the candidate product copywritings, and calculates the comprehensive score of the score and the number of votes by weighting. Select from the candidate product copywritings according to the comprehensive score as the product copywriting output by the evaluation module.

6. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 5, wherein In Step S4, if the evaluation module is the last round of iteration, select the candidate product copywriting with the highest comprehensive score as the product copywriting output by the evaluation module. If it is not the last round of iteration, adopt the greedy algorithm Top-K, retain the top K candidate product copywritings with the highest comprehensive score as the product copywriting output by the evaluation module, and enter the next round of iteration.

7. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 1, wherein The memory module in Step S5 uses the MySQL database, converts the user input and the generated product copywriting into the structured triple form of product-keyword-copywriting, and stores it in the MySQL database.

8. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 1, wherein The prompting engineering module in step S6, where the thought tree structure includes a root node, intermediate nodes, leaf nodes, and the number of thought tree layers; If it is the first iteration, the root node is the keyword and the implicit keyword. If it is a subsequent iteration, the root node is the product copywriting generated in the previous iteration and the keyword related to the focus of this iteration preset; The intermediate nodes are the list of candidate product copywritings in the current iteration; The leaf nodes are the final output of the product copywriting; The number of thought tree layers is the number of iterations, preset by the user.

9. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 8, characterized in that In the prompting engineering module in step S6, for each iteration, keywords related to the focus are preset, and the focus includes the product main body, user group, product features and attributes, main style and scenario.

10. The method for generating product copy based on conditional variational autoencoder and retrieval enhancement according to claim 8, wherein In the prompting engineering module in step S6, the steps for constructing and pruning the thought tree structure are as follows: Step S6.1: If it is the first iteration, input the keyword and the implicit keyword output by the reasoning module. If it is a subsequent iteration, call the product copywriting generated in the previous iteration in the memory module and input the keyword related to the focus of this iteration preset; Step S6.2: Call the generation module to generate the list of candidate product copywritings in the current iteration, corresponding to the labels {1, 2,..., N}, where N is the number of candidate product copywritings; Step S6.3: Call the evaluation module to obtain the comprehensive score of each candidate product copywriting; Step S6.4: Use the greedy algorithm Top-K to select the top K candidate product copywritings with the highest comprehensive scores to enter the next iteration; Step S6.5: After completing the preset number of iterations, select the product copywriting with the highest comprehensive score from the candidate product copywritings as the final output of the product copywriting.

Citation Information

Patent Citations

  • Inference method and device for multi-hop question and answer of knowledge graph and storage medium

    CN118503366A

  • Iterative design method and system for cultural and creative products based on user emotion feedback

    CN119295185A

  • Systems and methods using deep joint variational autoencoders

    US20230245204A1

Cited By

  • Multi-module collaborative agent for enhancing credibility of intelligence analysis, equipment and medium

    CN120880722A

  • Effective signal extraction method and system based on CVAE

    CN121765621A

  • Effective signal extraction method and system based on CVAE

    CN121765621B