A copywriting method and device, computer equipment and storage medium
By training a copywriting generation model and utilizing target and obfuscated sample data, the model automatically outputs copywriting that aligns with the target theme, thus solving the problem of low copywriting efficiency and achieving highly efficient copywriting generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-06
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, copywriting is inefficient, requiring professionals to manually write copy to match item attributes, resulting in low efficiency.
By using a pre-trained copywriting generation model, which is trained with target sample data and obfuscated sample data, the model automatically outputs copywriting based on the same target topic, improving writing efficiency and enhancing topic differentiation capabilities.
It achieves automated copy generation, improves copywriting efficiency, and ensures that the generated copy is consistent with the target theme, thereby enhancing the copy generation effect.
Smart Images

Figure CN116383386B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more specifically, to a document generation method, apparatus, computer device, and storage medium. Background Technology
[0002] With the rapid development of internet technology, the ways of promoting products have also changed. Besides traditional methods such as television, radio, and newspapers, promotion can now be done through websites and applications. When promoting products, appropriate copywriting is necessary; excellent copywriting usually achieves good promotional results.
[0003] In related technologies, in order to ensure that the copywriting matches the attributes of the product to be promoted, it is often necessary for professionals with copywriting experience to manually write the copy, which results in relatively low copywriting efficiency and needs to be improved. Summary of the Invention
[0004] This disclosure provides at least one document generation method, apparatus, computer device, and storage medium.
[0005] In a first aspect, embodiments of this disclosure provide a text generation method, including:
[0006] Obtain multiple texts to be selected;
[0007] The multiple texts to be selected are input into a pre-trained text generation model to obtain the target text generated by the text generation model based on the multiple texts to be selected corresponding to the same target theme.
[0008] The copy generation model is trained based on target sample data. Each target sample data is constructed based on a first sample copy and a confused sample copy corresponding to the first sample copy. The confused sample copy includes a second sample copy with the same topic type as the target topic and a third sample copy with a different topic type than the target topic.
[0009] In one possible implementation, the first sample text includes at least one short text.
[0010] The method further includes generating the first sample text according to the following steps:
[0011] Obtain the original long text to be segmented;
[0012] The original long text is segmented according to a preset text segmentation rule to obtain the first sample text after segmentation; or, the original long text is input into a pre-trained text segmentation model to obtain the first sample text after text segmentation of the original long text output by the text segmentation model.
[0013] In one possible implementation, the method further includes constructing the target sample data according to the following steps:
[0014] The first sample text, the second sample text with the same topic type as the first sample text, and the third sample text with a different topic type than the first sample text are combined to obtain a mixed sample text.
[0015] The order of the text in the mixed sample text is randomly adjusted to obtain the target sample data after random adjustment.
[0016] In one possible implementation, the topic type corresponding to the third sample text is a confusing topic type that is easily confused with the topic type of the target topic; the confusing topic type is determined based on a pre-trained topic type recognition model.
[0017] In one possible implementation, the step of inputting the plurality of texts to be selected into a pre-trained text generation model to obtain the target text generated by the text generation model based on the plurality of texts to be selected corresponding to the same target topic includes:
[0018] Based on the word vectors, position vectors, and importance vectors corresponding to each text to be selected, a model input vector corresponding to the multiple texts to be selected is constructed; wherein, the importance vector is determined based on the similarity between each text to be selected and an importance detection text used to characterize the importance of the attribute corresponding to the target topic; the importance detection text includes attribute description information corresponding to the target topic, and / or landing page information corresponding to the target topic;
[0019] The model input vectors corresponding to the multiple texts to be selected are input into the pre-trained text generation model to obtain the target text output by the text generation model.
[0020] In one possible implementation, the method further includes training the copy generation model according to the following steps:
[0021] Obtain target sample data and corresponding sample tags; wherein, the sample tags include the standard text generation results corresponding to the target sample data;
[0022] The target sample data is input into the feature encoding module of the copy generation model to obtain the initial feature vector output by the feature encoding module corresponding to the target sample data;
[0023] The initial feature vector is input into the feature decoding module of the copywriting generation model to obtain the sample copywriting generation result output by the feature decoding module corresponding to the target sample data;
[0024] Based on the sample copy generation results and the standard copy generation results, a first loss value is determined, and the network parameters of the copy generation model to be trained are adjusted based on the first loss value.
[0025] In one possible implementation, the sample label further includes a first sample label for indicating the retention status of each text that makes up the target sample data in the text generation result;
[0026] The method further includes:
[0027] The initial feature vector is input into the classification layer used during training to obtain the text retention results corresponding to each text that makes up the target sample data, as output by the classification layer.
[0028] Based on the text retention results and the first sample label, a second loss value is determined;
[0029] The adjustment of the network parameters of the text generation model to be trained based on the first loss value includes:
[0030] A target loss value is determined based on the first loss value and the second loss value, and the network parameters of the text generation model to be trained are adjusted based on the target loss value.
[0031] Secondly, this disclosure also provides a document generation apparatus, comprising:
[0032] The acquisition module is used to acquire multiple texts to be selected;
[0033] The generation module is used to input the multiple texts to be selected into a pre-trained text generation model to obtain the target text generated by the text generation model based on the multiple texts to be selected corresponding to the same target theme;
[0034] The copy generation model is trained based on target sample data. Each target sample data is constructed based on a first sample copy and a confused sample copy corresponding to the first sample copy. The confused sample copy includes a second sample copy with the same topic type as the target topic and a third sample copy with a different topic type than the target topic.
[0035] In one possible implementation, the first sample text includes at least one short text.
[0036] The generation module is further configured to generate the first sample text according to the following steps:
[0037] Obtain the original long text to be segmented;
[0038] The original long text is segmented according to a preset text segmentation rule to obtain the first sample text after segmentation; or, the original long text is input into a pre-trained text segmentation model to obtain the first sample text after text segmentation of the original long text output by the text segmentation model.
[0039] In one possible implementation, the generation module is further configured to construct the target sample data according to the following steps:
[0040] The first sample text, the second sample text with the same topic type as the first sample text, and the third sample text with a different topic type than the first sample text are combined to obtain a mixed sample text.
[0041] The order of the text in the mixed sample text is randomly adjusted to obtain the target sample data after random adjustment.
[0042] In one possible implementation, the topic type corresponding to the third sample text is a confusing topic type that is easily confused with the topic type of the target topic; the confusing topic type is determined based on a pre-trained topic type recognition model.
[0043] In one possible implementation, the generation module, when inputting the plurality of texts to be selected into a pre-trained text generation model to obtain the target text generated by the text generation model based on the plurality of texts to be selected corresponding to the same target topic, is used to:
[0044] Based on the word vectors, position vectors, and importance vectors corresponding to each text to be selected, a model input vector corresponding to the multiple texts to be selected is constructed; wherein, the importance vector is determined based on the similarity between each text to be selected and an importance detection text used to characterize the importance of the attribute corresponding to the target topic; the importance detection text includes attribute description information corresponding to the target topic, and / or landing page information corresponding to the target topic;
[0045] The model input vectors corresponding to the multiple texts to be selected are input into the pre-trained text generation model to obtain the target text output by the text generation model.
[0046] In one possible implementation, the generation module is further configured to train the copywriting generation model according to the following steps:
[0047] Obtain target sample data and corresponding sample tags; wherein, the sample tags include the standard text generation results corresponding to the target sample data;
[0048] The target sample data is input into the feature encoding module of the copy generation model to obtain the initial feature vector output by the feature encoding module corresponding to the target sample data;
[0049] The initial feature vector is input into the feature decoding module of the copywriting generation model to obtain the sample copywriting generation result output by the feature decoding module corresponding to the target sample data;
[0050] Based on the sample copy generation results and the standard copy generation results, a first loss value is determined, and the network parameters of the copy generation model to be trained are adjusted based on the first loss value.
[0051] In one possible implementation, the sample label further includes a first sample label for indicating the retention status of each text that makes up the target sample data in the text generation result;
[0052] The generation module is also used for:
[0053] The initial feature vector is input into the classification layer used during training to obtain the text retention results corresponding to each text that makes up the target sample data, as output by the classification layer.
[0054] Based on the text retention results and the first sample label, a second loss value is determined;
[0055] The generation module, when adjusting the network parameters of the text generation model to be trained based on the first loss value, is used for:
[0056] A target loss value is determined based on the first loss value and the second loss value, and the network parameters of the text generation model to be trained are adjusted based on the target loss value.
[0057] Thirdly, embodiments of this disclosure also provide a computer device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps of the first aspect above, or any possible implementation of the first aspect, are performed.
[0058] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the first aspect or any possible implementation thereof.
[0059] The copywriting generation method, apparatus, computer equipment, and storage medium provided in this disclosure, through pre-training a copywriting generation model, can automatically output target copywriting based on multiple texts to be selected, corresponding to the same target theme, based on multiple input texts to be selected. Compared with manual copywriting, this improves the efficiency of copywriting. On the other hand, during the training of the copywriting generation model, the model is trained using target sample data constructed based on a first sample text and obfuscated sample texts corresponding to the first sample text. This enables the trained copywriting generation model to have a strong theme discrimination ability, thereby accurately selecting multiple input texts to be selected and generating the target copywriting based on the selected texts. This ensures that the content of the final generated target copywriting corresponds to the same target theme, improving the copywriting generation effect.
[0060] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0061] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.
[0062] Figure 1 A flowchart of a text generation method provided by an embodiment of this disclosure is shown;
[0063] Figure 2 This diagram illustrates the construction of the model input vector in the copy generation method provided in this embodiment of the present disclosure;
[0064] Figure 3 The diagram illustrates the model architecture during the training process in the copy generation method provided in this embodiment of the present disclosure.
[0065] Figure 4 This diagram illustrates the architecture of a copy generation apparatus provided in an embodiment of the present disclosure;
[0066] Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation
[0067] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0068] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0069] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0070] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0071] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0072] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0073] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0074] Research has found that in order to ensure that the copywriting matches the attributes of the product being promoted, it often needs to be manually written by professionals with copywriting experience, resulting in relatively low copywriting efficiency, which needs to be improved.
[0075] Based on the above research, this disclosure provides a copywriting generation method, apparatus, computer device, and storage medium. By pre-training a copywriting generation model, the model can automatically output target copywriting based on multiple input copywritings to be selected, corresponding to the same target theme. Compared with manual copywriting, this improves the efficiency of copywriting. On the other hand, during the training of the copywriting generation model, target sample data constructed based on a first sample copywriting and confused sample copywriting corresponding to the first sample copywriting are used to train the model. This enables the trained model to have a strong theme discrimination ability, thereby accurately selecting multiple input copywritings to be selected and generating the target copywriting based on the selected copywritings. This ensures that the content of the final generated target copywriting corresponds to the same target theme, improving the copywriting generation effect.
[0076] To facilitate understanding of this embodiment, a detailed description of the text generation method disclosed in this disclosure is provided first. The execution entity of the text generation method provided in this disclosure is generally a computer device with a certain computing capability. This computer device may include, for example, a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), mobile device, user terminal, terminal, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. In some possible implementations, this text generation method can be implemented by a processor calling computer-readable instructions stored in memory.
[0077] Example 1
[0078] See Figure 1The diagram shows a flowchart of a text generation method provided in this embodiment of the present disclosure. The method includes steps S101 to S102, wherein:
[0079] S101: Obtain multiple texts to be filtered.
[0080] Here, the text to be screened has a corresponding theme, which can be understood as a product type. For example, the product type "shirt" is the theme of the text to be screened.
[0081] The text to be filtered can be extracted from the original long text corresponding to the product type. The original long text is the description of the product under the product type. For example, if the product type is "perfume", the description of "perfume" contains specific attribute information describing "perfume", such as the duration of release and the type of fragrance. In this case, one type of text to be filtered extracted from the description can include "duration of release" and another type of text to be filtered can include "fragrance type".
[0082] When extracting the text to be screened from the above introductory text, the introductory text can first be segmented, and then the resulting short texts can be screened to obtain the screened texts. The screening criteria can include: containing target keywords matching the corresponding product type, having a word count greater than a preset first word count and less than a preset second word count, and carrying corresponding text retention tags (which can be added by the user). In this way, by screening the short texts obtained after text segmentation, the number of texts to be screened can be effectively reduced, thereby improving the processing efficiency of the subsequent text generation model when generating target texts, and ultimately improving the overall text generation efficiency.
[0083] Furthermore, the specific process of segmenting the introductory text described above can be found in the subsequent description of the sample text generation process during model training, which will not be elaborated here.
[0084] S102: Input the multiple texts to be selected into a pre-trained text generation model to obtain the target text generated by the text generation model based on the multiple texts to be selected corresponding to the same target theme.
[0085] The copy generation model is trained based on target sample data. Each target sample data is constructed based on a first sample copy and a confused sample copy corresponding to the first sample copy. The confused sample copy includes a second sample copy with the same topic type as the target topic and a third sample copy with a different topic type than the target topic.
[0086] Here, the copy generation model can be a neural network model, specifically a Transformer model with serialization data processing capabilities.
[0087] The multiple candidate texts used to generate the target text correspond to the same target theme (i.e., the same product type). In other words, for the same product type, a unified target text can be obtained by integrating multiple candidate texts.
[0088] In S102, the theme type can be understood as the industry to which the product type belongs. For example, shirts and coats belong to the clothing industry. Then, the second sample copy with the same theme type as the target theme can be a sample copy whose corresponding product type belongs to the same industry as the product type corresponding to the target theme. For example, if the product type corresponding to the target theme is "shirt", then the second sample copy with the same theme type as the target theme can be a sample copy whose corresponding product type is "coat". Both "shirt" and "coat" belong to the "clothing" industry.
[0089] Similarly, a third sample copy with a different theme type than the target theme can be a sample copy whose corresponding product type belongs to a different industry than the product type corresponding to the target theme. For example, if the product type corresponding to the target theme is "shirt", then a third sample copy with a different theme type than the target theme can be a sample copy whose corresponding product type is "milk powder". "Shirt" belongs to the "clothing" industry, while "milk powder" belongs to the "maternal and infant" industry.
[0090] In the above embodiments, by pre-training the copywriting generation model, the model can automatically output target copy based on multiple input copywritings to be selected, corresponding to the same target theme. Compared with manual copywriting, this improves the efficiency of copywriting. On the other hand, during the training of the copywriting generation model, by training the model with target sample data constructed based on the first sample copy and the confused sample copy corresponding to the first sample copy, the trained copywriting generation model can have a strong theme discrimination ability. This allows it to accurately select multiple input copywritings to be selected and generate the target copy based on the selected copywritings. This ensures that the content of the final generated target copy corresponds to the same target theme, improving the copywriting generation effect.
[0091] Example 2
[0092] In Example 1 above, a copywriting generation model was used. During the training process of the copywriting generation model, various first sample copywritings are required. The method for generating the first sample copywritings is described below:
[0093] Method 1: Segment the text according to the preset text segmentation rules.
[0094] Here, the original long text to be segmented can be obtained, and the original long text can be segmented according to the preset text segmentation rules to obtain the first sample text after segmentation.
[0095] The text segmentation rules may include at least one of the following rules:
[0096] ①Segmented according to punctuation marks
[0097] Here, punctuation marks used for segmentation can include commas, periods, etc.
[0098] Specifically, when segmenting according to punctuation marks, the text between the Nth preset target punctuation mark used for segmentation and the (N-1)th preset target punctuation mark used for segmentation can be used as the Nth text after segmentation; among them, all text before the 1st preset target punctuation mark used for segmentation can be used as the 1st text after segmentation.
[0099] For example, taking the target punctuation marks "," and "." as examples, the text to be segmented, "AXXXXXX,BXXXXX.", can be segmented into "AXXXXXX" and "BXXXXX" according to the target punctuation marks.
[0100] ②Segment by capitalization
[0101] Here, since the first letter of the first word of a sentence in English text is often capitalized, while the English content in subsequent sentences is often lowercase, the text can be segmented according to the capitalization of the letters.
[0102] For example, taking the original long text as "You really made me feel at home. I am looking forward to hearing from you again soon.", it can be split according to capitalization to get "You really made me feel at home." and "I am looking forward to hearing from you again soon."
[0103] ③ Segmentation based on stop words
[0104] Here, the stop words are words that need to be filtered during natural language data processing, and may include pronouns, prepositions, conjunctions, articles, etc.
[0105] For example, taking the original long text as "Because CXXXXX, therefore DXXXXX", since "because" and "therefore" are conjunctions among stop words, the long text can be divided into "CXXXXX" and "DXXXXX".
[0106] ④ Segment according to conjunctions
[0107] Here, the conjunctions are used to connect sentence components in a sentence, and the conjunctions may be, for example, and, or.
[0108] For example, if the original long copy is "EXXXXandFXXXX", it can be divided into "EXXXX" and "FXXXX" according to the conjunctions.
[0109] In this way, by using the above-mentioned multiple text segmentation rules, after obtaining the original long text to be segmented, the text segmentation rule that matches the Chinese and English recognition results can be selected to segment the text based on the Chinese and English recognition results corresponding to the long text, thereby obtaining the first sample text after text segmentation processing.
[0110] Method 2: Use a text segmentation model for segmentation.
[0111] Here, the original long text to be segmented can be obtained and input into a pre-trained text segmentation model to obtain the first sample text output by the text segmentation model after text segmentation processing of the original long text.
[0112] Here, the text segmentation model can be a Natural Language Processing (NLP) model, which can perform text segmentation processing on the original long text based on the semantics of each word in the input original long text.
[0113] The text segmentation model can be trained using supervised training. During the training process, a sample long text can be input into the text segmentation model to obtain the sample segmentation result obtained by the text segmentation model after segmenting the sample long text. Then, based on the sample segmentation result and the sample label corresponding to the sample long text (the segmented sample short text corresponding to the sample long text), the loss value of this training is calculated, so that the network parameters of the text segmentation model can be adjusted based on the loss value.
[0114] Example 3
[0115] Regarding the training process of the above copywriting generation model, in addition to using the first sample copywriting involved in each embodiment, a second sample copywriting and a third sample copywriting are also used. The following describes the process of constructing target sample data containing the first sample copywriting, the second sample copywriting, and the third sample copywriting:
[0116] A1: Combine the first sample text, the second sample text with the same theme type as the first sample text, and the third sample text with a different theme type than the first sample text to obtain a mixed sample text.
[0117] A2: Randomly adjust the order of the text in the mixed sample text to obtain the randomly adjusted target sample data.
[0118] For example, taking the original long text as s0, the first sample text can be obtained by segmenting the original long text. The second sample text, which has the same topic type as the first sample text, can be The third sample text, which corresponds to a different topic type than the first sample text, can be By combining the first sample text, the second sample text with the same topic type as the first sample text, and the third sample text with a different topic type than the first sample text, a mixed sample text can be obtained. Furthermore, after randomly adjusting the text order in the mixed sample text, the randomly adjusted target sample data is obtained.
[0119] In this way, by combining the first sample text, the second sample text with the same topic type as the first sample text, and the third sample text with a different topic type than the first sample text, other sample texts that can be confused with the first sample text can be combined into the final target sample data. This allows the final target sample data to be used to train the text classification ability of the text generation model. On the other hand, by randomly adjusting the text order in the mixed sample text, the final target sample data, after being used to train the text generation model, enables the trained text generation model to have the ability to adjust the natural sentence order, thereby improving the model performance of the text generation model.
[0120] When constructing the target sample data, the topic type corresponding to the third sample text can be an easily confused topic type that is easily confused with the topic type of the target topic; the easily confused topic type is determined based on a pre-trained topic type recognition model.
[0121] Here, the topic type recognition model is used to determine the topic type (i.e., industry) corresponding to the input product attribute text. For example, based on the input product attribute text corresponding to the product type "shirt", the topic type corresponding to the product attribute text is determined to be "clothing". The topic type recognition model can be a neural network model, specifically an NLP model.
[0122] The topic type recognition model can be trained using supervised training. During the training process, sample attribute text can be input into the topic type recognition model to obtain the topic type recognition result obtained by the topic type recognition model after the sample attribute text is identified. Then, based on the topic type recognition result and the sample label (topic type annotation result corresponding to the sample attribute text) corresponding to the sample attribute text, the loss value of this training is calculated, so that the network parameters of the topic type recognition model can be adjusted based on the loss value.
[0123] In one possible implementation, the topic type identification result corresponding to the first sample text and the topic type identification result corresponding to the third sample text can meet a preset similarity condition.
[0124] The topic type identification result can be in the form of a probability distribution, which represents the probability value of each preset type in the topic type identification result. The preset similarity condition can be that the similarity between the corresponding probability distribution forms is greater than a preset similarity threshold.
[0125] In this way, by adding third sample texts corresponding to easily confused topic types that are easily confused with the target topic to the target sample data used to train the text generation model, the final target sample data can be used to train the text classification ability of the text generation model, so that the trained text generation model can distinguish easily confused topic types, thereby improving the model performance of the text generation model.
[0126] Example 4
[0127] In the above embodiments, the target text is obtained by inputting the multiple texts to be selected into a pre-trained text generation model. In specific implementation, feature vectors can be extracted from the texts to be selected first, and then the extracted feature vectors can be input into the text generation model. The specific process is as follows:
[0128] B1: Construct model input vectors for the multiple texts to be selected based on the word vectors, position vectors, and importance vectors corresponding to each text to be selected; wherein, the importance vector is determined based on the similarity between each text to be selected and the importance detection text used to characterize the importance of the attribute corresponding to the target topic; the importance detection text includes the attribute description information corresponding to the target topic, and / or the landing page information corresponding to the target topic.
[0129] Here, the model input vector can be composed of word vectors, position vectors, and importance vectors. The word vectors are obtained by extracting features from the text to be screened and are used to represent the semantics of the text content. Since the Transformer model cannot capture the sequence order information of the sequence data, position vectors representing the sequence order information need to be added to the sequence data during the input process. These position vectors represent the input order of each text to be screened. Therefore, the importance vector represents the importance of each text to be screened. The value of the importance vector can include "0" and "1", where "0" indicates that the corresponding text to be screened is not important and its content can be discarded (i.e., the generated target text may not contain the content corresponding to the text to be screened); "1" indicates that the corresponding text to be screened is important and its content needs to be retained (i.e., the generated target text needs to contain the content corresponding to the text to be screened).
[0130] Specifically, when determining the importance vector corresponding to each text to be screened, the similarity between each text to be screened and the importance detection text can be calculated; wherein, the similarity may include text similarity and / or semantic similarity.
[0131] For example, a schematic diagram of constructing the model input vector can be shown as follows: Figure 2 As shown, Figure 2 In the model, the importance vectors corresponding to each text to be selected are 0, 0, and 1; the position vectors corresponding to each text to be selected are 0, 1, and 2; and the word vectors corresponding to each text to be selected are Tok1, Tok2, and Tok3. For any text to be selected, the values of the importance vector, position vector, and word vector corresponding to the text to be selected can be added together and then concatenated according to the concatenation order corresponding to each text to be selected to obtain the model input vector.
[0132] Furthermore, the importance vector can be manually labeled, thereby allowing intervention in the output of the copywriting generation model to ensure that the generated target copy contains the set content that needs to be retained. This allows the output of the copywriting generation model to be adjusted according to actual needs, meeting the personalized needs of users in practical applications.
[0133] B2: Input the model input vectors corresponding to the multiple texts to be selected into the pre-trained text generation model to obtain the target text output by the text generation model.
[0134] Here, when inputting the model input vectors corresponding to the multiple texts to be screened into the pre-trained text generation model, the model input vectors can first be input into the feature encoding module of the text generation model to obtain the initial feature vector output by the feature encoding module; the initial feature vector is then input into the feature decoding module of the text generation model to obtain the target text output by the feature decoding module.
[0135] In this way, by introducing the importance vector into the model input vector of the constructed input to the copywriting generation model, the feature dimensions of the model input vector can be enriched. This allows the copywriting generation model to selectively retain the content of each copy to be screened based on the importance vector, thereby improving the correlation between the final generated target copy and the corresponding attributes of the target topic. On the other hand, by allowing the setting of specific parameters of the importance vector, users can also adjust the output of the copywriting generation model according to actual needs, meeting the personalized needs of users in practical applications.
[0136] Example 5
[0137] In the above embodiments, the construction of target sample data (first sample data, second sample data, and third sample data) has been completed. Based on this, the copywriting generation model can be specifically trained using the target sample data. The specific training process may include:
[0138] C1: Obtain the target sample data and the sample tags corresponding to the sample data; wherein, the sample tags include the standard text generation results corresponding to the target sample data.
[0139] C2: Input the target sample data into the feature encoding module of the copy generation model to obtain the initial feature vector output by the feature encoding module corresponding to the target sample data.
[0140] C3: Input the initial feature vector into the feature decoding module of the copywriting generation model to obtain the sample copywriting generation result output by the feature decoding module corresponding to the target sample data.
[0141] C4: Based on the sample copy generation results and the standard copy generation results, determine the first loss value, and adjust the network parameters of the copy generation model to be trained based on the first loss value.
[0142] Here, when determining the first loss value based on the sample copy generation result and the standard copy generation result, the first loss value generated based on the sample copy generation result and the standard copy generation result can be calculated according to a preset first loss function; wherein, the first loss function can be the negative log-likelihood (NLL) loss function.
[0143] For example, the first loss function can be:
[0144]
[0145] Where X represents the input to the copywriting generation model; θ is the hyperparameter of the copywriting generation model; y <j This refers to the text already generated during the process of generating the target text; y j This represents the text content (characters or words) generated in the j-th step of the text generation model; n represents the output length.
[0146] In this way, by using the generated copy to determine the supervision data for the next step of copy generation during the copy generation process of the copy generation model, the loss value corresponding to each step can be calculated based on the supervision data, and the first loss value can be calculated based on the loss value corresponding to each step.
[0147] Example 6
[0148] This embodiment introduces another loss function; that is, two loss functions are used to calculate the loss value during the training process of the copywriting generation model to improve the model training accuracy. The specific process of training the model based on the two loss functions includes:
[0149] D1: Obtain sample data and sample tags corresponding to the sample data; wherein, the sample tags include the standard copy generation result corresponding to the sample data, and a first sample tag used to indicate the retention status of each copy that makes up the sample data in the copy generation result.
[0150] D2: Input the sample data into the feature encoding module of the copywriting generation model to obtain the initial feature vector output by the feature encoding module corresponding to the sample data.
[0151] D3: Input the initial feature vector into the feature decoding module of the copywriting generation model to obtain the sample copywriting generation result corresponding to the sample data output by the feature decoding module; and input the initial feature vector into the classification layer used in the training process to obtain the copywriting retention result corresponding to each copywriting that makes up the sample data output by the classification layer.
[0152] D4: Based on the sample copy generation result and the standard copy generation result, determine the first loss value; and based on the copy retention result and the first sample label, determine the second loss value.
[0153] D5: Determine the target loss value based on the first loss value and the second loss value, and adjust the network parameters of the text generation model to be trained based on the target loss value.
[0154] Here, when determining the second loss value based on the text retention result and the first sample label, the second loss value generated based on the text retention result and the first sample label can be calculated according to a preset second loss function; wherein, the second loss function can be the cross entropy (CE) loss function.
[0155] For example, the second loss function can be:
[0156]
[0157] Where X represents the input of the copywriting generation model; Φ is the hyperparameter of the feature encoding module in the copywriting generation model; ci represents the classification category corresponding to the i-th copywriting to be selected in the input, which is the supervision data used in the training process; and m represents the number of copywritings to be selected in the input.
[0158] For example, the model architecture diagram during the training process can be as follows: Figure 3 As shown, Figure 3 After the sample data is input into the copywriting generation model, feature extraction processing can be performed based on the feature encoding module. The generated initial feature vector corresponding to the sample data is then input into the feature decoding module and the classification layer, respectively. This yields the sample copywriting generation result output by the feature decoding module and the copywriting retention result output by the classification layer. Based on the first loss value determined by the sample copywriting generation result and the second loss value determined by the copywriting retention result, the final target loss value used to adjust the network parameters of the copywriting generation model can be determined.
[0159] The copywriting generation method provided in this disclosure, by pre-training a copywriting generation model, can automatically output target copywriting based on multiple texts to be selected, corresponding to the same target theme, based on multiple input texts to be selected. Compared with manual copywriting, this improves the efficiency of copywriting. On the other hand, during the training of the copywriting generation model, the model is trained by constructing target sample data based on a first sample text and obfuscated sample texts corresponding to the first sample text. This enables the trained copywriting generation model to have a strong theme discrimination ability, thereby accurately selecting multiple input texts to be selected and generating the target copywriting based on the selected texts. This ensures that the content of the final generated target copywriting corresponds to the same target theme, improving the copywriting generation effect.
[0160] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0161] Based on the same inventive concept, this disclosure also provides a copywriting generation device corresponding to the copywriting generation method. Since the principle of the device in this disclosure for solving the problem is similar to the copywriting generation method described above in this disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0162] Reference Figure 4 The diagram shown is an architectural schematic of a document generation device provided in an embodiment of this disclosure. The device includes: an acquisition module 401 and a generation module 402; wherein,
[0163] Module 401 is used to retrieve multiple texts to be filtered;
[0164] The generation module 402 is used to input the multiple texts to be selected into a pre-trained text generation model to obtain the target text generated by the text generation model based on the multiple texts to be selected corresponding to the same target theme.
[0165] The copy generation model is trained based on target sample data. Each target sample data is constructed based on a first sample copy and a confused sample copy corresponding to the first sample copy. The confused sample copy includes a second sample copy with the same topic type as the target topic and a third sample copy with a different topic type than the target topic.
[0166] In one possible implementation, the first sample text includes at least one short text.
[0167] The generation module 402 is further configured to generate the first sample text according to the following steps:
[0168] Obtain the original long text to be segmented;
[0169] The original long text is segmented according to a preset text segmentation rule to obtain the first sample text after segmentation; or, the original long text is input into a pre-trained text segmentation model to obtain the first sample text after text segmentation of the original long text output by the text segmentation model.
[0170] In one possible implementation, the generation module 402 is further configured to construct the target sample data according to the following steps:
[0171] The first sample text, the second sample text with the same topic type as the first sample text, and the third sample text with the same topic type as the first sample text are combined to obtain a mixed sample text.
[0172] The order of the text in the mixed sample text is randomly adjusted to obtain the target sample data after random adjustment.
[0173] In one possible implementation, the topic type corresponding to the third sample text is a confusing topic type that is easily confused with the topic type of the target topic; the confusing topic type is determined based on a pre-trained topic type recognition model.
[0174] In one possible implementation, the generation module 402, when inputting the plurality of texts to be selected into a pre-trained text generation model to obtain the target text generated by the text generation model based on the plurality of texts to be selected corresponding to the same target topic, is used to:
[0175] Based on the word vectors, position vectors, and importance vectors corresponding to each text to be selected, a model input vector corresponding to the multiple texts to be selected is constructed; wherein, the importance vector is determined based on the similarity between each text to be selected and an importance detection text used to characterize the importance of the attribute corresponding to the target topic; the importance detection text includes attribute description information corresponding to the target topic, and / or landing page information corresponding to the target topic;
[0176] The model input vectors corresponding to the multiple texts to be selected are input into the pre-trained text generation model to obtain the target text output by the text generation model.
[0177] In one possible implementation, the generation module 402 is further configured to train the copywriting generation model according to the following steps:
[0178] Obtain target sample data and corresponding sample tags; wherein, the sample tags include the standard text generation results corresponding to the target sample data;
[0179] The target sample data is input into the feature encoding module of the copy generation model to obtain the initial feature vector output by the feature encoding module corresponding to the target sample data;
[0180] The initial feature vector is input into the feature decoding module of the copywriting generation model to obtain the sample copywriting generation result output by the feature decoding module corresponding to the target sample data;
[0181] Based on the sample copy generation results and the standard copy generation results, a first loss value is determined, and the network parameters of the copy generation model to be trained are adjusted based on the first loss value.
[0182] In one possible implementation, the sample label further includes a first sample label for indicating the retention status of each text that makes up the target sample data in the text generation result;
[0183] The generation module 402 is further configured to:
[0184] The initial feature vector is input into the classification layer used during training to obtain the text retention results corresponding to each text that makes up the target sample data, as output by the classification layer.
[0185] Based on the text retention results and the first sample label, a second loss value is determined;
[0186] The generation module 402, when adjusting the network parameters of the text generation model to be trained based on the first loss value, is used for:
[0187] A target loss value is determined based on the first loss value and the second loss value, and the network parameters of the text generation model to be trained are adjusted based on the target loss value.
[0188] The copywriting generation device provided in this embodiment pre-trains a copywriting generation model. This model can automatically output target copy based on multiple input copywriting samples corresponding to the same target theme, improving writing efficiency compared to manual copywriting. Furthermore, during the training of the copywriting generation model, target sample data constructed based on a first sample copy and corresponding obfuscated sample copy is used to train the model. This enables the trained model to have strong theme discrimination capabilities, accurately filtering multiple input copywriting samples and generating the target copy based on the filtered samples. This ensures that the content of the final target copy corresponds to the same target theme, improving the copywriting generation effect.
[0189] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.
[0190] Based on the same technical concept, this disclosure also provides a computer device. (See also...) Figure 5 The diagram shows the structure of a computer device 500 provided in this embodiment of the present disclosure, including a processor 501, a memory 502, and a bus 503. The memory 502 stores execution instructions and includes main memory 5021 and external memory 5022. The main memory 5021, also called internal memory, is used to temporarily store computational data in the processor 501 and data exchanged with external memory 5022 such as a hard disk. The processor 501 exchanges data with the external memory 5022 through the main memory 5021. When the computer device 500 is running, the processor 501 and the memory 502 communicate through the bus 503, causing the processor 501 to execute the following instructions:
[0191] Obtain multiple texts to be selected;
[0192] The multiple texts to be selected are input into a pre-trained text generation model to obtain the target text generated by the text generation model based on the multiple texts to be selected corresponding to the same target theme.
[0193] The copy generation model is trained based on target sample data. Each target sample data is constructed based on a first sample copy and a confused sample copy corresponding to the first sample copy. The confused sample copy includes a second sample copy with the same topic type as the target topic and a third sample copy with a different topic type than the target topic.
[0194] In one possible implementation, the instructions of the processor 501 include at least one short text.
[0195] It also includes generating the first sample text according to the following steps:
[0196] Obtain the original long text to be segmented;
[0197] The original long text is segmented according to a preset text segmentation rule to obtain the first sample text after segmentation; or, the original long text is input into a pre-trained text segmentation model to obtain the first sample text after text segmentation of the original long text output by the text segmentation model.
[0198] In one possible implementation, the instructions of the processor 501 further include constructing the target sample data according to the following steps:
[0199] The first sample text, the second sample text with the same topic type as the first sample text, and the third sample text with the same topic type as the first sample text are combined to obtain a mixed sample text.
[0200] The order of the text in the mixed sample text is randomly adjusted to obtain the target sample data after random adjustment.
[0201] In one possible implementation, in the instructions of the processor 501, the topic type corresponding to the third sample text is an easily confused topic type that is easily confused with the topic type of the target topic; the easily confused topic type is determined based on a pre-trained topic type recognition model.
[0202] In one possible implementation, the instructions of the processor 501, wherein inputting the plurality of texts to be selected into a pre-trained text generation model to obtain target texts generated by the text generation model based on the plurality of texts to be selected corresponding to the same target topic, includes:
[0203] Based on the word vectors, position vectors, and importance vectors corresponding to each text to be selected, a model input vector corresponding to the multiple texts to be selected is constructed; wherein, the importance vector is determined based on the similarity between each text to be selected and an importance detection text used to characterize the importance of the attribute corresponding to the target topic; the importance detection text includes attribute description information corresponding to the target topic, and / or landing page information corresponding to the target topic;
[0204] The model input vectors corresponding to the multiple texts to be selected are input into the pre-trained text generation model to obtain the target text output by the text generation model.
[0205] In one possible implementation, the instructions of the processor 501 further include training the copywriting generation model according to the following steps:
[0206] Obtain target sample data and corresponding sample tags; wherein, the sample tags include the standard text generation results corresponding to the target sample data;
[0207] The target sample data is input into the feature encoding module of the copy generation model to obtain the initial feature vector output by the feature encoding module corresponding to the target sample data;
[0208] The initial feature vector is input into the feature decoding module of the copywriting generation model to obtain the sample copywriting generation result output by the feature decoding module corresponding to the target sample data;
[0209] Based on the sample copy generation results and the standard copy generation results, a first loss value is determined, and the network parameters of the copy generation model to be trained are adjusted based on the first loss value.
[0210] In one possible implementation, the instructions of the processor 501 further include a first sample label for indicating the retention status of each text that makes up the target sample data in the text generation result;
[0211] Also includes:
[0212] The initial feature vector is input into the classification layer used during training to obtain the text retention results corresponding to each text that makes up the target sample data, as output by the classification layer.
[0213] Based on the text retention results and the first sample label, a second loss value is determined;
[0214] The adjustment of the network parameters of the text generation model to be trained based on the first loss value includes:
[0215] A target loss value is determined based on the first loss value and the second loss value, and the network parameters of the text generation model to be trained are adjusted based on the target loss value.
[0216] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the document generation method described in the above method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0217] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the text generation method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.
[0218] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0219] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0220] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0221] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0222] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0223] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. A method for generating copy, characterized in that, include: Obtain multiple texts to be selected; The multiple texts to be selected are input into a pre-trained text generation model to obtain the target text generated by the text generation model based on the multiple texts to be selected corresponding to the same target theme. The copy generation model is trained based on target sample data. Each target sample data is constructed from a first sample copy and a corresponding confusing sample copy. The confusing sample copy includes a second sample copy with the same topic type as the target topic, and a third sample copy with a different topic type. The topic type corresponding to the third sample copy is a easily confused topic type, which is easily confused with the topic type of the target topic. The easily confused topic type is determined based on a pre-trained topic type recognition model. The step of inputting the multiple texts to be selected into a pre-trained text generation model to obtain the target text generated by the text generation model based on the multiple texts to be selected corresponding to the same target topic includes: Based on the word vectors, position vectors, and importance vectors corresponding to each text to be selected, a model input vector corresponding to the multiple texts to be selected is constructed; wherein, the importance vector is determined based on the similarity between each text to be selected and an importance detection text used to characterize the importance of the attribute corresponding to the target topic; the importance detection text includes attribute description information corresponding to the target topic, and / or landing page information corresponding to the target topic; The model input vectors corresponding to the multiple texts to be selected are input into the pre-trained text generation model to obtain the target text output by the text generation model.
2. The method according to claim 1, characterized in that, The first sample copy includes at least one short copy; The method further includes generating the first sample text according to the following steps: Obtain the original long text to be segmented; The original long text is segmented according to a preset text segmentation rule to obtain the first sample text after segmentation; or, the original long text is input into a pre-trained text segmentation model to obtain the first sample text after text segmentation of the original long text output by the text segmentation model.
3. The method according to claim 1 or 2, characterized in that, The method further includes constructing the target sample data according to the following steps: The first sample text, the second sample text with the same topic type as the first sample text, and the third sample text with a different topic type than the first sample text are combined to obtain a mixed sample text. The order of the text in the mixed sample text is randomly adjusted to obtain the target sample data after random adjustment.
4. The method according to claim 1, characterized in that, The method further includes training the copy generation model according to the following steps: Obtain target sample data and corresponding sample tags; wherein, the sample tags include the standard text generation results corresponding to the target sample data; The target sample data is input into the feature encoding module of the copy generation model to obtain the initial feature vector output by the feature encoding module corresponding to the target sample data; The initial feature vector is input into the feature decoding module of the copywriting generation model to obtain the sample copywriting generation result output by the feature decoding module corresponding to the target sample data; Based on the sample copy generation results and the standard copy generation results, a first loss value is determined, and the network parameters of the copy generation model to be trained are adjusted based on the first loss value.
5. The method according to claim 4, characterized in that, The sample label also includes a first sample label used to indicate the retention status of each text that makes up the target sample data in the text generation result; The method further includes: The initial feature vector is input into the classification layer used during training to obtain the text retention results corresponding to each text that makes up the target sample data, as output by the classification layer. Based on the text retention results and the first sample label, a second loss value is determined; The adjustment of the network parameters of the text generation model to be trained based on the first loss value includes: A target loss value is determined based on the first loss value and the second loss value, and the network parameters of the text generation model to be trained are adjusted based on the target loss value.
6. A copywriting generation device, characterized in that, include: The acquisition module is used to acquire multiple texts to be selected; The generation module is used to input the multiple texts to be selected into a pre-trained text generation model to obtain the target text generated by the text generation model based on the multiple texts to be selected corresponding to the same target theme. The copy generation model is trained based on target sample data. Each target sample data is constructed from a first sample copy and a corresponding confusing sample copy. The confusing sample copy includes a second sample copy with the same topic type as the target topic, and a third sample copy with a different topic type. The topic type corresponding to the third sample copy is a easily confused topic type, which is easily confused with the topic type of the target topic. The easily confused topic type is determined based on a pre-trained topic type recognition model. The step of inputting the multiple texts to be selected into a pre-trained text generation model to obtain the target text generated by the text generation model based on the multiple texts to be selected corresponding to the same target topic includes: Based on the word vectors, position vectors, and importance vectors corresponding to each text to be selected, a model input vector corresponding to the multiple texts to be selected is constructed; wherein, the importance vector is determined based on the similarity between each text to be selected and an importance detection text used to characterize the importance of the attribute corresponding to the target topic; the importance detection text includes attribute description information corresponding to the target topic, and / or landing page information corresponding to the target topic; The model input vectors corresponding to the multiple texts to be selected are input into the pre-trained text generation model to obtain the target text output by the text generation model.
7. A computer device, characterized in that, include: The computer device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the document generation method as described in any one of claims 1 to 5 are performed.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the document generation method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Copywriting generation method and device, copywriting evaluation model training method and device, and equipment
CN112232067A