A text generation method and apparatus

By leveraging knowledge graphs and text generation models, along with short text and linking generation models, the problem of high marketing text generation costs has been solved. This enables large-scale customized generation of marketing texts with controllable quality, capturing social hotspots and reflecting consumer sentiment.

CN115345135BActive Publication Date: 2026-05-01BEIJING XUEZHITU NETWORK TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING XUEZHITU NETWORK TECH
Filing Date
2022-05-06
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Marketing text generation is costly and difficult to customize quickly and in large quantities; the application of AI text generation technology is insufficient.

Method used

By acquiring knowledge graphs and text generation models, and utilizing short text generation and linking generation models, customized marketing texts are generated, including the processing of entity information and sentiment classification, and the generation of linking sentences.

Benefits of technology

It enables the generation of large-scale, customized marketing texts, reduces labor costs, ensures controllable text quality, and captures social hot topics and reflects consumer sentiment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115345135B_ABST
    Figure CN115345135B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a text generation method and device, the method comprising: obtaining an established knowledge graph and a text generation model; the text generation model comprising: a short text generation model and a cohesion generation model; determining a subject of a generated text; selecting a corresponding entity and a sentiment classification corresponding to the entity according to entity information of the subject in the knowledge graph according to a preset rule; inputting the selected entity and the corresponding sentiment classification into the short text generation model, obtaining a plurality of short texts from the short text generation model; inputting adjacent short texts into the cohesion generation model in turn, outputting a cohesion sentence of adjacent short texts from the cohesion generation model; and inserting a corresponding cohesion sentence between a plurality of adjacent short texts to generate a text about the subject. Through the embodiment, a large number of customized required texts are generated, and the labor cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to text generation technology, and more particularly to a text generation method and apparatus. Background Technology

[0002] Marketing text is very common and important in marketing campaigns, but the cost of generating text manually is very high, and it cannot be generated quickly and in large quantities with customization. Therefore, using artificial intelligence for text generation has great application potential. Summary of the Invention

[0003] This application provides a text generation method and apparatus that can generate large quantities of customized text, reducing labor costs.

[0004] This application provides a text generation method, which may include:

[0005] The knowledge graph and text generation model are acquired and established; the text generation model includes: a short text generation model and a linking generation model.

[0006] Determine the main body of the generated text;

[0007] Based on the entity information of the subject in the knowledge graph, select the corresponding entity and the sentiment category corresponding to the entity according to preset rules;

[0008] The selected entity and its corresponding sentiment classification are input into the short text generation model, which then generates multiple short texts. The short text refers to a text with a word count within a preset range.

[0009] The adjacent short texts are sequentially input into the connection generation model, and the connection generation model outputs the connection sentences of the adjacent short texts.

[0010] By inserting corresponding connecting sentences between multiple adjacent short texts, text about the subject is generated.

[0011] In an exemplary embodiment of this application, the entity information may include any one or more of the following: the entity types contained in the subject;

[0012] One or more entities corresponding to each entity type;

[0013] The probability of occurrence for each entity; and,

[0014] Each entity corresponds to a sentiment category.

[0015] In an exemplary embodiment of this application, the preset rules may include: one or more first entity types contained in the short text to be generated, and the sorting of multiple first entity types;

[0016] The step of selecting the corresponding entity according to the entity information of the subject in the knowledge graph and according to preset rules may include:

[0017] Obtain the first entity type corresponding to the subject from the knowledge graph;

[0018] Extract entities whose occurrence probability is greater than or equal to a preset probability threshold from all entities corresponding to the first entity type.

[0019] In an exemplary embodiment of this application, the method may further include:

[0020] The generated text of the main body is then segmented into sentences;

[0021] Calculate the repetition rate of each sentence in the text;

[0022] When the repetition rate is greater than or equal to a preset repetition rate threshold, the sentence is deleted.

[0023] In an exemplary embodiment of this application, pre-establishing the knowledge graph may include: constructing knowledge graphs for different products on different platforms and / or in different industries using the following steps:

[0024] Obtain a first preset model; the first preset model may include: a named entity recognition model, an entity relationship recognition model, and an entity sentiment classification model;

[0025] Collect text input from preset entities on preset platforms and / or in preset industries into the first preset model;

[0026] The named entity recognition model outputs the entity type corresponding to the subject; the entity relationship recognition model outputs the relationship between entities; and the entity sentiment classification model outputs the sentiment classification corresponding to each entity.

[0027] The probability of occurrence of each entity is calculated, and the knowledge graph of the subject is constructed from the entity type of the subject, the entities corresponding to each entity type, and the occurrence probability and / or sentiment classification of each entity.

[0028] In an exemplary embodiment of this application, the method may further include: re-collecting the text of a preset subject of the preset platform and / or industry according to a preset period and inputting it into the first preset model to periodically update the knowledge graph.

[0029] In an exemplary embodiment of this application, pre-establishing the short text generation model may include:

[0030] First training data is obtained based on the first preset model;

[0031] The short text generation model is obtained by training the second preset model using the first training data;

[0032] The pre-established connection generation model may include:

[0033] Obtain the second training data;

[0034] The connection generation model is obtained by training the third preset model using the second training data.

[0035] In an exemplary embodiment of this application, obtaining the first training data based on the first preset model may include:

[0036] Filter the texts in the text library that meet the criteria based on keywords;

[0037] The analyzed text is input into a pre-built named entity recognition model, entity relationship recognition model, and entity sentiment classification model to obtain entities and corresponding sentiment classifications for one or more products.

[0038] The entities of the one or more products and their corresponding sentiment classifications are labeled as one or more short texts, forming the first training data.

[0039] In an exemplary embodiment of this application, obtaining the second training data may include:

[0040] Obtain multiple texts with a preset text structure; each text consists of a first short text, a second short text, and a third short text, wherein the first short text, the second short text, and the third short text are connected in sequence;

[0041] Each of the second short texts is marked as a connecting sentence, and all the marked texts are used as the second training data.

[0042] This application also provides a text generation apparatus, which may include a processor and a computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed by the processor, implement the text generation method.

[0043] Compared with related technologies, the embodiments of this application may include: acquiring an established knowledge graph and a text generation model; the text generation model includes a short text generation model and a cohesion generation model; determining the subject of the generated text; selecting the corresponding entity and its corresponding sentiment classification according to preset rules based on the entity information of the subject in the knowledge graph; inputting the selected entity and its corresponding sentiment classification into the short text generation model to obtain multiple short texts; sequentially inputting adjacent short texts into the cohesion generation model to output connecting sentences between adjacent short texts; inserting corresponding connecting sentences between multiple adjacent short texts to generate text about the subject. This embodiment achieves large-scale, customized generation of required text, reducing labor costs.

[0044] Other features and advantages of this application will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the application. Other advantages of this application can be realized and obtained by means of the solutions described in the description and the accompanying drawings. Attached Figure Description

[0045] The accompanying drawings are used to provide an understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.

[0046] Figure 1 This is a flowchart of a text generation method according to an embodiment of this application;

[0047] Figure 2 This is a schematic diagram of a text generation method according to an embodiment of this application;

[0048] Figure 3 This is a flowchart illustrating a method for pre-establishing the knowledge graph according to an embodiment of this application.

[0049] Figure 4 This is a block diagram of the text generation apparatus according to an embodiment of this application. Detailed Implementation

[0050] This application describes several embodiments, but these descriptions are exemplary and not restrictive, and it will be apparent to those skilled in the art that many more embodiments and implementations are possible within the scope of the embodiments described herein. Although many possible combinations of features are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically limited, any feature or element of any embodiment may be used in combination with, or may replace, any feature or element of any other embodiment.

[0051] This application includes and contemplates combinations of features and elements known to those skilled in the art. The embodiments, features, and elements disclosed in this application may also be combined with any conventional features or elements to form a unique inventive scheme as defined by the claims. Any feature or element of any embodiment may also be combined with features or elements from other inventive schemes to form another unique inventive scheme as defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in this application may be implemented individually or in any suitable combination. Therefore, the embodiments are not limited except by the limitations imposed by the appended claims and their equivalents. Furthermore, various modifications and changes may be made within the scope of the appended claims.

[0052] Furthermore, in describing representative embodiments, the specification may have presented methods and / or processes as a specific sequence of steps. However, the method or process should not be limited to the specific order of steps described herein, to the extent that it does not depend on such a specific order. As will be understood by those skilled in the art, other sequences of steps are also possible. Therefore, the specific order of steps set forth in the specification should not be construed as a limitation of the claims. Moreover, the claims concerning the method and / or process should not be limited to the steps performed in the written order, and those skilled in the art will readily understand that these orders can be varied and still remain within the spirit and scope of the embodiments of this application.

[0053] This application provides a text generation method, such as Figure 1 , Figure 2 As shown, the method includes steps S101-S105:

[0054] S101. Obtain the established knowledge graph and text generation model; the text generation model includes: a short text generation model and a linking generation model.

[0055] In an exemplary embodiment of this application, a pre-established knowledge graph and text generation model corresponding to the selected platform and / or industry can be selected based on the chosen platform and / or industry.

[0056] In the exemplary embodiments of this application, the platform may include, but is not limited to, any communication platform, shopping platform, learning platform, etc., and the industry may include, but is not limited to, the beauty industry, the home decoration and building materials industry, the food industry, etc. For example, this application may use communication platforms and the beauty industry as examples for illustration.

[0057] In an exemplary embodiment of this application, after identifying the communication platform and the beauty industry, a knowledge graph and text generation model about the beauty industry on the communication platform can be obtained.

[0058] In an exemplary embodiment of this application, knowledge graphs and text generation models for various industries can be pre-generated for different platforms.

[0059] In exemplary embodiments of this application, as Figure 3 As shown, pre-establishing the knowledge graph may include: constructing knowledge graphs for different products on different platforms and / or in different industries using the following steps S201-S204 respectively:

[0060] S201. Obtain a first preset model; the first preset model may include: a named entity recognition model, an entity relationship recognition model, and an entity sentiment classification model.

[0061] In an exemplary embodiment of this application, the named entity recognition model can be used for named entity recognition; the entity relationship recognition model can be used for entity relationship recognition; and the entity sentiment classification model can be used for entity sentiment classification.

[0062] In an exemplary embodiment of this application, analytical texts that meet certain conditions can be filtered from a massive text library based on keywords. For example, arbitrary keywords, time periods, data source platforms, etc., can be set. Analytical texts that meet the conditions can be selected on the data source platform by inputting keywords, time periods, and other information. These analytical texts can be used for training a first preset model.

[0063] In an exemplary embodiment of this application, for a named entity recognition model, a large amount of acquired analytical text can be labeled with different entity names, and the large amount of analytical text labeled with entity names can be used as training data to train a preset first network model to obtain the named entity recognition model.

[0064] In an exemplary embodiment of this application, for an entity relationship recognition model, a large amount of acquired analytical text can be labeled with relationships between different entities, and the large amount of analytical text labeled with relationships between different entities can be used as training data to train a preset second network model to obtain the entity relationship recognition model.

[0065] In an exemplary embodiment of this application, for an entity sentiment classification model, a large amount of acquired analytical text can be labeled with sentiment classifications of different entities, and the large amount of analytical text labeled with sentiment classifications of different entities can be used as training data to train a preset third network model to obtain the entity sentiment classification model.

[0066] S202. Collect text input from the preset model of the preset platform and / or the preset subject of the industry.

[0067] In an exemplary embodiment of this application, the preset subject may refer to different preset products, such as a product in the beauty industry obtained on a certain communication platform (e.g., which may include, but is not limited to, platforms such as XX Books, XX Music, etc.), such as XX Skin Care Essence.

[0068] In an exemplary embodiment of this application, text about the XX skincare essence can be collected from the communication platform, such as: "XX skincare essence is so good, it visibly repairs wrinkles! It smells like saliva." This text can then be input into the named entity recognition model, entity relationship recognition model, and entity sentiment classification model.

[0069] S203. The named entity recognition model outputs the entity type corresponding to the subject; the entity relationship recognition model outputs the relationship between entities; and the entity sentiment classification model outputs the sentiment classification corresponding to each entity.

[0070] In an exemplary embodiment of this application, the named entity recognition model can recognize the input text "XX skincare essence is so good, it visibly repairs wrinkles! It just smells like saliva" and obtain the following named entity recognition results:

[0071] "XX Skincare Essence Lotion", "Product"

[0072] "Repairs wrinkles", "Efficacy"

[0073] "The taste of saliva" or "flavor".

[0074] In an exemplary embodiment of this application, the identified “XX Skin Care Essence” is an entity, and the corresponding “product” is the entity type of that entity.

[0075] In an exemplary embodiment of this application, the entity relationship recognition model can obtain the following entity relationship recognition results from the input text "XX skincare essence is so good, it visibly repairs wrinkles! It just smells like saliva":

[0076] "XX skincare essence", "repairs wrinkles", "related"

[0077] "XX skincare essence," "saliva smell," "related to"

[0078] "Repairing wrinkles," "saliva smell," "irrelevant."

[0079] In an exemplary embodiment of this application, the entity sentiment classification model can perform entity sentiment classification from the input text "XX skincare essence is so good, it visibly repairs wrinkles! It just smells like saliva" to obtain the following entity sentiment classification results:

[0080] "SK-II Facial Treatment Essence", "Positive"

[0081] "Repair wrinkles", "Front view"

[0082] "The smell of saliva" is a negative connotation.

[0083] S204. Calculate the occurrence probability of each entity, and construct the knowledge graph of the subject by the entity type of the subject, the entities corresponding to each entity type, and the occurrence probability and / or sentiment classification of each entity.

[0084] In an exemplary embodiment of this application, the calculation of the occurrence probability of each entity can be illustrated by the following example: If, in the selected corpus, among all the "efficacy" entity types related to "XX skin care essence", the entity "repair wrinkles" appears 6 times, the entity "moisturizes" appears 6 times, the entity "maintains activity" appears 4 times, and the entities in other "efficacy" entity types appear 4 times, then the probability corresponding to the entity "repair wrinkles" is: 6 / (6+6+4+4) = 0.3, and the probability corresponding to the entity "maintains activity" is: 4 / (6+6+4+4) = 0.2.

[0085] In an exemplary embodiment of this application, based on the obtained entity types (efficacy, ingredients) and the obtained entities (wrinkle repair, moisturizing, maintaining activity, Pitera, collagen, etc.), and the probability of these entities appearing, a knowledge graph of the main product "XX skincare essence" in the beauty industry within the current preset platform can be constructed:

[0086] XX Skincare Essence:

[0087] effect:

[0088] [Repairs wrinkles, 0.3]

[0089] [Moisturizing, 0.3]

[0090] [Maintain activity, 0.2]

[0091]

[0092] Element:

[0093] [pitera, 0.3]

[0094] [Collagen, 0.3]

[0095]

[0096] Other entity types...

[0097] In an exemplary embodiment of this application, the method may further include: re-collecting the text of a preset subject of the preset platform and / or industry according to a preset period and inputting it into the first preset model to periodically update the knowledge graph.

[0098] In an exemplary embodiment of this application, after the knowledge graph of one or more subjects is constructed, the latest data and / or corpus about the subject can be filtered according to a preset period (e.g., one day, three days, four days, etc.), and steps S201-S204 can be repeated to update the knowledge graph of the subject.

[0099] In an exemplary embodiment of this application, pre-establishing the short text generation model may include:

[0100] First training data is obtained based on the first preset model;

[0101] The short text generation model is obtained by training the second preset model using the first training data.

[0102] In an exemplary embodiment of this application, obtaining the first training data based on the first preset model may include:

[0103] Filter the texts in the text library that meet the criteria based on keywords;

[0104] The analyzed text is input into a pre-built named entity recognition model, entity relationship recognition model, and entity sentiment classification model to obtain entities and corresponding sentiment classifications for one or more products.

[0105] The entities of the one or more products and their corresponding sentiment classifications are labeled as one or more short texts, forming the first training data.

[0106] In an exemplary embodiment of this application, analytical texts that meet certain conditions can be filtered from a massive text library based on keywords. For example, arbitrary keywords, time periods, data source platforms, etc., can be set. Analytical texts that meet the conditions can be selected on the data source platform by inputting keywords, time periods, and other information. These analytical texts can be input into a first preset model that has been established in advance. Specifically, named entity recognition model, entity relationship recognition model, and entity sentiment classification model can be input to obtain the entity type, entity, entity relationship, entity sentiment classification, etc. contained in these analytical texts, respectively.

[0107] In an exemplary embodiment of this application, the text "XX skincare essence is so good, it visibly repairs wrinkles! It just smells like saliva" is used as an example. The output of the first preset model can include information such as: XX skincare essence <positive>, wrinkle repair <positive>, and saliva smell <negative>. This information can be used as training data, and labels can be pre-set for this information. For example, it can be set as the label "XX skincare essence is so good, it visibly repairs wrinkles! It just smells like saliva" to indicate that XX skincare essence <positive>, wrinkle repair <positive>, and saliva smell <negative> can constitute the short text "XX skincare essence is so good, it visibly repairs wrinkles! It just smells like saliva". This labeled training data can then be used as the first training data to train the second preset model.

[0108] In an exemplary embodiment of this application, the second preset model can be a self-created neural network model or an existing model product, such as, but not limited to, the T5 model.

[0109] In an exemplary embodiment of this application, a short text generation model can be obtained by training a second preset model with first training data.

[0110] In an exemplary embodiment of this application, pre-establishing the connection generation model may include:

[0111] Obtain the second training data;

[0112] The connection generation model is obtained by training the third preset model using the second training data.

[0113] In an exemplary embodiment of this application, obtaining the second training data may include:

[0114] Obtain multiple texts with a preset text structure; each text consists of a first short text, a second short text, and a third short text, wherein the first short text, the second short text, and the third short text are connected in sequence;

[0115] Each of the second short texts is marked as a connecting sentence, and all the marked texts are used as the second training data.

[0116] In an exemplary embodiment of this application, texts that meet certain conditions can be filtered from a massive text library based on keywords. For example, arbitrary keywords, word count, structure, data source platform, etc., can be set. On the data source platform, texts that meet the conditions can be selected by inputting keywords, word count, structure, and other information. For example, texts with the structure of sentence 1 (first short text), sentence 2 (second short text), and sentence 3 (third short text) can be selected. That is, the text can be composed of sentence 1, sentence 2, and sentence 3 arranged in sequence.

[0117] In an exemplary embodiment of this application, the short text refers to text with a number of characters within a preset range. This preset range can be defined according to different needs, and the specific value may vary for different application scenarios. No detailed limitation is made on the specific value here. For example, the preset range may include 20-60 characters; such as the length of sentence 1 and sentence 3 being between 20 and 60 Chinese characters, and the length of sentence 2 being between 20 and 40 Chinese characters.

[0118] In an exemplary embodiment of this application, since sentence 2 connects sentence 1 and sentence 3 to form the entire text, sentence 2 can be marked as a connecting sentence.

[0119] In an exemplary embodiment of this application, for a large number of texts selected from a massive text library, these texts can be divided into at least three short text sentences, and the middle sentence is marked as a connecting sentence. These large number of sentences (here referring to short texts) containing the marked connecting sentences are used as second training data to train a third preset model to obtain the connecting generation model.

[0120] In an exemplary embodiment of this application, the linking generation model can connect multiple short texts into a single long text.

[0121] In an exemplary embodiment of this application, a knowledge graph and a text generation model can be created without the above content.

[0122] S102. Determine the main body of the generated text.

[0123] In an exemplary embodiment of this application, after obtaining the knowledge graph and text generation model of a certain industry on a certain preset platform, the subject of the text to be commonly used can be determined, that is, the specific product, such as the aforementioned "XX skin care essence" or "XX skin care water" in the beauty industry.

[0124] In an exemplary embodiment of this application, the subject can be identified by entering the subject name in a preset input window, or by selecting from a large number of subject selection windows.

[0125] S103. Based on the entity information of the subject in the knowledge graph, select the corresponding entity and the sentiment category corresponding to the entity according to preset rules.

[0126] In an exemplary embodiment of this application, the entity information may include any one or more of the following: the entity types contained in the subject;

[0127] One or more entities corresponding to each entity type;

[0128] The probability of occurrence for each entity; and,

[0129] Each entity corresponds to a sentiment category.

[0130] In an exemplary embodiment of this application, the preset rules may include: one or more first entity types contained in the short text to be generated, and the sorting of multiple first entity types.

[0131] In an exemplary embodiment of this application, for example, if the short text to be generated contains first entity types such as product, scenario, efficacy, and ingredients, the preset rule may include:

[0132] The first short text: Product, scenario, efficacy;

[0133] The second short text: Product, ingredients, efficacy;

[0134] The third short text:…

[0135] In an exemplary embodiment of this application, the preset rule can be determined by the first entity type input in the preset input window and the sorting of these first entity types, or by the selection result of the preset first entity type selection window and the first entity type sorting scheme selection window.

[0136] In an exemplary embodiment of this application, relevant entities can be selected from the knowledge graph according to the preset rules, and relevant entities can be selected based on the occurrence probability of different entities.

[0137] In an exemplary embodiment of this application, the step of selecting the corresponding entity according to a preset rule based on the entity information of the subject in the knowledge graph may include:

[0138] Obtain the first entity type corresponding to the subject from the knowledge graph;

[0139] Extract entities whose occurrence probability is greater than or equal to a preset probability threshold from all entities corresponding to the first entity type.

[0140] In an exemplary embodiment of this application, for example, when selecting relevant entities from a knowledge graph, in the second short text, 2 to 3 entities can be selected from the "Ingredients" type of entity related to the subject "XX Skin Care Water". The probability of each entity being selected in the "Ingredients" type of entity is the same as its probability of occurrence. That is, entities can be selected according to the probability of occurrence of each entity, and one or more entities with the highest probability of occurrence can be selected. Then, 2 to 3 relevant entities under the "Efficacy" and "Scenario" type of entity can be selected in the same way.

[0141] S104. Input the selected entity and the corresponding sentiment classification into the short text generation model, and the short text generation model obtains multiple short texts; the short text refers to a text with a number of characters within a preset range.

[0142] In an exemplary embodiment of this application, after obtaining the entity set through the above steps, each entity and its corresponding entity sentiment can be input into the short text generation model. For example, the following entities and their corresponding entity sentiments can be input into the short text generation model:

[0143] XX Skin Toner | Almond Extract | Soothes Skin | Brightens | Yeast | Repairs | Antioxidant

[0144] In an exemplary embodiment of this application, a short text generation model can output multiple short texts, such as: short text 11, short text 12, short text 13, ..., short text 1n.

[0145] S105. Input the adjacent short texts into the connection generation model in sequence, and output the connection sentence of the adjacent short texts through the connection generation model.

[0146] In an exemplary embodiment of this application, adjacent short texts can be input pairwise and concatenated to generate a model, for example:

[0147] Input: Short text 1, Short text 2; Output: Connecting sentences 1-2;

[0148] Input: Short text 2, Short text 3; Output: Connecting sentences 2-3;

[0149] Input: Short text 3, Short text 4; Output: Connecting sentences 3-4.

[0150] S106. Insert corresponding connecting sentences between multiple adjacent short texts to generate text about the subject.

[0151] In an exemplary embodiment of this application, for example, short texts and connecting sentences can be connected in the order of short text 1, connecting sentences 1-2, short text 2, connecting sentences 2-3, and short text 3 to obtain the text of the subject.

[0152] In an exemplary embodiment of this application, the method may further include:

[0153] The generated text of the main body is then segmented into sentences;

[0154] Calculate the repetition rate of each sentence in the text;

[0155] When the repetition rate is greater than or equal to a preset repetition rate threshold, the sentence is deleted.

[0156] In an exemplary embodiment of this application, the text about the subject obtained in the above steps is processed into sentences. If the repetition rate of a certain sentence with the previous sentence reaches a repetition rate threshold, such as 0.7, the sentence can be deleted.

[0157] In an exemplary embodiment of this application, the sentence repetition rate calculation method may include:

[0158] a = N1-2 / min(N1,N2);

[0159] Where a is the repetition rate, N1 is the number of characters in sentence 1 (after deduplication), N2 is the number of characters in sentence 2 (after deduplication), and N1-2 is the number of identical characters in sentences 1 and 2 after deduplication.

[0160] In exemplary embodiments of this application, the solutions of this application embodiments include at least the following advantages:

[0161] 1. Marketing texts can be generated in large quantities and customized to reduce labor costs;

[0162] 2. Capture comments from different platforms and industries, build corresponding knowledge graphs and text generation models, and generate texts of different styles for different platforms and industries;

[0163] 3. It can crawl public data daily and update the knowledge graph, enabling the generated text to capture social hot topics;

[0164] 4. When generating text, inputting entities from the knowledge graph ensures that the generated text avoids factual errors.

[0165] 5. When generating text, display the emotional information of the input consumer towards the entity, so that the generated text has the emotional inclination of the consumer;

[0166] 6. Since the existing technology has uncontrollable effects when generating long texts, this invention uses a short text generation model to generate short texts, and then uses a connecting sentence generation model to generate connecting sentences to connect the short texts. This hierarchical and structured approach to generating text makes the generation results more controllable.

[0167] It can generate text from the consumer's perspective, simulating consumer reviews.

[0168] This application also provides a text generation device 1, such as... Figure 4 As shown, the device may include a processor 11 and a computer-readable storage medium 12, wherein the computer-readable storage medium 12 stores instructions that, when executed by the processor 11, implement the text generation method.

[0169] In the exemplary embodiments of this application, any of the embodiments in the foregoing text generation method embodiments are applicable to the device embodiments, and will not be described in detail here.

[0170] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

Claims

1. A text generation method, characterized in that, The method includes: The knowledge graph and text generation model are acquired and established; the text generation model includes: a short text generation model and a linking generation model. Determine the main body of the text to be generated; Based on the entity information of the subject in the knowledge graph, select the corresponding entity and the sentiment category corresponding to the entity according to preset rules; The selected entity and its corresponding sentiment classification are input into the short text generation model, which then generates multiple short texts. The adjacent short texts are sequentially input into the connection generation model, and the connection generation model outputs the connection sentences of the adjacent short texts. By inserting corresponding connecting sentences between multiple adjacent short texts, text about the subject is generated; The connection generation model is obtained by acquiring second training data and training the model based on the acquired second training data; The acquisition of the second training data includes: Obtain multiple texts with a preset text structure; each text consists of a first short text, a second short text, and a third short text, wherein the first short text, the second short text, and the third short text are connected in sequence; Each of the second short texts is marked as a connecting sentence, and all the marked texts are used as the second training data.

2. The text generation method according to claim 1, characterized in that, The entity information includes any one or more of the following: the entity types contained in the subject; One or more entities corresponding to each entity type; The probability of occurrence for each entity; and, Each entity corresponds to a sentiment category.

3. The text generation method according to claim 2, characterized in that, The preset rules include: the sorting of one or more first entity types contained in the short text to be generated, and the sorting of multiple first entity types; The step of selecting corresponding entities according to preset rules based on the entity information of the subject in the knowledge graph includes: Obtain the first entity type corresponding to the subject from the knowledge graph; Extract entities whose occurrence probability is greater than or equal to a preset probability threshold from all entities corresponding to the first entity type.

4. The text generation method according to claim 1, characterized in that, The method further includes: The generated text of the main body is then segmented into sentences; Calculate the repetition rate of each sentence in the text; When the repetition rate is greater than or equal to a preset repetition rate threshold, the sentence is deleted.

5. The text generation method according to claim 2, characterized in that, The knowledge graph is pre-built, including the following steps: constructing knowledge graphs for different products on different platforms and / or in different industries: Obtain a first preset model; the first preset model includes: a named entity recognition model, an entity relationship recognition model, and an entity sentiment classification model; Collect text input from preset entities on preset platforms and / or in preset industries into the first preset model; The named entity recognition model outputs the entity type corresponding to the subject; the entity relationship recognition model outputs the relationship between entities; and the entity sentiment classification model outputs the sentiment classification corresponding to each entity. The probability of occurrence of each entity is calculated, and the knowledge graph of the subject is constructed from the entity type of the subject, the entities corresponding to each entity type, and the occurrence probability and / or sentiment classification of each entity.

6. The text generation method according to claim 5, characterized in that, The method further includes: re-collecting the text of the preset entities of the preset platform and / or industry according to a preset period and inputting it into the first preset model to periodically update the knowledge graph.

7. The text generation method according to claim 5, characterized in that, The short text generation model is pre-established, including: First training data is obtained based on the first preset model; The short text generation model is obtained by training the second preset model using the first training data; The pre-establishment of the connection generation model includes: Obtain the second training data; The connection generation model is obtained by training the third preset model using the second training data.

8. The text generation method according to claim 7, characterized in that, The step of obtaining the first training data based on the first preset model includes: Filter the texts in the text library that meet the criteria based on keywords; The analyzed text is input into a pre-built named entity recognition model, entity relationship recognition model, and entity sentiment classification model to obtain entities and corresponding sentiment classifications for one or more products. The entities of the one or more products and their corresponding sentiment classifications are labeled as one or more short texts, forming the first training data.

9. A text generation apparatus, comprising a processor and a computer-readable storage medium storing instructions, characterized in that, When the instructions are executed by the processor, the text generation method as described in any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Text generation method and device

    CN110489755A

  • Social media sentiment classification method and device based on knowledge graph

    CN111538835A

  • Long text generation method and device, equipment and storage medium

    CN112541348A