A text generation method, device, equipment, and storage medium

By using semantic vector model and clustering algorithm to mine homogeneous problems in the toB e-commerce scenario, extracting answers from product details with text classification models and generating Q&A text text, the accuracy and efficiency of Q&A text generation in the toB e-commerce scenario are solved, and automated answer provision and user experience improvement are achieved.

CN113886553BActive Publication Date: 2025-07-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111272247.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-29
Publication Date
2025-07-04
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

It is difficult for the existing technology to efficiently generate highly targeted Q&A text in toB e-commerce scenarios, and it is impossible to effectively use product details to provide accurate answers to user questions, and human resources are consumed relatively large.

Method used

By obtaining question-and-answer pairs and product details information, using semantic vector model and clustering algorithm to mine homogeneous problems, extract answers from product details in combination with text classification models, and generate question-and-answer text.

Benefits of technology

It automatically provides users with accurate answers, saves human resources, improves user experience, and is suitable for question-and-answer services in toB e-commerce scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113886553B_ABST
    Figure CN113886553B_ABST
Patent Text Reader

Abstract

The present disclosure provides a text generation method, apparatus, device, and storage medium, relating to the field of data processing, and particularly to fields such as information retrieval, intelligent search, and big data. The specific implementation solution is as follows: obtaining original materials, where the original materials include at least one question-and-answer pair and product detail information; extracting questions from at least one question-and-answer pair; clustering the questions in at least one question-and-answer pair to obtain at least one representative question; for each representative question, based on the product detail information, extracting an answer to answer the representative question, and forming a question-and-answer type text with the representative question and the answer. The present disclosure realizes the generation of question-and-answer type texts, can automatically answer questions raised by users, meet user needs, and can save human resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and in particular to the fields of information retrieval, intelligent search, big data, etc. Background Art

[0002] Text generation technology is an advanced technology that extracts specific valuable information from text corpus data using theories such as machine learning or deep learning. This technology can greatly save manpower and replace manual extraction of high-value content from massive text data, such as question-and-answer text generation, etc. Summary of the Invention

[0003] The present disclosure provides a text generation method, apparatus, device, and storage medium.

[0004] According to a first aspect of the present disclosure, there is provided a text generation method, including:

[0005] Obtain original materials, where the original materials include at least one question-and-answer pair and product detail information;

[0006] Extract the questions in the at least one question-and-answer pair;

[0007] Cluster the questions in the at least one question-and-answer pair to obtain at least one representative question;

[0008] For each representative question, based on the product detail information, extract an answer to answer the representative question, and form a question-and-answer text with the representative question and the answer.

[0009] According to a second aspect of the present disclosure, there is provided a text generation apparatus, including:

[0010] An acquisition module, configured to acquire original materials, where the original materials include at least one question-and-answer pair and product detail information;

[0011] A first extraction module, configured to extract the questions in the at least one question-and-answer pair;

[0012] A clustering module, configured to cluster the questions in the at least one question-and-answer pair to obtain at least one representative question;

[0013] A second extraction module, configured to, for each representative question, based on the product detail information, extract an answer to answer the representative question;

[0014] A composition module, configured to form a question-and-answer text with the representative question and the answer.

[0015] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect.

[0019] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the first aspect.

[0020] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the method described in the first aspect when executed by a processor.

[0021] The present disclosure realizes generating Q&A text, and based on the Q&A text, it can automatically answer questions raised by users and meet the needs of users.

[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0023] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0024] Figure 1 is a flowchart of a text generation method according to an embodiment of the present disclosure;

[0025] Figure 2 is a flowchart of clustering questions in at least one Q&A pair to obtain at least one representative question according to an embodiment of the present disclosure;

[0026] Figure 3 is a schematic diagram of obtaining a representative question through clustering according to an embodiment of the present disclosure;

[0027] Figure 4 is a flowchart of extracting an answer to answer the representative question based on the product detail information according to an embodiment of the present disclosure;

[0028] Figure 5 is a schematic diagram of a text generation process according to an embodiment of the present disclosure;

[0029] Figure 6 is a schematic structural diagram of a text generation device according to an embodiment of the present disclosure;

[0030] Figure 7 is another structural schematic diagram of a text generation device according to an embodiment of the present disclosure;

[0031] Figure 8 is a block diagram of an electronic device for implementing the text generation method according to an embodiment of the present disclosure. Detailed implementation manners

[0032] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0033] The related text generation mainly has the following several methods: (1) Template-based method, where some text templates are set manually, and the blank parts in the templates need to be filled according to actual information. (2) Text summary-based method, where a model is trained using a deep network structure, which can extract key information from text corpora to generate answer texts. (3) Knowledge graph-based method, based on the knowledge information extraction technology widely used in the current industry, mining commodity knowledge information from text corpora, drawing a commodity knowledge graph, and finally realizing the generation of answers.

[0034] For the template-based method, some fixed answer templates need to be set manually. Different types of questions correspond to different templates. When generating answers, only the blank positions in the templates need to be filled. The application scenarios of the template-based method can only be used in simple types of question-and-answer text generation scenarios, such as weather-related Q&A, etc. The application scope is relatively limited and it is difficult to apply it to the generation of question-and-answer texts in complex toB e-commerce scenarios.

[0035] For the text summary-based method, a natural language processing model is trained using a deep network, and then the answer text is extracted from the original corpus through the model. Although this method does not rely on templates and has great autonomy, the summary extracted from the original corpus may lack pertinence and cannot answer a certain question well. In many cases, the mined text has a poor correlation with the question and is off-topic.

[0036] For the knowledge graph-based method, first, knowledge information is mined from the original corpus to construct a knowledge graph, and then the answer text is generated with the help of this graph. The process of constructing a knowledge graph is relatively complex. Generally, for relatively simple scenarios, key information can be extracted from the original corpus, with strong generality and good pertinence. However, for more complex questions, it is difficult to aggregate multiple knowledge points for answering. That is, it is not suitable for complex scenarios such as toB e-commerce scenarios.

[0037] In today's Internet information retrieval, the application scenarios of Q&A text are very extensive. For example, in the toB e-commerce scenario, in order to understand products in detail, users generally ask a large number of questions. Among them, toB refers to a business model in enterprise business where the enterprise serves as the service provider to provide platforms, products or services for enterprise customers and earns profits. It can also be called enterprise service. Mining questions and generating corresponding answers for the questions are the key and difficult contents in the process of generating Q&A text.

[0038] Through analysis in the embodiments of the present disclosure, it is found that in the toB e-commerce scenario, the questions of users are homogeneous, which means that the questions raised by different users are mostly the same, and the difference is only some differences in adjectives or adverbs. For example, "How much is the excavator?" and "What is the price of the excavator?". And the answers to these questions are generally included in the product detail information. Text generation technology can be used to generate answers to the questions raised by users relying on corpus data such as product details. Based on this, in the embodiments of the present disclosure, the original corpus composed of product detail information and historical Q&A pairs is used to mine user questions, and the answers to the questions are extracted from the product detail information to generate Q&A pair text.

[0039] The text generation method provided by the embodiments of the present disclosure will be described in detail below.

[0040] The text generation method provided by the embodiments of the present disclosure can be applied to electronic devices. Specifically, the electronic devices can include servers, terminals, and so on.

[0041] The text generation method provided by the embodiments of the present disclosure may include:

[0042] Obtain original materials, where the original materials include at least one Q&A pair and product detail information;

[0043] Extract the questions in at least one Q&A pair;

[0044] Cluster the questions in at least one Q&A pair to obtain at least one representative question;

[0045] For each representative question, based on the product detail information, extract the answer to answer the representative question, and form Q&A text by combining the representative question and the answer.

[0046] In the embodiments of the present disclosure, Q&A text is generated according to the original corpus including multiple Q&A pairs and product detail information. In this way, when a user asks a question, the corresponding answer to the question raised by the user can be found based on this Q&A text, that is, the question raised by the user can be automatically answered to meet the user's needs. Furthermore, human resources can be saved and the user experience can be improved.

[0047] Figure 1It is a flowchart of the text generation method provided by an embodiment of the present disclosure. Refer to Figure 1 for a detailed description of the text generation method provided by an embodiment of the present disclosure.

[0048] S101, Obtain the original materials.

[0049] The original materials include at least one question-and-answer pair and product details information.

[0050] At least one question-and-answer pair may include a question-and-answer pair composed of a question raised by a user and an answer to that question. Among them, the question raised by the user may include questions raised by the user during the historical query process. For example, the accumulated question-and-answer data in the toB e-commerce scenario includes questions raised by users and answers from merchant customer service. One question in the question pair may correspond to multiple answers, and the content covered by each answer will vary. Or, one question corresponds to one answer.

[0051] The product details information represents information related to the product, which may include the content included in the product details page. For example, the price, size, style introduction of the product, and so on.

[0052] In one implementable manner, after obtaining the original materials, the original materials can be preprocessed. Among them, the preprocessing process can be understood as a process of standardizing the original corpus, which may include removing blank characters and illegal characters in the text of the original materials, correcting incorrect characters, and so on.

[0053] S102, Extract the questions in at least one question-and-answer pair.

[0054] Extract questions from the question-and-answer pair. For example, extract questions from the accumulated question-and-answer data in the toB e-commerce scenario. These questions are all questions actually raised by users. In this way, these questions can accurately reflect the actual needs of users in the toB e-commerce scenario, which helps the e-commerce system to understand users, capture users, and improve user activity and retention rate.

[0055] S103, Cluster the questions in at least one question-and-answer pair to obtain at least one representative question.

[0056] The questions raised by different users have homogeneity. Simply understood, the questions raised by different users are essentially the same. For example, if one user's question is: "How much is this piece of clothing?", and another user's question is: "What is the price of this piece of clothing?", then these questions with the same essence can be regarded as the same type of question, or it can be understood as a representative question.

[0057] In one implementation, the extracted questions can be classified. For example, "How much is an excavator?" belongs to the price category of questions, and "How to repair an excavator?" belongs to the repair category of questions. Specifically, sample data can be extracted in advance for manual annotation, that is, the category corresponding to each question in the question-and-answer pair is annotated, and the text classification model for determining the question category corresponding to the question is trained using the annotated sample data. For example, a convolutional neural network for text classification (Convolutional Neural Networks for Sentence Classification, TextCNN), and then, the trained TextCNN is used to determine the question category corresponding to each question. If there are many types of questions, a small number of annotated samples are difficult to cover the questions with many types. However, if there are too many annotated samples, it will consume more human resources. This method is suitable for scenarios with relatively few types of questions.

[0058] In another implementation, for relatively complex scenarios, such as the toB e-commerce scenario, questions can be automatically mined through clustering. As Figure 2 shown, S103 may include:

[0059] S201, determine a semantic vector for each question respectively.

[0060] A semantic vector model, such as the word2vec model, can be used to calculate a semantic-level vector for each question, that is, a semantic vector.

[0061] The purpose of this step is to place semantically related content in a similar numerical space. For example, price and how much are semantically related, so the distance between their semantic vectors will be very close, which is the basis for subsequent clustering operations.

[0062] S202, cluster the questions in at least one question-and-answer pair according to the distance between the semantic vectors of each question to obtain at least one representative question.

[0063] The distance between the semantic vectors of each question can be calculated. If the distance between the semantic vectors is less than the preset distance threshold, the questions corresponding to the semantic vectors can be clustered into a representative question.

[0064] Alternatively, a clustering algorithm can be used for adaptive clustering. For example, the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm can be used to cluster multiple problems. As the name implies, the DBSCAN algorithm will take clusters with a relatively high density within a certain distance near a certain center point as the same class. The selection of the center point is random, so it does not require specifying the number of classes in advance, which is suitable for scenarios with a relatively large number of problem types, such as the toB e-commerce scenario. In the embodiments of the present disclosure, multiple semantic vectors are randomly selected using DBSCAN. For each semantic vector, those within a preset distance threshold from the semantic vector are taken as a cluster to obtain a representative problem. In this way, at least one representative problem can be clustered.

[0065] As Figure 3 shown in an example, the problem pair includes problem A, problem B, and problem C proposed by the user. Problem A, problem B, and problem C are respectively input into the vector model, vectorized to obtain the semantic vector of problem A, the semantic vector of problem B, and the semantic vector of problem C. Then, through the clustering module, based on the semantic vector of problem A, the semantic vector of problem B, and the semantic vector of problem C, problem A, problem B, and problem C are clustered to obtain problem A and problem B. Problem A and problem B are the representative problems obtained by clustering.

[0066] The semantic vector of the problem text is calculated using word2vec, and then the DBSCAN algorithm is used to cluster the problem semantic vectors into multiple homogeneous problems, that is, at least one representative problem is clustered.

[0067] In the embodiments of the present disclosure, the vector representation of each problem is calculated through the semantic vector model, and then after iterative clustering algorithms, similar problems are aggregated together, and finally several typical problems are refined, that is, at least one representative problem is extracted. In this way, no annotation is required, and at least one representative problem is automatically mined through clustering, which can be applied to scenarios with a relatively large number of problem types and does not require excessive human resources.

[0068] S104. For each representative problem, based on the commodity detail information, extract the answer to the representative problem, and form a Q&A text with the representative problem and the answer.

[0069] The commodity detail information includes information related to the commodity. Generally, in the e-commerce scenario, the questions raised by users are all about the commodity. It can be understood that the answers to the questions raised by users are generally included in the commodity detail information. Based on this, in the embodiments of the present disclosure, corresponding answers are extracted from the commodity detail information for each representative problem.

[0070] In an alternative embodiment, in S104, for each representative question, based on the product details information, answers to the representative questions are extracted. For example, Figure 4 as shown, it may include:

[0071] S401, splitting the product details information into multiple paragraphs.

[0072] The product details information may include product detail text, and the product detail text can be split into different paragraphs. The paragraphs can be long or short. In one case, one paragraph contains at least one sentence.

[0073] S402, for each paragraph, based on the degree of association between the paragraph and each representative question respectively, the representative question with the highest degree of association with the paragraph is used as the representative question for answering the paragraph.

[0074] The degree of association between a paragraph and a representative question can also be understood as the degree of relevance of the paragraph to the representative question.

[0075] In one implementable manner, the paragraph can be input into a text classification model, and the representative question with the highest degree of association with the paragraph is output through the text classification model, and the representative question output by the text classification model is used as the representative question for answering the paragraph.

[0076] For example, the text classification model used to determine the representative question corresponding to the paragraph can be a TextCNN text classification model. The TextCNN text classification model converts the paragraph into a semantic vector, then extracts the key features of the semantic vector of the paragraph, and finally performs text classification discrimination, that is, determines the representative question corresponding to the paragraph.

[0077] Specifically, each paragraph can be loaded into the text classification model respectively. For each paragraph, the text classification model generates a set of scores. Each score in this set of scores represents the degree of association between the paragraph and a representative question. The higher the score, the higher the degree of association between the paragraph and the representative question. The higher the score, the more relevant it is to a certain category. The text classification model selects the highest one from this set of scores and outputs the representative question corresponding to the highest score. It can also be understood that the text classification model selects the highest one from the scores, that is, the one with the highest degree of relevance, as the category corresponding to this paragraph, that is, the question corresponding to the answer of this paragraph.

[0078] Among them, if the category of a certain paragraph is not clear, then filter out that paragraph because some paragraphs in the actual product details are redundant. For example, after selecting the highest score among the scores of paragraphs and each representative question, compare this score with a preset score. If this score is less than the preset score, the text classification model outputs "There is no corresponding representative question for the paragraph", and at this time, filter out this paragraph, that is, this paragraph is no longer used in subsequent calculation processes, where the preset score is determined according to actual needs.

[0079] Finally, each paragraph above the preset score will find the most matching representative question. Among them, each paragraph can determine a unique representative question. For example, paragraph 1 corresponds to representative question 1, paragraph 2 corresponds to representative question 2, paragraph 3 corresponds to representative question 1, paragraph 4 corresponds to representative question 3, and so on.

[0080] By adopting this embodiment, a text classification model can be pre-trained. The input of the text classification model is a text, and the output is the representative question corresponding to this text. In this way, using the pre-trained text classification model can more conveniently determine the representative questions corresponding to each paragraph.

[0081] The training of this text classification model can be achieved through the following steps:

[0082] For each question-and-answer pair, label the answer text in the sample question-and-answer pair and the representative question annotation information corresponding to the answer text; label the answer text in the sample product details information and the representative question annotation information corresponding to each answer text; use multiple answer texts and the representative question annotation information corresponding to each answer text to train the text classification model, and the multiple answer texts include the answer text in the sample question-and-answer pair and the answer text in the sample product details information.

[0083] The answer text in the sample question-and-answer pair and the representative question annotation information corresponding to the answer text, as well as the answer text in the sample product details information and the representative question annotation information corresponding to each answer text are the sample data for training the text classification model. An answer text and the corresponding representative question annotation information can be used as a sample pair for training.

[0084] An initial model can be obtained. For a sample pair, the answer text in the sample pair is input into the initial model, and the parameters of the initial model are adjusted so that the difference between the output of the initial model and the representative question corresponding to the question annotation information in the sample pair is less than a preset value. The preset value can be determined according to actual needs. For example, the preset value can be 0.1, 0.01, etc. Performing the above process on a sample pair once is called one iteration. The above steps are respectively performed on multiple sample pairs until the iteration end condition is met. For example, the number of iterations reaches a preset number of iterations, or the accuracy of the model reaches a preset accuracy. At this time, the training is completed, and a trained text classification model is obtained. Among them, the preset accuracy represents the difference between the output of the model and the representative question corresponding to the question annotation information, and can be determined according to actual needs.

[0085] Among them, the sample question-and-answer pairs can include the question-and-answer pairs in the above-mentioned original corpus, or can include question-and-answer pairs obtained in other scenarios. For example, question-and-answer pairs for another commodity, and the other commodity is different from the commodity targeted by the question-and-answer pairs in the above-mentioned original corpus. Similarly, the sample commodity detail information can include the commodity detail information in the above-mentioned original corpus, or can include commodity detail information obtained in other scenarios. For example, commodity detail information for another commodity.

[0086] In the embodiments of the present disclosure, the sample question-and-answer pairs and the sample commodity detail information are derived from actual data in multiple scenarios. The text classification model for determining the representative question corresponding to a paragraph is trained using multiple sample question-and-answer pairs and sample commodity detail information, which can more accurately reflect the correspondence between the paragraph and the representative question, making the determined representative question more matched with the paragraph, and determining a more accurate representative question for the paragraph.

[0087] S403, in response to the representative questions answered by multiple paragraphs being the same, integrate the multiple paragraphs with the same representative question answered, to obtain the multiple paragraphs with the same representative question answered and the answers to the representative questions they answered.

[0088] Each representative question may correspond to multiple paragraphs, that is, the representative questions corresponding to multiple paragraphs may be the same. In the embodiments of the present disclosure, one answer is extracted for one representative question. In this case, it is necessary to integrate multiple paragraphs with the same representative question answered. For example, paragraph 1 corresponds to representative question 1, paragraph 3 corresponds to representative question 1, and paragraphs 1 and 3 can also be integrated to obtain the answer corresponding to representative question 1.

[0089] When there is no redundancy in the multiple paragraphs with the same representative question answered and the sentences formed by directly splicing these multiple paragraphs are smooth, the multiple paragraphs with the same representative question answered can be directly spliced to obtain an answer to the representative question.

[0090] However, generally, there may be redundancy in multiple paragraphs with the same representative question in the answer. Additionally, directly splicing these multiple paragraphs simply may result in semantic incoherence and require adjustment of the order.

[0091] A text summary extraction model can be pre-trained. This text summary extraction model can be implemented based on a natural language processing framework. Input multiple paragraphs with the same representative question in the answer into this text summary extraction model. The text summary extraction model extracts the core content of these multiple paragraphs, which can also be understood as removing redundant information. At the same time, it can also adjust the semantic order and grammar to obtain a refined answer that meets the grammar and semantic order, thus completing the improvement of the answer.

[0092] By integrating multiple paragraphs with the same representative question in the answer, an answer with concise, smooth, and grammar-compliant sentence expressions can be obtained, improving the quality of the answer text.

[0093] In the embodiments of the present disclosure, a question-and-answer pair text is generated by comprehensively considering a corpus composed of question-and-answer pairs and merchant product detail information. Semantic vectors are generated for the questions in the corpus, and the semantic vectors of multiple questions are clustered to obtain at least one representative question, which can also be understood as clustering to obtain multiple question categories. A text classification model is trained using the corpus. The product detail information is divided into multiple paragraphs, and each paragraph is input into the text classification model respectively to obtain the representative question corresponding to the paragraph. After optimizing and adjusting the answer text of the representative question, the adjusted answer text is used as the answer to the representative question, completing the generation process of the question-and-answer text. In this way, based on this question-and-answer text, questions raised by users can be automatically answered, meeting user needs and saving a large amount of human resources.

[0094] At the same time, by comprehensively considering a corpus composed of question-and-answer pairs and merchant product detail information, and the product detail information contains the answers to the questions raised by users. Extracting the answers to the questions from the product detail information can improve the accuracy of the generated questions and answers. Additionally, all the content in the product detail information, that is, each paragraph obtained by splitting, is considered in the process of generating the answers. In this way, as much complete information as possible can be provided to users during the process of providing answers to users. During the process of answering users' questions, more accurate and complete answers can be provided to users, improving the user experience, etc.

[0095] In a specific embodiment, as Figure 5 shown, the text generation method provided by the embodiments of the present disclosure includes four stages: (1) preprocessing; (2) question mining; (3) answer generation; (4) answer integration.

[0096] The preprocessing stage can be understood as a process of standardizing the original corpus, which can specifically include filtering illegal characters, correcting incorrect characters, and so on.

[0097] The problem mining stage mainly includes a vector model, similarity clustering, and result output.

[0098] The vector model process may include: using a semantic vector model, such as the word2vec model, to calculate a semantic-level vector for each problem, that is, a semantic vector.

[0099] The similarity clustering process includes clustering the problems according to the distances of the semantic vectors of each problem to obtain at least one representative problem. For example, using DBSCAN to randomly select multiple semantic vectors, for each semantic vector, those with a distance within a preset distance threshold from this semantic vector are taken as a cluster to obtain a representative problem. Simply put, multiple problems with similarity are clustered into one representative problem.

[0100] The result output is to cluster to obtain at least one representative problem, which can also be understood as homogeneous types of problems.

[0101] The answer generation stage mainly includes text splitting, model training, and answer summarization.

[0102] Text splitting includes splitting the detailed product information into multiple paragraphs.

[0103] The model training in the answer generation stage includes training a text classification model using the answer text in the sample Q&A pairs and the representative problem annotation information corresponding to the answer text, as well as the answer text in the sample product details information and the representative problem annotation information corresponding to each answer text. The input of this text classification model is a text, and the output is the representative problem corresponding to this text.

[0104] Answer summarization can be understood as summarizing the paragraphs answering the same question.

[0105] The answer integration stage mainly includes model training and information extraction.

[0106] The answer integration stage trains a text summary extraction model. This text summary extraction model can be implemented based on a natural language processing framework. Inputting multiple paragraphs with the same representative problem answered into this text summary extraction model, this text summary extraction model extracts the core content of these multiple paragraphs, which can also be understood as removing redundant information. At the same time, it can also adjust the semantic order and grammar to obtain a refined answer that meets the grammar and semantic order, that is, to achieve information extraction.

[0107] The generated Q&A text can be applied to the toB e-commerce scenario. After generating the Q&A text, the question raised by the user is obtained; it is determined which representative question the question raised by the user belongs to; the answer corresponding to the representative question is obtained from the generated Q&A text, and the answer is fed back to the user. In this way, the answer corresponding to the question raised by the user can be automatically fed back to meet the user's needs.

[0108] In the embodiments of the present disclosure, homogeneous questions are mined through clustering, that is, at least one representative question is obtained. And answers are extracted for each representative question from the product detail information. Since the product detail information contains information related to the product, and the questions raised by users are generally also about the product, it can be understood that the product detail information provides a relatively accurate template for generating the answers to the questions. By extracting answers from the product detail information, more accurate answers can be obtained for the representative questions. The Q&A text obtained based on the representative questions and answers is more suitable for the e-commerce scenario, can meet the consultation of users in the e-commerce scenario, can accurately and quickly find the corresponding answers to the questions raised by users, and greatly reduces the labor cost and reduces the communication cost of customer service, improving the product experience.

[0109] The embodiments of the present disclosure also provide a text generation device, as Figure 6 shown, including:

[0110] An acquisition module 601, configured to acquire original materials, where the original materials include at least one Q&A pair and product detail information;

[0111] A first extraction module 602, configured to extract questions from at least one Q&A pair;

[0112] A clustering module 603, configured to cluster the questions in at least one Q&A pair to obtain at least one representative question;

[0113] A second extraction module 604, configured to, for each representative question, extract an answer to answer the representative question based on the product detail information;

[0114] A composition module 605, configured to compose the representative question and the answer into a Q&A text.

[0115] Optionally, the clustering module 603 is specifically configured to determine a semantic vector for each question respectively; cluster the questions in the at least one Q&A pair according to the distance between the semantic vectors of each question to obtain at least one representative question.

[0116] Optionally, the second extraction module 604 is specifically configured to split the product detail information into multiple paragraphs; for each paragraph, based on the degree of association between the paragraph and each representative question, take the representative question with the highest degree of association with the paragraph as the representative question for the paragraph's answer; in response to the same representative questions for the answers of multiple paragraphs, integrate the multiple paragraphs with the same representative questions for the answers to obtain multiple paragraphs with the same representative questions for the answers and the answers to the representative questions answered.

[0117] Optionally, the second extraction module 604 is specifically configured to classify the paragraph input text into a text classification model, output, through the text classification model, the representative question with the highest degree of association with the paragraph, and take the representative question output by the text classification model as the representative question for the paragraph's answer.

[0118] Optionally, as Figure 7 shown, the apparatus further includes:

[0119] The annotation module 701 is configured to, for each question-and-answer pair, annotate the answer text in the sample question-and-answer pair and the representative question annotation information corresponding to the answer text; annotate the answer text in the sample product detail information and the representative question annotation information corresponding to each answer text.

[0120] The training module 702 is configured to use multiple answer texts and the representative question annotation information corresponding to each answer text to train and obtain a text classification model, where the multiple answer texts include the answer texts in the sample question-and-answer pair and the answer texts in the sample product detail information.

[0121] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure, etc., of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0122] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0123] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0124] As Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 802 or computer programs loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0125] Multiple components in device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0126] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the text generation method. For example, in some embodiments, the text generation method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the text generation method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the text generation method in any other appropriate manner (e.g., by means of firmware).

[0127] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0128] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on the remote machine or server.

[0129] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0130] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0131] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0132] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0133] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0134] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A text generation method, comprising: Obtaining original materials, where the original materials include at least one question-and-answer pair and product details information; Extracting the questions in the at least one question-and-answer pair; Clustering the questions in the at least one question-and-answer pair to obtain at least one representative question; Splitting the product details information into multiple paragraphs; For each paragraph, based on the degree of association between the paragraph and each representative question, taking the representative question with the highest degree of association with the paragraph as the representative question for answering the paragraph; In response to the representative questions for answering multiple paragraphs being the same, integrating the multiple paragraphs with the same representative question for answering, obtaining multiple paragraphs with the same representative question for answering, the answers to the representative questions for answering, and forming a question-and-answer type text by combining the representative question and the answer.

2. The method according to claim 1, wherein, The clustering the questions in the at least one question-and-answer pair to obtain at least one representative question includes: Determining a semantic vector for each question respectively; Clustering the questions in the at least one question-and-answer pair according to the distances of the semantic vectors of each question to obtain at least one representative question.

3. The method according to claim 1, wherein The for each paragraph, based on the degree of association between the paragraph and each representative question, taking the representative question with the highest degree of association with the paragraph as the representative question for answering the paragraph includes: Inputting the paragraph into a text classification model, outputting, through the text classification model, the representative question with the highest degree of association with the paragraph, and taking the representative question output by the text classification model as the representative question for answering the paragraph.

4. According to the method described in claim 3, the method further includes: For each question-and-answer pair, annotating the answer text in the sample question-and-answer pair and the representative question annotation information corresponding to the answer text; Annotating the answer text in the sample product details information and the representative question annotation information corresponding to each answer text; Training the text classification model by using the multiple answer texts and the representative question annotation information corresponding to each answer text, where the multiple answer texts include the answer text in the sample question-and-answer pair and the answer text in the sample product details information.

5. A text generation device, comprising: An obtaining module, configured to obtain original materials, where the original materials include at least one question-and-answer pair and product details information; A first extraction module, configured to extract the questions in the at least one question-and-answer pair; A clustering module, configured to cluster the questions in the at least one question-and-answer pair to obtain at least one representative question; A second extraction module, configured to split the product details information into multiple paragraphs; For each paragraph, based on the degree of association between the paragraph and each representative question, taking the representative question with the highest degree of association with the paragraph as the representative question for answering the paragraph; In response to the representative questions for answering multiple paragraphs being the same, integrating the multiple paragraphs with the same representative question for answering, obtaining multiple paragraphs with the same representative question for answering, the answers to the representative questions for answering; A composition module, configured to form a question-and-answer type text by combining the representative question and the answer.

6. The device according to claim 5, wherein, The clustering module is specifically configured to determine semantic vectors for each question respectively; cluster the questions in the at least one question-answer pair according to the distances of the semantic vectors of the respective questions to obtain at least one representative question.

7. The apparatus according to claim 5, wherein, The second extraction module is specifically configured to input the paragraph into a text classification model, output, through the text classification model, the representative question with the highest degree of association with the paragraph, and use the representative question output by the text classification model as the representative question for the answer to the paragraph.

8. The apparatus according to claim 7, wherein the apparatus further comprises: A labeling module, configured to label the answer text in the sample question-answer pair and the representative question labeling information corresponding to the answer text for each question-answer pair; Label the answer text in the sample product detail information and the representative question labeling information corresponding to each answer text; A training module, configured to train the text classification model by using the multiple answer texts and the representative question labeling information corresponding to each answer text, where the multiple answer texts include the answer texts in the sample question-answer pair and the answer texts in the sample product detail information.

9. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-4.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.

11. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Intelligent network response method and device

    CN104933204A

  • Knowledge base construction method based on massive amounts of problems, electronic device and storage medium

    CN107784105A