An enterprise-level Q&A update method and system based on a language model

The method and system use a language model to generate multiple relevant questions and automate updates in enterprise question-answer databases, addressing issues of single-mindedness and manual complexity in existing technologies.

CN113934818BActive Publication Date: 2025-07-15BAIRONG FINANCIAL INFORMATION SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111191411.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-13
Publication Date
2025-07-15
Estimated Expiration
2041-10-13

AI Technical Summary

Technical Problem

The question and answer extracted from the prior art have single questions, unreasonable questions, and unsuitable answers to employees in the enterprise. At the same time, the human resources update the question bank is complicated and complicated.

Method used

Through the enterprise-level question-and-answer update method based on language model, regular expressions are used to extract document titles and content, combine domain keyword information and dependent syntax information, and generate similar question methods, and generate model output question sequences through similar question methods, and automatically update the question-and-answer library in combination with the update judgment unit.

Benefits of technology

It realizes the generation of multiple similar questions, the answer content is suitable for answering employee questions, and the automatic update of the question bank is realized, improving the accuracy and efficiency of the question-and-answer pairs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113934818B_ABST
    Figure CN113934818B_ABST
Patent Text Reader

Abstract

The present application discloses an enterprise-level Q&A update method and system based on a language model. The method includes: obtaining and converting the format of the first uploaded document to obtain a first converted document; using regular expressions to obtain the first title content and the first text content, and constructing domain keyword information and dependency syntax information according to the answer text set of all domain documents of the first enterprise; inputting the domain keyword information and the dependency syntax information into a similar question generation model; obtaining first output information; combining the information library generated according to the first similar question sequence and the first title content with the update judgment unit to judge whether a preset update trigger condition is satisfied. If satisfied, update the first Q&A library. This solves the technical problems in the prior art that the extracted Q&A pairs have single problems, unreasonable question forms, answers that are not suitable for answering the questions of employees in the enterprise, and at the same time, the manual update of the question bank is cumbersome and complex.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and in particular, to an enterprise-level question-and-answer update method and system based on a language model. Background Art

[0002] Large-scale question-and-answer pair data is crucial for promoting research in fields such as machine reading comprehension and question-and-answer. Currently, the methods for extracting question-and-answer pairs from documents can be summarized into three categories: The first category is the rule-based method, which means using rules to screen out the title as the question and the content corresponding to the title as the answer; the second category is the machine learning method, which means using a computer model to screen and determine the question and the answer, thereby forming a question-and-answer pair; the third category is the deep learning method to select possible answer segments from the document and then generate questions based on the answers. Extracting question-and-answer pairs based on rules has the drawback of a single type of question; extracting question-and-answer pairs based on the machine learning method has the problem of error propagation. However, in enterprise-level documents, the general content below the title is specific and can clearly elaborate on employees' questions, and its answer should be specific; extracting question-and-answer pairs based on deep learning ignores the connection between question generation and answer extraction, which may lead to incompatible generated questions and is also not very suitable for enterprise-level documents.

[0003] In the process of implementing the technical solution in the embodiment of this application, the inventors of this application found that the above technologies have at least the following technical problems:

[0004] The question-and-answer pairs extracted by the prior art have problems such as a single type of question, unreasonable question formulations, answers that are not suitable for answering employees' questions in enterprises, and at the same time, there are technical problems that it is cumbersome and complex to update the question bank manually. Summary of the Invention

[0005] The purpose of this application is to provide an enterprise-level question-and-answer update method and system based on a language model to solve the technical problems that the question-and-answer pairs extracted by the prior art have a single type of question, unreasonable question formulations, answers that are not suitable for answering employees' questions in enterprises, and at the same time, there are technical problems that it is cumbersome and complex to update the question bank manually.

[0006] In view of the above problems, the embodiments of this application provide an enterprise-level question-and-answer update method and system based on a language model.

[0007] First aspect, the present application provides an enterprise-level Q&A update method based on a language model, and the method is implemented through an enterprise-level Q&A update system based on a language model. Among them, the method includes: obtaining a first uploaded document of a first enterprise; obtaining a first converted document by converting the format of the first uploaded document, where the format of the first converted document is txt format; extracting the title and content of the first converted document respectively according to a regular expression to obtain a first title content and a first text content, where the first title content corresponds to the first text content; constructing domain keyword information and dependency syntax information according to the answer text set of all domain documents of the first enterprise; inputting the domain keyword information and the dependency syntax information into a similar question generation model, where the similar question generation model is trained to convergence through multiple sets of training data; obtaining a first similar question sequence according to the first output information output by the similar question generation model; judging whether the preset update trigger condition is satisfied by combining the information library generated by the first similar question sequence and the first title content with the update judgment unit, and if so, updating the first Q&A library.

[0008] On the other hand, the present application also provides an enterprise-level Q&A update system based on a language model, which is used to execute an enterprise-level Q&A update method based on a language model as described in the first aspect. Among them, the system includes: a first obtaining unit: the first obtaining unit is used to obtain a first uploaded document of a first enterprise; a second obtaining unit: the second obtaining unit is used to obtain a first converted document by converting the format of the first uploaded document, where the format of the first converted document is txt format; a third obtaining unit: the third obtaining unit is used to extract the title and content of the first converted document respectively according to a regular expression to obtain a first title content and a first text content, where the first title content corresponds to the first text content; a first constructing unit: the first constructing unit is used to construct domain keyword information and dependency syntax information according to the answer text set of all domain documents of the first enterprise; a first input unit: the first input unit is used to input the domain keyword information and the dependency syntax information into a similar question generation model, where the similar question generation model is trained to convergence through multiple sets of training data; a fourth obtaining unit: the fourth obtaining unit is used to obtain a first similar question sequence according to the first output information output by the similar question generation model; a first update unit: the first update unit is used to judge whether the preset update trigger condition is satisfied by combining the information library generated by the first similar question sequence and the first title content with the update judgment unit, and if so, updating the first Q&A library.

[0009] In a third aspect, an embodiment of the present application further provides an enterprise-level Q&A update system based on a language model, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method described in the first aspect above are implemented.

[0010] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0011] 1. Obtain the first uploaded document of the first enterprise; perform format conversion on the first uploaded document to obtain a first converted document, where the format of the first converted document is the txt format; extract the title and content of the first converted document according to a regular expression to obtain a first title content and a first text content, where the first title content corresponds to the first text content; construct domain keyword information and dependency syntax information according to the answer text set of all domain documents of the first enterprise; input the domain keyword information and the dependency syntax information into a similar question generation model, where the similar question generation model is trained to convergence through multiple sets of training data; obtain a first sequence of similar questions according to the first output information output by the similar question generation model; combine the information library generated according to the first sequence of similar questions and the first title content with the update judgment unit to judge whether a preset update trigger condition is satisfied. If satisfied, update the first Q&A library. The technical effect of generating multiple similar questions based on the title by using the model and retaining all the text under the title as the answer is achieved, so as to ensure that the answer content is suitable for answering employees' questions and realize the automatic update of the question library at the same time.

[0012] 2. By combining keyword mining and dependency syntax analysis theories, extract keyword information and dependency syntax information in the answer text, fuse the two kinds of information into the sentence vector of the title to enhance the semantic information of the title, and input them into the similar question generation model to generate similar questions at the same time. The multiple similar question forms corresponding to the title include both keyword question sentences and interrogative sentence questions. The similar question generation model has strong analysis and calculation capabilities, achieving accurate and efficient technical effects.

[0013] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0015] Figure 1 It is a schematic flowchart of a method for updating enterprise-level Q&A based on a language model according to an embodiment of the present application;

[0016] Figure 2 It is a schematic flowchart of constructing domain keyword information and dependency syntactic information according to the answer text set of all domain documents of the first enterprise in a method for updating enterprise-level Q&A based on a language model according to an embodiment of the present application;

[0017] Figure 3 It is another schematic flowchart of constructing domain keyword information and dependency syntactic information according to the answer text set of all domain documents of the first enterprise in a method for updating enterprise-level Q&A based on a language model according to an embodiment of the present application;

[0018] Figure 4 It is a schematic flowchart of determining whether the preset update trigger condition is satisfied by combining the information library generated from the first similar question sequence and the first in-title text in a method for updating enterprise-level Q&A based on a language model according to an embodiment of the present application, and if satisfied, updating the first Q&A library;

[0019] Figure 5 It is a schematic structural diagram of an enterprise-level Q&A update system based on a language model according to an embodiment of the present application;

[0020] Figure 6 It is a schematic structural diagram of an exemplary electronic device according to an embodiment of the present application.

[0021] Explanation of reference numerals:

[0022] The first acquisition unit 11, the second acquisition unit 12, the third acquisition unit 13, the first construction unit 14, the first input unit 15, the fourth acquisition unit 16, the first update unit 17, the bus 300, the receiver 301, the processor 302, the transmitter 303, the memory 304, the bus interface 305. Detailed implementation manners

[0023] Embodiments of the present application provide an enterprise-level Q&A update method and system based on a language model, which solve the technical problems in the prior art that the extracted Q&A pairs have single problems, unreasonable question formulations, answers that are not suitable for answering employees' questions in enterprises, and the cumbersome and complex manual update of the question bank. It achieves the technical effect of using the model to generate multiple similar question formulations based on the title and retaining all the text under the title as the answer, thereby ensuring that the answer content is suitable for answering employees' questions and realizing the automatic update of the question bank.

[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described herein. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application. Additionally, it should be noted that for the sake of description, only the parts related to the present application are shown in the accompanying drawings rather than all of them.

[0025] Application Overview

[0026] Large-scale Q&A pair data is crucial for promoting research in fields such as machine reading comprehension and Q&A. Currently, the methods for extracting Q&A pairs from documents can be summarized into three categories: The first category is the rule-based method, which means using rules to screen out the title as the question and the corresponding content under the title as the answer; the second category is the machine learning method, which means using a computer model to screen and determine the question and the answer to form a Q&A pair; the third category is the deep learning method to select possible answer segments from the document and then generate questions based on the answers. When extracting Q&A pairs based on rules, there is the drawback of single problems; when extracting Q&A pairs based on the machine learning method, there is the problem of error propagation. However, in enterprise-level documents, the general content under the title is specific and can clearly elaborate on employees' questions, and its answer should be specific; when extracting Q&A pairs based on deep learning, the connection between question generation and answer extraction is ignored, which may lead to incompatible generated question pairs and is also not very suitable for enterprise-level documents.

[0027] The Q&A pairs extracted in the prior art have the technical problems of single problems, unreasonable question formulations, answers that are not suitable for answering employees' questions in enterprises, and the cumbersome and complex manual update of the question bank.

[0028] In response to the above technical problems, the general idea of the technical solution provided by the present application is as follows:

[0029] The present application provides an enterprise-level Q&A update method based on a language model. The method is applied to an enterprise-level Q&A update system based on a language model. The method includes: obtaining a first uploaded document of a first enterprise; performing format conversion on the first uploaded document to obtain a first converted document, where the format of the first converted document is the txt format; extracting the title and content of the first converted document respectively according to a regular expression to obtain a first title content and a first text content, where the first title content corresponds to the first text content; constructing domain keyword information and dependency syntax information according to the answer text set of all domain documents of the first enterprise; inputting the domain keyword information and the dependency syntax information into a similar question generation model, where the similar question generation model is trained to convergence through multiple sets of training data; obtaining a first sequence of similar questions according to the first output information output by the similar question generation model; combining the information library generated according to the first sequence of similar questions and the first title content with the update judgment unit to judge whether a preset update trigger condition is satisfied. If satisfied, update the first Q&A library.

[0030] After introducing the basic principle of the present application, the various non-limiting implementation manners of the present application will be specifically introduced below with reference to the accompanying drawings of the specification.

[0031] Embodiment 1

[0032] Please refer to the atta Figure 1 chment. An embodiment of the present application provides an enterprise-level Q&A update method based on a language model. The method is applied to an enterprise-level Q&A update system based on a language model. The system includes an update unit. The method specifically includes the following steps:

[0033] Step S100: Obtain a first uploaded document of a first enterprise;

[0034] Specifically, the enterprise-level Q&A update method based on a language model is applied to the enterprise-level Q&A update system based on a language model. The title in the document can be directly used as a question, and at the same time, the model is used to generate multiple similar questions based on the title. By retaining all the text under the title as the answer, it is ensured that the answer content is suitable for answering employees' questions, and at the same time, the question library can be automatically updated. The first enterprise refers to any enterprise that generates question pairs using the enterprise-level Q&A update system based on a language model. The first uploaded document refers to any enterprise-level document of the first enterprise. First, the first uploaded document is uploaded to the enterprise-level Q&A update system based on a language model. For example, a document recording employee assessment, employee training, welfare allowances, etc. by the human resources department is uploaded to the system. It achieves the technical effect of providing a document data basis for generating Q&A pairs.

[0035] Step S200: Obtain a first converted document by converting the format of the first uploaded document, where the format of the first converted document is the txt format;

[0036] Specifically, the first uploaded document may be in formats such as PDF or Word. Use methods such as online parsing tools to parse the first uploaded document, extract the corresponding text information of the document, and convert it into a document in the txt format, which is the first converted document. By converting the format of the first uploaded document, the technical effect of facilitating the subsequent screening of titles and related content from the document is achieved.

[0037] Step S300: Extract the title and content from the first converted document according to a regular expression to obtain a first title content and a first text content, where the first title content corresponds to the first text content;

[0038] Specifically, the regular expression is a pattern used to describe the characteristics of a set of strings and is used to match specific strings. It is a tool for achieving text matching by describing patterns through special characters + ordinary characters. Use the regular expression to match and screen out all the titles and the text content below the titles in the first converted document, that is, the first title content and the first text content. Among them, the title is the general summary question, and the corresponding content below the title is the answer corresponding to the question. In addition, the title and the text answer correspond one by one, that is, the first title content corresponds to the first text content. By directly extracting the title and the corresponding text answer from the converted document without modifying the text, that is, the corresponding answer, the rigor and accuracy of the rules and regulations and technical operation processes in the enterprise are ensured.

[0039] Step S400: Construct domain keyword information and dependency syntactic information according to the answer text set of all domain documents of the first enterprise;

[0040] Specifically, all domain documents of the first enterprise form the answer text set. Further, based on the answer text set, screen out the keywords in all the answer texts and the dependency relationships between words, that is, obtain the domain keyword information and the dependency syntactic information. Crawl the answer texts from the domain documents through web crawler technology, extract the keywords in each domain document, and jointly form the keyword table with the keywords obtained by part-of-speech analysis and manual proofreading from the answer texts. Further, perform dependency syntactic analysis on the answer texts to obtain the dependency syntactic information, achieving the technical effect of providing corresponding keyword and dependency syntactic data bases for the intelligent generation of question-and-answer pairs in the enterprise-level Q&A update system based on the language model.

[0041] Step S500: Input the domain keyword information and the dependency syntax information into a similar question generation model, where the similar question generation model is trained with multiple groups of training data until convergence.

[0042] Step S600: Obtain a first sequence of similar questions according to the first output information output by the similar question generation model.

[0043] Specifically, the similar question generation model can continuously self-train and learn based on the training data. Each set of data in the multiple groups of training data includes the domain keyword information, the dependency syntax information, and identification information for identifying the first output information. The similar question generation model continuously corrects itself. When the output information of the similar question generation model reaches a predetermined accuracy / convergence state, the supervised learning process ends. The similar question generation model established based on the neural network model can output questions that are accurate and similar to the title. Through the entire iteration, multiple similar questions corresponding to the title are finally generated. Further, all automatically generated similar questions are screened to obtain reasonable similar question information for retention, which is the first sequence of similar questions. By training the data of the similar question generation model, the similar question generation model processes the input data more accurately, and thus the first output information output is also more accurate, achieving the technical effect of accurately obtaining data information and improving the intelligence of the output result.

[0044] Step S700: Combine the information library generated from the first sequence of similar questions and the first title content, and use the update judgment unit to determine whether the preset update trigger condition is met. If it is met, update the first Q&A library.

[0045] Specifically, the update judgment unit is embedded in the enterprise-level Q&A update system based on the language model, and can intelligently perform real-time dynamic updates on the Q&A pairs in the first Q&A library. The preset update trigger condition refers to the situation where there are no reserved reasonable similar question expressions and the content of the first title in the Q&A library. Using the update judgment unit, it is judged whether the preset update trigger condition is met. If the preset update trigger condition is met, the update judgment unit automatically updates the first Q&A library. Among them, the update means automatically accumulating similar question expressions, and the answers corresponding to the questions only change with the changes in the rules and regulations. That is to say, the question library of the first Q&A library is updated and maintained in real time automatically. The reserved reasonable similar question expressions and the titles in the document are used as questions, and the answer text is used as the final answer without any modification. Among them, the title is used as the general summary question, the similar question expressions are added to the similar question group, and the answers are added to the answer group. It is judged whether the question and the corresponding answer exist in the first Q&A library. If not, the Q&A pair in the first output information is added to the first Q&A library. If so, the answer in the first output information is updated. It achieves the technical effect that questions are continuously accumulated and answers only change with the changes in the rules and regulations of the document.

[0046] Further, as shown in the attached Figure 2 figure, step S400 of the embodiment of the present application further includes:

[0047] Step S410: Extract all the documents in the first enterprise based on the first part-of-speech extraction rule to form a first keyword table;

[0048] Step S420: Determine the first affiliated field by performing part-of-speech field feature analysis on the first keyword table;

[0049] Step S430: Obtain a first type of field word table in a similar field to the first keyword table by performing analysis on similar field words in the first affiliated field;

[0050] Step S440: Expand the first keyword table according to the first type of field word table to generate the field keyword information.

[0051] Specifically, for all the domain documents of the first enterprise, extraction is performed based on the first part-of-speech extraction rule to form the first keyword table. Among them, the first part-of-speech extraction rule refers to using pyltp for word segmentation and part-of-speech tagging, and selecting nouns, verbs, and gerunds in all domain documents. Based on the extracted keywords, manual proofreading is carried out to complete the construction of the keyword table. Further, the part of speech and domain of each keyword in the first keyword table are manually judged to determine the domain information involved in the keyword, that is, the first affiliated domain. Then, web crawler technology is used to crawl and find domain words semantically similar to each keyword in the first keyword table from related domains of the first affiliated domain, which is the first type of domain word table. Expanding the first type of domain word table into the first keyword table can generate the domain keyword information. It achieves the technical effect of intelligently obtaining keywords in all domain documents of the enterprise, and improves the technical effect of the number of ways to ask similar questions in the question and answer pairs.

[0052] Further, as shown in the appendix Figure 3 It is shown that step S400 of the embodiment of the present application further includes:

[0053] Step S450: Generate a first annotated text set by performing part-of-speech tagging on the answer text set;

[0054] Step S460: Based on dependency syntactic analysis of the dependency relationship between words in the first annotated text set, obtain a first dependency coefficient set;

[0055] Step S470: Obtain N dependency coefficients greater than or equal to a preset dependency coefficient in the first dependency coefficient set;

[0056] Step S480: Generate the dependency syntactic information according to the syntax corresponding to the N dependency coefficients.

[0057] Specifically, use ltp to perform part-of-speech tagging on all answer texts of each document in the answer text set to generate the first annotated text set. Further, based on dependency syntactic analysis of the dependency relationship between words in the first annotated text set, obtain the first dependency coefficient set. When the number of dependency coefficients in the first dependency coefficient set is greater than or equal to the preset dependency coefficient, obtain the corresponding N dependency coefficients, and generate the dependency syntactic information according to the syntax corresponding to the N dependency coefficients. It achieves the technical effect of obtaining dependency syntactic information, and lays a foundation for subsequent obtaining the dependency relationship between words in a sentence based on dependency syntactic analysis.

[0058] Further, step S300 of the embodiment of the present application further includes:

[0059] Step S310: Input the first converted document into the title answer screening module to obtain the first screening result of the title answer screening module, where the first screening result includes a first screening partition and a second screening partition;

[0060] Step S320: After encoding the first title content and the first text content correspondingly, store the first title content in the first screening partition, and screen the first text content into the second screening partition.

[0061] Specifically, the title answer screening module is embedded in an enterprise-level Q&A update system based on a language model, and can intelligently implement the screening of titles and corresponding text contents in a document. Input the enterprise document converted into txt format into the title answer screening module to automatically obtain the first screening result. Wherein, the first screening result includes a first screening partition and a second screening partition. By encoding the titles and corresponding texts in the converted document intelligently screened by the title answer screening module, that is, the first title content and the first text content correspondingly, and then store the first title content screened in the first screening partition, and store the first text content screened in the second screening partition. Through the title answer screening module, the technical effect of storing questions and corresponding answer texts in partitions is achieved, which is convenient for the system to intelligently manage Q&A pairs.

[0062] Further, as shown in Figure 4 the following, step S700 of the embodiment of the present application further includes:

[0063] Step S710: Determine whether the first similar question sentence is in the first Q&A library;

[0064] Step S720: If the first similar question sentence is not in the first Q&A library, add the first similar question sentence to the first Q&A library;

[0065] Step S730: If the first similar question sentence is in the first Q&A library, update the answer library of the first Q&A library based on the corresponding answer information of the first similar question sentence.

[0066] Specifically, an update judgment unit in the enterprise-level Q&A update system based on a language model is used to determine whether the Q&A pairs automatically generated by the similar question formulation generation model already exist in the first Q&A library. When the Q&A pairs automatically generated by the similar question formulation generation model, that is, the first similar question sentences, are not in the first Q&A library, the title mapping information of the first similar question sentences is added to the title library of the first Q&A library; when the Q&A pairs automatically generated by the similar question formulation generation model, that is, the first similar question sentences, already exist in the first Q&A library, the answer text mapping information of the first similar question sentences is updated to the answer library of the first Q&A library. Similarly, multiple other similar question sentences automatically generated by the similar question formulation generation model are judged in sequence, and finally, the technical effect of continuous accumulation of question formulations while the corresponding answers only change with the changes in document rules and regulations is achieved.

[0067] Further, step S500 of the embodiment of the present application further includes:

[0068] Step S510: Obtain a first training sample data set, where the first training sample data set includes multiple similar question sets;

[0069] Step S520: Construct the similar question formulation generation model according to the first training sample data set;

[0070] The similar question formulation generation model is obtained through training with multiple groups of training data, where the multiple groups of training data include the first training sample data set and the identification information identifying the first output information.

[0071] Specifically, the first training sample data set includes multiple similar question sets, and the similar question formulation generation model is constructed according to the first training sample data set. The similar question formulation generation model is obtained through training with multiple groups of training data, where the multiple groups of training data include the first training sample data set and the identification information identifying the first output information. By using similar question pairs and question sentence pattern question pairs, a similar question formulation generation model based on GPT2 is trained, and the similar question formulation generation model outputs accurate one-to-many mapping information of titles and answers, thus having strong analysis and calculation capabilities and achieving accurate and efficient technical effects.

[0072] Further, step S800 of the embodiment of the present application further includes:

[0073] Step S810: Obtain the first newly added department of the first enterprise;

[0074] Step S820: Determine the first department field according to the first newly added department;

[0075] Step S830: When the first department area is not within all the areas of the first enterprise, obtain a first new instruction;

[0076] Step S840: Generate a first area addition module according to the first new instruction, and implement enterprise area update based on the first area addition module.

[0077] Specifically, the first new department of the first enterprise refers to any newly added department of the first enterprise. According to the first new department, analyze and determine the area information involved in the first new department, that is, the first department area. Through intelligent judgment, when the first department area is not within all the existing areas of the first enterprise, the enterprise-level Q&A update system based on the language model automatically obtains a first new instruction, and then the system automatically generates a first area addition module, and realizes enterprise area update based on the first area addition module. It achieves the technical effect of intelligently and real-time updating the enterprise area and providing a document data basis for the update of Q&A pairs.

[0078] In summary, the enterprise-level Q&A update method based on the language model provided by the embodiments of the present application has the following technical effects:

[0079] 1. By obtaining the first uploaded document of the first enterprise; by performing format conversion on the first uploaded document to obtain a first converted document, where the format of the first converted document is txt format; extracting the title and content of the first converted document respectively according to regular expressions to obtain a first title content and a first text content, where the first title content corresponds to the first text content; constructing domain keyword information and dependency syntax information according to the answer text set of all area documents of the first enterprise; inputting the domain keyword information and the dependency syntax information into a similar question generation model, where the similar question generation model is trained to convergence through multiple groups of training data; obtaining a first similar question sequence according to the first output information output by the similar question generation model; combining the information library generated according to the first similar question sequence and the first title content with the update judgment unit to judge whether the preset update trigger condition is satisfied, and if so, update the first Q&A library. It achieves the technical effect of using the model to generate multiple similar questions based on the title and retaining all the text under the title as the answer, so as to ensure that the answer content is suitable for answering employees' questions and realize the automatic update of the question library at the same time.

[0080] 2. By combining keyword mining and dependency syntax analysis theories, keyword information and dependency syntax information in the answer text are extracted, and the two types of information are fused into the sentence vector of the title to enhance the semantic information of the title. Meanwhile, it is input into the similar question generation model to generate similar questions. Based on multiple similar question forms corresponding to the title, there are both keyword question sentences and interrogative sentence questions. The similar question generation model has strong analysis and calculation capabilities, achieving accurate and efficient technical effects.

[0081] Embodiment 2

[0082] Based on the same inventive concept as an enterprise-level Q&A update method based on a language model in the foregoing embodiments, the present invention further provides an enterprise-level Q&A update system based on a language model. Please refer to the attached Figure 5 , the system includes:

[0083] The first acquisition unit 11: The first acquisition unit 11 is used to acquire the first uploaded document of the first enterprise;

[0084] The second acquisition unit 12: The second acquisition unit 12 is used to obtain a first converted document by converting the format of the first uploaded document, wherein the format of the first converted document is the txt format;

[0085] The third acquisition unit 13: The third acquisition unit 13 is used to extract the title and content of the first converted document according to a regular expression to obtain a first title content and a first text content, wherein the first title content corresponds to the first text content;

[0086] The first construction unit 14: The first construction unit 14 is used to construct domain keyword information and dependency syntax information according to the answer text set of all domain documents of the first enterprise;

[0087] The first input unit 15: The first input unit 15 is used to input the domain keyword information and the dependency syntax information into a similar question generation model, wherein the similar question generation model is trained to convergence through multiple groups of training data;

[0088] The fourth acquisition unit 16: The fourth acquisition unit 16 is used to obtain a first sequence of similar questions according to the first output information output by the similar question generation model; The first update unit 17:

[0089] The first update unit 17 is used to determine whether the information library generated according to the first sequence of similar questions and the first title content meets a preset update trigger condition in combination with the update judgment unit. If it meets, the first Q&A library is updated.

[0090] Further, the system further includes:

[0091] A second construction unit, configured to extract all-domain documents of the first enterprise based on a first part-of-speech extraction rule to form a first keyword table;

[0092] A first determination unit, configured to determine a first affiliated domain by performing part-of-speech domain feature analysis on the first keyword table;

[0093] A fifth acquisition unit, configured to obtain a first type of domain word table in a domain similar to the first keyword table by performing similar-domain word analysis on the first affiliated domain;

[0094] A first generation unit, configured to expand the first keyword table according to the first type of domain word table to generate the domain keyword information.

[0095] Further, the system further includes:

[0096] A second generation unit, configured to generate a first annotated text set by performing part-of-speech annotation on the answer text set;

[0097] A sixth acquisition unit, configured to obtain a first dependency coefficient set based on the dependency relationship between words in the first annotated text set through dependency syntax analysis;

[0098] A seventh acquisition unit, configured to obtain N dependency coefficients greater than or equal to a preset dependency coefficient in the first dependency coefficient set;

[0099] A third generation unit, configured to generate the dependency syntax information according to the syntax corresponding to the N dependency coefficients.

[0100] Further, the system further includes:

[0101] An eighth acquisition unit, configured to input the first converted document into a title answer screening module to obtain a first screening result of the title answer screening module, where the first screening result includes a first screening partition and a second screening partition;

[0102] A first storage unit, configured to store the first title content in the first screening partition and screen the first text content into the second screening partition after encoding the first title content and the first text content correspondingly.

[0103] Further, the system further includes:

[0104] A second judgment unit, which is used to judge whether the first similar question sentence is in the first Q&A library;

[0105] A first addition unit, which is used to add the first similar question sentence to the first Q&A library if the first similar question sentence is not in the first Q&A library;

[0106] A second update unit, which is used to update the answer library of the first Q&A library based on the corresponding answer information of the first similar question sentence if the first similar question sentence is in the first Q&A library. Further, the system further includes:

[0107] A ninth acquisition unit, which is used to acquire a first training sample data set, wherein the first training sample data set includes multiple similar question sets;

[0108] A third construction unit, which is used to construct the similar question generation model according to the first training sample data set;

[0109] A tenth acquisition unit, which is used to obtain the similar question generation model through training with multiple groups of training data, wherein the multiple groups of training data include the first training sample data set and the identification information identifying the first output information.

[0110] Further, the system further includes:

[0111] An eleventh acquisition unit, which is used to acquire a first newly added department of the first enterprise;

[0112] A second determination unit, which is used to determine a first department field according to the first newly added department;

[0113] A twelfth acquisition unit, which is used to acquire a first new instruction when the first department field is not in all the fields of the first enterprise;

[0114] A fourth generation unit, which is used to generate a first field new module according to the first new instruction, and realize enterprise field update based on the first field new module.

[0115] The various embodiments in this specification are described in a progressive manner, and the key point of each embodiment is the difference from other embodiments. The foregoing Figure 1The enterprise-level Q&A update method and specific example in Embodiment 1 are equally applicable to the enterprise-level Q&A update system based on a language model in this embodiment. Through the foregoing detailed description of the enterprise-level Q&A update method based on a language model, those skilled in the art can clearly know the enterprise-level Q&A update system based on a language model in this embodiment. Therefore, for the sake of simplicity of the specification, it will not be described in detail here. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0116] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0117] Exemplary Electronic Device

[0118] Next, reference is made to Figure 6 to describe the electronic device according to an embodiment of the present application.

[0119] Figure 6 FIG. illustrates a schematic structural diagram of an electronic device according to an embodiment of the present application.

[0120] Based on the inventive concept of the enterprise-level Q&A update method based on a language model in the foregoing embodiment, the present invention further provides an enterprise-level Q&A update system based on a language model, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of any of the methods of the enterprise-level Q&A update method described above.

[0121] Among them, in Figure 6 the bus architecture (represented by bus 300), bus 300 may include any number of interconnected buses and bridges, and bus 300 links together various circuits including one or more processors represented by processor 302 and a memory represented by memory 304. Bus 300 may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and therefore, will not be further described herein. Bus interface 305 provides an interface between bus 300 and receiver 301 and transmitter 303. Receiver 301 and transmitter 303 may be the same element, i.e., a transceiver, providing a unit for communicating with various other devices over a transmission medium.

[0122] The processor 302 is responsible for managing the bus 300 and general processing, while the memory 304 can be used to store data used by the processor 302 during operation.

[0123] The present application provides an enterprise-level Q&A update method based on a language model. The method is applied to an enterprise-level Q&A update system based on a language model. The method includes: obtaining a first uploaded document of a first enterprise; obtaining a first converted document by converting the format of the first uploaded document, where the format of the first converted document is the txt format; extracting the title and content of the first converted document respectively according to a regular expression to obtain a first title content and a first text content, where the first title content corresponds to the first text content; constructing domain keyword information and dependency syntax information according to the answer text set of all domain documents of the first enterprise; inputting the domain keyword information and the dependency syntax information into a similar question generation model, where the similar question generation model is trained with multiple sets of training data until convergence; determining whether a preset update trigger condition is satisfied by combining an information library generated according to the first similar question sequence and the first title content with the update judgment unit. If the condition is satisfied, the first Q&A library is updated. This solves the technical problems in the prior art that the extracted Q&A pairs have single problems, unreasonable question forms, answers that are not suitable for answering questions of employees in the enterprise, and the cumbersome and complex manual update of the question bank. It achieves the technical effect of generating multiple similar question forms based on the title using the model and retaining all the text under the title as the answer, thereby ensuring that the answer content is suitable for answering employees' questions and realizing the automatic update of the question bank.

[0124] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, an apparatus, or a computer program product. Therefore, the present application can take the form of a complete software embodiment, a complete hardware embodiment, or an embodiment combining software and hardware aspects. In addition, the present application is in the form of a computer program product that can be implemented on one or more computer-usable storage media containing computer-usable program code. The computer-usable storage media includes, but is not limited to: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disk memories, compact disc read-only memories (CD-ROMs), optical memories, and other media that can store program code.

[0125] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be realized by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate a system for realizing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or one or more blocks.

[0126] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction system that realizes the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or one or more blocks.

[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or one or more blocks. Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0128] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. An enterprise-level Q&A update method based on a language model, wherein, The method is applied to an enterprise-level Q&A update system based on a language model. The system includes an update judgment unit, and the method includes: Obtain the first uploaded document of the first enterprise; Obtain a first converted document by performing format conversion on the first uploaded document, wherein the format of the first converted document is the txt format; Extract the title and content of the first converted document according to a regular expression to obtain a first title content and a first text content, wherein the first title content corresponds to the first text content; Construct domain keyword information and dependency syntax information corresponding to the answer text set according to the answer text set of all domain documents of the first enterprise. Among them, all domain documents of the first enterprise are extracted based on the first part-of-speech extraction rule, and the extracted keywords form a first keyword table; Determine the first affiliated domain by performing part-of-speech domain feature analysis on the first keyword table; Obtain a first type of domain word table in a domain similar to the first keyword table by performing similar domain word analysis on the first affiliated domain; Expand the first keyword table according to the first type of domain word table to generate the domain keyword information; Input the domain keyword information and the dependency syntax information into a similar question generation model, wherein the similar question generation model is trained with multiple groups of training data until convergence; Obtain a first similar question sentence sequence corresponding to the first title content according to the first output information output by the similar question generation model; Combine the information library generated according to the first similar question sentence sequence and the first title content with the update judgment unit to judge whether a preset update trigger condition is satisfied. The preset update trigger condition means that the first similar question sentence sequence and the first title content do not exist in the first Q&A library. If satisfied, update the first Q&A library.

2. The method according to claim 1, wherein for constructing the domain keyword information and the dependency syntax information according to the answer text set of all domain documents of the first enterprise, the method further includes: Generate a first annotated text set by performing part-of-speech annotation on the answer text set; Obtain a first dependency coefficient set based on analyzing the dependency relationship between words in the first annotated text set by dependency syntax; Obtain N dependency coefficients in the first dependency coefficient set that are greater than or equal to a preset dependency coefficient; Generate the dependency syntax information according to the syntax corresponding to the N dependency coefficients.

3. The method according to claim 1, wherein for extracting the title and content of the first converted document according to a regular expression to obtain a first title content and a first text content, the method further includes: Input the first converted document into a title answer screening module to obtain a first screening result of the title answer screening module, wherein the first screening result includes a first screening partition and a second screening partition; After encoding the first title content and the first text content correspondingly, store the first title content in the first screening partition, and screen the first text content into the second screening partition.

4. The method according to claim 1, wherein the information base generated according to the first similar question sequence and the first title content combines with the update judgment unit to judge whether a preset update trigger condition is satisfied. If satisfied, update the first Q&A library. The method further includes: Judge whether the first similar question is in the first Q&A library; If the first similar question is not in the first Q&A library, add the first similar question to the first Q&A library; If the first similar question is in the first Q&A library, update the answer library of the first Q&A library based on the corresponding answer information of the first similar question.

5. The method according to claim 1, wherein the domain keyword information and the dependency syntactic information are input into the similar question generation model, where The similar question generation model is trained to convergence through multiple groups of training data. The method further includes: Obtain a first training sample data set, wherein the first training sample data set includes multiple similar question sets; Construct the similar question generation model according to the first training sample data set; The similar question generation model is obtained by training with multiple groups of training data, wherein the multiple groups of training data include the first training sample data set and the identification information identifying the first output information.

6. The method according to claim 5, the method further includes: Obtain the first newly added department of the first enterprise; Determine the first department field according to the first newly added department; When the first department field is not in all the fields of the first enterprise, obtain a first new instruction; Generate a first field new module according to the first new instruction, and realize enterprise field update based on the first field new module.

7. An enterprise-level Q&A update system based on a language model, wherein, The system includes an update judgment unit, and the system includes: The first obtaining unit: The first obtaining unit is used to obtain the first uploaded document of the first enterprise; The second obtaining unit: The second obtaining unit is used to obtain a first converted document by converting the format of the first uploaded document, wherein the format of the first converted document is txt format; The third obtaining unit: The third obtaining unit is used to extract the title and content of the first converted document respectively according to a regular expression, and obtain the first title content and the first text content, wherein the first title content corresponds to the first text content; The first constructing unit: The first constructing unit is used to construct the domain keyword information and dependency syntax information corresponding to the answer text set according to the answer text set of all the domain documents of the first enterprise. Among them, all the domain documents of the first enterprise are extracted based on the first part-of-speech extraction rule, and the extracted keywords form a first keyword table; Determine the first affiliated field by performing part-of-speech domain feature analysis on the first keyword table; Obtain a first type of domain word table in a field similar to the first keyword table by performing similar field word analysis on the first affiliated field; Expand the first keyword list according to the first type of domain word list to generate the domain keyword information; First input unit: The first input unit is used to input the domain keyword information and the dependency syntax information into the similar question generation model, where the similar question generation model is trained to convergence through multiple sets of training data; Fourth acquisition unit: The fourth acquisition unit is used to obtain a first similar question sequence corresponding to the first title content according to the first output information output by the similar question generation model; First update unit: The first update unit is used to combine the information library generated according to the first similar question sequence and the first title content with the update judgment unit to judge whether the preset update trigger condition is satisfied, where the preset update trigger condition means that the first similar question sequence and the first title content do not exist in the first Q&A library. If satisfied, update the first Q&A library.

8. An enterprise-level Q&A update system based on a language model, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Question and answer pair generation method and device

    CN110196929A

  • Method for extracting question and answer pairs from semi-structured document based on machine learning

    CN111078875A