Method for generating large model question-answering knowledge base and electronic device

By preprocessing and customizing knowledge extraction rules for multi-source corpus data, the problem of poor multi-source corpus knowledge analysis and processing capabilities in the existing technology is solved, and efficient knowledge extraction and question-and-answer pair generation of multi-source corpus data is realized, which improves the context understanding and knowledge reasoning capabilities of the big model.

CN119719389BActive Publication Date: 2025-05-13ZHEJIANG UNIV HIGH-END EQUIP RES INST
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510231692.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-13
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

When the prior art generates a question-and-answer knowledge base for large-scale model training, the analysis and processing capabilities of multi-source corpus knowledge are poor, and it is particularly difficult to process multi-source and multi-format corpus data, and it is impossible to effectively generate multiple rounds of question-and-answer pairs.

Method used

Knowledge extraction is performed based on the data structure by preprocessing multi-source corpus data, including deletion of useless symbols, segmentation processing and custom knowledge extraction rules. For hierarchical table data, it is converted into a two-dimensional data table, parsing the node relationship and generating a data node relationship table. Convert multi-source corpus knowledge into question-and-answer pairs, and deduplication and conflict resolution are carried out to generate a large-scale Q&A knowledge base.

Benefits of technology

The efficiency and adaptability of knowledge extraction and question-and-answer pair generation of multi-source and multi-format corpus data is improved, and multiple rounds of question-and-answer pair generation of data with hierarchical relationships is realized, which improves the context understanding and knowledge reasoning ability of the big model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719389B_ABST
    Figure CN119719389B_ABST
Patent Text Reader

Abstract

The present application relates to a method for generating a large-model question-and-answer knowledge base and an electronic device. The method includes: preprocessing the acquired multi-source corpus data to obtain preprocessed multi-source corpus data; extracting knowledge from the preprocessed multi-source corpus data based on the data structure of the preprocessed multi-source corpus data to obtain multi-source corpus knowledge; converting the multi-source corpus knowledge into question-and-answer pairs; deduplicating and resolving conflicts on the question-and-answer pairs to generate a large-model question-and-answer knowledge base. The present application solves the technical problem that the current technology for generating a question-and-answer knowledge base for large-model training has poor parsing and processing capabilities for multi-source corpus knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing, and in particular to a method for generating a large-model question-and-answer knowledge base and an electronic device. Background Art

[0002] With the rapid development of artificial intelligence technology, large models based on natural language processing have demonstrated excellent performance in knowledge question answering, text generation and other tasks. In practical applications, the effectiveness and intelligence of such large models largely depend on high-quality question-answering knowledge bases. High-quality question-answering knowledge bases can significantly improve the performance of large models in question-answering tasks and provide important support for the intelligent application of large models. How to extract key knowledge from corpus data with diverse sources and complex formats and form high-quality question-answer pairs has become a key issue in industry research.

[0003] In terms of knowledge extraction, traditional knowledge extraction and fusion technologies are unable to efficiently process corpus data from different sources and formats. It is difficult to accurately understand the implicit semantics of unstructured text data that is long and has complex semantics. Text knowledge extraction methods are not very scalable. For structured data, especially structured data with hierarchical relationships, it is impossible to parse and maintain logical relationships.

[0004] In terms of question-answer pair generation, the existing large-model question-answer knowledge base can realize single-round question-answer pair generation based on corpus data, but has not yet achieved multi-round question-answer pair generation for data with hierarchical relationships.

[0005] To address the above problems, no effective solution has been proposed yet. Summary of the invention

[0006] The embodiments of the present application provide a method for generating a large-model question-and-answer knowledge base and an electronic device, which at least solve the problem that the current technology for generating a question-and-answer knowledge base for large-model training has poor parsing and processing capabilities for multi-source corpus knowledge.

[0007] According to one aspect of an embodiment of the present application, a method for generating a large-model question-and-answer knowledge base is provided, comprising: preprocessing the acquired multi-source corpus data to obtain preprocessed multi-source corpus data; extracting knowledge from the preprocessed multi-source corpus data based on the data structure of the preprocessed multi-source corpus data to obtain multi-source corpus knowledge; converting the multi-source corpus knowledge into question-and-answer pairs; deduplicating and resolving conflicts on the question-and-answer pairs to generate a large-model question-and-answer knowledge base.

[0008] Optionally, in the case that the preprocessed multi-source corpus data is text data, knowledge extraction is performed on the preprocessed multi-source corpus data based on the data structure of the preprocessed multi-source corpus data, including: deleting useless symbols in the text data, wherein the useless symbols include at least: blank characters and space characters; if the length of the text data after deleting the useless symbols exceeds a first preset length, segmenting the text data; and using a large model to extract knowledge from the text data based on customized large model knowledge extraction rules.

[0009] Optionally, in the case where the preprocessed multi-source corpus data is hierarchical tabular data, knowledge extraction is performed on the preprocessed multi-source corpus data based on the data structure of the preprocessed multi-source corpus data, including: converting the hierarchical tabular data into a two-dimensional data table, wherein each row of data in the two-dimensional data table includes: an index, a number and data content; traversing the two-dimensional data table, and determining the node position corresponding to each row of data and the relationship with other nodes according to the number of the row of data; associating the data with parent-child relationship in the two-dimensional data table through the index, and recording the parent node and child node of each row of data to generate a data node relationship table; and converting each row of data in the data node relationship table into data in JSON format.

[0010] Optionally, traverse the two-dimensional data table, and determine the node position corresponding to each row of data and the relationship with other nodes according to the number of each row of data, including: setting the name of the two-dimensional data table to the root node; converting the number included in each row of data in the two-dimensional data table into a string, if the string does not contain the target symbol, determine the row of data corresponding to the string as the first-level parent node, and record the root node as the parent node of the first-level parent node; if the string contains the target symbol, determine the row of data corresponding to the string as a child node, wherein the parent node of the child node is the row of data corresponding to the number of the row of data corresponding to the child node minus the lowest level; if the number contained in the string is the target number, determine the row of data corresponding to the target number as the data content of the parent node or the child node.

[0011] Optionally, before converting the hierarchical table data into a two-dimensional data table, the method further includes: if the length of the text content in the hierarchical table data exceeds a second preset length, using a preset model to simplify and refine the text content in the hierarchical table data.

[0012] Optionally, in the case where the preprocessed multi-source corpus data is non-hierarchical tabular data, knowledge extraction is performed on the preprocessed multi-source corpus data based on the data structure of the preprocessed multi-source corpus data, including: if the length of the text content in the non-hierarchical tabular data exceeds a third preset length, using a preset model to simplify and refine the text content in the non-hierarchical tabular data.

[0013] Optionally, if the multi-source corpus knowledge is non-hierarchical corpus knowledge, the multi-source corpus knowledge is converted into question-answer pairs, including: setting a question-answer pair template according to the source document content or source table content of the multi-source corpus knowledge; collecting multi-source corpus knowledge that can use the question-answer pair template based on the name of the source document or source table; mapping the collected multi-source corpus knowledge into the question-answer pair template to generate a question-answer pair.

[0014] Optionally, if the multi-source corpus knowledge is hierarchical corpus knowledge, the multi-source corpus knowledge is converted into question-answer pairs, including: setting a multi-round question-answer pair template based on the node relationship and node data content in the hierarchical corpus knowledge, wherein the multi-round question-answer pair template includes a plurality of hierarchical question-answer pair templates, and each hierarchical question-answer pair template includes a question template and an answer template; traversing from the root node of the hierarchical corpus knowledge, mapping each node to the question template of the hierarchical question-answer pair template where each node is located in turn, searching the child nodes of each node from the data in JSON format to generate the answer required in the answer template; setting a question-answer pair index for each generated question-answer pair, and recording the index of the question-answer pair at the previous level to maintain the logical relationship between the multi-round question-answer pairs.

[0015] Optionally, before deduplication and conflict resolution of question and answer pairs, the above method also includes: vectorizing the question, answer and document name or table name of each question and answer pair respectively to generate vector triples, wherein the vector triples include: a question semantic vector, an answer semantic vector and a document name vector or a table name vector; performing vector matching on the document name vectors or table name vectors of the question and answer pairs, and if the similarity between two document name vectors or table name vectors is greater than a preset similarity threshold, determining that the two documents or tables corresponding to the two document name vectors or table name vectors are related; if the similarity between the two document name vectors or table name vectors is lower than the preset similarity threshold, determining that the two documents or tables corresponding to the two document name vectors or table name vectors are unrelated.

[0016] Optionally, after performing vector matching on the document name vector or table name vector of the question-answer pair, the above method also includes: simultaneously performing a matching search on the question semantic vector and the answer semantic vector in the question-answer pair in which the document or table is a related document or a related table; when the similarity between the question semantic vector and the answer semantic vector of two question-answer pairs is simultaneously higher than a semantic overlap threshold, determining that the two question-answer pairs are overlapping question-answer pairs; when the similarity between the question semantic vectors of the two question-answer pairs is higher than a semantic overlap threshold, and the similarity between the answer semantic vectors is lower than a semantic conflict threshold, determining that the two question-answer pairs are conflicting question-answer pairs.

[0017] Optionally, duplicate removal and conflict resolution are performed on the question-answer pairs, including: retaining question-answer pairs with high answer coverage among overlapping question-answer pairs; retaining question-answer pairs with high credibility of knowledge source documents among conflicting question-answer pairs.

[0018] According to another aspect of an embodiment of the present application, there is also provided an electronic device, including: a processor, and a memory storing a program, wherein the program includes instructions, and when the instructions are executed by the processor, the processor executes the above method.

[0019] According to another aspect of an embodiment of the present application, a non-transitory machine-readable medium storing computer instructions is also provided, where the computer instructions are used to enable a computer to execute the above method.

[0020] Beneficial effects of the embodiments of the present application:

[0021] In an embodiment of the present application, the acquired multi-source corpus data is preprocessed to obtain preprocessed multi-source corpus data; knowledge is extracted from the preprocessed multi-source corpus data based on the data structure of the preprocessed multi-source corpus data to obtain multi-source corpus knowledge; the multi-source corpus knowledge is converted into question-answer pairs; the question-answer pairs are deduplicated and conflict resolved to generate a large-model question-answer knowledge base. By setting knowledge extraction rules based on the structured and unstructured characteristics of the multi-source corpus data, knowledge extraction and question-answer pair generation are achieved for the multi-source and multi-format corpus data, thereby achieving the technical effect of improving the adaptability and efficiency of data processing and enhancing the context understanding ability and knowledge reasoning ability of the large model.

[0022] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other embodiments can be obtained based on these drawings without creative work.

[0024] Figure 1 It is a flowchart of a method for generating a large model question-answering knowledge base according to an embodiment of the present application;

[0025] Figure 2 It is a structural schematic diagram of a system for generating a large model question-answering knowledge base according to an embodiment of the present application;

[0026] Figure 3 Schematic diagram of the structure of the electronic device of this embodiment. DETAILED DESCRIPTION

[0027] Embodiments of the present embodiment will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present embodiment are shown in the accompanying drawings, it should be understood that the present embodiment can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, which are instead provided for a more thorough and complete understanding of the present embodiment. It should be understood that the drawings and embodiments of the present embodiment are only for exemplary purposes and are not intended to limit the scope of protection of the present embodiment.

[0028] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:

[0029] Multi-source corpus refers to a collection of language data collected and integrated from multiple different sources, which may include but are not limited to books, news articles, social media posts, forum discussions, professional literature, government documents, emails, chat records, etc. The characteristic of multi-source corpus is its diversity, which is not only reflected in the breadth of sources, but also in language style, subject range, stylistic features and other aspects.

[0030] Big model: A big model based on natural language processing refers to a large neural network model trained by deep learning technology that can understand and generate human language.

[0031] In the related art, the technology for generating question-answering knowledge bases for large model training has poor parsing and processing capabilities for multi-source corpus knowledge. Some existing methods for generating large model question-answering knowledge bases are as follows:

[0032] 1. Rule-based knowledge extraction technology

[0033] Knowledge is extracted from corpus data through artificially designed rules such as keyword matching and regular expressions. This method is simple and easy to use, and is highly efficient for clear rule scenarios, but it lacks flexibility, is difficult to handle complex semantic relationships, cannot adapt to changes in data formats, and performs poorly for unstructured and long corpus data.

[0034] 2. Knowledge extraction technology based on traditional machine learning

[0035] Traditional machine learning methods identify entities, extract relationships, and generate structured knowledge from texts by preprocessing and feature engineering data, using classifiers or sequence annotation models. In the implementation process, the text is first segmented, annotated, and cleaned to extract features such as part of speech, context windows, and syntactic dependencies. Machine learning (such as support vector machines, random forests, or conditional random fields) is used to name and recognize entities, classify semantic relationships between entities, and train models and verify performance through annotated data. Finally, the extracted entities and relationships are integrated into structured data. This method can handle certain complex semantics and contexts based on supervised learning, and knowledge extraction is more accurate and efficient. However, this method relies on a large amount of annotated data as a training set, has high construction and maintenance costs, and has limited processing capabilities for domain knowledge and complex semantics, making it difficult to expand to multiple data forms.

[0036] 3. Knowledge base construction technology based on knowledge graph

[0037] The knowledge base construction technology based on knowledge graphs extracts entities, attributes and relationships from multi-source data and organizes knowledge into a structured graph network. First, data preprocessing is performed from structured, semi-structured and unstructured data, including cleaning, word segmentation and formatting. Then, the core entities in the data, their attributes and the semantic relationships between entities are extracted through technologies such as named entity recognition, dependency syntax analysis and relationship classification. Then, knowledge fusion methods such as entity alignment and conflict resolution are used to deduplicate and unify knowledge from different sources. Finally, these entities and relationships are stored in a graph database (such as Neo4j) in a graph structure. In this way, knowledge graphs can transform complex domain knowledge into intuitive and usable structured networks. However, the relationship definition of this method relies on a lot of manual participation, and its real-time update and scalability are somewhat complex.

[0038] In order to solve the above problems, relevant solutions are provided in the embodiments of the present application, which are described in detail below.

[0039] Figure 1 is a flow chart of a method for generating a large model question-answering knowledge base according to an embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps:

[0040] Step S102, preprocessing the acquired multi-source corpus data to obtain preprocessed multi-source corpus data.

[0041] In the embodiment of the present application, the user can select the corpus data that he wants to process and upload it. Specifically, the corpus data formats supported for uploading include: CSV, PDF, PPT, EXCEL, and WORD.

[0042] CSV is a file format used to store tabular data, such as numbers and text. Each record is represented by a line, and the different fields in each line are separated by commas or other delimiters. CSV files are often used to exchange data between different applications, especially when spreadsheets or databases are involved.

[0043] In this step, preprocessing of multi-source corpus data according to data format mainly refers to:

[0044] For CSV data and EXCEL data: delete the data with empty values ​​in the key fields.

[0045] For PPT data and PDF data: if the uploaded data is in PPT format, it will be converted into PDF format first; then the pdfplumber library in python will be called to separate the text and table in the PDF document respectively; the separated tables will be subjected to irregular table removal, cell merging and cross-page table processing.

[0046] Step S104: extracting knowledge from the preprocessed multi-source corpus data based on the data structure of the preprocessed multi-source corpus data to obtain multi-source corpus knowledge.

[0047] The preprocessed corpus data structure is mainly divided into three categories: text data, hierarchical table data, and non-hierarchical table data. Specifically, text data refers to word documents uploaded by users and text extracted from PDF documents; table data refers to CSV data, EXCEL data, and tables extracted from PDF documents. Hierarchical table data refers to table data with a tree structure between rows and rows, including parent-child relationships; and non-hierarchical table data refers to table data that does not form the above tree structure between rows. When extracting knowledge, different methods will be used according to the data structure of the corpus data.

[0048] Step S106, converting the multi-source corpus knowledge into question-answer pairs.

[0049] The corpus knowledge obtained in step S104 is converted into a question-answer pair format for use in fine-tuning training of industrial large models.

[0050] Step S108, deduplicate and resolve conflicts in question-answer pairs to generate a large model question-answer knowledge base.

[0051] In some optional embodiments of the present application, the question and answer knowledge is deduplicated and conflict resolved based on semantic vectorization technology and the FAISS matching algorithm to improve the quality of question and answer pairs in the knowledge base.

[0052] Semantic vectorization is a core technology in the field of natural language processing. It aims to convert text data into numerical vector representations so that these vectors can reflect the semantic similarities between words, phrases or sentences in geometric space.

[0053] FAISS (Facebook AI Similarity Search) is an efficient similarity search library developed by Facebook AI Research, specifically for fast Approximate Nearest Neighbor (ANN) search in large vector datasets. It is particularly suitable for processing high-dimensional sparse or dense vectors, and can provide sublinear query performance on very large datasets. FAISS is widely used in information retrieval, recommendation systems, image search, natural language processing and other fields.

[0054] The above method provided in the embodiment of the present application formulates knowledge extraction rules based on the structured and unstructured characteristics of multi-source corpus data, realizes knowledge extraction and question-answer pair generation for multi-source and multi-format corpus data, thereby achieving the technical effect of improving the adaptability and efficiency of data processing and enhancing the context understanding ability and knowledge reasoning ability of large models.

[0055] According to some optional embodiments of the present application, when the preprocessed multi-source corpus data is text data, step S104 is executed to extract knowledge from the preprocessed multi-source corpus data based on the data structure of the preprocessed multi-source corpus data, including the following steps: deleting useless symbols in the text data, wherein the useless symbols include at least blank characters and space characters; if the length of the text data after deleting the useless symbols exceeds a first preset length, segmenting the text data; and using a large model to extract knowledge from the text data based on customized large model knowledge extraction rules.

[0056] Knowledge extraction from text data is based on a large model. The specific steps are as follows:

[0057] 1) Delete useless symbols such as blank characters and spaces in text data.

[0058] 2) If the text is short, this step can be skipped; if the text is long, the text needs to be segmented. In this application, the text is segmented by identifying the text chapter titles. Specifically, a short text length means that the number of words in the text (including punctuation) is less than the number of tokens acceptable to the selected large model, and a long text length means that the number of words in the text (including punctuation) is more than the number of tokens acceptable to the selected large model.

[0059] Here, "token" usually refers to the basic unit when a language model processes text, which can be a word, a character, or a subword unit, depending on the word segmentation strategy used by the model.

[0060] 3) Formulate big model knowledge extraction rules, which refer to the prompts input to the big model, used to guide the big model to focus on specific content and output expected structured information when understanding and generating. "Prompt" refers to the input text or instructions provided to the model.

[0061] In this step, users can define large model knowledge extraction rules according to the specific content of the text. When formulating the rules, it is necessary to clarify the specific tasks and output format of the model to ensure the accuracy and effectiveness of knowledge extraction.

[0062] 4) The large model will understand the input text and extract knowledge according to the established knowledge extraction rules, and the final output results will be stored in the database in the form of a structured data table, as shown in Table 1.

[0063] Table 1

[0064]

[0065] According to other optional embodiments of the present application, when the preprocessed multi-source corpus data is hierarchical tabular data, step S104 is executed to extract knowledge from the preprocessed multi-source corpus data based on the data structure of the preprocessed multi-source corpus data, including the following steps: converting the hierarchical tabular data into a two-dimensional data table, wherein each row of data in the two-dimensional data table includes: an index, a number and data content; traversing the two-dimensional data table, and determining the node position corresponding to each row of data and the relationship with other nodes according to the number of the row of data; associating the data with parent-child relationship in the two-dimensional data table through the index, and recording the parent node and child node of each row of data to generate a data node relationship table; and converting each row of data in the data node relationship table into data in JSON format.

[0066] The key to processing hierarchical table data is to parse the logical relationship between rows in the table. The solution provided in this application identifies the number of each row of data and determines its node position and node relationship with other data.

[0067] The following example further illustrates the method for processing hierarchical table data in this application:

[0068] The existing one-level table data is shown in Table 2:

[0069] Table 2

[0070]

[0071] The processing steps are as follows:

[0072] 1) Use Python to convert the hierarchical table data in Table 2 into a Dataframe table (i.e. the two-dimensional data table mentioned above), and set an index for each row of data, as shown in Table 3 below:

[0073] Table 3

[0074]

[0075] In data analysis and data science, DataFrame is a two-dimensional, resizable, potentially heterogeneous data structure. It consists of rows and columns, each column can contain different types of values ​​(integers, strings, floating-point numbers, etc.), but the data types within the same column are usually the same.

[0076] 2) Traverse the Dataframe table and determine the node position and its relationship with other nodes based on the number of each row of data.

[0077] As some optional embodiments of the present application, a two-dimensional data table is traversed, and the node position corresponding to each row of data and the relationship between the row of data and other nodes are determined according to the number of each row of data, which is achieved by the following method: the name of the two-dimensional data table is set as the root node; the number included in each row of data in the two-dimensional data table is converted into a string, and if the string does not contain the target symbol, the row of data corresponding to the string is determined as the first-level parent node, and the root node is recorded as the parent node of the first-level parent node; if the string contains the target symbol, the row of data corresponding to the string is determined as a child node, wherein the parent node of the child node is the row of data corresponding to the number of the row of data corresponding to the child node minus the lowest level; if the number contained in the string is the target number, the row of data corresponding to the target number is determined to be the data content of the parent node or the child node.

[0078] In an embodiment of the present application, the table name is set as the root node, and then the number is converted into a string format. If the string does not contain "." (i.e., the target symbol mentioned above), the row of data is a first-level parent node, such as number 1 and number 2, and the root node is recorded as the parent node of the first-level parent node (i.e., the table name); if there is "." in the number after conversion into a string format, the row of data is a child node, and its parent node is the number after removing one level, for example, "1.2.1" is numbered "1.2" after removing one level, and "1.2" is the parent node of "1.2.1"; if the number is NAN (i.e., the target number mentioned above), it is the data content of the node.

[0079] Next, the data with "parent-child" relationship is associated through indexing, and the parent node and child node of each row of data are recorded to generate a data node relationship table as shown in Table 4 below.

[0080] Table 4

[0081]

[0082] 3) Based on the generated data node relationship table, the data is converted into JSON format. For each row of data, its index, number, node name, data content, parent node index, and child node index are recorded. Taking the data with index 1 as an example, the data converted into JSON format is as follows:

[0083] {

[0084] Index: "1"

[0085] Node name: "Component a.1"

[0086] Function: "Function a.1"

[0087] Parent node index: [0]

[0088] Child node index: [3]

[0089] }.

[0090] As some optional embodiments of the present application, before converting the hierarchical table data into a two-dimensional data table, if the length of the text content in the hierarchical table data exceeds a second preset length, a preset model is used to simplify and refine the text content in the hierarchical table data.

[0091] In the embodiment of the present application, if the text content in the hierarchical table data is long or the expression is relatively redundant and complex, the text content can be input into a large model for simplification and refinement. The large model here is a pre-trained neural network model.

[0092] Through the above method, it is possible to process long texts and corpus data with complex structures, and maintain the logical relationship between data while extracting text knowledge.

[0093] According to another optional embodiment of the present application, when the preprocessed multi-source corpus data is non-hierarchical tabular data, step S104 is executed to extract knowledge from the preprocessed multi-source corpus data based on the data structure of the preprocessed multi-source corpus data, including the following steps: if the length of the text content in the non-hierarchical tabular data exceeds a third preset length, a preset model is used to simplify and refine the text content in the non-hierarchical tabular data.

[0094] Since non-hierarchical table data has a simple structure and there is no association between rows, there is no need to parse the relationship between data like hierarchical table data. If the text content in such data is long or the expression is redundant and complex, the text content can be input into the large model for simplification and refinement. If the text content is already concise and clear enough, this step can be skipped.

[0095] In some optional embodiments of the present application, if the multi-source corpus knowledge is non-hierarchical corpus knowledge, step S106 is executed to convert the multi-source corpus knowledge into question-answer pairs, including the following steps: setting a question-answer pair template according to the source document content or source table content of the multi-source corpus knowledge; collecting multi-source corpus knowledge that can use the question-answer pair template based on the name of the source document or source table; mapping the collected multi-source corpus knowledge to the question-answer pair template to generate a question-answer pair.

[0096] In an embodiment of the present application, multiple rounds of question-answer pair generation are performed on corpus knowledge containing a hierarchical structure to improve the context understanding and knowledge reasoning capabilities of a large model.

[0097] A question-answer pair consists of two parts: a question and an answer. The question is the input provided to the model during large model training to help the model identify user intent, question type, and question background. The answer is a direct output of the question, used to guide the model to generate accurate answers, and serves as a supervisory signal during model training to optimize model parameters.

[0098] The method for generating knowledge question-answer pairs for non-hierarchical corpus is as follows:

[0099] 1) Create a question-answer pair template based on the document / table content of the corpus knowledge source. For example, if the document is a user manual for a product, the question-answer pair template can be: {Question: "Please tell me how to use {component name} of {product name}"; Answer: "How to use {component name} of {product name} is {related corpus knowledge content}"}.

[0100] 2) Based on the source document / table name, query and collect corpus knowledge that can use the question-answer pair template.

[0101] 3) Map the collected corpus knowledge one by one to the designed question-answer pair template to generate question-answer pairs.

[0102] In some other optional embodiments of the present application, if the multi-source corpus knowledge is hierarchical structure corpus knowledge, step S106 is executed to convert the multi-source corpus knowledge into question-answer pairs, which is achieved by the following method: setting a multi-round question-answer pair template based on the node relationship and node data content in the hierarchical structure corpus knowledge, wherein the multi-round question-answer pair template includes a plurality of hierarchical question-answer pair templates, and each hierarchical question-answer pair template includes a question template and an answer template; traversing from the root node of the hierarchical structure corpus knowledge, mapping each node to the question template of the hierarchical question-answer pair template where each node is located in turn, searching the child nodes of each node from the data in JSON format to generate the answer required in the answer template; setting a question-answer pair index for each generated question-answer pair, and recording the index of the question-answer pair of the previous level to maintain the logical relationship between the multi-round question-answer pairs.

[0103] The method for generating knowledge question-answer pairs for hierarchical corpus is as follows:

[0104] 1) Based on the node relationships and node data content in the corpus knowledge, a multi-round question-answering template is formulated. Taking the hierarchical table data (Table 2) in the above text as an example, the following multi-round question-answering template is formulated:

[0105] {

[0106] { Level: Level 1

[0107] Question: "Please tell me what {device name} is made of"

[0108] Answer: "{Device Name} consists of {System}"}

[0109] { Level: Level 2

[0110] Question: "Please tell me what {system} is made of and what each part does"

[0111] Answer: "{system} is made up of {components}, and its function is {function}"}

[0112] { Level: Level 3

[0113] Question: "Please tell me what {component} is made of, and what each unit does"

[0114] Answer: "{component} is composed of {units}, and its function is {function}"}

[0115] }.

[0116] 2) Each hierarchical structure corpus data to which the question-answer pair template can be applied is traversed starting from the root node. When processing each node, it is first mapped to the question template designed for that level based on the level it is in, and then the child nodes of the node are queried through the "" field in the JSON file to generate an answer, thereby continuously recursively realizing multiple rounds of question-answer pair generation.

[0117] 3) Assign a unique question-answer pair index to each generated question-answer pair, and record the index of its previous question-answer pair to maintain the logical relationship between multiple rounds of question-answer pairs.

[0118] Through the dynamic generation of multiple rounds of questions and answers, the large model question and answer knowledge base has stronger context understanding capabilities and continuous dialogue logic, meeting the needs of multiple rounds of dialogue in large models in applications. It provides high-quality training data for large model training and improves the knowledge adaptation ability and reasoning performance of large models in practical applications.

[0119] According to some optional embodiments of the present application, before executing step S108 to deduplicate and resolve conflicts with the question and answer pairs, it is also necessary to vectorize the questions, answers and document names or table names of each question and answer pair respectively to generate vector triples, wherein the vector triples include: a question semantic vector, an answer semantic vector and a document name vector or a table name vector; vector matching is performed on the document name vectors or table name vectors of the question and answer pairs, and if the similarity between the two document name vectors or table name vectors is greater than a preset similarity threshold, it is determined that the two documents or tables corresponding to the two document name vectors or table name vectors are related; if the similarity between the two document name vectors or table name vectors is lower than the preset similarity threshold, it is determined that the two documents or tables corresponding to the two document name vectors or table name vectors are unrelated.

[0120] The question-answer pairs obtained after step S106 are deduplicated and conflict resolved based on semantic vectorization and FAISS matching algorithm to further improve the quality of question-answer knowledge. The specific implementation steps are as follows:

[0121] 1) Use the BERT model to convert the generated question-answer pairs into semantic vectors. Specifically, the question, answer, and document / table name of each question-answer pair are vectorized, and finally a vector triple (question semantic vector, answer semantic vector, document / table name vector) is generated.

[0122] BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model proposed by Google in 2018, which has triggered major changes in the field of natural language processing. The main innovation of BERT lies in its bidirectional training mechanism, which enables the model to understand the context of the text from left to right and from right to left at the same time, thereby more accurately capturing the relationship between words.

[0123] 2) Considering the large number of generated question-answer pairs, we first perform vector matching on the text / table names to which the question-answer pairs belong when performing semantic matching. Since the text / table names to which the corpus knowledge belongs basically contain the entities and industry fields involved in the corpus knowledge, this step achieves coarse-grained screening of question-answer pairs, which can narrow the matching scope and improve matching efficiency.

[0124] Semantic matching is implemented based on the FAISS matching algorithm. The FAISS matching algorithm can calculate the similarity of a large number of semantic vectors by establishing an inner product index. A similarity threshold is set. When the vector representation similarity of the two document / table names is greater than this threshold, the two documents / tables are considered related. If the vector representation similarity of the two document / table names is less than this threshold, the two documents / tables are considered unrelated.

[0125] As an optional embodiment of the present application, after vector matching of the document name vector or the table name vector of the question-answer pair, a matching search is simultaneously performed on the question semantic vector and the answer semantic vector in the question-answer pair where the document or table is a related document or a related table; when the similarity of the question semantic vector and the answer semantic vector of two question-answer pairs is simultaneously higher than the semantic overlap threshold, the two question-answer pairs are determined to be overlapping question-answer pairs; when the similarity of the question semantic vectors of the two question-answer pairs is higher than the semantic overlap threshold, and the similarity of the answer semantic vectors is lower than the semantic conflict threshold, the two question-answer pairs are determined to be conflicting question-answer pairs.

[0126] 3) In question-answer pairs where the documents / tables are related documents / tables, the question semantic vector and the answer semantic vector are matched and searched simultaneously, which is still based on the FAISS matching algorithm. Two thresholds need to be set in this step: the semantic overlap threshold and the semantic conflict threshold. When the similarity between the question semantic vector and the answer semantic vector of two question-answer pairs is higher than the semantic overlap threshold at the same time, they are judged as overlapping question-answer pairs and need to be deduplicated; when the similarity between the question semantic vectors of two question-answer pairs is higher than the semantic overlap threshold and the similarity between the answer semantic vectors is lower than the semantic conflict threshold, they are judged as conflicting question-answer pairs and need to be resolved.

[0127] In an optional embodiment of the present application, step S108 is executed to deduplicate and resolve conflicts in question-answer pairs, which is achieved by: retaining question-answer pairs with high answer coverage among overlapping question-answer pairs; retaining question-answer pairs with high credibility of knowledge source documents among conflicting question-answer pairs.

[0128] 4) Process the question-and-answer knowledge that needs to be deduplicated or conflict resolved as follows:

[0129] Deduplication of question and answer knowledge: For question and answer knowledge with highly overlapping semantics, only one piece needs to be retained. Question and answer knowledge with higher answer coverage can be retained.

[0130] Question and answer knowledge conflict resolution: Questions with similar semantics but lower semantic similarity in answers will be judged as conflicting question and answer pairs. You can choose to retain the knowledge source document with higher credibility.

[0131] This application innovatively proposes a method for constructing the above-mentioned corpus question-answer knowledge base for large model training, which can realize knowledge extraction and question-answer pair generation for multi-source and multi-format data, formulate knowledge extraction rules based on the structured and unstructured characteristics of multi-source corpus data, and improve the adaptability and efficiency of data processing. In particular, when extracting knowledge from data with hierarchical relationships, the logical relationship between the data is extracted by identifying and parsing the corresponding numbers, and efficient maintenance of the data structure is achieved by recording the indexes of adjacent nodes. Based on this method, the generation of multiple rounds of question-answer pairs is also achieved by querying the node positions and node relationships of the corpus data, which improves the context understanding ability and knowledge reasoning ability of the large model.

[0132] Figure 2 is a structural diagram of a system for generating a large model question-answering knowledge base according to an embodiment of the present application, such as Figure 2 As shown, the system includes:

[0133] The data uploading module 20 is used to upload the multi-source corpus data to be processed.

[0134] In the embodiment of the present application, the user can select the corpus data that he wants to process and upload it. Specifically, the corpus data formats supported for uploading include: CSV, PDF, PPT, EXCEL, and WORD.

[0135] The data preprocessing module 21 is used to preprocess the uploaded multi-source corpus data.

[0136] In the embodiments of the present application, preprocessing the multi-source corpus data according to the data format mainly refers to:

[0137] For CSV data and EXCEL data: delete the data with empty values ​​in the key fields.

[0138] For PPT data and PDF data: if the uploaded data is in PPT format, it will be converted into PDF format first; then the pdfplumber library in python will be called to separate the text and table in the PDF document respectively; the separated tables will be subjected to irregular table removal, cell merging and cross-page table processing.

[0139] The knowledge extraction module 22 is used to extract knowledge from the pre-processed text data, hierarchical structure data, and non-hierarchical structure data.

[0140] The knowledge extraction rule management module 23 is used to manage the large model knowledge extraction rules defined by the user. The user can add, delete, modify and check the large model knowledge extraction rules in this module.

[0141] The question-answer pair generation module 24 is used to generate single-round question-answer pairs and multi-round question-answer pairs based on the extracted knowledge and the question-answer pair templates set by the user.

[0142] The question-answer template management module 25 is used to manage the question-answer templates prepared by the user. The user can add, delete, modify and check the question-answer templates in this module.

[0143] The question-answer knowledge fusion module 26 is used to remove duplication and resolve conflicts in the question-answer knowledge in the knowledge base based on semantic vectorization and FAISS matching algorithm.

[0144] It should be noted that Figure 2 The preferred implementation of the embodiment shown can be referred to Figure 1 The relevant description of the illustrated embodiment will not be repeated here.

[0145] The embodiment of the present application also provides an electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication. The memory stores a computer program executable by the at least one processor, and the computer program is used to enable the electronic device to perform the method of the embodiment of the present application when executed by the at least one processor.

[0146] An embodiment of the present application also provides a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiment of the present application.

[0147] The embodiment of the present application also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiment of the present application.

[0148] refer to Figure 3, the structural block diagram of the electronic device that can be used as the server or client of the embodiment of the present application will now be described, which is an example of the hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples, and are not intended to limit the implementation of the present application described and / or required herein.

[0149] like Figure 3 As shown, the electronic device includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device can also be stored. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0150] Multiple components in the electronic device are connected to the I / O interface 305, including: an input unit 306, an output unit 307, a storage unit 308, and a communication unit 309. The input unit 306 can be any type of device that can input information to the electronic device, and the input unit 306 can receive input digital or character information, and generate key signal input related to user settings and / or function control of the electronic device. The output unit 307 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 308 can include but is not limited to a disk, an optical disk. The communication unit 309 allows the electronic device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0151] The computing unit 301 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a CPU, a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 301 performs the various methods and processes described above. For example, in some embodiments, the method embodiments of the present application may be implemented as a computer program, which is tangibly contained in a machine-readable medium, such as a storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on an electronic device via a ROM 302 and / or a communication unit 309. In some embodiments, the computing unit 301 may be configured to perform the above method in any other appropriate manner (e.g., by means of firmware).

[0152] The computer program for implementing the method of the embodiment of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.

[0153] In the context of the embodiments of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0154] It should be noted that the term "including" and its variations used in the embodiments of the present application are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "multiple" mentioned in the embodiments of the present application are illustrative and not restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0155] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0156] The various steps described in the method implementation methods provided in the embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method implementation methods may include additional steps and / or omit the steps shown. The scope of protection of the present application is not limited in this respect.

[0157] The term "embodiment" in this specification refers to specific features, structures or characteristics described in conjunction with the embodiment that can be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. The various embodiments in this specification are described in a related manner, and the same or similar parts between the various embodiments refer to each other. In particular, for the device, equipment, and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiment.

[0158] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of patent protection. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the attached claims.

Claims

1. A method for generating a large model question-answering knowledge base, characterized in that: include: Preprocessing the acquired multi-source corpus data to obtain preprocessed multi-source corpus data; Extracting knowledge from the preprocessed multi-source corpus data based on the data structure of the preprocessed multi-source corpus data to obtain multi-source corpus knowledge; In the case where the preprocessed multi-source corpus data is hierarchical table data, knowledge extraction is performed on the preprocessed multi-source corpus data based on the data structure of the preprocessed multi-source corpus data, including: converting the hierarchical table data into a two-dimensional data table, wherein each row of data in the two-dimensional data table includes: an index, a number and data content; traversing the two-dimensional data table, and determining the node position corresponding to the row of data and the relationship with other nodes according to the number of each row of data; associating the data with parent-child relationship in the two-dimensional data table through the index, and recording the parent node and child node of each row of data to generate a data node relationship table; converting each row of data in the data node relationship table into data in JSON format respectively; Determine the node position corresponding to the row of data and the relationship with other nodes according to the number of each row of data, including: setting the name of the two-dimensional data table to the root node; converting the number included in each row of data in the two-dimensional data table into a character string respectively, if the character string does not contain the target symbol, determine the row of data corresponding to the character string as the first-level parent node, and record the root node as the parent node of the first-level parent node; if the character string contains the target symbol, determine the row of data corresponding to the character string as a child node, wherein the parent node of the child node is the row of data corresponding to the number of the row of data corresponding to the child node minus the lowest level; if the number included in the character string is the target number, determine the row of data corresponding to the target number as the data content of the parent node or the child node; Converting the multi-source corpus knowledge into question-answer pairs; The question and answer pairs are deduplicated and conflicts are resolved to generate a large model question and answer knowledge base.

2. The method according to claim 1, characterized in that In the case where the preprocessed multi-source corpus data is text data, knowledge extraction is performed on the preprocessed multi-source corpus data based on the data structure of the preprocessed multi-source corpus data, including: Deleting useless symbols in the text data, wherein the useless symbols include at least blank characters and space characters; If the length of the text data after deleting the useless symbols exceeds a first preset length, segmenting the text data; The big model is used to extract knowledge from the text data based on the customized big model knowledge extraction rules.

3. The method according to claim 1, characterized in that Before converting the hierarchical table data into a two-dimensional data table, the method further includes: If the length of the text content in the hierarchical table data exceeds a second preset length, a preset model is used to simplify and refine the text content in the hierarchical table data.

4. The method according to claim 1, characterized in that: In the case where the preprocessed multi-source corpus data is non-hierarchical table data, knowledge extraction is performed on the preprocessed multi-source corpus data based on the data structure of the preprocessed multi-source corpus data, including: If the length of the text content in the non-hierarchical table data exceeds a third preset length, a preset model is used to simplify and refine the text content in the non-hierarchical table data.

5. The method according to claim 1, characterized in that If the multi-source corpus knowledge is non-hierarchical structured corpus knowledge, converting the multi-source corpus knowledge into question-answer pairs includes: Setting a question-answer pair template according to the source document content or source table content of the multi-source corpus knowledge; Collect multi-source corpus knowledge using the question-answer pair template based on the name of the source document or source table; The collected multi-source corpus knowledge is mapped into the question-answer pair template to generate the question-answer pair.

6. The method according to claim 1, characterized in that If the multi-source corpus knowledge is hierarchical structured corpus knowledge, converting the multi-source corpus knowledge into question-answer pairs includes: A multi-round question-answer pair template is set based on the node relationship and node data content in the hierarchical structure corpus knowledge, wherein the multi-round question-answer pair template includes a plurality of hierarchical question-answer pair templates, and each hierarchical question-answer pair template includes a question template and an answer template; Traversing from the root node of the hierarchical structure corpus knowledge, mapping each node to the question template of the hierarchical question-answer pair template where each node is located, searching the child nodes of each node from the data in the JSON format to generate the answer required in the answer template; A question-answer pair index is set for each generated question-answer pair, and the index of its previous question-answer pair is recorded to maintain the logical relationship between multiple rounds of question-answer pairs.

7. The method according to claim 1, characterized in that Before removing duplication and resolving conflicts of the question-answer pairs, the method further includes: The question, answer and document name or table name of each question-answer pair are respectively vectorized to generate a vector triple, wherein the vector triple includes: a question semantic vector, an answer semantic vector and a document name vector or a table name vector; Performing vector matching on the document name vectors or table name vectors of the question-answer pair, and if the similarity between two document name vectors or table name vectors is greater than a preset similarity threshold, determining that the two documents or tables corresponding to the two document name vectors or table name vectors are related; If the similarity between two document name vectors or table name vectors is lower than the preset similarity threshold, it is determined that the two documents or tables corresponding to the two document name vectors or table name vectors are unrelated.

8. The method according to claim 7, characterized in that After performing vector matching on the document name vector or the table name vector of the question-answer pair, the method further includes: Simultaneously performing a matching search on the question semantic vector and the answer semantic vector in a question-answer pair in which the document or table is a related document or a related table; When the similarities of the question semantic vector and the answer semantic vector of two question-answer pairs are both higher than the semantic overlap threshold, the two question-answer pairs are determined to be overlapped question-answer pairs; When the similarity of the question semantic vectors of two question-answer pairs is higher than a semantic overlap threshold, and the similarity of the answer semantic vectors is lower than a semantic conflict threshold, the two question-answer pairs are determined to be conflicting question-answer pairs.

9. The method according to claim 8, characterized in that Deduplication and conflict resolution are performed on the question-answer pairs, including: Retain the question-answer pairs with high answer coverage among the overlapping question-answer pairs; The question-answer pairs with high credibility of the knowledge source documents among the conflicting question-answer pairs are retained.

10. An electronic device comprising: A processor, and a memory storing a program, wherein the program comprises instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 9.

11. A non-transitory machine-readable medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Question-answer-corpus construction method based on semantic analysis of neural network

    CN108345640A

  • Response method, device and equipment and computer readable storage medium

    CN114416950A

  • Data classification and grading field knowledge base construction method based on information extraction

    CN115292450A