A question and answer knowledge base, a method for establishing the same, and an intelligent question and answer method and system

By extracting the semantic structure of the text and keyword types to generate preprocessed data Yc, and combining it with question classification templates and answer matching templates, the problems of low data processing efficiency and insufficient accuracy in traditional question-and-answer systems are solved, thus realizing an efficient and intelligent question-and-answer system.

CN119848231BActive Publication Date: 2025-10-24NORTHWEST INST OF ECO ENVIRONMENT & RESOURCES CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411907913.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-10-24
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

Traditional question-and-answer systems suffer from low data processing efficiency, low accuracy, and a lack of effective data organization and storage mechanisms, making it difficult to provide comprehensive and in-depth answers and meet complex user needs.

Method used

The source data processing module extracts the semantic structure and keyword types of the text to generate preprocessed data Yc. The data organization module groups and stores the data chain. By combining question classification templates and answer matching templates, efficient matching, sorting, and secondary storage are achieved. Intelligent question answering methods are used for data processing and display.

Benefits of technology

It improves the efficiency and accuracy of data processing, provides more comprehensive and logical answers, reduces system resource consumption, and improves user experience and data reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119848231B_ABST
    Figure CN119848231B_ABST
Patent Text Reader

Abstract

The application relates to a question and answer knowledge base and a method for establishing the same, and an intelligent question and answer method and system, comprising a source data acquisition unit, which is used for acquiring data in a computer system and marking the data as source data; a feature extraction and classification unit, which extracts feature data such as text semantic structure, keyword type and information format of the source data to classify the source data, obtains distinguished data QB and marks the distinguished data, and can analyze associated data in the distinguished data QB to generate pretreatment data Yc. The application specifically relates to the technical field of question and answer knowledge base technology, and through a source data processing module, the application can accurately extract feature data such as text semantic structure, keyword type and information format of the source data, thereby realizing effective classification of the source data, obtaining distinguished data QB and marking the distinguished data. This is helpful for quickly positioning and screening data related to question and answer, improves the efficiency and accuracy of data processing, and lays a solid foundation for subsequent question and answer knowledge construction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of question and answer knowledge base, and particularly relates to a question and answer knowledge base, a method for establishing the same, and an intelligent question and answer method and system. BACKGROUND

[0002] With the rapid development of information technology, a large amount of data resources have been accumulated in computer systems. In various application scenarios such as customer service, intelligent assistants, knowledge retrieval, etc., how to efficiently utilize these data to accurately answer user questions has become a key challenge. Traditional question and answer systems often face problems such as low data processing efficiency, low question and answer accuracy, and lack of effective data organization and storage mechanisms. For example, simple keyword matching-based question and answer systems are difficult to understand the semantic connotation of user questions, and are prone to misjudgment or unable to provide accurate answers. Moreover, in the face of large-scale data, without reasonable classification, integration and storage strategies, it will lead to data query and invocation delay, affecting the overall performance of the system. In addition, the lack of deep mining and utilization of data association relationships makes it difficult for the question and answer system to provide comprehensive and in-depth answers, and cannot meet the increasingly complex needs of users. Therefore, it is particularly necessary to develop an efficient, intelligent and perfect data processing flow question and answer knowledge base. SUMMARY

[0003] The present application provides a question and answer knowledge base, a method for establishing the same, and an intelligent question and answer method and system, which solves the problems in the prior art.

[0004] The object of the present application can be achieved by the following technical solutions:

[0005] A question and answer knowledge base, comprising:

[0006] A source data processing module configured to collect data in a computer system and label it as source data, extract text semantic structure, keyword type and information format and other feature data of the source data to classify the source data, obtain differentiated data QB and mark it, and analyze associated data in the differentiated data QB to generate preprocessed data Yc, the associated data being text data with the same semantic theme or related keywords;

[0007] The data organization module integrates the preprocessed data Yc obtained by different classifications into a preprocessed data bundle containing multiple groups of question and answer knowledge related databases, wherein the database is composed of multiple length data chains covering question text, answer text and knowledge association label, and the module also performs grouping processing according to the data chain type, and transmits the same type data chain to the storage unit by means of the data recognition channel, the storage unit is provided with question classification templates and answer matching templates in advance, which are used for matching, sorting, re-matching and secondary storage of the data chain, and the data chain after matching is sequentially calculated, and the result is stored in the result template, if the same type of question and answer knowledge data group is encountered in subsequent calculation, the calculation is skipped and the result template data is read, and a matching template is generated and stored in the memory storage unit for calling;

[0008] The data storage and display module includes a storage module and a data display terminal, the storage module is used for storing and recording analysis calculation result template data and matching with the result template, and the data display terminal is used for displaying the question and answer data result processed by the storage module and the invalid question and answer data transmitted by the invalid data module, and the invalid data module is used for processing the input invalid question and answer data.

[0009] The application also discloses a method for establishing a question and answer knowledge base, which specifically comprises the following steps:

[0010] The data information obtained by the computer is collected and marked as source data, the feature data of the text semantic structure, keyword type and information format of the source data is extracted, and the source data is classified according to the feature data to obtain the difference data QB and is identified;

[0011] The associated data in the difference data QB, i.e. the text data of the same semantic theme or related keywords, is analyzed and processed to generate the preprocessed data Yc;

[0012] The preprocessed data Yc obtained by different classifications is bundled to form a preprocessed data bundle containing multiple groups of question and answer knowledge related databases, and the database is composed of multiple length data chains, the data chain type includes question text, answer text and knowledge association label, and the same type data chain is transmitted to the storage unit according to the data chain type and by means of the data recognition channel;

[0013] The storage unit matches the data chain by using the pre-set template, sorts and stores the data chain according to the number of question and answer knowledge or the length of the text and other factors from large to small, and then more accurately matches the sorted data chain according to the keyword range and semantic depth and feeds back to the storage unit for secondary storage;

[0014] The data chain which has completed matching is sequentially calculated from the smallest capacity, the calculation mode includes integration and correlation analysis of the question and answer knowledge, the calculation result is stored in the result template, if the same type of question and answer knowledge data group is encountered in subsequent calculation, the calculation is skipped and the result template data is read, and a matching template is generated and stored in the memory storage unit.

[0015] The application further discloses an intelligent question and answer method of a question and answer knowledge base, and specifically comprises the following steps.

[0016] Data acquisition operation: the computer acquires question and answer related data information;

[0017] Data processing operation: the data is transmitted to a data processing center, a data grouping unit of the data processing center groups the data according to the data chain type, a data sorting unit sorts the grouped data chain according to the capacity size, a data analysis and calculation unit calculates the data chain which has completed sorting and matching from the smallest capacity, the calculation mode includes integration and correlation analysis of the question and answer knowledge, if the same type of question and answer knowledge data group is encountered in subsequent calculation, the calculation is skipped and the result template data is read, and a matching template is stored in a memory storage unit for calling;

[0018] Result display operation: the data result is presented to an operator through a data display terminal, a storage module stores and records the data result and matches and checks the result template, an invalid data module processes invalid question and answer data and transmits the invalid question and answer data to the data display terminal for observation by the operator.

[0019] The application further discloses a system for establishing a question and answer knowledge base, comprising:

[0020] A source data acquisition unit is used for acquiring data in a computer system and marking the data as source data;

[0021] A feature extraction and classification unit is used for extracting feature data such as text semantic structure, keyword type and information format of the source data to classify the source data, obtaining and marking difference data QB, and being capable of analyzing associated data in the difference data QB to generate pretreatment data Yc;

[0022] A data integration and transmission unit is used for integrating the pretreatment data Yc obtained through different classification into a pretreatment data bundle package containing multiple groups of question and answer knowledge related databases;

[0023] A data chain processing unit, a storage unit is provided with a question classification template and an answer matching template, the unit is used for matching, sorting, re-matching and secondary storage of the data chain, and sequentially calculating the data chain which has completed matching, starting from the smallest capacity data chain, storing the result in a result template, if the same type of question and answer knowledge data group is encountered in subsequent calculation, the calculation is skipped and the result template data is read, and a matching template is generated and stored in a memory storage unit;

[0024] The storage and display unit includes a storage module and a data display terminal. The storage module is used to store record analysis calculation result template data and match with the result template. The data display terminal is used to display the question and answer data result processed by the storage module and invalid question and answer data transmitted by the invalid data module. The invalid data module is used to process the input invalid question and answer data.

[0025] The beneficial effects of the present application are as follows: through the source data processing module, the text semantic structure, keyword type, information format and other feature data of the source data can be accurately extracted, thereby realizing effective classification of the source data, obtaining the difference data QB and performing identification. This helps to quickly locate and filter the data related to the question and answer, improves the efficiency and accuracy of data processing, and lays a solid foundation for subsequent question and answer knowledge construction; the associated data in the difference data QB is analyzed in depth to generate the preprocessed data Yc, which can mine the internal relationship between the text data with the same semantic theme or related keywords. Such association analysis enables the question and answer knowledge base to provide more comprehensive and more logical answers, avoids simple fragmented information reply, and improves the quality and depth of user knowledge acquisition.

[0026] The data organization module integrates the preprocessed data Yc obtained by different classification into a preprocessed data bundle containing multiple groups of question and answer knowledge related databases, and performs grouping processing and intelligent transmission according to the data chain type. With the pre-set question classification template and answer matching template, efficient matching, sorting, re-matching and secondary storage of the data chain are realized. At the same time, through the unique calculation and result storage mechanism, such as starting calculation from the data chain with the smallest capacity and using the result template and memory storage unit, the repeated calculation is greatly reduced, the storage and retrieval efficiency is improved, and the system resource consumption is reduced.

[0027] In the intelligent question and answer method, the units of the data processing center work cooperatively, which can group, sort and analyze the question and answer data in detail, ensuring that answers can be quickly and accurately provided when facing different types and complex degrees of questions. The cooperation of the data display terminal and the storage module can not only clearly display the question and answer data result, but also effectively store and check the result, ensuring the reliability and traceability of the data. The setting of the invalid data module further improves the stability and user experience of the system, which can process invalid question and answer data in time and avoid interference with the normal question and answer process. BRIEF DESCRIPTION OF DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creating any creative labor.

[0029] Figure 1 Figure 1 is a flowchart of the present application. DETAILED DESCRIPTION

[0030] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.

[0031] Please refer to Figure 1 The present application works around the construction of the question and answer knowledge base and the intelligent question and answer process, and specifically includes the following embodiments.

[0032] Embodiment 1

[0033] The source data acquisition unit is connected with various data interfaces of the computer system, such as database connection interface, file reading interface, network data receiving interface, etc., to obtain the data in the computer system. These data sources are extensive, which can be system logs, user operation records, document materials, webpage information, etc. After the data is collected, it is immediately given a unique identification mark as source data for subsequent processing. For example, when collecting the data of the internal computer system of an enterprise, the business operation record data such as order processing data and customer information data is obtained from the database of the office software used by the employees, and the similar identification such as "source data-001", "source data-002", etc. is added to each record.

[0034] The feature extraction and classification unit conducts in-depth analysis on the labeled source data. First, natural language processing techniques and data format parsing algorithms are used to extract the text semantic structure of the source data. For example, for a text describing product usage instructions, the sentence structure, subject-predicate-object relationship, etc. are analyzed to understand the core semantics of the text. Then, key word types are identified, distinguishing between professional terms, general vocabulary, functional description vocabulary, and other types of key words. Based on these extracted feature data, pre-set classification rules and machine learning classification models (such as classification models trained based on decision trees, neural networks, etc.) are used to classify the source data into different categories, obtaining differentiated data QB, and adding classification identifiers to each QB data, such as "product knowledge class QB-01", "operation process class QB-02", etc. For associated data in the differentiated data QB, semantic similarity calculation algorithms (such as cosine similarity algorithm, etc.) are used to find text data with the same semantic theme or related keywords. For example, in the product knowledge class QB, all different description data about a certain product model are associated. Then the associated data is integrated and preprocessed to generate preprocessed data Yc, such as organizing related data into a unified data structure, extracting key information and performing standardization processing.

[0035] Embodiment 2

[0036] The data integration and transmission unit integrates the preprocessed data Yc obtained from different classifications according to certain logic and structure. A preprocessed data bundle containing multiple groups of question and answer knowledge related databases is created. The database is composed of multiple length data chains covering question text, answer text, and knowledge association labels. According to the data chain type, group processing is performed, and question text data chains, answer text data chains, and knowledge association label data chains are classified respectively. Then, with the help of data recognition channels, the same type of data chain is transmitted to the storage unit through the identification information of the data chain and the channel rules. For example, all question text data chains are transmitted to the corresponding area of the storage unit through a special question text data recognition channel.

[0037] The storage unit is pre-configured with question classification templates and answer matching templates. Upon receiving a data link, an initial matching process is first performed. For example, based on the question classification template, the question text data link is matched against existing question classifications to determine the problem domain to which it belongs. The data links are then sorted and stored from largest to smallest based on factors such as the amount of question-answer knowledge or text length. For example, question data links containing more knowledge points or longer text are prioritized for processing and retrieval. The storage unit is pre-configured with question classification templates and answer matching templates. Upon receiving a data link, an initial matching process is first performed. For example, based on the question classification template, the question text data link is matched against existing question classifications to determine the problem domain to which it belongs. The data links are then sorted and stored from largest to smallest based on factors such as the amount of question-answer knowledge or text length. For example, question data links containing more knowledge points or longer text are prioritized for processing and retrieval. The matched data links are then calculated sequentially, starting with the data link with the smallest capacity. The calculation method includes the integration of question-answer knowledge and correlation analysis. For example, for simple, smaller question data chains, direct answer matching and associated knowledge integration can be performed, such as associating relevant product manual snippets and operating instructions with the answer. The calculation results are stored in a result template. If a subsequent calculation encounters the same type of question-and-answer knowledge data set, the calculation is skipped and the result template data is read. A matching template is generated and stored in the memory storage unit for easy recall. For example, if a question-and-answer data set regarding the basic operation of a product has already been calculated, when a similar question is encountered again, the data in the result template can be directly read, improving processing efficiency.

[0038] Example 3

[0039] The storage module receives the analysis and calculation result template data from the data processing module and matches and verifies it with the result template. Data verification algorithms, such as hash verification and data structure comparison, are used to ensure the accuracy and integrity of the stored data. For example, when storing a question and answer result data, its data hash value is calculated and compared with the hash value of the corresponding data in the result template. If they are consistent, the storage is successful. Otherwise, an error prompt and data repair processing is performed. The stored data is backed up and organized regularly, and the data is stored in different storage media or storage areas based on factors such as the frequency of use and importance of the data. For example, frequently queried popular question and answer data is stored in the cache area, and historical data is backed up to a large-capacity disk array.

[0040] The data display terminal obtains the processed question and answer data results from the storage module and displays them to the operator in a friendly user interface. For example, the question and answer are displayed in a clear text format, a chart (such as a knowledge correlation diagram), and the like in the form of a web page interface, a desktop application interface, and the like. At the same time, the invalid data module processes the input invalid question and answer data. When receiving question and answer data that does not meet the system requirements or cannot be recognized, such as random code data, data unrelated to the knowledge base topic, and the like, the invalid data is marked and recorded, and is transmitted to the data display terminal for observation by the operator. The operator can adjust and optimize the data collection source or data processing flow according to the invalid data information.

[0041] Embodiment 4

[0042] According to the above implementation steps of the source data processing module, the data organization module, and the data storage and display module, the question and answer knowledge base is gradually constructed. First, the source data collection unit is started to collect a large amount of original data, and then a series of processes such as feature extraction and classification, data integration and transmission, storage unit processing, and the like are performed to continuously enrich and improve the question and answer knowledge data in the question and answer knowledge base. During the establishment process, the classification rules, matching templates, calculation algorithms, and the like can be adjusted according to actual needs and data characteristics to improve the quality and applicability of the knowledge base.

[0043] Data acquisition operation: the computer obtains question and answer related data information through various input methods (such as user input text, voice recognition input, and the like). For example, the user inputs the question "how to query the product order status?" in the input box of the intelligent customer service system.

[0044] Data processing operation: the data is transmitted to the data processing center, and the data grouping unit of the data processing center groups the data according to the data chain type, and identifies the question data as a question text data chain type. The data sorting unit sorts the grouped data chains by capacity size, and here, since it is a newly input question, there is no sorting comparison for the time being. The data analysis and calculation unit calculates the data chain after sorting and matching from the smallest capacity, and the calculation method includes integration and correlation analysis of question and answer knowledge. The question and answer knowledge data group related to the question is searched in the knowledge base, and if the same type of question and answer knowledge data group is found, the calculation is skipped and the result template data is read, such as a standard answer to query the product order status, which is directly read and returned to the user. The memory storage unit stores the matching template for subsequent calling, and records the related information of the current question and answer, such as the key words of the question, the asking time, and the like, for subsequent data analysis and knowledge base optimization.

[0045] Result display operation: the data result is presented to the operator through the data display terminal. For example, the intelligent customer service interface displays "you can log in to our official website, enter the order number and verification code on the order query page to query the order status". The storage module stores the data result record and matches the result template to check, ensuring the correct storage and traceability of the data. The invalid data module processes invalid question and answer data and transmits it to the data display terminal for the operator to observe. If the user inputs random code or a question unrelated to the product business, the corresponding prompt information will be displayed on the interface, such as "the question you entered cannot be recognized, please re-enter".

[0046] In the description of the specification, the description of the terms "one embodiment", "example", "specific example" and the like means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are contained in at least one embodiment or example of the present application. In the specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0047] The above is only an example and description of the concept of the present application, and those skilled in the art can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present claims, which shall belong to the protection scope of the present application.

Claims

1. A question and answer knowledge base, characterized by: Comprise: A source data processing module configured to collect data in a computer system and mark as source data, extract feature data such as text semantic structure, keyword type and information format of the source data to classify the source data, obtain distinguished data QB and identify, and can analyze the associated data in the distinguished data QB to generate pre-processing data Yc, the associated data is text data with the same semantic theme or related keywords; A data organization module integrates the pre-processing data Yc obtained by different classification into a pre-processing data bundle containing multiple groups of question and answer knowledge related databases, wherein the database is composed of multiple length data chains covering question text, answer text and knowledge association label, the module also groups according to data chain type, transmits data chains of the same type to the storage unit by means of data recognition channel, the storage unit is provided with question classification template and answer matching template in advance, which is used to match, sort, match again and store twice the data chain, and sequentially calculate the matched data chain, start from the smallest capacity data chain, store the result in the result template, if the same type of question and answer knowledge data group is encountered in subsequent calculation, skip the calculation and read the result template data, and generate a matching template stored in a memory storage unit for calling; A data storage and display module, including a storage module and a data display terminal, the storage module is used to store record analysis and calculation result template data and match with the result template, the data display terminal is used to display the question and answer data results processed by the storage module and the invalid question and answer data transmitted by the invalid data module, the invalid data module is used to process the input invalid question and answer data.

2. A method of building a question and answer knowledge base as claimed in claim 1, characterized in that: Specifically comprising the following steps: Collecting data information obtained by computer and marking as source data, extracting feature data such as text semantic structure, keyword type and information format of the source data, classifying the source data according to the feature data to obtain distinguished data QB and identification; Processing the associated data in the distinguished data QB, i.e. text data with the same semantic theme or related keywords, to generate pre-processing data Yc; Bundling the pre-processing data Yc obtained by different classification to form a pre-processing data bundle containing multiple groups of question and answer knowledge related databases, the database is composed of multiple length data chains, the data chain type includes question text, answer text and knowledge association label, grouping according to data chain type and transmitting data chains of the same type to the storage unit by means of data recognition channel; The storage unit matches the data chain by using the pre-set template, sorts and stores according to the number of question and answer knowledge or the length of text contained in the data chain, and more accurately matches the sorted data chain according to the keyword range and semantic depth and feeds back to the storage unit for secondary storage; Starting from the smallest capacity data chain, sequentially calculating the matched data chain, the calculation method includes integration and correlation analysis of question and answer knowledge, storing the calculation result in the result template, if the same type of question and answer knowledge data group is encountered in subsequent calculation, skip the calculation and read the result template data, and generate a matching template stored in a memory storage unit.

3. The method of intelligent question answering of a question-answer knowledge base as claimed in claim 1, wherein, Specifically comprising the following steps: Data acquisition operation: the computer acquires question and answer related data information; Data processing operation: the data is transmitted to the data processing center, the data grouping unit of the data processing center groups the data according to the data chain type, the data sorting unit sorts the grouped data chain according to the capacity size, the data analysis and calculation unit starts to calculate from the smallest capacity of the data chain after sorting and matching, and the calculation method includes the integration and correlation analysis of question and answer knowledge. If the same type of question and answer knowledge data group is encountered, the calculation is skipped and the result template data is read. The memory storage unit stores the matching template for calling; Result display operation: the data result is presented to the operator through the data display terminal, the storage module stores and records the data result and matches the result template, the invalid data module processes the invalid question and answer data and transmits it to the data display terminal for the operator to observe.

4. The system for establishing a question and answer knowledge base according to claim 1, characterized by: It includes: Source data acquisition unit, used for collecting data in computer system and marking as source data; Feature extraction and classification unit, which extracts the text semantic structure, keyword type and information format of source data to classify source data, obtains and identifies difference data QB, and can analyze the associated data in difference data QB to generate pre-processing data Yc; Data integration and transmission unit, which integrates the pre-processing data Yc obtained by different classification into a pre-processing data bundle containing multiple groups of question and answer knowledge related database; Data chain processing unit, the storage unit is provided with question classification template and answer matching template, which is used for matching, sorting, re-matching and secondary storage of data chain, and sequentially calculating the matched data chain. When calculating, start from the smallest capacity of the data chain, store the result in the result template, if the same type of question and answer knowledge data group is encountered during subsequent calculation, skip the calculation and read the result template data, and generate a matching template stored in the memory storage unit; Storage and display unit, including storage module and data display terminal, the storage module is used for storing and recording analysis calculation result template data and matching with result template, the data display terminal is used for displaying question and answer data result processed by the storage module and invalid question and answer data transmitted by the invalid data module, and the invalid data module is used for processing invalid question and answer data input.

Citation Information

Patent Citations

  • Intelligent question-answering system and method, and related equipment

    CN112579666A

  • Multi-engine intelligent question answering system for multi-type knowledge base

    CN115238101A