Knowledge base management system based on NLP and machine learning

By adopting NLP and machine learning technology in the knowledge base management system, the problems of low automation degree and poor data processing efficiency of traditional knowledge base management methods are solved, efficient processing and intelligent recommendation of massive knowledge data are achieved, and the timeliness and personalized management effect of the knowledge base is improved.

CN120104794AActive Publication Date: 2025-06-06BEIJING JINGNENG TENDERING & COLLECTIVE PROCUREMENT CENT CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510176234.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-06
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

Traditional knowledge base management methods have low degree of automation, poor data processing efficiency, and excessive manual intervention, which makes them unable to efficiently process large amounts of data and changing information needs. The existing knowledge base management system has not been fully met in terms of dynamic updates of information, intelligent recommendations, sentiment analysis, etc.

Method used

A knowledge base management system based on NLP and machine learning is adopted, including a knowledge acquisition module, a knowledge processing and classification module, a knowledge storage and index construction module, and a knowledge update and maintenance module. Through text preprocessing, text vectorization and classification, knowledge index construction, and dynamic update based on user search behavior, efficient management and intelligent recommendation of the knowledge base are achieved.

Benefits of technology

It realizes the automated processing of massive raw knowledge data, improving data processing efficiency and accuracy; by quickly storing and retrieving classified data, improving storage and index search efficiency; through real-time updates and personalized recommendations, the timeliness and personalized management of the knowledge base is ensured, and users' query efficiency and knowledge acquisition accuracy are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104794A_ABST
    Figure CN120104794A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of knowledge base management, in particular to a knowledge base management system based on NLP and machine learning, and the system comprises a knowledge collection module which is used for receiving original knowledge data, executing text preprocessing, and generating preprocessed knowledge data; the knowledge processing and classification module is used for receiving the preprocessed knowledge data, executing text vectorization and classification and generating classified knowledge data; the knowledge storage and index construction module is used for storing the classified knowledge data and constructing a knowledge index based on the classified knowledge data; and the knowledge updating and maintaining module is used for screening the original knowledge data stored in the historical time period based on the search word input by the user and the historical search frequency, determining the optimal original knowledge data and dynamically adjusting the knowledge index and recommendation strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge base management, and in particular to a knowledge base management system based on NLP and machine learning. Background Art

[0002] Traditional knowledge base management methods face problems such as low automation, poor data processing efficiency, and excessive human intervention. Specifically, the collection, classification, and storage of knowledge usually rely on manual input and traditional rules, which makes it impossible to efficiently process large amounts of data and ever-changing information needs.

[0003] Although existing knowledge base management systems are capable of basic data storage and retrieval, they have not yet fully met the needs for dynamic information updates, intelligent recommendations, sentiment analysis, etc. Summary of the invention

[0004] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a knowledge base management system based on NLP and machine learning.

[0005] In order to achieve the above object, the main technical solutions adopted by the present invention include:

[0006] An embodiment of the present invention provides a knowledge base management system based on NLP and machine learning, the system comprising:

[0007] The knowledge acquisition module is used to receive original knowledge data and perform text preprocessing to generate preprocessed knowledge data;

[0008] The knowledge processing and classification module is used to receive pre-processed knowledge data, perform text vectorization and classification, and generate classified knowledge data;

[0009] A knowledge storage and index building module, used for storing classified knowledge data and building a knowledge index based on the classified knowledge data;

[0010] The knowledge updating and maintenance module is used to filter the original knowledge data stored in the historical time period based on the search terms input by the user and their historical search times, and determine the optimal original knowledge data.

[0011] Preferably, the knowledge acquisition module includes:

[0012] A data input unit, used to receive original knowledge data, wherein the original knowledge data includes documents uploaded by a user;

[0013] The data preprocessing unit is used to perform preprocessing on the original knowledge data, wherein the preprocessing includes:

[0014] Remove HTML tags, advertisements, garbled characters, and repeated text to obtain cleaned text data;

[0015] For the cleaned text data, the BERT attention mechanism combined with the TF-IDF algorithm is used to calculate the word weights to extract the keyword set;

[0016] Use the LSTM-CRF model to extract sentiment features from the cleaned text data and generate sentiment score data;

[0017] Process the cleaned text data based on the BART pre-trained model, generate text summaries and output summary data;

[0018] The data output unit is used to output the preprocessed knowledge data and provide it to the knowledge processing and classification module.

[0019] Preferably, for the cleaned text data, the BERT attention mechanism is combined with the TF-IDF algorithm to calculate the word weights to extract the keyword set, specifically including:

[0020] Formula (1) is used to calculate the value of each term t in the cleaned text data. i TF-IDF weight; the formula (1) is:

[0021]

[0022] Among them, TF·IDF(t i ) is the term t i TF-IDF weight of

[0023] f i For term t i The number of occurrences in the cleaned text data;

[0024] N is the total number of cleaned text data in the corpus; wherein the corpus includes all the original knowledge data received in the historical time period;

[0025] n i To contain the term t i The number of cleaned text data;

[0026] TTF represents the total number of occurrences of all terms in the cleaned text data;

[0027] Based on the pre-trained BERT model, the cleaned text data in the pre-processed text data is processed by the Transformer structure to obtain the attention matrix, and the formula (2) is used to calculate each term t i The average attention score across all layers of the BERT model;

[0028] The formula (2) is:

[0029]

[0030] Among them, S i For term t i The average attention score across all layers of the BERT model;

[0031] A h,i is the hth layer of the BERT model for term t i Attention score;

[0032] H is the total number of layers of the BERT model;

[0033] Formula (3) is used to calculate each term t i The position information P in the cleaned text data i , normalized to get term t i Position weight;

[0034] The formula (3) is:

[0035]

[0036] Among them, R i For term t i Position weight; P i For term t i The sum of the rankings in the cleaned text data; L is the total number of terms in the cleaned text data;

[0037] Based on the term t i TF-IDF weight of term t i Average attention scores and term t across all layers of the BERT model i Position weight, calculated using formula (4) for each term t i The final keyword weight W i ;

[0038] The formula (4) is:

[0039] W i =γ 1 ×TF·IDF(t i )+γ 2 ×S i +γ 3 ×R i ;

[0040] Among them, γ 1 is the first weight adjustment parameter; 2 is the second weight adjustment parameter; 3 is the third weight adjustment parameter; 1 +γ 2 +γ 3=1;

[0041] According to the calculated comprehensive keyword weight W i , select the first K terms with the highest keyword weights to form the final keyword set, and output the keyword set.

[0042] Preferably, the knowledge processing and classification module includes:

[0043] A text vectorization unit, used for receiving preprocessed knowledge data and converting the cleaned text data into text vector data;

[0044] A classification unit, used to perform text classification based on the text vector data and generate classification knowledge data;

[0045] The data output unit is used to output the classified knowledge data and transmit it to the knowledge storage and index construction module.

[0046] Preferably, the classification unit is used to receive text vector data and perform text classification to obtain classified knowledge data, specifically including:

[0047] The classification unit uses a preset support vector machine model to perform text classification on the text vector data, and after classification, assigns corresponding classification knowledge data and the keyword set to each text data;

[0048] The preset support vector machine model is obtained by pre-training with text vector data of the historical stage marked with classification knowledge data labels;

[0049] The classification knowledge data includes classification labels and semantic vectors.

[0050] Preferably, the knowledge storage and index building module includes:

[0051] A storage unit, used for storing classified knowledge data in a specified storage method;

[0052] The specified storage methods include: MongoDB and PostgreSQL;

[0053] An index building unit, used for building a knowledge index based on the classified knowledge data;

[0054] The knowledge index includes an inverted index and an HNSW index;

[0055] The data output unit is used to output the knowledge index and transmit it to the knowledge update and maintenance module.

[0056] Preferably, the knowledge updating and maintenance module comprises: an event-driven updating unit, which is used to calculate the correlation score S between any original knowledge data stored in the historical time period and the search term q according to the search term q input by the user and the historical search times of the search term q;

[0057] According to the relevance score S, the best original knowledge data is selected from the original knowledge data received in the historical time period;

[0058] The optimal original knowledge data is the original knowledge data corresponding to the highest relevance score S.

[0059] Preferably, the correlation score S between any original knowledge data stored in the historical time period and the search term q is calculated using formula (5) based on the search term q input by the user and the historical search times of the search term q;

[0060] The formula (5) is:

[0061]

[0062] S is the relevance score between the search term q and any original knowledge data;

[0063] f q is the number of times the search term q appears in any original knowledge data;

[0064] M is the total number of original knowledge data stored in the historical time period;

[0065] n q is the number of original knowledge data containing the search term q stored in the historical time period;

[0066] T q is the total number of searches for the search term q in the historical time period;

[0067] T is the total number of searches for all keywords;

[0068] α is a preset balance factor.

[0069] Preferably, the format of the original knowledge data is any one of Word, PDF, TXT, and HTML.

[0070] Preferably, 0<α<1.

[0071] The beneficial effects of the present invention are as follows: a knowledge base management system based on NLP and machine learning of the present invention, due to the use of a knowledge acquisition module to perform text preprocessing on original knowledge data, including the steps of removing irrelevant information, keyword extraction, sentiment analysis, summary generation, etc., combined with machine learning for text vectorization and classification, ensures efficient data processing and classification, and compared with the prior art, it can automatically process massive amounts of original knowledge data and improve data processing efficiency and accuracy; through the knowledge storage and index construction module, it can quickly store classified data and generate accurate knowledge indexes, and compared with traditional manual processing methods, it achieves the effect of greatly improving storage efficiency and index retrieval efficiency; and the knowledge update and maintenance module can update the knowledge base content in real time and optimize the recommendation strategy by screening and dynamically adjusting based on user search behavior and historical data, ensuring the timeliness and personalized recommendation of the knowledge base, achieving a more accurate and personalized knowledge management effect, and improving the user's query efficiency and the accuracy of knowledge acquisition. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 A schematic diagram of the structure of a knowledge base management system based on NLP and machine learning of the present invention;

[0073] Figure 2 This is a flow chart of a knowledge base management method based on NLP and machine learning provided in Example 4 of the present invention. DETAILED DESCRIPTION

[0074] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation modes in conjunction with the accompanying drawings.

[0075] In order to better understand the above technical solution, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0076] Embodiment 1

[0077] Figure 1 As shown, this embodiment provides a structural block diagram of a knowledge base management system based on NLP and machine learning. The system includes a knowledge acquisition module, a knowledge processing and classification module, a knowledge storage and index construction module, and a knowledge update and maintenance module. Each module plays a different role in the system, and together realizes efficient and intelligent knowledge base management.

[0078] Knowledge acquisition module, including:

[0079] Data input unit: This unit receives raw knowledge data from different sources, which may include documents uploaded by users (such as Word, PDF, TXT format, etc.), data from external databases, text data from social media, and real-time streaming data, etc. Before entering the system, the raw data is directly input into the data input unit without any processing.

[0080] Data preprocessing unit: After the data input unit receives the original knowledge data, the data preprocessing unit further processes the original data. First, by removing irrelevant information such as HTML tags, advertisements, garbled characters, and repeated text, clean and usable cleaned text data is obtained. Then, for the cleaned text data, the BERT attention mechanism is combined with the TF-IDF algorithm to calculate the weights of words, and the keyword set in the text is extracted through these weights. Next, the LSTM-CRF model is used to extract the sentiment features in the text and generate sentiment score data for each document. In addition, based on the BART pre-trained model, the cleaned text data is processed to generate text summaries, and these summaries are used as part of the output data for use by subsequent modules.

[0081] Data output unit: The processed data, including keyword sets, sentiment score data and summary data, will be output through the data output unit and provided to the knowledge processing and classification module for further processing.

[0082] Knowledge processing and classification module, including:

[0083] After receiving the preprocessed knowledge data from the knowledge acquisition module, the knowledge processing and classification module performs text vectorization on the data. Specifically, first, the cleaned text data is vectorized for text representation using a pre-trained model based on BERT, converting the text into a high-dimensional vector representation. Next, the text vectorized data is classified using a machine learning algorithm (such as a support vector machine (SVM) or XGBoost classifier) ​​to generate classified knowledge data. This classified knowledge data contains different categories and labels divided according to the text content and features. The classified knowledge data is further transmitted to the knowledge storage and index construction module.

[0084] The knowledge storage and index construction module is mainly used to store the classified knowledge data and construct the corresponding knowledge index. The classified knowledge data is stored by using a distributed database (such as MongoDB, Cassandra, etc.). In order to improve the efficiency of subsequent knowledge retrieval, the inverted index and hierarchical navigation small world graph (HNSW) are used for index optimization to ensure that retrieval and query can be performed quickly and accurately in a large-scale data environment. The constructed knowledge index will further provide data support for the knowledge update and maintenance module.

[0085] The knowledge update and maintenance module filters the original knowledge data stored in the system based on historical data according to the search terms entered by the user and their historical search times, so as to filter out the best original knowledge data within the historical time period. The screening process is derived by calculating user interests and search relevance, ensuring that users obtain the most relevant and high-quality knowledge content when searching. The filtered data is transmitted to the knowledge index for update, and the recommendation strategy is optimized in combination with user feedback to ensure the accuracy and timeliness of the system's recommendations.

[0086] Through the above-mentioned system and method of the embodiment of the present invention, knowledge data can be processed and managed efficiently and intelligently, and real-time updating, dynamic adjustment and personalized recommendation can be achieved, which significantly improves the efficiency and accuracy of knowledge base management, and especially provides strong support for data management and retrieval in large-scale knowledge base environments.

[0087] Embodiment 2

[0088] Suppose a company wants to build a customer service knowledge base based on natural language processing, which can automatically push corresponding solutions based on the questions raised by customers. The system will execute the following steps:

[0089] The company collected customer service-related document data from multiple departments, including customer problem records, technical support documents, and training materials. The data input unit inputs these documents into the system as raw knowledge data. The data preprocessing unit removes HTML tags and duplicate content from the documents, extracts keywords through the BERT attention mechanism combined with the TF-IDF algorithm, and extracts sentiment features through the LSTM-CRF model. The sentiment score data and summary information generated in the end will be used as data output and enter the knowledge processing and classification module.

[0090] In the knowledge processing and classification module, the system will vectorize the text data and classify it through the SVM classifier.

[0091] The classified knowledge data is stored in a distributed database and optimized using inverted index and HNSW index technology. Through index optimization, the system can quickly match user queries and return relevant solutions. The constructed knowledge index also supports efficient personalized recommendations in future retrieval.

[0092] As customer questions are constantly updated and feedback is received, the system will dynamically update the knowledge base based on the user's search history data (such as the number of user clicks) to ensure that customers can obtain the latest and most relevant solutions. At the same time, the recommendation strategy is adjusted based on user feedback, so that the system can provide more personalized recommendations and further improve the user experience.

[0093] Through these operations, the company can use NLP and machine learning technologies to significantly improve the efficiency and accuracy of customer service, while ensuring that the knowledge base is always up to date and can provide personalized knowledge recommendations based on actual needs.

[0094] Embodiment 3

[0095] See also Figure 1 This embodiment provides a knowledge base management system based on NLP and machine learning, the system comprising:

[0096] The knowledge acquisition module is used to receive original knowledge data and perform text preprocessing to generate preprocessed knowledge data;

[0097] The knowledge acquisition module in this embodiment includes:

[0098] A data input unit, used to receive original knowledge data, wherein the original knowledge data includes documents uploaded by a user;

[0099] The data preprocessing unit is used to perform preprocessing on the original knowledge data, wherein the preprocessing includes:

[0100] Remove HTML tags, advertisements, garbled characters, and repeated text to obtain cleaned text data;

[0101] For the cleaned text data, the BERT attention mechanism combined with the TF-IDF algorithm is used to calculate the word weights to extract the keyword set;

[0102] Use the LSTM-CRF model to extract sentiment features from the cleaned text data and generate sentiment score data;

[0103] Process the cleaned text data based on the BART pre-trained model, generate text summaries and output summary data;

[0104] The data output unit is used to output the preprocessed knowledge data and provide it to the knowledge processing and classification module.

[0105] The knowledge processing and classification module is used to receive pre-processed knowledge data, perform text vectorization and classification, and generate classified knowledge data;

[0106] A knowledge storage and index building module, used for storing classified knowledge data and building a knowledge index based on the classified knowledge data;

[0107] The knowledge updating and maintenance module is used to filter the original knowledge data stored in the historical time period based on the search terms input by the user and their historical search times, and determine the optimal original knowledge data.

[0108] In this embodiment, for the cleaned text data, the BERT attention mechanism is combined with the TF-IDF algorithm to calculate the word weights to extract the keyword set, specifically including:

[0109] Formula (1) is used to calculate the value of each term t in the cleaned text data. i TF-IDF weight; the formula (1) is:

[0110]

[0111] Among them, TF·IDF(t i ) is the term t i TF-IDF weight of

[0112] f i For term t i The number of occurrences in the cleaned text data directly affects the calculation of the TF value, can reflect the importance of the term in the document, provides the necessary data for calculating TF-IDF, and ensures the rationality of subsequent steps.

[0113] N is the total number of cleaned text data in the corpus; wherein the corpus includes all original knowledge data received within a historical period of time; by considering the number of all documents in the corpus, the rarity of terms in the entire data set can be effectively evaluated, thereby improving the ability to screen keywords.

[0114] n i To contain the term t i The number of cleaned text data; in this embodiment, n i The definition of the term t i The number of documents in which a term appears can measure the prevalence of the term, thus affecting the IDF value.

[0115] TTF represents the total number of occurrences of all terms in the cleaned text data. TFF can reflect the prevalence of a term in a document to a certain extent. It is one of the key parameters when calculating TF-IDF and ensures the rationality of term weight calculation.

[0116] In this embodiment, the weight of each term is calculated by TF-IDF. This helps to measure the importance of words in the text, especially to reduce the influence of common words, highlight key content, and help optimize subsequent text analysis and information retrieval. In this embodiment, the calculation results of TF-IDF are clarified to quantify the importance of each term in the text, providing reliable data support for subsequent steps (such as classification, indexing, etc.).

[0117] Based on the pre-trained BERT model, the cleaned text data in the pre-processed text data is processed by the Transformer structure to obtain the attention matrix, and the formula (2) is used to calculate each term t i The average attention score across all layers of the BERT model;

[0118] The formula (2) is:

[0119]

[0120] Among them, S i For term t i The average attention score across all layers of the BERT model;

[0121] A h,i is the hth layer of the BERT model for term t i Attention score;

[0122] H is the total number of layers of the BERT model;

[0123] In this embodiment, the attention mechanism of the BERT model can further enhance the ability to extract keywords and semantic features in the text, and use the average attention score to evaluate the importance of terms. This operation provides an effective basis for further screening and classification of keywords and optimizes the accuracy of keyword recognition. These steps achieve efficient extraction of keywords and features from text data by introducing a combination of technologies such as BERT, TF-IDF, and LSTM-CRF, which helps to improve the intelligence and accuracy of the knowledge base management system, especially in automatic classification, indexing, and recommendation. It has innovative effects.

[0124] Formula (3) is used to calculate each term t i The position information P in the cleaned text data i , normalized to get term t i Position weight;

[0125] The formula (3) is:

[0126]

[0127] Among them, R i For term t i Position weight; P i For term t i The sum of the rankings in the cleaned text data; L is the total number of terms in the cleaned text data;

[0128] The position information of the term in the cleaned text data is calculated by formula (3) and unified to obtain the position information weight of the term. By considering the position of the term in the text, the accuracy of keyword extraction can be improved. As an additional feature, position information helps to improve the depth of text understanding, thereby having higher efficiency in text analysis and information retrieval. This formula (3) clarifies how to calculate the position information weight of the term. The position of the term is standardized relative to the total number of words L in the text where it is located through the calculation formula of the position weight. The standardized method helps to balance the position effect of each term in the text, thereby ensuring that the contribution of the position to text analysis will not be biased due to different document lengths.

[0129] Based on the term t i TF-IDF weight of term t i Average attention scores and term t across all layers of the BERT model i Position weight, calculated using formula (4) for each term t i The final keyword weight W i ;

[0130] The formula (4) is:

[0131] W i =γ 1 ×TF·IDF(t i )+γ 2 ×S i +γ 3 ×R i ;

[0132] Among them, γ 1 is the first weight adjustment parameter; 2 is the second weight adjustment parameter; 3 is the third weight adjustment parameter; 1 +γ 2 +γ 3 =1;

[0133] According to the calculated comprehensive keyword weight W i , select the first K terms with the highest keyword weights to form the final keyword set, and output the keyword set.

[0134] In this embodiment, TF-IDF is used to calculate the basic weight of each term, and the attention weight, position weight and other information of each term in the BERT model are combined to obtain the final keyword weight. Through weighted fusion, this method can more accurately capture the most relevant keywords in the text. By using TF-IDF and BERT's attention, position and other information at the same time, the importance of each term can be more comprehensively measured to enhance the performance of the model. By combining the contextual information of the BERT model and the statistical information of TF-IDF, the truly important keywords in the text can be better identified, rather than relying on a single method. By adjusting the weighting parameter (γ 1 , γ 2 , γ 3 ), you can adjust the sensitivity of the model according to different needs and optimize the keyword extraction effect.

[0135] In this embodiment, the knowledge processing and classification module includes:

[0136] A text vectorization unit, used for receiving preprocessed knowledge data and converting the cleaned text data into text vector data;

[0137] A classification unit, used to perform text classification based on the text vector data and generate classification knowledge data;

[0138] The data output unit is used to output the classified knowledge data and transmit it to the knowledge storage and index construction module.

[0139] In this embodiment, it is assumed that the input is an article about artificial intelligence, such as "The Application of Transformer in NLP". The system will perform preprocessing steps such as word segmentation, stop word removal, and stemming on the article, and then convert it into a numerical vector (such as TF-IDF, Word2Vec, BERT embedding, etc.). It can be used for tasks such as news classification and automatic paper summarization. Based on the vectorized text data, use machine learning or deep learning (such as SVM, CNN, LSTM, Transformer) to classify the text. For example, the article is automatically classified into the "artificial intelligence" category instead of the "biotechnology" or "blockchain" category. The classified knowledge data can be stored in a database for user queries. For example, when a user searches for "natural language processing", the system can return all articles classified as "NLP" and make intelligent recommendations.

[0140] In this embodiment, the system can quickly process a large number of documents and improve classification efficiency through text vectorization. Combining NLP with machine learning can reduce the misjudgment of traditional keyword matching and improve the intelligence level of text classification.

[0141] Specifically, the classification unit is used to receive text vector data and perform text classification to obtain classified knowledge data, which specifically includes:

[0142] The classification unit uses a preset support vector machine model to perform text classification on the text vector data, and after classification, assigns corresponding classification knowledge data and the keyword set to each text data;

[0143] The preset support vector machine model is obtained by pre-training with text vector data of the historical stage marked with classification knowledge data labels;

[0144] The classification knowledge data includes classification labels and semantic vectors.

[0145] In this embodiment, by training models such as support vector machines, text categories can be identified more accurately than traditional keyword matching. Combined with NLP technology, the system can automatically classify and store texts, reducing the workload of manual sorting. Through semantic vector representation (such as BERT, Word2Vec), the system can understand the deep meaning of the text, not just simple keyword matching.

[0146] In this embodiment, the knowledge storage and index building module includes:

[0147] A storage unit, used for storing classified knowledge data in a specified storage method;

[0148] The specified storage methods include: MongoDB and PostgreSQL;

[0149] An index building unit, used for building a knowledge index based on the classified knowledge data;

[0150] The knowledge index includes an inverted index and an HNSW index;

[0151] The data output unit is used to output the knowledge index and transmit it to the knowledge update and maintenance module.

[0152] The detailed description is as follows: The storage unit is used to store classified knowledge data to ensure efficient management and call of data.

[0153] Specify storage method:

[0154] MongoDB: It is a document database based on NoSQL and is suitable for storing unstructured or semi-structured data, such as text and JSON format data. Its advantages are fast query speed, strong scalability, and support for complex queries and full-text search.

[0155] PostgreSQL: It is a SQL-based relational database suitable for storing structured data and supports ACID transaction processing. Its advantages are high data consistency, support for complex queries, full-text search, and extended functions (such as JSONB storage).

[0156] The index building unit builds an index based on classified knowledge data to improve query efficiency.

[0157] Knowledge index type:

[0158] Inverted Index: The principle is to map the keywords in the document to a list of documents containing these keywords, which is used to quickly query text data.

[0159] HNSW (Hierarchical Navigable Small World) index: The principle is an indexing method based on approximate nearest neighbor (ANN) search, which is suitable for high-dimensional data (such as vectors).

[0160] Applications: semantic search, recommendation system, knowledge graph query.

[0161] The data output unit is used to output the constructed knowledge index and transmit it to the knowledge update and maintenance module to ensure the accessibility and maintainability of the data.

[0162] In this embodiment, MongoDB and PostgreSQL are used to meet the requirements of unstructured text storage and structured data management, taking into account query efficiency and data consistency. Inverted index accelerates text keyword retrieval and is suitable for full-text search and knowledge retrieval. HNSW index supports high-dimensional vector search, improving the accuracy and speed of semantic search and similarity matching.

[0163] In this embodiment, the knowledge updating and maintenance module includes: an event-driven updating unit, which is used to calculate the correlation score S between any original knowledge data stored in the historical time period and the search term q according to the search term q input by the user and the historical search times of the search term q;

[0164] According to the relevance score S, the best original knowledge data is selected from the original knowledge data received in the historical time period;

[0165] The optimal original knowledge data is the original knowledge data corresponding to the highest relevance score S.

[0166] In this embodiment, the matching degree between the original knowledge data and the search term q is calculated by the relevance score S, which can ensure that the most relevant knowledge is returned first when the user queries, avoiding interference from irrelevant content. Using the number of historical searches as weights can make popular search results easier to retrieve, allowing users to find valuable information faster and reduce unnecessary query costs.

[0167] Specifically, the correlation score S between any original knowledge data stored in the historical time period and the search term q is calculated using the search term q input by the user and the historical search times of the search term q using formula (5);

[0168] The formula (5) is:

[0169]

[0170] S is the relevance score between the search term q and any original knowledge data; f q is the number of times the search term q appears in any original knowledge data; M is the total number of original knowledge data stored in the historical time period; n q is the amount of original knowledge data containing search term q stored in the historical period; T q is the total number of searches for the search term q in the historical period; T is the total number of searches for all keywords; α is a pre-set balance factor.

[0171] In this embodiment, the relevance score S is calculated by formula (5), which can quantify the matching degree between the search term q and the knowledge data, so that the system can preferentially return the most relevant knowledge data and improve the accuracy of the retrieval results. Formula (5) uses: q (The number of occurrences of search term q in the original knowledge data) to ensure that high-frequency matching data is given priority. q (the number of original knowledge data containing the search term q), avoiding deviations caused by a certain knowledge data containing q too many times, and enhancing generalization capabilities. α can adjust the impact of search popularity on relevance scores, avoiding some accidental high-frequency searches from affecting knowledge ranking, and enhancing system stability.

[0172] The format of the original knowledge data is any one of Word, PDF, TXT, and HTML. In this embodiment, multiple formats such as Word, PDF, TXT, and HTML are supported, so that the knowledge base can process text data from different sources, improving the scalability and applicability of the system. This compatibility ensures that the knowledge base can automatically parse and learn knowledge data in different formats, improving the intelligence of the system.

[0173] Among them, 0<α<1. The design of dynamic parameter α enables the knowledge base to maintain the stability of knowledge while adapting to the needs of knowledge updating, thereby improving intelligence and controllability.

[0174] Embodiment 4

[0175] See also Figure 2 In this embodiment, a knowledge base management method based on NLP and machine learning is also provided, including:

[0176] S1, receiving original knowledge data, and performing text preprocessing to generate preprocessed knowledge data;

[0177] S2, performing text vectorization and classification on the preprocessed knowledge data to generate classified knowledge data;

[0178] S3, storing classified knowledge data and building a knowledge index based on the classified knowledge data;

[0179] S4. Based on the search terms input by the user and their historical search times, the original knowledge data stored in the historical time period is screened to determine the optimal original knowledge data.

[0180] In this embodiment, S1 specifically includes:

[0181] S11, data input: receiving original knowledge data, wherein the original knowledge data includes documents uploaded by users, database records or other text data;

[0182] S12, preprocessing: performing preprocessing on the original knowledge data, the preprocessing includes:

[0183] Remove HTML tags, advertisements, garbled characters, and repeated text to obtain cleaned text data;

[0184] Use NLP engine to parse content and extract keyword sets from documents;

[0185] Generate document summaries to quickly understand the document content;

[0186] Use machine learning models to perform sentiment analysis and identify text sentiment features;

[0187] In this embodiment, by removing HTML tags, advertisements, garbled characters, and repeated text, noise can be reduced, data quality can be improved, and the knowledge base can be made more refined. The NLP engine performs content analysis and can extract keywords to make knowledge more searchable and improve the accuracy of subsequent classification and indexing. Automatically generating document summaries helps users quickly understand documents and improve information utilization efficiency. Using machine learning for sentiment analysis, the emotional tendency of the text can be analyzed to provide users with more intuitive knowledge classification (such as positive, negative, and neutral). Data preprocessing is a key step in NLP tasks, which can improve the clarity and consistency of data and lay a solid foundation for subsequent classification, indexing, and retrieval. NLP processing methods such as keyword extraction, summary generation, and sentiment analysis can effectively reduce redundant information, make knowledge more structured, and facilitate classification and query.

[0188] S13. Data output: Output preprocessed knowledge data and provide it to the knowledge processing and classification steps.

[0189] In this embodiment, S2 specifically includes:

[0190] S21. Machine learning classification: Apply supervised learning, unsupervised learning or semi-supervised learning algorithms to automatically classify pre-processed knowledge data;

[0191] S22, Topic Identification: Classify documents into predefined categories or generate new categories based on their content and features;

[0192] S23, index building: create an index for the classified documents to support fast retrieval, and output the classification and index results.

[0193] In this embodiment, supervised learning, unsupervised learning or semi-supervised learning algorithms are used to automatically identify and classify knowledge data, reduce manual intervention, and improve classification efficiency and accuracy. Topic identification can automatically classify knowledge data and even discover new knowledge categories, making the structure of the knowledge base more dynamic and intelligent. Index construction can improve retrieval efficiency, support rapid query of massive data, and enhance user experience. The classification ability of machine learning far exceeds the traditional rule-based classification method, and can adapt to the ever-changing knowledge system and achieve adaptive classification. Topic identification combined with NLP models such as LDA, BERT or Transformer can automatically identify knowledge fields and enhance the semantic understanding ability of knowledge. Establishing efficient indexes, such as inverted indexes or vector-based ANN (Approximate Nearest Neighbors) retrieval, makes queries faster and more accurate.

[0194] In this embodiment, S3 specifically includes:

[0195] S31, Distributed database storage: Stores data from NLP engines and machine learning classifiers, supporting structured, semi-structured, and unstructured data management;

[0196] S32, Intelligent Indexing: Build a multi-dimensional index structure based on document content and user behavior data to support fast retrieval and improve the efficiency of users in finding information;

[0197] S33. Personalized recommendations: Push relevant content based on the user’s historical behavior and preferences.

[0198] In this embodiment, the distributed database storage supports structured, semi-structured and unstructured data management, ensures the scalability of the knowledge base, and adapts to large-scale data growth. The intelligent indexing system combines NLP with user behavior data to build a multi-dimensional index, improve query efficiency, and reduce the delay of knowledge retrieval. Personalized recommendations are based on user interests and behavior history, actively push relevant knowledge, and improve knowledge utilization.

[0199] Distributed databases (such as Elasticsearch, MongoDB, and HBase) can support PB-level data storage and ensure the scalability and high availability of the system. Intelligent indexing combined with TF-IDF, Word2Vec, and BERT semantic vectors can optimize search results and improve the accuracy of user queries. Personalized recommendations use collaborative filtering, deep learning (such as Transformer), and other methods to dynamically adjust recommended content to meet user personalized needs.

[0200] In this embodiment, S4 specifically includes:

[0201] S41, Event-driven update: Real-time connection with business systems and external data sources through API interfaces, monitoring data change events, and automatically triggering the knowledge base update process;

[0202] S42, user feedback processing: provide functions such as user evaluation, annotation, and error correction, collect and analyze user feedback to optimize the content and structure of the knowledge base;

[0203] S43 Continuous monitoring and optimization: Set knowledge base health indicators, including visit volume, update frequency, and user satisfaction, regularly evaluate the performance of the knowledge base, and adjust management strategies based on the evaluation results.

[0204] In this embodiment, event-driven updates can monitor external data changes in real time to ensure that the knowledge base is always up to date. The user feedback system allows users to evaluate, annotate and correct knowledge content, improve knowledge quality, and make the knowledge base more in line with user needs. Continuous monitoring and optimization regularly optimize the knowledge base through indicators such as visit volume, update frequency and user satisfaction to improve overall system performance. Event-driven models (such as Kafka and RabbitMQ) can achieve real-time data updates and avoid lagging knowledge base information. The user feedback system combined with the active learning mechanism of NLP allows the model to continuously learn user feedback, thereby improving the accuracy of knowledge classification and retrieval.

[0205] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0206] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0207] In the present invention, unless otherwise clearly specified and limited, when a first feature is “on” or “below” a second feature, it may be that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Moreover, when a first feature is “above”, “above” or “above” a second feature, it may be that the first feature is directly above or obliquely above the second feature, or it may simply mean that the first feature is higher in level than the second feature. When a first feature is “below”, “below” or “below” a second feature, it may be that the first feature is directly below or obliquely below the second feature, or it may simply mean that the first feature is lower in level than the second feature.

[0208] In the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.

[0209] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may alter, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A knowledge base management system based on NLP and machine learning, characterized in that: The system comprises: The knowledge acquisition module is used to receive original knowledge data and perform text preprocessing to generate preprocessed knowledge data; The knowledge processing and classification module is used to receive pre-processed knowledge data, perform text vectorization and classification, and generate classified knowledge data; A knowledge storage and index building module, used to store classified knowledge data and build a knowledge index based on the classified knowledge data; The knowledge updating and maintenance module is used to filter the original knowledge data stored in the historical time period based on the search terms input by the user and their historical search times, and determine the optimal original knowledge data.

2. The knowledge base management system based on NLP and machine learning according to claim 1, characterized in that: The knowledge acquisition module includes: A data input unit, used to receive original knowledge data, wherein the original knowledge data includes documents uploaded by a user; The data preprocessing unit is used to perform preprocessing on the original knowledge data, wherein the preprocessing includes: Remove HTML tags, advertisements, garbled characters, and repeated text to obtain cleaned text data; For the cleaned text data, the BERT attention mechanism combined with the TF-IDF algorithm is used to calculate the word weights to extract the keyword set; Use the LSTM-CRF model to extract sentiment features from the cleaned text data and generate sentiment score data; Process the cleaned text data based on the BART pre-trained model, generate text summaries and output summary data; The data output unit is used to output the preprocessed knowledge data and provide it to the knowledge processing and classification module.

3. The knowledge base management system based on NLP and machine learning according to claim 2, characterized in that: For the cleaned text data, the BERT attention mechanism combined with the TF-IDF algorithm is used to calculate the word weights to extract the keyword set, including: Formula (1) is used to calculate the value of each term t in the cleaned text data. i TF-IDF weight; the formula (1) is: Among them, TF·IDF(t i ) is the term t i TF-IDF weight of f i For term t i The number of occurrences in the cleaned text data; N is the total number of cleaned text data in the corpus; wherein the corpus includes all the original knowledge data received in the historical time period; n i To contain the term t i The number of cleaned text data; TTF represents the total number of occurrences of all terms in the cleaned text data; Based on the pre-trained BERT model, the cleaned text data in the pre-processed text data is processed by the Transformer structure to obtain the attention matrix, and the formula (2) is used to calculate each term t i The average attention score across all layers of the BERT model; The formula (2) is: Among them, S i For term t i The average attention score across all layers of the BERT model; A h,i is the hth layer of the BERT model for term t i Attention score; H is the total number of layers of the BERT model; Formula (3) is used to calculate each term t i The position information P in the cleaned text data i , normalized to get term t i Position weight; The formula (3) is: Among them, R i For term t i Position weight; P i For term t i The sum of the rankings in the cleaned text data; L is the total number of terms in the cleaned text data; Based on the term t i TF-IDF weight of term t i Average attention scores and term t across all layers of the BERT model i Position weight, calculated using formula (4) for each term t i The final keyword weight W i ; The formula (4) is: W i =γ1×TF·IDF(t i )+γ2×S i +γ3×R i ; Wherein, γ1 is the first weight adjustment parameter; γ2 is the second weight adjustment parameter; γ3 is the third weight adjustment parameter; γ1+γ2+γ3=1; According to the calculated comprehensive keyword weight W i , select the first K terms with the highest keyword weights to form the final keyword set, and output the keyword set.

4. The knowledge base management system based on NLP and machine learning according to claim 3, characterized in that: The knowledge processing and classification module includes: A text vectorization unit, used for receiving preprocessed knowledge data and converting the cleaned text data into text vector data; A classification unit, used to perform text classification based on the text vector data and generate classification knowledge data; The data output unit is used to output the classified knowledge data and transmit it to the knowledge storage and index construction module.

5. The knowledge base management system based on NLP and machine learning according to claim 4, characterized in that: The classification unit is used to receive text vector data and perform text classification to obtain classified knowledge data, including: The classification unit uses a preset support vector machine model to perform text classification on the text vector data, and after classification, assigns corresponding classification knowledge data and the keyword set to each text data; The preset support vector machine model is obtained by pre-training with text vector data of the historical stage marked with classification knowledge data labels; The classification knowledge data includes classification labels and semantic vectors.

6. The knowledge base management system based on NLP and machine learning according to claim 5, characterized in that: The knowledge storage and index building module includes: A storage unit, used for storing classified knowledge data in a specified storage method; The specified storage methods include: MongoDB and PostgreSQL; An index building unit, used for building a knowledge index based on the classified knowledge data; The knowledge index includes an inverted index and an HNSW index; The data output unit is used to output the knowledge index and transmit it to the knowledge update and maintenance module.

7. The knowledge base management system based on NLP and machine learning according to claim 6, characterized in that: The knowledge updating and maintenance module includes: an event-driven updating unit, which is used to calculate the correlation score S between any original knowledge data stored in the historical time period and the search term q according to the search term q input by the user and the historical search times of the search term q; According to the relevance score S, the best original knowledge data is selected from the original knowledge data received in the historical time period; The optimal original knowledge data is the original knowledge data corresponding to the highest relevance score S.

8. The knowledge base management system based on NLP and machine learning according to claim 7, characterized in that: Based on the search term q input by the user and the historical search times of the search term q, the correlation score S between any original knowledge data stored in the historical time period and the search term q is calculated using formula (5); The formula (5) is: S is the relevance score between the search term q and any original knowledge data; f q is the number of times the search term q appears in any original knowledge data; M is the total number of original knowledge data stored in the historical time period; n q is the number of original knowledge data containing the search term q stored in the historical time period; T q is the total number of searches for the search term q in the historical time period; T is the total number of searches for all keywords; α is a preset balance factor.

9. The knowledge base management system based on NLP and machine learning according to claim 8, characterized in that: The format of the original knowledge data is any one of Word, PDF, TXT, and HTML.

10. The knowledge base management system based on NLP and machine learning according to claim 9, characterized in that: in, 0<α<1。

Citation Information

Patent Citations

  • Expert system knowledge base construction method and system

    CN112836509A

  • Mathematical knowledge marking method based on large language model

    CN118394942A

  • Knowledge graph construction method and system for chemical waste treatment technology

    CN118734958A

  • Power system vector knowledge base construction method based on text vectorization

    CN118964695A

  • A method to provide comprehensive key vocabulary for search

    WO2025027623A1