A method for constructing a knowledge service system

Through multi-dimensional classification and similarity calculation, the knowledge service system is built, the superior label is generated, and the tags and associations are dynamically adjusted in combination with user needs, which solves the problem of inefficient search in the traditional knowledge management system, and achieves personalized knowledge recommendation and user experience improvement.

CN119322858BActive Publication Date: 2025-07-08STATE GRID SHANDONG ELECTRIC POWER CO MARKETING SERVICE CENT (MEASURING CENT) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411864751.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-07-08
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

The traditional knowledge management system is inefficient in knowledge retrieval, it is difficult to quickly find comprehensive and accurate information, and lacks personalized and intelligent knowledge services.

Method used

Through multi-dimensional classification, similarity calculation and association steps, a knowledge service system is built, primary and superior labels are generated, and tags and associations are dynamically adjusted in accordance with user needs to provide personalized knowledge recommendations.

Benefits of technology

It improves the efficiency and accuracy of knowledge retrieval, provides personalized knowledge services, enhances user experience and knowledge sharing capabilities, and promotes innovation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119322858B_ABST
    Figure CN119322858B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of knowledge base construction, and particularly relates to a method for constructing a knowledge service system, which includes a knowledge acquisition step for collecting knowledge data from knowledge sources in multiple fields; a preprocessing step for preprocessing the knowledge data; a classification step for classifying the acquired knowledge data; a similarity calculation step for calculating the similarity between knowledge bases of different categories; a first judgment step for judging whether the similarity is greater than a first threshold; an association step for establishing associations between knowledge according to the similarity; a second judgment step for judging whether the similarity is greater than a second threshold; a labeling step for constructing new upper-level labels; a label directory generation step for establishing a label directory; a query requirement acquisition step for acquiring the user's knowledge retrieval requirement and generating a query request; and a knowledge retrieval step for performing knowledge recommendation according to the query request. The present application has the effect of improving the search efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of knowledge base construction, and in particular to a method for constructing a knowledge service system. Background Art

[0002] Currently, with the rapid development of information technology and the rise of the knowledge economy, knowledge service has become an important force driving innovation and development in all industries. With the rapid development of technologies such as big data and artificial intelligence, the generation, dissemination, and application speed of knowledge data have accelerated unprecedentedly, and the sources of knowledge have become increasingly diverse, covering multiple fields such as science, technology, economy, electricity, and culture. Knowledge data has shown an explosive growth, and the cross-integration of knowledge in multiple fields has become increasingly frequent.

[0003] Most traditional knowledge management systems are based on a single field or simple classification criteria, and organize knowledge in the form of a tree structure or a flat list. Although this system facilitates the storage and retrieval of knowledge to a certain extent, it has obvious limitations. First of all, the classification of knowledge is often too broad, resulting in large differences in the knowledge content under the same category. When users retrieve information, the information is too redundant, making it difficult to quickly find comprehensive and accurate information, and resulting in poor retrieval efficiency. Summary of the Invention

[0004] In order to improve the retrieval efficiency, this application provides a method for constructing a knowledge service system.

[0005] A method for constructing a knowledge service system provided by this application adopts the following technical solutions:

[0006] A method for constructing a knowledge service system includes the following steps:

[0007] Knowledge acquisition: Collect knowledge data from knowledge sources in multiple fields;

[0008] Preprocessing: Preprocess the knowledge data to eliminate noise;

[0009] Classification: Classify the obtained knowledge data, generate corresponding primary labels for each category, and store the classified data separately to form a knowledge base;

[0010] Similarity calculation: Calculate the similarity between knowledge bases of different categories;

[0011] First judgment: Judge whether the similarity is greater than a first threshold. If so, execute the association step;

[0012] Association: Establish associations between knowledge according to the similarity;

[0013] Second judgment: Judge whether the similarity is greater than a second threshold. If so, execute the labeling step;

[0014] Labeling: Analyze their common features and construct new upper-level labels to generalize and reflect the common attributes of this knowledge;

[0015] Label directory generation: Establish a label directory based on the knowledge base category, primary labels, and upper-level labels;

[0016] Retrieval requirement acquisition: Obtain the user's knowledge retrieval requirement, determine keywords or labels according to the retrieval requirement, and generate a query request;

[0017] Knowledge retrieval: According to the query request, combined with upper-level labels and knowledge associations, screen and sort the results and make knowledge recommendations.

[0018] By adopting the above technical solution, through classification, primary labels are generated for each category, and a knowledge base is formed, which helps users locate the required knowledge more quickly; this technical solution introduces a similarity calculation step, and by calculating the similarity between knowledge bases of different categories, potential knowledge associations can be discovered; through the first and second judgment steps, this technical solution can further analyze the common features between knowledge, construct new upper-level labels to generalize and reflect the common attributes of this knowledge. This helps users understand the knowledge content more deeply and simplifies the retrieval process. In addition, the generation of the label directory also improves the navigability and usability of the knowledge system.

[0019] Combined with upper-level labels and knowledge associations, this technical solution can more accurately screen and sort retrieval results and make knowledge recommendations, which provides a more personalized and intelligent user experience compared with simple keyword matching or content-based retrieval methods in the prior art;

[0020] In summary, through classification, similarity calculation, and association steps, this technical solution can construct a more complete and coherent knowledge system, which helps users understand the knowledge content more comprehensively; through the combination of upper-level labels and knowledge associations, this technical solution can more accurately screen and sort retrieval results, facilitate users to quickly find the knowledge they want to know, reduce the steps of self-screening from redundant knowledge, and improve the retrieval efficiency and accuracy.

[0021] The generation of the label directory and the intelligent knowledge recommendation function enable users to locate the required knowledge more quickly and obtain a personalized learning experience; this technical solution helps to promote knowledge sharing and innovation by constructing a systematic knowledge service system, providing strong support for academic research and business development.

[0022] Optionally, the classification step further includes: performing multi-dimensional classification based on the theme, content, and formal features of knowledge, and assigning one or more labels to the knowledge under each dimension.

[0023] By adopting the above technical solution, multi-dimensional classification can capture the characteristics of knowledge more comprehensively. It not only considers the theme and content of knowledge but also its formal features (such as document type, media format, etc.). This classification method is more accurate and detailed than single-dimensional classification, which helps users find the required knowledge more precisely.

[0024] Assigning one or more labels to the knowledge under each dimension not only increases the quantity and variety of labels but also enhances the correlation between labels. This label assignment method enables each knowledge point in the knowledge system to be interconnected through multiple labels, forming a more compact and complex knowledge network.

[0025] Multi-dimensional classification and a rich label system provide users with more retrieval entry points and paths. Users can, according to their own needs and preferences, select different dimensions and labels for combined queries, thereby obtaining the required knowledge more flexibly. At the same time, this retrieval method also improves the user's retrieval efficiency, enabling users to find relevant information faster.

[0026] Multi-dimensional classification and the label system help to discover the knowledge correlations and intersections between different fields and themes. This relevance and intersectionality provide the possibility for the cross-integration and innovation of knowledge, which helps to stimulate new ideas, viewpoints, and methods.

[0027] By classifying knowledge multi-dimensionally and assigning labels, the knowledge system can be managed and organized more effectively. This management method not only improves the accessibility and comprehensibility of knowledge but also contributes to the long-term preservation and inheritance of knowledge.

[0028] Optionally, the labeling step further includes: analyzing the commonalities and differences of similar knowledge bases, as well as their positions and roles in the knowledge system, to determine the naming and hierarchical structure of the upper-level labels.

[0029] By adopting the above technical solution, through in-depth analysis of the commonalities and differences of similar knowledge bases, the internal connections and essential characteristics between them can be grasped more accurately. This analysis helps to determine the naming of the upper-level labels to ensure that the labels can accurately reflect the common attributes of the relevant knowledge. At the same time, considering the position and role of knowledge in the knowledge system, the hierarchical structure and logical relationship of the upper-level labels can be further ensured to be reasonable, thereby enhancing the representativeness of the labels.

[0030] This refinement operation in the labeling step helps to clarify the hierarchical and logical relationships between knowledge. By determining the hierarchical structure of the upper-level labels, the subordinate relationships and association paths between knowledge can be clearly shown. The optimization of this hierarchical structure and logical relationship makes the knowledge system more well-organized and easy to understand, which helps users better grasp the overall structure and internal logic of the knowledge system.

[0031] An accurate and hierarchical superordinate tag system provides users with more retrieval paths and options; users can, according to their needs, combine queries by selecting superordinate tags at different levels, so as to locate the required knowledge more quickly; at the same time, since superordinate tags can accurately reflect the common attributes of relevant knowledge, the accuracy and relevance of retrieval results can be improved.

[0032] By analyzing the commonalities and differences of similar knowledge bases, as well as their positions and roles in the knowledge system, the knowledge associations and intersections between different fields and different topics can be discovered. This relevance and intersectionality make it possible for in-depth mining and integration of knowledge, and contribute to the formation of a more comprehensive and in-depth knowledge system.

[0033] The optimized superordinate tag system provides strong support for the intelligence and personalization of knowledge services; based on information such as users' historical retrieval records and interest preferences, superordinate tags and knowledge system paths related to users' needs can be intelligently recommended, so as to provide more accurate and personalized knowledge services.

[0034] Optionally, after the query requirement acquisition step, a tag circumscription step is also set;

[0035] Tag circumscription: According to the query request, initially recommend knowledge base categories and superordinate tags for the querier to select, and enrich the query request.

[0036] By adopting the above technical solution, the tag circumscription step provides a more definite query direction for the querier by initially recommending knowledge base categories and superordinate tags. The querier can select the tags and categories most relevant to their needs according to the recommendation, so as to quickly narrow down the query scope and improve the query efficiency; at the same time, since the recommendation is based on the intelligent analysis of the query request, the accuracy of the query can also be improved.

[0037] The tag circumscription step allows the querier to select tags and categories according to their own needs. This interactive query method enhances the user's sense of participation and satisfaction; users can make personalized selections according to their interests and preferences, so as to obtain knowledge recommendations more in line with their own needs.

[0038] Through the tag circumscription step, the querier can further enrich and improve their query request; they can add or modify query keywords according to the recommended tags and categories, so as to express their needs more accurately; this optimized query request can enhance the matching ability of the retrieval system and improve the quality and relevance of retrieval results.

[0039] The label circumscription step not only helps the inquirer quickly find the required knowledge, but also guides them to discover more knowledge in related fields; by recommending relevant knowledge base categories and upper-level labels, the inquirer can expand their knowledge horizons and discover more potential valuable knowledge resources.

[0040] By intelligently analyzing the query request and recommending relevant knowledge base categories and upper-level labels, the system can provide users with more intelligent and personalized knowledge services; this kind of intelligent service not only improves the query efficiency, but also can be continuously optimized and improved according to the user's feedback, enhancing the overall service quality.

[0041] Optionally, in the knowledge retrieval step, it also includes: dynamically adjusting the weights of upper-level labels and knowledge associations according to the user's retrieval history, preferences and behavior patterns, and optimizing the knowledge recommendation ranking.

[0042] By adopting the above technical solutions, by analyzing the user's retrieval history, preferences and behavior patterns, the system can more accurately understand the user's real needs; on this basis, dynamically adjusting the weights of upper-level labels and knowledge associations, making the recommended knowledge more in line with the user's personalized needs; this personalized recommendation method can significantly improve user satisfaction and loyalty.

[0043] The system dynamically adjusts the weights of knowledge associations according to the user's preferences and behavior patterns, so that the knowledge most relevant to the user's interests can be displayed first; this helps to reduce the user's invalid clicks and improve the retrieval efficiency; the user can find the required knowledge in a shorter time, thus enhancing the overall user experience.

[0044] By dynamically adjusting the weights of upper-level labels, the system can more accurately evaluate the relevance and importance of knowledge; on this basis, sorting the knowledge so that the knowledge most relevant and important to the user's needs can be ranked in the front; this helps to improve the accuracy and quality of knowledge recommendation, enabling users to more easily obtain valuable information.

[0045] Personalized recommendation and optimized knowledge ranking can enhance user stickiness, making users more willing to stay and explore in the system; this helps to promote the sharing and dissemination of knowledge, forming a more active and rich knowledge community.

[0046] By continuous learning and optimization, the system can more accurately grasp the user's preferences and behavior patterns, thus providing more intelligent knowledge services; the improvement of this intelligent level helps to enhance the competitiveness and market position of the system.

[0047] As the needs and interests of users change, the system can dynamically adjust the weights of upper-level tags and knowledge associations, thus ensuring that the recommended knowledge always aligns with the users' needs; this helps to promote the update and maintenance of knowledge, and maintain the timeliness and accuracy of the knowledge system.

[0048] Optionally, a deviation acquisition step is also set after the query requirement acquisition step;

[0049] Deviation acquisition: includes a requirement similarity calculation step, a third judgment step, a difference statistics step, a difference acquisition step, a difference judgment step, a correction step, a fourth judgment step, a supplementary calculation step, a supplementary judgment step, and a supplementary label step;

[0050] Requirement similarity calculation: Calculate multiple query requests from the same query requester, and calculate the similarity between the previous and subsequent query requests;

[0051] Third judgment: Judge whether the similarity is greater than a preset third threshold. If so, execute the difference statistics step; otherwise, execute the knowledge retrieval step;

[0052] Difference statistics: Count the number of difference similarities, and judge whether it is greater than a fourth threshold. If so, execute the difference judgment step;

[0053] Difference acquisition: Extract the words modified multiple times, denoted as difference words, obtain the common points between multiple difference words, denoted as real meaning words, calculate the association degree between the real meaning words and the upper-level tags, and then execute the difference judgment step;

[0054] Difference judgment: Judge whether the association degree is less than a preset fifth threshold. If so, execute the correction step; otherwise, execute the fourth judgment step;

[0055] Correction: Intervene in the labeling step and modify and update the upper-level tags;

[0056] Fourth judgment: Judge whether the number of difference similarities is greater than a sixth threshold. If so, execute the supplementary calculation step;

[0057] Supplementary calculation: Calculate the similarity between upper-level tags, and denote it as synonymous similarity;

[0058] Supplementary judgment: Judge whether the synonymous similarity is greater than a seventh threshold. If so, execute the supplementary label step;

[0059] Supplementary label: Set pop-ups for multiple upper-level tags for the query requester to make additional selections again to correct the query request.

[0060] By adopting the above technical solution, the demand similarity calculation step can identify and compare multiple query requests of the same query requester. By calculating the similarity between the previous and subsequent query requests, it helps to discover minor changes or continuous concerns in the query requests; when the similarity is higher than the preset threshold, the difference statistics step can further analyze the differences between the query requests. Especially when the number of similar differences exceeds the threshold, it indicates that the query requester may encounter a certain query obstacle or the demand is not fully met;

[0061] The difference acquisition step can intelligently identify the true intention of the query requester by extracting the words modified multiple times (difference words) and finding the common points (substantive words) between these words; the difference judgment step can judge whether the current label accurately reflects the needs of the query requester by calculating the correlation degree between the substantive words and the upper-level label; if the correlation degree is lower than the threshold, the labeling step is intervened through the correction step to ensure the accuracy and effectiveness of the label; when the number of similar differences exceeds another threshold, the supplementary calculation step and the supplementary judgment step can identify the similarity between the upper-level labels, and then discover possible synonymous labels or label redundancy;

[0062] The supplementary label step provides a pop-up window of multiple upper-level labels for the query requester to select again, which not only enhances the interactivity of the query, but also provides an opportunity to correct the query request, thus improving the satisfaction of the query results; by introducing multiple judgment steps and threshold settings, this technical solution can flexibly respond to different types of query demand changes, avoiding the limitations that may be brought by a single query path; by dynamically adjusting the upper-level labels and label correlation degrees in the query process, the query process is optimized, and the query efficiency and accuracy are improved.

[0063] The correction step and the supplementary label step not only help to solve the problems of the current query request, but also can promote the update and maintenance of the knowledge system; by continuously adjusting and optimizing the upper-level labels, it can ensure that the knowledge system is synchronized with the needs of the query requester.

[0064] Optionally, a user participation step, an artificial correction step, an associated label matching step, a feedback statistics step, and an update step are also provided between the difference judgment step and the correction step;

[0065] User participation: Recommend the substantive words to the query requester, judge whether the substantive words meet the needs of the query requester. If so, execute the correction step and the associated label matching step; otherwise, execute the artificial correction step;

[0066] Artificial correction: The query requester corrects the substantive words, and then executes the associated label matching step;

[0067] Associated tag matching: Match relevant upper-level tags through content words and push them to the query requester. According to the feedback of the query requester, execute the feedback statistics step;

[0068] Feedback statistics: Obtain the number of times the relevance between a content word and an upper-level tag is corrected, and determine whether it is greater than the eighth threshold. If so, execute the update step;

[0069] Update: Update the relevance between the upper-level tag and the content word for subsequent query request matching.

[0070] By adopting the above technical solution, in the user participation step, content words are directly recommended to the query requester and they are asked to judge whether these words meet the requirements, significantly enhancing the user's participation and interactivity, which helps to ensure that the query results are closer to the user's true intention.

[0071] When the user confirms that the content words meet the requirements, relevant upper-level tags are further pushed through the associated tag matching step, which not only improves the accuracy of the query but also enhances the user's satisfaction; the manual correction step allows the query requester to correct the content words, which provides the user with an opportunity to directly intervene in the query process, thus helping to correct possible misunderstandings or deviations and improving the accuracy of the query results; the associated tag matching step further refines the query results through the matching of content words and upper-level tags, making the results more accurate and targeted.

[0072] The feedback statistics step can identify which tag associations may have problems or need improvement by collecting the number of times the user corrects the relevance between content words and upper-level tags; the update step updates the relevance between the upper-level tag and the content word according to the results of the feedback statistics, which helps to optimize the knowledge system to better meet the user's query needs and habits.

[0073] By introducing the user feedback and manual correction mechanisms, it helps to improve the intelligence and adaptability of the system, enabling it to better handle future query needs; steps such as user participation, manual correction, and feedback statistics enhance the interaction between the user and the system, making the user feel the system's attention and respect for their needs; this helps to establish a trust relationship between the user and the system and improve the user's loyalty and satisfaction with the system.

[0074] Optionally, after the associated tag matching step, a requirement judgment step and a single association step are also set;

[0075] Requirement judgment: Judge whether the generated tags meet the needs of the query requester. If so, execute the feedback statistics step; otherwise, execute the single association step;

[0076] Single Association: The query requester inputs custom tags on their own, matches them from the upper-level tags, and establishes a single association for the query request of the query requester in this instance, and then performs the knowledge retrieval step.

[0077] By adopting the above technical solution, the requirement judgment step provides an additional verification link by judging whether the generated tags meet the requirements of the query requester; this helps to ensure that the tags recommended by the system highly match the actual needs of the user, thereby improving the accuracy and relevance of the query results.

[0078] When the tags generated by the system do not meet the user's requirements, the single association step allows the user to input custom tags on their own and match them from the upper-level tags, which provides a way for the user to directly express their needs and further enhances the precise satisfaction of the user's needs.

[0079] The single association step allows the user to input custom tags on their own, which significantly improves the user's participation. The user can flexibly select or create tags according to their own understanding and needs, so as to more accurately describe their query intention; by establishing a single association, the system can respond to the specific needs of the query requester in this instance and provide a more personalized and customized query service.

[0080] The introduction of the requirement judgment step and the single association step makes the query process more efficient and flexible. When the user finds that the tags generated by the system do not meet their requirements, they can immediately correct them through the single association step without having to repeat the entire query process; this not only improves the query efficiency, but also reduces the user's waiting time and frustration, thereby enhancing the user experience.

[0081] After the single association step, the knowledge retrieval step is executed. The system can perform more precise knowledge retrieval based on the user-defined tags, which helps to dig deeper information and associations and provide more comprehensive and in-depth query results for the user.

[0082] Through the tag association established by the single association step, the system can also better understand and process the user's query intention, so as to provide more accurate and relevant knowledge retrieval results; the introduction of the requirement judgment step and the single association step enables the system to better adapt to the query needs and habits of different users, which helps to enhance the adaptability and scalability of the system and enables it to handle more complex and diverse query scenarios; at the same time, by allowing the user to input custom tags on their own, the system can also continuously learn and accumulate new tags and association information, thereby continuously optimizing its query algorithm and knowledge system.

[0083] Optionally, after the single association step, an association statistics step is also set;

[0084] Association statistics: Count the number of times a single association is established for the same upper-level tag, and determine whether it is greater than the ninth threshold. If so, execute the tag conversion step;

[0085] Tag conversion: Convert the custom tag to an upper-level tag and establish an association with the knowledge base.

[0086] By adopting the above technical solution, the association statistics step can identify which upper-level tags frequently appear in the user's query process and establish a single association with them by counting the number of times a single association is established for the same upper-level tag; this helps the system better understand the user's query needs and habits, thereby optimizing tag management and the knowledge system.

[0087] When the number of single associations of a certain upper-level tag exceeds the set ninth threshold, convert the custom tag to an upper-level tag through the tag conversion step and establish an association with the knowledge base; this can not only reduce the redundancy and confusion of tags, but also improve the accuracy and authority of tags, making the knowledge system clearer and more orderly.

[0088] The tag conversion step converts frequently used custom tags into upper-level tags and establishes an association with the knowledge base, which helps the system identify and match relevant tags faster in future query processes, thereby improving query efficiency;

[0089] At the same time, since upper-level tags usually have broader meanings and higher generality, using upper-level tags for queries can reduce ambiguity and misunderstanding and improve query accuracy; the introduction of the association statistics step and the tag conversion step enables the system to make adaptive adjustments and optimizations according to the user's query behavior and needs; this helps improve the system's adaptability and learning ability, enabling it to better handle future possible query needs.

[0090] By continuously learning and accumulating new tag and association information, the system can continuously improve its tag system and knowledge system, thereby providing more intelligent and efficient query services; the introduction of the association statistics step and the tag conversion step enables the system to more accurately understand the user's query intent and needs, thereby providing more personalized and customized query results; this helps improve the user experience and satisfaction, and enhances the user's trust and loyalty to the system

[0091] Optionally, after the knowledge retrieval step, a recommendation statistics step is also set;

[0092] Recommendation statistics: Obtain the click-through rate of the recommended knowledge by the query requester, determine the matching degree, and correct the sorting of subsequent similar query requirements.

[0093] By adopting the above technical solutions, the recommendation statistics step can more accurately understand users' preferences and satisfaction with recommended content by collecting click-through rate data of users on recommended knowledge; by analyzing the click-through rate data, the system can evaluate the matching degree and effectiveness of different recommended content. This helps the system continuously optimize its recommendation algorithm and improve the accuracy and relevance of recommendations. For example, the system can adjust the weights, sorting, or display methods of recommended content according to the click-through rate data to better meet users' needs.

[0094] The recommendation statistics step enables the system to more deeply understand users' personalized needs. By collecting and analyzing users' click behaviors, the system can identify users' interest points, preferences, and habits, thereby providing more personalized recommendation services, which helps enhance users' satisfaction and loyalty.

[0095] For subsequent similar query requirements, the system can correct the sorting according to the results of the recommendation statistics step, which means that the system can find the content that best meets users' needs faster, reduce the time for users to browse and filter, and improve the query efficiency.

[0096] The recommendation statistics step not only focuses on the satisfaction of the current query requirements but also guides future knowledge updates and iterations by collecting and analyzing data. When the system finds that the click-through rate of some recommended content is low, it can consider updating or replacing these contents to better adapt to the changes in users' needs; by introducing the recommendation statistics step, the system can continuously learn and adapt to users' query behaviors and needs; this helps improve the intelligence of the system, enabling it to more accurately understand users' intentions and provide more intelligent and efficient query services.

[0097] In summary, this application includes at least one of the following beneficial technical effects:

[0098] 1. Through the classification, similarity calculation, and association steps, this technical solution can construct a more complete and coherent knowledge system, which helps users more comprehensively understand the knowledge content; through the combination of upper-level tags and knowledge association, this technical solution can more accurately screen and sort retrieval results, facilitating users to quickly find the knowledge they want to know, reducing the steps of self-screening from redundant knowledge, and improving the retrieval efficiency and accuracy;

[0099] 2. By introducing multiple judgment steps and threshold settings, it can flexibly cope with changes in different types of query requirements, avoiding the limitations that may be brought by a single query path; by dynamically adjusting the upper-level tags and tag association degrees in the query process, the query process is optimized, and the query efficiency and accuracy are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0100] Figure 1 is a flowchart of the method for constructing a knowledge service system in Embodiment 1 of this application;

[0101] Figure 2 is the flowchart of the positional relationship of the deviation acquisition step in Embodiment 2 of the present application;

[0102] Figure 3 is the flowchart of the sub-steps of the deviation acquisition step in Embodiment 2 of the present application;

[0103] Figure 4 is the flowchart from the user participation step to the update step in Embodiment 2 of the present application;

[0104] Figure 5 is the flowchart from the requirement judgment step to the single association step in Embodiment 2 of the present application. Detailed implementation manners

[0105] The following is combined with Figures 1 to 5 to further elaborate on the present application in detail.

[0106] This embodiment discloses a method for constructing a knowledge service system.

[0107] Embodiment 1: Refer to Figure 1 , the method for constructing a knowledge service system includes the following steps:

[0108] Knowledge acquisition: Collect knowledge data from knowledge sources in multiple fields;

[0109] Specifically, select authoritative knowledge sources covering multiple fields, such as academic paper databases, professional books, online courses, industry reports, etc.; use web crawler technology, API interfaces or manual entry, etc. to batch collect knowledge data from the above-mentioned knowledge sources, including various forms such as text, pictures, and videos.

[0110] Preprocessing: Preprocess the knowledge data to eliminate noise;

[0111] Specifically, the preprocessing includes data cleaning, text processing, and data format unification; among them, data cleaning includes removing duplicate data, invalid data (such as blanks, garbled characters) and advertising information, and cleaning the original data set through a data cleaning tool; text processing includes natural language processing operations such as word segmentation, stop word removal, and stemming, and natural language processing libraries (such as NLTK, jieba) can be used for processing; data format unification includes converting data in different formats into a unified storage format for subsequent processing, and converting text data into JSON or XML format.

[0112] Classification: Classify the obtained knowledge data, generate corresponding primary labels for each category, and store the classified data separately to form a knowledge base;

[0113] Specifically, the acquired knowledge is initially classified by domain, and then a machine learning algorithm is selected for model training. The model can adopt a Naive Bayes classifier. During the training process, a dataset with known classifications (also known as the training set) is required to train the classification model. This dataset contains a large number of knowledge data samples, and each sample has been manually labeled with the correct classification label. By enabling the model to learn the relationship between the features and labels of these samples, the model can gradually learn how to classify new knowledge data.

[0114] Then, personnel with a professional knowledge background conduct manual review to check the results of automatic classification one by one, ensuring that each piece of knowledge data is correctly classified into the corresponding category. If classification errors are found, the model needs to be adjusted and optimized to improve the classification accuracy. After classification, a descriptive primary label needs to be generated for each classification, which should accurately reflect the theme and content of the knowledge data included in this classification. The process of generating the primary label can be automated or manual. The methods for automated generation of primary labels usually include natural language processing techniques such as text summarization and keyword extraction. These techniques can extract key information from the knowledge data in the classification and generate a concise label to describe this classification.

[0115] For example, in the power industry, the initial domain can be divided into:

[0116] Power load forecasting: This domain mainly focuses on the changing trends and forecasting of power loads. The data may include historical load data, weather forecasts, holiday information, economic indicators, etc., which are used to train the model to predict future power demand.

[0117] Power market trading: This domain involves trading data and price information in the power market. The data may include power plant quotes, power trading volumes, power prices, etc., which are used to analyze market dynamics, formulate trading strategies, etc.

[0118] Power grid operation and maintenance: This domain focuses on the stable operation and fault maintenance of the power grid. The data may include power grid topology, equipment status monitoring data, fault records, etc., which are used to monitor the power grid status, predict faults, and take corresponding maintenance measures.

[0119] Power energy management: This domain involves the production, distribution, and use of energy. The data may include the production capacity of power plants, energy transmission losses, user electricity consumption behaviors, etc., which are used to optimize energy allocation and improve energy utilization efficiency.

[0120] Power Safety and Emergency: This area focuses on the safety and emergency response capabilities of power systems. Data may include power grid fault simulation data, emergency response plans, accident records, etc., which are used to evaluate the safety risks of the power grid and formulate emergency response strategies.

[0121] New Power Technologies and Applications: This area focuses on new technologies and innovative applications in the power field. Data may include smart grid technologies, electric vehicle charging station data, distributed energy system data, etc., which are used to explore the application potential and effects of new technologies in power systems.

[0122] Similarity Calculation: Calculate the similarity between different categories of knowledge bases;

[0123] First Judgment: Judge whether the similarity is greater than the first threshold. If so, execute the association step;

[0124] Association: Establish associations between knowledge based on similarity;

[0125] Specifically, similarity calculation is a quantitative indicator to measure the similarity degree between two or more knowledge items. Similarity calculation methods may include cosine similarity, Jaccard similarity, and semantic similarity;

[0126] Cosine Similarity: Knowledge items can be represented as vectors, and each dimension of the vector represents a feature (such as keywords, topics, etc.). Cosine similarity = (A·B) / (|A|*|B|), where A and B respectively represent the vector representations of two knowledge items, · represents the dot product of vectors, and |A| and |B| respectively represent the norms of vectors A and B, which is applicable to measuring the similarity degree between knowledge items based on feature vectors;

[0127] Jaccard Similarity: Knowledge items can be represented as sets, and the elements in the set represent the features of the knowledge items (such as keywords, concepts, etc.). Jaccard similarity = |A∩B| / |A∪B|, where A and B respectively represent the set representations of two knowledge items, ∩ represents the intersection, and ∪ represents the union, which is applicable to measuring the similarity degree between knowledge items based on set representations.

[0128] Semantic Similarity: Semantic similarity measures the similarity degree between two knowledge items by calculating the distance between them in the semantic space. This usually requires the use of natural language processing techniques and semantic models (such as word embeddings, knowledge graphs, etc.) to achieve; Pre-trained semantic models can be used to calculate the semantic distance between knowledge items, such as cosine distance, Euclidean distance, etc., and then the distance is converted into a similarity score, usually using the reciprocal or negative exponential function of the distance, which is applicable to measuring the similarity degree between knowledge items based on semantic understanding.

[0129] Then, based on the comparison and screening of the similarity calculation results with the set first threshold, the association relationship between knowledge items is established. The greater the similarity, the closer the association. After exceeding the set first threshold, the association between the two is established, which can be achieved by creating association tables, relationship diagrams in the database or using graph database technologies, etc.

[0130] Verify and optimize the established association relationship, which can be achieved through manual review, user feedback or statistical-based methods.

[0131] Among them, the determination of the first threshold can be based on past project experience or expert opinions, and a fixed similarity threshold is directly set.

[0132] Or by analyzing a large amount of data, calculating the similarity distribution, and then setting the first threshold according to the distribution characteristics. For example, a certain percentile in the similarity distribution can be selected as the threshold, such as the 90% or 95% percentile.

[0133] Or dynamically adjust the first threshold according to the characteristics of the current data set. Machine learning algorithms such as support vector machines (SVM) or decision trees can be used to learn and predict the optimal threshold based on data features.

[0134] Second judgment: Determine whether the similarity is greater than the second threshold. If so, perform the labeling step;

[0135] Labeling: Analyze their common features and construct a new upper-level label to summarize and reflect the common attributes of this knowledge;

[0136] Label directory generation: Establish a label directory according to the knowledge base category, primary label and upper-level label;

[0137] Retrieval requirement acquisition: Obtain the user's knowledge retrieval requirement, determine keywords or labels according to the retrieval requirement, and generate a query request;

[0138] Knowledge retrieval: According to the query request, combined with the upper-level label and knowledge association, screen and sort the results to perform knowledge recommendation.

[0139] Specifically, set the second threshold, screen out highly similar knowledge items, manually or automatically analyze the common features of these knowledge items, construct upper-level labels according to the common features, and assign them to relevant knowledge items; design the hierarchical structure and logical relationship of the label directory, organize various labels and knowledge items according to the designed structure, construct a hierarchical label directory according to the knowledge base category, primary label and upper-level label, and display the label directory in the form of a graphical interface or web page for easy user browsing and querying.

[0140] Provide a user input interface for the inquirer, receive the user's knowledge retrieval requirements, generate a query request based on the keywords entered by the user or the selected tags, and combine the upper-level tags and knowledge associations to screen and sort the retrieval results to generate a recommended list.

[0141] Among them, the second threshold can be determined by statistical analysis method, experimental method and expert consultation method. For the statistical analysis method: analyze the similarity distribution between knowledge items, observe the concentrated area and dispersion degree of similarity, and set the second threshold according to the statistical characteristics of the similarity distribution (such as mean, standard deviation, percentile, etc.); for example, a higher percentile (such as 95% or 99%) can be selected as the second threshold to screen out highly similar knowledge items.

[0142] Experimental method: Test the screening effect under different thresholds through experiments. Manual evaluation or machine learning algorithms can be used to evaluate the accuracy and relevance of the screening results, and adjust the second threshold according to the experimental results until a satisfactory screening effect is achieved.

[0143] Expert consultation method: Invite domain experts or practitioners with rich experience to participate in the setting of the second threshold, and use the experts' professional knowledge and experience to judge which knowledge items are highly similar and set the second threshold accordingly.

[0144] For example, when performing power load forecasting, for the specific field of "power load forecasting", we can collect relevant data from the following knowledge sources:

[0145] Academic paper databases: Search databases such as IEEE Xplore and CNKI to find academic papers on power load forecasting, and understand the latest forecasting models, algorithms and case studies.

[0146] Professional books: Consult classic works in the fields of electrical engineering and power system analysis to understand the basic principles, methods and historical development of power load forecasting.

[0147] Online courses: Learn professional courses related to power load forecasting on platforms such as China University MOOC and NetEase Cloud Classroom to master skills such as the construction of forecasting models, parameter optimization and result evaluation.

[0148] Industry reports: Read the annual reports of power enterprises such as State Grid and China Southern Power Grid, as well as market analysis reports released by power industry associations and research institutions to understand the application status and trends of power load forecasting in the power industry.

[0149] The collected power load forecasting data may contain noise and outliers, and preprocessing is required to improve the data quality. The specific steps include:

[0150] Data cleaning: Remove duplicate data, invalid data (such as blanks, garbled characters), and outliers (such as extreme load values) to ensure the accuracy and validity of the data.

[0151] Text processing (if the data contains text information): Perform natural language processing operations such as word segmentation and stop word removal on the text data for subsequent analysis.

[0152] Unify data formats: Convert data in different formats into a unified storage format, such as JSON or CSV, for easy subsequent processing and storage.

[0153] Data normalization: Normalize the load data to eliminate the influence of different dimensions on the data and improve the prediction performance of the model.

[0154] Classify the preprocessed power load prediction data according to different features, such as time scale (daily load, weekly load, monthly load, etc.), geographical scope (city, region, province, etc.), and load type (residential load, industrial load, commercial load, etc.). Generate corresponding primary labels for each category, such as "Daily load prediction - Urban residents", "Monthly load prediction - Industrial area", etc.

[0155] Calculate the similarity between different categories of power load prediction data and establish associations between knowledge based on the similarity. The specific steps include:

[0156] Selection of similarity calculation method: Divide the power load subsequences:

[0157] Divide two power load curves S and S' to be compared into equal-length and equally-spaced power load subsequences S(i) and S'(i) according to a preset same rule, where i = 1, 2,..., n, and n is the number of power load subsequences.

[0158] The divided power load subsequences S(i) and S'(i) have the same sampling interval, that is, the number of elements m(i) and m'(i) of the power load subsequences S(i) and S'(i) are the same.

[0159] Set the weights of the power load subsequences: The weight w(i) of each power load subsequence S(i) is set according to the formula w(i) = P(i) / P, where P(i) is the sum of the load values of each element in the power load subsequence S(i), and P is the total load value of the power load curve S.

[0160] Calculate the distance between power load subsequences: Use, for example, the Euclidean distance to calculate the distance between corresponding power load subsequences.

[0161] According to the distances between the power load subsequences and the weights of the power load subsequences, after obtaining the distances between each subsequence, multiply these distances by the corresponding weights to obtain weighted distances; then, sum or average all the weighted distances to obtain a total similarity score, and calculate the similarity between two power load curves to be compared.

[0162] Similarity calculation and comparison: Calculate the similarity between different load prediction data and compare it with a set first threshold. If the similarity is greater than the first threshold, establish an association between the two. For example, there may be an association between "Daily load prediction - Urban residents" and "Daily load prediction - Urban commerce".

[0163] Association relationship verification and optimization: Verify the accuracy and effectiveness of the association relationship through manual review or statistical methods and optimize it.

[0164] Select highly similar power load prediction data, analyze their common features, and construct new upper-level tags. The specific steps include:

[0165] Setting of the second threshold: Set an appropriate second threshold according to methods such as statistical analysis method, experimental method, and expert consultation method.

[0166] Construction of upper-level tags: Select data with a similarity greater than the second threshold, analyze their common features (such as time scale, geographical scope, load type, etc.), and construct new upper-level tags, such as "Urban resident load prediction", "Industrial area load prediction", etc.

[0167] Generation of tag directory: Based on the knowledge base category, primary tags, and upper-level tags, establish a hierarchical tag directory for easy browsing and querying by users.

[0168] According to the user's knowledge retrieval requirements, combine upper-level tags and knowledge associations to screen and sort power load prediction data for knowledge recommendation. The specific steps include:

[0169] Design of user input interface: Design a user input interface to receive the user's knowledge retrieval requirements for power load prediction, such as entering a keyword "Daily load prediction" or selecting a tag "Urban resident load prediction".

[0170] Generation of query request: Generate a query request according to the keyword entered by the user or the selected tag.

[0171] Knowledge retrieval and screening: Combine upper-level tags and knowledge associations to retrieve and screen power load prediction data, such as screening out all data related to "Daily load prediction - Urban residents".

[0172] Knowledge recommendation: Generate a recommendation list to display relevant power load prediction data, upper-level tags, and associated knowledge to users, helping users quickly find the knowledge they need.

[0173] In the embodiments of the present application, through the refinement of upper-level tags, screening options are added to facilitate narrowing down the search scope. The requester can quickly query based on the upper-level tags, reducing the interference of redundant information and quickly obtaining relevant knowledge.

[0174] Embodiment 2: Refer to Figure 2 , the difference between this embodiment and Embodiment 1 is that in the classification step, it further includes: performing multi-dimensional classification based on the theme, content, and formal features of knowledge, and assigning one or more tags to the knowledge under each dimension.

[0175] Specifically, for each knowledge item, relevant information on its theme, content, and formal features needs to be extracted, which can be achieved through natural language processing (NLP) techniques, image recognition techniques, and metadata extraction methods. For example, for text-based knowledge items, NLP techniques can be used to extract keywords, topic sentences, paragraph structures, etc.; for image or video-based knowledge items, image recognition techniques can be used to extract objects, scenes, colors, etc. Based on the extracted feature information, a classification system is constructed, which includes determining the classification criteria, classification levels, and classification relationships under each dimension. For example, under the theme dimension, a classification tree of subject fields can be constructed to hierarchically classify knowledge items according to subject fields; under the content dimension, knowledge items can be classified according to classification criteria such as knowledge points, cases, and theories.

[0176] According to the classification system, assign one or more classification tags to each knowledge item, compare the feature information of the knowledge item with the classification criteria in the classification system to find the most suitable classification tag. At the same time, multi-label classification techniques can also be considered to assign multiple relevant tags to the knowledge item to more comprehensively describe its attributes and features.

[0177] Before label assignment, a reasonable label set needs to be designed, and labels are established based on the theme, content, and formal features of knowledge items. At the same time, there should be a certain degree of independence and mutual exclusivity between labels to avoid redundancy and repetition.

[0178] Label assignment can adopt a combination of automatic assignment and manual review. Automatic assignment can be achieved through methods such as rule-based matching algorithms and machine learning algorithms to automatically assign labels according to the feature information of knowledge items; manual review is to verify and correct the results of automatic assignment to ensure the accuracy and rationality of labels.

[0179] As the knowledge base is continuously updated and expanded, the tags also need to be updated and maintained accordingly. This includes operations such as adding new tags, deleting obsolete tags, and merging duplicate tags. At the same time, it is also necessary to monitor and analyze the usage of tags to promptly discover and solve problems.

[0180] For example, data item: the power load data of a certain area on a certain day, including residential load, industrial load, and commercial load.

[0181] Feature extraction: Theme feature: power load data; Content features: residential load, industrial load, commercial load, load value, load change trend; Form feature: table form

[0182] Classification system: Theme dimension: power load data; Content dimension: classified by user type (residential, industrial, commercial); Tag assignment: Theme tag: power load data;

[0183] Content tags: residential load analysis, industrial load analysis, commercial load analysis, load value analysis, load change trend analysis;

[0184] Tag update and maintenance: As the power data knowledge base is continuously updated and expanded, the tags also need to be updated and maintained accordingly. This includes: adding new tags to cover new data types or analysis requirements; deleting obsolete tags to avoid redundancy and confusion; merging duplicate tags to maintain the simplicity and consistency of the tag set.

[0185] In other embodiments, the labeling step further includes: analyzing the commonalities and differences of similar knowledge bases, as well as their positions and roles in the knowledge system, to determine the naming and hierarchical structure of the upper-level tags.

[0186] Specifically, similarity calculation methods (such as cosine similarity, Jaccard similarity, etc.) are used to calculate the similarity of knowledge items between different knowledge bases. Based on the similarity threshold, highly similar knowledge items are screened out to form a knowledge item set. The knowledge item set is analyzed in terms of subject, content, form, etc. to find out their similarities (such as common themes, similar descriptions, etc.) and differences (such as different emphases, differences in details, etc.), and the analysis results are recorded. Based on the data and analysis results of the existing knowledge base, a knowledge system framework is constructed. The framework should include the hierarchical structure of the knowledge base, the relationship between knowledge items, etc. Each knowledge item is positioned in the knowledge system framework, and its subject, category, and level are clarified. The location information of the knowledge item is recorded and analyzed. The role of the item in the knowledge system, such as whether it is core knowledge, whether it is related to other knowledge items, etc.; according to the results of the role analysis, adjust the knowledge system framework to ensure the accuracy and completeness of the knowledge items; according to the results of the analysis of the similarities and differences of the knowledge items, as well as their position and role in the knowledge system, determine the name of the superordinate label, and according to the knowledge system framework and the positioning information of the knowledge items, use different hierarchical structures such as tree structure and network structure to construct the hierarchical structure of the superordinate label; optimize and adjust the naming and hierarchical structure of the superordinate labels to ensure that they can accurately reflect the characteristics of the knowledge base and the structure of the knowledge system; invite domain experts or experienced practitioners to review and provide feedback to improve the accuracy and rationality of the superordinate labels.

[0187] According to the determined upper-level label naming and hierarchical structure, the knowledge items in the knowledge base are labeled. During the labeling process, attention should be paid to maintaining the consistency and accuracy of the labels. The labeling results should be verified and evaluated to ensure the accuracy and rationality of the labels; based on the verification results, necessary optimization and adjustment should be made to the upper-level labels and hierarchical structure.

[0188] Reference Figure 2 In other implementations, after the query requirement acquisition step, a label delineation step is also provided;

[0189] Tag delineation: Based on the query request, preliminary knowledge base categories and upper-level tags are recommended for the queryer to choose from, and the query request is enriched.

[0190] Specifically, the system first receives the user's query request and parses the request. The parsing process includes identifying keywords, phrases, and possible contextual information in the query. Through natural language processing technology (such as word segmentation, part-of-speech tagging, named entity recognition, etc.), the system can more accurately understand the user's query intention.

[0191] Based on the parsed query request, the system filters out relevant categories from the knowledge base. The category recommendations are made based on methods such as keyword matching and semantic similarity calculation. For example, if the query request contains the keyword "smart grid", the system can recommend knowledge base categories related to artificial intelligence, such as "power grid dispatching" and "distributed energy management".

[0192] While recommending knowledge base categories, the system can also recommend relevant upper-level tags based on the hierarchical structure and association relationships in the tag system according to the content of the query request. For example, if the query request is about "wind power generation technology", the system can recommend upper-level tags such as "renewable energy" and "power generation technology".

[0193] To improve the accuracy of the recommendations, the system can also consider information such as the user's historical query records and preference settings. The system displays the recommended knowledge base categories and upper-level tags to the user for selection; the user can select relevant knowledge base categories and upper-level tags according to their own needs to further enrich the query request.

[0194] The user's selection and feedback can also be used to optimize the system's recommendation algorithm and improve the accuracy of subsequent recommendations; according to the knowledge base categories and upper-level tags selected by the user, the system can further enrich the content of the query request; for example, the system can automatically add keywords, phrases, or query conditions related to the selected categories and tags to improve the accuracy and comprehensiveness of the query.

[0195] In other embodiments, in the knowledge retrieval step, it further includes: dynamically adjusting the weights of upper-level tags and knowledge associations according to the user's retrieval history, preferences, and behavior patterns to optimize the knowledge recommendation ranking.

[0196] Specifically, the system needs to collect the user's retrieval history records, including the query keywords entered by the user, the search results clicked, the browsing time, etc.; the user can set their own preferences in the system, such as the topics, fields, and knowledge types of interest; the system analyzes indicators such as the user's click behavior, browsing time, and return rate to infer the user's behavior patterns and potential needs; the system calculates the association degree between each upper-level tag and the user's query keywords, historical retrieval records, and preference settings; according to the calculation results of the association degree, the system dynamically adjusts the weight of each upper-level tag; the higher the association degree of the tag, the greater its weight, and thus it occupies a more important position in knowledge recommendation.

[0197] The system calculates the association degree between each knowledge item and the user's query keywords, historical retrieval records, preference settings, and upper-level tags; similarly, according to the calculation results of the association degree, the system dynamically adjusts the weight of each knowledge item; the higher the association degree of the knowledge item, the greater its weight, and the more likely it is to be recommended to the user.

[0198] Use sorting algorithms (such as PageRank, TF-IDF, etc.), combined with the upper-level tags and knowledge association weights after dynamic adjustment, to sort the knowledge items; according to the sorting results, the system generates a personalized knowledge recommendation list for the user; the knowledge items in the recommendation list are arranged in descending order of weight to ensure that the user can first see the knowledge items that best meet their needs.

[0199] It is necessary to collect user feedback on the recommendation results, such as satisfaction, comments, etc.; according to the user feedback, the system adjusts the weights of the upper-level tags and knowledge associations, as well as the parameters of the sorting algorithm in real time to continuously optimize the effect of knowledge recommendation.

[0200] Refer to Figure 2 and Figure 3 , in other embodiments, a deviation acquisition step is further provided after the query requirement acquisition step;

[0201] Deviation acquisition: includes a requirement similarity calculation step, a third judgment step, a difference statistics step, a difference acquisition step, a difference judgment step, a correction step, a fourth judgment step, a supplementary calculation step, a supplementary judgment step, and a supplementary label step;

[0202] Requirement similarity calculation: Calculate multiple query requests of the same query requester, and calculate the similarity of the previous and subsequent query requests;

[0203] Third judgment: Judge whether the similarity is greater than a preset third threshold. If so, execute the difference statistics step; otherwise, execute the knowledge retrieval step;

[0204] Difference statistics: Count the number of difference similarities, and judge whether it is greater than a fourth threshold. If so, execute the difference judgment step;

[0205] Difference acquisition: Extract the words modified multiple times, denoted as difference words, obtain the common points between multiple difference words, denoted as semantic words, calculate the correlation degree between the semantic words and the upper-level tags, and then execute the difference judgment step;

[0206] Difference judgment: Judge whether the correlation degree is less than a preset fifth threshold. If so, execute the correction step; otherwise, execute the fourth judgment step;

[0207] Correction: Intervene in the labeling step and modify and update the upper-level tags;

[0208] Fourth judgment: Judge whether the number of difference similarities is greater than a sixth threshold. If so, execute the supplementary calculation step;

[0209] Supplementary calculation: Calculate the similarity between the upper-level tags and denote it as the synonymous similarity;

[0210] Supplemental Judgment: Determine whether the synonym similarity is greater than the seventh threshold. If so, perform the supplemental tagging step;

[0211] Supplemental Tagging: Pop up multiple upper-level tags for the query requester to make additional selections again to correct the query request.

[0212] Specifically, first, collect all query request records of the same query requester within a certain period of time. Perform text preprocessing on each query request, using NLP libraries (such as NLTK, spaCy, Porter Stemmer) to remove stop words, punctuation marks, perform stemming or lemmatization, etc.; use models such as TF-IDF, word vectors (such as Word2Vec), or BERT to extract the feature vectors of the query requests; calculate the cosine similarity or other similarity metrics (Euclidean distance, Manhattan distance, Jaccard similarity) between adjacent query request feature vectors as their similarity; preset a third threshold, which represents the standard for query requests to be similar enough and is determined based on historical data and experiments. In this embodiment, the value of the third threshold is 0.75; if the calculated similarity is greater than or equal to the third threshold, it is considered that there is a significant similarity between the query requests and further analysis of the differences is required; otherwise, it is considered that the query requests have changed significantly and the knowledge retrieval step is directly executed.

[0213] For similar query requests, record the differences between them (such as keyword changes, changes in query intent, etc.); count the number of times differences occur in similar query requests, that is, the difference similarity count; preset a fourth threshold, which is determined based on historical data and experiments. In this embodiment, the fourth threshold is preferably 5 times; if the difference similarity count is greater than or equal to the fourth threshold, perform the difference acquisition step; otherwise, it may indicate that the change in the query request is not significant enough to attract attention and can continue to be observed or other processing can be performed.

[0214] Extract the modified or replaced words from the differences as the difference words; analyze the difference words to find the commonalities or patterns between them, and these commonalities may reflect the core or essential words of the query request; calculate the correlation between the essential words and the current upper-level tag (i.e., the topic or category tag to which the query request is classified), and a semantic-based similarity calculation method can be used; preset a fifth threshold, which represents the lowest acceptable level of the correlation between the essential words and the upper-level tag, and the fifth threshold is determined based on historical data, experiments, or user feedback. In this embodiment, the value of the fifth threshold is 0.8; if the correlation is less than the fifth threshold, it is considered that the current upper-level tag may be inaccurate and needs to be corrected; otherwise, perform the fourth judgment step.

[0215] According to the results of the difference analysis, manually or automatically adjust the upper-level tags of the query request to more accurately reflect the query intention; save the updated upper-level tags to the database for subsequent queries; preset a sixth threshold, where the sixth threshold is determined based on historical data, experiments, or user feedback. In this embodiment, the value of the sixth threshold is 7 times; it is used to determine whether the number of similar differences has reached the level that requires more in-depth analysis. If the number of similar differences is greater than or equal to the sixth threshold, perform supplementary calculation steps; otherwise, it may indicate that the changes in the query request are still within the acceptable range and can continue to be observed.

[0216] Calculate the similarity between the current upper-level tag and other relevant or potential upper-level tags, that is, the synonym similarity; you can use a lexical-based similarity calculation method (such as Jaccard similarity, cosine similarity) or a semantic-based similarity calculation method (such as WordNet, semantic vector space model); preset a seventh threshold, indicating the lowest acceptable level of synonym similarity, where the seventh threshold is determined based on historical data and experiments. In this embodiment, the value of the seventh threshold is 0.75; if the synonym similarity is greater than or equal to the seventh threshold, it is considered that there are multiple upper-level tags that may be applicable to the current query request, and supplementary tag steps are required; otherwise, it may indicate that the current upper-level tag is already accurate enough and no further supplementation is needed.

[0217] Display multiple possible upper-level tags to the query requester in the form of a pop-up window for their selection or confirmation; allow the query requester to make supplementary selections or modifications based on the tags in the pop-up window to further refine the query intention; update the upper-level tags of the query request according to the user's selection and save them to the database.

[0218] For example, user A entered the following query requests in the system:

[0219] First query: "What are the methods of electric load forecasting?";

[0220] Second query (later): "What are the models of electric load forecasting?";

[0221] Third query (even later): "What is the algorithm principle of the electric load forecasting model?";

[0222] The system first preprocesses the query requests using an NLP library (such as spaCy) to extract feature vectors. Then, calculate the cosine similarity between adjacent query requests:

[0223] Similarity between the first and second queries: 0.85;

[0224] Similarity between the second and third queries: 0.78;

[0225] The preset third threshold is 0.75. Since the similarity between the first and second queries (0.85) is greater than the third threshold, the system believes there is a significant similarity between them and needs to enter the difference statistics step.

[0226] The system records the different words between these two queries: "method" and "model". The number of similar differences is 1 (because there is only one significant difference).

[0227] The system extracts the different words "method" and "model" and calculates their association degrees with the current upper-level label "power load forecasting". Since both "method" and "model" are core words in the field of power load forecasting and have a high association degree with the "power load forecasting" label (assumed to be 0.9 and 0.95), they do not meet the correction conditions (that is, the association degree is less than the fifth threshold of 0.8).

[0228] However, considering that the user may be gradually exploring different aspects of power load forecasting in depth, continue to observe subsequent queries.

[0229] When user A makes the third query, the system notices that the number of similar differences has increased to 2 (although in this example it has not reached the sixth threshold of 7 times, but to show the complete process, we assume that there will be more similar queries in the future so that the number of similar differences reaches 7 times).

[0230] At this time, the system calculates the synonymous similarity between the current upper-level label "power load forecasting" and other potential upper-level labels (such as "power market trading", "power system operation"); assume that the synonymous similarity with "power market trading" is 0.65 (lower than the seventh threshold of 0.75), and the synonymous similarity with "power system operation" is 0.5 (also lower than the seventh threshold).

[0231] Since there is no other upper-level label with a high enough synonymous similarity with "power load forecasting", the system does not perform the supplementary label step temporarily. However, to cope with more similar queries that may occur in the future, the system decides to optimize the label system and considers introducing more detailed labels (such as "power load forecasting method", "power load forecasting model", etc.) at an appropriate time.

[0232] Pop-up display and user selection (assuming that the supplementary tag condition is met in subsequent queries): Assume that in subsequent queries, the number of similar differences continues to increase, and the system discovers that tags such as "power load forecasting model" and "power load forecasting algorithm" have a high synonymous similarity. At this time, the system will display these possible upper-level tags to user A in the form of a pop-up window: Pop-up window title: "Please select or confirm your query topic"; Tag options: "Power load forecasting", "Power load forecasting model", "Power load forecasting algorithm"; User A can select or confirm a more specific tag according to their own needs to further refine the query intention. The system updates the upper-level tags of the query request according to the user's selection and saves them to the database.

[0233] Refer to Figure 4 , in other embodiments, a user participation step, an artificial correction step, an associated tag matching step, a feedback statistics step, and an update step are further provided between the difference determination step and the correction step;

[0234] User participation: Recommend real sense words to the query requester, determine whether the real sense words meet the needs of the query requester. If so, execute the correction step and the associated tag matching step; otherwise, execute the artificial correction step;

[0235] Artificial correction: The query requester corrects the real sense words, and then executes the associated tag matching step;

[0236] Associated tag matching: Match relevant upper-level tags through real sense words and push them to the query requester. According to the feedback of the query requester, execute the feedback statistics step;

[0237] Feedback statistics: Obtain the number of times the association degree between a certain real sense word and the upper-level tag is corrected, and determine whether it is greater than the eighth threshold. If so, execute the update step;

[0238] Update: Update the association degree between the upper-level tag and the real sense word for subsequent matching of query requests.

[0239] Specifically, the system recommends the extracted content words to the query requester and asks whether these content words accurately reflect their query intent; if the query requester confirms that these content words meet the requirements, the system executes the associated tag matching step; if the query requester believes that these content words are inaccurate or do not fully meet the requirements, it enters the manual correction step; the query requester corrects the content words recommended by the system, adding or deleting words to more accurately express their query intent; the system receives the corrected content words and re-executes the associated tag matching step; the system uses an algorithm or model to match relevant upper-level tags based on the corrected content words; the system pushes the matched upper-level tags to the query requester and asks about their satisfaction; according to the feedback of the query requester, the system executes the feedback statistics step, and the system records the number of times the association degree between each content word and the upper-level tag is corrected.

[0240] The system determines whether the number of times the association degree between a certain content word and the upper-level tag is greater than the eighth threshold, where the determination of the eighth threshold is based on historical data and experiments, and the value of the eighth threshold in this embodiment is 3 times; if it is greater than the eighth threshold, it indicates that the association degree between the content word and the upper-level tag is not stable or accurate enough and needs to be updated, then the update step is executed; if it is not greater than the eighth threshold, the system continues to process the next query request; the system updates the association degree between the upper-level tag and the content word according to the feedback statistics result, and the updated association degree is used for subsequent query request matching to improve the accuracy and efficiency of the matching.

[0241] For example, user B enters the following query request in the system: "What are the influencing factors of power load forecasting?" The system first calculates the similarity between this query and previous query requests and finds no significant similarity, so it directly enters the knowledge retrieval step. However, during the knowledge retrieval process, the system extracts the content words "power load forecasting" and "influencing factors".

[0242] It is judged whether these content words meet the query requirements, but due to the lack of sufficient context information, the system cannot determine whether these words are completely accurate. Therefore, it is decided to enter the user participation step; the system recommends the extracted content words "power load forecasting" and "influencing factors" to user B and asks whether these words accurately reflect their query intent. After seeing this, user B believes that "power load forecasting" is accurate, but "influencing factors" may not be specific enough. What he hopes to know more is "which factors will affect the accuracy of power load forecasting". Therefore, user B chooses to correct the content words and changes "influencing factors" to "factors that affect the accuracy of power load forecasting". The system receives the corrected content words and is ready to perform the associated tag matching.

[0243] The system uses algorithms or models to match relevant upper-level tags based on the revised real semantic terms "electric load forecasting" and "factors affecting the accuracy of electric load forecasting". The system has found upper-level tags associated with these real semantic terms, such as "electric load forecasting technology", "electric power market analysis", etc. (Note: This is just an example, and the actual matching may be more precise).

[0244] The system pushes the matched upper-level tags to User B and asks for their satisfaction. After seeing them, User B thinks that the tag "electric load forecasting technology" is relatively close, but hopes more for a dedicated tag about "factors affecting the accuracy of electric load forecasting". Therefore, User B gives feedback and suggests that the system add or optimize relevant tags.

[0245] The system records this feedback and counts the number of times the association degree between the real semantic term "factors affecting the accuracy of electric load forecasting" and the upper-level tags is revised. Since this is the first revision, the number of revisions is 1, and it has not reached the eighth threshold (assumed to be 3 times).

[0246] Update step (assuming the update condition is met in subsequent queries): In subsequent queries, if User B or other users revise the association degree between the real semantic term "factors affecting the accuracy of electric load forecasting" and the upper-level tags multiple times (for example, the number of revisions reaches 3 or more), the system will determine that the association degree between this real semantic term and the upper-level tags is not stable or accurate enough and needs to be updated.

[0247] The system updates the association degree between the upper-level tags and the real semantic terms according to the feedback statistics results. For example, the system can add a new upper-level tag "factors affecting the accuracy of electric load forecasting" and associate it with the real semantic term "factors affecting the accuracy of electric load forecasting". The updated association degree will be used for subsequent query request matching to improve the accuracy and efficiency of matching.

[0248] Refer to Figure 5 , in other embodiments, after the associated tag matching step, a requirement judgment step and a single association step are also provided;

[0249] Requirement judgment: Judge whether the generated tags meet the needs of the query requester. If so, execute the feedback statistics step; otherwise, execute the single association step;

[0250] Single association: The query requester inputs a custom tag by themselves, matches it from the upper-level tags, and establishes a single association for the query request of this query requester, and then executes the knowledge retrieval step.

[0251] Specifically, first, the system receives a query request from the query requester and analyzes the query content through natural language processing technology to generate a series of initial tags. Then, the system uses a pre-trained tag association model to match these initial tags with the upper-level tags in the knowledge base and find the most relevant set of upper-level tags. After the associated tag matching is completed, the system enters the requirement judgment step. The purpose of this step is to determine whether the generated set of tags accurately reflects the actual requirements of the query requester. The system can perform requirement judgment in the following ways: Based on user feedback: The system displays the matched tags to the query requester and asks if they are satisfied. The user can provide feedback by clicking "Yes" or "No". The system uses machine learning algorithms, combined with the query requester's historical query records and the context information of the current query, to automatically judge the accuracy of the tags. The system presets a series of rules, such as tag quantity thresholds, tag relevance scores, etc., and judges whether the set of tags meets the requirements according to these rules.

[0252] If the system determines that the tags meet the requirements, it enters the manual input step, allowing the query requester to further refine or correct the tags. If the requirements are not met, the system returns to the tag generation step to regenerate or adjust the tags.

[0253] After the requirement judgment step, if the query requester chooses to enter custom tags manually, the system enters the single association step. This step allows the query requester to enter specific custom tags according to their own needs. The system matches these custom tags with the upper-level tags in the knowledge base, finds the most suitable upper-level tags, and establishes a single association. Single association means that these custom tags are only used for this query request and will not be permanently stored in the system's tag association model.

[0254] After establishing the single association, the system retrieves relevant information from the knowledge base based on the associated upper-level tags and custom tags. The retrieval results are sorted according to the relevance scores and displayed to the query requester. The query requester can view detailed information, perform further filtering, or submit a new query request as needed.

[0255] For example, the query requester wants to query information related to "the application of smart grid technology in power distribution". The system first analyzes the query content and generates initial tags such as "smart grid technology", "power distribution", and "application". Then, the system uses a pre-trained tag association model to match these initial tags with the upper-level tags in the knowledge base and finds relevant upper-level tags such as "energy technology" and "power system management".

[0256] In the requirement judgment step, the system presents the matched tags to the query requester and asks if they are satisfied. The query requester finds that "smart grid technology" and "power distribution" are accurate, but the tag "application" is also too broad and not specific enough. Therefore, the query requester chooses to enter the single association step and inputs the custom tags "energy saving effect" and "fault prediction" according to their own needs.

[0257] The system matches these custom tags with the upper-level tags in the knowledge base and finds the most suitable upper-level tags "energy efficiency improvement" and "preventive maintenance". Then, based on these associated upper-level tags and custom tags, the system retrieves relevant information from the knowledge base, such as how smart grid technology improves energy efficiency in power distribution and how to predict faults through smart grid technology, and sorts and displays it to the query requester according to the relevance score.

[0258] Refer to Figure 5 , in other embodiments, after the single association step, an association statistics step is also set;

[0259] Association statistics: Count the number of times of single association established with the same upper-level tag, and judge whether it is greater than the ninth threshold. If so, execute the tag conversion step;

[0260] Tag conversion: Convert the custom tag into an upper-level tag and establish an association with the knowledge base.

[0261] Specifically, in the single association step, the system allows users to input custom tags according to the query requirements, matches these tags with the upper-level tags in the knowledge base, and establishes a single association. This association is only applicable to the current query and will not permanently change the system's tag system.

[0262] After the single association step, the system enters the association statistics step. The system maintains an association statistics table to record the number of associations between each upper-level tag and the custom tag, traverses the single association records, and counts the number of associations between each upper-level tag and the custom tag; Set a ninth threshold to determine whether the association between a certain upper-level tag and the custom tag is frequent enough to perform tag conversion, where the ninth threshold is determined based on historical data or experiments.

[0263] When the number of associations between a certain upper-level tag and a certain custom tag exceeds the ninth threshold, the system enters the tag conversion step. The purpose of this step is to convert the frequently used custom tag into an upper-level tag and establish a permanent association with the knowledge base, thereby optimizing the tag system and improving the retrieval efficiency.

[0264] In other embodiments, after the knowledge retrieval step, a recommendation statistics step is also set;

[0265] Recommendation statistics: Obtain the click-through rate of the recommended knowledge for the query requester, determine the matching degree, and correct the sorting of subsequent similar query requirements.

[0266] Specifically, record the click behavior of users on the recommended knowledge, including the number of clicks, click time, etc., calculate the click-through rate of each recommended knowledge, that is, the ratio of the number of clicks to the number of displays, and determine the matching degree between the recommended knowledge and the user's query requirements according to the click-through rate. The higher the click-through rate, the higher the matching degree; store the matching degree information for subsequent correction of the sorting of similar query requirements; based on the results of the recommendation statistics, the system corrects the sorting algorithm for subsequent similar query requirements, and the basis for the correction is the click-through rate of the user on the recommended knowledge before, that is, the matching degree; when the system receives a new similar query request, adjust the sorting of the retrieval results according to the previous recommendation statistics results; for the knowledge items with a high matching degree, give a higher sorting priority to make them easier to be discovered by users; for the knowledge items with a low matching degree, reduce their sorting priority or exclude them from the retrieval results.

[0267] The above are all preferred embodiments of this application. The protection scope of this application is not limited by this. Therefore, all equivalent changes made according to the structure, shape, and principle of this application should be covered within the protection scope of this application.

Claims

1. A method for constructing a knowledge service system, characterized in that: Including the following steps: Knowledge acquisition: Collect knowledge data from knowledge sources in multiple fields; Preprocessing: Preprocess the knowledge data to eliminate noise; Classification: Classify the acquired knowledge data, generate corresponding primary labels for each category, and store the classified data separately to form a knowledge base; Similarity calculation: Calculate the similarity between knowledge bases of different categories; First judgment: Judge whether the similarity is greater than a first threshold. If so, execute the association step; Association: Establish associations between knowledge according to the similarity; Second judgment: Judge whether the similarity is greater than a second threshold. If so, execute the labeling step; Labeling: Analyze their common features, construct new upper-level labels to summarize and reflect the common attributes of this knowledge; Label directory generation: Establish a label directory based on the knowledge base category, primary label, and upper-level label; Query requirement acquisition: Obtain the user's knowledge retrieval requirement, determine keywords or labels according to the retrieval requirement, and generate a query request; Knowledge retrieval: According to the query request, combined with the upper-level label and knowledge association, screen and sort the results to perform knowledge recommendation.

2. The method for constructing a knowledge service system according to claim 1, wherein: Among them, the classification step also includes: Perform multi-dimensional classification based on the theme, content, and form features of knowledge, and assign one or more labels to the knowledge under each dimension.

3. The method for constructing a knowledge service system according to claim 1, wherein: The labeling step also includes: Analyze the common points and differences of similar knowledge bases, as well as their positions and roles in the knowledge system, to determine the naming and hierarchical structure of the upper-level label.

4. The method for constructing a knowledge service system according to claim 1, wherein: After the query requirement acquisition step, a label circumscription step is also set; Label circumscription: According to the query request, initially recommend the knowledge base category and upper-level label for the querier to select, and enrich the query request.

5. The method for constructing a knowledge service system according to claim 1, wherein: In the knowledge retrieval step, it also includes: Dynamically adjust the weights of the upper-level label and knowledge association according to the user's retrieval history, preferences, and behavior patterns, and optimize the knowledge recommendation sorting.

6. The method for constructing a knowledge service system according to any one of claims 1-5, characterized in that: After the query requirement acquisition step, a deviation acquisition step is also set; Deviation acquisition: Includes a demand similarity calculation step, a third judgment step, a difference statistics step, a difference acquisition step, a difference judgment step, a correction step, a fourth judgment step, a supplementary calculation step, a supplementary judgment step, and a supplementary label step; Demand similarity calculation: Calculate multiple query requests of the same query requester, and calculate the similarity of the previous and subsequent query requests; Third judgment: Judge whether the similarity is greater than a preset third threshold. If so, execute the difference statistics step, otherwise, execute the knowledge retrieval step; Difference statistics: Statistically count the number of difference similarities, and judge whether it is greater than a fourth threshold. If so, execute the difference judgment step; Difference acquisition: Extract the repeatedly modified words, record them as difference words, obtain the common points between multiple difference words, record them as semantic words, calculate the association degree between the semantic words and the upper-level label, and then execute the difference judgment step; Difference judgment: Judge whether the association degree is less than a preset fifth threshold. If so, execute the correction step, otherwise, execute the fourth judgment step; Correction: Intervene in the labeling step and modify and update the upper-level label; Fourth judgment: Judge whether the number of difference similarities is greater than a sixth threshold. If so, execute the supplementary calculation step; Supplemental calculation: Calculate the similarity between upper-level tags and record it as the synonym similarity; Supplemental judgment: Judge whether the synonym similarity is greater than the seventh threshold. If so, execute the supplemental tag step; Supplemental tag: Set pop-ups for multiple upper-level tags for the query requester to make additional selections again to correct the query request.

7. The method for constructing a knowledge service system according to claim 6, wherein: A user participation step, a manual correction step, a related tag matching step, a feedback statistics step, and an update step are also set between the difference judgment step and the correction step; User participation: Recommend content words to the query requester, judge whether the content words meet the needs of the query requester. If so, execute the correction step and the related tag matching step; Otherwise, execute the manual correction step; Manual correction: The query requester corrects the content words and then executes the related tag matching step; Related tag matching: Match relevant upper-level tags through the content words and push them to the query requester. According to the feedback of the query requester, execute the feedback statistics step; Feedback statistics: Obtain the number of times the association degree between a certain content word and the upper-level tag is corrected, judge whether it is greater than the eighth threshold. If so, execute the update step; Update: Update the association degree between the upper-level tag and the content word for subsequent query request matching.

8. The method for constructing a knowledge service system according to claim 7, wherein: A requirement judgment step and a single association step are also set after the related tag matching step; Requirement judgment: Judge whether the generated tag meets the needs of the query requester. If so, execute the feedback statistics step. Otherwise, execute the single association step; Single association: The query requester inputs a custom tag by himself, matches it from the upper-level tags, and establishes a single association for the query request of this query requester, and then executes the knowledge retrieval step.

9. The method for constructing a knowledge service system according to claim 8, wherein: An association statistics step is also set after the single association step; Association statistics: Count the number of times a single association is established for the same upper-level tag, judge whether it is greater than the ninth threshold. If so, execute the tag conversion step; Tag conversion: Convert the custom tag into an upper-level tag and establish an association with the knowledge base.

10. The method for constructing a knowledge service system according to claim 1, wherein: A recommendation statistics step is also set after the knowledge retrieval step; Recommendation statistics: Obtain the click-through rate of the knowledge recommended by the query requester, determine the matching degree, and correct the sorting of subsequent similar query requirements.

Citation Information

Patent Citations

  • Text data processing method and device

    CN114090777A

  • Method for establishing offshore oil and gas exploration and development data label system

    CN118861321A