A chronic disease dietary knowledge graph driven intelligent question and answer method

CN122838534APending Publication Date: 2026-09-29JIANGSU PROVINCE HOSPITAL (THE FIRST AFFILIATED HOSPITAL OF NANJING MEDICAL UNIVERSITY)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610886361.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0006]本发明针对现有慢病饮食知识图谱中知识条目可靠性低、时效性差、与临床指南一致性难以保障的问题,提供一种能够动态更新、证据驱动、一致性可量化的智能问答方法

Benefits of technology

[0108]以上述依据本发明的实施例为启示,通过上述的说明内容,相关工作人员完全可以在不偏离本项发明技术思想的范围内,进行多样的变更以及修改。本项发明的技术性范围并不局限于说明书上的内容,必须要根据权利要求范围来确定其技术性范围。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838534A_ABST
    Figure CN122838534A_ABST
Patent Text Reader

Abstract

The application discloses a kind of chronic disease diet knowledge graph driven intelligent question and answer method, it is related to diet knowledge graph technical field, including according to the evidence grade annotation in original data, using analytic hierarchy process to knowledge entry priority ordering, pairwise comparison value is jointly assigned by evidence grade quantization function and timeliness factor;Through fusing clinical guideline consistency graph embedding similarity formula calculation node association;Extract inconsistent part and use random forest classification bias type;Fusion external verification source data, using weighted penalty formula calculation consistency score;If score is lower than threshold value, through L1 regularization objective function iterative adjustment node attribute, obtain optimized graph version;Generate diet suggestion query interface, and use probability formula quantization real-time response performance.The application solves the problems of missing evidence level management, poor consistency with clinical guidelines and update lag in knowledge graph, and improves the authority and timeliness of intelligent question and answer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of chronic disease dietary knowledge graph technology, and in particular to an intelligent question-answering method driven by chronic disease dietary knowledge graph. Background Technology

[0002] Dietary management for patients with chronic diseases is directly related to disease control and quality of life improvement, holding crucial significance in the healthcare field. With the continuous rise in chronic disease incidence, knowledge graph-driven intelligent question-answering systems have become important tools supporting patients' daily dietary decisions. These systems can provide personalized food choices and nutritional advice based on the patient's condition. However, existing methods often face challenges in practical applications, such as ensuring the reliability and timeliness of the knowledge content. This can lead to dietary guidance that is inconsistent with the latest medical evidence or contains biases in key recommendations.

[0003] The core content of a knowledge graph, such as dietary restrictions, recommended daily intake of nutrients, and food exchange portions, requires strict reliance on medical evidence. Without effective management of evidence levels, knowledge entries are susceptible to the influence of inconsistent source quality, leading to inconsistencies in recommendations for the same food or nutrient in different patient scenarios. Further complicating matters, clinical guidelines and research findings are constantly being updated. If existing knowledge is not adjusted promptly after new evidence is released, the knowledge graph content will gradually deviate from actual medical consensus. This intertwining of update lag and difficulty in evidence traceability makes it difficult for a knowledge graph to maintain its authority in the long run.

[0004] Especially during the construction and maintenance process, the consistency verification of dietary recommendations involves the fusion and comparison of multi-source data. Without a systematic review mechanism, the accuracy of key knowledge items cannot be guaranteed. For example, when a new study adjusts the intake limits of a certain type of food for diabetic patients, if the contraindications in the original atlas are not verified one by one by medical experts and updated accordingly, the system may continue to output outdated recommendations, posing potential risks to patients.

[0005] How to establish strict control over the level of medical evidence, the traceability of data sources, and consistency with clinical guidelines throughout the entire process of knowledge graph construction and updating, while ensuring that key knowledge can be adjusted in a timely manner with the latest research results, has become a key issue in ensuring the authority and timeliness of the dietary guidance output by the intelligent question-and-answer system. Summary of the Invention

[0006] This invention addresses the problems of low reliability, poor timeliness, and difficulty in ensuring consistency with clinical guidelines in existing chronic disease dietary knowledge graphs by providing an intelligent question-answering method that is dynamically updated, evidence-driven, and whose consistency can be quantified.

[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: A knowledge graph-driven intelligent question-answering method for chronic disease diets includes the following steps: Step S101: Extract the latest clinical guidelines and research findings from the medical database using a pre-defined evidence collection module to obtain an original dataset containing evidence level labels. Step S102: Based on the evidence level labels in the original dataset, use the analytic hierarchy process (AHP) to prioritize each knowledge item and determine a list of high-priority items. Step S103: Obtain details of items in the high-priority list. If the evidence level is higher than a preset threshold, trigger an update mechanism to identify the knowledge graph nodes that need updating. Step S104: For the knowledge graph nodes that need updating, calculate the similarity between nodes using a graph embedding algorithm to obtain a set of associations consistent with clinical guidelines. Step S105: Extract inconsistencies from the association set and classify them using a random forest algorithm to determine deviation type labels. Step S106: Based on the deviation type labels, obtain the data fusion results from external validation sources and determine the consistency score after fusion. Step S107: If the consistency score is lower than a preset threshold, iteratively adjust node attributes to obtain an optimized version of the knowledge graph. Step S108: Generate a dietary advice query interface through an optimized knowledge graph version to obtain a real-time output set for patient decision-making.

[0008] As a further preferred embodiment of the intelligent question-answering method driven by a knowledge graph of chronic disease diet according to the present invention, the step of extracting the latest clinical guidelines and research results from a medical database through a preset evidence collection module to obtain a raw data set containing evidence level labels includes: obtaining the latest clinical guidelines and research results from a medical database through a preset module to construct a raw data set containing evidence level labels; scanning the medical database using an automated crawling tool to obtain the latest information related to clinical guidelines and research results; if the latest information obtained meets the preset evidence level standards, it is classified as valid data to obtain a preliminary filtered data set; for the preliminary filtered data set, text parsing technology is applied to extract the evidence level information to determine the structured data with level labels; according to the level labels in the structured data, clinical guidelines and research results of different evidence levels are classified and organized to obtain hierarchically organized data groups; through the hierarchically organized data groups, a raw data set containing evidence level labels is constructed, and its completeness and consistency are judged; after obtaining the raw data set, it is saved to a designated database using a storage management tool to obtain the final data resources available for subsequent analysis.

[0009] As a further preferred embodiment of the intelligent question-answering method driven by a knowledge graph of chronic disease diet according to the present invention, the step of prioritizing each knowledge item and determining a list of high-priority items by using the analytic hierarchy process (AHP) based on the evidence level labels in the original dataset includes: constructing pairwise comparison relationships for each knowledge item based on the evidence level labels in the original dataset; constructing a judgment matrix using the AHP, filling in the pairwise comparison values ​​to obtain an initial judgment matrix; calculating the feature vector through the judgment matrix, and obtaining the weight value of each knowledge item after normalization. After obtaining the weight values, the largest eigenvalue of the judgment matrix is ​​calculated, and the consistency index is determined in combination with the matrix order. If the ratio of the consistency index to the random consistency index is less than the preset threshold, the consistency test is passed and the weight values ​​are retained. The weight values ​​that pass the consistency test are sorted and all knowledge item weight values ​​are arranged in descending order. Based on the sorted weight value sequence, the knowledge items at the top are extracted to determine the list of high-priority items. The pairwise comparison values ​​are assigned based on the following formula to determine the level of evidence and timeliness: In the formula, Representing knowledge entries Relative to the entry Importance ratio; , Each item and entries The level of evidence quantification value, through Calculation, where For the original level of evidence, As a preset baseline level, Sensitivity coefficient; , The year the entry was published; The maximum time span among all entries; This is a timeliness adjustment coefficient; In the consistency test of the judgment matrix of the analytic hierarchy process, if the consistency ratio... The test passes; when the number of knowledge items... At that time, hierarchical clustering was used to group items according to disease categories. Pairwise comparisons were performed within each group, and comparisons were made between groups using class centroids. The computational complexity was reduced from... Down to ,in This represents the number of categories.

[0010] As a further preferred embodiment of the intelligent question-answering method driven by a knowledge graph for chronic disease diet according to the present invention, the step of obtaining the details of items in the high-priority item list, and triggering an update mechanism if the evidence level is higher than a preset threshold to determine the knowledge graph nodes that need to be updated, includes: obtaining detailed information of high-priority items from the item list, performing preliminary screening based on the evidence level of each item to obtain a set of items that meet the conditions; for the selected set of items, if the evidence level is higher than a preset threshold, triggering an update process to determine the range of knowledge graph nodes to be processed; based on the determined range of knowledge graph nodes, obtaining the current associated data of the nodes, determining whether there are inconsistent link relationships, and obtaining a list of nodes to be adjusted; for the list of nodes to be adjusted, using pre-established mapping rules to update the link relationships between nodes, and determining the updated node structure; through the updated node structure, obtaining detailed information of associated items, determining whether they meet the priority level requirements, and obtaining the final adjustment result; based on the final adjustment result, using a batch processing tool to synchronize the adjusted node data to the knowledge graph, completing the entire update process.

[0011] As a further preferred embodiment of the intelligent question-answering method driven by a knowledge graph of chronic disease diet according to the present invention, the step of calculating the similarity between nodes through a graph embedding algorithm to obtain a set of association relationships consistent with clinical guidelines for the knowledge graph nodes that need to be updated includes: For graph nodes in the knowledge graph, a graph embedding algorithm is used to analyze node similarity, obtaining preliminary node association data and obtaining node similarity results related to clinical guidelines. Based on the preliminary node similarity results, the node association data is filtered to obtain a set of associations that conform to guideline consistency, determining the priority range of nodes to be processed. According to the determined node range, detailed node information related to clinical guidelines is extracted from the knowledge graph, and potential connections between nodes are determined, resulting in a list of relationships to be processed. For the list of relationships to be processed, a pre-established rule base is used for comparison; if the relationship conforms to guideline consistency requirements, it is retained, obtaining the final node association structure. Based on the final node association structure, relevant node data is obtained from the knowledge graph, and the logical connections between the data are organized to determine the complete set of associations. Based on the organized set of associations, a batch processing tool is used to synchronize the node association data to the knowledge graph, completing the node similarity and clinical guideline consistency update process. For the synchronized knowledge graph data, the updated node association information is obtained, and the data integrity is checked using an internal verification tool to obtain the final processing result. The graph embedding algorithm uses the following similarity formula, which incorporates consistency with clinical guidelines, to calculate the similarity between nodes: In the formula, For nodes With nodes Overall similarity between them; , They are nodes and Graph embedding vectors; Indicates cosine similarity; For nodes and The set of existing relationship types between them; This is the set of standard relation types defined in clinical guidelines; This is a balancing factor used to adjust the weights of structural similarity and guideline consistency. Among them, when the degree of the knowledge graph node to be updated is lower than a preset threshold In option 3, semantic distance based on medical ontology is used as the initial similarity: ,in For nodes and The shortest path length in SNOMED CT or UMLS ontology; after the node degree is restored, switch to the graph embedding similarity formula.

[0012] As a further preferred embodiment of the intelligent question-answering method driven by a knowledge graph for chronic disease diets according to the present invention, the method involves extracting inconsistent parts from the set of association relationships and classifying these inconsistent parts using a random forest algorithm. The random forest algorithm employs an active learning strategy: calculating the prediction entropy of each decision tree for the same relation sample; if the entropy value is greater than a preset threshold, the sample is pushed to manual review, and the review results are added as new training data. The random forest is retrained after every 1000 samples. Determining the deviation type label includes: extracting association relationship data from the knowledge graph; performing preliminary screening on parts with low consistency to obtain a set of association relationships that do not meet preset standards; and using the random forest algorithm to classify the deviation type for the selected set of association relationships. The process involves classifying and assigning deviation type labels to each relationship. Based on these labels, information is compared against anomaly markers to identify category identifiers that do not match the pre-established rule base. Using the comparison results, nodes are grouped and data is layered to obtain a layered set of relationship data. For this layered set, relationships are filtered using information comparison; if a relationship's category identifier matches the rule base, that relationship is retained, resulting in a filtered list of relationships. Based on this list, labels are assigned to the layered data to determine the final classification structure. Finally, using this final structure, the remaining relationship data is updated and stored to obtain a complete record of deviation type processing.

[0013] As a further preferred embodiment of the intelligent question-answering method driven by a knowledge graph of chronic disease diet according to the present invention, the step of obtaining the data fusion result of the external verification source based on the deviation type label and judging the consistency score after fusion includes: External validation sources are categorized and filtered using deviation type labels; raw data is extracted from the corresponding external validation sources based on the categorization and filtering results, and a weighted average method is used to perform fusion processing on the extracted raw data to obtain the fusion result; the numerical distribution of each dimension is extracted from the fusion result; the absolute value of the deviation from the standard reference value is calculated for the numerical distribution of each dimension; if the absolute value of the deviation is less than a preset threshold, the dimension is determined to be consistent; otherwise, it is determined to be inconsistent; the final consistency score is determined based on the proportion of consistent dimensions to the total number of dimensions. The consistency score is calculated using a weighted penalty formula: In the formula, For final consistency score; This represents the total number of dimensions. For the first The absolute value of the dimensional deviation; For the first Dimensional deviation threshold; For the first The weights of the dimensions are obtained using the analytic hierarchy process (AHP). This is an indicator function that takes the value 1 when the condition is true and 0 otherwise. Before fusion, Laplace noise is added to the values ​​of each dimension of the original data from the external verification source to satisfy differential privacy. ,in For query sensitivity, A value of 0.1 is preferred for the privacy budget; the merged result retains only the statistics and does not store the original case data.

[0014] As a further preferred embodiment of the intelligent question-answering method driven by a knowledge graph of chronic disease diet according to the present invention, if the consistency score is lower than a preset threshold, the node attributes are iteratively adjusted to obtain an optimized knowledge graph version, including: External validation sources are divided into multiple subsets by deviation type labels. Original indicator data are extracted from each subset to form indicator data groups. The indicator data groups are processed by weighted averaging to obtain fused indicator values. The specific values ​​of each dimension are separated from the fused indicator values ​​to form a numerical sequence. The specific values ​​of each dimension in the numerical sequence are compared with the standard reference value to obtain the absolute deviation. If the absolute deviation is lower than the preset threshold, the corresponding dimension is marked as a consistent state to obtain a consistent state set. The consistency score is determined by calculating the ratio of the number of marked dimensions in the consistent state set to the total number of dimensions. The iterative adjustment of node attributes is achieved by minimizing the following objective function: In the formula, Adjust the vector for node attributes; The adjusted consistency score; Score for consistency with objectives; express Norms are used to encourage sparse adjustments; , This is the balance coefficient.

[0015] As a further preferred embodiment of the intelligent question-answering method driven by a knowledge graph for chronic disease diet according to the present invention, the step of generating a dietary advice query interface through an optimized knowledge graph version to obtain a real-time output set for patient decision-making includes: extracting structured information related to dietary advice through the constructed knowledge graph, classifying the information using a pre-established classification model to obtain a preliminary set of advice categories; based on the preliminary set of advice categories and combined with the specific conditions required for patient decision-making, using a filtering mechanism to select a subset of categories that meet the conditions to determine the appropriate advice range; and for the appropriate advice range, using the query interface to obtain the corresponding detailed data content from the knowledge graph to obtain a personalized information combination related to patient decision-making. By combining personalized information, the system invokes pre-defined logical rules in the interface development. If the personalized information combination meets the pre-defined matching conditions, corresponding dietary advice items are generated, resulting in a targeted advice list. Based on the targeted advice list, a real-time processing mechanism is implemented, transmitting the list content to the feedback set through a query interface to determine whether it meets the time constraints of real-time feedback and obtain the final feedback result. From the final feedback result, content directly related to decision support is extracted, and data processing tools are used for structured processing to determine the final output real-time feedback set. For the real-time feedback set, a version-optimized knowledge graph update mechanism is used to associate the feedback result with the graph version, resulting in continuously updated graph data content. The determination of whether the real-time feedback time constraint is met is based on the following probability formula to quantify the real-time performance of the interface: In the formula, The probability of satisfying the time constraint; To query the total number of samples; For the first The actual response time for this query; Maximum allowable response time; This is an indicator function that takes the value 1 when the condition is true and 0 otherwise. The dietary advice query interface also includes a caching mechanism: for patients with the same or similar decision criteria, i.e., feature vector cosine similarity. The query results are cached for 24 hours; an "useful / useless" feedback button is embedded in the interface. When a suggestion is marked as "useless" more than 5 times, the backend is triggered to re-evaluate the priority and consistency of the knowledge entries corresponding to the suggestion and update the knowledge graph version.

[0016] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This invention discloses a knowledge graph optimization and dietary suggestion generation method based on medical evidence. Addressing the unique challenges of frequent clinical guideline updates, difficulty in ensuring knowledge graph consistency, and low efficiency in generating personalized suggestions for patients in the medical field, this method solves these problems through multi-layered technology integration. It acquires the latest medical data through an evidence collection module, selects high-priority knowledge items using the analytic hierarchy process (AHP), and triggers a knowledge graph node update mechanism based on evidence level. Subsequently, it uses graph embedding and random forest algorithms to identify deviations in node associations and integrates external validation source data to optimize consistency scores. Finally, it generates a real-time dietary suggestion query interface to provide patients with precise decision support. This invention ensures dynamic optimization of the knowledge graph by iteratively adjusting node attributes, significantly improving the accuracy of medical knowledge updates and the personalization level of suggestion generation, achieving intelligent management of the entire chain from data extraction to application output. (See attached figures.)

[0017] Figure 1 This is a flowchart of an intelligent question-answering method driven by a knowledge graph of dietary knowledge for chronic diseases, according to the present invention. Figure 2 This is a schematic diagram illustrating evidence collection and prioritization driven by a knowledge graph of dietary knowledge for chronic diseases, as described in this invention. Figure 3 This is a schematic diagram illustrating the generation of a knowledge graph optimization and query interface driven by a chronic disease diet knowledge graph according to the present invention. Detailed implementation method.

[0018] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] like Figures 1-3 As shown, this embodiment of a chronic disease diet knowledge graph-driven intelligent question-answering method may specifically include: like Figure 2As shown in the figure, this embodiment illustrates the process of extracting clinical guidelines and research results from a medical database, determining a list of high-priority items through the analytic hierarchy process, and triggering a knowledge graph node update mechanism based on the level of evidence.

[0020] Step S101: Extract the latest clinical guidelines and research findings from the medical database through the preset evidence collection module to obtain a raw data set containing evidence level labels.

[0021] The latest clinical guidelines and research findings are retrieved from medical databases using pre-defined modules to construct a raw dataset containing evidence level labels. An automated web crawling tool scans the medical database to obtain the latest information related to clinical guidelines and research findings. If the retrieved information meets the pre-defined evidence level criteria, it is classified as valid data, resulting in a pre-selected dataset. For this pre-selected dataset, text parsing techniques are applied to extract the evidence level information, identifying structured data with level labels. Based on the level labels in the structured data, clinical guidelines and research findings at different evidence levels are categorized and organized to obtain hierarchically organized data groups. Using these hierarchically organized data groups, a raw dataset containing evidence level labels is constructed, and its completeness and consistency are assessed. After obtaining the raw dataset, a storage management tool is used to save it to a designated database, yielding the final data resources available for subsequent analysis.

[0022] For example, when constructing a raw dataset with evidence level annotations, the latest clinical guidelines and research findings can be obtained from medical databases such as PubMed or UpToDate. For the implementation of automated web scraping tools, customized scripts can be used to regularly scan these databases, setting keywords such as "clinical guidelines 2023" or "randomized controlled trials," retrieving approximately 100 recent articles daily to ensure the timeliness of the information. The benefit of this approach is the ability to promptly capture cutting-edge research developments, providing fresh data for subsequent analysis.

[0023] In one possible implementation, when initially screening the dataset, the evidence level standard can be set based on the internationally accepted GRADE system, retaining only documents with an evidence level of "high" or "medium".

[0024] For example, out of 100 crawled documents, 30 meet the criteria and are classified as valid data. This screening mechanism helps to eliminate low-quality research, improve data reliability, and lay a solid foundation for subsequent analysis.

[0025] Specifically, text parsing technology can be applied to extract evidence level information from document abstracts. For example, if a document abstract explicitly states "this study has a GRADE (High) level of evidence," the system can identify and extract this information using natural language processing, creating structured data fields such as "Document ID: 001, Evidence Level: High." The advantage of this method is that it transforms unstructured text into actionable data, facilitating classification and retrieval.

[0026] For example, when categorizing and organizing clinical guidelines and research findings at different levels of evidence, the data can be divided into three tiers: "high," "medium," and "low." Assuming that after screening there are 20 high-level evidence guidelines and 10 medium-level research findings, the system automatically generates tiered data groups. This tiered approach helps researchers quickly locate high-quality evidence, improving decision-making efficiency.

[0027] In one possible implementation, when determining the integrity and consistency of the original data set, one can check whether there are missing fields or contradictory annotations in the data group.

[0028] For example, if the evidence level rating of a document is inconsistent across different databases, the system will flag it and prompt for manual verification to ensure data quality. This verification mechanism effectively avoids data bias and guarantees the reliability of the analysis results.

[0029] Specifically, storage management tools can be implemented using relational databases such as MySQL, and the final data resources can be partitioned and stored according to evidence level and publication time.

[0030] For example, 30 valid data entries can be stored separately in a "high-level partition" and a "medium-level partition," with a daily backup mechanism set up. This storage method not only facilitates data retrieval but also ensures data security, providing stable support for subsequent analysis.

[0031] For example, the collaborative work of the aforementioned stages, from crawling to storage, forms a complete data processing chain, ensuring the efficiency of transformation from raw information to final resources. Through automation and structuring, the cost of manual intervention is significantly reduced, while data accuracy and usability are improved, providing strong technical support for clinical research and decision support.

[0032] Step S102: Based on the evidence level labels in the original dataset, the hierarchical analysis method is used to prioritize each knowledge item and determine the list of high-priority items.

[0033] Based on the evidence level labels in the original dataset, pairwise comparison relationships are constructed for each knowledge item. A judgment matrix is ​​constructed using the analytic hierarchy process (AHP), and the pairwise comparison values ​​are filled in to obtain an initial judgment matrix. Eigenvectors are calculated from the judgment matrix, and after normalization, the weight values ​​for each knowledge item are obtained. After obtaining the weight values, the largest eigenvalue of the judgment matrix is ​​calculated, and a consistency index is determined based on the matrix order. If the ratio of the consistency index to the random consistency index is less than a preset threshold, the consistency test passes, and the weight values ​​are retained. The weight values ​​that pass the consistency test are then sorted, with all knowledge item weight values ​​arranged in descending order. Based on the sorted weight value sequence, the top-ranking knowledge items are extracted to determine a high-priority item list.

[0034] The pairwise comparison values ​​are assigned based on the following formula to determine the level of evidence and timeliness: In the formula, Representing knowledge entries Relative to the entry Importance ratio; , Each item and entries The level of evidence quantification value, through Calculation, where For the original level of evidence, As a preset baseline level, Sensitivity coefficient; , The year the entry was published; The maximum time span among all entries; This is the timeliness adjustment coefficient.

[0035] In the consistency test of the judgment matrix of the analytic hierarchy process, if the consistency ratio... The test passes; when the number of knowledge items... At that time, hierarchical clustering was used to group items according to disease categories. Pairwise comparisons were performed within each group, and comparisons were made between groups using class centroids. The computational complexity was reduced from... Down to ,in This represents the number of categories.

[0036] For example, when constructing pairwise comparison relationships between knowledge items based on the evidence level labels in the original dataset, clinical guidelines and research findings in the same disease area can be transformed into independent knowledge items, and then their importance can be compared pairwise.

[0037] In one possible implementation, for the field of hypertension management, five core knowledge items are extracted from the previously screened 30 valid data points, such as "lifestyle intervention", "preferred ACEI drugs", "combination therapy", "target organ protection strategy" and "adjustment of follow-up frequency", and the relative importance of each pair of items is systematically assessed.

[0038] Specifically, when constructing the judgment matrix using the analytic hierarchy process, the values ​​are filled using a 1-9 scale, where 1 indicates that both are equally important, 3 indicates that the former is slightly more important, and 9 indicates that the former is absolutely important.

[0039] For example, comparing "lifestyle intervention" with "ACEI drugs as the first choice," if the clinical guidelines emphasize drug intervention more in high-level evidence items, then 1 / 3 can be filled in, indicating that lifestyle intervention is slightly less important; otherwise, 3 can be filled in. Similarly, all pairwise combinations are filled in to form a 5-order judgment matrix. This scaling method directly reflects the differences in the level of evidence; high-level evidence items tend to receive higher scores in the comparison, thus reflecting objective priority.

[0040] In one possible implementation, the weight values ​​of each knowledge item are obtained by calculating the eigenvectors through the judgment matrix and then normalizing them.

[0041] For example, the calculation results show that "ACEI drugs as the first choice" has a weight of 0.35, "combination therapy" has a weight of 0.28, "target organ protection strategy" has a weight of 0.20, "lifestyle intervention" has a weight of 0.12, and "adjustment of follow-up frequency" has a weight of 0.05. These weight values ​​quantify the relative contribution of each item, which facilitates subsequent prioritization.

[0042] For example, after obtaining the weight values, the largest eigenvalue of the judgment matrix is ​​calculated, and the consistency index is determined in combination with the matrix order. If the consistency ratio is less than 0.1, the matrix is ​​considered to have satisfactory consistency, and the obtained weight values ​​are retained.

[0043] Specifically, assuming the calculated largest eigenvalue is 5.18 and the consistency index is 0.045, compared to the 5th order matrix random consistency index of 1.12, the ratio is approximately 0.04, far below the 0.1 threshold, thus passing the consistency test. This testing mechanism effectively avoids subjective bias and ensures that the weight allocation is reasonable and reliable.

[0044] In one possible implementation, the weights that pass the consistency test are used for sorting, forming a sequence in descending order. In the example above, the sorting is: "ACEI drugs as the first choice," "combination therapy," "target organ protection strategies," "lifestyle interventions," and "adjustment of follow-up frequency." Based on the sorted weight sequence, the top items are extracted to determine a high-priority list; for example, the top three items with a cumulative weight of 70% are selected as high-priority items. This extraction method highlights the most critical knowledge, directly supports the optimization of clinical decision-making, and improves the relevance and efficiency of guideline application.

[0045] Specifically, the entire process progresses step-by-step from pairwise comparisons to the formation of a high-priority list, ensuring that the weights reflect the strength of the actual evidence. Through quantitative ranking, researchers can quickly focus on core recommendations, reduce information overload, and improve the accuracy and timeliness of decision-making.

[0046] In a preferred embodiment, when the number of knowledge items exceeds 100, to improve the computational efficiency of the analytic hierarchy process (AHP), the items are first clustered according to disease categories (e.g., diabetes, hypertension, kidney disease). Within each cluster, pairwise comparisons are performed to construct a judgment matrix, and the intra-cluster weights are calculated. Then, each cluster is treated as a whole, and the inter-cluster weights are obtained through comparisons between items at the cluster centroids. Finally, the global weight of each item is the intra-cluster weight multiplied by the inter-cluster weight. This method reduces the number of comparisons from... Reduced to approximately This significantly reduces human intervention.

[0047] Step S103: Obtain the details of the entries in the high-priority entry list. If the evidence level is higher than the preset threshold, trigger the update mechanism to determine the knowledge graph nodes that need to be updated.

[0048] The process begins by retrieving detailed information from a list of high-priority entries and performing initial screening based on the evidence level of each entry, resulting in a set of entries that meet the criteria. For the selected set of entries, if the evidence level exceeds a preset threshold, an update process is triggered to determine the range of knowledge graph nodes requiring processing. Based on this range, the current associated data of the nodes is retrieved, and any inconsistent links are identified, resulting in a list of nodes to be adjusted. For this list, pre-established mapping rules are used to update the links between nodes, determining the updated node structure. The updated node structure is then used to retrieve detailed information about associated entries, assessing their compliance with priority requirements to obtain the final adjustment result. Finally, a batch processing tool is used to synchronize the adjusted node data to the knowledge graph, completing the entire update process.

[0049] For example, in the field of hypertension management, when retrieving detailed information on high-priority items from a list of items, the focus can be on items with higher weight in the previous ranking, such as "ACEIs as the first choice" and "combination therapy regimens." An initial screening is performed based on the evidence level of these items. Assuming a preset threshold of evidence level A, if an item cites multiple authoritative international guidelines, it is included in the set of eligible items. After screening, if the evidence level of "ACEIs as the first choice" is A+, which is higher than the threshold, an update process is triggered to determine the scope of relevant nodes in the knowledge graph, such as all sub-nodes under the drug treatment category.

[0050] For example, after determining the scope of knowledge graph nodes, when acquiring current associated data, it can be found that the "ACEI drug first choice" node is linked to the "lifestyle intervention" node. However, data shows that the two are not directly related in certain clinical scenarios, and this inconsistent link relationship needs to be adjusted. By analyzing the list of nodes to be adjusted and using pre-established mapping rules, such as redefining link priorities based on the level of evidence, a closer link is established between the "ACEI drug first choice" and "combination therapy" nodes, while the association with other low-priority nodes is weakened, forming an updated node structure.

[0051] For example, when retrieving detailed information about related entries for the updated node structure, if the "combination therapy" entry has a weight of 0.28 and meets the priority requirements, its core position in the knowledge graph is retained. The final adjustment results show that the links between drug-related nodes are more consistent with clinical practice logic. Subsequently, batch processing tools are used to synchronize the adjusted node data to the knowledge graph to ensure data consistency. This approach improves the accuracy of the knowledge graph and provides a more reliable basis for clinical decision-making.

[0052] For example, the combination of evidence level screening and node link adjustment throughout the process ensures that high-priority items in the knowledge graph are highlighted. Taking hypertension management as an example, the updated graph more clearly demonstrates the core role of drug treatment and reduces interference from low-evidence items. This method helps optimize the information structure and improve data retrieval efficiency. Simultaneously, the application of batch synchronization tools enables rapid large-scale data updates, saving manual proofreading time and ensuring the real-time nature and usability of the knowledge graph.

[0053] Figure 3 This diagram illustrates the knowledge graph optimization and query interface generation process in the method of the present invention, corresponding to steps S104 to S108. It shows the process of the graph embedding algorithm calculating node similarity, the random forest algorithm classifying inconsistent parts, the fusion of external verification source data and consistency score judgment, the iterative adjustment of node attributes, and finally generating a dietary suggestion query interface and outputting a real-time feedback set. Step S104: For the knowledge graph nodes that need to be updated, the similarity between nodes is calculated using a graph embedding algorithm to obtain a set of association relationships consistent with clinical guidelines.

[0054] For graph nodes in the knowledge graph, a graph embedding algorithm is used to analyze node similarity, obtaining preliminary node association data and obtaining node similarity results related to clinical guidelines. Based on the preliminary node similarity results, the node association data is filtered to obtain a set of associations that conform to guideline consistency, determining the priority range of nodes to be processed. According to the determined node range, detailed information about nodes related to clinical guidelines is extracted from the knowledge graph, and potential connections between nodes are determined, resulting in a list of associations to be processed. For the list of associations to be processed, a pre-established rule base is used for comparison. If an association meets the guideline consistency requirements, the relationship is retained, obtaining the final node association structure. Based on the final node association structure, relevant node data is obtained from the knowledge graph, and the logical connections between the data are organized to determine the complete set of associations. Based on the organized set of associations, a batch processing tool is used to synchronize the node association data to the knowledge graph, completing the node similarity and clinical guideline consistency update process. For the synchronized knowledge graph data, the updated node association information is obtained, and the data integrity is checked using an internal verification tool to obtain the final processing result.

[0055] The graph embedding algorithm uses the following similarity formula, which incorporates consistency with clinical guidelines, to calculate the similarity between nodes: In the formula, For nodes With nodes Overall similarity between them; , They are nodes and Graph embedding vectors; Indicates cosine similarity; For nodes and The set of existing relationship types between them; This is the set of standard relation types defined in clinical guidelines; This is a balancing factor used to adjust the weights of structural similarity and guideline consistency. Among them, when the degree of the knowledge graph node to be updated is lower than a preset threshold In option 3, semantic distance based on medical ontology is used as the initial similarity: ,in For nodes and The shortest path length in SNOMED CT or UMLS ontology; after the node degree is restored, switch to the graph embedding similarity formula.

[0056] For example, in the field of hypertension management, when performing similarity analysis on graph nodes in a knowledge graph, graph embedding algorithms can be used to transform the nodes into vector form, capturing the potential relationships between them. In principle, graph embedding algorithms analyze the position and connectivity of nodes in the graph, transforming the complex graph structure into a low-dimensional vector space, thus facilitating the calculation of similarity between nodes. Suppose that during the analysis, for the nodes "calcium channel blockers" and "β-blockers," a relatively short vector distance is found, indicating that they may have a high correlation in clinical applications. This preliminary node association data provides a foundation for subsequent screening.

[0057] For example, when screening preliminary results for similar nodes, the consistency requirements of clinical guidelines can be considered, with a focus on nodes related to hypertension treatment. Assuming the screening criteria are based on the treatment priorities recommended in the guidelines, if the "calcium channel blockers" node is listed as the first-line drug for elderly patients in the guidelines, and the "beta-blockers" node is suitable for patients with tachycardia, then both are included in the consistent association set and identified as nodes to be prioritized. This screening approach ensures that subsequent treatment focuses on core aspects of clinical practice.

[0058] For example, when extracting detailed node information and identifying potential connections, an indirect association may be found between the "calcium channel blockers" node and the "dietary control" node. However, guideline analysis reveals that there is no direct causal relationship between the two in certain scenarios, so this node is added to the list of relationships to be processed. For the relationships in the list, a pre-established rule base is used for comparison. If the rule base explicitly states that drug treatment and non-drug interventions should be handled separately, the association is removed, and node relationships that meet the guideline requirements are retained, forming the final association structure.

[0059] For example, when organizing the logical connections of the final node association structure, a strong clinical relevance was found between the "calcium channel blockers" node and the "combination therapy regimen" node, with a weight value assumed to be 0.85, significantly higher than other nodes. Therefore, it was listed as a core association. Subsequently, batch processing tools were used to synchronize these association data to the knowledge graph to ensure timely updates. After synchronization, internal validation tools were used to check data integrity, such as verifying whether the "combination therapy regimen" node was correctly linked to all relevant drug nodes. If any missing data was found, it was promptly added to ensure the logical consistency of the knowledge graph.

[0060] For example, from another perspective, in the synchronous update process, for the knowledge graph in the field of hypertension management, nodes related to first-line treatment can be prioritized, such as the combination relationship between "ACEI drugs" and "diuretics." By analyzing the recommended scenarios in clinical guidelines, the link strength between the two can be adjusted. This adjustment ensures the prominence of key treatment options in the graph, providing more accurate data support for subsequent applications.

[0061] In the early stages of knowledge graph construction or when the number of nodes in a particular disease area is small (e.g., degree), When graph embedding algorithms fail to obtain stable vectors, semantic distance based on medical ontology is used as a similarity metric. Taking SNOMED CT as an example, each node corresponds to a concept ID, and the shortest path length is calculated through the "is-a" relationship between concepts in the ontology. Similarity is defined as When the node degree grows to exceed a threshold (e.g., 10), it automatically switches back to the graph embedding similarity formula, thus achieving a smooth transition from the cold start to the hot phase.

[0062] Step S105: Extract the inconsistent parts from the association set, classify the inconsistent parts using the random forest algorithm, and determine the deviation type label.

[0063] Relationship data is extracted from the knowledge graph. Initial screening is performed on parts with low consistency, resulting in a set of relationships that do not meet preset criteria. For this set of filtered relationships, a random forest algorithm is used to classify the deviation types. This random forest algorithm employs an active learning strategy: it calculates the prediction entropy of each decision tree for the same relationship sample; if the entropy value exceeds a preset threshold, the sample is pushed to manual review, and the review results are added as new training data. The random forest is retrained after every 1000 samples. A deviation type label is determined for each relationship. Based on the determined deviation type labels, information comparison is performed on the anomaly markers to obtain category identifiers that do not match the pre-established rule base. Based on the category identifier comparison results, nodes are grouped and data is stratified to obtain a stratified set of relationship data. For this stratified set of relationship data, information comparison is used to filter relationships. If the category identifier of a relationship matches the rule base, the relationship is retained, resulting in a filtered list of relationships. Based on the filtered list of relationships, labels are assigned to the data stratification results to determine the final classification processing structure. Through the final classification and processing structure, the remaining relational data is stored and updated to obtain a complete record of deviation type processing.

[0064] For example, in the knowledge graph of hypertension management, the extraction and processing of relational data can be analyzed and implemented from multiple perspectives. For the initial screening of parts with low consistency, preset criteria can be used to evaluate the relationships between nodes in the knowledge graph. The assumed criteria are based on the consistency of recommendations from clinical guidelines; if a relationship does not reach a preset threshold, such as a correlation score below 0.8, it is categorized into the set that does not meet the criteria. This process lays the foundation for subsequent classification.

[0065] For example, when classifying bias types using the random forest algorithm for a selected set of associations, biases can be categorized into three types: missing data, logical conflicts, and inapplicable scenarios. Suppose that when analyzing the relationship between the "diuretics" node and the "dietary intervention" node, it is found that the two are incorrectly associated in certain scenarios; the system would mark this as an inapplicable scenario. This classification helps to accurately pinpoint the root cause of the problem.

[0066] For example, during the information comparison phase, by matching anomaly markers with the rule base, discrepancies can be found in certain category identifiers. Suppose the rule base specifies that the "ACEI drugs" and "renal function monitoring" nodes should have a strong correlation, but this is not reflected in the actual data; the system will mark it as a mismatch category. This comparison result provides a basis for subsequent stratified processing.

[0067] For example, for hierarchical processing of data grouped by nodes, relational data can be divided into two layers: priority processing and secondary processing, based on the type of deviation. Suppose that the nodes "calcium channel blockers" and "heart rate control" are assigned to the priority layer due to logical conflict, while other non-critical relationships are assigned to the secondary layer. This hierarchical approach ensures efficient processing.

[0068] For example, during the relationship screening phase, relationships that meet the requirements can be retained by checking the consistency between the category identifier and the rule base. For instance, if the relationship between the nodes "β-blockers" and "tachycardia treatment" conforms to guideline recommendations, it will be retained in the final list. This screening ensures the clinical relevance of the data.

[0069] For example, labeling the results of data stratification can provide a clear identifier for the final classification structure. Assuming a relationship in a priority layer is marked as high priority, the system will assign it a red label to indicate its importance. This labeling facilitates subsequent tracking and management.

[0070] For example, during the storage update phase, remaining relationship data can be synchronized to the knowledge graph database using a batch upload tool. Assuming that after the update, the system generates complete records of deviation type processing, including the processing time and results for each relationship, data traceability is ensured. This approach significantly improves the maintenance efficiency of the knowledge graph and provides a reliable basis for clinical decision support.

[0071] To reduce the cost of manual labeling of bias types, an active learning strategy is adopted. The random forest algorithm outputs the predicted category of each tree for each unlabeled association sample and calculates the prediction entropy. ,in For category The predicted probability. If If the model is unsure about a particular sample, it is added to the review queue. After expert review, the sample is used as a new training sample, and the random forest is retrained every 1000 new samples. This method can improve classification accuracy by approximately 30% with the same annotation budget.

[0072] Step S106: Based on the deviation type label, obtain the data fusion result of the external verification source and determine the consistency score after fusion.

[0073] External validation sources are categorized and filtered using deviation type labels. Raw data is extracted from the corresponding external validation sources based on the categorization results. A weighted average method is used to perform fusion processing on the extracted raw data to obtain the fusion result. The numerical distribution of each dimension is extracted from the fusion result. The absolute value of the deviation from the standard reference value is calculated for each dimension's numerical distribution. If the absolute value of the deviation is less than a preset threshold, the dimension is considered consistent; otherwise, it is considered inconsistent. The final consistency score is determined based on the proportion of consistent dimensions to the total number of dimensions.

[0074] The consistency score is calculated using a weighted penalty formula: In the formula, For final consistency score; This represents the total number of dimensions. For the first The absolute value of the dimensional deviation; For the first Dimensional deviation threshold; For the first The weights of the dimensions are obtained using the analytic hierarchy process (AHP). This is an indicator function that takes the value 1 when the condition is true and 0 otherwise. Before fusion, Laplace noise is added to the values ​​of each dimension of the original data from the external verification source to satisfy differential privacy. ,in For query sensitivity, A value of 0.1 is preferred for the privacy budget; the merged result retains only the statistics and does not store the original case data.

[0075] For example, in the construction of a knowledge graph in the field of hypertension management, external validation sources can be categorized and screened based on deviation type labels, starting with different sources of clinical data. Assuming external validation sources include clinical trial data, patient medical records, and expert guidelines, the system will classify these sources into high-confidence and low-confidence categories based on deviation type labels, such as missing data or logical conflicts. Data from high-confidence sources will be extracted first, while low-confidence sources will be used as supplementary references. This classification method helps to focus on a more reliable data foundation.

[0076] For example, when extracting raw data from the classification and screening results, one can extract efficacy data on a certain antihypertensive drug, such as ACEI drugs, in a specific population from highly reliable clinical trial data.

[0077] Specifically, the extracted data is assumed to include multiple indicators such as drug dosage, blood pressure control rate, and adverse reaction rate to ensure data comprehensiveness. For sources of low reliability, such as partially incomplete medical records, only key fields are extracted to reduce the introduction of errors.

[0078] For example, when using a weighted average method to fuse extracted raw data, weights can be assigned to data from different sources. Assuming a weight of 0.7 for clinical trial data, 0.2 for expert guidelines, and 0.1 for medical records, a weighted average can be used to calculate the overall effect of a certain drug on blood pressure control. This method can balance the reliability differences of data from different sources, resulting in a fusion result that is closer to reality.

[0079] For example, when extracting the numerical distribution of each dimension from the fusion results, we can analyze the distribution of blood pressure control rates among patients of different age groups. Suppose the fusion results show that the control rate for patients aged 40-60 is concentrated between 75% and 85%, while that for patients over 60 is between 60% and 70%. This difference in distribution provides data support for subsequent bias analysis.

[0080] For example, when calculating the absolute value of the deviation from the standard reference value for the distribution of values ​​in each dimension, suppose the standard reference value is a blood pressure control rate of 80%, while the actual value for a certain age group is 70%, resulting in an absolute deviation of 10%. If the preset threshold is 5%, then that dimension is considered inconsistent. This method of judgment can intuitively reflect the gap between the data and the standard.

[0081] For example, when determining the final consistency score based on the proportion of consistent judgments across all dimensions to the total number of dimensions, assuming there are 5 dimensions and 3 are judged to be consistent, the consistency score is 60%. This score can serve as a measure of the quality of related relationship data in the knowledge graph, providing a reference direction for subsequent optimization. Through the above series of processes, the relevance and reliability of data processing can be effectively improved, ensuring the application value of knowledge graphs in the field of hypertension management.

[0082] Before fusing external validation sources (such as hospital medical records and clinical trial data), differential privacy protection is employed to prevent patient privacy breaches. Specifically, for each dimension (such as blood pressure and blood glucose levels), noise following a Laplace distribution is added when calculating its statistics (mean and variance). ,in This represents the range of values ​​for that dimension. A privacy budget of 0.1 is used. The fusion result only outputs the statistics after adding noise; the original data does not leave the local server and does not store any information traceable to the individual.

[0083] In step S107, if the consistency score is lower than the preset threshold, the node attributes are iteratively adjusted to obtain an optimized version of the knowledge graph.

[0084] External validation sources are divided into multiple subsets by deviation type labels. Original indicator data are extracted from each subset to form indicator data groups. The indicator data groups are processed using a weighted average method to obtain fused indicator values. Specific values ​​for each dimension are separated from the fused indicator values ​​to form a numerical sequence. The specific value of each dimension in the numerical sequence is compared with the standard reference value to obtain the absolute deviation. If the absolute deviation is lower than a preset threshold, the corresponding dimension is marked as a consistent state, resulting in a consistent state set. The consistency score is determined by calculating the ratio of the number of marked dimensions in the consistent state set to the total number of dimensions.

[0085] The iterative adjustment of node attributes is achieved by minimizing the following objective function: In the formula, Adjust the vector for node attributes; The adjusted consistency score; Score for consistency with objectives; express Norms are used to encourage sparse adjustments; , This is the balance coefficient.

[0086] The dietary advice query interface also includes a caching mechanism: for patients with the same or similar decision criteria, i.e., feature vector cosine similarity. The query results are cached for 24 hours; an "useful / useless" feedback button is embedded in the interface. When a suggestion is marked as "useless" more than 5 times, the backend is triggered to re-evaluate the priority and consistency of the knowledge entries corresponding to the suggestion and update the knowledge graph version.

[0087] For example, in the construction of a knowledge graph in the field of hypertension management, the processing and consistency assessment of external validation sources can start with the diversity of clinical data and be further analyzed in conjunction with deviation type labels. Dividing external validation sources into multiple subsets based on deviation type labels can be understood as classifying data sources according to reliability or completeness.

[0088] For example, clinical trial data is categorized as a high-reliability subset due to its rigor, while some patient-reported data is categorized as a low-reliability subset due to potential subjective bias. This categorization helps to make subsequent data processing more targeted.

[0089] For example, when extracting raw indicator data from multiple subsets to form an indicator data set, key indicators such as blood pressure control rate and drug side effect incidence can be extracted first from high-reliability subsets, while only auxiliary indicators such as patient compliance feedback can be extracted from low-reliability subsets. This hierarchical extraction method ensures the quality of the core information in the data set. The weighted average method is then used to process the indicator data set to obtain the fused indicator value, which can be understood as assigning different levels of importance to data from different sources.

[0090] For example, the weight of high-reliability data is set to 0.8, and the weight of low-reliability data is set to 0.2. Through comprehensive calculation, a more realistic fusion index value is obtained, such as the overall effectiveness of a certain drug.

[0091] For example, when separating the specific values ​​of each dimension from the fusion index values ​​to form a numerical sequence, the effectiveness can be subdivided into the performance of different age groups, such as the specific values ​​of three dimensions: under 40 years old, 40 to 60 years old, and over 60 years old. This separation method facilitates subsequent analysis one by one. Comparing the specific value of each dimension in the numerical sequence with the standard reference value yields the absolute deviation, which can be understood as measuring the gap between the actual data and the ideal target.

[0092] For example, suppose the standard reference value is 85% effectiveness, while the actual value for the 40 to 60 age group is 80%, with a deviation of 5%. If the preset threshold is 3%, this dimension is marked as inconsistent.

[0093] For example, after obtaining a consistent set of states, the consistency score can be determined by calculating the ratio of the number of labels to the total number of dimensions, which can intuitively reflect the overall quality of the data.

[0094] For example, if there are 5 dimensions, and 2 of them are marked as consistent, then the consistency score is 40%. This scoring method provides a reference direction for subsequent data optimization.

[0095] It should be noted that the determination of absolute deviation and the marking of consistency status reflect the refinement of data analysis and help identify potential problems.

[0096] For example, in hypertension management, if a certain dimension shows a significant deviation, the issue can be traced back to the reliability of the data source, allowing for adjustments to the weights or the addition of higher-quality data sources. This approach effectively improves the accuracy of knowledge graphs.

[0097] In one possible implementation, for dimensions with low consistency scores, expert opinions can be introduced as a supplementary verification source to further calibrate the fusion index values.

[0098] For example, if the efficacy rate deviation is significant in the group aged 60 and above, data weights or indicator definitions can be adjusted by incorporating specific recommendations for elderly patients from expert guidelines. This supplementary approach ensures the comprehensiveness and reliability of the data, providing a more solid foundation for decision support in the field of hypertension management. Through the above multi-dimensional analysis and processing, the targeting of data processing is enhanced, laying the foundation for constructing a high-quality knowledge graph.

[0099] During iterative optimization, multiple convergence conditions are set: 1) Master condition: ;2) Auxiliary conditions: 3) Maximum number of iterations: 200. The Adam optimizer is used to adaptively adjust the learning rate, with an initial learning rate of 0.001, which does not require manual setting. The iteration is terminated early when the consistency score improvement is less than 0.005 after 5 consecutive iterations to avoid invalid computation. Testing shows that for a knowledge graph containing 5000 nodes, convergence occurs on average within 15-30 iterations, with each iteration taking less than 2 seconds.

[0100] Step S108: Generate a dietary advice query interface through an optimized knowledge graph version to obtain a real-time output set for patient decision-making.

[0101] By constructing a knowledge graph, structured information related to dietary recommendations is extracted. This information is then categorized using a pre-established classification model to obtain a preliminary set of recommendation categories. Based on this preliminary set, and considering the specific conditions required for patient decision-making, a filtering mechanism is used to select a subset of categories that meet the criteria, determining the appropriate recommendation range. For this appropriate range, a query interface is used to retrieve corresponding detailed data from the knowledge graph, obtaining personalized information combinations relevant to patient decision-making. These personalized information combinations are then used to invoke pre-defined logical rules in the interface development. If the personalized information combinations meet the pre-defined matching conditions, corresponding dietary recommendation items are generated, resulting in a targeted recommendation list. Based on this targeted recommendation list, a real-time processing mechanism is implemented, transmitting the list content to the feedback set via the query interface to determine if it meets the time constraints for real-time feedback and obtain the final feedback result. From the final feedback result, content directly related to decision support is extracted and structured using data processing tools to determine the final output real-time feedback set. For this real-time feedback set, a version-optimized knowledge graph update mechanism is used to associate the feedback result with the graph version, resulting in continuously updated graph data content.

[0102] The determination of whether the real-time feedback time constraint is met is based on the following probability formula to quantify the real-time performance of the interface: In the formula, The probability of satisfying the time constraint; To query the total number of samples; For the first The actual response time for this query; Maximum allowable response time; This is an indicator function that takes the value 1 when the condition is true and 0 otherwise.

[0103] For example, in the application of knowledge graphs in the field of hypertension management, after constructing the graph, structured information related to dietary recommendations can be extracted. Starting with patients' daily eating habits and nutritional needs, data nodes related to sodium intake, fat ratios, etc., can be identified. This information is then categorized using a pre-established classification model to form a preliminary set of recommendation categories, such as low-sodium diets and balanced nutrition. For each category set, combined with the patient's specific conditions such as age and blood pressure control, a filtering mechanism is used to select a suitable subset of categories, for example, prioritizing low-sodium, low-fat recommendations for middle-aged and elderly patients.

[0104] For example, based on the appropriate recommendation range, the system uses a query interface to retrieve detailed data from a knowledge graph, which may include recommended daily sodium intake values, recommended food types, etc., forming a personalized information combination. Suppose a patient needs to control their daily sodium intake below 2000 mg; the system will first extract a list of low-sodium foods. Then, through preset logical rules, if the patient's information meets the matching criteria and there are no other complications, specific dietary recommendations are generated, such as prioritizing brown rice as the staple food and choosing spinach instead of pickled foods as vegetables, forming a targeted recommendation list.

[0105] For example, in the real-time processing mechanism, the suggestion list is transmitted to the feedback set via a query interface to determine if it meets time constraints, such as requiring feedback to be completed within 5 minutes of a patient's query. If it does, content relevant to decision support, such as specific food pairing suggestions, is extracted, structured using data processing tools, and the real-time feedback set is output. For the feedback set, a version-optimized update mechanism links the feedback results to the atlas version, ensuring continuous updates to the atlas data, such as adding the latest nutritional data for a particular type of vegetable.

[0106] Specifically, the classification process of the classification model can be understood as dividing dietary recommendations into multiple levels based on historical data and expert rules, such as basic recommendations and personalized adjustment recommendations, to ensure coverage of different patient needs. When retrieving data from the query interface, high-priority data nodes, such as core nutritional indicators, are called first. The generation of personalized information combinations relies on the comparative analysis of real-time data input by patients and long-term data stored in the knowledge graph to ensure that the recommendations are realistic. The time constraint judgment of real-time feedback reflects the system's emphasis on response speed, while the version update mechanism ensures the timeliness of the knowledge graph content. The close collaboration of these links makes the generation of dietary recommendations more accurate, provides reliable decision support for patients, and also enhances the practical value of knowledge graphs in hypertension management.

[0107] To improve real-time responsiveness, the query interface implements a two-level cache: the first-level cache stores the 1000 most recent queries and their results, valid for 24 hours; the second-level cache stores aggregated results of high-frequency queries (such as "breakfast recommendations for diabetic patients"), valid for 7 days, and automatically expires when the knowledge graph version is updated. Simultaneously, an implicit feedback collection function is embedded in the interface: after receiving suggestions, users can click the "helpful" or "useless" button, and the system records the feedback. When a suggestion is marked "useless" more than 5 times, an asynchronous background task is triggered: the evidence level quantification value and node similarity of the corresponding knowledge item are recalculated. If the consistency score drops by more than 0.1, the node attributes are adjusted and a new version is released. Through this closed loop, the system can continuously optimize itself based on actual user experience.

[0108] Based on the embodiments of the present invention described above, and through the above description, those skilled in the art can make various changes and modifications without departing from the technical concept of the present invention. The technical scope of the present invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.

Claims

1. A knowledge graph-driven intelligent question-answering method for chronic disease diets, characterized in that, Specifically, the following steps are included: Step S101: Extract the latest clinical guidelines and research findings from the medical database using a pre-defined evidence collection module to obtain an original dataset containing evidence level labels. Step S102: Based on the evidence level labels in the original dataset, use the analytic hierarchy process (AHP) to prioritize each knowledge item and determine a list of high-priority items. Step S103: Obtain details of items in the high-priority list. If the evidence level is higher than a preset threshold, trigger an update mechanism to identify the knowledge graph nodes that need updating. Step S104: For the knowledge graph nodes that need updating, calculate the similarity between nodes using a graph embedding algorithm to obtain a set of associations consistent with clinical guidelines. Step S105: Extract inconsistencies from the association set and classify them using a random forest algorithm to determine deviation type labels. Step S106: Based on the deviation type labels, obtain the data fusion results from external validation sources and determine the consistency score after fusion. Step S107: If the consistency score is lower than a preset threshold, iteratively adjust node attributes to obtain an optimized version of the knowledge graph. Step S108: Generate a dietary advice query interface through an optimized knowledge graph version to obtain a real-time output set for patient decision-making.

2. The intelligent question-answering method driven by a knowledge graph of dietary knowledge for chronic diseases according to claim 1, characterized in that, The process of extracting the latest clinical guidelines and research findings from a medical database using a pre-defined evidence collection module to obtain a raw data set containing evidence level labels includes: obtaining the latest clinical guidelines and research findings from a medical database using a pre-defined module to construct a raw data set containing evidence level labels; scanning the medical database using an automated crawling tool to obtain the latest information related to clinical guidelines and research findings; if the latest information obtained meets the pre-defined evidence level standards, it is classified as valid data to obtain a preliminary filtered data set; for the preliminary filtered data set, text parsing technology is applied to extract the evidence level information to determine the structured data with level labels; based on the level labels in the structured data, clinical guidelines and research findings of different evidence levels are classified and organized to obtain hierarchically organized data groups; through the hierarchically organized data groups, a raw data set containing evidence level labels is constructed, and its completeness and consistency are judged; after obtaining the raw data set, a storage management tool is used to save it to a designated database to obtain the final data resources available for subsequent analysis.

3. The intelligent question-answering method driven by a knowledge graph of dietary knowledge for chronic diseases according to claim 1, characterized in that, The step involves prioritizing each knowledge item based on the evidence level labels in the original dataset using the analytic hierarchy process (AHP) to determine a list of high-priority items, including: Based on the evidence level labels in the original dataset, pairwise comparison relationships are constructed for each knowledge item; a judgment matrix is ​​constructed using the analytic hierarchy process (AHP), and the pairwise comparison values ​​are filled in to obtain the initial judgment matrix; eigenvectors are calculated from the judgment matrix, and the weight values ​​of each knowledge item are obtained after normalization. After obtaining the weight values, the largest eigenvalue of the judgment matrix is ​​calculated, and the consistency index is determined in combination with the matrix order. If the ratio of the consistency index to the random consistency index is less than the preset threshold, the consistency test is passed and the weight values ​​are retained. The weight values ​​that pass the consistency test are sorted and all knowledge item weight values ​​are arranged in descending order. Based on the sorted weight value sequence, the knowledge items at the top are extracted to determine the list of high-priority items. The pairwise comparison values ​​are assigned based on the following formula to determine the level of evidence and timeliness: In the formula, Representing knowledge entries Relative to the entry Importance ratio; , Each item and entries The level of evidence quantification value, through Calculation, where For the original level of evidence, As a preset baseline level, Sensitivity coefficient; , The year the entry was published; The maximum time span among all entries; This is a timeliness adjustment coefficient; In the consistency test of the judgment matrix of the analytic hierarchy process, if the consistency ratio... The test passes; when the number of knowledge items... At that time, hierarchical clustering was used to group items according to disease categories. Pairwise comparisons were performed within each group, and comparisons were made between groups using class centroids. The computational complexity was reduced from... Down to ,in This represents the number of categories.

4. The intelligent question-answering method driven by a knowledge graph of dietary knowledge for chronic diseases according to claim 1, characterized in that, The process of obtaining details of entries in the high-priority entry list, and triggering an update mechanism if the evidence level is higher than a preset threshold, to determine the knowledge graph nodes that need to be updated, includes: obtaining detailed information of high-priority entries from the entry list; performing preliminary screening based on the evidence level of each entry to obtain a set of entries that meet the conditions; for the selected set of entries, if the evidence level is higher than a preset threshold, triggering an update process to determine the range of knowledge graph nodes to be processed; based on the determined range of knowledge graph nodes, obtaining the current associated data of the nodes, determining whether there are inconsistent link relationships, and obtaining a list of nodes to be adjusted; for the list of nodes to be adjusted, using pre-established mapping rules to update the link relationships between nodes, and determining the updated node structure; through the updated node structure, obtaining detailed information of associated entries, determining whether they meet the priority level requirements, and obtaining the final adjustment result; based on the final adjustment result, using batch processing tools to synchronize the adjusted node data to the knowledge graph, completing the entire update process.

5. The intelligent question-answering method driven by a knowledge graph of dietary knowledge for chronic diseases according to claim 1, characterized in that, For the knowledge graph nodes that need updating, the similarity between nodes is calculated using a graph embedding algorithm to obtain a set of associations consistent with clinical guidelines, including: For graph nodes in the knowledge graph, a graph embedding algorithm is used to analyze node similarity, obtaining preliminary node association data and obtaining node similarity results related to clinical guidelines. Based on the preliminary node similarity results, the node association data is filtered to obtain a set of associations that conform to guideline consistency, determining the priority range of nodes to be processed. According to the determined node range, detailed node information related to clinical guidelines is extracted from the knowledge graph, and potential connections between nodes are determined, resulting in a list of relationships to be processed. For the list of relationships to be processed, a pre-established rule base is used for comparison; if the relationship conforms to guideline consistency requirements, it is retained, obtaining the final node association structure. Based on the final node association structure, relevant node data is obtained from the knowledge graph, and the logical connections between the data are organized to determine the complete set of associations. Based on the organized set of associations, a batch processing tool is used to synchronize the node association data to the knowledge graph, completing the node similarity and clinical guideline consistency update process. For the synchronized knowledge graph data, the updated node association information is obtained, and the data integrity is checked using an internal verification tool to obtain the final processing result. The graph embedding algorithm uses the following similarity formula, which incorporates consistency with clinical guidelines, to calculate the similarity between nodes: In the formula, For nodes With nodes Overall similarity between them; , They are nodes and Graph embedding vectors; Indicates cosine similarity; For nodes and The set of existing relationship types between them; This is the set of standard relation types defined in clinical guidelines; This is a balancing factor used to adjust the weights of structural similarity and guideline consistency. Among them, when the degree of the knowledge graph node to be updated is lower than a preset threshold In option 3, semantic distance based on medical ontology is used as the initial similarity: ,in For nodes and The shortest path length in SNOMED CT or UMLS ontology; after the node degree is restored, switch to the graph embedding similarity formula.

6. The intelligent question-answering method driven by a knowledge graph of dietary knowledge for chronic diseases according to claim 1, characterized in that, The process involves extracting inconsistencies from the set of relationships and classifying these inconsistencies using a random forest algorithm. This random forest algorithm employs an active learning strategy: it calculates the prediction entropy of each decision tree for the same relationship sample; if the entropy value exceeds a preset threshold, the sample is sent for manual review, and the review results are added as new training data. The random forest is retrained after every 1000 samples. The process also involves determining the deviation type label, including: extracting relationship data from the knowledge graph; initially screening for parts with low consistency to obtain a set of relationships that do not meet preset standards; and then using the random forest algorithm to classify the deviation types of the selected set of relationships to determine the deviation type for each relationship. Labeling; Based on the determined deviation type labels, information comparison is performed on the anomaly markers to obtain category identifiers that do not match the pre-established rule base; Based on the comparison results of the category identifiers, the nodes are grouped and data is layered to obtain a layered relational data set; For the layered relational data set, relation filtering is performed by information comparison, and if the category identifier of a certain relation is consistent with the rule base, the relation is retained, resulting in a filtered list of association relations; Based on the filtered list of association relations, labels are assigned to the data layering results to determine the final classification processing structure; Based on the final classification processing structure, the remaining association relation data is stored and updated to obtain a complete deviation type processing record.

7. The intelligent question-answering method driven by a knowledge graph of dietary knowledge for chronic diseases according to claim 1, characterized in that, The step of obtaining the data fusion result from the external verification source based on the deviation type label and determining the consistency score after fusion includes: External validation sources are categorized and filtered using deviation type labels; raw data is extracted from the corresponding external validation sources based on the categorization and filtering results, and a weighted average method is used to perform fusion processing on the extracted raw data to obtain the fusion result; the numerical distribution of each dimension is extracted from the fusion result; the absolute value of the deviation from the standard reference value is calculated for the numerical distribution of each dimension; if the absolute value of the deviation is less than a preset threshold, the dimension is determined to be consistent; otherwise, it is determined to be inconsistent; the final consistency score is determined based on the proportion of consistent dimensions to the total number of dimensions. The consistency score is calculated using a weighted penalty formula: In the formula, For final consistency score; This represents the total number of dimensions. For the first The absolute value of the dimensional deviation; For the first Dimensional deviation threshold; For the first The weights of the dimensions are obtained using the analytic hierarchy process (AHP). This is an indicator function that takes the value 1 when the condition is true and 0 otherwise. Before fusion, Laplace noise is added to the values ​​of each dimension of the original data from the external verification source to satisfy differential privacy. ,in For query sensitivity, A value of 0.1 is preferred for the privacy budget; the merged result retains only the statistics and does not store the original case data.

8. The intelligent question-answering method driven by a knowledge graph of dietary knowledge for chronic diseases according to claim 1, characterized in that, If the consistency score is lower than a preset threshold, the node attributes are iteratively adjusted to obtain an optimized knowledge graph version, including: External validation sources are divided into multiple subsets by deviation type labels. Original indicator data are extracted from each subset to form indicator data groups. The indicator data groups are processed by weighted averaging to obtain fused indicator values. The specific values ​​of each dimension are separated from the fused indicator values ​​to form a numerical sequence. The specific values ​​of each dimension in the numerical sequence are compared with the standard reference value to obtain the absolute deviation. If the absolute deviation is lower than the preset threshold, the corresponding dimension is marked as a consistent state to obtain a consistent state set. The consistency score is determined by calculating the ratio of the number of marked dimensions in the consistent state set to the total number of dimensions. The iterative adjustment of node attributes is achieved by minimizing the following objective function: In the formula, Adjust the vector for node attributes; The adjusted consistency score; Score for consistency with objectives; express Norms are used to encourage sparse adjustments; , This is the balance coefficient; The convergence condition for iteratively adjusting node attributes is as follows: or Or the number of iterations reaches The learning rate uses the Adam adaptive optimizer with an initial learning rate of 0.001; the iteration terminates early when the consistency score improvement is less than 0.005 over 5 consecutive iterations.

9. The intelligent question-answering method driven by a knowledge graph of dietary knowledge for chronic diseases according to claim 1, characterized in that, The process of generating a dietary advice query interface using an optimized knowledge graph version to obtain a real-time output set for patient decision-making includes: extracting structured information related to dietary advice from the constructed knowledge graph; classifying the information using a pre-established classification model to obtain a preliminary set of advice categories; based on the preliminary set of advice categories and the specific conditions required for patient decision-making, using a filtering mechanism to select a subset of categories that meet the conditions to determine the appropriate advice range; for the appropriate advice range, using the query interface to obtain corresponding detailed data content from the knowledge graph to obtain personalized information combinations related to patient decision-making; and using the personalized information combinations to call the interface development... The system uses pre-defined logical rules. If the personalized information combination meets the pre-defined matching conditions, corresponding dietary suggestion items are generated, resulting in a targeted suggestion list. Based on this targeted suggestion list, a real-time processing mechanism is implemented, transmitting the list content to the feedback set via a query interface to determine if it meets the time constraints for real-time feedback and obtain the final feedback result. From the final feedback result, content directly related to decision support is extracted and structured using data processing tools to determine the final output real-time feedback set. For the real-time feedback set, a version-optimized knowledge graph update mechanism is used to associate the feedback result with the graph version, resulting in continuously updated graph data content. The determination of whether the real-time feedback time constraint is met is based on the following probability formula to quantify the real-time performance of the interface: In the formula, The probability of satisfying the time constraint; To query the total number of samples; For the first The actual response time for this query; Maximum allowable response time; This is an indicator function that takes the value 1 when the condition is true and 0 otherwise. The dietary advice query interface also includes a caching mechanism: for patients with the same or similar decision criteria, i.e., feature vector cosine similarity. The query results are cached for 24 hours; an "useful / useless" feedback button is embedded in the interface. When a suggestion is marked as "useless" more than 5 times, the backend is triggered to re-evaluate the priority and consistency of the knowledge entries corresponding to the suggestion and update the knowledge graph version.