Method and system for hierarchical construction and classified management of medical insurance knowledge base
Through context-aware semantic slicing technology, hierarchical nested clustering algorithms and entropy-driven adaptive classification algorithms, combined with large language models, a dynamically adaptable medical insurance knowledge base is built, solving the shortcomings of traditional medical insurance knowledge bases in data processing and management, and achieving efficient and intelligent medical insurance knowledge management.
Patent Information
- Application Number
- CN202510363513.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
When facing complex and changing data needs and large-scale data processing, the existing medical insurance knowledge base management methods have problems such as insufficient dynamic adaptability, insufficient data heterogeneity and noise processing, insufficient semantic analysis, low retrieval efficiency and limited intelligent question-and-answer capabilities.
The context-aware semantic slicing technology, hierarchical nested clustering algorithm, entropy-driven adaptive classification algorithm and homoethic continuous algorithm are used to build a medical insurance knowledge base with strong dynamic adaptability, accurate data processing, profound semantic understanding, and intelligent classification management, and combine large language models to conduct intelligent Q&A.
It significantly improves the data processing and classification management efficiency of the medical insurance knowledge base, enhances the adaptability and practicality of the knowledge base, improves the search accuracy and user experience, and ensures the real-time and accuracy of the knowledge base.
Smart Images

Figure CN120277185A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of application of large language model knowledge bases, and particularly to a method and system for hierarchical construction and classification management of a medical insurance knowledge base. Background Art
[0002] With the continuous advancement of medical informatization, data management and application in the field of medical insurance are facing unprecedented challenges. As an important tool for medical insurance management and decision-making, the medical insurance knowledge base undertakes important responsibilities such as data collection, collation, storage, and retrieval. However, existing medical insurance knowledge base management methods have obvious deficiencies in dealing with complex and changing data requirements and large-scale data processing.
[0003] Firstly, traditional medical insurance knowledge bases usually rely on fixed classification systems and structures, lacking dynamic adaptability. With the continuous update of medical insurance policies and the sharp increase in the amount of medical data, traditional static classification methods are difficult to meet the ever-changing needs. This method often cannot flexibly adjust classification criteria and structures, resulting in a lag in the update of the knowledge base content and an inability to promptly reflect the latest policies and medical information, thereby affecting the accuracy of medical insurance management and decision-making.
[0004] Secondly, in terms of data processing and collation, traditional medical insurance knowledge base management systems generally have problems of insufficient handling of data heterogeneity and noise. The sources of medical insurance data are complex, involving heterogeneous data in various formats and structures, such as text, images, tables, etc. Existing technologies often lack effective standardization and normalization means when dealing with these heterogeneous data, resulting in difficult data integration, uneven data quality, and further exacerbating the complexity of the knowledge base in retrieval and application. In addition, the existence of noise data is also likely to lead to a decline in the accuracy and reliability of the knowledge base content, affecting the actual application effect of the system.
[0005] Furthermore, the existing technical means for semantic analysis and information extraction in medical insurance knowledge bases are relatively single, and cannot effectively capture and understand complex medical semantic relationships. Traditional semantic analysis methods often rely on simple rules or models, lacking in-depth understanding of context and semantic relationships. This method is prone to problems such as incomplete information extraction or inaccurate classification when dealing with complex medical text data, resulting in a large amount of redundant information or incorrect information in the construction and management of the knowledge base, further reducing the application value of the knowledge base.
[0006] In addition, in retrieval-augmented generation, the query input by the user is usually vectorized and embedded, matched to approximate documents, and then sentences suspected to be relevant are further recalled. Such a retrieval process is slow, inaccurate, and requires a high degree of knowledge intensity. The knowledge classification system of this method is not unified, resulting in low storage and retrieval efficiency of knowledge; and the relevance and accuracy of retrieval results for professional terms are not high, and the user experience is poor.
[0007] Finally, the performance of traditional medical insurance knowledge bases in intelligent question answering and information recall is also relatively limited. When dealing with user queries, existing systems often rely on simple keyword matching and rule retrieval, unable to fully understand the query intent of users, and thus unable to provide accurate answers. With the rapid development of big data technology and artificial intelligence, existing systems are unable to cope and are difficult to make full use of massive data resources to improve the user experience.
[0008] Therefore, how to provide a method and system for hierarchical construction and classification management of a medical insurance knowledge base is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0009] An object of the present invention is to propose a method and system for hierarchical construction and classification management of a medical insurance knowledge base. The present invention details the whole process from user requirement collection, data sorting and classification, knowledge base construction to an intelligent question answering system by applying context-aware semantic slicing technology, hierarchical nested clustering algorithm, entropy-driven adaptive classification algorithm and homotopy continuous algorithm. This method has the advantages of strong dynamic adaptability, precise data processing, profound semantic understanding, intelligent classification management and high retrieval accuracy, effectively improving the efficiency and accuracy of large-scale data management and application in the field of medical insurance.
[0010] According to an embodiment of the present invention, a method and system for hierarchical construction and classification management of a medical insurance knowledge base includes the following steps:
[0011] S1. Collect and analyze user requirements, and set the goals, uses, content scopes, classification criteria and structural requirements of the knowledge base;
[0012] S2. Obtain medical insurance-related data that matches user requirements from different sources, and perform sorting, screening, standardization and normalization processing on the data to generate a cleaned data set, eliminating heterogeneity and noise;
[0013] S3. Apply context-aware semantic slicing technology to decompose the cleaned data set into multiple semantically related segments, perform semantic matching and information extraction in combination with the context information of each segment, extract key feature data, and perform preliminary classification on the data according to the set classification criteria, and then use bce-embedding-base v1. The model vectorizes and embeds this classified knowledge;
[0014] S4. Apply the hierarchical nested clustering algorithm to the vectorized and embedded data for organization, construct a multi-level medical insurance knowledge base, refine the hierarchical structure of the data through layer-by-layer clustering, and store it according to the hierarchical relationship. The clustering result of each layer is associated with a specific medical insurance theme or domain;
[0015] S5. Apply the entropy-driven adaptive classification algorithm to the classified data stored in the knowledge base for dynamic management and automatic classification. Based on the entropy value of the data, adjust the classification criteria, data labels, and weights in real time to adapt to user needs and policy changes;
[0016] S6. Conduct cross-domain association analysis on the classified and managed data in the knowledge base by applying the homotopy continuous algorithm to construct knowledge links across categories and levels;
[0017] S7. Use the bge-large-zh model to recall knowledge from the data in the knowledge base. By analyzing the user's query content and the generated knowledge links, identify the knowledge points most relevant to the user's needs, and use the bge-reranker-large model to re-rank the recall results. Convert the optimized retrieval results into prompts, and finally generate intelligent answers through the Qwen-14B-chat large language model;
[0018] S8. Regularly audit and update the content of the knowledge base. Combine the latest policies, regulations, and technological development trends to optimize the content, classification system, and management methods of the knowledge base, so that the knowledge base can continuously reflect the latest industry information.
[0019] Optionally, S3 includes the following steps:
[0020] S31. Preprocess the cleaned data set to generate a set of data fragments through word segmentation and part-of-speech tagging;
[0021] S32. Apply context-aware semantic slicing technology to decompose the preprocessed set of data fragments into multiple semantic slices. Combine time series analysis methods to capture the time series semantic trends between data fragments and generate a set of semantic slices containing time features;
[0022] S33. Combine context information and time series analysis to perform semantic matching and information extraction on the set of semantic slices, extract key feature data, and form feature data containing the time dimension;
[0023] S34. Classify the feature data according to the set classification criteria to construct a time-weighted semantic feature matrix W1(t):
[0024]
[0025] Among them, represents the weighted semantic association degree of classification label i with respect to feature data j at time t, and f ij (t) represents the direct association degree of label i with feature data j at time t, T represents the total number of context slices, is the semantic similarity between label i and context slice k, and g kj (t) is the semantic association degree of context slice k with feature data j at time t, γ is the attenuation coefficient, and α and β are parameters for adjusting weights, is the time decay factor, where t k is the timestamp of context slice k;
[0026] S35. Input the time-weighted semantic feature matrix W1(t) into the bce-embedding-base v 1 model for vectorized embedding to generate a set of knowledge representation vectors for the time series;
[0027] S36. Use the set of knowledge representation vectors for index construction and knowledge base organization to support hierarchical management and classification retrieval in the time dimension.
[0028] Optionally, S4 includes the following steps:
[0029] S41. Take the generated time-weighted semantic feature matrix W1(t) and the set of vectorized embedding data as inputs, apply the hierarchical clustering algorithm to organize the data, perform clustering at different levels, and generate clustering results at multiple levels;
[0030] S42. At each level, use the semantically consistent-driven hierarchical clustering algorithm to construct a hierarchical structure tree, where each hierarchical structure tree corresponds to a specific medical insurance theme or domain, represents the clustering result at that level, and perform clustering based on the complex semantic consistency measure C(v i , v j , t):
[0031]
[0032] Among them, C(v i , v j , t) represents the complex semantic consistency measure of feature vectors v i (t) and v i (t) at time t, w ik (t) and w jk (t) represent the weights of the i-th and j-th feature vectors in the k-th dimension of the time-weighted semantic feature matrix W(t), is the time decay factor, γ is the attenuation coefficient, sik and s jk represents the feature vector v i (t) and v j (t) at the k-th context slice's semantic similarity;
[0033] S43. According to the calculated complex semantic consistency measure C(v i , v j , t), construct the clustering results layer by layer. Each clustering result generates a corresponding hierarchical structure tree and maps it to a specific medical insurance theme or field to form a hierarchical knowledge base structure, and dynamically adjust the clustering parameters to make the clustering results at each level have high semantic consistency;
[0034] S44. Integrate the clustering results with the existing knowledge base structure, update the hierarchical structure of the knowledge base, make the newly added clustering results compatible with the existing system, and finally form a multi-level medical insurance knowledge base structure;
[0035] S45. Build an index for the updated knowledge base structure to provide basic data support for knowledge retrieval and management.
[0036] Optionally, the said S5 includes the following steps:
[0037] S51. Take the generated clustering results and the time-weighted semantic feature matrix W(t) as inputs, apply the entropy-driven adaptive classification algorithm to dynamically manage these data, and calculate the weighted comprehensive semantic entropy H k (t):
[0038]
[0039] where w ij (t) is the weight of the i-th class on the j-th feature in the time-weighted semantic feature matrix W(t), C(v i , v j , t) is the complex semantic consistency measure used to measure the semantic consistency between the feature vectors v i (t) and v j (t), and p ik (t) represents the data occurrence probability of the i-th class on the feature dimension k;
[0040] S52. Calculate the overall weighted entropy change rate ΔH(t) for dynamic entropy regulation, and adjust the classification criterion λ(t) and the classification weight w o (t) based on this regulation mechanism:
[0041]
[0042] Among them, ΔH(t) represents the overall weighted entropy change rate, which measures the change in semantic complexity at the current time t. λ0 is the initial classification criterion, and β is a regulation coefficient used to control the influence of the entropy change rate on the classification criterion.
[0043] S53. Reclassify the data using the dynamically adjusted classification criterion λ(t), and dynamically adjust the data labels and classification weights w i (t) according to the real-time monitored user needs and policy changes, so that the weights are dynamically adjusted with time and data changes;
[0044] S54. Apply the updated classification results to the knowledge base, and organize and allocate the data in the knowledge base according to the classification weights w i (t), and update the knowledge base index at the same time to make it reflect the latest classification structure and data weights.
[0045] Optionally, the S6 includes the following steps:
[0046] S61. Extract the feature data of various categories and levels from the classified dataset, initialize the multi-dimensional homotopy path, and generate the initial path matrix and weight matrix on multiple data dimensions;
[0047] S62. In each iteration process, dynamically adjust the weight matrix W(t) and the path matrix P(t) based on the change of data features, and apply the multi-dimensional homotopy path tracking algorithm:
[0048]
[0049] Among them, ΔP ij (t) represents the adjustment amount of the i-th and j-th categories in the path matrix at time t. η(t) is the adaptive step size, which is adjusted based on the entropy gradient of the path w ij (t) is the weight of the i-th and j-th categories in the time-weighted semantic feature matrix W(t). C(v i ,v j ,t) is the complex semantic consistency metric, which combines time decay and semantic similarity to measure the semantic consistency between the feature vectors v i (t) and v j (t), is the time decay factor, which controls the influence of the semantic features far from the current time t on the current calculation. γ is the time decay coefficient, is the logarithmic term based on entropy;
[0050] S63. Combine the results of multi-dimensional path tracking, calculate and update the association degrees between various categories and levels, iteratively update the association matrix until the matrix change amount is less than the preset threshold, terminate the iteration, and generate the final cross-category and cross-level association matrix.
[0051] S64. Construct a cross - category and hierarchical knowledge link network in the knowledge base according to the final association matrix, integrate these links with the existing knowledge base hierarchy to form a cross - domain knowledge link network;
[0052] S65. Integrate the knowledge link network with the hierarchy of the existing knowledge base, update the knowledge base hierarchy to reflect the latest associations between different categories and levels, and generate corresponding indexes to support cross - domain knowledge retrieval and application.
[0053] Optionally, the said S7 includes the following steps:
[0054] S71. Receive the query input by the user, tokenize it, remove stop words and perform part - of - speech tagging to generate a query vector q(t);
[0055] S72. Input the query vector q(t) into the distributed knowledge recall and question - answering system, parallel - process the knowledge base data set through multiple distributed computing nodes, and apply the bge - large - zh model for knowledge recall:
[0056]
[0057] Among them, S m (q(t), v i (t)) represents the multi - level weighted semantic similarity between the query vector q(t) and the knowledge vector v i (t), w ik (t) is the weight of the i - th class on the k - th feature in the time - weighted semantic feature matrix, p ok (t) represents the data occurrence probability of the i - th class on the feature dimension k, C(q k (t), v ik (t)) is the semantic consistency measure between the query vector and the knowledge vector on the feature k, is the time decay factor;
[0058] S73. Based on the calculated multi - level weighted semantic similarity S m (q(t), v i (t)), recall the most relevant knowledge points to the query content on each node to form a preliminary knowledge vector set;
[0059] S74. Through the coordination mechanism of the distributed knowledge recall and question - answering system, summarize the preliminary knowledge vector sets of each node and input them into the bge - reranker - large model for rearrangement to generate an optimized knowledge vector set;
[0060] S75. Convert the optimized knowledge vector set into an input prompt and send it to the Qwen-14B-chat large language model to generate an intelligent answer that matches the user's query;
[0061] S76. According to the answer generated by Qwen-14B-chat, combined with the cross-category and hierarchical knowledge link network in the knowledge base, return the generated answer to the user, and dynamically update the content and structure of the knowledge base according to the user feedback, and synchronously update the knowledge base content and index structure through distributed nodes.
[0062] A hierarchical construction and classification management system for a medical insurance knowledge base, including:
[0063] Data collection module: used to collect medical insurance-related data from multiple data sources, support the collection of multi-modal data such as text, images, audio, and video, and integrate it into the system through an interface;
[0064] Data preprocessing module: perform denoising, standardization, and normalization on the collected raw data, eliminate data heterogeneity and noise, extract key features at the same time, and decompose the data into semantic fragments to generate a cleaned data set for subsequent processing steps;
[0065] Semantic slicing and information extraction module: apply context-aware semantic slicing technology to perform semantic slicing on the preprocessed data set, perform semantic matching and information extraction on the data fragments combined with context information, generate key feature data with time series information, and embed it into the vector space;
[0066] Hierarchical knowledge base construction module: use the hierarchical nested clustering algorithm to organize the vectorized and embedded data according to the hierarchical classification structure, refine the hierarchical structure of the data layer by layer, and store it in the knowledge base. Each layer is associated with a specific medical insurance theme or field;
[0067] Dynamic classification and management module: dynamically manage and automatically classify the classified data stored in the knowledge base through an entropy-driven adaptive classification algorithm, and adjust the classification criteria, data labels, and weights in real time to cope with user needs and policy changes;
[0068] Cross-domain association analysis module: based on the homotopy continuous algorithm, perform cross-domain association analysis on the classified and managed data, construct cross-category and hierarchical knowledge links, form a cross-domain knowledge network, and optimize the relevance and comprehensive utilization value of the knowledge base;
[0069] Intelligent Knowledge Recall and Q&A Module: It converts the query input by the user into a query vector, and through a distributed knowledge recall and Q&A system, it parallelly processes the knowledge base data on multiple distributed computing nodes, applies a semantic similarity calculation model for knowledge recall and rearrangement, generates an intelligent answer most relevant to the user's query, and provides personalized Q&A services through the Qwen-14B-chat large language model;
[0070] Security Protection Module: In all processing stages of the system, including data collection, processing, storage, and transmission, it applies security policies such as data encryption algorithms, access privilege management, and audit log recording to optimize data security and privacy protection.
[0071] The beneficial effects of the present invention are as follows:
[0072] (1) By combining context-aware semantic slicing technology, hierarchical nested clustering algorithm, and entropy-driven adaptive classification algorithm, the present invention significantly improves the data processing and classification management efficiency of the medical insurance knowledge base. Especially in processing heterogeneous data, achieving dynamic classification, and semantic understanding, the present invention overcomes the limitations of traditional static knowledge base systems, can accurately extract and classify medical insurance data, and form a hierarchical and adaptable knowledge base structure.
[0073] (2) By applying the homotopy continuation algorithm for cross-domain correlation analysis, the present invention constructs a knowledge link network across categories and levels. This technology shows a high level of intelligence in correlation analysis and knowledge link generation, can effectively manage and integrate large-scale medical insurance data, support complex multi-domain data associations, and thus improve the adaptability and practicality of the knowledge base in different application scenarios.
[0074] (3) The present invention introduces an intelligent knowledge recall and Q&A system, which uses distributed knowledge recall technology and large language models, and through a multi-level optimization and rearrangement process, enhances the understanding and response ability to user queries. This system is outstanding in retrieval accuracy and user interaction experience, can generate accurate and intelligent answers according to user needs, and significantly improves the query efficiency and user satisfaction of the medical insurance knowledge base.
[0075] (4) By comprehensively applying the above advanced technologies, the present invention provides an efficient, dynamic, and adaptive medical insurance knowledge base management solution. The system can automatically process and update medical insurance data, optimize the structure and content of the knowledge base, and at the same time ensure the real-time nature and accuracy of the knowledge base under the background of rapidly changing policies and requirements. Description of the Drawings
[0076] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0077] Figure 1 is a flowchart of a method for hierarchical construction and classification management of a medical insurance knowledge base proposed by the present invention;
[0078] Figure 2 is a working flowchart of the context-aware semantic slicing technology proposed by the present invention. Detailed implementation manners
[0079] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0080] Refer to Figure 1-2 , a method and system for hierarchical construction and classification management of a medical insurance knowledge base, including the following steps:
[0081] S1. Collect and analyze user requirements, and set the goals, uses, content scopes, classification criteria, and structural requirements of the knowledge base;
[0082] S2. Obtain medical insurance-related data that matches the user requirements from different sources, and perform sorting, screening, standardization, and normalization processing on the data to generate a cleaned data set, eliminating heterogeneity and noise;
[0083] S3. Apply the context-aware semantic slicing technology to decompose the cleaned data set into multiple semantically related segments, perform semantic matching and information extraction by combining the context information of each segment, extract key feature data, and perform preliminary classification on the data according to the set classification criteria, and then use the bce-embedding-base v 1 model to vectorize and embed these classified knowledge;
[0084] S4. Organize the vectorized and embedded data by applying a hierarchical nested clustering algorithm to construct a multi-level medical insurance knowledge base, refine the hierarchical structure of the data through layer-by-layer clustering, and store it according to the hierarchical relationship. The clustering result of each layer is associated with a specific medical insurance theme or field;
[0085] S5. Apply an entropy-driven adaptive classification algorithm to dynamically manage and automatically classify the classified data stored in the knowledge base, and adjust the classification criteria, data labels, and weights in real time based on the entropy value of the data to adapt to user requirements and policy changes;
[0086] S6. Apply the homotopy continuation algorithm to the classified and managed data in the knowledge base for cross - domain correlation analysis, and construct knowledge links across categories and hierarchies;
[0087] S7. Use the bge - large - zh model to recall knowledge from the data in the knowledge base. By analyzing the user's query content and the generated knowledge links, identify the knowledge points most relevant to the user's needs, and use the bge - reranker - large model to re - rank the recall results. Convert the optimized retrieval results into prompts, and finally generate intelligent answers through the Qwen - 14B - chat large language model;
[0088] S8. Regularly review and update the content of the knowledge base. Combine the latest policies, regulations, and technological development trends to optimize the content, classification system, and management methods of the knowledge base, so that the knowledge base can continuously reflect the latest industry information.
[0089] In this embodiment, S3 includes the following steps:
[0090] S31. Pre - process the cleaned data set, and generate a set of data fragments through word segmentation and part - of - speech tagging;
[0091] S32. Apply context - aware semantic slicing technology to decompose the pre - processed set of data fragments into multiple semantic slices. Combine time - series analysis methods to capture the time - series semantic trends between data fragments, and generate a set of semantic slices containing time features;
[0092] S33. Combine context information and time - series analysis to perform semantic matching and information extraction on the set of semantic slices, extract key feature data, and form feature data containing the time dimension;
[0093] S34. Classify the feature data according to the set classification criteria, and construct a time - weighted semantic feature matrix W1(t):
[0094]
[0095] where, represents the weighted semantic association degree of classification label i with feature data j at time t, f ij (t) represents the direct association degree of label i with feature data j at time t, T represents the total number of context slices, is the semantic similarity between label i and context slice k, g kj (t) is the semantic association degree of context slice k with feature data j at time t, γ is the attenuation coefficient, α and β are parameters for adjusting weights, is the time decay factor, where t k is the timestamp of context slice k;
[0096] S35. Input the time-weighted semantic feature matrix W1(t) into the bce-embedding-base v 1 model for vectorized embedding to generate a set of knowledge representation vectors for the time series;
[0097] S36. Use the set of knowledge representation vectors for index construction and knowledge base organization to support hierarchical management and classification retrieval in the time dimension.
[0098] In this embodiment, S4 includes the following steps:
[0099] S41. Take the generated time-weighted semantic feature matrix W1(t) and the vectorized embedding data set as inputs, apply the hierarchical clustering algorithm to organize the data, perform clustering at different levels, and generate clustering results at multiple levels;
[0100] S42. At each level, use the semantically consistent-driven hierarchical clustering algorithm to construct a hierarchical structure tree, where each hierarchical structure tree corresponds to a specific medical insurance theme or domain, representing the clustering result at that level, and perform clustering based on the complex semantic consistency measure C(v i ,v j ,t):
[0101]
[0102] Among them, C(v i ,v j ,t) represents the complex semantic consistency measure of the feature vectors v i (t) and v j (t) at time t, w ik (t) and w jk (t) represent the weights of the i-th and j-th feature vectors in the k-th dimension of the time-weighted semantic feature matrix W(t), is the time decay factor, γ is the decay coefficient, s ik and s jk represent the semantic similarity of the feature vectors v i (t) and v j (t) in the k-th context slice;
[0103] S43. According to the calculated complex semantic consistency measure C(v i ,v j ,t), construct the clustering results layer by layer. Each clustering result generates a corresponding hierarchical structure tree and maps it to a specific medical insurance theme or domain to form a hierarchical knowledge base structure, and dynamically adjust the clustering parameters to make the clustering results at each level highly consistent semantically;
[0104] S44. Integrate the clustering results with the existing knowledge base structure, update the hierarchical structure of the knowledge base, make the newly added clustering results compatible with the existing system, and finally form a multi-level medical insurance knowledge base structure;
[0105] S45. Build an index for the updated knowledge base structure to provide basic data support for knowledge retrieval and management.
[0106] In this embodiment, the said S5 includes the following steps:
[0107] S51. Take the generated clustering results and the time-weighted semantic feature matrix W(t) as inputs, apply the entropy-driven adaptive classification algorithm to dynamically manage these data, and calculate the weighted comprehensive semantic entropy H k (t) of each feature dimension k in the clustering:
[0108]
[0109] where w ij (t) is the weight of the i-th class on the j-th feature in the time-weighted semantic feature matrix W(t), and C(v i , v j , t) is the complex semantic consistency measure, used to measure the semantic consistency between the feature vectors v i (t) and v j (t), and p ik (t) represents the data occurrence probability of the i-th class on the feature dimension k;
[0110] S52. Calculate the overall weighted entropy change rate ΔH(t) for dynamic entropy regulation, and adjust the classification criterion λ(t) and the classification weight w i (t) based on this regulation mechanism:
[0111]
[0112] where ΔH(t) represents the overall weighted entropy change rate, measuring the change in semantic complexity at the current time t, λ0 is the initial classification criterion, and β is the adjustment coefficient, used to control the influence of the entropy change rate on the classification criterion;
[0113] S53. Reclassify the data using the dynamically adjusted classification criterion λ(t), and dynamically adjust the data labels and the classification weight w i (t) according to the real-time monitored user needs and policy changes, so that the weight dynamically adjusts with time and data;
[0114] S54. Apply the updated classification results to the knowledge base, and according to the classification weight w i(t) Organize and distribute the data in the knowledge base, and update the knowledge base index simultaneously to make it reflect the latest classification structure and data weights.
[0115] In this embodiment, S6 includes the following steps:
[0116] S61. Extract the feature data of each category and level from the classified dataset, initialize the multi-dimensional homotopy path, and generate the initial path matrix and weight matrix on multiple data dimensions;
[0117] S62. In each iteration process, dynamically adjust the weight matrix W(t) and the path matrix P(t) based on the change of data features, and apply the multi-dimensional homotopy path tracking algorithm:
[0118]
[0119] where, ΔP ij (t) represents the adjustment amount of the i-th and j-th categories in the path matrix at time t, η(t) is the adaptive step size, and it is adjusted based on the entropy gradient of the path w ij (t) is the weight of the i-th and j-th categories in the time-weighted semantic feature matrix W(t), C(v i , v j , t) is the complex semantic consistency measure, which combines time decay and semantic similarity and is used to measure the semantic consistency between the feature vectors v i (t) and v j (t), is the time decay factor, which controls the influence of semantic features far from the current time t on the current calculation, γ is the time decay coefficient, is the logarithmic term based on entropy;
[0120] S63. Combine the results of multi-dimensional path tracking, calculate and update the association degree between each category and level, iteratively update the association matrix until the matrix change amount is less than the preset threshold, terminate the iteration and generate the final cross-category and cross-level association matrix;
[0121] S64. According to the final association matrix, construct a cross-category and cross-level knowledge link network in the knowledge base, integrate these links with the existing knowledge base hierarchy to form a cross-domain knowledge link network;
[0122] S65. Integrate the knowledge link network with the existing knowledge base hierarchy, update the knowledge base hierarchy to make it reflect the latest association between different categories and levels, and generate the corresponding index to support cross-domain knowledge retrieval and application.
[0123] In this embodiment, S7 includes the following steps:
[0124] S71. Receive the query input by the user, tokenize it, remove stop words, and perform part-of-speech tagging to generate the query vector q(t);
[0125] S72. Input the query vector q(t) into the distributed knowledge recall and question-answering system, and parallelly process the knowledge base data set through multiple distributed computing nodes, and apply the bge-large-zh model for knowledge recall:
[0126]
[0127] Among them, S m (q(t), v i (t)) represents the multi-level weighted semantic similarity between the query vector q(t) and the knowledge vector v i (t), w ik (t) is the weight of the i-th category on the k-th feature in the time-weighted semantic feature matrix, p ik (t) represents the data occurrence probability of the i-th category on the feature dimension k, C(q k (t), v ik (t)) is the semantic consistency measure between the query vector and the knowledge vector on the feature k, is the time decay factor;
[0128] S73. Based on the calculated multi-level weighted semantic similarity S m (q(t), v i (t)), recall the most relevant knowledge points to the query content on each node to form a preliminary knowledge vector set;
[0129] S74. Through the coordination mechanism of the distributed knowledge recall and question-answering system, aggregate the preliminary knowledge vector sets of each node and input them into the bge-reranker-large model for rearrangement to generate an optimized knowledge vector set;
[0130] S75. Convert the optimized knowledge vector set into an input prompt and send it to the Qwen-14B-chat large language model to generate an intelligent answer matching the user's query;
[0131] S76. According to the answer generated by Qwen-14B-chat, combined with the cross-category and hierarchical knowledge link network in the knowledge base, return the generated answer to the user, and dynamically update the content and structure of the knowledge base according to the user feedback, and synchronously update the knowledge base content and index structure through distributed nodes.
[0132] A hierarchical construction and classification management system for a medical insurance knowledge base, including:
[0133] Data Acquisition Module: Used to collect medical insurance-related data from multiple data sources, support the acquisition of multi-modal data such as text, images, audio, and video, and integrate it into the system through an interface;
[0134] Data Preprocessing Module: Denoise, standardize, and normalize the collected raw data, eliminate data heterogeneity and noise, extract key features, and decompose the data into semantic segments to generate a cleaned dataset for subsequent processing steps;
[0135] Semantic Slicing and Information Extraction Module: Apply context-aware semantic slicing technology to perform semantic slicing on the preprocessed dataset, perform semantic matching and information extraction on data segments in combination with context information, generate key feature data with time series information, and embed it into the vector space;
[0136] Hierarchical Knowledge Base Construction Module: Use hierarchical nested clustering algorithms to organize the vectorized and embedded data according to a hierarchical classification structure, refine the hierarchical structure of the data layer by layer, and store it in the knowledge base. Each layer is associated with a specific medical insurance theme or domain;
[0137] Dynamic Classification and Management Module: Dynamically manage and automatically classify the classified data stored in the knowledge base through an entropy-driven adaptive classification algorithm, and adjust classification criteria, data labels, and weights in real time to cope with user needs and policy changes;
[0138] Cross-Domain Association Analysis Module: Based on the homotopy continuation algorithm, perform cross-domain association analysis on the classified and managed data, construct knowledge links across categories and hierarchies, form a cross-domain knowledge network, and optimize the relevance and comprehensive utilization value of the knowledge base;
[0139] Intelligent Knowledge Retrieval and Q&A Module: Convert the user input query into a query vector, and through a distributed knowledge retrieval and Q&A system, parallel process the knowledge base data on multiple distributed computing nodes, apply a semantic similarity calculation model for knowledge retrieval and re-ranking, generate an intelligent answer most relevant to the user query, and provide personalized Q&A services through the Qwen-14B-chat large language model;
[0140] Security Protection Module: Apply security policies such as data encryption algorithms, access rights management, and audit log recording in all processing stages of the system, including data acquisition, processing, storage, and transmission, to optimize data security and privacy protection.
[0141] Example 1:
[0142] To verify the feasibility of the present invention in implementation, the present invention is applied to a large comprehensive hospital located in Shanghai, China, which faces challenges in managing and updating the medical insurance policy knowledge base. Due to the frequent changes in medical insurance policies and the complexity of medical information involved, the traditional manual management method can no longer meet the hospital's needs. Therefore, this embodiment introduces a method and system for hierarchical construction and classification management of a medical insurance knowledge base proposed by the present invention to improve the work efficiency and accuracy of the hospital in medical insurance policy management.
[0143] First, the hospital's data management team cleaned and standardized the existing medical and medical insurance data. Using the method of the present invention, the system automatically removed the noise and heterogeneity in the data and generated a unified standardized data set. This data includes the diagnosis information, treatment plans, medication records, and relevant medical insurance reimbursement situations of patients. After cleaning and standardizing, the consistency and accuracy of the data have been significantly improved, providing a reliable basis for subsequent data processing.
[0144] Next, the system applies context-aware semantic slicing technology to deeply analyze the cleaned data set. This technology can decompose the data set into multiple semantically related segments and perform semantic matching and information extraction by combining the context information of each segment, extracting key feature data from it. For example, when processing case data related to "diabetes", the system can automatically extract key information related to the latest medical insurance policy, providing a scientific basis for doctors when prescribing treatment plans and medical insurance reimbursement.
[0145] Subsequently, the system uses a hierarchical nested clustering algorithm to organize the extracted feature data and construct a multi-level medical insurance knowledge base. This algorithm refines and classifies the data according to different dimensions through layer-by-layer clustering. For example, when processing medical data related to "hypertension", the system can cluster and store policy documents from different years according to the time dimension, providing doctors with a comparison of historical and current policies to help them make more reasonable treatment decisions.
[0146] During the actual operation of the system, the hospital's data analysis team dynamically manages the classified data in the knowledge base through an entropy-driven adaptive classification algorithm. The system can automatically adjust the classification criteria, data labels, and weights according to the change in the data entropy value monitored in real time to adapt to the changing needs of the hospital and new medical insurance policies. For example, when the government issues a new medical insurance policy, the system can automatically classify and adjust the relevant data and update the content in the medical insurance knowledge base to ensure the timeliness and accuracy of the policy.
[0147] To further enhance the application effect of medical insurance policies, the system applies the homotopy continuation algorithm to conduct cross-domain correlation analysis on the classified and managed data. The system can automatically identify and construct knowledge links across categories and levels. For example, it correlates data related to "cardiovascular diseases" with that of "diabetes" and links relevant medical insurance policies to the data of these diseases, thus providing comprehensive medical insurance information for doctors when dealing with complex medical conditions.
[0148] In actual operation, when a doctor or administrator enters a query, the system, through the distributed knowledge recall and question-answering system, quickly identifies knowledge points related to the user's query. The system matches the query content with the data in the medical insurance knowledge base and generates the optimal retrieval results through an intelligent knowledge recall mechanism and provides them to the user. For example, when a doctor queries the "medical insurance reimbursement policy for hypertension in 2023", the system can not only provide the current policy details but also relevant historical data and policy comparisons to support the doctor's decision-making.
[0149] Table 1: Evaluation of the Implementation Effect of the Hospital Medical Insurance Management System
[0150]
[0151] By evaluating the actual application effect of the system, we found that after the hospital introduced the method of the present invention, the processing efficiency of medical insurance policies increased by approximately 40%, and the policy update response time was shortened by 50%. Within the first quarter of implementation, the error rate of medical insurance policy application in the hospital decreased by 60%. Specific data shows that when facing complex medical conditions, the system can automatically provide relevant policy links, significantly improving the accuracy of doctors' decision-making. At the same time, the satisfaction of patients with the hospital has also increased significantly, and the work efficiency and service quality of the hospital have been comprehensively improved.
[0152] Generally speaking, the implementation of the present invention effectively solves many problems faced in traditional medical insurance policy management, significantly improves the data processing efficiency and accuracy, enhances the hospital's ability to handle complex policies, and provides better medical services for patients.
[0153] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A method for hierarchical construction and classification management of a medical insurance knowledge base, characterized in that It includes the following steps: S1. Collect and analyze user requirements, and set the goals, uses, content scope, classification criteria, and structural requirements of the knowledge base; S2. Obtain medical insurance-related data that matches user requirements from different sources, and organize, screen, standardize, and normalize the data to generate a cleaned dataset, eliminating heterogeneity and noise; S3. Apply the context-aware semantic slicing technology to decompose the cleaned dataset into multiple semantically related segments, perform semantic matching and information extraction by combining the context information of each segment, extract key feature data, and preliminarily classify the data according to the set classification criteria. Then use the bce-embedding-base v 1 model to vectorize and embed this classified knowledge; S4. Apply the hierarchical nested clustering algorithm to organize the vectorized and embedded data, construct a multi-level medical insurance knowledge base, refine the hierarchical structure of the data through layer-by-layer clustering, and store it according to the hierarchical relationship. The clustering results at each level are associated with specific medical insurance topics or domains; S5. Apply the entropy-driven adaptive classification algorithm to dynamically manage and automatically classify the classified data stored in the knowledge base, and adjust the classification criteria, data labels, and weights in real time based on the entropy value of the data to adapt to user needs and policy changes; S6. Conduct cross-domain association analysis on the classified and managed data in the knowledge base by applying the homotopy continuous algorithm to construct knowledge links across categories and levels; S7. Use the bge-large-zh model to recall knowledge from the data in the knowledge base. By analyzing the user's query content and the generated knowledge links, identify the knowledge points most relevant to the user's needs, and use the bge-reranker-large model to re-rank the recall results. Convert the optimized retrieval results into prompts, and finally generate intelligent answers through the Qwen-14B-chat large language model; S8. Regularly review and update the content of the knowledge base, and optimize the content, classification system, and management methods of the knowledge base in combination with the latest policies, regulations, and technological development trends, so that the knowledge base can continuously reflect the latest industry information.
2. The method for hierarchical construction and classification management of a medical insurance knowledge base according to claim 1, characterized in that, The specific content of S3 includes: S31. Preprocess the cleaned dataset to generate a set of data segments through word segmentation and part-of-speech tagging; S32. Apply context-aware semantic slicing technology to decompose the preprocessed set of data segments into multiple semantic slices, and combine time series analysis methods to capture the time series semantic trends between data segments, generating a set of semantic slices containing time features; S33. Combine context information and time series analysis to perform semantic matching and information extraction on the set of semantic slices, extract key feature data, and form feature data containing the time dimension; S34. Classify the feature data according to the set classification criteria to construct a time-weighted semantic feature matrix W1(t): Among them, represents the weighted semantic association degree of classification label i with feature data j at time t, and f ij (t) represents the direct association degree of label i with feature data j at time t, T represents the total number of context slices, is the semantic similarity between label i and context slice k, and g kj (t) is the semantic association degree of context slice k with feature data j at time t, γ is the decay coefficient, and α and β are parameters for adjusting weights, is the time decay factor, where t k is the timestamp of context slice k; S35. Input the time-weighted semantic feature matrix W1(t) into the bce-embedding-base v 1 model for vectorized embedding to generate a set of knowledge representation vectors for the time series; S36. Use the set of knowledge representation vectors for index construction and knowledge base organization to support hierarchical management and classification retrieval in the time dimension.
3. A method for hierarchical construction and classification management of a medical insurance knowledge base according to claim 1, characterized in that, The specific content of S4 includes: S41. Use the generated time-weighted semantic feature matrix W1(t) and the set of vectorized and embedded data as inputs, apply the hierarchical clustering algorithm to organize the data, perform clustering at different levels, and generate clustering results at multiple levels; S42. At each level, a hierarchical structure tree is constructed using a semantic consistency-driven hierarchical clustering algorithm, where each hierarchical structure tree corresponds to a specific medical insurance theme or domain, representing the clustering result at that level, and clustering is performed based on the complex semantic consistency measure C(v i , v j , t) between feature vectors: Among them, C(v i , v i , t) represents the complex semantic consistency measure of the feature vectors v i (t) and v j (t) at time t, w ik (t) and w jk (t) represent the weights of the i-th and j-th feature vectors in the k-th dimension of the time-weighted semantic feature matrix W(t), is the time decay factor, γ is the decay coefficient, s ik and s jk represent the semantic similarity of the feature vectors v i (t) and v j (t) in the k-th context slice; S43. According to the calculated complex semantic consistency metric C(v i , v j , t), construct the clustering results layer by layer. Each clustering result generates a corresponding hierarchical structure tree, and maps it to a specific medical insurance theme or field to form a well-structured knowledge base. Dynamically adjust the clustering parameters to make the clustering results at each level have high semantic consistency; S44. Integrate the clustering results with the existing knowledge base structure, update the hierarchical structure of the knowledge base, and make the newly added clustering results compatible with the existing system, finally forming a multi-level medical insurance knowledge base structure; S45. Index the updated knowledge base structure to provide basic data support for knowledge retrieval and management.
4. A method for hierarchical construction and classification management of a medical insurance knowledge base according to claim 1, characterized in that, The S5 specifically includes: S51. Taking the generated clustering results and the time-weighted semantic feature matrix W(t) as inputs, apply an entropy-driven adaptive classification algorithm to dynamically manage this data, and calculate the weighted comprehensive semantic entropy H of each feature dimension k in the clustering: k (t): where, w ij (t) is the weight of the i-th class on the j-th feature in the time-weighted semantic feature matrix W(t), C(v i , v j , t) is the complex semantic consistency measure, which is used to measure the semantic consistency between the feature vectors v i (t) and v j (t), and p ik (t) represents the data occurrence probability of the i-th class on the feature dimension k; S52. Calculate the overall weighted entropy change rate ΔH(t) for dynamic entropy regulation, and adjust the classification criterion λ(t) and classification weight w i (t) based on this regulation mechanism: Among them, ΔH(t) represents the overall weighted entropy change rate, which measures the change in semantic complexity at the current time t, λ0 is the initial classification standard, and β is the adjustment coefficient, which is used to control the impact of the entropy change rate on the classification standard; S53. Reclassify the data using the dynamically adjusted classification criterion λ(t), and dynamically adjust the data labels and classification weights w i (t) according to the real-time monitored user requirements and policy changes, so that the weights are dynamically adjusted with time and data changes; S54. Apply the updated classification results to the knowledge base and organize and allocate the data in the knowledge base according to the classification weight w i (t), while updating the knowledge base index to reflect the latest classification structure and data weights.
5. A method for hierarchical construction and classification management of a medical insurance knowledge base according to claim 1, characterized in that, The S6 specifically includes: S61, extracting feature data of each category and level from the classified data set, initializing a multi-dimensional homology path, and generating an initial path matrix and a weight matrix in multiple data dimensions; S62. In each iteration, based on the changes in data features, the weight matrix W(t) and the path matrix P(t) are dynamically adjusted, and a multi-dimensional homotopy path tracking algorithm is applied: Among them, ΔP ii (t) represents the adjustment amount of the i-th and j-th categories in the path matrix at time t. η(t) is the adaptive step size, which is adjusted based on the entropy gradient of the path . w ij (t) is the weight of the i-th and j-th categories in the time-weighted semantic feature matrix W(t). C(v i , v j , t) is the complex semantic consistency metric, which combines time decay and semantic similarity and is used to measure the semantic consistency between the feature vectors v i (t) and v j (t). is the time decay factor that controls the influence of semantic features far from the current time t on the current calculation. γ is the time decay coefficient, is the logarithmic term based on entropy; S63, combining the results of multi-dimensional path tracking, calculating and updating the correlation between each category and level, iteratively updating the correlation matrix until the matrix change is less than a preset threshold, terminating the iteration and generating the final cross-category and cross-level correlation matrix; S64. Based on the final association matrix, a knowledge link network across categories and levels is constructed in the knowledge base, and these links are integrated with the existing knowledge base hierarchy to form a cross-domain knowledge link network; S65. Integrate the knowledge link network with the hierarchical structure of the existing knowledge base, update the hierarchical structure of the knowledge base to reflect the latest associations between different categories and levels, and generate corresponding indexes to support cross-domain knowledge retrieval and application.
6. A method for hierarchical construction and classification management of a medical insurance knowledge base according to claim 1, characterized in that, The S7 specifically includes: S71, receiving a query input by a user, segmenting it, removing stop words and performing part-of-speech tagging to generate a query vector q(t); S72. Input the query vector q(t) into the distributed knowledge recall and question-answering system, process the knowledge base data set in parallel through multiple distributed computing nodes, and apply the bge-large-zh model for knowledge recall: Among them, S m (q(t), v i (t)) represents the multi-level weighted semantic similarity between the query vector q(t) and the knowledge vector v i (t), w ik (t) is the weight of the i-th category on the k-th feature in the time-weighted semantic feature matrix, p ok (t) represents the data occurrence probability of the i-th category on the feature dimension k, C(q k (t), v ik (t)) is the semantic consistency measure between the query vector and the knowledge vector on the feature k, is the time decay factor; S73. Based on the calculated multi-level weighted semantic similarity S n (q(t), v i (t)), recall the most relevant knowledge points to the query content on each node to form a preliminary knowledge vector set; S74, through the coordination mechanism of distributed knowledge recall and question-answering system, the preliminary knowledge vector set of each node is summarized and input into the bge-reranker-large model for rearrangement to generate an optimized knowledge vector set; S75, converting the optimized knowledge vector set into an input prompt and sending it to the Qwen-14B-chat large language model to generate an intelligent answer that matches the user query; S76. Based on the answers generated by Qwen-14B-chat, combined with the cross-category and hierarchical knowledge link network in the knowledge base, the generated answers are returned to the user, and the knowledge base content and structure are dynamically updated based on user feedback, and the knowledge base content and index structure are synchronously updated through distributed nodes.
7. A hierarchical construction and classification management system for a medical insurance knowledge base, characterized in that, include: Data collection module: used to collect medical insurance related data from various data sources, supports the collection of multimodal data such as text, image, audio and video, and integrates it into the system through interfaces; Data Preprocessing Module: Denoise, standardize, and normalize the collected raw data to eliminate data heterogeneity and noise. Meanwhile, extract key features and decompose the data into semantic segments to generate a cleaned dataset for subsequent processing steps; Semantic Slicing and Information Extraction Module: Apply context-aware semantic slicing technology to perform semantic slicing on the preprocessed dataset. Match and extract information by combining data segments with context information to generate key feature data with time series information and embed it into the vector space; Hierarchical Knowledge Base Construction Module: Use the hierarchical nested clustering algorithm to organize the vectorized and embedded data according to the hierarchical classification structure, refine the hierarchical structure of the data layer by layer, and store it in the knowledge base. Each layer is associated with a specific medical insurance theme or domain; Dynamic Classification and Management Module: Dynamically manage and automatically classify the classified data stored in the knowledge base through an entropy-driven adaptive classification algorithm, and adjust the classification criteria, data labels, and weights in real time to cope with user needs and policy changes; Cross-Domain Association Analysis Module: Based on the homotopy continuation algorithm, perform cross-domain association analysis on the classified and managed data, construct knowledge links across categories and hierarchies, form a cross-domain knowledge network, and optimize the relevance and comprehensive utilization value of the knowledge base; Intelligent Knowledge Retrieval and Q&A Module: Convert the user input query into a query vector. Through a distributed knowledge retrieval and Q&A system, parallelly process the knowledge base data on multiple distributed computing nodes, apply a semantic similarity calculation model for knowledge retrieval and rearrangement, generate an intelligent answer most relevant to the user query, and provide a personalized Q&A service through the Qwen-14B-chat large language model; Security Protection Module: Apply security policies such as data encryption algorithms, access privilege management, and audit log recording in all processing stages of the system, including data collection, processing, storage, and transmission, to optimize data security and privacy protection.
Citation Information
Cited By
Hierarchical tree and auto-reflection-based retrieval enhancement generation method and device and medium
CN121051186A
Method, device and medium for search enhancement generation based on hierarchical tree and self-reflection
CN121051186B