Similar case recommendation method and system based on knowledge graph
By constructing a knowledge graph of medical information, using semantic correlations between cases, and determining similar nodes, it solves the problem that existing recommendation systems are difficult to identify semantic correlations of medical terms, and achieves efficient and accurate recommendations of similar cases.
Patent Information
- Application Number
- CN202411942336.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-06-06
AI Technical Summary
Existing recommendation systems are difficult to identify semantic associations of medical terms, and cannot identify synonyms or multi-level semantic associations, resulting in a decrease in comprehensiveness and accuracy of recommendation results.
Using the similar case recommendation method based on the knowledge graph, by constructing a historical case knowledge graph, medical information is used as nodes and attribute relationships are used as edges, and similar nodes between the case data to be queried and the historical case knowledge graph are determined.
A comprehensive capture of complex relationships between cases is achieved, the efficiency and accuracy of recommendations for similar cases is improved, and complex medical terms and nuances can be better handled.
Smart Images

Figure CN120108615A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical data processing, and in particular to a similar case recommendation method and system based on knowledge graph. Background Art
[0002] With the rapid development of deep learning technology, more and more deep learning models are being applied to various fields, especially in the medical field, such as medical image analysis, disease prediction, diagnostic decision support, etc. Such applications not only improve the accuracy of diagnosis, but also significantly reduce the workload of medical staff. In actual medical scenarios, doctors usually need to consult case data similar to the current patient's condition in order to have a more comprehensive understanding of the disease evolution and effective treatment plans. This demand has led to the gradual attention of case recommendation systems.
[0003] At present, the application of various recommendation systems in many fields has been relatively mature. There are mainly two types of systems. The first is a retrieval system based on keyword matching. The core of this recommendation system is to retrieve relevant medical record information by matching keywords with text in the medical record database. By extracting keywords from the medical record content, using keywords to index the medical record data, the user enters relevant keywords, and the system matches by calculating the similarity between keywords and terms in the medical record data (such as cosine similarity, Jaccard similarity, etc.). The system will sort the retrieved medical records and display the most relevant results to doctors or patients. The second is a recommendation system based on tags and classifications. Doctors can filter out relevant medical records by selecting tags. By establishing a medical record category model, medical records are classified according to different characteristics (such as disease type, treatment method, etc.). Users narrow the search scope by selecting tags or categories (such as "heart disease" and "headache"). According to the tags or categories selected by the doctor, the system finds matching medical records from the database.
[0004] The above methods meet the needs of medical record recommendation to a certain extent. The traditional recommendation system based on keyword matching and label classification is difficult to identify the semantic association of medical terms, synonyms or multi-level semantic associations, and has obvious limitations. In medicine, many diseases or symptoms have different expressions. For example, "hypertension" and "high blood pressure" essentially describe the same disease, but the retrieval system based on keyword matching is difficult to judge the different meanings expressed by the same keywords in different contexts simply through keyword matching. They may be regarded as different search items. The recommendation system based on labels and classification relies on artificially predefined labels. The selection and definition of these labels are very critical. If the labels do not cover a wide enough range of semantic levels or synonyms, the classification effect of the system will be limited. For multidimensional case data in the real world, there is a certain relationship network, and the traditional recommendation model may only obtain the semantic information in the data, but cannot fully capture the complex relationship between cases. Therefore, a more comprehensive method is still needed to fully explore and model the complex relationship in multidimensional case data. Summary of the invention
[0005] To this end, the technical problem to be solved by the present invention is to overcome the defects of existing recommendation systems that are difficult to identify the semantic associations of medical terms, cannot identify synonyms or multi-level semantic associations, and reduce the comprehensiveness and accuracy of recommendation results.
[0006] In order to solve the above technical problems, the present invention provides a similar case recommendation method based on knowledge graph, comprising the following steps:
[0007] Obtain and preprocess the historical case data set. Based on the preprocessed historical case data set, take various types of medical information as nodes and the attribute relationships between various types of medical information as edges to construct a historical case knowledge graph.
[0008] Obtain the case data to be queried, and determine a preset number of similar nodes related to the case data to be queried based on the case data to be queried and the historical case knowledge graph;
[0009] Each similar node is segmented, each segmentation result is mapped to a Token, and each Token is normalized to obtain the Token embedding representation of each similar node;
[0010] Construct an attention mask with the same length as the Token embedding representation of each similar node, mark each element in the Token embedding representation of each similar node through the attention mask of each similar node, and perform average pooling to obtain the target embedding representation of each similar node;
[0011] The case data to be queried is segmented, each segmentation result is mapped to a Token, and each Token is normalized to obtain the Token embedding representation of the case to be queried;
[0012] Construct an attention mask with the same length as the Token embedding representation of the case to be queried, mark the Token embedding representation of the case to be queried with the attention mask of the case to be queried, and perform average pooling to obtain the target embedding representation of the case to be queried;
[0013] By calculating the similarity between the target embedding representation of the case to be queried and the target embedding representation of each similar node, a recommendation list of similar cases is generated.
[0014] Preferably, an attention mask with the same length as the Token embedding representation of each similar node is constructed, and the Token embedding representation of each similar node is marked by the attention mask of each similar node, and average pooling is performed to obtain the target embedding representation of each similar node, including:
[0015]
[0016]
[0017] Among them, M i is the attention mask of the i-th similar node, n is the length of the Token embedding representation of the i-th similar node, is the nth element of the attention mask of the i-th similar node, It is the j-th element of the Token embedding representation of the i-th similar node. valid means valid, invalid means invalid, and padding means padding.
[0018] Preferably, based on the case data to be queried and the historical case knowledge graph, a preset number of similar nodes related to the case data to be queried is determined, including:
[0019] Each type of medical information in the case data to be queried is used as a query node, and the relevant nodes corresponding to each query node are identified in the historical case knowledge graph;
[0020] The nodes connected to the related nodes corresponding to each query node are used as the adjacent nodes corresponding to each query node;
[0021] Based on the adjacent nodes corresponding to each query node, the indirect related nodes corresponding to each query node are obtained through multi-hop reasoning;
[0022] Calculate the correlation weights of each query node and its corresponding adjacent nodes and indirectly related nodes respectively;
[0023] According to the correlation weights between each query node and its corresponding adjacent nodes and indirectly related nodes, a preset number of similar nodes related to the case data to be queried are determined.
[0024] Preferably, the medical information includes: name, disease name, symptoms, diagnosis time, and treatment plan.
[0025] Preferably, the historical case data set is preprocessed, and the preprocessing includes: data cleaning and formatting.
[0026] Preferably, each similar node is segmented or the case data to be queried is segmented, and the segmentation method is any one of Workpiece, NLTK, Spacy, and Jieba.
[0027] Preferably, each Token is normalized, and the normalization method is any one of L2 regularization, Min-Max normalization, and Z-score normalization.
[0028] Preferably, the similarity between the target embedding representation of the case to be queried and the target embedding representation of each historical case is calculated, and the similarity is any one of Euclidean distance, Manhattan distance, and cosine similarity.
[0029] Preferably, the cosine similarity between the target embedding representation of the case to be queried and the target embedding representation of each similar node is calculated using the following formula:
[0030]
[0031] Among them, Similarity(A,B i ) is the cosine similarity between the target embedding representation of the case to be queried and the target embedding representation of the i-th similar node, A is the target embedding representation of the case to be queried, B i is the target embedding representation of the i-th similar node, Similarity(.) is the cosine similarity, and ∥.∥ is the modulus.
[0032] The present invention also provides a similar case recommendation system based on knowledge graph, comprising:
[0033] The knowledge graph construction module is used to obtain and preprocess the historical case data set. Based on the preprocessed historical case data set, various types of medical information are used as nodes, and the attribute relationships between various types of medical information are used as edges to construct a historical case knowledge graph.
[0034] A similar node confirmation module is used to obtain the case data to be queried, and determine a preset number of similar nodes related to the case data to be queried based on the case data to be queried and the historical case knowledge graph;
[0035] The first Token embedding representation acquisition module is used to segment each similar node. Each segmentation result is mapped to a Token. Each Token is normalized to obtain the Token embedding representation of each similar node.
[0036] The first target embedding representation acquisition module is used to construct an attention mask with the same length as the Token embedding representation of each similar node. After marking each element in the Token embedding representation of each similar node through the attention mask of each similar node, average pooling is performed to obtain the target embedding representation of each similar node;
[0037] The second Token embedding representation acquisition module is used to segment the case data to be queried. Each segmentation result is mapped to a Token. Each Token is normalized to obtain the Token embedding representation of the case to be queried.
[0038] The second target embedding representation acquisition module is used to construct an attention mask with the same length as the Token embedding representation of the case to be queried, mark the Token sequence of the case to be queried with the attention mask of the case to be queried, and perform average pooling to obtain the target embedding representation of the case to be queried;
[0039] The generation module is used to generate a recommendation list of similar cases by calculating the similarity between the target embedding representation of the case to be queried and the target embedding representation of each similar node.
[0040] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0041] The present invention describes a similar case recommendation method and system based on a knowledge graph, which is based on a preprocessed historical case data set, takes various types of medical information as nodes, and the attribute relationships between various types of medical information as edges to construct a historical case knowledge graph; both nodes and edges can carry rich semantic information, and utilize contextual semantic associations between cases to allow a relationship network to be formed between different types of medical information, thereby enhancing the semantic depth of the data and being able to fully capture the complex relationships between cases. Based on the case data to be queried and the knowledge graph of historical cases, a preset number of similar nodes related to the case data to be queried is determined, which can effectively improve processing efficiency and reduce computing resources; the case data to be queried and each similar node are segmented, each segmentation result is mapped to a Token, and each Token is normalized to obtain the Token embedding representation of each similar node and the Token embedding representation of the case data to be queried; the deviation caused by vocabulary differences, spelling errors, etc. is reduced, and the model's ability to understand different expressions is improved. At the same time, the Token embedding representation can retain the deep semantics in the context, making the similarity calculation more accurate, and can better handle complex medical terms and subtle differences. After constructing an attention mask to mark the Token embedding representation, it can focus on meaningful information and improve the accuracy and rate of recommendation; by calculating the similarity between the target embedding representation of the case to be queried and the target embedding representation of each similar node, a similar case recommendation list is generated. The present invention can realize comprehensive capture of the complex relationship between cases and improve the efficiency and accuracy of similar case recommendation. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below according to specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:
[0043] Figure 1 It is a schematic diagram of the overall framework of a similar case recommendation method based on knowledge graph of the present invention.
[0044] Figure 2 It is a step flow chart of a similar case recommendation method based on knowledge graph of the present invention. DETAILED DESCRIPTION
[0045] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.
[0046] like Figure 1 As shown, Figure 1 Schematic diagram of the overall framework of a similar case recommendation method based on knowledge graph in the present invention.
[0047] Reference Figure 2 As shown, Figure 2 This is a flowchart of the steps of a similar case recommendation method based on knowledge graph of the present invention.
[0048] Embodiment 1 of the present invention provides a similar case recommendation method based on a knowledge graph, comprising the following steps:
[0049] Step S1: Obtain a historical case data set and preprocess it. Based on the preprocessed historical case data set, take various types of medical information as nodes and the attribute relationships between various types of medical information as edges to construct a historical case knowledge graph;
[0050] In this embodiment, preferably, the historical case data set is preprocessed, and the preprocessing includes: data cleaning and formatting. Data cleaning and formatting are crucial for constructing a historical case knowledge graph. Data cleaning can remove missing values, outliers, and duplicate data to ensure the accuracy and completeness of the data, thereby avoiding the impact of errors or deviations on the construction of the knowledge graph. Formatting helps to unify the data format, ensure that case data from different sources can be effectively integrated under the same framework, and provide a consistent and standardized data foundation for the construction of the knowledge graph. Through these preprocessing steps, data quality can be improved, noise interference can be reduced, and reliable support can be provided for the subsequent construction of the knowledge graph.
[0051] The historical case data set is preprocessed and the expression formula is:
[0052] D CSV =ExportToCSV(P θ *A θ ),
[0053] Among them, D CSV is the preprocessed historical case data set, ExportToCSV(.) is the preprocessing, P θ The patient set in the historical case data set, θ is the total number of patients in the historical case data set, A θ is the attribute set in the historical case dataset.
[0054] Based on the preprocessed historical case data set, various types of medical information are used as nodes, and the attribute relationships between various types of medical information are used as edges to construct a historical case knowledge graph. The expression formula is:
[0055] G = GraphConstruction(D CSV ),
[0056] Among them, G is the historical case knowledge graph, GraphConstruction(.) is to build the graph structure, and D CSV It is the historical case data set after preprocessing.
[0057] In this embodiment, specifically, the medical information includes: name, disease name, symptoms, diagnosis time, and treatment plan.
[0058] Based on the preprocessed historical case data set, the present invention takes various types of medical information as nodes and the attribute relationships between various types of medical information as edges to construct a historical case knowledge graph; both nodes and edges can carry rich semantic information, and utilize the contextual semantic associations between cases to allow the formation of a relationship network between different types of medical information, thereby enhancing the semantic depth of the data and being able to comprehensively capture the complex relationships between cases.
[0059] Step S2: Obtain the case data to be queried, and determine a preset number of similar nodes related to the case data to be queried based on the case data to be queried and the historical case knowledge graph;
[0060] In this embodiment, preferably, based on the case data to be queried and the historical case knowledge graph, a preset number of similar nodes related to the case data to be queried are determined, including:
[0061] Each type of medical information in the case data to be queried is used as a query node, and the relevant nodes corresponding to each query node are identified in the historical case knowledge graph;
[0062] The nodes connected to the related nodes corresponding to each query node are used as the adjacent nodes corresponding to each query node;
[0063] Based on the adjacent nodes corresponding to each query node, the indirect related nodes corresponding to each query node are obtained through multi-hop reasoning;
[0064] Calculate the correlation weights of each query node and its corresponding adjacent nodes and indirectly related nodes respectively;
[0065] According to the correlation weights between each query node and its corresponding adjacent nodes and indirectly related nodes, a preset number of similar nodes related to the case data to be queried are determined.
[0066] The present invention determines a preset number of similar nodes related to the case data to be queried based on the case data to be queried and the historical case knowledge graph, which can effectively improve processing efficiency and reduce computing resources.
[0067] Step S3: Segment each similar node, map each segmentation result into a Token, normalize each Token, and obtain the Token embedding representation of each similar node;
[0068] Each similar node is segmented, and each segmentation result is mapped to a Token, expressed as:
[0069]
[0070] Among them, T i is the token list after the segmentation of the i-th similar node, It is the nth Token after the segmentation of the i-th similar node.
[0071] Each Token is normalized to obtain the Token embedding representation of each similar node. The representation formula is:
[0072]
[0073] in, is the jth element of the Token embedding representation of the i-th similar node, It is the jth Token after the segmentation of the i-th similar node. LayerNorm(.) is the normalization process.
[0074] In this embodiment, specifically, each Token is normalized, and the normalization method is any one of L2 regularization, Min-Max normalization, and Z-score normalization.
[0075] In this embodiment, specifically, each similar node is segmented or the case data to be queried is segmented, and the segmentation method is any one of Workpiece, NLTK, Spacy, and Jieba.
[0076] Step S4: construct an attention mask with the same length as the Token embedding representation of each similar node, mark each element in the Token embedding representation of each similar node through the attention mask of each similar node, and perform average pooling to obtain the target embedding representation of each similar node;
[0077] Among them, the attention mask is expressed as:
[0078]
[0079]
[0080] Among them, M i is the attention mask of the i-th similar node, n is the length of the Token embedding representation of the i-th similar node, is the nth element of the attention mask of the i-th similar node, It is the j-th element of the Token embedding representation of the i-th similar node. valid means valid, invalid means invalid, and padding means padding.
[0081] The present invention focuses on the validity of each word or clause. The input sentence is divided into words or subwords and converted into corresponding tokens, which can handle different languages, dialects and new words, making the model more flexible. By marking valid tokens and invalid tokens, it can focus on meaningful information and avoid unnecessary interference, thereby improving the performance of the similar case recommendation system.
[0082] After marking each element in the Token embedding representation of each similar node through the attention mask of each similar node, average pooling is performed to obtain the target embedding representation of each similar node; the expression of average pooling is:
[0083]
[0084] Among them, B is the target embedding representation of the current similar node, m is the number of valid tokens of the current similar node, The kth valid token embedding represented by the Token embedding of the current similar node.
[0085] Step S5: Segment the case data to be queried, and each segmentation result is mapped to a Token. Each Token is normalized to obtain the Token embedding representation of the case to be queried;
[0086] Step S6: construct an attention mask with the same length as the Token embedding representation of the case to be queried, mark the Token embedding representation of the case to be queried with the attention mask of the case to be queried, and perform average pooling to obtain the target embedding representation of the case to be queried;
[0087] The present invention converts the case to be queried and each similar node into a corresponding target embedding representation, thereby reducing the deviation caused by vocabulary differences, spelling errors, etc., and improving the model's ability to understand different expressions. At the same time, the Token embedding representation can retain the deep semantics in the context, making the similarity calculation more accurate and better able to handle complex medical terms and subtle differences. After constructing an attention mask to mark the Token embedding representation, it can focus on meaningful information and improve the accuracy and rate of similar case recommendations.
[0088] Step S7: Generate a recommendation list of similar cases by calculating the similarity between the target embedding representation of the case to be queried and the target embedding representation of each similar node.
[0089] In this embodiment, specifically, the similarity between the target embedding representation of the case to be queried and the target embedding representation of each historical case is calculated, and the similarity is any one of Euclidean distance, Manhattan distance, and cosine similarity.
[0090] By calculating the cosine similarity between the target embedding representation of the case to be queried and the target embedding representation of each similar node, the calculation formula is:
[0091]
[0092] Among them, Similarity(A,B i ) is the cosine similarity between the target embedding representation of the case to be queried and the target embedding representation of the i-th similar node, A is the target embedding representation of the case to be queried, B i is the target embedding representation of the i-th similar node, Similarity(.) is the cosine similarity, and ∥.∥ is the modulus.
[0093] This second embodiment provides a similar case recommendation system based on knowledge graph, including:
[0094] The knowledge graph construction module is used to obtain and preprocess the historical case data set. Based on the preprocessed historical case data set, various types of medical information are used as nodes, and the attribute relationships between various types of medical information are used as edges to construct a historical case knowledge graph.
[0095] A similar node confirmation module is used to obtain the case data to be queried, and determine a preset number of similar nodes related to the case data to be queried based on the case data to be queried and the historical case knowledge graph;
[0096] The first Token embedding representation acquisition module is used to segment each similar node. Each segmentation result is mapped to a Token. Each Token is normalized to obtain the Token embedding representation of each similar node.
[0097] The first target embedding representation acquisition module is used to construct an attention mask with the same length as the Token embedding representation of each similar node. After marking each element in the Token embedding representation of each similar node through the attention mask of each similar node, average pooling is performed to obtain the target embedding representation of each similar node;
[0098] The second Token embedding representation acquisition module is used to segment the case data to be queried. Each segmentation result is mapped to a Token. Each Token is normalized to obtain the Token embedding representation of the case to be queried.
[0099] The second target embedding representation acquisition module is used to construct an attention mask with the same length as the Token embedding representation of the case to be queried, mark the Token sequence of the case to be queried with the attention mask of the case to be queried, and perform average pooling to obtain the target embedding representation of the case to be queried;
[0100] The generation module is used to generate a recommendation list of similar cases by calculating the similarity between the target embedding representation of the case to be queried and the target embedding representation of each similar node.
[0101] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0102] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0103] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0105] Obviously, the above embodiments are merely examples for the purpose of clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the present invention.
Claims
1. A similar case recommendation method based on knowledge graph, characterized in that: The following steps are involved: Obtain and preprocess the historical case data set. Based on the preprocessed historical case data set, take various types of medical information as nodes and the attribute relationships between various types of medical information as edges to construct a historical case knowledge graph. Obtain the case data to be queried, and determine a preset number of similar nodes related to the case data to be queried based on the case data to be queried and the historical case knowledge graph; Each similar node is segmented, each segmentation result is mapped to a Token, and each Token is normalized to obtain the Token embedding representation of each similar node; Construct an attention mask with the same length as the Token embedding representation of each similar node, mark each element in the Token embedding representation of each similar node through the attention mask of each similar node, and perform average pooling to obtain the target embedding representation of each similar node; The case data to be queried is segmented, each segmentation result is mapped to a Token, and each Token is normalized to obtain the Token embedding representation of the case to be queried; Construct an attention mask with the same length as the Token embedding representation of the case to be queried, mark the Token embedding representation of the case to be queried with the attention mask of the case to be queried, and perform average pooling to obtain the target embedding representation of the case to be queried; By calculating the similarity between the target embedding representation of the case to be queried and the target embedding representation of each similar node, a recommendation list of similar cases is generated.
2. According to claim 1, a similar case recommendation method based on knowledge graph is characterized in that: Construct an attention mask with the same length as the Token embedding representation of each similar node, mark the Token embedding representation of each similar node through the attention mask of each similar node, perform average pooling, and obtain the target embedding representation of each similar node, including: Among them, M i is the attention mask of the i-th similar node, n is the length of the Token embedding representation of the i-th similar node, is the nth element of the attention mask of the i-th similar node, It is the j-th element of the Token embedding representation of the i-th similar node. valid means valid, invalid means invalid, and padding means padding.
3. According to claim 1, a similar case recommendation method based on knowledge graph is characterized in that: Based on the case data to be queried and the historical case knowledge graph, a preset number of similar nodes related to the case data to be queried are determined, including: Each type of medical information in the case data to be queried is used as a query node, and the relevant nodes corresponding to each query node are identified in the historical case knowledge graph; The nodes connected to the related nodes corresponding to each query node are used as the adjacent nodes corresponding to each query node; Based on the adjacent nodes corresponding to each query node, the indirect related nodes corresponding to each query node are obtained through multi-hop reasoning; Calculate the correlation weights of each query node and its corresponding adjacent nodes and indirectly related nodes respectively; According to the correlation weights between each query node and its corresponding adjacent nodes and indirectly related nodes, a preset number of similar nodes related to the case data to be queried are determined.
4. The method for recommending similar cases based on knowledge graph according to claim 1, characterized in that: Medical information includes: name, disease name, symptoms, diagnosis time, and treatment plan.
5. The method for recommending similar cases based on knowledge graph according to claim 1, characterized in that: The historical case data set is preprocessed, and the preprocessing includes: data cleaning and formatting.
6. The similar case recommendation method based on knowledge graph according to claim 1, characterized in that: Segment each similar node or segment the case data to be queried using any of the following segmentation methods: Workpiece, NLTK, Spacy, or Jieba.
7. The similar case recommendation method based on knowledge graph according to claim 1, characterized in that: Each Token is normalized using any of the following methods: L2 regularization, Min-Max normalization, or Z-score normalization.
8. The similar case recommendation method based on knowledge graph according to claim 1, characterized in that: Calculate the similarity between the target embedding representation of the case to be queried and the target embedding representation of each historical case. The similarity is any one of the Euclidean distance, Manhattan distance, and cosine similarity.
9. A similar case recommendation method based on knowledge graph according to claim 8, characterized in that: By calculating the cosine similarity between the target embedding representation of the case to be queried and the target embedding representation of each similar node, the calculation formula is: Among them, Similarity(A,B i ) is the cosine similarity between the target embedding representation of the case to be queried and the target embedding representation of the i-th similar node, A is the target embedding representation of the case to be queried, B i is the target embedding representation of the i-th similar node, Similarity(.) is the cosine similarity, and ∥.∥ is the modulus.
10. A similar case recommendation system based on knowledge graph, characterized in that: include: The knowledge graph construction module is used to obtain and preprocess the historical case data set. Based on the preprocessed historical case data set, various types of medical information are used as nodes, and the attribute relationships between various types of medical information are used as edges to construct a historical case knowledge graph. A similar node confirmation module is used to obtain the case data to be queried, and determine a preset number of similar nodes related to the case data to be queried based on the case data to be queried and the historical case knowledge graph; The first Token embedding representation acquisition module is used to segment each similar node. Each segmentation result is mapped to a Token. Each Token is normalized to obtain the Token embedding representation of each similar node. The first target embedding representation acquisition module is used to construct an attention mask with the same length as the Token embedding representation of each similar node. After marking each element in the Token embedding representation of each similar node through the attention mask of each similar node, average pooling is performed to obtain the target embedding representation of each similar node; The second Token embedding representation acquisition module is used to segment the case data to be queried. Each segmentation result is mapped to a Token. Each Token is normalized to obtain the Token embedding representation of the case to be queried. The second target embedding representation acquisition module is used to construct an attention mask with the same length as the Token embedding representation of the case to be queried, mark the Token sequence of the case to be queried with the attention mask of the case to be queried, and perform average pooling to obtain the target embedding representation of the case to be queried; The generation module is used to generate a recommendation list of similar cases by calculating the similarity between the target embedding representation of the case to be queried and the target embedding representation of each similar node.