Knowledge graph construction method for PCOS auxiliary diagnosis

By building a knowledge graph for PCOS-assisted diagnosis, using the BERT-BiLSTM-CRF model for data processing and relationship extraction, the problem of low accuracy of PCOS diagnosis in the prior art is solved, and more accurate diagnostic and therapeutic support is achieved.

CN120012898AInactive Publication Date: 2025-05-16WEST CHINA HOSPITAL SICHUAN UNIV +1

Patent Information

Application Number
CN202410121010.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The lack of methods for constructing knowledge maps for polycystic ovary syndrome (PCOS) in the prior art has led to lower diagnostic accuracy of PCOS, especially in adolescent patients.

Method used

PCOS-related data is collected through crawler tools, preprocessed and structured, and named entity recognition and relationship extraction are used to construct a knowledge graph for PCOS-assisted diagnosis.

Benefits of technology

It improves the accuracy of PCOS diagnosis, can more accurately identify different types of PCOS, assists in artificial diagnosis, promotes precise treatment, and improves patient prognosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012898A_ABST
    Figure CN120012898A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence diagnosis, and particularly relates to a knowledge graph construction method for PCOS auxiliary diagnosis. The method for constructing the knowledge graph for PCOS auxiliary diagnosis comprises the following steps: step 1, collecting PCOS related data by using a crawler tool to obtain original data; 2, preprocessing the original data to obtain structured data; 3, named entity recognition and relation extraction are carried out, and the relation between the medical entities is obtained; and step 4, constructing a knowledge graph according to the relationship between the medical entities. By using the knowledge graph and combining a line graph neural network, the relationship between the detection index and different types of PCOS can be further obtained. Therefore, manual diagnosis of PCOS is assisted, diagnosis accuracy is improved, follow-up precise treatment is facilitated, patient prognosis is improved, and good application prospects are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence diagnosis, and specifically relates to a knowledge graph construction method for PCOS auxiliary diagnosis. Background Art

[0002] Polycystic Ovary Syndrome (PCOS) is one of the most common endocrine diseases in women of childbearing age. Its main clinical features include rare or no ovulation, clinical or biochemical hyperandrogenism, and polycystic ovarian changes. It often occurs after menarche in adolescence. If it is not diagnosed and treated promptly and accurately, it can increase the patient's long-term complications, such as reproductive dysfunction, endometrial cancer, type 2 diabetes, etc., seriously affecting the quality of life, endangering the patient's life and health, and even shortening his life. Therefore, timely and accurate diagnosis of PCOS is conducive to subsequent precise treatment to improve the patient's prognosis.

[0003] At present, accurate diagnosis of PCOS, especially the diagnosis of PCOS in adolescence, is challenging. Different androgen phenotypes of PCOS patients with hyperandrogenemia need to be distinguished in the diagnosis, and clinical characteristics should be combined. In addition, the clinical manifestations of PCOS are diverse, and the diagnostic criteria and differential diagnosis are relatively complex. When performing manual diagnosis, the accuracy of PCOS diagnosis may be affected by differences in personal experience levels. Therefore, there is an urgent need to develop new diagnostic technologies in this field to improve the accuracy of PCOS diagnosis.

[0004] Knowledge graph is a semantic network that reveals the relationship between entities. It uses entities (concepts) as nodes and relationships as edges. It can be represented by a graph structure and used to store structured semantic knowledge bases to realize concept retrieval based on reasoning. The core function of knowledge graph is to convert the scattered and structured information on the Internet into a regular form that is closer to the human brain after unified processing, so as to store and manage the existing and constantly increasing knowledge and information on the Internet. Knowledge graph is now being integrated into the development trend of multiple fields, and continues to shine in areas such as search, intelligent question-answering systems, and recommendation systems. Technologies such as knowledge graph and deep learning are constantly driving computer technology to develop in a more intelligent direction, and are expected to become a basic facility for the Internet to provide knowledge storage and reasoning expression in the future. With the increase in the amount of medical data, it is of great significance to discover new knowledge from medical entities such as diseases, drugs, treatments, and genes, and to mine the implicit knowledge between medical data to assist in disease diagnosis. Knowledge graph technology has become an important technical support for knowledge question-answering and domain knowledge discovery. Combining medical domain knowledge to build a medical knowledge graph is the driving force for the development of intelligent medicine in the future.

[0005] It can be seen that if the auxiliary means of knowledge graph can be introduced in the diagnosis of PCOS, it will have a positive significance for improving the accuracy of diagnosis and achieving precise treatment. However, for literature in different fields, due to differences in characteristics such as the degree of language structuring, the methods of constructing knowledge graphs are also quite different. Therefore, for specific fields, specific methods need to be used to construct knowledge graphs. At present, there is a lack of relevant research in the prior art, and no knowledge graph construction method for PCOS is provided. Summary of the invention

[0006] In view of the problems in the prior art, the object of the present invention is to provide a knowledge graph construction method for auxiliary diagnosis of PCOS.

[0007] A knowledge graph construction method for PCOS auxiliary diagnosis includes the following steps:

[0008] Step 1, using crawler tools to collect PCOS related information and obtain raw data;

[0009] Step 2: preprocess the original data to obtain structured data;

[0010] Step 3: Use the BERT-BiLSTM-CRF model to perform named entity recognition on the structured data obtained in step 2, and use the extraction model that integrates BERT-BiLSTM-CRF and multi-head selection to perform relationship extraction to obtain the relationship between medical entities;

[0011] Step 4: construct a knowledge graph based on the medical entities and the relationships between entities.

[0012] Preferably, in step 1, the crawler tool is a distributed crawler framework Scrapy;

[0013] And / or, in step 1, the PCOS-related data include medical literature, medical dictionaries, electronic medical records, medical guidelines and expert consensus; the specific contents of the PCOS-related data include clinical diagnosis, clinical symptoms and androgen test results.

[0014] Preferably, in step 2, referring to the Unified Medical Language System and the International Classification of Diseases, unique concept identifiers are used to encode concepts from different vocabulary sources but with the same vocabulary, and RDF triples are used to represent the semantics of the identifiers and the associations between different identifiers to obtain structured data.

[0015] Preferably, in step 4, the knowledge graph is constructed using the graph-based database Neo4j.

[0016] Preferably, the method further includes a step of updating the knowledge graph, including:

[0017] Step 5, regularly use crawler tools to collect PCOS-related information to obtain updated raw data;

[0018] Step 6, preprocessing the updated original data to obtain structured data;

[0019] Step 7: Use the BERT-BiLSTM-CRF model to perform named entity recognition on the data obtained in step 6, and use the extraction model integrating BERT-BiLSTM-CRF and multi-head selection to perform relationship extraction to obtain new medical entities and the relationships between entities.

[0020] Step 8: merge the new medical entities and the relationships between entities obtained in step 7 into the previous knowledge graph to obtain an updated knowledge graph.

[0021] The present invention also provides a method for obtaining PCOS auxiliary diagnosis information, comprising the following steps:

[0022] Step A, constructing a knowledge graph for auxiliary diagnosis of PCOS according to the knowledge graph construction method according to any one of claims 1 to 5;

[0023] Step B, using the data in the knowledge graph as a training set to train a line graph neural network, and obtaining the relationship between the detection indicators and different types of PCOS through the line graph neural network.

[0024] Preferably, in step B, two-thirds of the knowledge graph is used to construct a training set;

[0025] And / or, in step B, the line graph neural network gradually extracts high-level features of the nodes in the knowledge graph by continuously updating and aggregating the nodes in the knowledge graph, and uses them to infer the relationship between the detection indicators and different types of PCOS.

[0026] Preferably, it also includes:

[0027] Step C, input the detection indicators of the real case into the knowledge graph for query, and obtain the PCOS type predicted by the computer;

[0028] Step D, manually comparing and correcting the PCOS types of real cases and the PCOS types predicted by computer, adjusting the relevant parameters of the line graph neural network, updating the data mining algorithm, and retraining to obtain an updated database, and using this database to update the knowledge graph.

[0029] The present invention also provides a PCOS auxiliary diagnosis information acquisition system, comprising:

[0030] A knowledge graph storage module, used to store the knowledge graph for PCOS auxiliary diagnosis constructed according to the above-mentioned knowledge graph construction method;

[0031] A line graph neural network training module, used to use the data in the knowledge graph as a training set to train the line graph neural network, and obtain the relationship between the detection index and different types of PCOS through the line graph neural network;

[0032] The knowledge graph updating module is used to update the knowledge graph according to the relationship between the detection indicators and different types of PCOS.

[0033] The present invention also provides a computer-readable storage medium, on which is stored: a computer program for implementing the above-mentioned method for constructing a recognition spectrum, or a method for obtaining PCOS auxiliary diagnosis information.

[0034] In the process of constructing a knowledge graph for auxiliary diagnosis of PCOS, the present invention introduces a character-based BERT pre-trained language model on the basis of the BiLSMT-CRF model, successfully solving the problem that the word vector in the current mainstream named entity recognition method BiLSMT-CRF model does not consider the context and cannot construct a word multi-vector. At the same time, the model can better combine the contextual semantics of the text and mine and discover the characteristic expressions in PCOS-related medical literature.

[0035] In the process of knowledge graph construction, the encoding of the BiLSTM layer is connected in series with the predicted label as the input of the relationship extraction task; in the relationship extraction part, the sigmoid layer is used to make relationship predictions to determine whether two entities have a relationship. At the same time, this method solves the problem of entity overlap in sentences.

[0036] Furthermore, the present invention adopts a line graph neural network (LGNN), which can utilize the relationship between nodes and context information to better model and predict graph structure data, thereby accurately predicting the relationship between detection indicators and different types of PCOS.

[0037] The present invention makes full use of the feature of knowledge graph that can clearly display the relationship between entities, develops the application of knowledge graph in the diagnosis process of PCOS, assists in the manual diagnosis of PCOS, improves the accuracy of diagnosis, facilitates subsequent precise treatment, and improves patient prognosis.

[0038] Obviously, according to the above contents of the present invention, in accordance with common technical knowledge and customary means in the art, without departing from the above basic technical ideas of the present invention, other various forms of modification, replacement or change may be made.

[0039] The above contents of the present invention are further described in detail below through specific implementation methods in the form of embodiments. However, this should not be understood as the scope of the above subject matter of the present invention being limited to the following examples. All technologies realized based on the above contents of the present invention belong to the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 The PCOS entity recognition model based on BERT-BiLSTM-CRF in Example 1;

[0041] Figure 2 The BERT pre-trained language model in Example 1;

[0042] Figure 3 The LSTM encoding unit in Example 1;

[0043] Figure 4 This is the relation extraction model that integrates BERT-BiLSTM-CRF and multi-head selection in Example 1. DETAILED DESCRIPTION

[0044] It should be noted that the algorithms of data collection, transmission, storage and processing steps not specifically described in the embodiments, as well as the hardware structure, circuit connection, etc. not specifically described can all be implemented through the contents disclosed in the prior art.

[0045] Example 1: Knowledge graph construction method for PCOS auxiliary diagnosis

[0046] This embodiment provides a method for constructing and updating a knowledge graph for PCOS auxiliary diagnosis. In order to construct a knowledge graph for PCOS auxiliary diagnosis, based on the data of PCOS clinical diagnosis, clinical symptoms, androgen detection in medical literature, medical dictionaries, electronic medical records, various medical guidelines, and expert consensus, combined with natural language processing technology, the knowledge graph is constructed through four parts: medical entity representation, named entity recognition, relationship extraction, and knowledge graph construction. Specifically, the following steps are included:

[0047] Step 1: Use the distributed crawler framework Scrapy to crawl and collect non- / semi-structured data such as various literature repositories and databases.

[0048] Step 2: Data cleaning, data labeling, and data integration;

[0049] Step 3: Refer to the Unified Medical Language System (UMLS), International Classification of Diseases (ICD-10), etc., use unique concept identifiers to encode concepts from different vocabulary sources but with the same vocabulary, and use RDF triples to represent the semantics of the identifiers and the association between different identifiers to obtain structured data;

[0050] In addition, RDF can also be represented by a graph model consisting of nodes and relationships, where nodes represent entities and attribute values, and lines represent the relationships between nodes.

[0051] Step 4: Build a BERT-BiLSTM-CRF model for named entity recognition. The overall structure of the BERT-BiLSTM-CRF model is as follows: Figure 1 As shown:

[0052] The model is mainly divided into three layers. The first layer is the BERT layer. Through the BERT pre-trained language model, each word in the sentence is converted into a low-dimensional word vector. The second layer is the BiLSTM layer. The BiLSTM neural network is used to automatically extract sentence features. The third layer is the CRF layer. The CRF algorithm is used to sequence the sentences. The models used in each layer are introduced in detail below:

[0053] BERT pre-trained model: The structure of the BERT model is as follows Figure 2 As shown in the figure, the model can simultaneously obtain information in both the front and back directions of a sentence through the bidirectional Transformer encoder. The main feature of this model is the use of the Transformer structure and the unsupervised learning pre-training method.

[0054] BiSLTM pre-trained language model: The BiLSTM layer models sentences based on the BERT model. Using LSTM to model sentences can only obtain one-way information, while BiLSTM can better capture the bidirectional semantic dependencies from front to back and back to front, thereby effectively combining contextual information. Its unit structure is as follows Figure 3 shown.

[0055] CRF model: The BiLSTM layer fully considers the context information, but does not consider the dependency information between labels. For example, the first word in a sentence should start with the label "B" or "O" instead of "I-"; in a segment of entity labels "B-label1, I-label2, I-label3...", label1, label2, label3 should be entities of the same type; "B-PersonI-Person" is a legal sequence, but "B-Person I-Organization" is an illegal label sequence.

[0056] Lafferty first proposed the Conditional Random Field (CRF) model in 2001. In this named entity recognition task, the CRF layer can add some constraints to the predicted labels to ensure the legitimacy of the predicted labels. In the process of training data, these constraints can be automatically learned by the CRF layer.

[0057] The parameter of the CRF layer is a (k+2)×(k+2) matrix A. The score of the token sequence y of sentence x can be calculated by the following formula:

[0058]

[0059] score(x,y) represents the score of the tag sequence y for sentence x, where A is the transformation matrix, yi represents the i-th tag of the tag sequence y, and P i,yi represents the score of the yi-th label of the character, and n represents the length of sentence x.

[0060] Use Softmax to get the normalized probability:

[0061]

[0062] The probability of the label sequence y is calculated by the above formula. The numerator represents the natural exponent of the score of the label sequence y, and the denominator represents the sum of the natural exponent of the scores of all label sequences y'. The set with the maximum probability is selected as the final label sequence of the CRF layer.

[0063] Step 5: Input the structured data into the BERT-BiLSTM-CRF model, extract the effective information from the encoded identifiers and RDF triples, and convert them into a series of feature vectors.

[0064] The relationship extraction model is constructed by integrating BERT-BiLSTM-CRF and multi-head selection. The specific structure of the model is as follows Figure 4 shown.

[0065] Step 6: Through the two subtasks of named entity recognition in step 4 and relationship extraction in step 5, the relationship between medical entities is obtained, which contains a lot of medical information. Then a knowledge graph is constructed based on the graph database Neo4j.

[0066] The drawing of knowledge graph includes the following three steps:

[0067] Organize the data and save the extracted knowledge into files according to different entity types and relationship types;

[0068] Data import: convert the organized data files into csv format, and import the nodes and relationships into the graph library;

[0069] Graph viewing: View the entity nodes and relationships between entities in the graph through the query language Cypher.

[0070] As a preferred embodiment, after constructing the knowledge graph for PCOS auxiliary diagnosis according to the above method, you can also try to update the data of the knowledge graph. Specifically, the steps include:

[0071] Continuously crawl data from medical databases and a large amount of literature to continuously update the data set.

[0072] According to the method of step 3 to step 5 above, the updated structured data is obtained.

[0073] The graph-based database Neo4j combines the newly obtained structured data with the original knowledge graph data to achieve the effect of continuously updating the knowledge graph.

[0074] Example 2 PCOS auxiliary diagnosis information acquisition system and method

[0075] This embodiment provides a system and method for obtaining PCOS auxiliary diagnosis information, the system comprising:

[0076] A knowledge graph storage module, used to store the knowledge graph for PCOS auxiliary diagnosis constructed according to the knowledge graph construction method of Example 1;

[0077] A line graph neural network (LGNN) training module, used to use the data in the knowledge graph as a training set to train the line graph neural network, and obtain the relationship between the detection index and different types of PCOS through the line graph neural network;

[0078] The knowledge graph updating module is used to update the knowledge graph according to the relationship between the detection index and different types of PCOS. During actual diagnosis, the detection index can be input into the knowledge graph to obtain the diagnosis result.

[0079] The method of obtaining PCOS auxiliary diagnosis information using the above system includes:

[0080] The line graph neural network is trained using 2 / 3 of the knowledge graph as the training set. The line graph neural network gradually extracts the high-level features of the nodes in the graph by continuously updating and aggregating the nodes in the graph, and is used to infer the relationship between multiple detection indicators and related diseases.

[0081] Finally, we obtained the complete relationship between multiple detection indicators and different types of PCOS.

[0082] It is worth noting that the LGNN used in this embodiment is a deep learning model for processing graph structure data. Compared with the traditional graph convolutional network (Graph Convolutional Network), the line graph neural network pays more attention to the modeling of the edges in the graph (i.e., the relationship between the connected nodes). LGNN updates the representation of the node by performing information transfer and aggregation operations on each node. This information transfer and aggregation process is usually carried out in an iterative manner. In each round of iteration, the network updates the representation of the node and improves the quality of the node representation by aggregating the information of neighboring nodes. In each round of iteration, LGNN combines the features of the node with the features of its neighboring nodes and applies nonlinear transformations to generate new node representations. In this way, the representation of the node will gradually integrate the information of the surrounding nodes to form a richer representation.

[0083] The training process of LGNN usually includes two stages: forward propagation and back propagation. In forward propagation, the network calculates the representation of the node based on the current parameters and the node's neighbor information. Then, in back propagation, the network updates the parameters according to the gradient of the loss function so that the node representation can better predict the label or output of the target task. The training process usually includes multiple rounds of iterations. In each round of iteration, the network updates the representation of the node and improves the quality of the node representation by aggregating the information of neighboring nodes. Through repeated iterations, the line graph neural network can gradually extract the high-level features of the nodes in the graph and use them to predict the output.

[0084] This embodiment uses LGNN to learn the representation of nodes and edges in the graph. Specifically, the nodes in the graph are represented as vectors, and these vectors are updated and aggregated through the line graph neural network, so that the relationship between the nodes can be used to infer the relationship between multiple detection indicators and related diseases.

[0085] By introducing LGNN, this embodiment can more accurately predict the relationship between multiple detection indicators and different types of PCOS, and provide more accurate support for the diagnosis and treatment of related diseases.

[0086] As a preferred embodiment, this embodiment can further use real cases to verify the accuracy of LGNN prediction results and modify LGNN.

[0087] The specific steps include:

[0088] The detection indicators of real cases are input into the knowledge graph for query, and the PCOS types predicted by the computer are obtained;

[0089] The PCOS types of real cases and the PCOS types predicted by computer were manually compared and corrected, the relevant parameters of the line graph neural network were adjusted, the data mining algorithm was updated and retrained to obtain an updated database, which was used to update the knowledge graph built based on the Neo4j graph database.

[0090] Through the above embodiments, it can be seen that the present invention constructs a method that can automatically capture PCOS-related literature and public information and construct them into a knowledge graph. The obtained knowledge graph can be used to obtain the relationship between various detection indicators and PCOS types. Therefore, the present invention can assist in the manual diagnosis of PCOS, improve the accuracy of diagnosis, facilitate subsequent precise treatment, improve patient prognosis, and has a good application prospect.

Claims

1. A knowledge graph construction method for PCOS auxiliary diagnosis, characterized in that: The steps include: Step 1, using crawler tools to collect PCOS related information and obtain raw data; Step 2: preprocess the original data to obtain structured data; Step 3: Use the BERT-BiLSTM-CRF model to perform named entity recognition on the structured data obtained in step 2, and use the extraction model that integrates BERT-BiLSTM-CRF and multi-head selection to perform relationship extraction to obtain the relationship between medical entities; Step 4: construct a knowledge graph based on the medical entities and the relationships between entities.

2. The knowledge graph construction method according to claim 1, characterized in that: In step 1, the crawler tool is the distributed crawler framework Scrapy; And / or, in step 1, the PCOS-related data include medical literature, medical dictionaries, electronic medical records, medical guidelines and expert consensus; the specific contents of the PCOS-related data include clinical diagnosis, clinical symptoms and androgen test results.

3. The knowledge graph construction method according to claim 1, characterized in that: In step 2, referring to the Unified Medical Language System and the International Classification of Diseases, unique concept identifiers are used to encode concepts with the same vocabulary from different vocabulary sources, and RDF triples are used to represent the semantics of the identifiers and the associations between different identifiers to obtain structured data.

4. The knowledge graph construction method according to claim 1, characterized in that: In step 4, the knowledge graph is constructed using the graph-based database Neo4j.

5. The knowledge graph construction method according to claim 1, characterized in that: The method also includes the step of updating the knowledge graph, including: Step 5, regularly use crawler tools to collect PCOS-related information to obtain updated raw data; Step 6, preprocessing the updated original data to obtain structured data; Step 7: Use the BERT-BiLSTM-CRF model to perform named entity recognition on the data obtained in step 6, and use the extraction model integrating BERT-BiLSTM-CRF and multi-head selection to perform relationship extraction to obtain new medical entities and the relationships between entities. Step 8: merge the new medical entities and the relationships between entities obtained in step 7 into the previous knowledge graph to obtain an updated knowledge graph.

6. A method for obtaining PCOS auxiliary diagnosis information, characterized in that: The method comprises the following steps: Step A, constructing a knowledge graph for auxiliary diagnosis of PCOS according to the knowledge graph construction method according to any one of claims 1 to 5; Step B, using the data in the knowledge graph as a training set to train a line graph neural network, and obtaining the relationship between the detection indicators and different types of PCOS through the line graph neural network.

7. The PCOS auxiliary diagnosis information acquisition method according to claim 6, characterized in that: In step B, two-thirds of the knowledge graph is used to construct the training set; And / or, in step B, the line graph neural network gradually extracts high-level features of the nodes in the knowledge graph by continuously updating and aggregating the nodes in the knowledge graph, and uses them to infer the relationship between the detection indicators and different types of PCOS.

8. The PCOS auxiliary diagnosis information acquisition method according to claim 6, characterized in that: Also includes: Step C, input the detection indicators of the real case into the knowledge graph for query, and obtain the PCOS type predicted by the computer; Step D, manually comparing and correcting the PCOS types of real cases and the PCOS types predicted by computer, adjusting the relevant parameters of the line graph neural network, updating the data mining algorithm, and retraining to obtain an updated database, and using this database to update the knowledge graph.

9. A PCOS auxiliary diagnosis information acquisition system, characterized in that: include: A knowledge graph storage module, used to store a knowledge graph for PCOS auxiliary diagnosis constructed according to the knowledge graph construction method according to any one of claims 1 to 5; A line graph neural network training module, used to use the data in the knowledge graph as a training set to train the line graph neural network, and obtain the relationship between the detection index and different types of PCOS through the line graph neural network; The knowledge graph updating module is used to update the knowledge graph according to the relationship between the detection indicators and different types of PCOS.

10. A computer-readable storage medium, characterized in that: Stored thereon is: a computer program for implementing the method for constructing a recognition spectrum as described in any one of claims 1 to 5, or the method for obtaining PCOS auxiliary diagnosis information as described in any one of claims 6 to 8.

Citation Information

Patent Citations

  • Knowledge graph construction method for auxiliary diagnosis

    CN115269865A

  • Disease prediction system and device based on clinical data screening and medical knowledge graph

    CN115862848A

  • Resource allocation method and system based on sparse knowledge graph link prediction

    CN116302481A

  • Disease prediction system for unbalanced data

    CN116936108A

Cited By

  • Diagnosis and treatment sample construction method and device and auxiliary diagnosis and treatment large model training method and device

    CN120930758A