Artificial Intelligence-Based Popular Science Method for Medical Data

By constructing a disease tree database and optimal allocation function based on ICD, medical literature is stored in the disease tree database, which solves the problem that users find it difficult to obtain disease-related literature, and realizes the function of quickly obtaining relevant literature after users ask questions, expanding users' medical knowledge.

CN119884328BActive Publication Date: 2025-06-20SICHUAN CANCER HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510362858.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-20
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

When users query disease-related information on the Internet, it is difficult for users to find a large number of reliable documents, resulting in limited understanding of the disease, and it is difficult for the existing technology to quickly recommend relevant documents through user questions.

Method used

Using a popular medical data science method based on artificial intelligence, medical-related literature is obtained through multiple data sources, a disease tree database based on International Disease Classification (ICD) is constructed, and the literature is stored in the closest node through the optimal allocation function, so as to realize the function of quickly retrieving relevant literature after user questions.

Benefits of technology

It has enabled users to query relevant literature from a wide range of disease types or from a detailed range of details, expand users' understanding of medical knowledge and improve users' efficiency in obtaining relevant literature.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119884328B_ABST
    Figure CN119884328B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data analysis and processing, and discloses a medical data popularization method based on artificial intelligence, including the steps of: obtaining medical-related literature through various data sources; constructing a disease tree database with reference to the disease type classification method of the International Classification of Diseases (ICD), where the disease tree database contains multiple nodes; classifying and organizing the medical-related literature by disease type and storing it in the corresponding nodes of the disease tree database; receiving a query statement input by the user, extracting the disease type from the query statement, and retrieving the medical-related literature from the nodes corresponding to the disease type and returning it to the user. The present invention constructs a multi-level disease tree database with reference to the disease classification method of ICD. When a user queries the relevant literature of a certain disease, the user can query the relevant literature of the disease type from broad to detailed or from detailed to broad, so as to popularize more medical knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis and processing, and particularly to a medical data popularization method based on artificial intelligence. Background Art

[0002] After a user or patient contracts a certain disease and wants to query relevant information about the disease through the Internet, the most authentic ones are published papers, periodicals and other literatures. However, the user cannot search for many literatures, and even the search for public literatures is limited. Therefore, the user's understanding of disease-related information is limited. How to quickly recommend relevant literatures of the disease mentioned by the user to the user by means of user questions to achieve the purpose of disease popularization is a problem that needs to be solved urgently at present. Summary of the Invention

[0003] The purpose of the present invention is to improve the deficiencies in the prior art and provide a medical data popularization method based on artificial intelligence.

[0004] To achieve the above-mentioned invention purpose, the embodiments of the present invention provide the following technical solutions:

[0005] A medical data popularization method based on artificial intelligence, comprising the following steps:

[0006] Step 1, obtaining medical-related literatures through multiple data sources;

[0007] Step 2, constructing a disease tree database with reference to the disease type classification method of the International Classification of Diseases (ICD), and the disease tree database contains multiple nodes;

[0008] Step 3, classifying and sorting the medical-related literatures by disease type and storing them in the corresponding nodes of the disease tree database;

[0009] Step 4, receiving a question statement input by the user, extracting the disease type from the question statement, and retrieving medical-related literatures from the nodes corresponding to the disease type and returning them to the user.

[0010] Compared with the prior art, the beneficial effects of the present invention:

[0011] The present invention constructs a multi-level disease tree database with reference to the disease classification method of ICD. When a user queries relevant literatures of a certain disease, the user can query the relevant literatures of the disease type from broad to fine or from fine to broad, so as to popularize more medical knowledge.

[0012] The present invention stores the literatures in the closest nodes based on the extracted literature word segmentation through the created optimal allocation function, so that when the user queries the disease type, the literatures in the corresponding nodes can be recommended to the user. Brief Description of the Drawings

[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0014] Figure 1 It is a flowchart of the method of the present invention;

[0015] Figure 2 It is the classification method of "05 Endocrine, nutritional or metabolic diseases" in the International Classification of Diseases ICD. Specific embodiments

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Usually, the components of the embodiments of the present invention described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but only represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0017] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance, or implying any such actual relationship or order between these entities or operations. In addition, terms such as "connected" and "coupled" can be directly connected between elements or indirectly connected via other elements.

[0018] The present invention is implemented through the following technical solutions, as Figure 1 shown, a medical data popularization method based on artificial intelligence, including the following steps:

[0019] Step 1, obtaining medical-related literature through multiple data sources.

[0020] Access the internal database of this hospital to obtain all medical-related literature. The medical data of this hospital are all published by the doctors of this hospital and are more closely related to the diagnosis and treatment processes of the patients in this hospital. Use a Web crawler tool to crawl a number of medical-related literature from publicly available network resources on the Internet, including various publicly available network resources such as academic databases, journal websites, institutional libraries, professional association institutions, and research institution reports. When crawling publicly available network resources, medical literature with a large number of citations and reproductions is given priority. This embodiment does not limit the sources of obtaining medical data literature, and as long as the data sources are reasonable, they can be obtained, so as to form a large-scale medical-related literature, which is used for subsequent sorting and classification to provide targeted medical popularization for patients.

[0021] Step 2, construct a disease tree database with reference to the disease type classification method of the International Classification of Diseases (ICD). The disease tree database contains multiple nodes.

[0022] Construct a disease tree database with reference to the International Classification of Diseases (ICD) formulated and published by the World Health Organization (WHO). ICD is classified into 26 chapters according to causes, locations, etc., and each chapter is further classified in detail until there are final codes. For example, Chapter 5 is "05 Endocrine, nutritional or metabolic diseases"; it is further divided into "Endocrine diseases, nutritional disorders, metabolic disorders, post-operative endocrine or metabolic disorders"; among them, "Endocrine diseases" are further divided into "Disorders of the thyroid or thyroid hormone system, diabetes, other glucose regulation or pancreatic endocrine disorders, disorders of the parathyroid or parathyroid hormone system, disorders of the pituitary hormone system, disorders of the adrenal or adrenal hormone system, etc."; among them, "Diabetes" is further divided into "5A10 Type 1 diabetes, 5A11 Type 2 diabetes, 5A12 Malnutrition-related diabetes, 5A13 Diabetes (other specified types), 5A14 Diabetes (unspecified type), Acute complications of diabetes". It can be seen that, as Figure 2 shown, "5A11" etc. are the final codes. For specific details, please refer to the official website https: / / icd11.pumch.cn, and the ICD classification method will not be elaborated in this solution.

[0023] This solution constructs a disease tree database by referring to the classification method of ICD, selects the classification methods of the first 24 chapters, uses the chapters "01 Certain infectious diseases or parasites, 02 Tumors, 03 Diseases of the blood or hematopoietic organs,..., 23 External causes of diseases or death, 24 Factors affecting health status or contact with health care institutions" as the root nodes, and then extends the descendant nodes downward, but only extends to the 4-digit code position. For example, it can extend to the level of the leaf node "5A10 Type 1 diabetes", which just has the 4-digit code "5A10".

[0024] It should be noted that due to the different classification results of each chapter, the degree, branching degree, number of descendant nodes, height, etc. of each tree are different. This solution constructs 24 disease tree databases completely according to the classification method of ICD, that is, there are 24 root nodes in total, and then further extends its descendant nodes according to the classification of each chapter. It is easy to understand that this tree construction method can be realized based on the existing technology, and this solution does not improve it. The node extension is based on the ICD classification method.

[0025] Step 3: Classify and organize the disease types in medical-related literature and store them in the corresponding nodes of the disease tree database.

[0026] In this step, the disease type of the literature is extracted through the name and keywords of the medical-related literature. Usually, a literature will record keywords, or the main idea has been pointed out in the literature name. For example, for a literature named "Efficacy and Safety of Shuxuening Injection in the Adjuvant Treatment of Diabetic Patients", the disease type of "diabetes" can be quickly extracted. Therefore, this literature can be stored in the disease tree database corresponding to 05 Endocrine, Nutrition or Metabolic Diseases. Specifically, it is stored in the descendant node "diabetes". Since the node "diabetes" is not a leaf node and has sub-nodes such as "5A10 Type 1 diabetes, 5A11 Type 2 diabetes, etc.", which are the leaf nodes of this tree, the content of this literature is further queried. If there is a disease type corresponding to the leaf node, the literature is stored in that leaf node. For example, if the literature "Efficacy and Safety of Shuxuening Injection in the Adjuvant Treatment of Diabetic Patients" contains the content of "Type 1 diabetes", the literature is stored in the leaf node 5A10; if there is also the content of "Type 2 diabetes", the literature is also stored in the leaf node 5A11; if no content represented by any leaf node can be queried in this literature, the literature is stored in the node corresponding to "diabetes". It is easy to understand that each parent node contains the addresses of all its descendant nodes. Therefore, the literatures stored in all its descendant nodes can be queried in the parent node.

[0027] This solution extracts the word segmentation of the literature and determines which descendant node a literature is stored in through the created optimal allocation function. The created optimal allocation function is improved based on the Support Vector Machine (SVM). Suppose the word segmentation result is obtained after word segmentation of the name and keywords of a literature, and after word vector encoding of the word segmentation result, a word vector encoding set X = {x1, x2,..., x n} is obtained, where x i represents the encoding value of the i-th word segmentation, i = 1, 2,..., n, and n represents the number of word segmentations of the name and keywords. The created optimal allocation function is:

[0028]

[0029] Among them, f(x i ) represents the optimal allocation function; x i represents the encoding value of the i-th word segment; y j represents the encoding value of the j-th node, where j = 1, 2,..., K and K represents the number of nodes; sgn represents the step function; represents the kernel function parameter; represents the Gaussian radial basis parameter, ; represents the bias term; h i represents the weight coefficient of the i-th word segment, which is related to the number of times the i-th word segment appears in the literature. It is easy to understand that the encoding values x i and y j are obtained through the word2vec word vector encoding technology.

[0030] When this solution assigns a literature to the nodes of the disease tree database, the following situations exist:

[0031] (1) If a literature has been stored in a non-leaf node according to the word segments of its name and keywords, then further query whether there are sub-nodes or leaf nodes with matching word segments in the content of this literature. If there is only a matching sub-node, then store this literature in the matching sub-node. If there is a matching leaf node, then store this literature in the matching leaf node.

[0032] (2) If there are no matching nodes for the word segments of the name and keywords of a literature, then further query whether there are nodes or leaf nodes with matching word segments in the content of this literature. If there is only a matching non-leaf node, then store this literature in the matching non-leaf node. If there is a matching leaf node, then store this literature in the matching leaf node.

[0033] (3) If a literature has been stored in a certain leaf node according to the word segments of its name and keywords, then no further matching is performed based on the content of this literature.

[0034] (4) If there are no matching nodes for the word segments of the name, keywords, and content of a literature, this literature is not stored.

[0035] It can be seen that it is necessary to store the literature in sub-nodes with a higher degree as much as possible, and it is best to store it in leaf nodes. This solution uses the optimal allocation function twice. The first time is to query the matching situation between the name and keywords of the literature and all nodes, and the second time is to query the matching situation between the content of the literature and all nodes. There are two differences in using the optimal allocation function:

[0036] 1. The first vector encoding set X = {x1, x2,..., x n} represents the encoded values of the name of the reference and the word segmentation of the keywords, and the second vector encoding set X = {x1, x2,..., x n} represents the encoded values of the word segmentation of the content of the reference.

[0037] 2. The first K value represents the total number of all nodes in 24 disease tree databases, while the second K value only represents the number of descendant nodes downward from the nodes stored in the first time.

[0038] This solution uses the optimal allocation function twice. The first time, it allocates according to the name and keywords of the reference. Since the name and keywords of the reference are the main ideas of the reference, it can find out which disease tree database the reference belongs to. If the node stored in the first time is a non-leaf node, then a second match is made, and only the descendant nodes of the node stored in the first time are matched during the second match, which has greatly reduced the computational amount. If the reference has been stored in a leaf node in the first time, then there is no need to perform a second match, and there is no need to analyze the content of the reference. The length of the content is much larger than that of the name and keywords, so the computational amount is further reduced.

[0039] Step 4, receive the question sentence input by the user, extract the disease type from the question sentence, and retrieve the medical-related references from the nodes corresponding to the disease type and return them to the user.

[0040] For example, if the user inputs the question sentence "How to treat hypothyroidism", it is easy to extract the disease type of "hypothyroidism", and then return the medical-related references in the leaf node "5A00 Hypothyroidism" from the disease tree database of 05 Endocrine, nutritional or metabolic diseases to the user.

[0041] When feedbacking the references, sort the references in the leaf node "5A00 Hypothyroidism" according to the recommendation value, and the references with higher recommendation values are presented to the user first. The recommendation value of a reference is calculated by the following formula:

[0042]

[0043] Among them, SID represents the recommendation value; QSC represents the sharing and citation index of the reference; ASC represents the interactive comment index of the reference; E samp represents the page scrolling index within the unit time T. If the page scrolling speed is faster within the unit time T, it means that the user dislikes the reference more; w1 is the weight of QSC, w2 is the weight of ASC, and w3 is the weight.

[0044] In addition, the user can be prompted to view the documents stored in the ancestor nodes of the leaf node "5A00 Hypothyroidism", so that the user can further learn about the medical-related documents on "Thyroid or Thyroid Hormone System Diseases", "Endocrine Diseases", and even "Endocrine, Nutritional or Metabolic Diseases". It can be seen that this solution divides the disease types from broad to detailed according to the disease classification method of ICD, and the parent node has the addresses of its descendant nodes, so that the user can query the broader or more detailed documents of related diseases up or down.

[0045] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A medical data popularization method based on artificial intelligence, characterized by: The following steps are involved: Step 1: Obtain medical-related literature through multiple data sources; Step 2, constructing a disease tree database according to the disease type classification method of the International Classification of Diseases (ICD), wherein the disease tree database includes multiple nodes; In the step 2, disease types corresponding to the first 24 chapters of the International Classification of Diseases (ICD) are selected to construct 24 disease tree databases, and the disease types corresponding to the chapters are used as root nodes of the disease tree databases; each disease tree database extends descendant nodes according to the classification method of the International Classification of Diseases (ICD), and the disease types with 4-bit codes are used as leaf nodes; Step 3, classify and sort the medical-related literature by disease type and store them in the corresponding nodes of the disease tree database; The step 3 specifically comprises the following steps: Step 3-1, after word segmentation of the document name and keywords, perform word vector encoding, and use the optimal allocation function to store the document in the node that matches the word segmentation. If the node is a non-leaf node, execute step 3-2. If the node is a leaf node, end; Step 3-2, after the content of the document is segmented, the word vector encoding is performed. Among the grandchild nodes of the node stored in step 3-1, the optimal allocation function is used to identify the node that matches the segmentation, and the document is stored in the node that matches the segmentation; The optimal allocation function is: Among them, f(x i ) represents the optimal allocation function; x i Represents the encoding value of the i-th word; y j represents the encoding value of the jth node, j=1,2,...,K, K represents the number of nodes; sgn represents the step function; Represents kernel function parameters; represents the Gaussian radial basis parameters, ; represents the bias term; h i Represents the weight coefficient of the i-th participle; Step 4, receiving the question sentence input by the user, extracting the disease type from the question sentence, and retrieving medical related literature from the node corresponding to the disease type and returning it to the user.

2. The medical data popularization method based on artificial intelligence according to claim 1, characterized in that: In step 4, when retrieving medical related documents from the node corresponding to the disease type and returning them to the user, the recommendation value of the document in the node is calculated: Among them, SID represents the recommendation value; QSC represents the document's shared citation index; ASC represents the document's interactive comment index; E samp represents the page scrolling index within the unit time T. The faster the page scrolling speed within the unit time T, the less the user likes the document. w1 is the weight of QSC, w2 is the weight of ASC, and w3 is The weight of .

3. The medical data popularization method based on artificial intelligence according to claim 1, characterized in that: In step 4, when medical related documents are retrieved from the node corresponding to the disease type and returned to the user, the user is prompted to query the documents in the ancestor node or descendant node of the node.

Citation Information

Patent Citations

  • Disease coding method and system based on original diagnosis data and case history file data

    CN107731269A

  • Medical literature retrieval method and device, electronic equipment and storage medium

    CN112885478A