Medical information extraction method and device, electronic equipment and storage medium

By segmenting and classifying medical literature, extracting medical entities and entity relationships, and constructing a high-quality medical database, the problem of storing irrelevant information in existing technologies is solved, and data quality and retrieval efficiency are improved.

CN120804165APending Publication Date: 2025-10-17ZHONGDIAN DATA IND CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511303191.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies that vectorize entire medical documents result in medical databases storing a large amount of irrelevant information, reducing data quality and user retrieval efficiency.

Method used

By segmenting and classifying medical literature, extracting medical entities and entity relationships, and building a high-quality medical database, only valuable medical information is stored.

Benefits of technology

It improves the data quality and user retrieval efficiency of medical databases, and realizes the structured storage and efficient retrieval of medical information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804165A_ABST
    Figure CN120804165A_ABST
Patent Text Reader

Abstract

The invention relates to a medical information extraction method and device, electronic equipment and a storage medium, and can obtain medical information valuable for constructing a medical database, so that the data quality of the medical database can be improved, and the medical information retrieval efficiency of a user can be improved. The method comprises the following steps: acquiring medical literatures; segmenting the medical literature to obtain segmented literature statements; based on knowledge rules in the medical field, the segmented literature statements are classified to obtain classification tags of the literature statements, and the classification tags of the literature statements are medical information or non-medical information; and performing information extraction on the literature statement marked with the medical information to obtain a medical entity and an entity relationship in the literature statement marked with the medical information, and storing the medical entity and the entity relationship in the literature statement marked with the medical information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of information processing, and particularly relates to a medical information extraction method and device, electronic equipment and storage medium. BACKGROUND

[0002] Up to now, the medical field has accumulated a large amount of medical literature, including electronic medical records, medical book literature and medical papers. These medical literature contains rich medical knowledge, and a medical database constructed by using these medical literature can provide a channel for users to quickly search for medical information.

[0003] After obtaining a medical literature according to the related technology, the medical literature is vectorized, and the obtained vector is stored in a database, which can be used as a medical database. Since the related technology saves the content of the entire medical literature in the medical database, the medical database saves some medical irrelevant information, which reduces the data quality of the medical database and also reduces the efficiency of searching for medical information from the medical database (i.e., reduces the efficiency of searching for medical information). SUMMARY

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a medical information extraction method and device, electronic equipment and storage medium.

[0005] In a first aspect of the embodiments of the present disclosure, a medical information extraction method is provided, which includes: obtaining a medical literature; segmenting the medical literature to obtain segmented literature sentences; classifying the segmented literature sentences based on knowledge rules in the medical field to obtain classification labels of the literature sentences, the classification labels of the literature sentences being medical information or non-medical information; performing information extraction on the literature sentences marked with medical information to obtain medical entities and entity relationships in the literature sentences marked with medical information, and saving the medical entities and entity relationships in the literature sentences marked with medical information in a medical database.

[0006] In a second aspect of the embodiments of the present disclosure, a medical information extraction device is provided, which includes: an obtaining module configured to obtain a medical literature; a segmentation module configured to segment the medical literature to obtain segmented literature sentences; a classification module configured to classify the segmented literature sentences based on knowledge rules in the medical field to obtain classification labels of the literature sentences, the classification labels of the literature sentences being medical information or non-medical information; and an information extraction module configured to perform information extraction on the literature sentences marked with medical information to obtain medical entities and entity relationships in the literature sentences marked with medical information, and save the medical entities and entity relationships in the literature sentences marked with medical information in a medical database.

[0007] In a third aspect, the present disclosure provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the medical information extraction method according to the first aspect.

[0008] In a fourth aspect, the present disclosure provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the medical information extraction method according to the first aspect.

[0009] In a fifth aspect, the present disclosure provides a computer program product, wherein the computer program product comprises a computer program, and when the computer program product is executed on a processor, the computer program causes the processor to execute the computer program to implement the medical information extraction method according to the first aspect.

[0010] In a sixth aspect, the present disclosure provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is configured to execute program instructions to implement the medical information extraction method according to the first aspect.

[0011] The technical solution provided by the embodiments of the present disclosure has the following advantages compared with the prior art: the entire medical literature is segmented to obtain segmented literature sentences, all the segmented literature sentences comprehensively cover the content of the entire medical literature, and the omission of medical information in the medical literature can be avoided. Then, the segmented literature sentences are classified into medical information and non-medical information, the literature sentences irrelevant to medical information (i.e., the literature sentences marked as non-medical information) are filtered out and are not subjected to further information extraction, and the processing efficiency of the medical literature is improved. The literature sentences relevant to medical information (i.e., the literature sentences marked as medical information) are subjected to information extraction, and medical entities and entity relationships are obtained. The medical entities and entity relationships are both medical information of value, and the present solution realizes the acquisition of medical information of value for constructing a medical database, thereby improving the data quality of the medical database for storing the medical entities and entity relationships, and improving the retrieval efficiency of medical information in the medical database by users. In addition, the medical entities and entity relationships are stored in the medical database in a structured manner, and the retrieval of medical information by users in the structured medical entities and entity relationships also improves the retrieval efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0012] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, further serve to explain the principles of the present disclosure.

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required by the embodiments or the prior art description. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without any creative effort.

[0014] Figure 1 A flowchart of a medical information extraction method provided by an embodiment of the present disclosure; Figure 2 A flowchart of a medical information extraction method provided by an embodiment of the present disclosure; Figure 3 A flowchart of a medical information extraction method provided by an embodiment of the present disclosure; Figure 4 A structural block diagram of a medical information extraction device provided by an embodiment of the present disclosure; Figure 5 A structural block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required by the embodiments or the prior art description. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without any creative effort.

[0016] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, not all the embodiments.

[0017] The terms "first", "second", and the like in the specification and claims of the present disclosure are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually a class, not limited to the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in an "or" relationship.

[0018] The following first explains some nouns or terms involved in the claims and the specification of the present disclosure.

[0019] The following first explains some nouns or terms involved in the claims and the specification of the present disclosure.

[0020] (1) Text Vectorization: converting text into computer-processable vectors for storage and retrieval in a vector database.

[0021] (2) Entities: nouns with specific meanings identified from text. For example, medical entities identified from medical literature can include disease names, drug names, symptom descriptions, etc.

[0022] (3) Entity Relationships: used to describe how two or more entities are related to each other. For example, entity relationships in medical literature can include the relationship between drug A and disease B is treatment, and the relationship between disease C and symptom D is cause.

[0023] (4) Triples: a structure that clearly expresses a piece of knowledge with three parts, like the "subject-predicate-object" of a sentence. The structure of triples can be entity 1-relationship-entity 2, where entity 1 is the subject and entity 2 is the object. For example, triples extracted from medical literature can include "cold-cause-cough" and "cold medicine-treat-influenza".

[0024] When building a medical database using medical literature, related technologies perform word segmentation and word vector conversion on the entire medical literature, and then store the converted vectors in the medical database. However, in addition to some segments related to medical information (e.g., segments related to treatment, disease, symptoms, diagnosis, and drugs), the entire medical literature also includes some segments unrelated to medicine (e.g., literature review, medical case background introduction, etc.), which are not medically valuable information for building a medical database. Related technologies vectorize the entire medical literature, which includes both segments related to medical information and segments unrelated to medicine. The vectorization of segments unrelated to medicine generates information that is not medically valuable, and storing information that is not medically valuable in the medical database reduces the quality of the data stored in the medical database. Moreover, storing information that is not medically valuable in the medical database also reduces the efficiency and accuracy of users' retrieval of medical information.

[0025] To address the above problems, the present disclosure provides a medical information extraction method that extracts medical entities and entity relationships between medical entities from medical literature by segmenting and classifying the entire medical literature. The extracted medical information is used to build a medical database, which can improve the quality of the data stored in the medical database. The high-quality medical database obtained by the present disclosure can also improve the efficiency and accuracy of users' retrieval of medical information.

[0026] The electronic devices in the embodiments of the present disclosure may be mobile electronic devices or non-mobile electronic devices. Mobile electronic devices may be mobile phones, tablet computers, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs); non-mobile electronic devices may be personal computers (PCs), etc.; these embodiments are not specifically limited.

[0027] The execution subject of the medical information extraction method provided in the embodiments of the present disclosure can be the above-mentioned electronic device (including mobile electronic devices and non-mobile electronic devices), or it can be a functional module and / or functional entity in the electronic device that can implement the method. The specific execution subject can be determined according to actual usage requirements and is not limited by the embodiments of the present disclosure.

[0028] The following describes in detail a medical information extraction method provided by an embodiment of the present disclosure through specific embodiments and application scenarios in conjunction with the accompanying drawings.

[0029] like Figure 1 As shown, an embodiment of the present disclosure provides a medical information extraction method, which may include the following steps 101 to 107.

[0030] 101. Access medical literature using electronic devices.

[0031] Electronic devices can use web search and web crawler technology to obtain medical literature published on the Internet. Then, the electronic device can use the obtained medical literature to build a medical database. The types of medical literature can include medical papers and medical guidelines.

[0032] In some embodiments, the electronic device may obtain one or more medical documents. If the electronic device obtains multiple medical documents, each of the multiple medical documents may be used to construct a medical database, and the electronic device may construct the same medical database using the multiple medical documents. Optionally, after the electronic device constructs a medical database using the first medical document, the process of constructing the medical database using the next medical document may be referred to as updating the medical database.

[0033] It should be noted that, in the following embodiments, a medical document is used as an example to schematically illustrate the process of constructing a medical database.

[0034] In some embodiments, the electronic device may periodically acquire medical documents published on the Internet, or may receive and respond to a trigger operation input by a user to acquire medical documents published on the Internet, wherein the trigger operation is used to trigger the construction of a medical database.

[0035] Optionally, the electronic device can further receive the medical literature input by the user in response to the trigger operation input by the user.

[0036] 102. The electronic device performs article structure segmentation on the medical literature to obtain structure units marked with structure identifiers.

[0037] Generally, a medical literature can be divided into multiple structure units, and the content topics recorded in the multiple structure units are different. Some structure units record content irrelevant to medical information, and other structure units record more medical information. For example, a medical paper can include multiple structure units such as an abstract, an introduction, chapters, and references, wherein the abstract and the chapters record more medical information, and the introduction and the references record less medical information. For another example, a medical journal paper can include multiple structure units such as a title, a preface, and a main text, wherein the title and the preface record less medical information, and the main text records more medical information.

[0038] Based on this, the embodiments of the present disclosure propose that the article structure segmentation can be performed on the medical literature based on the content topics recorded in each paragraph in the medical literature to obtain at least one structure unit marked with a structure identifier, and the structure units marked with the same structure identifier record the same topic content. In this way, the entire medical literature is divided into structure units marked with different structure identifiers.

[0039] Optionally, the structure identifier can include a title label, an abstract label, an introduction label, a main text label, a chapter label, and a reference label. The structure unit marked with the title label generally includes one sentence, and the structure unit marked with the abstract label, the introduction label, the main text label, or the chapter label generally includes at least one paragraph.

[0040] It should be noted that each structure unit (for example, a title, an abstract, a chapter, or a reference) in the medical literature refers to the specific content included in the structure unit. For example, the abstract in the medical literature refers to all the content in the medical literature that belongs to the abstract, not the name “abstract” in the medical literature; the chapter in the medical literature refers to all the content in the medical literature that belongs to the chapter, not the name “Chapter 1” or “Chapter 2” in the medical literature.

[0041] In some embodiments, as shown in FIG. 2, Figure 2 Step 102 can include steps 201 and 202.

[0042] 201. The electronic device performs article structure segmentation on the medical literature by using a text parsing technology in natural language processing (NLP) to obtain structure units marked with structure identifiers.

[0043] Each structural unit in the medical literature can contain some landmark elements, for example, the medical literature can include "abstract", "introduction", "chapter 1" or "1.1" and the like specific names, therefore, the electronic device can adopt a text parsing technology to perform article structure segmentation on the medical literature based on these landmark elements.

[0044] However, if the format of the medical literature is not standardized and may not contain these landmark elements, the electronic device may fail to perform article structure segmentation on the medical literature by using the text parsing technology. For the medical literature that fails to be segmented by using the text parsing technology, the electronic device can perform article structure segmentation on the medical literature by using a text classification model, that is, step 202 is performed.

[0045] 202. If the medical literature fails to be segmented by using the text parsing technology, perform article structure segmentation on the medical literature by using the text classification model to obtain structural units marked with structural identifiers.

[0046] If the medical literature fails to be segmented by using the text parsing technology, it indicates that the format of the medical literature is not standardized, and the text classification model can be used to perform article structure segmentation on the medical literature.

[0047] Optionally, the electronic device can directly use the text classification model to perform article structure segmentation on the medical literature without using the text parsing technology, to obtain structural units marked with structural identifiers.

[0048] In some embodiments, the electronic device can train the first neural network model by using a first training sample to obtain the text classification model before step 202. The first training sample can include medical literature samples marked with structural identifiers. Training the model by using the medical literature samples marked with structural identifiers can enable the text classification model to learn the text features of each structural unit in the medical literature samples, for example, the abstract in the medical literature samples usually contains the summary of the research purpose, method and conclusion, and the text or chapter in the medical literature samples usually contains the specific operation process and setting parameters. Further, the text classification model can perform article structure segmentation on the medical literature based on the text features of each structural unit in the medical literature.

[0049] Further, it can be known that the text classification model trained by the electronic device has the function of extracting the text features of the medical literature and performing article structure segmentation on the medical literature based on the extracted text features.

[0050] 103. The electronic device performs sentence segmentation on the structural units marked with structural identifiers to obtain literature sentences marked with structural identifiers.

[0051] After the electronic device cuts the medical literature into different structural units, the electronic device can further cut each structural unit into one or more literature sentences. Each literature sentence is marked with the same structural identifier as the structural unit to which the literature sentence belongs.

[0052] It can be understood that the electronic device cuts the medical literature from the whole to the sentence in multiple levels, and the obtained cut literature sentences (i.e., the literature sentences marked with the structural identifier) comprehensively cover the content of the whole medical literature, which can avoid missing or incorrect extraction of medical information in the medical literature due to improper cutting granularity, thereby improving the quality of the constructed medical database.

[0053] In some embodiments, the electronic device can perform sentence cutting on each structural unit to obtain the literature sentences marked with the structural identifier, taking a period as a delimiter. In this way, each literature sentence obtained by cutting includes only one sentence. A sentence refers to a paragraph containing only one period.

[0054] In other embodiments, the electronic device can perform semantic cutting on each structural unit to obtain the literature sentences marked with the structural identifier by using a semantic analysis technique. In this way, each literature sentence obtained by cutting can include one or more sentences.

[0055] Exemplarily, the electronic device can perform semantic cutting on each structural unit by using a semantic analysis technique, which can include: first, performing word segmentation on the structural unit by using a word segmentation technique in natural language processing to obtain a plurality of words in the structural unit; and then, cutting the plurality of words in the structural unit according to semantic coherence in combination with punctuation marks in the structural unit to obtain the literature sentences marked with the structural identifier.

[0056] It can be understood that each literature sentence obtained by the electronic device by performing semantic cutting on each structural unit is a semantically complete sentence. Compared with a semantically incomplete literature sentence, the electronic device has higher accuracy in classifying a semantically complete literature sentence.

[0057] 104、The electronic device classifies the literature sentences marked with the structural identifier based on knowledge rules in the medical field to obtain classification tags of the literature sentences.

[0058] After obtaining the sentence-level literature sentences, the electronic device can first classify the medical information and non-medical information of each literature sentence marked with a structural identifier based on the knowledge rules in the medical field to obtain a classification label of the literature sentence. The classification label of the literature sentence is medical information (or knowledge information) or non-medical information (or non-knowledge information). If the classification label of the literature sentence is medical information, the literature sentence is further processed; if the classification label of the literature sentence is non-medical information, the literature sentence can be deleted. Alternatively, the literature sentence with the classification label of medical information can be referred to as the literature sentence marked with medical information, and the literature sentence with the classification label of non-medical information can be referred to as the literature sentence marked with non-medical information.

[0059] It can be understood that by performing the binary classification of medical information and non-medical information on each literature sentence, the electronic device can exclude some obvious interference information unrelated to medicine (i.e., the literature sentence marked with non-medical information), thereby improving the processing efficiency and quality of medical literature.

[0060] In some embodiments, the knowledge rules in the medical field can include medical terms (for example, disease names, drug names, symptom descriptions, and treatment methods, etc.) and sentence grammars used in the medical field.

[0061] In some embodiments, step 104 can include: the electronic device taking the knowledge rules in the medical field as prompt instructions (or referred to as prompt instructions) required by large language models (LLM); and then using the large language models and the prompt instructions to classify each literature sentence marked with a structural identifier to obtain a classification label of the literature sentence.

[0062] The prompt instruction is a text instruction or question input to the large language model, which is used to guide the large language model to generate a specific and expected output.

[0063] In other embodiments, since the structural units marked with different structural identifiers in the medical literature usually have different text features, the electronic device can set different prompt instructions for the structural units marked with different structural identifiers. For example, the prompt instruction set for the structural units marked with the text label or the chapter label can include the numerical unit of the medical parameter (for example, the unit of blood pressure), and the description words of the body symptoms (for example, dizziness, headache, and bleeding, etc.). Alternatively, the prompt instruction set for each structural unit marked with a structural identifier can be referred to as the prompt instruction corresponding to the structural identifier.

[0064] Further, the electronic device classifies the document sentence marked with the structure identifier by using the large language model and a prompt instruction corresponding to the structure identifier, to obtain a classification label of the document sentence.

[0065] In some embodiments, the electronic device can train the large language model before classifying the document sentence by using the large language model, and use the trained large language model to classify the document sentence.

[0066] 105、The electronic device performs entity recognition on the document sentence marked with medical information to obtain a document sentence marked with medical entities, wherein the document sentence marked with medical entities can include at least one marked word, and each marked word is a medical entity.

[0067] The electronic device can use a named entity recognition (NER) technology in natural language processing to perform entity recognition on the document sentence marked with medical information to obtain a document sentence marked with medical entities. The medical entities can include drug names, disease names, symptom description words, diagnosis results, and the like.

[0068] In some embodiments, the electronic device can mark the medical entities in the document sentence by using a preset marking method (e.g., bold font or a preset font color).

[0069] In some embodiments, the electronic device can train the second neural network model by using the second training sample to obtain a named entity recognition model, and then perform entity recognition on the document sentence marked with medical information by using the named entity recognition model to obtain a document sentence marked with medical entities. The second training sample can include a plurality of medical literature samples marked with medical entities, and the plurality of medical literature samples marked with medical entities can include medical entities of a plurality of medical entity types.

[0070] In some embodiments, the plurality of medical entity types can include drugs, diseases, symptoms, treatment methods, and diagnosis results, etc. Optionally, the plurality of medical entity types can be further refined. Specifically, the drugs can be divided into a plurality of drug categories (e.g., hypertension drugs, cold drugs, and health care drugs, etc.) according to different functions, the diseases can be divided into a plurality of disease types (e.g., common cold, influenza, and pneumonia, etc.), the treatment methods can be divided into interventional treatment and non-interventional treatment according to different invasiveness, and the symptoms can be divided into a plurality of levels of symptoms (e.g., mild symptoms, moderate symptoms, and severe symptoms) according to severity.

[0071] It can be understood that the plurality of medical literature samples of the marked medical entities include medical entities related to a plurality of medical entity types, and the named entity recognition module trained by using the plurality of medical literature samples has the function of recognizing medical entities in the plurality of medical entity types.

[0072] In some embodiments, the second neural network model can be a recurrent neural network (RNN), a long short-term memory (LSTM), or a transformer model.

[0073] Exemplarily, the second neural network model can include an embedding layer, a bidirectional long short-term memory layer (BiLSTM), and a conditional random field (CRF). The embedding layer is used to convert each word input into the second neural network model into a word vector; the BiLSTM layer is used to form a word vector sequence from the word vector from the embedding layer, and generate a context-fused feature representation for each word in the word vector sequence, and generate a possibility score of each medical entity type for each word; the CRF is used to determine the possibility score of each medical entity type corresponding to each word based on the possibility score of each medical entity type corresponding to each word output by the BiLSTM layer, and the constraint condition between medical entities, and determine the possibility score of each medical entity type corresponding to each word again, and take the medical entity type with the highest possibility score as the output of the word. For example, the constraint condition between medical entities can include: a medical noun cannot be connected to a medical verb, a medical treatment method cannot be connected to another medical treatment method, and a drug name cannot be connected to another drug name.

[0074] Further, the named entity recognition model obtained by the electronic device by training the second neural network model composed of the embedding layer, the BiLSTM layer and the CRF layer is also composed of the embedding layer, the BiLSTM layer and the CRF layer.

[0075] It can be understood that the accuracy of each medical entity type entity in the literature sentence determined by the CRF layer based on the constraint condition between medical entities is higher, that is, the electronic device uses the named entity recognition model including the CRF layer to perform entity recognition on the literature sentence marked with medical information, which can improve the accuracy of medical entity recognition.

[0076] In addition, the named entity recognition model including the BiLSTM has a high recall rate for medical entity recognition.

[0077] Optionally, the electronic device can also refer to the structural identification to which the literature sentence belongs when performing entity recognition on each literature sentence marked with medical information. That is, the electronic device can perform entity recognition on the literature sentence based on the structural identification to which the literature sentence belongs, to obtain the literature sentence marked with medical entities.

[0078] It can be understood that the electronic device combines the structural identification to which each literature sentence belongs to perform entity recognition on the literature sentence, which can improve the accuracy of entity recognition.

[0079] 106、The electronic device performs relationship prediction on the literature sentence marked with medical entities to obtain the entity relationship between each group of medical entities in the literature sentence. Each group of medical entities can include two medical entities.

[0080] The electronic device can use a relationship classification model based on a neural network to perform relationship prediction on each literature sentence marked with medical entities to obtain the entity relationship between each group of medical entities in the literature sentence. The entity relationship between a group of medical entities can be, for example, triggering, causing, or treatment, etc. For example, the entity relationship between a cold and a cough is causing, and the entity relationship between a cold medicine and influenza is treatment.

[0081] In some embodiments, the relationship classification model can include an input layer, multiple hidden layers, and an output layer. The electronic device inputs the literature sentence marked with medical entities into the relationship classification model, the input layer in the relationship classification model first performs vector conversion on the literature sentence marked with medical entities to obtain a semantic vector corresponding to the literature sentence marked with medical entities; then, the multiple hidden layers perform relationship prediction on the semantic vector to obtain probability values of various entity relationships corresponding to each group of medical entities in the literature sentence; finally, the output layer can output the entity relationship with the highest probability value as the output of the group of medical entities.

[0082] In some embodiments, the electronic device can use a remote supervision learning technique to train the third neural network model using the third training sample to obtain the relationship classification model. The third training sample can include multiple medical literature samples marked with entity relationships. For example, each two associated medical entities in each medical literature sample marked with entity relationships can be marked with an entity relationship.

[0083] 107、The electronic device uses each group of medical entities in the literature sentence marked with medical entities and the entity relationship between each group of medical entities to form a triple and save it.

[0084] The electronic device can use each group of medical entities in the literature sentence marked with medical entities and the entity relationship between each group of medical entities to form a triple and save it to a medical database. For example, the triple can include: "cold - cause - cough", and "cold medicine - treatment - influenza".

[0085] Furthermore, the electronic device may also save all the obtained triples into a preset database to obtain a medical database.

[0086] As you can understand, the electronic device extracts medical entities and their relationships from medical literature, and uses these relationships to form triples, which represent medical information related to evidence-based medicine. Using this triplet of evidence-based medical information, users can quickly retrieve disease-related medications and treatments, providing accurate decision-making guidance.

[0087] In some embodiments, in addition to saving the medical entities and entity relationships in each document sentence that marks the medical entity in the form of triples, the electronic device can also save the medical entities and entity relationships in each document sentence that marks the medical entity in the form of a medical knowledge graph.

[0088] Optionally, before saving the medical entities and entity relationships in each document sentence that labels a medical entity in the form of a medical knowledge graph, the electronic device may first construct a medical knowledge graph using multiple medical entity types. The electronic device then saves the medical entities and entity relationships in the document sentence to the medical knowledge graph, or in other words, updates the medical knowledge graph using the medical entities and entity relationships in the document sentence.

[0089] Before the electronic device saves the medical entities and entity relationships in the document statement into a medical knowledge graph including multiple medical entity types, the electronic device can first identify the entity type of each medical entity in the document statement, obtain the medical entity type to which each medical entity belongs, and then save each medical entity in the document statement into the medical knowledge graph under the medical entity type to which the medical entity belongs.

[0090] It should be noted that the details of the various medical entity types used to construct the medical knowledge graph can be referred to the introduction of the various medical entity types in the above step 105, which will not be repeated here.

[0091] For example, Figure 3 As shown, the process of the electronic device saving the medical entities and entity relationships in each document sentence that marks the medical entity in the form of a medical knowledge graph may include steps 301 to 303.

[0092] 301. The electronic device matches the document sentence marking the medical entity with the multiple medical entity types included in the medical knowledge graph to determine the medical entity type to which the medical entity in the document sentence belongs.

[0093] The electronic device can employ a semantic matching technology in natural language processing to match each medical entity in each medical entity-labeled literature sentence with all medical entity types included in the medical knowledge graph, and obtain a medical entity type to which each medical entity belongs.

[0094] In some embodiments, if the literature sentence includes a medical entity that is semantically ambiguous or complex, and the electronic device fails to match the medical entity with multiple medical entity types included in the medical knowledge graph using the semantic matching technology, the electronic device can employ a semantic understanding model to identify the entity type of the literature sentence, and obtain a medical entity type to which the medical entity belongs. That is, for a literature sentence that fails to match, the electronic device can employ a semantic understanding model to identify the entity type of the literature sentence, i.e., perform step 302. The literature sentence that fails to match can refer to at least one medical entity that fails to match with multiple medical entity types in the medical knowledge graph (or referred to as a medical entity that fails to match).

[0095] 302、If the medical entity-labeled literature sentence fails to match, employ a semantic understanding model to identify the entity type of the literature sentence, and obtain a medical entity type to which the medical entity in the literature sentence belongs.

[0096] The electronic device inputs the literature sentence into the semantic understanding model, performs semantic encoding and analysis, and outputs a medical entity type to which a medical entity that fails to match in the literature sentence belongs.

[0097] Alternatively, the electronic device can directly employ a semantic understanding model to identify the entity type of the literature sentence, and obtain a medical entity type to which the medical entity in the literature sentence belongs, without employing the semantic matching technology.

[0098] In some embodiments, the electronic device can train the fourth neural network model using fourth training samples before step 302 to obtain the semantic understanding model. The fourth training samples can include a plurality of literature sentence samples labeled with medical entity types. Each literature sentence sample labeled with a medical entity type means that the medical entity in the literature sentence sample is labeled with a corresponding medical entity type.

[0099] For example, the fourth neural network model can be a bidirectional encoder representations from transformers (BERT) model or a transformer model.

[0100] Optionally, when performing entity type identification on the literature sentence marking the medical information, the electronic device can also refer to the structural identification to which the literature sentence belongs. That is, the electronic device can perform entity type identification on the literature sentence based on the structural identification to which the literature sentence belongs, to obtain the medical entity type to which the medical entity in the literature sentence belongs.

[0101] It can be understood that the electronic device performs entity type identification on each literature sentence in combination with the structural identification to which the literature sentence belongs, which can improve the accuracy of entity type identification.

[0102] 303、The electronic device saves the medical entity in the literature sentence marking the medical entity, the medical entity type to which the medical entity belongs, and the entity relationship between each group of medical entities into the medical knowledge graph.

[0103] The electronic device extracts medical information such as medical entities, medical entity types, and entity relationships between medical entities from medical literature, and saves the medical information into a medical knowledge graph, thereby realizing the structuring of medical information in medical literature. The medical knowledge graph is used to store the structured medical information.

[0104] Further, when a user needs to query any medical information (for example, the treatment effect of any drug, the treatment method of any disease), the electronic device inputs the keyword to be searched. The electronic device can quickly search for medical information related to the keyword to be searched from the structured medical information stored in the medical knowledge graph. That is, the medical knowledge graph constructed by the electronic device can provide an efficient and fast search channel for the user.

[0105] Exemplarily, the structure of the medical knowledge graph can be a tree structure, based on which the electronic device can save various medical entity types as root nodes of the medical knowledge graph. Then, after predicting the entity relationship and identifying the medical entity type for each literature sentence marking the medical entity, the electronic device saves each medical entity in the literature sentence as a child node of the root node to which the medical entity belongs to the medical knowledge graph, and saves each entity relationship in the literature sentence as an association relationship between two child nodes corresponding to the entity relationship to the medical knowledge graph. The root node to which each medical entity belongs is the medical entity type to which the medical entity belongs.

[0106] It can be known that the root node of the medical knowledge graph obtained by the electronic device is a plurality of medical entity types, the child node of the medical knowledge graph is a medical entity in the literature sentence marking the medical entity, and the association relationship between two child nodes of the medical knowledge graph is an entity relationship in the literature sentence marking the medical entity.

[0107] It is understandable that the electronic device classifies medical entities in medical literature according to medical entity type and stores them in a medical knowledge graph. This medical knowledge graph records the associations between different medical information. Furthermore, when a user needs to query any medical information, they can not only retrieve the medical information from the medical knowledge graph, but also retrieve other medical information related to the medical information. In other words, the user can retrieve more useful medical information, improving retrieval efficiency. For example, multiple medical documents may record different drug names for the same cold medicine E. If the electronic device classifies the medical entities in these multiple medical documents and stores them in the medical knowledge graph, the medical information of the cold medicine E in these multiple medical documents can be associated with the same medical entity type (e.g., cold medicine). Furthermore, when a user searches for any of the drug names of the cold medicine E, they can obtain the medical information of the cold medicine E in these multiple medical documents, improving retrieval efficiency.

[0108] In some embodiments, unlike the description of step 303 above, the electronic device may store only medical entities in the medical knowledge graph, without storing the entity relationships between medical entities. Specifically, the electronic device stores the medical entities in the literature sentences that tagged the medical entities, as well as the medical entity types to which the medical entities belong, in the medical knowledge graph. Furthermore, users can use these triples and the medical knowledge graph simultaneously for searches.

[0109] Figure 4 This is a structural block diagram of a medical information extraction device shown in an embodiment of the present disclosure. Figure 4 As shown, it includes an acquisition module 401, a segmentation module 402, a classification module 403 and an information extraction module 404.

[0110] Among them, the acquisition module 401 is used to acquire medical literature. The segmentation module 402 is used to segment the medical literature to obtain segmented literature sentences. The classification module 403 is used to classify the segmented literature sentences based on knowledge rules in the medical field to obtain classification labels for the literature sentences, where the classification labels of the literature sentences are medical information or non-medical information. The information extraction module 404 is used to extract information from the literature sentences marked with medical information, obtain medical entities and entity relationships in the literature sentences marked with medical information, and store the medical entities and entity relationships in the literature sentences marked with medical information in the medical database.

[0111] Optionally, the segmentation module 402 is specifically configured to: segment the medical literature according to article structure to obtain structural units marked with structural identifiers, wherein the structural identifiers include at least one of the following: a title label, an abstract label, a brief introduction label, a main text label, a chapter label, and a reference label; and segment the structural units marked with the structural identifiers into sentences to obtain literature sentences marked with the structural identifiers, which are segmented literature sentences.

[0112] Optionally, the segmentation module 402 is specifically configured to: segment the medical literature according to article structure by using a text parsing technology in natural language processing to obtain structural units marked with structural identifiers; or segment the medical literature according to article structure by using a text classification model to obtain the structural units marked with the structural identifiers.

[0113] Optionally, the segmentation module 402 is specifically configured to: segment the structural units into sentences by using a period as a delimiter to obtain the literature sentences marked with the structural identifiers; or segment each structural unit by using a semantic analysis technology to obtain the literature sentences marked with the structural identifiers.

[0114] Optionally, the information extraction module 404 is specifically configured to: perform entity recognition on the literature sentences marked with medical information to obtain literature sentences marked with medical entities, wherein the literature sentences marked with the medical entities include at least one labeled medical entity; and perform relationship prediction on the literature sentences marked with the medical entities to obtain an entity relationship between each group of medical entities in the literature sentences marked with the medical entities, wherein each group of medical entities includes two medical entities.

[0115] Optionally, the information extraction module 404 is specifically configured to: form a triple with each group of medical entities in the literature sentences marked with the medical entities and the entity relationship between each group of medical entities, and save the triple in a medical database.

[0116] Optionally, the information extraction module 404 is further configured to: determine a medical entity type to which a medical entity in the literature sentences marked with the medical entities belongs; and save the medical entity in the literature sentences marked with the medical entities, the medical entity type to which the medical entity belongs, and the entity relationship between each group of medical entities into a medical knowledge graph.

[0117] Optionally, the information extraction module 404 is specifically configured to: match the literature sentences marked with the medical entities with a plurality of preset medical entity types to determine the medical entity type to which a medical entity in the literature sentences marked with the medical entities belongs; or use a semantic understanding model to identify the medical entity type to which a medical entity in the literature sentences marked with the medical entities belongs.

[0118] Optionally, the medical knowledge graph is tree-structured, and the root node of the medical knowledge graph includes multiple medical entity types. Information extraction module 404 is specifically configured to: store the medical entity in the literature sentence labeled with the medical entity as a child node of the root node to which the medical entity belongs in the medical knowledge graph, where the root node to which the medical entity belongs is the medical entity type to which the medical entity belongs; and store the entity relationship in the literature sentence labeled with the medical entity as an association relationship between two child nodes corresponding to the entity relationship in the medical knowledge graph.

[0119] Optionally, the multiple medical entity types include: multiple drug categories, multiple disease types, multiple levels of symptoms and multiple treatment methods.

[0120] Optionally, the classification module 403 is specifically used to: generate prompt instructions required by the large language model based on knowledge rules in the medical field; use the large language model and prompt instructions to classify document sentences marked with structural identifiers to obtain classification labels for the document sentences.

[0121] Optionally, the classification module 403 is specifically used to: generate prompt instructions corresponding to multiple structural identifiers based on knowledge rules in the medical field; use a large language model and prompt instructions corresponding to the structural identifiers marked by the document sentences marked with structural identifiers to classify the document sentences and obtain classification labels for the document sentences.

[0122] In the embodiments of the present disclosure, each module can implement the medical information extraction method provided by the above method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described here.

[0123] Figure 5 The structural diagram of an electronic device provided in the embodiment of the present disclosure is used to exemplify the electronic device that implements any medical information extraction method in the embodiment of the present disclosure, and should not be understood as a specific limitation of the embodiment of the present disclosure.

[0124] like Figure 5 As shown, electronic device 500 may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 502 or programs loaded from storage device 508 into random access memory (RAM) 503. RAM 503 also stores various programs and data required for the operation of electronic device 500. Processor 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0125] Generally, the following devices can be connected to the I / O interface 505: input devices 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 508 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 509. The communication devices 509 can allow the electronic device 500 to communicate wirelessly or wired with other devices to exchange data. While the electronic device 500 is shown with various devices, it is understood that all of the shown devices are not required to be implemented or possessed. More or less devices can alternatively be implemented or possessed.

[0126] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 509, or installed from the storage devices 508, or installed from the ROM 502. When the computer program is executed by the processor 501, the functions defined in any of the medical information extraction methods provided by embodiments of the present disclosure can be performed.

[0127] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. EPROM or flash memory. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, a RF (radio frequency), or the like, or any suitable combination of the above.

[0128] In some embodiments, the client, server can communicate using any currently known or future developed network protocol, such as hyperText transfer protocol (HTTP), and can be interconnected with digital data communication (e.g., communication network) of any form or medium (e.g., communication network). Examples of communication networks include local area networks (LAN), wide area networks (WAN), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.

[0129] The above computer-readable medium can be contained in the above electronic device; or can exist separately without being assembled into the electronic device.

[0130] The computer readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to implement any of the medical information extraction methods provided in the embodiments of the present disclosure.

[0131] In the embodiments of the present disclosure, the computer program code for performing the operations of the present disclosure can be written in one or more programming languages or a combination thereof, including an object-oriented programming language, such as Java, Smalltalk, C++, and a conventional procedural programming language, such as "C" language or a similar programming language. The program code can be executed entirely on a computer, partially on a computer, as an independent software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or a server. In the case involving a remote computer, the remote computer can be connected to the computer through any kind of network, including a LAN or a WAN, or can be connected to an external computer (for example, through the Internet by using an Internet service provider).

[0132] The flowcharts and block diagrams in the drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment, or a part of code containing one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different orders than those noted in the drawings. For example, two blocks represented in succession can actually be executed substantially in parallel, and sometimes in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and the combination of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0133] The units involved in the embodiments of the present disclosure can be implemented in a software manner or in a hardware manner. In some cases, the name of the unit does not constitute a limitation on the unit itself.

[0134] The functionality described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), etc.

[0135] In the context of the present disclosure, a computer-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disks, RAM, ROM, EPROM (or flash memory), optical fiber, CD-ROMs, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0136] The above description is only preferred embodiments of the present disclosure and a description of the principles of the technology used. It should be understood by those skilled in the art that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.

[0137] Moreover, while operations are depicted in a particular order, this should not be understood as requiring such an order nor infringing on the scope of the disclosure. Certain of the operations described in the discussion are combinable into a single operation, and certain operations can be separated into several operations. In some embodiments, the operations described in the discussion can be performed in an order different than presented in the discussion. In some embodiments, the operations described in the discussion can be performed concurrently. Also, while several specific implementation details are discussed in the discussion, these should not be interpreted as limiting the scope of the disclosure. Rather, certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.

[0138] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A medical information extraction method, characterized in that: The method comprises: access to medical literature; Segmenting the medical literature to obtain segmented literature sentences; Based on the knowledge rules in the medical field, the segmented literature sentences are classified to obtain classification labels of the literature sentences, where the classification labels of the literature sentences are medical information or non-medical information; Information extraction is performed on the literature sentences marking the medical information to obtain the medical entities and entity relationships in the literature sentences marking the medical information, and the medical entities and entity relationships in the literature sentences marking the medical information are stored in a medical database.

2. The method according to claim 1, characterized in that The segmenting of the medical literature to obtain segmented literature sentences includes: Segmenting the medical literature into article structures to obtain structural units marked with structural identifiers, wherein the structural identifiers include at least one of the following: a title tag, an abstract tag, an introduction tag, a body tag, a chapter tag, and a reference tag; Sentence segmentation is performed on the structural unit marked with the structural identifier to obtain a document sentence marked with the structural identifier, and the document sentence marked with the structural identifier is the segmented document sentence.

3. The method according to claim 2, characterized in that The sentence segmentation of the structural unit marked with the structural identifier to obtain the document sentence marked with the structural identifier includes: Using a period as a separator, the structural unit is segmented into sentences to obtain the document sentences marked with the structural identifier; Alternatively, a semantic analysis technique is used to perform semantic segmentation on each structural unit to obtain the document sentences marked with the structural identifier.

4. The method according to claim 2, characterized in that The segmented literature sentences are classified based on the knowledge rules in the medical field to obtain classification labels for the literature sentences, including: Based on the knowledge rules in the medical field, generating prompt instructions corresponding to the plurality of structure identifiers; The document sentences are classified using a large language model and a prompt instruction corresponding to the structural identifier marked by the document sentence marked with the structural identifier to obtain a classification label of the document sentence.

5. The method according to any one of claims 1 to 4, characterized in that The extracting information from the document sentences of the marked medical information to obtain medical entities and entity relationships in the document sentences of the marked medical information includes: Performing entity recognition on the document sentences of the marked medical information to obtain document sentences of marked medical entities, wherein the document sentences of marked medical entities include at least one marked medical entity; Relationship prediction is performed on the literature sentences of the marked medical entities to obtain entity relationships between each group of medical entities in the literature sentences of the marked medical entities, where each group of medical entities includes two medical entities.

6. The method according to claim 5, characterized in that The medical entities and entity relationships in the document sentences of the marked medical information stored in the medical database include: Each group of medical entities in the document sentence of the marked medical entity and the entity relationship between each group of medical entities are used to form a triple, and the triple is stored in the medical database.

7. The method according to claim 5, characterized in that The method further comprises: Determining the medical entity type to which the medical entity in the document sentence of the marked medical entity belongs; The medical entities in the literature sentences marking the medical entities, the medical entity types to which the medical entities belong, and the entity relationships between each group of medical entities are saved in the medical knowledge graph.

8. The method according to claim 7, characterized in that The medical knowledge graph is a tree structure, and the root node of the medical knowledge graph includes multiple medical entity types; The step of storing the medical entities in the literature sentences marking the medical entities, the medical entity types to which the medical entities belong, and the entity relationships between each group of medical entities in the medical knowledge graph includes: Saving the medical entity in the document sentence marking the medical entity as a child node of a root node to which the medical entity belongs in the medical knowledge graph, where the root node to which the medical entity belongs is the medical entity type to which the medical entity belongs; The entity relationship in the literature sentence of the marked medical entity is saved in the medical knowledge graph as the association relationship between two child nodes corresponding to the entity relationship.

9. An electronic device, characterized in that: include: A memory and a processor, the memory is used to store a computer program; the processor is used to execute the medical information extraction method according to any one of claims 1 to 8 when calling the computer program.

10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the medical information extraction method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Cause knowledge graph event detection method fusing extended features

    CN112241457A

  • Method and device for extracting cell marker genes in literature

    CN119150861A

  • Biomedical knowledge graph

    US20250246318A1

  • Machine learning-based medicine recognition method and related device

    WO2021174695A1