Information processing method and device and storage medium
By determining the target label of disease entities in the medical knowledge graph, the problem of low accuracy of medical content labeling in the prior art is solved, and a higher degree of label labeling accuracy and automation is achieved.
Patent Information
- Application Number
- CN202311473539.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-07
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has accuracy problems in the annotation of disease in medical content, which is affected by manual level, number of training samples and equipment calculation capabilities.
When receiving the medical tag labeling instruction, the disease entity is acquired and its corresponding target label is determined in the medical knowledge graph, thereby automatically labeling the label of the disease entity.
Improve the accuracy of labeling and reduce dependence on manual level, training sample number and device computing capabilities.
Smart Images

Figure CN119993522A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information processing technology, and in particular to an information processing method and device, and a storage medium. Background Art
[0002] Medical and health content is an important source for obtaining medical knowledge and learning about health and disease. In particular, searching for relevant medical and health content based on diseases is one of the important paths. Therefore, accurate identification and labeling of diseases in medical content, as well as accurate association labeling of other disease information such as the department to which the disease belongs and applicable drugs are particularly critical.
[0003] In related technologies, medical content is labeled manually, or machine learning or deep learning algorithms are used to train a large amount of labeled data to allow the system to automatically learn to classify and label different types of content. However, due to the influence of subjective factors due to manual labor issues, the accuracy is affected. Automatic labeling through training systems requires a large amount of training data. In the case of insufficient training samples or insufficient computing power of the equipment, the accuracy of labeling will also be affected, thereby reducing the accuracy of labeling. Summary of the invention
[0004] In order to solve the above technical problems, the embodiments of the present application hope to provide an information processing method and device, and a storage medium, which can improve the accuracy of labeling.
[0005] The technical solution of this application is implemented as follows:
[0006] The present invention provides an information processing method, which includes:
[0007] When receiving a medical label marking instruction, obtaining a disease entity according to the medical label marking instruction;
[0008] In the medical knowledge graph, determine the target label corresponding to the disease entity;
[0009] The disease entity is labeled according to the target label.
[0010] The present application provides an information processing device, the device comprising:
[0011] An acquisition unit, configured to acquire a disease entity according to a medical labeling instruction when a medical labeling instruction is received;
[0012] A determination unit, used to determine a target label corresponding to the disease entity in the medical knowledge graph;
[0013] A labeling unit is used to label the disease entity according to the target label.
[0014] The present application provides an information processing device, the device comprising:
[0015] A memory, a processor and a communication bus, wherein the memory communicates with the processor via the communication bus, and the memory stores an information processing program executable by the processor. When the information processing program is executed, the information processing method described above is executed by the processor.
[0016] An embodiment of the present application provides a storage medium having a computer program stored thereon, which is applied to an information processing device, and is characterized in that the computer program implements the above-mentioned information processing method when executed by a processor.
[0017] The embodiment of the present application provides an information processing method and device, and a storage medium. The information processing method includes: upon receiving a medical labeling instruction, obtaining a disease entity according to the medical labeling instruction; determining a target label corresponding to the disease entity in a medical knowledge graph; and labeling the disease entity according to the target label. The above method is implemented in that the information processing device obtains the disease entity and uses the medical knowledge graph to determine the target label corresponding to the disease entity, that is, the target label corresponding to the disease entity can be found in the medical knowledge graph, and the disease entity is labeled according to the target label, so that the labeling process will not be affected by the problem of manual level, the number of training samples, and the computing power of the equipment, thereby improving the accuracy of labeling. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A flow chart of an information processing method provided in an embodiment of the present application;
[0019] Figure 2 A schematic diagram of constructing multiple disease index vectors provided in an exemplary embodiment of the present application;
[0020] Figure 3 An exemplary medical knowledge graph provided for an embodiment of the present application;
[0021] Figure 4 An exemplary information processing block diagram provided for an embodiment of the present application;
[0022] Figure 5 A schematic diagram of the structure of an information processing device provided in an embodiment of the present application Figure 1 ;
[0023] Figure 6 A schematic diagram of the structure of an information processing device provided in an embodiment of the present application Figure 2 . DETAILED DESCRIPTION
[0024] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. It should be understood that the specific embodiments described here are only used to explain the present application and are not used to limit the present application.
[0025] The present application provides an information processing method, which is applied to an information processing device. Figure 1 A flow chart of an information processing method provided in an embodiment of the present application, such as Figure 1 As shown, the information processing method may include:
[0026] S101. Upon receiving a medical labeling instruction, obtaining a disease entity according to the medical labeling instruction.
[0027] An information processing method provided in an embodiment of the present application is suitable for the scenario of labeling disease entities.
[0028] In the embodiments of the present application, the information processing device can be implemented in various forms. For example, the information processing device described in the present application may include devices such as mobile phones, cameras, tablet computers, laptop computers, PDAs, portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, and devices such as digital TVs, desktop computers, and servers.
[0029] In an embodiment of the present application, the medical label marking instruction can be an instruction transmitted by a user to an information processing device; it can also be an instruction transmitted by other devices to the information processing device. The specific way in which the information processing device obtains the medical label marking instruction can be determined based on actual conditions, and the embodiment of the present application does not limit this.
[0030] In an embodiment of the present application, the information processing device can obtain the disease entity from the medical label marking instruction; it can also obtain the disease instance according to the disease entity acquisition method indicated by the medical label marking instruction; the specific way in which the information processing device obtains the disease entity can be determined according to actual conditions, and the embodiment of the present application does not limit this.
[0031] In the embodiment of the present application, the disease entity includes the name of the disease, etc.
[0032] In the embodiment of the present application, the number of disease entities can be one or more. The specific number of disease entities can be determined based on actual conditions, and the embodiment of the present application does not limit this.
[0033] In an embodiment of the present application, the process of an information processing device obtaining disease entities according to medical labeling instructions includes: obtaining a medical file to be labeled from the medical labeling instructions; and extracting disease entities from the medical file to be labeled using a language learning model.
[0034] In an embodiment of the present application, the medical file to be annotated may be a comma-separated values (CSV) file, a text document (TXT) file, or a file in other forms. The specific medical file to be annotated may be determined based on actual conditions and is not limited in the embodiment of the present application.
[0035] It should be noted that the medical label marking instruction carries the medical file to be marked.
[0036] In an embodiment of the present application, a language learning model (LLM) may be a model configured in an information processing device, or a model transmitted to the information processing device by other devices, or a model obtained by the information processing device through other methods. The specific method in which the information processing device obtains the language learning model may be determined based on actual conditions, and the embodiment of the present application does not limit this.
[0037] In an embodiment of the present application, it may also be a model trained by an information processing device for identifying disease entities in a medical file to be annotated.
[0038] In an embodiment of the present application, the process of extracting disease entities from medical documents to be annotated using a language learning model can be to use the language learning model to search for disease entities in the medical documents to be annotated and obtain the found disease entities; disease entities can also be extracted from medical documents to be annotated using a language learning model in other ways; the specific implementation process can be determined based on actual conditions, and the embodiment of the present application is not limited to this.
[0039] S102. In the medical knowledge graph, determine the target label corresponding to the disease entity.
[0040] In an embodiment of the present application, after the information processing device obtains the disease entity according to the medical labeling instruction, it can determine the target label corresponding to the disease entity in the medical knowledge graph.
[0041] In an embodiment of the present application, the process of an information processing device determining a target label corresponding to a disease entity in a medical knowledge graph includes: matching the disease entity with multiple disease index vectors in a disease database to obtain a target disease index; and determining a target label that matches the target disease index in the medical knowledge graph.
[0042] In an embodiment of the present application, the number of target disease indexes may be one or more. The specific number of target disease indexes may be determined based on actual conditions, and the embodiment of the present application does not limit this.
[0043] In the embodiment of the present application, if one disease entity can correspond to one target disease index; one disease entity can also correspond to multiple target disease indexes, the specific ones can be determined according to actual conditions and are not limited in the embodiment of the present application.
[0044] In an embodiment of the present application, in a medical knowledge graph, the process of determining a target label that matches a target disease index can be to compare the target disease index with the information at each node in the medical knowledge graph to obtain multiple comparison results; obtain a target comparison result that is greater than or equal to a preset threshold among the multiple comparison results; determine the associated node information at the target node corresponding to the target comparison result, and use the associated node information as the target label.
[0045] In an embodiment of the present application, the information processing device matches the disease entity with multiple disease index vectors in the disease database to obtain a target disease index, including: performing vector conversion on the disease entity to obtain a disease entity vector; determining the similarity between the disease entity vector and each of the multiple disease index vectors in turn to obtain multiple similarities; based on the multiple similarities, obtaining a preset number of target index vectors from the multiple disease index vectors; and determining the target disease index based on the target index vector.
[0046] In an embodiment of the present application, a method for performing vector transformation on a disease entity to obtain a disease entity vector can be to embed the disease entity into a vector to obtain a disease entity vector; or the disease entity can be vector transformed by other vector transformation methods to obtain a disease entity vector; the specific implementation method can be determined based on actual conditions, and the embodiment of the present application is not limited to this.
[0047] In an embodiment of the present application, the information processing device can determine the cosine similarity between the disease entity vector and each vector in a plurality of disease index vectors in turn to obtain multiple similarities; it can also determine the log-likelihood similarity between the disease entity vector and each vector in a plurality of disease index vectors in turn to obtain multiple similarities; it can also determine the cosine similarity between the disease entity vector and each vector in a plurality of disease index vectors in turn through other similarity algorithms to obtain multiple similarities; specifically, the method of determining the cosine similarity between the disease entity vector and each vector in a plurality of disease index vectors in turn to obtain multiple similarities can be determined according to actual conditions, and the embodiment of the present application does not limit this.
[0048] In an embodiment of the present application, the process of an information processing device obtaining a preset number of target index vectors from multiple disease index vectors based on multiple similarities includes: sorting multiple similarities according to their similarity values to obtain a similarity sequence; starting from the position of the maximum similarity value in the similarity sequence, screening a preset number of target similarities; and determining the target index vector corresponding to the target similarity in multiple disease index vectors.
[0049] In the embodiment of the present application, the process of sorting multiple similarities according to their similarity values to obtain a similarity sequence may be sorting multiple similarities in descending order of similarity values to obtain a similarity sequence.
[0050] It should be noted that the preset number can be the number configured in the information processing device, the number transmitted to the information processing device by other devices, or the number obtained by the information processing device in other ways. The specific way in which the information processing device obtains the preset number can be determined based on actual conditions, and the embodiments of the present application do not limit this.
[0051] Exemplarily, the value of the preset number can be 1 or 2. The specific value of the preset number can be determined based on actual conditions, and the embodiments of the present application do not limit this.
[0052] In the embodiment of the present application, the process of screening a preset number of target similarities starting from the position of the maximum similarity value in the similarity sequence can be screening a preset number of target similarities starting from the similarity with the largest similarity value in the similarity sequence (starting from the sequence head of the similarity sequence).
[0053] In an embodiment of the present application, the process of automatic content labeling (disease entity labeling) includes: extracting disease entities from medical content based on a large model (language learning model) (obtaining disease entities according to medical labeling instructions), similarity matching of disease entity label vector index library data (matching disease entities with multiple disease index vectors in the disease database to obtain target disease indexes), knowledge graph information query and extraction of related label information (determining target labels that match target disease indexes in the medical knowledge graph), label annotation supplementation and vector matching (labeling disease entities according to target labels) four parts.
[0054] Exemplary process of extracting disease entities from medical content based on the LLM big model: Receive Prompt instruction: The content of the Prompt instruction is to request medical NLP medical experts to help extract disease entities from the article. The entity is required to be strongly related to the context of the article. The number is at most 1. It is not necessary to just list the entities. The content of the article is in {content}.
[0055] Exemplarily, the process of similarity matching of disease entities with label vector index library data is as follows: similarity matching is performed between disease entities extracted based on the LLM large model and disease index vectors in the vector library to obtain an entity with the highest similarity (matching the disease entity with multiple disease index vectors in the disease database to obtain a target disease index).
[0056] The specific implementation method of matching similarity based on the index library FAISS can be:
[0057] results_with_scores=docsearch.similarity_search_with_score(query=word,k=1)
[0058] For example, if the disease entity extracted by LLM is type 2 diabetes, the most relevant entity is determined after vector similarity matching: diabetes.
[0059] The specific steps of similarity matching are as follows:
[0060] Document representation: First, each document in the document collection is converted into a representation vector (multiple medical disease metadata are vectorized to obtain multiple disease index vectors). Documents can be converted into numerical vector representations using a text representation method (such as bag-of-words, TF-IDF, Word2Vec, etc.). The vector representation of each document captures the key features of the document.
[0061] Query representation: Convert the input query string into a representation vector (convert the disease entity into a vector to obtain the disease entity vector). Similar to document representation, the same text representation method is used to convert the query string into a vector representation.
[0062] Similarity calculation: The cosine similarity similarity measurement algorithm is used to calculate the similarity score between the query vector and each document vector (the similarity between the disease entity vector and each of the multiple disease index vectors is determined in turn to obtain multiple similarities).
[0063] Sorting and returning: Sort the document collection according to the similarity score, put the documents most relevant to the query string in front, and return the highest-ranked document as the search result (according to multiple similarities, obtain a preset number of target index vectors from multiple disease index vectors; determine the target disease index according to the target index vector.).
[0064] Exemplary process of knowledge graph information query and extraction of related label information: query the medical knowledge graph based on the disease entity with the highest similarity extracted in the previous stage (in the medical knowledge graph, determine the target label that matches the target disease index), and obtain the labels that need to be labeled, such as: department, symptoms, etc.
[0065] For example, query based on Neo4j graph database:
[0066] MATCH(d:Disease{name:'Adult respiratory disease'})-[:BELONGS_TO]->(dpt:Department)RETURN d.name AS disease,dpt.name AS department
[0067] For example: Based on adult respiratory diseases, the matching department can be queried from the knowledge graph database (medical knowledge graph) as the respiratory department.
[0068] Exemplary, the process of label annotation supplementation and vector matching: labels in specific scenarios may be missing in the medical atlas, labeling is performed based on the LLM large model, the labeled label execution path is matched, and TOPN accurate label information is obtained (labeling disease entities according to target labels).
[0069] In an embodiment of the present application, the information processing device matches the disease entity with multiple disease index vectors in the disease database, and before obtaining the target disease index, it also obtains multiple medical disease metadata; performs vector conversion processing on the multiple medical disease metadata to obtain multiple disease index vectors; and adds the multiple disease index vectors to the disease database.
[0070] In an embodiment of the present application, the information processing device can obtain multiple medical disease metadata from a medical-related database; it can also obtain multiple medical disease metadata from information input by a user; the specific way in which the information processing device obtains multiple medical disease metadata can be determined based on actual conditions, and the embodiment of the present application does not limit this.
[0071] It should be noted that multiple medical disease metadata include diseases, symptoms, departments, etc.
[0072] In an embodiment of the present application, a method for performing vector conversion processing on multiple medical disease metadata to obtain multiple disease index vectors can be to embed multiple medical disease metadata into vectors to obtain multiple disease index vectors, or to perform vector conversion processing on multiple medical disease metadata through other vector conversion methods to obtain multiple disease index vectors; the specific implementation method can be determined according to actual conditions, and the embodiment of the present application is not limited to this.
[0073] In the embodiment of the present application, the method of performing vector conversion on the disease entity to obtain the disease entity vector may be the same as the method of performing vector conversion processing on multiple medical disease metadata to obtain multiple disease index vectors.
[0074] In the embodiment of the present application, the content label metadata (multiple medical disease metadata) is embedded based on the LLM model to form an index vector (multiple disease index vectors) and stored in a vector database (disease database, such as FAISS). The vector database then compares the indexed query label vector with the index label vector in the data set, and can find the nearest label based on the similarity measurement (determine the similarity between the disease entity vector and each of the multiple disease index vectors in turn, thereby determining the target disease index in the multiple disease index vectors). The specific process is as follows Figure 2 As shown: obtaining multiple medical disease metadata (vector); performing vector conversion processing (indexing) on multiple medical disease metadata to obtain multiple disease index vectors; adding multiple disease index vectors to a disease database (vectorDB), and upon receiving a medical labeling instruction, obtaining a disease entity according to the medical labeling instruction (Querying); post-processing (Post Processing) the disease entity and multiple disease index vectors in the disease database to determine a target disease index.
[0075] In the embodiment of the present application, based on the python language langchain and the index library FAISS, vectors are indexed. An exemplary Index process may be:
[0076] Step a: Introduction of dependent projects
[0077] from langchain.chains import LLMChain
[0078] from langchain.embeddings import OpenAIEmbeddings
[0079] from langchain.vectorstores import FAISS
[0080] from langchain.chat_models import ChatOpenAI
[0081] Step b: LLM call, embedding and index library FAISS initial index
[0082] llm=ChatOpenAI(model_name="gpt-3.5-turbo-16k",max_tokens=4096,temperature=0.1)
[0083] embeddings=OpenAIEmbeddings()
[0084] docsearch=FAISS.from_documents(content,embeddings)
[0085] Step c: Data storage after FAISS indexing (you can specify the storage path and index name), and user subsequent index querying process
[0086] docsearch.save_local(folder_path=index_file_path,index_name='faissIndex')
[0087] In an embodiment of the present application, the information processing device obtains initial medical information from a medical database before determining a target tag that matches a target disease index in a medical knowledge graph; performs standard processing on the initial medical information using preset data standardization processing rules to obtain a medical data set; and establishes medical associations between the data in the medical data set to obtain a medical knowledge graph.
[0088] In the embodiment of the present application, the medical database may be an open source data set or a knowledge base within the medical system; the specific medical database may be determined based on actual conditions, and the embodiment of the present application does not limit this.
[0089] It should be noted that the initial medical information includes information such as diseases, symptoms, departments, diagnostic examination items, commonly used drugs for the disease, examinations required for the disease, and recommended drugs for the disease.
[0090] In an embodiment of the present application, the preset data specification processing rules may be rules configured in the information processing device, or may be rules transmitted to the information processing device from other devices, or may be rules obtained by the information processing device in other ways. The specific way in which the information processing device obtains the preset data specification processing rules may be determined based on actual conditions, and the embodiment of the present application does not limit this.
[0091] It should be noted that the preset data standard processing rules can form standard standard rules for ETL (Extract-Transform-Load) processes such as data cleaning and extraction, or they can be other standard processing rules. The specific preset data standard processing rules can be determined according to actual conditions, and the embodiments of the present application are not limited to this.
[0092] In an embodiment of the present application, the information processing device establishes a medical association relationship between data in a medical data set to obtain a medical knowledge graph, including: matching related data in the medical data set to obtain the association relationship between the data; and establishing a medical knowledge graph based on the association relationship and the data in the medical data set.
[0093] In an embodiment of the present application, the information processing device can use the MATCH statement to find the entity to be associated, and use the CREATE statement to create a relationship between two nodes (associated entities), thereby obtaining an association relationship.
[0094] In the embodiment of the present application, a standard data set (medical data set) can be formed based on an internal knowledge base or an open source data set (medical database) through ETL processes such as data cleaning and extraction, and stored in a graph database such as Neo4J, NebulaGraph, etc. for the subsequent construction of a medical knowledge graph. The medical knowledge graph mainly includes medical entities such as diseases, symptoms, departments, diagnostic examination items, etc., as well as entity relationships such as commonly used drugs for diseases, required examinations for diseases, and recommended drugs for diseases. The medical knowledge graph after construction is as follows Figure 3 As shown: Taking adult respiratory disease as an example, patients with the disease are recommended to eat lily sugar porridge, lily porridge, etc. The patient must eat lotus seeds, etc. Items that need to be checked include pulmonary capillaries, etc.
[0095] It should be noted that the construction process of the medical knowledge graph includes two parts: medical entity construction and entity relationship construction. The specific process of building a medical knowledge graph based on Neo4j graph data (medical data set) includes:
[0096] Exemplarily, the process of constructing medical entities is to collect csv files and import them into the software (or database) used to construct the knowledge graph:
[0097] LOAD CSV WITH HEADERS FROM'file: / / / medical_entities.csv'AS row CREATE(entity:MedicalEntity{id:row.id,name:row.name,type:row.type})
[0098] Assume there is a CSV file named "medical_entities.csv" which contains data of medical entities. The file includes column headers such as id, name and type.
[0099] Exemplarily, the process of constructing entity relationships may be:
[0100] MATCH(entity1:MedicalEntity{name:'Adult respiratory disease'})MATCH(entity2:MedicalEntity{name:'Lotus seeds'})CREATE(entity1)-[:RELATIONSHIP_TYPE]->(entity2)
[0101] MATCH(entity1:MedicalEntity{name:'Adult Respiratory Disease'})MATCH(entity2:MedicalEntity{name:'Respiratory Medicine'})CREATE(entity1)-[:RELATIONSHIP_TYPE]->(entity2)
[0102] Medical entity data has been imported into the software (or database) used to build the knowledge graph. Each entity has an attribute called "name". To create a relationship between entities, use the MATCH statement to find the entity to be associated, and use the name attribute of matching entity 1 and entity 2 to get the node. Then, use the CREATE statement to create a relationship between the two nodes to build a medical knowledge graph.
[0103] S103, labeling the disease entity according to the target label.
[0104] In an embodiment of the present application, after the information processing device determines the target label corresponding to the disease entity in the medical knowledge graph, it can label the disease entity according to the target label.
[0105] In the embodiment of the present application, an exemplary information processing device is as follows: Figure 4As shown: the information processing device obtains initial medical information in a medical database (open source data set or web); uses preset data standard processing rules (ETL) to standardize the initial medical information and obtain a medical data set (GraphDB); matches related data in the medical data set to obtain the association relationship between the data; establishes a medical knowledge graph based on the association relationship and the data in the medical data set. The information processing device obtains multiple medical disease metadata; performs vector conversion processing (metadata vectorization) on multiple medical disease metadata to obtain multiple disease index vectors (LLM-Embedding); adds multiple disease index vectors to the disease database (vectorDB or faiss / chroma). When receiving a medical labeling instruction, the information processing device obtains a disease entity according to the medical labeling instruction (LLM-disease entity extraction); performs vector conversion on the disease entity to obtain a disease entity vector; sequentially determines the similarity between the disease entity vector and each of a plurality of disease index vectors (entity similarity matching) to obtain a plurality of similarities; obtains a preset number of target index vectors from a plurality of disease index vectors according to the plurality of similarities; determines a target disease index according to the target index vector; determines a target label that matches the target disease index in the medical knowledge graph; and labels the disease entity according to the target label (LLM-Chat index labeling, entity-Index similarity matching).
[0106] It can be understood that the information processing device obtains the disease entity and uses the medical knowledge graph to determine the target label corresponding to the disease entity. That is, the target label corresponding to the disease entity can be found in the medical knowledge graph, and the disease entity is labeled according to the target label, so that the labeling process will not be affected by manual level issues, the number of training samples, and the computing power of the equipment, thereby improving the accuracy of labeling.
[0107] Based on the same inventive concept as the above-mentioned information processing method, the embodiment of the present application provides an information processing device 1 corresponding to an information processing method; Figure 5 A schematic diagram of the structure of an information processing device provided in an embodiment of the present application Figure 1 , the information processing device 1 may include:
[0108] An acquisition unit 11 is used to acquire a disease entity according to the medical label labeling instruction when a medical label labeling instruction is received;
[0109] A determination unit 12, used to determine a target label corresponding to the disease entity in the medical knowledge graph;
[0110] The labeling unit 13 is used to label the disease entity according to the target label.
[0111] In some embodiments of the present application, the device further comprises a matching unit;
[0112] The matching unit is used to match the disease entity with a plurality of disease index vectors in the disease database to obtain a target disease index;
[0113] The determination unit 12 is used to determine a target tag matching the target disease index in the medical knowledge graph.
[0114] In some embodiments of the present application, the device further comprises a conversion unit;
[0115] The conversion unit is used to perform vector conversion on the disease entity to obtain a disease entity vector;
[0116] The determining unit 12 is used to sequentially determine the similarity between the disease entity vector and each of the multiple disease index vectors to obtain multiple similarities; determine the target disease index according to the target index vector;
[0117] The acquisition unit 11 is used to acquire a preset number of target index vectors from the multiple disease index vectors according to the multiple similarities.
[0118] In some embodiments of the present application, the device further includes a processing unit and an adding unit;
[0119] The acquisition unit 11 is used to acquire multiple medical disease metadata;
[0120] The processing unit is used to perform vector conversion processing on the plurality of medical disease metadata to obtain the plurality of disease index vectors;
[0121] The adding unit is used to add the multiple disease index vectors to the disease database.
[0122] In some embodiments of the present application, the apparatus further comprises an establishing unit;
[0123] The acquisition unit 11 is used to acquire initial medical information from a medical database;
[0124] The processing unit is used to perform standard processing on the initial medical information using preset data standard processing rules to obtain a medical data set;
[0125] The establishing unit is used to establish the medical association relationship between the data in the medical data set to obtain the medical knowledge graph.
[0126] In some embodiments of the present application, the matching unit is used to match the associated data in the medical data set to obtain the association relationship between the data;
[0127] The establishing unit is used to establish the medical knowledge graph based on the association relationship and the data in the medical data set.
[0128] In some embodiments of the present application, the device further comprises an extraction unit;
[0129] The acquisition unit 11 is used to acquire the medical file to be annotated from the medical label annotation instruction;
[0130] The extraction unit is used to extract the disease entity from the medical document to be annotated by using a language learning model.
[0131] It should be noted that, in actual applications, the above-mentioned acquisition unit 11, determination unit 12 and labeling unit 13 can be implemented by a processor 14 on the information processing device 1, specifically a CPU (Central Processing Unit), an MPU (Microprocessor Unit), a DSP (Digital Signal Processing) or a field programmable gate array (FPGA); the above-mentioned data storage can be implemented by a memory 15 on the information processing device 1.
[0132] The present application also provides an information processing device 1, such as Figure 6 As shown, the information processing device 1 includes: a processor 14, a memory 15 and a communication bus 16. The memory 15 communicates with the processor 14 via the communication bus 16. The memory 15 stores a program executable by the processor 14. When the program is executed, the information processing method described above is executed by the processor 14.
[0133] In practical applications, the memory 15 may be a volatile memory, such as a random access memory (RAM); or a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk (HDD) or a solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 14.
[0134] An embodiment of the present application provides a computer-readable storage medium having a computer program thereon, and when the program is executed by the processor 14, the information processing method as described above is implemented.
[0135] It can be understood that the information processing device obtains the disease entity and uses the medical knowledge graph to determine the target label corresponding to the disease entity. That is, the target label corresponding to the disease entity can be found in the medical knowledge graph, and the disease entity is labeled according to the target label, so that the labeling process will not be affected by manual level issues, the number of training samples, and the computing power of the equipment, thereby improving the accuracy of labeling.
[0136] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.
[0137] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0138] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0139] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0140] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application.
Claims
1. An information processing method, characterized in that: The method comprises: When receiving a medical label marking instruction, obtaining a disease entity according to the medical label marking instruction; In the medical knowledge graph, determine the target label corresponding to the disease entity; The disease entity is labeled according to the target label.
2. The method according to claim 1, characterized in that Determining the target label corresponding to the disease entity in the medical knowledge graph includes: Matching the disease entity with multiple disease index vectors in a disease database to obtain a target disease index; In the medical knowledge graph, a target label matching the target disease index is determined.
3. The method according to claim 2, characterized in that The matching of the disease entity with a plurality of disease index vectors in a disease database to obtain a target disease index includes: Performing vector transformation on the disease entity to obtain a disease entity vector; Determining the similarity between the disease entity vector and each of the multiple disease index vectors in sequence to obtain multiple similarities; According to the multiple similarities, obtaining a preset number of target index vectors from the multiple disease index vectors; The target disease index is determined according to the target index vector.
4. The method according to claim 2, characterized in that: Before matching the disease entity with a plurality of disease index vectors in a disease database to obtain a target disease index, the method further includes: Get metadata for multiple medical diseases; Performing vector conversion processing on the plurality of medical disease metadata to obtain the plurality of disease index vectors; The plurality of disease index vectors are added to the disease database.
5. The method according to claim 2, characterized in that: Before determining the target tag matching the target disease index in the medical knowledge graph, the method further includes: Obtaining initial medical information in a medical database; Using preset data standard processing rules, the initial medical information is standardized and processed to obtain a medical data set; A medical association relationship is established between the data in the medical data set to obtain the medical knowledge graph.
6. The method according to claim 5, characterized in that The step of establishing the medical association relationship between the data in the medical data set to obtain the medical knowledge graph includes: Matching the associated data in the medical data set to obtain the association relationship between the data; The medical knowledge graph is established based on the association relationship and the data in the medical dataset.
7. The method according to claim 1, characterized in that The step of obtaining the disease entity according to the medical labeling instruction includes: Acquire the medical file to be annotated from the medical label annotation instruction; The disease entity is extracted from the medical document to be annotated using a language learning model.
8. An information processing device, characterized in that: The device comprises: An acquisition unit, configured to acquire a disease entity according to a medical labeling instruction when a medical labeling instruction is received; A determination unit, used to determine a target label corresponding to the disease entity in the medical knowledge graph; A labeling unit is used to label the disease entity according to the target label.
9. An information processing device, characterized in that: The device comprises: A memory, a processor and a communication bus, wherein the memory communicates with the processor via the communication bus, the memory stores an information processing program executable by the processor, and when the information processing program is executed, the method according to any one of claims 1 to 7 is performed by the processor.
10. A storage medium having a computer program stored thereon, applied to an information processing device, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.