Medical business knowledge graph generation and question and answer processing method, device and equipment

By generating a medical business knowledge graph, the problems of low efficiency and inconsistent data management of traditional medical business are solved, and more efficient and accurate medical business data management and intelligent question-and-answer processing are achieved.

CN120072333APending Publication Date: 2025-05-30BEIJING UNITED FAMILY HOSPITAL CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510071441.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional medical business information management methods are inefficient, and due to heterogeneous data sources, information is inconsistent and inaccurate, affecting the accuracy of intelligent retrieval and intelligent question-and-answer.

Method used

By obtaining medical business texts from multiple different sources, identifying business entities and extracting entity description information and entity relationship information, deduplication processing and merging, generating a medical business knowledge graph, and processing medical business Q&A based on this graph.

Benefits of technology

It improves the efficiency and accuracy of medical business data management, enhances the accuracy of intelligent retrieval and intelligent question-and-answer, and uses structured information of the knowledge graph to improve the accuracy of querying complex relationships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072333A_ABST
    Figure CN120072333A_ABST
Patent Text Reader

Abstract

The invention relates to a medical business knowledge graph generation and question and answer processing method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring a plurality of medical business texts from different sources; business entities are identified from the medical business texts, and entity description information and entity relation information corresponding to the business entities are extracted according to the identified business entities; performing de-duplication processing on the plurality of service entities identified in the plurality of different medical service texts to generate new service entities and new entity description information corresponding to the new service entities; and merging the entity relationship information according to the new business entity and the new entity description information to obtain new entity relationship information, and generating a medical business knowledge graph according to the new business entity, the new entity description information and the new entity relationship information. By adopting the method, the medical business data management efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a method for generating a medical business knowledge graph, a method for processing medical business question and answer, an apparatus, a computer device, and a storage medium. Background Art

[0002] With the development of artificial intelligence technology, how to effectively utilize and manage information resources using artificial intelligence technology has become an important challenge faced by enterprises and research institutions. Especially in the medical field, hospital business knowledge is complex and scattered, lacking a unified knowledge system.

[0003] Traditional medical business information management methods mainly rely on keyword retrieval. When faced with a large amount of data, this method is inefficient and prone to missing relevant information. In addition, information such as hospital business and doctor introductions is maintained by different departments or individuals. Due to reasons such as maintenance personnel and maintenance timeliness, the same content may have problems such as conflicts and omissions, thus affecting the authenticity and integrity of medical business-related information, resulting in low efficiency of medical business information management, and further leading to low accuracy of feedback results such as intelligent retrieval or intelligent question and answer processing related to medical business data. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a method for generating a medical business knowledge graph that can improve the efficiency of medical business data management, a method for processing medical business question and answer that can improve the accuracy of intelligent question and answer processing results related to medical business, as well as its apparatus, computer device, and storage medium.

[0005] In a first aspect, a method for generating a medical business knowledge graph is provided, and the method includes:

[0006] Obtain medical business texts from multiple different sources;

[0007] Identify business entities from each medical business text, and extract entity description information and entity relationship information corresponding to each business entity according to the identified business entities;

[0008] Perform deduplication processing on the multiple business entities identified in multiple different medical business texts to generate new business entities and new entity description information corresponding to the new business entities; and

[0009] Merge the entity relationship information according to the new business entities and the new entity description information to obtain new entity relationship information, and generate a medical business knowledge graph according to the new business entities, the new entity description information, and the new entity relationship information.

[0010] In some embodiments, duplicate removal processing is performed on multiple business entities identified in multiple different medical business texts to generate new business entities and new entity description information corresponding to the new business entities, including:

[0011] Select a target business entity from multiple business entities, and calculate the similarity between other business entities and the target business entity among the multiple business entities;

[0012] Use the business entities with similarity greater than a preset threshold among other business entities as similar entities of the target business entity;

[0013] Determine the confidence weights of each similar entity; and

[0014] Determine duplicate entities from multiple similar entities, and generate new business entities and new entity description information corresponding to the new business entities according to the confidence weights of the duplicate entities and the confidence weights of the target business entity.

[0015] In some embodiments, calculating the similarity between other business entities and the target business entity among multiple business entities includes:

[0016] Input the target business entity and its corresponding entity description information, and each other business entity and its corresponding entity description information into a trained text vectorization model to obtain the target text vector corresponding to the target business entity output by the text vectorization model, and other text vectors corresponding to each other business entity; and

[0017] Perform vector inner product calculation on the target text vector and each other text vector respectively to obtain the similarity between each other business entity and the target business entity.

[0018] In some embodiments, the method further includes:

[0019] Fine-tune and train the text vectorization model using the following loss function:

[0020]

[0021] Among them, f(x) is used to represent the embedding features of the target sample, f(x + ) is used to represent the embedding features of the positive sample, f(x i ) is used to represent the embedding features of the negative sample, and τ represents the temperature parameter, which is used to control the distribution of similarity.

[0022] In some embodiments, determining the confidence weights of each similar entity includes:

[0023] Determine the medical business text from which each similar entity comes;

[0024] Determine the meta-information of the source of the characterization information of each similar entity according to the medical business text from which each similar entity comes; and

[0025] Determine the confidence weights of each similar entity according to the weight setting rules and the meta-information of each similar entity.

[0026] In some embodiments, determine duplicate entities from multiple similar entities, and generate a new business entity and corresponding new entity description information for the new business entity according to the confidence weights of the duplicate entities and the confidence weight of the target business entity, including:[[]]

[0027] Adopt at least one of a large language model, an entity duplicate classification model based on deep learning, and an entity description generation model;

[0028] Among them, the method of adopting a large language model includes: constructing duplicate entity judgment and entity summary prompt words, so that the large language model performs the following steps according to the duplicate entity judgment and entity summary prompt words:

[0029] Analyze the entity description information corresponding to the target business entity, and determine the entity type and key information of the target business entity;

[0030] Match duplicate entities from multiple similar entities according to the entity type and key information, generate a general name according to the name of the target business entity and the name of the duplicate entity, use the general name as the name of the new business entity, and retain the entity type of the target entity; and

[0031] Integrate the entity description information of the target business entity and the entity description information of the duplicate entity, and generate new entity description information.

[0032] In a second aspect, a medical business question and answer processing method is provided, and the method includes:

[0033] Receive question consultation information related to medical business;

[0034] Match corresponding relevant business entities from the medical business knowledge graph according to the question consultation information; wherein, the medical business knowledge graph is generated according to the medical business knowledge graph generation method of any one of the first aspects;

[0035] Obtain entity context information corresponding to the relevant business entity according to the relevant business entity; wherein, the entity context information includes entity description information, medical business text, entity relationship information, and entity description information of the first-level relationship entity of the relevant business entity;

[0036] Generate a large language model answer prompt word according to the question consultation information, the relevant business entity, and the entity context information, and input the large language model answer prompt word into the answer generation large language model to generate answer information corresponding to the question consultation information.

[0037] In a third aspect, a medical business Q&A processing device is provided, which includes:

[0038] A text acquisition module that acquires medical business texts from multiple different sources;

[0039] A text processing module that respectively identifies business entities from each medical business text, and extracts entity description information and entity relationship information corresponding to each business entity according to the identified business entities; performs deduplication processing on the multiple business entities identified in the medical business texts from multiple different sources to generate new business entities and new entity description information corresponding to the new business entities, and merges the entity relationship information according to the new business entities and the new entity description information to obtain new entity relationship information; and

[0040] A knowledge graph generation module that generates a medical business knowledge graph according to the new business entities, the new entity description information, and the new entity relationship information;

[0041] A question receiving module that receives question consultation information related to medical business;

[0042] An entity matching module that matches corresponding relevant business entities from the medical business knowledge graph according to the question consultation information;

[0043] A context retrieval module that is used to obtain entity context information corresponding to the relevant business entities according to the relevant business entities; wherein, the entity context information includes entity description information, medical business texts, entity relationship information, and entity description information of the first-level relationship entities of the relevant business entities; and

[0044] An answer generation module that is used to generate a large language model answer prompt according to the question consultation information, the relevant business entities, and the entity context information, and input the large language model answer prompt into the answer generation large language model to generate answer information corresponding to the question consultation information.

[0045] In a fourth aspect, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the steps of any one of the methods in the first aspect and the second aspect are implemented.

[0046] In a fifth aspect, a computer-readable storage medium has a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of any one of the methods in the first aspect and the second aspect are implemented.

[0047] The above-mentioned method for generating a medical business knowledge graph extracts business entities from heterogeneous medical business texts with different sources and performs deduplication processing, thereby constructing a medical business knowledge graph. Since the business entities extracted from heterogeneous medical business texts with different sources are deduplicated and integrated to generate new business entities with unified expressions, new entity description information, and new entity relationship information, the adoption of this solution can solve the problems of inconsistent and inaccurate medical business data caused by different data sources, thereby improving the efficiency and accuracy of medical data management. The above-mentioned medical business question-answering processing method, device, computer device, and storage medium, when receiving a user's question consultation, call the medical business knowledge graph constructed by the method according to the above-mentioned embodiments, and retrieve relevant business entities and their corresponding entity context information that match the question consultation information based on the medical business knowledge graph, thereby being able to generate prompt words more efficiently and accurately, and further improving the accuracy of the answers returned by the large language model for answer generation. In addition, since the medical business knowledge graph integrates medical business data from different sources, the retrieval and generation capabilities of the model are enhanced, and the structured information of the knowledge graph can be used to improve the accuracy of queries for multi-level and complex relationships. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 FIG. is an application environment diagram of the medical business knowledge graph generation method and the medical business question-answering processing method in some embodiments;

[0049] Figure 2 FIG. is a schematic flowchart of the medical business knowledge graph generation method in some embodiments;

[0050] Figure 3 FIG. is a schematic flowchart of the medical business question-answering processing method in some embodiments;

[0051] Figure 4 FIG. is a structural block diagram of the medical business question-answering processing device in some embodiments;

[0052] Figure 5 FIG. is an internal structure diagram of a computer device in some embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0054] The medical business knowledge graph generation method and the medical business question-answering processing method provided by the present application can be applied to an application environment as shown in Figure 1 FIG. Among them, the terminal 102 communicates with the server 104 through the network.

[0055] Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0056] In some embodiments, in the medical business knowledge graph generation stage:

[0057] The server 104 obtains medical business texts from multiple different sources; respectively identifies business entities from each medical business text, and extracts entity description information and entity relationship information corresponding to each business entity according to the identified business entities; performs deduplication processing on the multiple business entities identified in the multiple different medical business texts to generate new business entities and new entity description information corresponding to the new business entities, and merges the entity relationship information according to the new business entities and the new entity description information to obtain new entity relationship information; and generates a medical business knowledge graph according to the new business entities, the new entity description information, and the new entity relationship information.

[0058] In the medical business question-answering processing stage:

[0059] The server 104 receives question consultation information related to medical business initiated by the terminal 102, matches corresponding relevant business entities from the medical business knowledge graph according to the question consultation information, obtains entity context information corresponding to the relevant business entities according to the relevant business entities, where the entity context information includes entity description information, medical business texts, entity relationship information, and entity description information of the first-level relationship entities of the relevant business entities, generates large language model answer prompts according to the question consultation information, the relevant business entities, and the entity context information, and inputs the large language model answer prompts into an answer generation large language model to generate answer information corresponding to the question consultation information. Further, the answer information can be returned to the terminal 102.

[0060] In some embodiments, as Figure 2 shown, a medical business knowledge graph generation method is provided. Taking the method applied to the Figure 1 server as an example for illustration, it may include the following steps:

[0061] Step S202: Obtain medical business texts from multiple different sources.

[0062] Specifically, the server can collect introduction document data, picture data, etc. of departments, doctors, hospital areas, and medical items related to medical operations from various data sources such as various systems within the hospital and / or department systems through APIs (Application Programming Interfaces), file reading, etc. as the original medical operation data.

[0063] Furthermore, the server can use natural language processing technology, document analysis technology, OCR (Optical Character Recognition) technology to preprocess the original medical operation data such as documents and pictures obtained from different data sources, so as to obtain text content related to medical operations such as departments, doctors, hospital areas, and medical items as medical operation texts.

[0064] Step S204: Identify business entities from each medical operation text respectively, and extract entity description information and entity relationship information corresponding to each business entity according to the identified business entities.

[0065] Specifically, the server can identify business entities from each medical operation text respectively, where the business entities can include but are not limited to department names, doctor names, diagnosis and treatment item names, etc. The extraction methods of business entities, entity description information, and entity relationship information can be to complete the extraction task by constructing business entity extraction prompt words and using a large language model, or other extraction methods such as those based on deep learning, statistics, or rules for business entity names, entity description information, and entity relationship information.

[0066] Step S206: Perform deduplication processing on multiple business entities identified in multiple different medical operation texts to generate new business entities and new entity description information corresponding to the new business entities.

[0067] Specifically, for medical operation texts with different sources containing heterogeneous data, there are problems of data inconsistency such as the same business entity name but different actual contents, and / or different business entity names but the same actual contents. Therefore, the server can verify and deduplicate the extracted business entities, so as to integrate multiple business entities to generate new business entities and new entity description information corresponding to the new business entities.

[0068] Exemplarily, similar business entities can be screened from multiple business entities for merging and deduplication processing. For example, but not limited to, a text vectorization model based on pre-training or fine-tuning training can be used to screen similar business entities, or the BM25 algorithm based on word frequency statistics or other algorithms can be used.

[0069] Step S208: Merge the entity relationship information according to the new business entity and the new entity description information to obtain new entity relationship information, and generate a medical business knowledge graph according to the new business entity, the new entity description information, and the new entity relationship information.

[0070] Specifically, the server can merge the original entity relationship information according to the new business entity and the new entity description information regenerated after deduplication. Since multiple relationships (connection relationships) may exist between two entities due to the merger of business entities after deduplication, the original entity relationship information can be merged to obtain a new entity connection relationship that conforms to the connection logic between the new business entities. Then, a medical business knowledge graph can be constructed with the new business entity as the node, the new entity description information as the node information, and the new entity relationship information as the edge.

[0071] Exemplarily, the entity relationship summary prompt can be constructed, and the large language model can complete the summary of the entity relationship information to obtain the new entity relationship information. In other examples, other deep learning models can also complete the summary and merger processing of the entity relationship information and generate the new entity relationship information, etc.

[0072] Furthermore, the constructed medical business knowledge graph can be stored in the graph database for invocation in response to the request instructions of intelligent retrieval or intelligent question-answering processing.

[0073] In some embodiments, generating a medical business knowledge graph according to the new business entity, the new entity description information, and the new entity relationship information includes: performing vectorization processing on the new business entity and the new entity description information based on a text vectorization model to obtain the new business entity vector and its corresponding new entity description vector; and generating a medical business knowledge graph according to the new business entity vector, the new entity description vector, and the new entity relationship information. In this embodiment, by first performing vectorization processing on the new business entity and the new entity description information and then constructing the medical business knowledge graph, the efficiency and accuracy of constructing the medical business knowledge graph can be improved, the volume of the medical business knowledge graph can be compressed, and further, the efficiency of invoking the medical business knowledge graph during intelligent retrieval or intelligent question-answering processing and the accuracy of result feedback can be improved.

[0074] The above-mentioned medical business knowledge graph generation method extracts business entities from heterogeneous medical business texts with different sources and performs deduplication processing, thereby constructing a medical business knowledge graph. Since the business entities extracted from heterogeneous medical business texts with different sources are deduplicated and integrated to generate new business entities, new entity description information, and new entity relationship information, the adoption of this solution can solve the problems of inconsistent and inaccurate medical business data caused by different data sources, thereby improving the efficiency and accuracy of medical data management.

[0075] In some embodiments, the deduplication processing of multiple business entities identified in multiple different medical business texts to generate new business entities and new entity description information corresponding to the new business entities includes: selecting a target business entity from the multiple business entities, and calculating the similarity between other business entities in the multiple business entities and the target business entity; taking the business entities with similarity greater than a preset threshold among the other business entities as similar entities of the target business entity; determining the confidence weights of the respective similar entities; and determining duplicate entities from the multiple similar entities, and generating new business entities and new entity description information corresponding to the new business entities according to the confidence weights of the duplicate entities and the confidence weights of the target business entity.

[0076] In this embodiment, after selecting a business entity as the target business entity, the similarity between other business entities and the selected target business entity can be calculated, and then the entities with similarity higher than the preset threshold can be taken as the similar entities of the target business entity. Among them, the screening of similar entities can be performed using a text vectorization model that has been pre-trained or pre-fine-tuned, or using the BM25 algorithm or other algorithms based on word frequency statistics. After the screening of similar entities is completed, a list of similar entities can be generated. Further, confidence weights can also be calculated for each similar entity. The confidence weight can be used to represent the proportion of the entity description information corresponding to it in the summary processing of the description information after entity deduplication. Finally, duplicate entities can be determined from the list of similar entities, and then according to the confidence weights of the duplicate entities and the confidence weights of the target business entity, the entity name and entity description information of the new business entity can be summarized again. In this embodiment, by screening similar entities and configuring execution weights, the accuracy and confidence of the generated new business entities and new entity description information can be improved.

[0077] In some embodiments, calculating the similarity between other business entities and a target business entity among multiple business entities includes: inputting the target business entity and its corresponding entity description information, as well as each of the other business entities and their corresponding entity description information into a trained text vectorization model to obtain the target text vector corresponding to the target business entity output by the text vectorization model, and the other text vectors corresponding to each of the other business entities; and performing vector inner product calculations between the target text vector and each of the other text vectors to obtain the similarity between each of the other business entities and the target business entity.

[0078] In this embodiment, a trained text vectorization model can be used to calculate the similarity between other business entities and a target business entity among multiple business entities, thereby improving the accuracy of similarity calculation and further improving the accuracy of knowledge graph generation.

[0079] More specifically, the text vectorization model can be trained in the following way in advance:

[0080] First, training data and validation data can be prepared as sample data. Among them, the sample data can include a dataset of labeled positive samples and negative samples. Among them, a positive sample is a sample pair composed of a query and a document semantically related or similar to the query, while a negative sample is a sample pair composed of a query and a document semantically unrelated to the query. Vertical domain data of the medical business can be used for model training to better capture semantic similarities in the professional field. The input of the text vectorization model is a piece of text, which can be the query or the document in the positive sample or negative sample, and the output is a corresponding text vector. Specifically, it can be referred to as follows:

[0081] f(Q)∈R1×kf(Q)∈R1×k

[0082] f(D)∈R1×kf(D)∈R1×k

[0083] Among them, Q is the query text, D is the document text, k is the dimension of the vector output by the model, and f is the vector representation function of the model.

[0084] After obtaining the vectorized representations of the query and the document, the text similarity between the query and the document can be calculated through vector inner product. Specifically, it can be referred to as follows:

[0085] S=sim(f(Q),f(D))

[0086] Among them, sim is the similarity calculation function. Alternatively, it can also be the dot product (inner product) or cosine similarity, etc.

[0087] In some embodiments, in order to further improve the accuracy of similarity calculation of the text vectorization model, the text vectorization model can be fine-tuned. The goal of fine-tuning the text vectorization model is to increase the text similarity of positive samples and decrease the text similarity of negative samples.

[0088] More specifically, the following loss function can be used to fine-tune and train the text vectorization model:

[0089]

[0090] where f(x) is used to represent the embedding features of the target sample, f(x + ) is used to represent the embedding features of positive samples, f(x i ) is used to represent the embedding features of negative samples, and τ represents the temperature parameter, which is used to control the distribution of similarity. Exemplarily, when fine-tuning the model, the Adam W optimizer can be used, with a relatively small learning rate, such as 5e-6, and the CosineLRScheduler.

[0091] In some embodiments, determining the confidence weights of each similar entity includes: determining the medical business text from which each similar entity comes; determining the meta-information of the source of the representation information of each similar entity according to the medical business text from which each similar entity comes; and determining the confidence weights of each similar entity according to the weight setting rules and the meta-information of each similar entity.

[0092] In this embodiment, the confidence weights can be assigned by analyzing the meta-information of the medical business text from which each similar entity comes, and the corresponding weight setting rules can be configured according to the requirements, so as to improve the accuracy of weight setting, and further improve the accuracy and rationality of the generation of new business entities and new description information.

[0093] Exemplarily, the meta-information can include but is not limited to the release time, release method, etc.; the weight setting rules can include but are not limited to: the confidence weights of business entities extracted from documents with a newer release time and / or a higher release formality are higher; therefore, the confidence weights of business entities extracted from newer documents are higher than those of business entities extracted from older documents; the confidence weights of business entities extracted from business knowledge documents are higher than those of business entities extracted from shift handover record documents.

[0094] In some embodiments, determining duplicate entities from multiple similar entities and generating new business entities and new entity description information corresponding to the new business entities according to the confidence weights of the duplicate entities and the confidence weights of the target business entities includes: using at least one of a large language model, an entity duplicate classification model based on deep learning, and an entity description generation model.

[0095] In some embodiments, a deep learning-based entity duplication classification model can be adopted. The training method of the entity duplication classification model can refer to the following:

[0096] The key of the method based on the entity duplication classification model lies in the implementation of the classification task. The BERT (Bidirectional Encoder Representations from Transformers) model can be used as the base model. By fine-tuning the BERT model and using the cross-entropy loss function, the performance of the model in the entity duplication classification task can be effectively improved, so as to obtain an entity duplication classification model that can more accurately determine duplicate entities from multiple similar entities.

[0097] Exemplarily, the BERT model used as the base model can include but is not limited to the bert-base-chinese or bert-large-uncased model.

[0098] Exemplarily, the following method can be used to train the entity duplication classification model:

[0099] First, training data and validation data can be prepared as sample data. The sample data can contain a labeled entity duplication classification data set, which usually includes positive samples and negative samples. Among them, positive samples refer to texts that are duplicate entities, and negative samples refer to texts that are non-duplicate entities. Sample data can be generated using data in the vertical field of the medical business for model training, so as to better capture the classification requirements of the professional field.

[0100] Hereinafter, taking the BERT model as the base model as an example for illustration, the input of the BERT model is usually a text sequence, which can be a single text or a text pair (such as a query and a document). When inputting, the Tokenizer of the BERT model converts the input text into a token sequence that the BERT model can understand. The token sequence is input into the BERT model again to obtain the vector representation corresponding to each token output by the BERT model. In the text classification task (classifying similar entities into two categories: duplicate entities and non-duplicate entities), when the entity duplication classification model based on the BERT model outputs, the output of the [CLS] token output by the BERT model can be used as the classification result, so as to accurately determine which or which similar entities are duplicate entities.

[0101] In some embodiments, in order to further improve the classification accuracy of the entity duplication classification model, the entity duplication classification model can be fine-tuned. The goal of fine-tuning is to increase the output value of positive samples and decrease the output value of negative samples.

[0102] Exemplarily, Binary Cross-Entropy Loss can be adopted during the fine-tuning process:

[0103] L = -[y·log(p) + (1 - y)·log(1 - p)];

[0104] where y is the one-hot encoding of the true label; p is the class probability predicted by the model. When fine-tuning the model, the Adam W optimizer can usually be used, with a relatively small learning rate, such as 5e-6, and the CosineLRScheduler.

[0105] In some embodiments, the method of the large language model can be adopted to determine duplicate entities from multiple said similar entities, and generate new business entities and the corresponding new entity description information of the new business entities according to the confidence weights of the duplicate entities and the confidence weights of the target business entities.

[0106] In this embodiment, duplicate entity judgment and entity summary prompts can be constructed to enable the large language model to perform the following steps according to the duplicate entity judgment and entity summary prompts:

[0107] Analyze the entity description information corresponding to the target business entity, and determine the entity type and key information of the target business entity;

[0108] Match duplicate entities from multiple similar entities according to the entity type and key information, generate a general name based on the name of the target business entity and the name of the duplicate entity, use the general name as the name of the new business entity, and retain the entity type of the target entity; and

[0109] Integrate the entity description information of the target business entity and the entity description information of the duplicate entity, and generate new entity description information.

[0110] In this embodiment, by constructing duplicate entity judgment and entity summary prompts, the large language model can complete the judgment of duplicate entities and the renaming of new business entities and the summary of new entity description information. The effect of duplicate entity and new business entity summary can be improved by adding Chain of Thought prompts including Think Protocol and few shot examples. Thus, the large language model can implement the steps of the above Work flow, improving the accuracy of duplicate entity judgment and the accuracy of integrating entity description information.

[0111] In some embodiments, the present application also provides a medical business question-answering processing method. Referring to Figure 3 as shown, this method may include:

[0112] Step S302: Receive question consultation information related to medical business.

[0113] Specifically, the server can receive the question consultation information initiated by the user through the terminal in various ways, such as web pages, mobile applications, chatbots, etc.

[0114] Step S304: Match the corresponding relevant business entities from the medical business knowledge graph according to the question consultation information.

[0115] The medical business knowledge graph is generated according to the method of any one or more of the above embodiments, and will not be elaborated here.

[0116] Specifically, the question consultation information can be matched with the entity description information of the business entities in the medical business knowledge graph to obtain the relevant business entities. More specifically, the text vector corresponding to the question consultation information can be calculated first according to the question consultation information, and then the semantic similarity comparison is made between the text vector and the text vector corresponding to the entity description information in the knowledge graph, and the business entity with the highest similarity is selected as the relevant business entity.

[0117] Step S306: Obtain the entity context information corresponding to the relevant business entity according to the relevant business entity. The entity context information includes entity description information, medical business text, entity relationship information, and entity description information of the first-level relationship entities of the relevant business entity.

[0118] Specifically, based on the at least one list of relevant business entities screened out, the server integrates the user's question consultation information and the entity context information, and then generates a prompt word. More specifically, the entity description information of the relevant business entity determined from the knowledge graph, the medical business text from which it comes, and the entity description information of each business entity within one hop of the entity relationship can be obtained.

[0119] Step S308: Generate a large language model answer prompt word according to the question consultation information, the relevant business entity, and the entity context information, and input the large language model answer prompt word into the answer generation large language model to generate an answer information corresponding to the question consultation information.

[0120] In the above medical business Q&A processing method, when receiving the user's question consultation, by invoking the medical business knowledge graph constructed according to the method of the above embodiments, and retrieving the relevant business entities and their corresponding entity context information that match the question consultation information based on the medical business knowledge graph, it is possible to generate prompt words more efficiently and accurately, thereby improving the accuracy of the answers returned by the answer generation large language model. In addition, since different sources of medical business data are integrated in the medical business knowledge graph, the retrieval and generation capabilities of the model are enhanced, and the structured information of the knowledge graph can be used to improve the accuracy of queries for multi-level and complex relationships.

[0121] It should be understood that although Figures 2 to 3 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figures 2 to 3 at least a part of the steps in

[0122] In some embodiments, as Figure 4 shown, a medical business Q&A processing device is provided, including: a text acquisition module 410, a text processing module 420, a knowledge graph generation module 430, a question reception module 440, an entity matching module 450, a context retrieval module 460, and an answer generation module 470; where:

[0123] The text acquisition module 410 acquires medical business texts from multiple different sources;

[0124] The text processing module 420 respectively identifies business entities from each medical business text, and extracts entity description information and entity relationship information corresponding to each business entity according to the identified business entities; performs deduplication processing on the multiple business entities identified in the medical business texts from multiple different sources to generate new business entities and new entity description information corresponding to the new business entities, and merges the entity relationship information according to the new business entities and the new entity description information to obtain new entity relationship information; and

[0125] The knowledge graph generation module 430 generates a medical business knowledge graph according to the new business entities, the new entity description information, and the new entity relationship information;

[0126] The question reception module 440 receives question consultation information related to medical business;

[0127] The entity matching module 450 matches corresponding relevant business entities from the medical business knowledge graph according to the question consultation information;

[0128] The context retrieval module 460 is used to obtain entity context information corresponding to the relevant business entities according to the relevant business entities; where the entity context information includes entity description information, medical business texts, entity relationship information, and entity description information of the first-level relationship entities of the relevant business entities; and

[0129] The answer generation module 470 is configured to generate large language model answer prompts based on the question consultation information, relevant business entities, and entity context information, and input the large language model answer prompts into the answer generation large language model to generate answer information corresponding to the question consultation information.

[0130] In some embodiments, the text processing module 420 selects a target business entity from multiple business entities, calculates the similarity between other business entities in the multiple business entities and the target business entity; uses the business entities with similarity greater than a preset threshold among the other business entities as similar entities of the target business entity; determines the confidence weights of the respective similar entities; and determines duplicate entities from the multiple similar entities, and generates new business entities and new entity description information corresponding to the new business entities according to the confidence weights of the duplicate entities and the confidence weight of the target business entity.

[0131] In some embodiments, the text processing module 420 inputs the target business entity and its corresponding entity description information, and each other business entity and its corresponding entity description information into a trained text vectorization model, to obtain the target text vector corresponding to the target business entity output by the text vectorization model, and the other text vectors corresponding to the respective other business entities; and performs vector inner product calculations on the target text vector and each of the other text vectors respectively, to obtain the similarity between each of the other business entities and the target business entity.

[0132] In some embodiments, the text processing module 420 is further configured to fine-tune and train the text vectorization model using the following loss function:

[0133]

[0134] where f(x) is used to represent the embedding feature of the target sample, f(x + ) is used to represent the embedding feature of the positive sample, f(x i ) is used to represent the embedding feature of the negative sample, and τ represents the temperature parameter, which is used to control the distribution of the similarity.

[0135] In some embodiments, the text processing module 420 determines the medical business text from which each similar entity comes; determines the meta-information of the source of the characterization information of each similar entity according to the medical business text from which each similar entity comes; and determines the confidence weight of each similar entity according to the weight setting rule and the meta-information of each similar entity.

[0136] In some embodiments, the text processing module 420 performs processing using at least one of a large language model, an entity duplicate classification model based on deep learning, and an entity description generation model.

[0137] In some embodiments, the text processing module 420 constructs duplicate entity judgment and entity summary prompting words, so that the large language model performs the following steps according to the duplicate entity judgment and entity summary prompting words: analyzing the entity description information corresponding to the target business entity, and determining the entity type and key information of the target business entity; matching duplicate entities from multiple similar entities according to the entity type and key information, generating a general name based on the name of the target business entity and the name of the duplicate entity, using the general name as the name of the new business entity, and retaining the entity type of the target entity; and integrating the entity description information of the target business entity with the entity description information of the duplicate entity, and generating new entity description information.

[0138] For the specific limitations of the medical business Q&A processing device, reference can be made to the limitations of the medical business Q&A processing method in the foregoing text, which will not be elaborated here. Each module in the above medical business Q&A processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or independent of it, or stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0139] In some embodiments, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store medical business knowledge graph data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a medical business knowledge graph generation method or a medical business Q&A processing method.

[0140] Those skilled in the art can understand that Figure 5 the structure shown in

[0141] In some embodiments, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: obtaining medical service texts from multiple different sources; respectively identifying business entities from each medical service text, and extracting entity description information and entity relationship information corresponding to each business entity according to the identified business entities; performing deduplication processing on the multiple business entities identified in the multiple different medical service texts to generate new business entities and new entity description information corresponding to the new business entities; and merging the entity relationship information according to the new business entities and the new entity description information to obtain new entity relationship information, and generating a medical service knowledge graph according to the new business entities, the new entity description information, and the new entity relationship information.

[0142] In some embodiments, when the processor executes the computer program, the following steps are further implemented: selecting a target business entity from the multiple business entities, and calculating the similarity between other business entities in the multiple business entities and the target business entity; using the business entities with a similarity greater than a preset threshold among the other business entities as similar entities of the target business entity; determining the confidence weights of the respective similar entities; and determining duplicate entities from the multiple similar entities, and generating new business entities and new entity description information corresponding to the new business entities according to the confidence weights of the duplicate entities and the confidence weight of the target business entity.

[0143] In some embodiments, when the processor executes the computer program, the following steps are further implemented: inputting the target business entity and its corresponding entity description information, and each other business entity and its corresponding entity description information into a trained text vectorization model to obtain a target text vector corresponding to the target business entity output by the text vectorization model, and other text vectors corresponding to each other business entity; and performing vector inner product calculation on the target text vector and each other text vector respectively to obtain the similarity between each other business entity and the target business entity.

[0144] In some embodiments, when the processor executes the computer program, the following steps are further implemented: fine-tuning and training the text vectorization model by using the following loss function:

[0145]

[0146] wherein, f(x) is used to represent the embedding feature of the target sample, f(x + ) is used to represent the embedding feature of the positive sample, f(x i ) is used to represent the embedding feature of the negative sample, and τ represents a temperature parameter, which is used to control the distribution of the similarity.

[0147] In some embodiments, when the processor executes the computer program, the following steps are further implemented: determining the medical business texts from which each similar entity originates; determining the meta-information of the source of the representation information of each similar entity according to the medical business texts from which each similar entity originates; and determining the confidence weights of each similar entity according to the weight setting rules and the meta-information of each similar entity.

[0148] In some embodiments, when the processor executes the computer program, the following steps are further implemented: determining duplicate entities from multiple similar entities by using at least one of a large language model, a deep learning-based entity duplication classification model, and an entity description generation model, and generating a new business entity and new entity description information corresponding to the new business entity according to the confidence weights of the duplicate entities and the confidence weights of the target business entities.

[0149] In some embodiments, when the processor executes the computer program, the following steps are further implemented: constructing duplicate entity judgment and entity summary prompt words, so that the large language model performs the following steps according to the duplicate entity judgment and entity summary prompt words: analyzing the entity description information corresponding to the target business entity, and determining the entity type and key information of the target business entity; matching duplicate entities from multiple similar entities according to the entity type and key information, generating a general name according to the name of the target business entity and the name of the duplicate entity, using the general name as the name of the new business entity, and retaining the entity type of the target entity; and integrating the entity description information of the target business entity and the entity description information of the duplicate entity, and generating new entity description information.

[0150] In some embodiments, when the processor executes the computer program, the following steps are further implemented: receiving question consultation information related to medical business; matching corresponding relevant business entities from the medical business knowledge graph according to the question consultation information; obtaining entity context information corresponding to the relevant business entities according to the relevant business entities; wherein the entity context information includes entity description information, medical business texts, entity relationship information, and entity description information of the first-level relationship entities of the relevant business entities; generating a large language model answer prompt word according to the question consultation information, the relevant business entities, and the entity context information, and inputting the large language model answer prompt word into an answer generation large language model to generate answer information corresponding to the question consultation information.

[0151] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining medical business texts from multiple different sources; respectively identifying business entities from each medical business text, and extracting entity description information and entity relationship information corresponding to each business entity according to the identified business entities; performing deduplication processing on the multiple business entities identified in the multiple different medical business texts to generate new business entities and new entity description information corresponding to the new business entities; and merging the entity relationship information according to the new business entities and the new entity description information to obtain new entity relationship information, and generating a medical business knowledge graph according to the new business entities, the new entity description information, and the new entity relationship information.

[0152] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: selecting a target business entity from the multiple business entities, and calculating the similarity between the other business entities in the multiple business entities and the target business entity; using the business entities with a similarity greater than a preset threshold among the other business entities as similar entities of the target business entity; determining the confidence weights of the respective similar entities; and determining duplicate entities from the multiple similar entities, and generating new business entities and new entity description information corresponding to the new business entities according to the confidence weights of the duplicate entities and the confidence weight of the target business entity.

[0153] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: inputting the target business entity and its corresponding entity description information, and each other business entity and its corresponding entity description information into a trained text vectorization model to obtain a target text vector corresponding to the target business entity output by the text vectorization model, and other text vectors corresponding to each other business entity; and performing vector inner product calculation on the target text vector and each other text vector respectively to obtain the similarity between each other business entity and the target business entity.

[0154] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: fine-tuning and training the text vectorization model by using the following loss function:

[0155]

[0156] where f(x) is used to represent the embedding feature of the target sample, f(x + ) is used to represent the embedding feature of the positive sample, f(x i ) is used to represent the embedding feature of the negative sample, and τ represents a temperature parameter, which is used to control the distribution of the similarity.

[0157] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: determining the medical business texts from which each similar entity is derived; determining the meta-information of the source of the characterization information of each similar entity according to the medical business texts from which each similar entity is derived; and determining the confidence weights of each similar entity according to the weight setting rules and the meta-information of each similar entity.

[0158] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: determining duplicate entities from multiple similar entities by using at least one of a large language model, a deep learning-based entity duplication classification model, and an entity description generation model, and generating a new business entity and new entity description information corresponding to the new business entity according to the confidence weights of the duplicate entities and the confidence weights of the target business entities.

[0159] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: constructing duplicate entity judgment and entity summary prompt words, so that the large language model performs the following steps according to the duplicate entity judgment and entity summary prompt words: analyzing the entity description information corresponding to the target business entity, and determining the entity type and key information of the target business entity; matching duplicate entities from multiple similar entities according to the entity type and key information, generating a general name according to the name of the target business entity and the name of the duplicate entity, using the general name as the name of the new business entity, and retaining the entity type of the target entity; and integrating the entity description information of the target business entity with the entity description information of the duplicate entity, and generating new entity description information.

[0160] In some embodiments, when the computer program is executed by a processor, the following steps are further implemented: receiving problem consultation information related to medical services; matching corresponding relevant business entities from a medical service knowledge graph according to the problem consultation information; obtaining entity context information corresponding to the relevant business entities according to the relevant business entities; wherein the entity context information includes entity description information, medical service text, entity relationship information, and entity description information of the first-level relationship entities of the relevant business entities; generating large language model answer prompt words according to the problem consultation information, the relevant business entities, and the entity context information, and inputting the large language model answer prompt words into an answer generation large language model to generate answer information corresponding to the problem consultation information. Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0161] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0162] In addition, the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the characters in this article generally represent that the associated objects before and after are in an "or" relationship.

[0163] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

[0164] It should be noted that in the embodiments of the present application, for data related to user information or user data, etc., it is necessary to obtain and process it after obtaining the authorization and consent of the user. When the embodiments of the present application are applied to specific products or technologies, it is necessary to obtain the permission or consent of the user, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

Claims

1. A method for generating a medical business knowledge graph, the method comprising: Obtain medical business texts from multiple different sources; Respectively identifying business entities from each of the medical business texts, and extracting entity description information and entity relationship information corresponding to each of the business entities according to the identified business entities; Deduplication processing is performed on the multiple business entities identified in multiple different medical business texts to generate new business entities and new entity description information corresponding to the new business entities; as well as The entity relationship information is merged according to the new business entity and the new entity description information to obtain new entity relationship information, and a medical business knowledge graph is generated according to the new business entity, the new entity description information and the new entity relationship information.

2. The method according to claim 1, characterized in that The deduplication processing of the plurality of business entities identified in the plurality of different medical business texts to generate a new business entity and new entity description information corresponding to the new business entity includes: Selecting a target business entity from the plurality of business entities, and calculating similarities between other business entities in the plurality of business entities and the target business entity; Taking business entities whose similarity among the other business entities is greater than a preset threshold as similar entities to the target business entity; determining a confidence weight for each of the similar entities; and A repeated entity is determined from the multiple similar entities, and a new business entity and new entity description information corresponding to the new business entity are generated according to the confidence weight of the repeated entity and the confidence weight of the target business entity.

3. The method according to claim 2, characterized in that The calculating the similarity between other business entities in the plurality of business entities and the target business entity comprises: Inputting the target business entity and its corresponding entity description information, as well as each of the other business entities and its corresponding entity description information into a trained text vectorization model, to obtain a target text vector corresponding to the target business entity and other text vectors corresponding to each of the other business entities output by the text vectorization model; and The target text vector is respectively calculated with each of the other text vectors as a vector inner product, so as to obtain the similarity between each of the other business entities and the target business entity.

4. The method according to claim 3, characterized in that The method further comprises: The text vectorization model is fine-tuned using the following loss function: Among them, f(x) is used to represent the embedded features of the target sample, f(x + ) is used to represent the embedding features of positive samples, f(x i ) is used to represent the embedding features of negative samples, and τ represents the temperature parameter, which is used to control the distribution of similarity.

5. The method according to claim 2, characterized in that: The determining of the confidence weight of each of the similar entities comprises: Determine the medical business text from which each of the similar entities comes; Determining meta information representing the source of information of each of the similar entities according to the medical business text from which each of the similar entities comes; and The confidence weight of each of the similar entities is determined according to a weight setting rule and the meta information of each of the similar entities.

6. The method according to claim 2, characterized in that The determining of a repeated entity from the plurality of similar entities, and generating a new business entity and new entity description information corresponding to the new business entity according to the confidence weight of the repeated entity and the confidence weight of the target business entity, comprises: Using at least one of a large language model, a deep learning-based entity duplication classification model, and an entity description generation model; The method of using a large language model includes: constructing repeated entity judgments and entity summary prompt words, so that the large language model performs the following steps according to the repeated entity judgments and entity summary prompt words: Analyze the entity description information corresponding to the target business entity, and determine the entity type and key information of the target business entity; Matching a duplicate entity from a plurality of similar entities according to the entity type and the key information, generating a summary name according to the name of the target business entity and the name of the duplicate entity, using the summary name as the name of the new business entity, and retaining the entity type of the target entity; and The entity description information of the target business entity is integrated with the entity description information of the repeated entity to generate new entity description information.

7. A method for processing medical business questions and answers, the method comprising: Receive consultation information on medical business-related issues; Matching corresponding relevant business entities from the medical business knowledge graph according to the question consultation information; wherein the medical business knowledge graph is generated according to the medical business knowledge graph generation method according to any one of claims 1 to 7; Acquire entity context information corresponding to the related business entity according to the related business entity; wherein the entity context information includes entity description information, medical business text, entity relationship information, and entity description information of a first-level relationship entity of the related business entity; A large language model answer prompt word is generated according to the question consultation information, the relevant business entity and the entity context information, and the large language model answer prompt word is input into the answer to generate a large language model to generate answer information corresponding to the question consultation information.

8. A medical business question and answer processing device, characterized in that: The device comprises: The text acquisition module acquires medical business texts from multiple different sources; a text processing module, which identifies business entities from each of the medical business texts, and extracts entity description information and entity relationship information corresponding to each of the business entities according to the identified business entities; performs deduplication processing on multiple business entities identified in medical business texts from multiple different sources, generates new business entities and new entity description information corresponding to the new business entities, and merges the entity relationship information according to the new business entities and the new entity description information to obtain new entity relationship information; and A graph generation module, generating a medical business knowledge graph according to the new business entity, the new entity description information and the new entity relationship information; Question receiving module, receiving consultation information related to medical business; An entity matching module matches corresponding relevant business entities from the medical business knowledge graph according to the question consultation information; A context retrieval module, used to obtain entity context information corresponding to the related business entity according to the related business entity; wherein the entity context information includes entity description information, medical business text, entity relationship information and entity description information of the first-level relationship entity of the related business entity; and An answer generation module is used to generate a large language model answer prompt word based on the question consultation information, the relevant business entity and the entity context information, and input the large language model answer prompt word into the answer to generate a large language model to generate answer information corresponding to the question consultation information.

9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Text generation method and device based on industrial knowledge graph, equipment and medium

    CN121210675A

  • Decision-making knowledge graph-oriented entity relationship identification method and device

    CN122088506A