Diagnostic multiple-write detection method and apparatus, electronic device, and storage medium

CN116153449BActive Publication Date: 2026-09-11BEIJING HUIJI ZHIYI TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310028186.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2026-09-11
Estimated Expiration
2043-01-09

AI Technical Summary

Technical Problem

[0004]本发明提供一种诊断多写检测方法、装置、电子设备和存储介质,用以解决现有技术中进行诊断多写检测时可能会出现检测方式不灵活,导致遗漏或者误判诊断多写的情况,或者,只能针对主要诊断进行判断是否多写,对大概率出现诊断多写的非主要诊断不能进行精准的判断是否存在诊断多写的缺陷

Benefits of technology

[0042]本发明提供的一种诊断多写检测方法、装置、电子设备和存储介质,通过检测诊断疾病名所对应的病历文本中包含的疾病名,并基于与诊断疾病名相关的疾病名在病历文本中所处的片段,以及从病历文本中检索得到的与诊断疾病名相关的片段,确定诊断疾病名的相关片段集合,确保了病历文本中相关片段召回的全面性。随后基于诊断疾病名的相关片段集合中的各片段分别与诊断疾病名之间的相关度,进行诊断多写检测,实现了不遗漏且不误判的相对精准的诊断多写检测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116153449B_ABST
    Figure CN116153449B_ABST
Patent Text Reader

Abstract

The application provides a diagnosis multi-write detection method and device, electronic equipment and a storage medium, wherein the method comprises the following steps: obtaining a diagnosis disease name to be detected; detecting a disease name contained in a medical record text corresponding to the diagnosis disease name, a segment in which a disease name related to the diagnosis disease name is located in the medical record text, and a segment related to the diagnosis disease name retrieved from the medical record text, and determining a related segment set of the diagnosis disease name; and performing diagnosis multi-write detection based on the correlation between each segment in the related segment set of the diagnosis disease name and the diagnosis disease name. The method, device, electronic equipment and storage medium provided by the application ensure the comprehensiveness of the related segments in the medical record text by determining the related segment set of the diagnosis disease name. Then, the diagnosis multi-write detection is performed based on the correlation between each segment in the related segment set of the diagnosis disease name and the diagnosis disease name, so that the diagnosis multi-write detection is not missed and misjudged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, and in particular to a diagnostic method, apparatus, electronic device, and storage medium for detecting multiple write errors. Background Technology

[0002] Current methods for detecting multiple diagnoses primarily rely on knowledge base-assisted detection. For each disease name in the diagnosis list, relevant elements of the diagnosis name are searched throughout the entire medical record text to determine if multiple diagnoses have been written. Alternatively, detection can be performed using disease prediction methods. By inputting the medical record text, deep learning techniques are used to perform end-to-end prediction of the diseases corresponding to the text to determine if multiple diagnoses have been written.

[0003] However, using knowledge base-assisted detection can lead to inflexible detection methods, resulting in omissions or misjudgments of multiple diagnoses. On the other hand, using disease prediction methods to detect multiple diagnoses can only predict the primary diagnosis and it is difficult to predict non-primary diagnoses. Therefore, even if a certain disease name is not in the prediction results, it is often difficult to determine whether multiple diagnoses exist. Summary of the Invention

[0004] This invention provides a diagnostic overwrite detection method, apparatus, electronic device, and storage medium to address the shortcomings of existing technologies in diagnostic overwrite detection, which may result in inflexible detection methods, leading to omissions or misjudgments of diagnostic overwrite, or the inability to accurately determine whether overwrite exists for non-primary diagnoses that are likely to result in overwrite.

[0005] This invention provides a diagnostic method for multiple write detection, comprising:

[0006] Obtain the name of the disease to be diagnosed;

[0007] The disease name contained in the medical record text corresponding to the diagnosed disease name is detected, and based on the segment in the medical record text where the disease name related to the diagnosed disease name is located, and the segment related to the diagnosed disease name retrieved from the medical record text, the set of related segments of the diagnosed disease name is determined.

[0008] Based on the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name, a diagnostic multi-write detection is performed.

[0009] According to a diagnostic multiple-write detection method provided by the present invention, the step of determining the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name includes:

[0010] Based on the disease elements and / or disease knowledge of the diagnosed disease name, calculate the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name.

[0011] According to a diagnostic multi-write detection method provided by the present invention, the step of calculating the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name based on the disease elements and / or disease knowledge of the diagnosed disease name includes:

[0012] Based on the disease elements and / or disease knowledge of the diagnosed disease name, construct a disease graph for the diagnosed disease name;

[0013] Based on the correlation between the words in each segment and the entities in each segment, a segment graph for each segment is constructed.

[0014] Based on the fragment graph of each fragment and the disease graph of the diagnosed disease name, the correlation between each fragment and the diagnosed disease name is calculated.

[0015] According to a diagnostic multiple-write detection method provided by the present invention, the step of calculating the correlation between each fragment and the diagnosed disease name based on the fragment graph of each fragment and the disease graph of the diagnosed disease name includes:

[0016] Based on the previous segment graph representation of any segment, intra-graph information transfer is performed on the segment graph of the any segment to obtain the current segment graph representation; and based on the previous disease graph representation of the disease graph, intra-graph information transfer is performed on the disease graph to obtain the current graph representation.

[0017] Based on the correlation between the current fragment graph representation and the current disease graph representation, update the current fragment graph representation and the current disease graph representation, and use the updated current fragment graph representation and the current disease graph representation as the previous fragment graph representation and the previous disease graph representation, respectively, until the number of updates reaches a preset threshold;

[0018] Based on the correlation between the current fragment graph representation and the current disease graph representation when the number of updates reaches a preset threshold, the correlation between any fragment and the diagnosed disease name is determined.

[0019] According to a diagnostic multi-write detection method provided by the present invention, the step of determining a set of related segments of the diagnosed disease name based on the segment of the disease name related to the diagnosed disease name in the medical record text and the segments related to the diagnosed disease name retrieved from the medical record text includes:

[0020] The name of the diagnosed disease contained in the medical record text is determined from the disease names contained in the medical record text.

[0021] Based on the segment of the currently diagnosed disease name related to the diagnosed disease name in the medical record text, and the segments related to the diagnosed disease name retrieved from the medical record text, a set of related segments for the diagnosed disease name is determined.

[0022] According to a diagnostic multiple-write detection method provided by the present invention, the step of determining the name of the currently diagnosed disease contained in the medical record text from the disease names contained in the medical record text includes:

[0023] Retrieve the disease names contained in the medical record text;

[0024] Based on the context of the disease name in the medical record text, determine whether the disease name is the name of the disease diagnosed in this case.

[0025] According to the present invention, a diagnostic multiple-write detection method is provided, wherein the diagnostic multiple-write detection is performed based on the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name, including:

[0026] The maximum correlation is determined from the correlation between each fragment and the name of the diagnosed disease.

[0027] Based on a preset relevance threshold and the maximum relevance, it is determined whether the diagnosed disease name belongs to multiple diagnostic entries.

[0028] According to the present invention, a diagnostic multiple-write detection method is provided, wherein the diagnostic multiple-write detection is performed based on the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name, including:

[0029] Based on the detection model, the correlation between each segment and the diagnosed disease name is calculated, and the correlation between each segment and the diagnosed disease name is used to perform diagnostic multiple writing detection.

[0030] The detection model is trained based on the first negative sample and / or the second negative sample, as well as the positive sample.

[0031] The positive example sample includes the sample diagnosis disease name corresponding to the sample medical record, and a set of related fragments of the sample diagnosis disease name;

[0032] The first negative example sample includes a first disease name that belongs to the same category of diseases as the disease name diagnosed in the sample, and a set of related fragments of the first disease name determined based on the medical records of the sample;

[0033] The second negative sample includes a randomly determined second disease name and a set of related fragments of the second disease name determined based on the sample medical records.

[0034] The present invention also provides a diagnostic multiple write detection device, comprising:

[0035] The acquisition unit retrieves the name of the diagnostic disease to be detected;

[0036] The recall unit detects the disease name contained in the medical record text corresponding to the diagnosed disease name, and determines the set of related fragments of the diagnosed disease name based on the segment of the disease name related to the diagnosed disease name in the medical record text and the segments related to the diagnosed disease name retrieved from the medical record text.

[0037] The detection unit performs diagnostic multi-write detection based on the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name.

[0038] The present invention also provides an electronic device, comprising:

[0039] A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the diagnostic multi-write detection method as described in any of the preceding claims.

[0040] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the diagnostic multi-write detection method as described above.

[0041] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the diagnostic multi-write detection method as described above.

[0042] This invention provides a diagnostic multi-write detection method, apparatus, electronic device, and storage medium. By detecting the disease name contained in the medical record text corresponding to the diagnosed disease name, and based on the segment of the medical record text containing the disease name related to the diagnosed disease name, as well as the segments related to the diagnosed disease name retrieved from the medical record text, a set of relevant segments for the diagnosed disease name is determined, ensuring the comprehensiveness of the recall of relevant segments in the medical record text. Subsequently, based on the relevance between each segment in the set of relevant segments for the diagnosed disease name and the diagnosed disease name, diagnostic multi-write detection is performed, achieving relatively accurate diagnostic multi-write detection that is neither omitted nor misjudged. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0044] Figure 1 This is one of the flowcharts of the diagnostic multiple-write detection method provided by the present invention;

[0045] Figure 2 This is one of the flowcharts illustrating the relevance calculation method provided by the present invention;

[0046] Figure 3 This is a schematic diagram of node connections in the disease graph provided by the present invention;

[0047] Figure 4 This is a schematic diagram of node connections in a fragment diagram provided by the present invention;

[0048] Figure 5 This is the second flowchart illustrating the relevance calculation method provided by the present invention;

[0049] Figure 6 This is a schematic diagram of the similarity matching layer provided by the present invention;

[0050] Figure 7 This is a schematic diagram of the diagnostic multi-write detection process based on each fragment and the name of the diagnosed disease provided by the present invention;

[0051] Figure 8 This is the second flowchart of the diagnostic multi-write detection method provided by the present invention;

[0052] Figure 9 This is a schematic diagram of the structure of the diagnostic multi-write detection device provided by the present invention;

[0053] Figure 10 This is a schematic diagram of the physical structure of the electronic device provided by the present invention. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0055] When performing diagnostic overwrite detection using related technologies, the detection methods may be inflexible, leading to omissions or misjudgments of diagnostic overwrite. Alternatively, it may only be able to determine whether overwrite exists for primary diagnoses, and cannot accurately determine whether overwrite exists for non-primary diagnoses that are likely to be overwritten.

[0056] To address the aforementioned problems, this invention provides a diagnostic multiple-write detection method for use in medical record quality inspection. Figure 1 This is one of the flowcharts of the diagnostic multi-write detection method provided by the present invention, such as... Figure 1 As shown, the method includes:

[0057] Step 110: Obtain the name of the disease to be diagnosed.

[0058] Here, the diagnosed disease name refers to the disease name in the diagnosis list, which is the object to be checked for multiple diagnoses in the medical record text. It is understood that the diagnosis list can contain one or more diagnosed disease names. In the following steps, the detection of multiple diagnoses for each diagnosed disease name can be performed separately.

[0059] Step 120: Detect the disease name contained in the medical record text corresponding to the diagnosed disease name, and determine the relevant fragment set of the diagnosed disease name based on the segment of the disease name related to the diagnosed disease name in the medical record text, and the segments related to the diagnosed disease name retrieved from the medical record text.

[0060] Specifically, the medical record text corresponds to the diagnosis list containing the diagnosed disease name, and the medical record text is the textual representation of the medical record corresponding to the diagnosis list.

[0061] Considering that in related solutions, the method of searching for relevant elements of disease diagnosis names from medical record text based on knowledge base-assisted detection is inevitably prone to false alarms if the element description method changes, leading to multiple false alarms. Conversely, relaxing the standards to avoid false alarms could result in a large number of missed detections. To address this issue, this embodiment of the invention employs a combination of forward and reverse recall when acquiring relevant fragments of the diagnosed disease name, thereby avoiding situations where relevant fragments are not recalled due to changes in the element description method.

[0062] Here, disease names related to the diagnosed disease name are detected in the medical record text, and the fragments containing these disease names within the text are obtained as relevant fragments for the corresponding diagnosed disease name; this is called reverse recall. Fragments related to the diagnosed disease name are directly detected in the medical record; this is called forward recall. The relevant fragments obtained through forward recall and reverse recall are then combined to obtain a complete set of relevant fragments for the diagnosed disease name. Here, relevant fragments refer to text fragments related to the diagnosed disease name; specifically, relevant fragments can be descriptive fragments of the disease elements of the diagnosed disease name. It can be understood that relevant fragments can serve as the diagnostic basis for the corresponding diagnosed disease name.

[0063] For reverse recall, we can first obtain the disease names contained in the medical record text. Specifically, this can be achieved through disease vocabulary matching methods or NER (Named Entity Recognition) tools on the medical record text, recalling all disease names in the medical record text. Then, we can calculate the relevance between the recalled disease names and the diagnosed disease names in the diagnosis list, thereby determining whether the recalled disease names and the diagnosed disease names are related, thus obtaining the disease names related to the diagnosed disease names. For example, the disease name with the highest relevance to the diagnosed disease name among the recalled disease names can be taken as the disease names related to the diagnosed disease name. This relevance can be achieved by calculating edit distance, or by using text representation learning methods such as BERT (Bidirectional Encoder Representations from Transformer), autoencoders, and contrastive learning. After determining the disease names related to the diagnosed disease names, the segments of the relevant disease names in the medical record text can be included as segments in the set of relevant segments for the diagnosed disease names, i.e., the relevant segments for the diagnosed disease names.

[0064] For positive recall, a disease-specific name-based retrieval (NER) tool can be used to extract disease elements from medical record text. These disease elements can include etiology, pathology, location, and clinical manifestations. After extraction, the fragments containing these disease elements within the medical record text can be retrieved as fragments in the relevant fragment set for diagnosing the disease name. Alternatively, key attributes of each disease in a disease knowledge base can be used to detect and retrieve fragments containing these key attributes from the medical record text, also as fragments in the relevant fragment set for diagnosing the disease name. The NER tool can be implemented using mature named entity recognition methods such as BiLSTM+CRF (Bi-directional Long Short-Term; Conditional Random Field) and BERT+CRF. Furthermore, any high-quality knowledge base can be used; this embodiment of the invention does not impose specific limitations.

[0065] Understandably, by combining forward and reverse recall, we can compensate for the shortcomings of not being able to recall valuable relevant fragments of a diagnosis name through forward recall alone, because in some cases the description of relevant elements of a diagnosis name in the medical record text differs greatly from the description of relevant elements of the same diagnosis name in the knowledge base, or the knowledge base of a certain diagnosis name is incomplete. This results in a more complete set of relevant fragments of the diagnosis name.

[0066] Step 130: Perform diagnostic multi-write detection based on the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name.

[0067] Specifically, after obtaining the set of relevant segments for the diagnosed disease name in the medical record, it is necessary to calculate the relevance between each segment in the relevant segment set and its corresponding diagnosed disease name. This is to assess whether there are any segments in the medical record text that can serve as diagnostic evidence for the diagnosed disease name, thereby determining whether there is a problem of over-writing the diagnosed disease name. Understandably, the greater the relevance between a segment in the relevant segment set retrieved from the medical record and its corresponding diagnosed disease name, the more likely it is to be the diagnostic evidence for that disease name, compared to segments with low relevance, thus proving that there is no problem of over-writing the diagnosed disease name. Conversely, if the relevance between each segment in the relevant segment set and the diagnosed disease name is relatively low, meaning that none of the segments can serve as diagnostic evidence for that disease name, then there may be a problem of over-writing the diagnosed disease name.

[0068] For example, a medical record might read: "Patient Information: Female, 74 years old; Department: Geriatrics; Diagnosis List: Coronary atherosclerosis, sclerotic heart disease; Admission: The patient presented with recurrent dizziness and chest tightness for over 10 years, accompanied by worsening shortness of breath and coughing for 2 days. Upon admission, physical examination revealed: clear consciousness, poor mental state; Auxiliary Examinations: Head and chest CT scans, sinus rhythm; complete left bundle branch block; Treatment Process: After admission, relevant examinations were completed. Based on the patient's symptoms, signs, imaging examinations, and past medical history, the diagnosis was confirmed. Comprehensive treatment was administered targeting cardiovascular and cerebrovascular diseases, including improving cardiovascular and cerebrovascular circulation." The corresponding diagnoses in this medical record are coronary atherosclerosis and sclerotic heart disease. The correlation between the phrase "recurrent dizziness and chest tightness for over 10 years, accompanied by worsening shortness of breath and coughing for 2 days" and the diagnosis of coronary atherosclerosis is sufficient to confirm that coronary atherosclerosis is not a case of misdiagnosis; similarly, the correlation between the phrase "sinus rhythm" and the diagnosis of sclerotic heart disease is sufficient to confirm that sclerotic heart disease is not a case of misdiagnosis.

[0069] The method provided in this invention detects the disease name contained in the medical record text corresponding to the diagnosed disease name, and determines the set of relevant segments for the diagnosed disease name based on the segment of the medical record text containing the disease name related to the diagnosed disease name, as well as the segments related to the diagnosed disease name retrieved from the medical record text. This ensures the comprehensiveness of the recall of relevant segments in the medical record text. Subsequently, based on the relevance between each segment in the set of relevant segments for the diagnosed disease name and the diagnosed disease name, multiple diagnostic write detection is performed, achieving relatively accurate multiple diagnostic write detection without omissions or misjudgments.

[0070] Based on the above embodiments, step 130, the step of determining the relevance between each segment in the relevant segment set of the diagnosed disease name and the diagnosed disease name, includes:

[0071] Based on the disease elements and / or disease knowledge of the diagnosed disease name, calculate the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name.

[0072] Generally speaking, the diagnostic basis for a disease name in a medical record text usually involves disease elements or disease knowledge related to that disease name. Accordingly, multiple write detection of diagnoses can be performed based on the correlation between each fragment in the calculated set of relevant fragments of the disease name and the disease name characterized by disease elements and / or disease knowledge.

[0073] Here, the disease elements of a diagnosed disease name can include information such as the location and clinical manifestations of the diagnosed disease, and the disease elements can be extracted from the diagnosed disease name itself; the disease knowledge of a diagnosed disease name can include medical knowledge related to the diagnosis of the diagnosed disease name, such as information from multiple dimensions such as symptoms, signs, patient population attributes, onset time, and causative factors.

[0074] To obtain the relevance between each segment in the set of relevant segments for a diagnosed disease name and the diagnosed disease name itself, feature encoding can be performed on the diagnosed disease name based on its disease elements and / or disease knowledge. Specifically, disease elements and / or disease knowledge are used to characterize the features of the diagnosed disease name; the encoded features are then considered as the feature codes of the diagnosed disease name itself. Alternatively, feature encoding can be performed on each segment in the set of relevant segments for the diagnosed disease name. Subsequently, based on the feature codes of the diagnosed disease name and the segment feature codes, similarity calculations can be performed between the feature codes, thereby determining the relevance between the segment and the diagnosed disease name.

[0075] Based on any of the above embodiments Figure 2 This is one of the flowcharts illustrating the relevance calculation method provided by the present invention, such as... Figure 2 As shown, the step of calculating the relevance between each segment in the relevant segment set of the diagnosed disease name and the diagnosed disease name based on the disease elements and / or disease knowledge of the diagnosed disease name includes:

[0076] Step 210: Based on the disease elements and / or disease knowledge of the diagnosed disease name, construct a disease graph for the diagnosed disease name;

[0077] Here, for a diagnosed disease name, disease elements extracted from the diagnosed disease name can be directly applied, and / or entities and elements related to the diagnosed disease name extracted from the disease knowledge of the diagnosed disease name can be applied to construct a disease graph for the diagnosed disease name.

[0078] Understandably, a disease graph is a graph structure representing the disease elements of a diagnosed disease. In a disease graph, disease elements can be used as nodes, and the existence of connections between corresponding nodes is determined by whether there are relationships between disease elements.

[0079] For example, a disease map can be constructed using the following steps:

[0080] First, the NER tool is used to extract key disease elements from the diagnosed disease name, denoted as K, such as location and clinical manifestations. Then, relevant disease knowledge is extracted from a disease knowledge base, denoted as T. The diagnosed disease name itself is denoted as 1. Therefore, for a given diagnosed disease name, the number of nodes in its disease graph is ultimately K + T + 1.

[0081] Next, connect the nodes in the disease graph. Figure 3 This is a schematic diagram of node connections in the disease graph provided by the present invention, such as... Figure 3 As shown: In the disease graph, all nodes contained in disease element K and disease knowledge T are connected to the disease itself node 1. It can be understood that there is more than one disease element and / or disease knowledge for diagnosing a disease name, so there are also more than one corresponding node. That is, disease element K contains nodes k1, k2, k3…, and disease knowledge T contains nodes t1, t2, t3…. Any two nodes k1 and k2 in node K, if they are connected end-to-end in the diagnosed disease name, are connected as k1 and k2.

[0082] Step 220: Based on the correlation between the word segments in each segment and the entities in each segment, construct the segment graph for each segment.

[0083] Here, for each segment in the set of relevant segments for diagnosing disease names, a segment graph can be constructed for each segment. It can be understood that the segment graph of any segment is a graph structure representation of the information contained in that segment. In the segment graph, each word in the segment can be used as a node, and the existence of connecting edges between corresponding nodes can be determined by whether there is a correlation between the word segments. Furthermore, all entities in the segment can also be used as nodes, and for the nodes corresponding to the word segments belonging to entities in the segment, connecting edges can be established between those nodes and the nodes corresponding to the word segments that intersect with each other in the segment.

[0084] For example, a fragment graph can be constructed using the following steps:

[0085] First, any segment is segmented into words to obtain all the words in the segment, and then the word nodes of each word are established, denoted as W; then, the NER tool is used to extract all entities in the segment to obtain the entity nodes, denoted as N.

[0086] Next, the nodes in the obtained fragment graph are connected. Figure 4 This is a schematic diagram of node connections in a fragment diagram provided by the present invention, such as... Figure 4As shown: For word nodes in word segmentation, word nodes of each word segment in word node W are connected sequentially according to the sequence order of the segments. The connection edges between word nodes can be measured by calculating the PMI (Pointwise Mutual Information) of the two word nodes to measure the relevance between the two word segments. Specifically, if the PMI value of the two nodes is greater than a certain preset threshold, the two word nodes are connected. For example, calculate the PMI(w1,w2) of any two word nodes w1 and w2 in all segments. If PMI(w1,w2) > a certain preset threshold, then word node w1 and word node w2 are connected. For each node in entity node N, such as n1, any entity node is connected to word nodes that have intersection with the segments. The intersection here can be semantic or logical intersection.

[0087] Step 230: Based on the segment graph of each segment and the disease graph of the diagnosed disease name, calculate the correlation between each segment and the diagnosed disease name.

[0088] Specifically, after completing the graph construction for each segment and the diagnosed disease name, the segment graph and the disease graph can be encoded separately, and the correlation can be calculated using the encoded graph representation to obtain the correlation between the segment graph and the disease graph, that is, the correlation between each segment and the corresponding diagnosed disease name.

[0089] Each node in the disease graph is encoded, and word vectors are obtained by encoding the words within each node. Node encoding can use methods such as word2vec and BERT to generate word vectors from the words within a node. It is understood that a single node can include one or more words. When a node includes a single word, its node vector can be obtained from the word vector of that single word; when a node includes multiple words, its node vector is obtained by averaging the word vectors of all the words. Furthermore, the disease itself has a zero-vector node vector.

[0090] Each node in the fragment graph is encoded using node encoding, and word vectors are obtained by encoding the word segments within the nodes. It's understandable that entities are a type of segmentation with attributes, so both word nodes obtained from segmentation and node nodes obtained from entities can be encoded into word vectors using methods such as word2vec and BERT.

[0091] The method provided in this invention determines the correlation between disease names and fragments based on a graph structure, effectively combining disease-related knowledge and related fragment information, thus helping to improve the accuracy and reliability of diagnostic multi-segment detection. Based on any of the above embodiments, Figure 5 This is the second flowchart illustrating the relevance calculation method provided by the present invention, as shown below. Figure 5 As shown, step 230 includes:

[0092] Step 510: Based on the previous segment graph representation of any segment, perform intra-graph information transfer on the segment graph of any segment to obtain the current segment graph representation; and based on the previous disease graph representation of the disease graph, perform intra-graph information transfer on the disease graph to obtain the current graph representation.

[0093] Step 520: Based on the correlation between the current fragment graph representation and the current disease graph representation, update the current fragment graph representation and the current disease graph representation, and use the updated current fragment graph representation and the current disease graph representation as the previous fragment graph representation and the previous disease graph representation, respectively, until the number of updates reaches a preset threshold.

[0094] Specifically, for a diagnosed disease name and any fragment in the set of related fragments for that diagnosed disease name, in order to obtain the correlation between the two, it is necessary to perform feature extraction and joint feature interaction on the disease image of the diagnosed disease name and the fragment image of the fragment. The feature extraction performed separately on the disease image and the fragment image corresponds to step 510, and the joint feature interaction on the disease image and the fragment image corresponds to step 520. Furthermore, feature extraction and feature interaction can be multi-round; that is, the fragment image representation and disease image representation obtained after the feature interaction in the previous round can be used as the basis for feature extraction in the current round.

[0095] Furthermore, in step 510, to achieve feature extraction, intra-graph message passing can be performed in the fragment graph of any segment and in the disease graph. Here, intra-graph message passing can be implemented using a Gated GNN (Graph Neural Network) to update the previous graph representation of each graph, thereby obtaining the current graph representation of the corresponding graph.

[0096] For example, the adjacency matrix of a single graph and the representation of the previous graph can be input into a model responsible for implementing intra-graph message passing, such as a Gated GNN. The Gated GNN then calculates the updated graph node representation using the following formula, thus obtaining the current graph representation. The calculation formula is as follows:

[0097] T t =AH t-1 +b t

[0098] z t =σ(T) t W z,t +H t-1 U z,t )

[0099] rt =σ(T) t W r,t +H t-1 U r,t )

[0100]

[0101]

[0102] Among them, H t ∈R N×e In this context, N represents the number of nodes, e represents the vector dimension, and H represents the vector dimension. t H represents the encoding of the t-level node. t-1 This is the initial node encoding. A∈R N×N Let b be the adjacency matrix of the graph. t ∈R e W z,t ∈R e×e W r,t ∈R e×e U z,t ∈R e×e U r,t ∈R e×e W g,t ∈R e×e U g,t ∈R e×e These are trainable parameters. t ∈R N×e r t ∈R N×e These are the update gate and the forget gate. σ represents the sigmoid function, and ⊙ represents dot product. The update and forget gates control information transfer in a GRU-like manner.

[0103] In step 520, the updated graph representation obtained in step 510 is used as the current graph representation, i.e., the current disease graph representation and the current fragment graph representation. Information can be exchanged based on the correlation between the two, and the updated current fragment graph representation and current disease graph representation after information exchange are used as the previous fragment graph representation and previous disease graph representation for use in the next round of feature extraction.

[0104] The feature interaction here specifically uses the current disease graph representation and the current fragment graph as input for the attention interaction. Through attention matching, updated representations of the current disease graph and the current fragment graph are obtained. The attention matching is calculated using the following formula:

[0105] S d,t =relu(H d,t W d,t +b d,t )

[0106] S p,t =relu(H p,t W p,t +b p,t )

[0107]

[0108] H d,t+1 =att t S p,t +S d,t

[0109] Where H d,t ∈R N×e H p,t ∈R M×e , respectively, represent the relevant node vectors in the disease graph representation at layer t, and the relevant fragment graph node vectors in the fragment graph representation. N and M represent the number of disease graph nodes and the number of fragment graph nodes, respectively. t ∈R N×M This indicates the attention score. W d,t ∈R e×e W p,t ∈R e×e b d,t ∈R e b p,t ∈R e These are trainable parameters.

[0110] Overall, steps 510 and 520 above can be regarded as the similarity matching layer required for relevance calculation. Figure 6 This is a schematic diagram of the similarity matching layer provided by the present invention, as shown below. Figure 6 As shown, a Gated GNN and attention interaction are set as a group. The similarity matching layer can include multiple cascaded Gated GNNs and attention interactions to achieve multiple rounds of feature extraction and feature interaction until the number of feature interactions, i.e. the number of updates, reaches a preset threshold.

[0111] Step 530: Based on the correlation between the current fragment graph representation and the current disease graph representation when the number of updates reaches a preset threshold, determine the correlation between any fragment and the diagnosed disease name.

[0112] Specifically, step 510 involves message passing within the disease graph and the fragment graph of the related fragment, respectively, and step 520 involves attention matching on the disease graph and the fragment graph of the related fragment, updating the graph representations of the disease graph and the fragment graph of the related fragment. If the number of updates does not reach a preset threshold, steps 510 and 520 are continued until the number of updates reaches the preset threshold. The current disease graph representation and fragment graph representation obtained at this point can be used to calculate the correlation between the diagnosed disease name and the fragment. Here, the value of the correlation can be the correlation applied when the features are interacted with in step 520, or it can be calculated based on a similarity matching algorithm. This embodiment of the invention does not specifically limit this value.

[0113] The method provided in this invention achieves reliable and accurate correlation determination through multiple rounds of feature extraction and feature interaction, which helps to improve the reliability of diagnostic multi-write detection.

[0114] Based on any of the above embodiments, step 120 includes:

[0115] The name of the diagnosed disease contained in the medical record text is determined from the disease names contained in the medical record text.

[0116] Based on the segment of the currently diagnosed disease name related to the diagnosed disease name in the medical record text, and the segments related to the diagnosed disease name retrieved from the medical record text, a set of related segments for the diagnosed disease name is determined.

[0117] Specifically, considering that directly retrieving all disease names from the medical record text ensures that no disease name is missed, the resulting disease names may include records of past medical history or preventative medical advice for potential future illnesses. In other words, the disease names retrieved from the medical record text may include preventative, differential, or suspected diseases not present in the current case. Therefore, it is necessary to further select confirmed disease names from all disease names retrieved from the medical record text. Here, a confirmed disease name refers to the currently diagnosed disease present in the medical record text. For the selection of confirmed disease names, the segment containing each disease name can be searched throughout the entire medical record text, and its context can be used to determine whether the disease name is a confirmed disease name.

[0118] After obtaining the confirmed disease name, during reverse recall, related confirmed disease names can be retrieved, along with the fragments containing those related confirmed disease names within the medical record text. These fragments are then used as the relevant fragments for the related confirmed disease names. It is understandable that this reverse recall, compared to retrieving related disease names from all disease names appearing in the medical record text and then performing fragment recall, avoids the risk of falsely recalling fragments not related to the current confirmed disease, further ensuring the reliability of multiple diagnostic writes.

[0119] Based on any of the above embodiments, determining the name of the currently diagnosed disease contained in the medical record text from the disease names contained in the medical record text includes:

[0120] Retrieve the disease names contained in the medical record text;

[0121] Based on the context of the disease name in the medical record text, determine whether the disease name is the name of the disease diagnosed in this case.

[0122] Specifically, considering that all disease names retrieved directly from the medical record text can be retrieved without missing any, the process involves determining whether a disease name is a confirmed diagnosis based on its context. It's understandable that disease names in the medical record text might contain records of past medical history or preventative medical advice for potential future illnesses. Therefore, the disease names retrieved from the medical record text may include preventative, diagnostic, or suspected diseases, and thus not necessarily present in the current case. Consequently, it's necessary to further select confirmed disease names from all the disease names retrieved from the medical record text. For the selection of confirmed disease names, the segment containing each disease name can be searched throughout the entire medical record text, and its context can be used to determine whether the disease name is a confirmed diagnosis.

[0123] Here, the context of a fragment can be the sentence containing the fragment, or the sentence containing the fragment along with the preceding and following sentences. This embodiment of the invention does not impose specific limitations on this. After obtaining the context of the fragment, the semantics of the context can be combined to determine whether the disease name in the fragment is the currently diagnosed disease. It is understood that the diagnosed disease name is more likely to have a greater correlation with the diagnosed disease names in the diagnosis list than other disease names in the medical record text. Therefore, the process of determining the diagnosed disease name from all disease names contained in the medical record text is equivalent to performing a rough screening of the correlation between all disease names contained in the medical record text and the disease names in the diagnosis list, avoiding irrelevant fragments that are not diagnosed in the current medical record text, and further improving the reliability of the diagnosis overwrite detection.

[0124] Based on any of the above embodiments, step 130 includes

[0125] The maximum correlation is determined from the correlation between each fragment and the name of the diagnosed disease.

[0126] Based on a preset relevance threshold and the maximum relevance, it is determined whether the diagnosed disease name belongs to multiple diagnostic entries.

[0127] Specifically, after calculating the correlation between each segment and its corresponding diagnostic disease name, the maximum correlation between each segment and its corresponding diagnostic disease name can be obtained, i.e., the maximum correlation.

[0128] Understandably, since some diagnosed disease names are not the primary diagnoses in the medical record text, there may only be one or two fragments related to the diagnosed disease name in the medical record text. The purpose of diagnosis overwriting detection is to determine whether there are overwritten diagnosed disease names. In order to avoid the same problem as diagnosis overwriting detection in disease prediction methods, which mistakenly regards non-primary diagnoses as diagnosis overwriting, as long as the correlation between a fragment and the diagnosed disease name is greater than the preset correlation threshold, it can be determined that the diagnosed disease name has a diagnostic basis and there is no diagnosis overwriting.

[0129] Therefore, it is possible to compare only the preset relevance threshold with the maximum relevance, without having to compare the preset relevance threshold with the relevance between each segment and the diagnosed disease name.

[0130] Here, the maximum relevance logit can be calculated using the following formula:

[0131]

[0132]

[0133] Z i =relu(Z′) i W′+b′)

[0134] logit=max(σ([Z1, Z2,...,Z i ]W″′+b″))

[0135] in The node vectors for the final fragment graph representation of each fragment. This represents the final disease graph node vector corresponding to each segment. Z′ i ∈R 4e Z i ∈R e This is the intermediate processing result. N and M represent the number of nodes in the disease graph and the number of nodes in the related fragment graph, respectively. W∈R 4e×4e U∈R4e×4e W′∈R 4e×e W″∈R e×1 b∈R 4e c∈R 4e b′∈R e b″∈R 1 All are trainable parameters. logit∈(0,1) represents the maximum correlation between each segment and the corresponding diagnosed disease name.

[0136] For example, Figure 7 This is a schematic diagram of the diagnostic multi-write detection process based on each fragment and the name of the diagnosed disease provided by the present invention, such as... Figure 7 As shown, the relevance between each segment output by the similarity matching layer and the diagnosed disease name is input into the max layer. The max layer then yields the maximum relevance among these relevance values. Comparing the maximum relevance with a preset relevance threshold yields the result of the diagnostic multiple-write detection.

[0137] The method provided in this invention performs diagnostic multiple-write detection based on the maximum correlation between relevant fragments and corresponding diagnostic disease names. Compared with the current diagnostic multiple-write detection through disease prediction methods, it can not only detect diagnostic multiple-write for primary diagnoses, but also for non-primary diagnoses that are prone to diagnostic multiple-write misjudgments, which is more in line with the actual scenarios in which diagnostic multiple-write occurs.

[0138] Based on any of the above embodiments, step 130 includes:

[0139] Based on the detection model, the correlation between each segment and the diagnosed disease name is calculated, and the correlation between each segment and the diagnosed disease name is used to perform diagnostic multiple writing detection.

[0140] The detection model is trained based on the first negative sample and / or the second negative sample, as well as the positive sample.

[0141] The positive example sample includes the sample diagnosis disease name corresponding to the sample medical record, and a set of related fragments of the sample diagnosis disease name;

[0142] The first negative example sample includes a first disease name that belongs to the same category of diseases as the disease name diagnosed in the sample, and a set of related fragments of the first disease name determined based on the medical records of the sample;

[0143] The second negative sample includes a randomly determined second disease name and a set of related fragments of the second disease name determined based on the sample medical records.

[0144] Here, the calculation of the relevance between the aforementioned fragments and the diagnosed disease names, as well as the detection of multiple diagnostic entries based on relevance, can both be implemented using a detection model. This detection model can be obtained through supervised learning using pre-labeled positive and negative samples. It can be understood that a positive sample in the conventional sense includes the diagnosed disease name of a sample that confirms no multiple diagnostic entries, and the set of relevant fragments for that sample's diagnosed disease name; a negative sample in the conventional sense is the diagnosed disease name of a sample that confirms multiple diagnostic entries, and the set of relevant fragments for that sample's diagnosed disease name. However, obtaining both positive and negative samples requires extensive labeling of large amounts of data by numerous professionals, which is time-consuming, labor-intensive, and prone to human error.

[0145] In view of this problem, embodiments of the present invention provide a method for obtaining samples without manual labeling.

[0146] Positive examples include the sample medical record's corresponding diagnostic disease name and a set of related fragments for that diagnostic disease name. It is understood that, based on experience, most sample medical record texts correspond to definite diagnostic diseases, meaning that most sample diagnostic disease names and their related fragment sets can be considered positive examples. Although a very small number of sample diagnostic disease names may not be explicitly described, i.e., a very small number of sample diagnostic disease names have issues with multiple diagnoses, considering the extremely small number of these sample diagnostic disease names, to reduce the cost of constructing training samples, this embodiment of the invention, without annotation, directly considers the sample medical record's corresponding diagnostic disease name and its related fragment set as positive examples.

[0147] In addition, regarding the construction of negative examples, two types of negative examples can be specifically obtained:

[0148] Here, the first negative example sample includes a first disease name belonging to the same disease category as the sample diagnosis disease name in the positive example sample, and a set of related fragments of the first disease name determined based on the sample medical record. For example, in ICD (International Classification of Diseases) coding, a disease name with the same 4-digit code but a different 6-digit code as the sample diagnosis disease name in the positive example sample can be selected as the first disease name, and the first disease name is different from the sample diagnosis disease name. It is understood that the first disease name itself is not the diagnosis disease name corresponding to the sample medical record text. Although a set of related fragments of the first disease name can be retrieved from the sample medical record text, there are no fragments in this set of related fragments that match the first disease name. Therefore, the first disease name and its set of related fragments can be used as the first negative example sample.

[0149] Additionally, the second negative sample includes a randomly determined second disease name and a set of related fragments for the second disease name determined based on the sample medical record. For example, a disease name can be randomly selected from the ICD encoding as the second disease name, and this second disease name is different from the sample diagnosed disease name. It is understood that the second disease name itself is not the diagnosed disease name corresponding to the sample medical record text. Although a set of related fragments for the second disease name can be retrieved from the sample medical record text, this set of related fragments does not contain any fragments matching the second disease name. Therefore, the second disease name and its set of related fragments can be used as the second negative sample.

[0150] Here, the selection of the first and second negative samples does not require additional manual labeling, greatly reducing the cost of constructing training samples. Furthermore, since the first disease name in the first negative sample belongs to the same disease category as the disease name in the positive sample, training the detection model using both the first and positive samples can further improve the model's ability to distinguish between multiple diagnoses of the same disease. Specifically, in model training, the positive samples obtained through the unlabeled method described above, along with the first and / or second negative samples, can be used to train the detection model; alternatively, the positive samples obtained through the unlabeled method, along with the first and / or second negative samples, can be used as sample A, and the training samples obtained through the labeled method can be used as sample B. First, a large-scale sample A is used to train the detection model, and then a small-scale sample B is used to fine-tune the model, resulting in a detection model with better performance in detecting multiple diagnoses.

[0151] Based on any of the above embodiments Figure 8 This is a second flowchart illustrating the diagnostic multi-write detection method provided by the present invention, as shown below. Figure 8 As shown:

[0152] First, the diagnostic disease names in the diagnostic list are obtained. Then, using these diagnostic disease names, fragments related to the diagnosed disease names are retrieved throughout the entire medical record text as a positive recall. Conversely, based on all disease names contained in the medical record text and the fragments containing the disease names related to the diagnosed disease names, a negative recall is generated. The fragments obtained from both positive and negative recall constitute a set of fragments related to all diagnostic disease names in the diagnostic list. The content of these related fragment sets can serve as diagnostic criteria for the corresponding disease names, such as etiology, pathology, location, and clinical manifestations, specifically "pulmonary inflammatory changes," "pulmonary discomfort," and "skin infection." Next, each fragment in the set of related fragments for the diagnosed disease names, along with the corresponding diagnosed disease names, is input into a disease description comparison model based on a knowledge graph neural network to obtain the final output of the multi-write diagnostic detection. Through the above process, this embodiment of the invention provides a relatively accurate multi-write detection method that avoids omissions and misjudgments.

[0153] The disease description comparison model based on knowledge graph neural networks consists of a similarity matching layer and a max layer. The similarity matching layer is handled by a Gated GNN for intra-graph message passing and an attention interaction layer that outputs updated fragment graph representations of each segment and disease graph representations of the corresponding diagnosed disease names through attention matching. When the number of updates reaches a preset threshold, the final fragment graph representations of each segment and the corresponding disease graph representations of the diagnosed disease names are input into the max layer to calculate the maximum correlation between each segment and the corresponding diagnosed disease name. The maximum correlation is then compared with a preset correlation threshold to perform diagnostic multiple write detection.

[0154] Based on any of the above embodiments Figure 9 This is a schematic diagram of the structure of the diagnostic multi-write detection device provided by the present invention, as shown below. Figure 9 As shown, the device includes:

[0155] Unit 910 retrieves the name of the diagnostic disease to be detected;

[0156] The recall unit 920 detects the disease name contained in the medical record text corresponding to the diagnosed disease name, and determines the set of related fragments of the diagnosed disease name based on the segment of the disease name related to the diagnosed disease name in the medical record text and the segments related to the diagnosed disease name retrieved from the medical record text.

[0157] The detection unit 930 performs diagnostic multi-write detection based on the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name.

[0158] The apparatus provided in this invention acquires the name of a diagnosed disease to be detected, detects the disease names contained in the corresponding medical record text, and determines a set of related fragments for the diagnosed disease name based on the segment of the medical record text containing the disease names related to the diagnosed disease name, and the segments related to the diagnosed disease name retrieved from the medical record text. Compared with current detection methods, this method does not miss any segments related to the diagnosed disease name. Finally, based on the relevance between each segment in the set of related fragments for the diagnosed disease name and the diagnosed disease name, a diagnostic multiple-write detection is performed, providing a relatively accurate multiple-write detection method that does not miss any and does not make false judgments, especially capable of detecting common cases of diagnostic multiple-write for non-primary diagnoses.

[0159] Based on any of the above embodiments, the detection unit is further configured to:

[0160] Based on the disease elements and / or disease knowledge of the diagnosed disease name, calculate the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name.

[0161] Based on any of the above embodiments, the detection unit is further configured to:

[0162] Based on the disease elements and / or disease knowledge of the diagnosed disease name, construct a disease graph for the diagnosed disease name;

[0163] Based on the correlation between the words in each segment and the entities in each segment, a segment graph for each segment is constructed.

[0164] Based on the fragment graph of each fragment and the disease graph of the diagnosed disease name, the correlation between each fragment and the diagnosed disease name is calculated.

[0165] Based on any of the above embodiments, the detection unit is further configured to:

[0166] Based on the previous segment graph representation of any segment, intra-graph information transfer is performed on the segment graph of the any segment to obtain the current segment graph representation; and based on the previous disease graph representation of the disease graph, intra-graph information transfer is performed on the disease graph to obtain the current graph representation.

[0167] Based on the correlation between the current fragment graph representation and the current disease graph representation, update the current fragment graph representation and the current disease graph representation, and use the updated current fragment graph representation and the current disease graph representation as the previous fragment graph representation and the previous disease graph representation, respectively, until the number of updates reaches a preset threshold;

[0168] Based on the correlation between the current fragment graph representation and the current disease graph representation when the number of updates reaches a preset threshold, the correlation between any fragment and the diagnosed disease name is determined.

[0169] Based on any of the above embodiments, the recall unit is further configured to:

[0170] The name of the diagnosed disease contained in the medical record text is determined from the disease names contained in the medical record text.

[0171] Based on the segment of the currently diagnosed disease name related to the diagnosed disease name in the medical record text, and the segments related to the diagnosed disease name retrieved from the medical record text, a set of related segments for the diagnosed disease name is determined.

[0172] Based on any of the above embodiments, the recall unit is further configured to:

[0173] Retrieve the disease names contained in the medical record text;

[0174] Based on the context of the disease name in the medical record text, determine whether the disease name is the name of the disease diagnosed in this case.

[0175] Based on any of the above embodiments, the detection unit is further configured to:

[0176] The maximum correlation is determined from the correlation between each fragment and the name of the diagnosed disease.

[0177] Based on a preset relevance threshold and the maximum relevance, it is determined whether the diagnosed disease name belongs to multiple diagnostic entries.

[0178] Based on any of the above embodiments, the detection unit further includes a detection model training unit, which is used for:

[0179] Based on the detection model, the correlation between each segment and the diagnosed disease name is calculated, and the correlation between each segment and the diagnosed disease name is used to perform diagnostic multiple writing detection.

[0180] The detection model is trained based on the first negative sample and / or the second negative sample, as well as the positive sample.

[0181] The positive example sample includes the sample diagnosis disease name corresponding to the sample medical record, and a set of related fragments of the sample diagnosis disease name;

[0182] The first negative example sample includes a first disease name that belongs to the same category of diseases as the disease name diagnosed in the sample, and a set of related fragments of the first disease name determined based on the medical records of the sample;

[0183] The second negative sample includes a randomly determined second disease name and a set of related fragments of the second disease name determined based on the sample medical records.

[0184] Figure 10 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 10 As shown, the electronic device may include a processor 1010, a communications interface 1020, a memory 1030, and a communication bus 1040, wherein the processor 1010, the communications interface 1020, and the memory 1030 communicate with each other via the communication bus 1040. The processor 1010 can call logical instructions in the memory 1030 to execute a diagnostic multi-write detection method. This method includes: obtaining the name of a diagnosed disease to be detected; detecting the disease names contained in the medical record text corresponding to the diagnosed disease name, and the segments of disease names related to the diagnosed disease name in the medical record text, as well as the segments related to the diagnosed disease name retrieved from the medical record text, determining a set of related segments for the diagnosed disease name; and performing diagnostic multi-write detection based on the relevance between each segment in the set of related segments for the diagnosed disease name and the diagnosed disease name.

[0185] Furthermore, the logical instructions in the aforementioned memory 1030 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0186] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the diagnostic multi-write detection method provided by the above methods. The method includes: obtaining a diagnostic disease name to be detected; detecting the disease name contained in the medical record text corresponding to the diagnostic disease name, and the segment in the medical record text where the disease name related to the diagnostic disease name is located, as well as the segments related to the diagnostic disease name retrieved from the medical record text, and determining a set of related segments of the diagnostic disease name; and performing diagnostic multi-write detection based on the correlation between each segment in the set of related segments of the diagnostic disease name and the diagnostic disease name.

[0187] On another front, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the diagnostic multi-write detection method provided by the above-described methods. This method includes: acquiring a diagnostic disease name to be detected; detecting disease names contained in the medical record text corresponding to the diagnostic disease name, and the segments in the medical record text containing disease names related to the diagnostic disease name, as well as segments related to the diagnostic disease name retrieved from the medical record text, determining a set of related segments for the diagnostic disease name; and performing diagnostic multi-write detection based on the correlation between each segment in the set of related segments for the diagnostic disease name and the diagnostic disease name. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0188] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A diagnostic method for multiple write detection, characterized in that, include: Obtain the name of the disease to be diagnosed; The disease name contained in the medical record text corresponding to the diagnosed disease name is detected, and based on the segment in the medical record text where the disease name related to the diagnosed disease name is located, and the segment related to the diagnosed disease name retrieved from the medical record text, the set of related segments of the diagnosed disease name is determined. Based on the disease elements and / or disease knowledge of the diagnosed disease name, construct a disease graph for the diagnosed disease name; Based on the correlation between each word segment in each segment of the relevant segment set of the diagnosed disease name, and the entity of each segment, construct the segment graph of each segment; Based on the previous segment graph representation of any segment, intra-graph information transfer is performed on the segment graph of the any segment to obtain the current segment graph representation; and based on the previous disease graph representation of the disease graph, intra-graph information transfer is performed on the disease graph to obtain the current disease graph representation. Based on the correlation between the current fragment graph representation and the current disease graph representation, update the current fragment graph representation and the current disease graph representation, and use the updated current fragment graph representation and the current disease graph representation as the previous fragment graph representation and the previous disease graph representation, respectively, until the number of updates reaches a preset threshold; Based on the correlation between the current fragment graph representation and the current disease graph representation when the number of updates reaches a preset threshold, the correlation between any fragment and the diagnosed disease name is determined; Based on the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name, a diagnostic multi-write detection is performed.

2. The diagnostic multiple-write detection method according to claim 1, characterized in that, The step of determining the set of relevant segments for the diagnosed disease name based on the segment of the disease name associated with the diagnosed disease name in the medical record text, and the segments related to the diagnosed disease name retrieved from the medical record text, includes: The name of the diagnosed disease contained in the medical record text is determined from the disease names contained in the medical record text. Based on the segment of the currently diagnosed disease name related to the diagnosed disease name in the medical record text, and the segments related to the diagnosed disease name retrieved from the medical record text, a set of related segments for the diagnosed disease name is determined.

3. The diagnostic multiple-write detection method according to claim 2, characterized in that, The step of determining the name of the currently diagnosed disease contained in the medical record text from the disease names contained in the medical record text includes: Retrieve the disease names contained in the medical record text; Based on the context of the disease name in the medical record text, determine whether the disease name is the name of the disease diagnosed in this case.

4. The diagnostic multiple-write detection method according to claim 1, characterized in that, The diagnostic multi-write detection is performed based on the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name, including: The maximum correlation is determined from the correlation between each fragment and the name of the diagnosed disease. Based on a preset relevance threshold and the maximum relevance, it is determined whether the diagnosed disease name belongs to multiple diagnostic entries.

5. The diagnostic multiple-write detection method according to claim 1, characterized in that, The diagnostic multi-write detection is performed based on the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name, including: Based on the detection model, the correlation between each segment and the diagnosed disease name is calculated, and the correlation between each segment and the diagnosed disease name is used to perform diagnostic multiple writing detection. The detection model is trained based on the first negative sample and / or the second negative sample, as well as the positive sample. The positive example sample includes the sample diagnosis disease name corresponding to the sample medical record, and a set of related fragments of the sample diagnosis disease name; The first negative example sample includes a first disease name that belongs to the same category of diseases as the disease name diagnosed in the sample, and a set of related fragments of the first disease name determined based on the medical records of the sample; The second negative sample includes a randomly determined second disease name and a set of related fragments of the second disease name determined based on the sample medical records.

6. A diagnostic multi-write detection device, characterized in that, include: The acquisition unit retrieves the name of the diagnostic disease to be detected; The recall unit detects the disease name contained in the medical record text corresponding to the diagnosed disease name, and determines the set of related fragments of the diagnosed disease name based on the segment of the disease name related to the diagnosed disease name in the medical record text and the segments related to the diagnosed disease name retrieved from the medical record text. The detection unit constructs a disease graph for the diagnosed disease name based on disease elements and / or disease knowledge; constructs a segment graph for each segment based on the correlation between word segments in each segment of the relevant segment set of the diagnosed disease name, and the entities of each segment; performs intra-graph information transfer on the segment graph of any segment based on the previous segment graph representation of any segment to obtain the current segment graph representation; and performs intra-graph information transfer on the disease graph based on the previous disease graph representation of the disease graph to obtain the current disease graph representation. Based on the correlation between the current fragment graph representation and the current disease graph representation, update the current fragment graph representation and the current disease graph representation, and use the updated current fragment graph representation and the current disease graph representation as the previous fragment graph representation and the previous disease graph representation, respectively, until the number of updates reaches a preset threshold; Based on the correlation between the current fragment graph representation and the current disease graph representation when the number of updates reaches a preset threshold, the correlation between any fragment and the diagnosed disease name is determined; Based on the correlation between each fragment in the relevant fragment set of the diagnosed disease name and the diagnosed disease name, a diagnostic multi-write detection is performed.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the diagnostic multi-write detection method as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the diagnostic multi-write detection method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Unreasonable disease diagnosis and detection method and device applied to clinical decision support system

    CN110033863A

  • ICD encoding method and device, electronic equipment and storage medium

    CN112183026A