A system for obtaining a disease database based on an NLP system
By introducing an NLP system-based method in the disease database construction, the problems of low processing efficiency and poor data accuracy in the existing technology are solved, and more efficient and accurate disease database construction is achieved.
Patent Information
- Application Number
- CN202410488593.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-23
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-04-23
AI Technical Summary
In the prior art, the use of rules to build and manage disease databases has problems such as low processing efficiency, poor data accuracy, and difficulty in adapting to large-scale data processing.
Using an NLP system-based method, the initial medical record text set is obtained, converted into an intermediate medical record text set, and input it into a preset NLP system to generate an intermediate event knowledge graph. Combining the preset intermediate rule list and knowledge graph, a disease database is constructed.
By improving the accuracy and efficiency of obtaining disease databases, the NLP system can more accurately understand and parse data in medical record text, providing more rich and accurate information for the rule engine.
Smart Images

Figure CN118299061B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text processing, and particularly to a system for obtaining a disease database based on an NLP system. Background Art
[0002] With the continuous development of Internet technology, medical record texts have become electronic. As important documents for judging diseases, treatment processes, and detecting the development status of diseases, these medical record texts are generated and stored. They contain rich event information, such as the occurrence, development, treatment process of diseases, and changes in the health status of patients. Effectively extracting and utilizing this event information and processing this time information to construct a disease database has become a popular research direction. In the prior art, many methods use rules to construct and manage disease databases, and these methods have problems such as low processing efficiency, poor data accuracy, and difficulty in adapting to large-scale data processing. Summary of the Invention
[0003] In view of the above technical problems, the technical solution adopted by the present invention is as follows: A system for obtaining a disease database based on an NLP system, the system includes: a processor and a memory storing a computer program. When the computer program is executed by the processor, the following steps are implemented:
[0004] S100, obtaining an initial medical record text set, where the initial medical record text set includes a plurality of initial medical record texts, and the initial medical record texts are medical record texts obtained from a preset database.
[0005] S200, obtaining an intermediate medical record text set according to the initial medical record text set, where the intermediate medical record text set includes a plurality of intermediate medical record texts, and the intermediate medical record texts are the initial medical record texts corresponding to each user obtained from the initial medical record text set.
[0006] S300, inputting the intermediate medical record text set into a preset NLP system to obtain an intermediate event knowledge graph corresponding to the intermediate medical record text set, where the intermediate event knowledge graph includes a plurality of target triples.
[0007] S400, obtaining a disease database according to a preset intermediate rule list and the intermediate event knowledge graph, where the disease database includes content data included in the processing nodes involved in the diseases.
[0008] The present invention has obvious beneficial effects compared with the prior art. By means of the above technical solution, the system for obtaining a disease database based on an NLP system provided by the present invention can achieve considerable technical progressiveness and practicality, and has wide industrial utilization value. It has at least the following beneficial effects:
[0009] The present invention relates to a system for obtaining a disease database based on an NLP system. The system includes a processor and a memory storing a computer program. When the computer program is executed by the processor, the following steps are implemented: obtaining an initial medical record text set, obtaining an intermediate medical record text set based on the initial medical record text set, inputting the intermediate medical record text set into a preset NLP system to obtain an intermediate event knowledge graph corresponding to the intermediate medical record text set, and obtaining a disease database according to a preset intermediate rule list and the intermediate event knowledge graph. As described above, by combining the NLP system with the rules, the data in the medical record text can be understood and analyzed more accurately, improving the accuracy and efficiency of obtaining the disease database. The NLP system can provide an understanding and analysis of the medical record text data, providing richer and more accurate information for the rule engine, resulting in a relatively high accuracy of the obtained disease database. At the same time, based on the text type corresponding to the medical record text, different methods are used to obtain the knowledge graph, and multiple models are combined to obtain the entities in the medical record text. The parameters of each model are continuously adjusted based on the sample data, improving the accuracy of obtaining the entities in the medical record text, and thus resulting in a relatively high accuracy of the obtained knowledge graph. The trained model is used to obtain the entity relationships in the medical record text, and different methods are used to adjust the parameters of the model based on the matching situation during the model training process, improving the accuracy of obtaining the entity relationships in the medical record text, and thus resulting in a relatively high accuracy of the obtained disease database.
[0010] The above description is only an overview of the technical solution of the present invention. In order to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. In order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following preferred embodiments are specifically described in detail in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a flowchart implemented when the processor of a system for obtaining a disease database based on an NLP system according to an embodiment of the present invention executes a computer program. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0013] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0014] Embodiment
[0015] This embodiment provides a system for obtaining a disease database based on an NLP system. The system includes: a processor and a memory storing a computer program. When the computer program is executed by the processor, the following steps are implemented, as Figure 1 shown:
[0016] S100, obtaining an initial medical record text set, where the initial medical record text set includes a number of initial medical record texts, and the initial medical record texts are medical record texts obtained from a preset database.
[0017] Specifically, those skilled in the art know that the database for obtaining medical record texts can be selected according to actual needs, and all fall within the protection scope of the present invention, which will not be elaborated here.
[0018] S200, obtaining an intermediate medical record text set according to the initial medical record text set, where the intermediate medical record text set includes a number of intermediate medical record texts, and the intermediate medical record texts are the initial medical record texts corresponding to each user obtained from the initial medical record text set.
[0019] Specifically, it can be understood that: the intermediate medical record text set is a multi-dimensional data set corresponding to a number of patients obtained by splitting the initial medical record text set according to the user, that is, the patient dimension.
[0020] Specifically, the data format corresponding to the intermediate medical record text is the JSON format. Among them, those skilled in the art know that any method of converting text into the JSON data format in the prior art falls within the protection scope of the present invention, which will not be elaborated here.
[0021] S300, inputting the intermediate medical record text set into a preset NLP system to obtain an intermediate event knowledge graph corresponding to the intermediate medical record text set, where the intermediate event knowledge graph includes a number of target triples.
[0022] Specifically, in S300, the target triple is obtained through the following steps:
[0023] S301, obtain the target medical record text label corresponding to the intermediate medical record text, where the target medical record text label is the text type corresponding to the intermediate medical record text obtained based on the content and structure corresponding to several titles in the intermediate medical record text.
[0024] Specifically, the system also includes several first medical record text labels and several second medical record text labels.
[0025] Further, the first medical record text label is the type of medical record text involved in the process nodes during the processing of the user when the user's body shows abnormalities, such as: physical signs, examinations, medications, etc. as the first medical record text labels.
[0026] Further, the second medical record text label is the type of medical record text corresponding to the characteristics of the user's differences that are irrelevant to the process nodes involved in the processing of the user when the user's body shows abnormalities, such as smoking history, menstrual history, allergy history, etc. as the second medical record text labels.
[0027] S302, when the target medical record text label is consistent with the first medical record text label, obtain the intermediate entity list corresponding to the intermediate medical record text and the intermediate entity label list corresponding to the intermediate entity list.
[0028] Specifically, the system also includes N preset entity labels, where the preset entity label is the label corresponding to the entity set in advance, and the label corresponding to the entity is the label corresponding to the word representing the user's physical state. Those skilled in the art know that the preset entity labels can be set according to actual needs, and all fall within the protection scope of the present invention and will not be elaborated here. For example: examination methods, observation results, etc. as preset entity labels.
[0029] Specifically, S302 also includes the following steps:
[0030] S3021, obtain the initial feature vector list A = {A 1 , ……, A i , ……, A n} corresponding to the intermediate medical record text, A i = (A i1 , ……, A ij , ……, A im ), A i is the initial feature vector corresponding to the i-th character in the intermediate medical record text, and A ij is A iThe bit value of the j-th position in it, where j = 1... m, m is the dimension of the initial feature vector, and i = 1... n, n is the number of Chinese character characters in the intermediate medical record text.
[0031] Specifically, the initial feature vector is a feature vector generated by concatenating the vectors generated based on the content information and position information corresponding to the Chinese character characters in the intermediate medical record text. Among them, those skilled in the art know that any method of converting text into a vector based on the content information and position information of the text in the prior art falls within the protection scope of the present invention and will not be elaborated here.
[0032] Further, m = m 1 +m 2 , where m 1 is the dimension of the vector generated based on the content information corresponding to the Chinese character characters in the intermediate medical record text, and m 2 is the dimension of the vector generated based on the position information corresponding to the Chinese character characters in the intermediate medical record text.
[0033] Preferably, the value of m 1 is 128.
[0034] Preferably, the value of m 2 is 64.
[0035] Specifically, the intermediate medical record text is the medical record text of the entity to be obtained and the entity corresponding label.
[0036] S3022, input A into the preset CNN model, and obtain the list of intermediate feature vectors B = {B 1 , ……, B i , ……, B n} corresponding to the intermediate medical record text, B i =(B i1 , ……, B ir , ……, B is ), B i is the intermediate feature vector corresponding to the i-th Chinese character character in the intermediate medical record text, and B ir is the bit value of the r-th position in B i , where r = 1... s, s is the dimension of the intermediate feature vector. Among them, in S3022, s is obtained through the following steps:
[0037] S30221, obtain the first intermediate parameter η 1 of the preset CNN model, where the first intermediate parameter η 1 is the number of types of convolution kernels in the preset CNN model.
[0038] Specifically, in S30221, η 1 is obtained through the following steps:
[0039] S1. Obtain a sample medical record text set, where the sample medical record text set includes a number of sample medical record texts, and the sample medical record texts are medical record texts for entity annotation.
[0040] S2. Obtain a sample entity set according to the sample medical record text set, where the sample entity set includes a number of sample entities, and the sample entities are entities annotated in the sample medical record texts.
[0041] S3. Obtain a first sample quantity list and a second sample quantity list F = {F 1 , ……, F t , ……, F f} corresponding to the sample entity type list, where F t is the second sample quantity corresponding to the t-th sample entity type, t = 1 …… f, and f is the number of second sample quantities in the second sample quantity list.
[0042] Specifically, the sample entity type list includes a number of sample entity types, where the sample entity types are entity types classified based on the number of literal characters corresponding to the sample entities. It can be understood that: those with 2 literal characters in the sample entities are classified into one category, such as skin, chest tube, number, etc.
[0043] Specifically, the first sample quantity list includes a number of first sample quantities, where the first sample quantity is the number of literal characters corresponding to the sample entity.
[0044] Specifically, the second sample quantity is the number of sample entities corresponding to each sample entity type in the sample entity set.
[0045] S4. Obtain an intermediate entity type list E = {E 1 , ……, E v , ……, E b} and a third sample quantity list E 0 = {E 0 1 , ……, E 0 v , ……, E 0 b} corresponding to E, where E v is the v-th intermediate entity type, and E 0 v is the third sample quantity corresponding to E v , v = 1 …… b, b is the number of intermediate entity types, where is the largest integer not exceeding (α × f), and α is a preset percentage.
[0046] Specifically, the intermediate entity type is the sample entity type corresponding to the second sample quantities that are ranked in the top α in descending order of the second sample quantities obtained from the sample entity type list.
[0047] Further, the third sample quantity is the first sample quantity corresponding to the intermediate entity type obtained from the first sample quantity list.
[0048] Specifically, the value range of α is 80% - 90%. Among them, those skilled in the art know that α can be selected according to actual needs, and all are within the protection scope of the present invention, which will not be elaborated here.
[0049] S5. According to E 0 , obtain the fourth sample quantity list E 1 = {E 1 1 , ……, E 1 u , ……, E 1 w}, where E 1 u is the u-th fourth sample quantity, u = 1 …… w, w = b. Among them, the fourth sample quantity is the third sample quantity after re - sorting the third sample quantities in E 0 in ascending order.
[0050] S6. When there is E 1 in E 1 u - E 1 u-1 ≠ 1, then α = α - α 0 , and repeat S4 - S5 until the preset loop termination condition is met. Among them, the preset loop termination condition is: E 1 u - E 1 u-1 = 1, where α 0 is the preset parameter threshold, and E 1 u-1 is the (u - 1)-th fourth sample quantity.
[0051] Specifically, the value range of α 0 is 0.1 - 0.5. Among them, those skilled in the art know that α 0 can be selected according to actual needs, and all are within the protection scope of the present invention, which will not be elaborated here.
[0052] S7. When E 1 u - E 1u-1 When it is equal to 1, obtain η 1 = b.
[0053] Based on the relevant sample data involved in the sample entities in the sample medical record text set, continuously adjust the parameters of the preset CNN model, so as to obtain the feature vectors of the corresponding dimensions of the intermediate medical record text. Determine the type of convolution kernel in the preset CNN model based on the number of text characters corresponding to the entities in the sample medical record text, ensure that the corresponding size of the convolution kernel is coherent, and improve the accuracy of obtaining the entities and entity labels in the medical record text.
[0054] S30223, according to η 1 , obtain the second intermediate parameter η corresponding to the preset CNN model 2 , where the second intermediate parameter η 2 meets the following conditions:
[0055] m / η 1 < η 2 ≤ (m × μ) / η 1 and η 2 = a × M, where μ is a preset parameter, a is the number of attention heads in the preset Transformer, and M is any positive integer.
[0056] Specifically, the second intermediate parameter η 2 is the number of convolution kernels of each type in the preset CNN model.
[0057] Specifically, the value range of μ is 5 to 302. Those skilled in the art know that μ can be selected according to actual needs, and all fall within the protection scope of the present invention, which will not be elaborated here.
[0058] Based on the above, set the second intermediate parameter corresponding to the preset CNN model within the corresponding range, avoiding the situation that the dimension of the feature vector corresponding to the intermediate medical record text output by the model is small due to too small parameters, resulting in the inability to obtain comprehensive information and making it difficult to find a good solution during training, making the model difficult to converge. At the same time, avoid the situation that the dimension of the feature vector corresponding to the intermediate medical record text output by the model is large due to too large model parameters, resulting in a decrease in the operating efficiency of the model, so that the accuracy of obtaining the entities and entity labels in the medical record text is relatively high.
[0059] S30225, according to η 1 and η 2 , obtain the dimension s of the intermediate feature vector, where the dimension s of the intermediate feature vector meets the following conditions:
[0060] s = 2 × η 1 × η 2 .
[0061] S3023, input B into a preset Transformer model to obtain a list of intermediate feature vectors C = {C 1 , ……, C i , ……, C n} corresponding to the intermediate medical record text. C i = (C i1 , ……, C ih , ……, C ig ), where C i is the intermediate feature vector corresponding to the i-th character in the intermediate medical record text, and C ih is the bit value at the h-th position in C i . h = 1 …… g, and g is the dimension of the intermediate feature vector. Among them, g meets the following conditions:
[0062] g = (N + 1) × 4, where N is the number of preset entity labels.
[0063] S3024, according to C, obtain the intermediate entity list corresponding to the intermediate medical record text and the intermediate label list corresponding to the intermediate entity list. Among them, the intermediate entity list includes several intermediate entities, and the intermediate entity is an entity in the intermediate medical record text recognized by combining a preset CNN model and a preset Transformer model. The intermediate label list includes several intermediate labels, and the intermediate label is a preset entity label classified by the intermediate entity.
[0064] Specifically, those skilled in the art know that any method in the prior art for combining a CNN model and a Transformer model to obtain entities and entity - corresponding labels in a text falls within the protection scope of the present invention and will not be elaborated here.
[0065] S303, according to the intermediate entity list and the intermediate entity label list, obtain the first intermediate entity relationship graph corresponding to the intermediate medical record text.
[0066] Specifically, in S303, the first intermediate entity relationship graph is obtained through the following steps:
[0067] S3031, input the intermediate entity list and the intermediate label list into a preset neural network model to obtain the intermediate entity relationship graph corresponding to the intermediate medical record text. The intermediate entity relationship graph is a tree - shaped structure graph of the relationship between intermediate entities constructed based on the intermediate entity list. The intermediate entity relationship graph includes an intermediate root node, several intermediate child nodes, and intermediate edges.
[0068] Specifically, in S3031, the intermediate entity relationship graph is obtained through the following steps:
[0069] S30311, when the intermediate entity in the intermediate entity list does not depend on any intermediate entity other than this intermediate entity in the intermediate entity list, connect the intermediate child node corresponding to this intermediate entity to the intermediate root node, and the connection direction between the intermediate child node corresponding to this intermediate entity and the intermediate root node is from the intermediate root node to the intermediate child node corresponding to this intermediate entity.
[0070] Specifically, the intermediate root node is a self-set node. Among them, those skilled in the art know that any method of using symbols to represent nodes in the prior art falls within the protection scope of the present invention and will not be elaborated here. For example, -1 is used to represent the intermediate root node.
[0071] Specifically, the intermediate child node includes an intermediate entity and an intermediate label corresponding to the intermediate entity.
[0072] Specifically, the intermediate edge represents the connection relationship between intermediate entities.
[0073] S30313, when the intermediate entity in the intermediate entity list depends on a certain intermediate entity other than this intermediate entity in the intermediate entity list, connect the intermediate child node corresponding to this intermediate entity to the intermediate child node corresponding to the dependent intermediate entity.
[0074] S30315, based on the connection between the intermediate root node and the intermediate child node and the connection between the intermediate child nodes to form intermediate edges, to obtain the first intermediate entity relationship graph.
[0075] Specifically, in S3031, the preset neural network model is obtained through the following steps:
[0076] S11, according to the key medical record text set, obtain the key entity list set U = {U 1 , ……, U d , ……, U z} corresponding to the key medical record text set and the key label list set corresponding to the key entity list set, where U d is the d-th key entity list, d = 1 …… z, and z is the number of key entity lists.
[0077] Specifically, the key medical record text set includes several key medical record texts, and the key medical record texts are medical record texts used to train the initial neural network model.
[0078] Specifically, the key label list set includes several key label lists.
[0079] Furthermore, each key medical record text corresponds to a key entity list, and each key entity list corresponds to a key label list.
[0080] Further, the method for obtaining the key entity list is the same as that for obtaining the intermediate entity list in S100, and the method for obtaining the key label list is the same as that for obtaining the intermediate label list in S100. Reference can be made to S3021 to S3024.
[0081] S12. Input U and the key label list set into the first neural network model to obtain the key entity relationship graph corresponding to U.
[0082] Specifically, the key entity relationship graph is a tree structure graph of the relationships between key entities constructed based on the key entity list.
[0083] Further, the method for obtaining the key entity relationship graph is the same as that for obtaining the intermediate entity relationship graph. Reference can be made to S30311 to S30315.
[0084] Specifically, those skilled in the art know that the neural network can be selected according to actual needs, all of which fall within the protection scope of the present invention and will not be elaborated herein.
[0085] S13. According to the key relationship graph, obtain the key parameter ζ corresponding to the first neural network model, where ζ meets the following conditions:
[0086] σ 0 d is the number of key entities that are the same and the relationships between the two key entities are the same in the key entity relationship graph corresponding to the d-th key medical record text and the true entity relationship graph corresponding to the d-th key medical record text, and is the number of relationships formed by two key entities in the true entity relationship graph corresponding to the d-th key medical record text.
[0087] S14. Continuously adjust the key parameter ζ until ζ ≤ ζ 0 to obtain the preset neural network model, where ζ 0 is the preset key parameter threshold.
[0088] Specifically, the value range of ζ 0 is 0.1 to 0.3. Those skilled in the art know that ζ 0 can be selected according to actual needs, all of which fall within the protection scope of the present invention and will not be elaborated herein.
[0089] S3033, process the intermediate entity relationship diagram based on a preset rule list to obtain a first intermediate entity relationship diagram. Among them, the preset rule list includes several preset rules, and the preset rules are rules set based on the semantic information of the intermediate entity and the information included in the intermediate label corresponding to the intermediate entity. The first intermediate entity relationship diagram is an intermediate entity relationship diagram obtained by determining the connection direction between the intermediate sub-nodes corresponding to the intermediate entities in the intermediate entity relationship diagram and the dependent intermediate entities based on the preset rule list on the basis of the intermediate entity relationship diagram.
[0090] Specifically, for example: classify the intermediate labels, and different types of intermediate labels have different priority orders. The priority order is: attribute < auxiliary class < location class < detection content class < pathological manifestation class < treatment class < diagnosis class < examination method class. The lower-priority ones depend on the higher-priority ones. That is, when the priority corresponding to an intermediate label is lower than the priority corresponding to another intermediate label, the intermediate sub-node corresponding to this intermediate label points to the intermediate sub-node corresponding to the other intermediate label in the intermediate entity relationship diagram.
[0091] As described above, by combining multiple models to obtain entities in the medical record text and continuously adjusting the parameters of each model based on sample data, the accuracy of obtaining entities in the medical record text is improved. Furthermore, the accuracy of obtaining entity relationships is relatively high. By combining the models with rules and presenting them in the form of a diagram, the efficiency of obtaining entity relationships in the medical record text is improved. First, use the models and then use the rules set based on feature information such as the semantic information of the entities in the medical record text to process the entities in the medical record text, so that the accuracy of the obtained medical record text entity relationship diagram is relatively high.
[0092] S304, according to the intermediate entity list, obtain the second intermediate entity relationship list corresponding to the intermediate medical record text.
[0093] Specifically, the system also includes β preset entity relationship labels. Among them, the preset entity relationship labels are labels for the relationships between entities set in advance. The labels for the relationships between entities are labels for the corresponding relationships between words representing the user's physical state. Those skilled in the art know that the preset entity relationship labels can be set according to actual needs, and all fall within the protection scope of the present invention and will not be elaborated here. For example: entity relationship labels such as occurrence location and confirmation relationship.
[0094] In a specific embodiment, the second intermediate entity relationship list is obtained in S304 through the following steps:
[0095] S3041, according to the candidate medical record text set, obtain the first candidate label list set corresponding to the candidate medical record text set. Among them, the candidate medical record text set includes several candidate medical record texts.
[0096] Specifically, the candidate medical record text set includes a number of candidate medical record texts, where the candidate medical record text is a medical record text with a preset entity relationship label between entities.
[0097] Specifically, the first candidate label list set includes a number of first candidate label lists. Each candidate medical record text corresponds to a first candidate label list. The first candidate label list includes a number of first candidate labels, and the first candidate label is the preset entity relationship label corresponding to two entities that actually exist in each candidate medical record text.
[0098] S3042. According to the candidate medical record text set, obtain the candidate entity list set corresponding to the candidate medical record text set, where the candidate entity list set includes a number of candidate entity lists, and the candidate entity list includes a number of candidate entities.
[0099] Specifically, the obtaining method of the candidate entity is the same as the obtaining method of the above intermediate entity, and steps S3021 to S3024 can be referred to.
[0100] S3043. Input the candidate entity list set into the initial neural network model to obtain the second candidate label list set, where the second candidate label list set includes a number of second candidate label lists, the second candidate label list includes a number of second candidate labels, and the second candidate label is the preset entity relationship label corresponding to two candidate entities in each candidate entity list obtained based on the initial neural network model.
[0101] S3044. According to the first candidate label list set and the second candidate label list set, obtain the target parameter θ corresponding to the initial neural network model.
[0102] Specifically, in S3044, θ is obtained through the following steps:
[0103] S30441. According to the first candidate label list set, obtain the first target quantity list Y = {Y 1 , ……, Y α , ……, Y β} corresponding to the preset entity relationship label list. Y α is the first target quantity corresponding to the α-th preset entity relationship label, α = 1 …… β, where the first target quantity is the quantity of each preset entity relationship label included in the first candidate label list set.
[0104] S30443. According to the second candidate label list set, obtain the second target quantity list R = {R 1 , ……, R α , ……, Rβ}, R α is the second target quantity corresponding to the α-th preset entity relationship label, where the second target quantity is the quantity of each preset entity relationship label included in the second candidate label list set.
[0105] S30445. According to the sample medical record text set, obtain the third target quantity list P = {P 1 , ……, P α , ……, P β}, P α is the third target quantity corresponding to the α-th preset entity relationship label, where the third target quantity is the quantity of each preset entity relationship label included in the sample medical record text, the sample medical record text set includes several sample medical record texts, and the sample medical record text is a medical record text for annotating entities.
[0106] S30447. According to Y, R, and P, obtain θ, where θ meets the following conditions:
[0107]
[0108] S3045. Continuously adjust the target parameter θ until the preset target condition is met to obtain the target entity relationship model, where the preset target condition is: θ ≤ θ 0 , θ 0 is the preset target parameter threshold.
[0109] Specifically, the value range of θ 0 is 0.1 to 0.3. Those skilled in the art know that θ 0 can be selected according to actual needs, and all fall within the protection scope of the present invention, which will not be elaborated here.
[0110] Specifically, the window corresponding to the target entity relationship model is obtained based on the number of text characters corresponding to the intermediate entities in the obtained intermediate entity list.
[0111] S3046. Input the intermediate entity list into the target entity relationship model to obtain the second intermediate entity relationship list corresponding to the intermediate medical record text, where the second intermediate entity relationship list includes several second intermediate entity relationships, and the second intermediate entity relationship includes the preset entity relationship label matched based on the target entity relationship model and the two intermediate entities corresponding to the matched preset entity label.
[0112] In another specific embodiment, the second intermediate entity relationship list is obtained through the following steps in S304:
[0113] S341. Obtain a set of candidate entity lists corresponding to the candidate medical record text set according to the candidate medical record text set, where the set of candidate entity lists includes a number of candidate entity lists, and each candidate entity list includes a number of candidate entities.
[0114] Specifically, the method for obtaining the candidate entities is the same as that for obtaining the candidate entities in the above embodiments.
[0115] S342. Input the set of candidate entity lists into the initial neural network model to obtain a second set of candidate label lists and a second set of candidate priorities G = {G 1 , ……, G x , ……, G p}, where G x is the second candidate priority corresponding to the x-th second candidate label, x = 1 …… p, and p is the number of second candidate priorities.
[0116] Specifically, the second set of candidate label lists includes a number of second candidate label lists, each second candidate label list includes a number of second candidate labels, and the second candidate labels are preset entity relationship labels corresponding to two candidate entities existing in each candidate entity list obtained based on the initial neural network model.
[0117] Furthermore, the second candidate priority is the matching degree of the relationship between the candidate entities obtained based on the initial neural network model and the preset entity relationship label.
[0118] S343. When ε / p ≥ F 0 , adopt the first processing method to obtain the key parameter γ corresponding to the initial neural network model, where ε is the number of second candidate priorities in G that are not 1, and F 0 is a preset priority threshold.
[0119] Specifically, the value range of F 0 is 0.5 - 0.7. Those skilled in the art know that F 0 can be selected according to actual needs, and all are within the protection scope of the present invention, which will not be elaborated here.
[0120] Specifically, in S343, γ is obtained through the following steps:
[0121] S3431. Obtain a first set of candidate label lists corresponding to the candidate medical record text set according to the candidate medical record text set, where the candidate medical record text set includes a number of candidate medical record texts.
[0122] Specifically, the candidate medical record text set includes a number of candidate medical record texts, where the candidate medical record texts are medical record texts in which there are preset entity relationship labels between entities.
[0123] Specifically, the first candidate label list set includes a number of first candidate label lists. Among them, each candidate medical record text corresponds to a first candidate label list, and the first candidate label list includes a number of first candidate labels. The first candidate labels are preset entity relationship labels corresponding to two entities actually existing in each candidate medical record text.
[0124] Specifically, the preset entity relationship label is a label for the relationship between entities set in advance. Among them, the label for the relationship between entities is a label for the corresponding relationship between words representing the user's physical state. Those skilled in the art know that the preset entity relationship label can be set according to actual needs, and all fall within the protection scope of the present invention, so no further elaboration is provided here. For example, entity relationship labels such as occurrence location and diagnosis relationship.
[0125] S3432. Input the candidate entity list set into the initial neural network model to obtain a second candidate label list set. Among them, the second candidate label list set includes a number of second candidate label lists, and the second candidate label list includes a number of second candidate labels. The second candidate labels are preset entity relationship labels corresponding to two candidate entities in each candidate entity list obtained based on the initial neural network model.
[0126] Specifically, those skilled in the art know that any method for training a model based on a neural network model to obtain labels in the prior art falls within the protection scope of the present invention, so no further elaboration is provided here. For example, neural network models such as the pcnn model.
[0127] S3433. According to the first candidate label list set, obtain the first target quantity list Y = {Y 1 , ……, Y α , ……, Y β} corresponding to the preset entity relationship label list. Y α is the first target quantity corresponding to the α-th preset entity relationship label, where α = 1 …… β. Among them, the first target quantity is the quantity of each preset entity relationship label included in the first candidate label list set.
[0128] S3434. According to the second candidate label list set, obtain the second target quantity list R = {R 1 , ……, R α , ……, R β} corresponding to the preset entity relationship label list. R α is the second target quantity corresponding to the α-th preset entity relationship label. Among them, the second target quantity is the quantity of each preset entity relationship label included in the second candidate label list set.
[0129] S3435. Obtain γ according to Y and R, where γ meets the following conditions:
[0130]
[0131] S344. When ε / p < F 0 Adopt the second processing method to obtain the key parameter γ corresponding to the initial neural network model.
[0132] Specifically, the obtaining method of γ in step S344 is the same as the obtaining method of θ in the above embodiment.
[0133] S345. Continuously adjust the key parameter γ until the preset target condition is met to obtain the target entity relationship model, where the preset target condition is: γ ≤ γ 0 γ 0 is the preset key parameter threshold.
[0134] S346. Input the intermediate entity list into the target entity relationship model to obtain the second intermediate entity relationship list corresponding to the intermediate medical record text, where the second intermediate entity relationship list includes several second intermediate entity relationships, and the second intermediate entity relationship includes a preset entity relationship label matched based on the target entity relationship model and two intermediate entities corresponding to this preset entity label.
[0135] As described above, entities in the medical record text are obtained by combining multiple models, and the parameters of each model are continuously adjusted based on sample data, which improves the accuracy of obtaining entities in the medical record text. Furthermore, the accuracy of obtaining entity relationships is relatively high. Based on the matching situation during the model training process, different methods are used to adjust the parameters of the model, which improves the accuracy of obtaining entity relationships in the medical record text.
[0136] S305. According to the first intermediate entity relationship graph and the second intermediate entity relationship list, obtain the target triple list corresponding to the intermediate medical record text. The target triple list includes several target triples, where the target triple includes two connected intermediate entities obtained from the first intermediate entity relationship graph and the connection relationship between these two intermediate entities, and two intermediate entities obtained from the second intermediate entity relationship list and the entity relationship label corresponding to these two intermediate entities.
[0137] S306. When the target medical record text label is the same as the second medical record text label, obtain the target triple list corresponding to the candidate medical record text. The target triple list includes several target triples, and the target triple includes the trigger word corresponding to the event obtained from the candidate medical record text based on the event extraction model and the two arguments corresponding to this trigger word.
[0138] Specifically, those skilled in the art are aware that the method of extracting any event from a text by any event extraction model in the prior art falls within the protection scope of the present invention and will not be elaborated herein.
[0139] Specifically, the argument is an element participating in the occurrence of an event and is composed of entities in the candidate medical record text.
[0140] As described above, different methods are used to obtain events in the medical record text based on the text type corresponding to the medical record text, so that the accuracy of the events in the obtained medical record text is relatively high. At the same time, multiple models are combined to obtain entities in the medical record text, and the parameters of each model are continuously adjusted based on sample data, improving the accuracy of the entities obtained in the medical record text. Furthermore, the accuracy of the obtained entity relationships is relatively high. The trained model is used to obtain the entity relationships in the medical record text, and different methods are used to adjust the parameters of the model based on the matching situation during the model training process, improving the accuracy of the entity relationships obtained in the medical record text, and thus improving the accuracy of the events obtained in the medical record text.
[0141] S400. Obtain a disease database according to a preset intermediate rule list and an intermediate event knowledge graph, where the disease database includes the content data included in the processing nodes involved in the disease.
[0142] Specifically, the preset intermediate rule list includes several preset intermediate rules, where the preset intermediate rule is a rule for judging entities in the intermediate event knowledge graph based on the characteristics of the disease.
[0143] As described above, by combining the NLP system with the rules, the data in the medical record text can be understood and parsed more accurately, improving the accuracy and efficiency of obtaining the disease database. The NLP system can provide the understanding and analysis of the medical record text data, providing richer and more accurate information for the rule engine, making the accuracy of the obtained disease database relatively high.
[0144] A system for obtaining a disease database based on an NLP system provided in this embodiment. The system includes: a processor and a memory storing a computer program. When the computer program is executed by the processor, the following steps are implemented: obtaining an initial medical record text set, obtaining an intermediate medical record text set according to the initial medical record text set, inputting the intermediate medical record text set into a preset NLP system to obtain an intermediate event knowledge graph corresponding to the intermediate medical record text set, and obtaining a disease database according to a preset intermediate rule list and the intermediate event knowledge graph. As mentioned above, by combining the NLP system with the rules, the data in the medical record text can be understood and analyzed more accurately, improving the accuracy and efficiency of obtaining the disease database. The NLP system can provide an understanding and analysis of the medical record text data, providing richer and more accurate information for the rule engine, resulting in a relatively high accuracy of the obtained disease database. At the same time, based on the text type corresponding to the medical record text, different methods are used to obtain the knowledge graph, and multiple models are combined to obtain the entities in the medical record text. The parameters of each model are continuously adjusted based on the sample data, improving the accuracy of obtaining the entities in the medical record text, and thus resulting in a relatively high accuracy of the obtained knowledge graph. The trained model is used to obtain the entity relationships in the medical record text, and different methods are used to adjust the parameters of the model based on the matching situation during the model training process, improving the accuracy of obtaining the entity relationships in the medical record text, and thereby resulting in a relatively high accuracy of the obtained disease database.
[0145] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. A system for acquiring a disease database based on an NLP system, characterized in that: The system comprises: a processor and a memory storing a computer program, and when the computer program is executed by the processor, the following steps are implemented: S100, obtaining an initial medical record text set, wherein the initial medical record text set includes a plurality of initial medical record texts, and the initial medical record texts are medical record texts obtained from a preset database; S200, obtaining an intermediate medical record text set according to the initial medical record text set, wherein the intermediate medical record text set includes a plurality of intermediate medical record texts, and the intermediate medical record texts are initial medical record texts corresponding to each user obtained from the initial medical record text set; S300, inputting the intermediate medical record text set into the preset NLP system, obtaining the intermediate event knowledge graph corresponding to the intermediate medical record text set, wherein the intermediate event knowledge graph includes a plurality of target triples, and obtaining the target triples in S300 through the following steps: S301, obtaining a target medical record text label corresponding to the intermediate medical record text, wherein the target medical record text label is a text type corresponding to the intermediate medical record text obtained based on the content and structure corresponding to a plurality of titles in the intermediate medical record text; S302, when the target medical record text label is consistent with the first medical record text label, obtaining an intermediate entity list corresponding to the intermediate medical record text and an intermediate entity label list corresponding to the intermediate entity list; S303, acquiring a first intermediate entity relationship graph corresponding to the intermediate medical record text according to the intermediate entity list and the intermediate entity label list; S304, obtaining a second intermediate entity relationship list corresponding to the intermediate medical record text according to the intermediate entity list; S305, according to the first intermediate entity relationship graph and the second intermediate entity relationship list, obtaining a target triple list corresponding to the intermediate medical record text, wherein the target triple list includes a plurality of target triples, wherein the target triple includes two connected intermediate entities obtained from the first intermediate entity relationship graph and the connection relationship corresponding to the two intermediate entities and two intermediate entities obtained from the second intermediate entity relationship list and the entity relationship labels corresponding to the two intermediate entities; S306, when the target medical record text label is consistent with the second medical record text label, obtain a target triple list corresponding to the medical record text to be selected, wherein the target triple list includes a plurality of target triples, and the target triples include a trigger word corresponding to an event obtained from the medical record text to be selected based on the event extraction model and two arguments corresponding to the trigger word; S400, obtaining a disease database according to a preset intermediate rule list and an intermediate event knowledge graph, wherein the disease database includes content data included in processing nodes involved in the disease.
2. The system for acquiring a disease database based on an NLP system according to claim 1, characterized in that: The first medical record text label is the type of medical record text involved in the process node of treating the user when the user has physical abnormalities.
3. The system for acquiring a disease database based on an NLP system according to claim 1, characterized in that: The second medical record text label is a type of medical record text corresponding to a feature of a user that is different from the user and is irrelevant to a process node involved in treating the user when the user has a physical abnormality.
4. The system for acquiring a disease database based on an NLP system according to claim 1, characterized in that: S302 also includes the following steps: S3021, obtain the initial feature vector list A corresponding to the intermediate medical record text = {A1, ..., A i , ..., A n }, A i =(A i1 , ..., A ij , ..., A im ), A i is the initial feature vector corresponding to the i-th character in the intermediate medical record text, A ij A i The j-th bit value in , j = 1 ... m, m is the dimension of the initial feature vector, i = 1 ... n, n is the number of characters in the intermediate medical record text; S3022, input A into the preset CNN model, and obtain the intermediate feature vector list B corresponding to the intermediate medical record text = {B1, ..., B i , ..., B n }, B i =(B i1 , ..., B ir , ..., B is ), B i is the intermediate feature vector corresponding to the i-th character in the intermediate medical record text, B ir For B i The rth bit value in , r = 1...s, s is the dimension of the intermediate feature vector, S3023, input B into the preset Transformer model, and obtain the intermediate feature vector list C={C1, ..., C i , ..., C n }, C i =(C i1 , ..., C ih , ..., C ig ), C i is the intermediate feature vector corresponding to the i-th character in the intermediate medical record text, C ih C i The bit value of the hth bit in , h = 1...g, g is the dimension of the intermediate feature vector, where g meets the following conditions: g=(N+1)×4, N is the number of preset entity tags; S3024, according to C, obtain an intermediate entity list corresponding to the intermediate medical record text and an intermediate label list corresponding to the intermediate entity list, wherein the intermediate entity list includes a plurality of intermediate entities, the intermediate entities are entities in the intermediate medical record text identified by combining a preset CNN model and a preset Transformer model, and the intermediate label list includes a plurality of intermediate labels, the intermediate labels are preset entity labels to which the intermediate entities are classified.
5. The system for acquiring a disease database based on an NLP system according to claim 1, characterized in that: In S304, the second intermediate entity relationship list is obtained through the following steps: S3041, obtaining a first candidate label list set corresponding to the candidate medical record text set according to the candidate medical record text set, wherein the candidate medical record text set includes a plurality of candidate medical record texts; S3042, according to the candidate medical record text set, obtaining a candidate entity list set corresponding to the candidate medical record text set, wherein the candidate entity list set includes a plurality of candidate entity lists, and the candidate entity list includes a plurality of candidate entities; S3043, inputting the candidate entity list set into the initial neural network model to obtain a second candidate label list set, wherein the second candidate label list set includes a plurality of second candidate label lists, the second candidate label lists include a plurality of second candidate labels, and the second candidate labels are preset entity relationship labels corresponding to two candidate entities in each candidate entity list obtained based on the initial neural network model; S3044, according to the first candidate label list set and the second candidate label list set, obtain the target parameter θ corresponding to the initial neural network model, wherein θ is obtained in S3044 by the following steps: S30441, according to the first candidate label list set, obtain a first target quantity list Y={Y1, ..., Y α , ..., Y β },Y α is the first target number corresponding to the αth preset entity relationship tag, α=1…β, wherein the first target number is the number of each preset entity relationship tag included in the first candidate tag list set; S30443, according to the second candidate label list set, obtain a second target quantity list R={R1, ..., R α , ..., R β },R α is the second target number corresponding to the αth preset entity relationship tag, wherein the second target number is the number of each preset entity relationship tag included in the second candidate tag list set; S30445, according to the sample medical record text set, obtain a third target quantity list P={P1, ..., P α , ..., P β },P α is the third target number corresponding to the αth preset entity relationship label, wherein the third target number is the number of each preset entity relationship label included in the sample medical record text, the sample medical record text set includes a plurality of sample medical record texts, and the sample medical record texts are medical record texts used to annotate entities; S30447, obtain θ according to Y, R and P, where θ meets the following conditions: ; S3045, continuously adjusting the target parameter θ until a preset target condition is met to obtain a target entity relationship model, wherein the preset target condition is: θ≤θ 0 ,θ 0 is the preset target parameter threshold; S3046, input the intermediate entity list into the target entity relationship model, and obtain the second intermediate entity relationship list corresponding to the intermediate medical record text, wherein the second intermediate entity relationship list includes a plurality of second intermediate entity relationships, and the second intermediate entity relationship includes a preset entity relationship label matched based on the target entity relationship model and two intermediate entities corresponding to the preset entity label matched.
Citation Information
Patent Citations
Text recognition method and device, equipment and storage medium
CN111814472A
Auxiliary disease differential diagnosis system based on causality-containing medical knowledge graph
CN113871003A