Medical entity recognition method, device and equipment and storage medium
By using two target models to identify non-pharmacological treatment behaviors, the problem of complex identification and limited accuracy in existing technologies is solved, achieving efficient and accurate identification of non-pharmacological treatment behavior categories, supporting doctors' diagnosis and treatment.
Patent Information
- Application Number
- CN202210554276.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-20
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-05-20
AI Technical Summary
Existing methods for classifying non-pharmacological treatment behaviors are complex and have limited accuracy, typically relying on manual annotation to summarize identification rules.
Two target models (a first target model and a second target model) are used to identify non-pharmaceutical medical entities. The first target model is used to identify non-pharmaceutical medical entities, and the second target model is used to identify their categories. The identification accuracy is improved by preprocessing and training data.
It enables accurate identification of non-pharmacological treatment behavior categories, reduces identification complexity, provides an effective data foundation, and supports doctors in diagnosing and treating patients.
Smart Images

Figure CN114944211B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of medical entity recognition, and in particular to a method, apparatus, device and storage medium for medical entity recognition. Background Technology
[0002] In the field of medical technology, hospital clinical data centers (CDRs) record clinical data for each patient, such as medical orders, medical records, laboratory data, electrocardiograms, ultrasounds, and pathology data. Regarding a patient's condition, the clinical data records information such as whether treatment was administered through medication, non-pharmacological treatments, or a combination of both. Clinically, non-pharmacological treatments include various categories such as surgery, procedures, and radiotherapy. In related technologies, the common approach to identifying the category of non-pharmacological treatments involves: based on medical rules, medical personnel manually annotate the medical data; based on the annotation results, feasible identification rules are derived; and based on these rules, the category of non-pharmacological treatment is identified. This identification method is relatively complex and has limited accuracy. Summary of the Invention
[0003] This disclosure provides a method, apparatus, device, and storage medium for medical entity identification, to at least solve the above-mentioned technical problems existing in the prior art.
[0004] According to a first aspect of this disclosure, a method for identifying medical entities is provided, the method comprising:
[0005] Acquire target medical data;
[0006] The target medical data is input into a first target model to obtain the desired data output by the first target model; the desired data is the non-pharmaceutical medical entity in the target medical data.
[0007] The desired data is input into the second target model to obtain the categories of non-pharmaceutical medical entities in the desired data output by the second target model;
[0008] The categories of non-pharmaceutical medical entities include surgery, procedures, and radiotherapy.
[0009] In the above scheme, the expected data includes the full-word data of non-pharmaceutical medical entities in the target medical data;
[0010] The desired data is input into the second target model to obtain the categories of non-drug medical behaviors described by the desired data output by the second target model, including:
[0011] The full-word data and the segmented data of the full-word data are input into the second target model to obtain the category of the non-drug medical entity output by the second target model.
[0012] In the above scheme, the expected data is obtained by the first target model identifying non-pharmaceutical medical entities in the target medical data based on the first reference information and the second reference information;
[0013] The first reference information is the fields that constitute non-pharmaceutical medical entities in medical data and their field types; the second reference information is used to characterize the sorting of fields of different field types.
[0014] In the above scheme, the second target model is obtained by training a preset model using target training data;
[0015] The target training data includes at least full-word training text for non-pharmaceutical medical entities and word-segmented training text for the full-word training text.
[0016] In the above scheme, the whole-word training text and the word-segmented training text are obtained by the first target model based on the identification of non-pharmaceutical medical entities in the predetermined medical text.
[0017] In the above scheme, acquiring the target medical data includes:
[0018] Acquire medical data to be identified;
[0019] The medical data to be identified is preprocessed to obtain the target medical data.
[0020] According to a second aspect of this disclosure, a medical entity recognition device is provided, the device comprising:
[0021] The first acquisition unit is used to acquire target medical data;
[0022] The first identification unit is used to input the target medical data into the first target model to obtain the expected data output by the first target model; the expected data is a non-drug medical entity in the target medical data.
[0023] The second identification unit is used to input the expected data into the second target model to obtain the category of non-pharmaceutical medical entities in the expected data output by the second target model;
[0024] The categories of non-pharmacological medical procedures include surgery, manipulation, and radiotherapy.
[0025] In the above scheme, the expected data is obtained by the first target model identifying non-pharmaceutical medical entities in the target medical data based on the first reference information and the second reference information;
[0026] The first reference information is a description of the fields that constitute the non-pharmaceutical medical entity in the medical data and their field types; the second reference information is used to characterize the sorting of fields of different field types.
[0027] According to a third aspect of this disclosure, an electronic device is provided, comprising:
[0028] At least one processor; and
[0029] A memory communicatively connected to the at least one processor; wherein,
[0030] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods described in this disclosure.
[0031] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the methods described in this disclosure.
[0032] The medical entity recognition method, apparatus, device, and storage medium disclosed herein, compared with the recognition of non-drug treatment behavior categories based on recognition rules in related technologies, can accurately identify the category of non-drug treatment entities without summarizing recognition rules based on manual annotation results and based on two target models (a first target model and a second target model), thereby reducing the complexity of recognition.
[0033] The identification of data described as non-pharmacological medical behaviors (non-pharmacological medical entities) in target medical data can provide an effective data foundation for the classification of non-pharmacological treatment behaviors (non-pharmacological medical entities), ensuring the accuracy of category identification.
[0034] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0035] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which:
[0036] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts.
[0037] Figure 1 This illustration shows the implementation flow of the medical entity recognition method according to an embodiment of the present disclosure. Figure 1 ;
[0038] Figure 2 This illustration shows the implementation flow of the medical entity recognition method according to an embodiment of the present disclosure. Figure 2 ;
[0039] Figure 3 A diagram illustrating the recognition mechanism implemented using two target models according to an embodiment of this disclosure is shown.
[0040] Figure 4 A schematic diagram of training a second target model according to an embodiment of this disclosure is shown;
[0041] Figure 5 A schematic diagram of the composition of a medical entity recognition device according to an embodiment of this disclosure is shown. Figure 1 ;
[0042] Figure 6 A schematic diagram of the composition of a medical entity recognition device according to an embodiment of this disclosure is shown. Figure 1 ;
[0043] Figure 7 A schematic diagram of the composition structure of an electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation
[0044] To make the objectives, features, and advantages of this disclosure more apparent and understandable, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0045] Before describing the technical solution of this disclosure, the technical terms that may be involved will be explained first.
[0046] 1) Medical Data: This includes medical orders, medical records, laboratory test results, ECG data, ultrasound data, pathology data, and other data generated related to a patient's illness. Considering the reference value of medical data, it can be stored for future reference. For example, it can be stored in a CorelDRAW (Card Receiver) and retrieved or accessed when needed.
[0047] 2) Medical behavior: In practical application, medical behavior can be divided into drug medical behavior and non-drug medical behavior.
[0048] In this scheme, any words, phrases, or sentences in medical data that can reflect non-pharmacological medical behaviors can be considered as non-pharmacological medical entities. The categories of non-pharmacological medical entities can be considered as categories of non-pharmacological treatment behaviors.
[0049] Considering that non-pharmacological treatments include categories such as surgery, procedures, and radiotherapy, this disclosed solution aims to achieve the identification of non-pharmacological medical entity categories through a low-complexity method. Specifically, compared to related technologies, it eliminates the need to derive identification rules based on manual annotation results; instead, it achieves accurate identification of non-pharmacological medical entity categories based on two target models (a first target model and a second target model).
[0050] Both the first and second objective models in this scheme exhibit strong stability and robustness. The first objective model can accurately identify non-pharmaceutical medical entities in the target medical data. The second objective model, based on the non-pharmaceutical medical entities identified by the first objective model, can accurately classify these entities. The identification of non-pharmaceutical medical entities in the target medical data provides an effective guarantee for accurately classifying these entities.
[0051] This solution enables the identification of non-pharmaceutical medical entities and their categories, and can also be seen as enabling the identification of non-pharmaceutical treatment behaviors and their categories.
[0052] That is, this disclosure provides a category recognition scheme with high accuracy and low complexity, and the recognition accuracy is high.
[0053] The medical entity identification method disclosed herein can be applied to any reasonable device, such as a server or terminal. The server can be a regular server, a cloud server, etc. The terminal includes mobile phones, tablets, and self-service terminals such as hospital self-service terminals.
[0054] Figure 1 This illustration shows the implementation flow of the medical entity recognition method according to an embodiment of the present disclosure. Figure 1 .like Figure 1 As shown, the method includes:
[0055] S(Step) 101: Obtain target medical data;
[0056] In this step, medical data can be any medical data that requires identification of non-pharmacological treatment behaviors. This includes data such as medical orders, medical records, laboratory test results, ECG data, ultrasound data, and pathology data stored in the CDR.
[0057] In practical applications, if medical data that requires identification of non-pharmacological treatment behaviors is considered as medical data to be identified, then the target medical data can be the medical data to be identified, which can be obtained by calling or reading the data stored in the CDR.
[0058] The target medical data can also be obtained by preprocessing the medical data to be identified. Based on this, S101 can be implemented through the following scheme: acquiring the medical data to be identified; preprocessing the medical data to be identified to obtain the target medical data.
[0059] The process involves retrieving or reading data stored in the CDR (Catalogue Record). It's understood that in practical applications, medical data such as medical records and lab reports may contain English names for diseases, or multiple different names for the same disease. The preprocessing in this solution includes, but is not limited to, format and / or formal standardization. Format standardization includes converting English disease names to Chinese names and unifying the use of Chinese or English punctuation. Formal standardization includes unifying multiple different names for the same disease into a single name and correcting spelling errors, all to facilitate subsequent identification.
[0060] The target medical data in this disclosure can be text data. For example, text recorded in medical records, text in laboratory images, and text in ultrasound images.
[0061] S102: Input the target medical data into the first target model to obtain the expected data output by the first target model, wherein the expected data is the non-drug medical entity in the target medical data;
[0062] In this step, any data in the medical data that reflects medical behavior as surgery, operation, or radiotherapy can be considered as expected data.
[0063] For example, data in the target medical data that can be described as treatment using surgical, operational, or radiotherapy methods, such as "tumor resection", "cyst removal", or "radiotherapy", can be considered as non-pharmaceutical medical entities.
[0064] The first target model can identify data in the target medical data that describes the medical treatment as being performed through surgery, manipulation, or radiation therapy.
[0065] The first target model can be any model capable of identifying non-pharmaceutical medical entities from the target medical data. Examples include neural network models, statistical models, and Aho-Corasick automata.
[0066] In practical applications, neural network models, statistical models, and Aho-Corasick automata exhibit strong stability and robustness. Identification of desired data within target medical data based on these models ensures the accuracy of identifying non-pharmaceutical medical entities within the target medical data. This, in turn, provides assurance for the accurate identification of the categories of non-pharmaceutical medical entities.
[0067] Furthermore, using the first target model to identify the desired data is highly feasible and has low complexity.
[0068] S103: Input the desired data into the second target model to obtain the categories of non-drug medical entities in the desired data output by the second target model; wherein, the categories of non-drug medical entities include surgery, manipulation and radiotherapy.
[0069] The second target model in this step is able to identify the categories of non-pharmaceutical medical entities in the desired data.
[0070] For example, if the expected data identified by the first target model is "lung tumor resection", then the second target model identifies the non-pharmacological medical entity as the category of surgery for that expected data. If the expected data identified by the first target model is "skin cyst removal", then the second target model identifies the non-pharmacological medical entity as the category of operation for that expected data.
[0071] In practical implementation, the second objective model can be any model capable of identifying categories of non-pharmacological treatment behaviors, such as neural network models or statistical models. Given the strong stability and robustness of the second objective model, class identification based on it can guarantee accuracy.
[0072] In S101–S103, both the first and second target models exhibit strong stability and robustness. Using the first target model to identify non-pharmaceutical medical entities in the target medical data ensures the accuracy of the identification of the desired data. Then, using the second target model to identify the categories of non-pharmaceutical medical entities ensures the accuracy of category identification. Compared to related technologies that rely on rule-based identification of non-pharmaceutical treatment behaviors, this approach eliminates the need to derive identification rules from manually labeled results. Using two target models (the first and second target models) not only effectively reduces the complexity of identification but also achieves accurate identification of the categories of non-pharmaceutical treatment behaviors.
[0073] Furthermore, the data identified in the target medical data, namely the expected data, is data described as non-pharmacological treatment behaviors. The identification of data described as non-pharmacological treatment behaviors in the target medical data can provide an effective and direct data basis for the classification of non-pharmacological treatment behaviors, and can ensure the accuracy of the classification of non-pharmacological treatment behaviors.
[0074] Accurate identification of non-pharmacological medical behaviors can provide strong support for decision-makers such as physicians in further diagnosis and treatment of patients. Furthermore, using a second-objective model for category identification is highly feasible and has low complexity.
[0075] In this disclosed solution, a first reference information and a second reference information are pre-set. The first reference information describes the fields that constitute non-pharmaceutical medical entities in medical data and their field types; the second reference information is used to characterize the sorting of fields of different field types.
[0076] In the scheme of inputting target medical data into a first target model to obtain the desired data output by the first target model, the principle for obtaining the desired data is as follows: the first target model can identify non-pharmaceutical medical entities in the target medical data based on first and second reference information to obtain the desired data. The accuracy of the identification of the desired data is guaranteed by using these two reference information. Please refer to the subsequent related descriptions for detailed explanations.
[0077] As an optional implementation, the desired data includes the full-word data of non-pharmaceutical medical entities in the target medical data. Thus, the aforementioned scheme of inputting the desired data into the second target model to obtain the category of medical behavior described by the desired data output by the second target model can be implemented in the following way: Figure 2 As shown, S103 becomes S103': Input the full word data and the segmented data of the full word data into the second target model to obtain the category of non-drug medical entities output by the second target model.
[0078] For example, "lung tumor resection" and "leg cyst removal" in the target medical data can be used as full-word data representing non-drug treatment entities. Among them, "lung," "tumor," and "resection" are three word segmentation data for the full-word data of "lung tumor resection." Similarly, "leg," "cyst," and "removal" can be used as three word segmentation data for the full-word data of "leg cyst removal."
[0079] The full-word data can be obtained by identifying non-pharmaceutical medical entities in the target medical data using the first target model. Word segmentation data can be obtained in two ways: first, by segmenting the full-word data; second, by identifying the individual word segments that make up the full-word data using the first target model. In the second method of obtaining word segmentation data, it's equivalent to both the full-word data and the word segmentation data being obtained through the identification process of the first target model. Obtaining both types of data (full-word data and word segmentation data) simultaneously using the same target model can significantly improve data acquisition efficiency and shorten the time required.
[0080] The second target model is a scheme for identifying the categories of non-pharmaceutical medical entities based on full-word data and word segmentation data. It is highly feasible and has low complexity. Please refer to the subsequent related descriptions for details.
[0081] It is understood that in this embodiment of the disclosure, the second target model is obtained by training a preset model using target training data; wherein, the target training data includes full-word training text representing non-pharmaceutical medical entities and word-segmented training text of the full-word training text.
[0082] The preset model can be any reasonable model from neural network models or statistical models, such as the Bayesian model in statistical models. Texts representing non-pharmaceutical medical entities, such as "lung tumor resection" and "leg cyst removal," can all be used as whole-word training texts. "Lung," "tumor," and "resection" can be used as word segmentation training texts for the whole-word training text "lung tumor resection." Similarly, "leg," "cyst," and "removal" can be used as word segmentation training texts for the whole-word training text "leg cyst removal."
[0083] It is understandable that the above whole-word training texts and word-segmentation training texts are merely specific examples.
[0084] In practical applications, considering the accuracy of training the pre-defined model, the training text should be diverse, covering any possible diseases and medical procedures requiring non-pharmacological treatment. This allows for the training of a more accurate secondary target model, thereby ensuring accurate category recognition.
[0085] In practical applications, target training data can be obtained through pre-setting. For example, the full-word training text and segmented training text that can be used as target training data can be pre-recorded, and the pre-recorded information can be read when needed. Alternatively, the full-word training text can be pre-recorded, and when needed, it can be read and segmented into words to obtain the segmented training text.
[0086] In practical applications, besides the methods mentioned above for obtaining target training data, target training data can also be obtained in the following way: Both the full-word training text and the segmented training text in the target training data are obtained by the first target model based on the identification of non-pharmaceutical medical entities in a predetermined medical text. That is, the text describing non-pharmaceutical medical behaviors identified by the first target model in the predetermined medical text is used as the full-word training text. The segmented text obtained by segmenting the full-word text is used as the segmented training text, or the text obtained by identifying each segmented text constituting the full-word text is used as the segmented training text. The predetermined medical text can be any medical data stored in CDR.
[0087] In practical applications, the target training data includes at least full-word training text that describes non-pharmacological medical behaviors, as well as segmented training text of the full-word training text. In simpler terms, in addition to the aforementioned full-word and segmented training texts, the target training data also includes the label data of the full-word training text. This label data indicates the category of the non-pharmacological treatment behavior described by the full-word training text.
[0088] In the aforementioned scheme, the method of obtaining (target) training data of the second target model based on the first target model can obtain training data simply and directly, reducing the complexity of obtaining training data.
[0089] The full-word data and full-word training text in the aforementioned scheme can be the same data or different data, depending on the actual use case.
[0090] The following is combined with Figure 3 and Figure 4 The technical solutions of the embodiments of this disclosure will be further described in detail below.
[0091] In this embodiment, the medical entity recognition method of this disclosure is implemented by applying a first target model and a second target model. The first target model is an Aucma automata model. The second target model is a Bayesian model in statistical modeling, specifically a trained or fully trained Bayesian model. It can be understood that the second target model is a model obtained by training a preset model, such as a (to-be-trained) Bayesian model, using target training data. Specifically, the second target model is a Bayesian model that has been trained or fully trained from the Bayesian model to be trained.
[0092] To better understand this disclosure, let’s first look at how the first and second reference information were obtained.
[0093] It is understood that in practical applications, non-pharmacological treatments mainly include surgical treatments and procedural treatments. This embodiment uses surgery and procedural procedures as examples of non-pharmacological treatments.
[0094] In practical applications, descriptions of surgeries or procedures in medical texts generally include the surgical method / operation, surgical site / operation approach, surgical lesion / operation, and surgical instruments. Based on this, the design of the first reference information—that is, the fields in the medical text describing non-pharmacological medical procedures and their types—mainly includes:
[0095] (1) Core words: words in medical texts that describe surgical methods / operations, such as excision, removal, and excision.
[0096] (2) Anatomical (location) terms: These are terms used in medical texts to describe body parts that are being treated; for example, heart, left kidney, eye, etc.
[0097] (3) Disease term / clinical discovery term: the target of treatment, that is, what disease has occurred in a part of the body, such as tumor or cyst; or what clinical symptoms have occurred in a part of the body, such as swelling or inflammation.
[0098] (4) Descriptive words: words that modify the surgery / operation, such as entering the surgical / operation area through the spleen, tumor size, sutures required, etc.
[0099] Due to space limitations, it is impossible to list all fields under the above field types. Any reasonable field and field type falls within the protection scope of this disclosure.
[0100] It is understandable that, in the medical field, fields from the aforementioned field types can constitute non-pharmacological treatment entities in medical text. For example, the anatomical location term "lung," the lesion term "tumor," and the core term "resection" constitute the non-pharmacological treatment entity "lung tumor resection." This entity can describe the medical procedure appearing in the target medical text as a non-pharmacological treatment.
[0101] In practical applications, primary reference information can be recorded in the form of a thesaurus. Fields represented as core terms, anatomical terms, etc., are recorded in the thesaurus according to certain rules. For example, the thesaurus includes field types and field content. The types include core terms, anatomical (location) terms, lesion / clinical finding terms, descriptive terms, etc. The field content under each field type includes fields that may appear in that field category in medical text. For example, the field content under the core term field type includes, but is not limited to, terms such as resection, extraction, and removal.
[0102] It is understood that the first reference information can be used to identify the fields in the target medical data that can constitute textual data describing non-pharmacological medical behaviors. For these fields to constitute textual descriptions of non-pharmacological medical behaviors, their order must conform to certain sorting rules to align with medical language. In this embodiment, the predefined sorting rules are: anatomical terms can be followed by lesion terms or core terms; lesion terms can be followed by core terms or descriptive terms; and descriptive terms can be followed by anatomical terms. This sorting rule can be considered as the second reference information, used to characterize the ordering of fields of different field types.
[0103] The design or setting of the second reference information is to constrain the positional ordering relationship between the fields identified based on the first reference information. For example, assuming that the fields identified by the first target model based on the first reference information include "lung" (anatomical location term), "tumor" (lesion term), and "resection" (core term), the text (expected data) composed of these three fields, as indicated by the second reference information, is "lung tumor resection".
[0104] In practical applications, the first reference information and the second reference information are designed or set so that the AC automaton model can accurately identify the expected data.
[0105] Combination Figure 3 As shown, the scheme for implementing the medical entity recognition method of this disclosure using AC automata and a trained Bayesian model includes:
[0106] It should be noted that both the medical data to be identified and the target medical data are text data. For ease of description, the medical data to be identified will be referred to as the medical text to be identified, and the target medical data will be referred to as the target medical text.
[0107] S301: Read or retrieve medical text (hereinafter referred to as the medical text to be identified) from the CDR that requires category identification of non-pharmacological treatment behaviors.
[0108] In practical applications, patient A's historical medical records can be retrieved from the CDR, and the text data recorded in the medical records can be used as the medical text to be identified.
[0109] For example, the patient A's medical history:
[0110] Patient name: A;
[0111] Gender: Male;
[0112] Age: 40;
[0113] An ultrasound scan in March 2021 revealed a shadow in the lungs, which was later confirmed to be cancer after further lung imaging analysis. The patient is scheduled to undergo cancer removal surgery at our hospital on April 2nd.
[0114] S302: Preprocess the medical text to be identified to obtain the target medical text.
[0115] In this step, the word "cancer" in two instances in the medical record will be changed to "tumor". Please refer to the aforementioned instructions on pretreatment for details.
[0116] S303: Input the target medical text into the AC automaton model, and the AC automaton model will identify the full-word data of non-drug medical entities in the target medical text based on the first reference information and the second reference information.
[0117] In practical applications, the aforementioned sorting principles can be represented as state transition relationships in the AC automaton model. These state transition relationships include, but are not limited to, the following: the state of an anatomical word can transition to the state of a lesion word or a core word; the state of a lesion word can transition to the state of a core word or a descriptive word; and the state of a descriptive word can transition to the state of an anatomical word.
[0118] For example, the fields identified from the target medical text by the AC automaton model based on the first reference information include the anatomical term "lung," the lesion term "tumor," and the core term "resection." Based on the aforementioned state transition relationship, the anatomical term can be transferred to the lesion term, and the lesion term can be transferred to the core term. Therefore, the aforementioned fields identified from the target medical text constitute the expected data "lung tumor resection" that conforms to medical language.
[0119] In practical applications, when dealing with target medical text, the AC automata model may produce results indicating that the expected data has been identified, or it may produce results indicating that the expected data has not been identified, depending on the actual content of the target medical text.
[0120] The AC automaton model can produce the aforementioned recognition results through at least one round of recognition of the target medical text.
[0121] It is understandable that the AC automaton identifies each character or word in the target medical text in the order of their appearance.
[0122] Taking the first round of recognition using the AC automaton model as an example, text recognition is performed on the first character or word appearing in the target medical text. This character or word serves as the starting word for the first round of recognition. According to the vocabulary, the fields and their types are identified. Based on the state transition relationship, it is determined whether the character or word identified as belonging to that field type can be transferred to the character or word following it in the target medical text (the second character or word in the target medical text). If it is identified as transferable, it is then determined whether the character or word identified as belonging to that field type (i.e., the second character or word in the target medical text) can be transferred to the character or word following it in the target medical text (the third character or word), and so on.
[0123] It should be noted that, upon identifying a character or word and its field type, if the state transition relationship records a transition relationship where the word can be transferred to a subsequent character or word, then the character or word is identified as transferable to the target medical text where it appears as a subsequent character or word. If the state transition relationship does not record a transition relationship between the character or word and a subsequent character or word (in the target medical text), or if a transition relationship between the character or word and a subsequent character or word is not permitted, then the character or word is identified as not transferable to the target medical text where it appears as a subsequent character or word, and the identification process stops.
[0124] If the fields constituting non-drug medical behaviors can be selected from the identified characters or words in the target medical text before recognition stops, the desired data can be obtained.
[0125] If, before stopping recognition, it is impossible to select the fields constituting non-drug medical behavior from the identified characters or words in the target medical text, then a second round of recognition will be performed starting from the character or word that stopped recognition, using that character or word as the starting word for the remaining characters or words in the target medical text. The process for the second round of recognition is the same as described above for the first round of recognition; repeated parts will not be repeated. This process continues until a result that either identifies the desired data or does not provide the desired data is given.
[0126] Specifically, if among the identified characters or words there is a core word representing the category of surgery or operation, or a core word and an anatomical location word, or a core word and a lesion word, or a core word, an anatomical location word, and a lesion word, then it is considered that the fields constituting non-drug medical behavior can be selected. Otherwise, it is considered that the fields constituting non-drug medical behavior cannot be selected.
[0127] Specifically, taking the target medical text "lung tumor resection surgery" as an example, the AC automaton model identifies the expected data as follows: In the first round of identification, "surgery" is identified as a descriptive word, followed by the anatomical location word "lung." The state transition relationship records the relationship that the descriptive word can be transferred to the anatomical location word, and identification continues. Following "lung" is identified as the lesion word "tumor." The state transition relationship records the relationship that the anatomical location word can be transferred to the lesion word, and identification continues. Following "tumor" is identified as the core word "resection surgery." The state transition relationship records the relationship that the anatomical location word can be transferred to the core word, and the text ends, stopping identification. Among the identified fields, the text "lung tumor resection surgery," which includes the anatomical location word, lesion word, and core word, is taken as the expected data.
[0128] In practical applications, the AC automaton model can identify the desired data in the first round of recognition, and also in the second or third round of recognition. The technical solution disclosed herein designs the AC automaton model to perform at least one round of recognition, which can greatly avoid the problem of missing the desired data caused by performing only one round of recognition, thereby achieving accurate recognition of the desired data appearing in the target medical text.
[0129] If the AC automaton model can identify the desired data, then the desired data output by the AC automaton model can be obtained by one or at least two rounds of identification according to the order of occurrence of each word or phrase in the target medical text, the vocabulary, and the state transition relationships. See the aforementioned relevant explanations for details.
[0130] The AC automaton model identifies fields and their types based on the aforementioned first reference information. Constrained by state transition relationships, and utilizing its built-in tree structure and pointers, it can accurately identify text data in target medical texts describing non-pharmacological treatment behaviors. The tree structure and pointers of the AC automaton model can be found in the model's documentation and will not be elaborated upon here.
[0131] The medical record data of the aforementioned patient A, after being identified by the AC automata model, can identify the text data of non-drug treatment behavior in the target medical text, that is, the expected data is "lung tumor resection".
[0132] In the example above, the expected data for the whole word is "lung tumor resection"; the word segmentation data for the whole word is "lung", "tumor" and "resection".
[0133] Another example, taking the target medical text "The patient's family indicated that coronary angiography should be performed first, followed by stent implantation if necessary," the AC automata model identifies the following: characters 8 to 13 of the text, "coronary angiography," are the full-word data, while characters 17 to 22, "stent implantation," are another full-word data. The word segmentation data for the full-word data "coronary angiography" consists of "coronary artery" (anatomical location) and "angiography" (core word). The full-word data "stent implantation," as the core word, can be considered as its word segmentation data.
[0134] S304: Input the full word data and its segmented data into a trained or completed Bayesian model, which will then provide the category of the non-drug treatment behavior described by the full word data.
[0135] Among them, the trained or completed Bayesian model uses Bayesian statistical principles to determine the category of the non-pharmacological treatment behavior described by the full-word data. The following formulas (1) to (3) are used in the Bayesian statistical principles:
[0136] P(Ci|B)=P(A1,A2,...An|Ci)*P(Ci) / P(A1,A2,...An) (1)
[0137] In formula (1), Ci represents the category of non-pharmacological treatment behavior, which includes but is not limited to surgery, procedures, or others (such as radiotherapy). B represents the whole word data. A1, A2...An are the segmented data that constitute the whole word data B, and n is a positive integer greater than 1. P(Ci|B) is the probability that the medical behavior described by the whole word data belongs to category Ci.
[0138] Where P(Ci) represents the probability of a full-word data item of category Ci appearing in the target training data of the Bayesian model, calculated as the number of labeled full-word data items of category Ci divided by the total number of full-word data items. In the target training data of the Bayesian model, the category of the medical behavior described by each full-word data item was manually labeled; the labeled category can be considered as the label data.
[0139] Assuming that each word segmentation is independent, then we have
[0140] P(A1,A2,...An|Ci)=P(A1|Ci)*P(A2|Ci)*...*P(An|Ci) (2)
[0141] P(A1,A2,...An)=P(A1)*P(A2)*...*P(An) (3)
[0142] Then, P(A1|Ci) is the probability that the segmented data A1 belongs to the Ci category. It is calculated by dividing the number of full-word data of type Ci containing segmented A1 by the total number of full-word data of type Ci.
[0143] P(A1) represents the probability of A1 appearing in the segmented data. It is calculated by dividing the number of full-word data containing A1 in the labeled full-word data by the total number of full-word data.
[0144] Substituting Formula 2 and Formula 3 into Formula 1, we obtain the following calculation formula (4).
[0145] P(Ci|B)=P(A1|Ci)*P(A2|Ci)*...*P(An|Ci)*P(Ci) / (P(A1)*P(A2)*...*P(An)); (4)
[0146] Based on the aforementioned formula, the probability P(Ci|B) of the medical behavior described by the full-word data B as each category Ci is calculated. The probabilities of each category are compared, and the largest Ci is selected. The category corresponding to the largest Ci is the category of the medical behavior described by the full-word data B.
[0147] For example, based on Bayesian statistical principles, the Bayesian statistical model calculates the probability that the non-pharmacological treatment described by the full-word data "lung tumor resection" is surgery, the probability that it is a procedure, and the probability that it is another type of treatment such as radiotherapy, using the aforementioned formula. If the probability of surgery is greater than the probabilities of procedure and radiotherapy, the non-pharmacological treatment described by the full-word data is determined to be a surgery. If the probability of procedure is greater than the probabilities of surgery and radiotherapy, the non-pharmacological treatment described by the full-word data is determined to be a procedure. This achieves the identification of the category of non-pharmacological treatment.
[0148] The preceding content uses the calculation of the probability of a non-pharmacological treatment described by the full-word data being surgery, a procedure, or another type such as radiotherapy as an example. It is also possible to calculate only the probability of a non-pharmacological treatment described by the full-word data being surgery and a procedure. If the probability of surgery is greater than the probability of a procedure, the non-pharmacological treatment described by the full-word data is determined to be a surgery. Otherwise, it is determined to be a procedure.
[0149] Figure 3 The target category refers to the category of non-pharmaceutical medical entities in the expected data, that is, which category among several non-pharmaceutical treatment behaviors the medical behavior described by the expected data belongs to.
[0150] The aforementioned scheme not only identifies the text data describing non-pharmacological treatment behaviors from the target medical text (a coarse-grained identification), but also identifies the specific categories of the non-pharmacological treatment behaviors described by the text data (a fine-grained identification). This is equivalent to performing overall identification of the text data describing non-pharmacological treatment behaviors and detailed identification of the categories of non-pharmacological treatment behaviors, making it a comprehensive and novel identification scheme.
[0151] Among them, the identification of non-pharmacological treatment behaviors in target medical texts provides an effective guarantee for accurately identifying the categories of non-pharmacological treatment behaviors.
[0152] Furthermore, the identification based on two models (the first target model and the second target model) reduces the complexity of identification compared to the identification of non-drug treatment behavior categories based on identification rules in related technologies. This is because it eliminates the need to derive identification rules from manual annotation results.
[0153] Furthermore, since both models are robust and stable, category recognition based on these two robust and stable models can guarantee accuracy.
[0154] In summary, the embodiments of this disclosure provide a category recognition scheme with low complexity and high accuracy.
[0155] Identifying categories of non-pharmacological treatment behaviors can support decision-makers, such as physicians, in making further diagnoses and treatments for patients. For example, given that patient A has a history of surgery for lung tumors, and considering the long interval between that surgery and their most recent consultation, coronary stent surgery could be considered.
[0156] The identification of non-pharmacological treatment behaviors disclosed herein can display the identification results to doctors, eliminating the need for doctors to manually review patients' medical history and treatment records, thus improving doctors' consultation efficiency.
[0157] Combination Figure 4 The training process of the second target model, such as the Bayesian model, is illustrated below.
[0158] S401: Read or retrieve medical text from a CDR;
[0159] Considering that diverse training data during training will lead to more accurate model training, various types of medical texts, such as medical records, ultrasound images, electrocardiograms, and laboratory images, can be read or retrieved from the CDR. Multiple medical texts of the same type can also be read or retrieved.
[0160] S402: Preprocess the medical text to obtain the preprocessed medical text;
[0161] Please refer to the aforementioned instructions for the preprocessing procedure; repeated details will not be elaborated upon.
[0162] The preprocessed medical text can be used as the pre-selected medical text.
[0163] S403: Input the preprocessed medical text into the AC automaton model, and the AC automaton model will identify and output the full word data and word segmentation data of non-drug medical entities in each preprocessed medical text;
[0164] The process of the AC automaton model in this step to identify whole word data and word segmentation data is similar to the related aspects mentioned above, and will not be repeated here.
[0165] The full-word data output by the AC automaton model is used as the full-word training text, and the segmented data is used as the segmented training text. The categories of non-pharmacological treatment behaviors described in the full-word training text are manually labeled to obtain the label data for the full-word training text. The full-word training text, the segmented training text, and the label data of the full-word training text are used as the target training data for training the preset model.
[0166] For example, the label data for the full-word training text "lung tumor resection" is categorized as "surgery".
[0167] S404: Input the target training data into the preset model to train the preset model using the target training data, so as to obtain a trained or completed Bayesian model.
[0168] It's understandable that the preset model contains initial model parameters, and training the preset model aims to adjust these initial parameters to obtain more stable parameters. The preset model is iteratively computed using the target training data, with the loss function calculated in each iteration. When the loss function of the preset model is less than or equal to a preset value, the model parameters tend to stabilize, and training of the preset model stops.
[0169] The loss function is a function relating to the actual category of the non-pharmacological treatment behavior described in the full-word training text (derived by the pre-defined model based on its own input) and the label data of the full-word training text. For example, if the actual category and label data are used to construct the mean squared error function, the constructed mean squared error function can be used as the loss function.
[0170] For details regarding the model parameters and loss function of the Bayesian model, please refer to the relevant explanations.
[0171] A Bayesian model whose model parameters tend to be stable is a well-trained or fully trained Bayesian model, which can be used as the aforementioned second target model to perform the medical entity recognition method of this disclosure.
[0172] The above description describes a scheme for obtaining (target) training data for a preset model based on the AC automata model. This approach allows for simple and direct acquisition of training data, reducing the complexity of obtaining training data, improving efficiency, and shortening the time required.
[0173] Furthermore, using the output of the AC automaton model as training data for the pre-defined model enriches and diversifies the training data, thereby making the trained model more accurate. An accurate model ensures accurate category identification.
[0174] This disclosure also provides an embodiment of a medical entity recognition device, such as... Figure 5 As shown, the device includes:
[0175] The first acquisition unit 501 is used to acquire target medical data;
[0176] The first identification unit 502 is used to input the target medical data into the first target model to obtain the expected data output by the first target model; the expected data is a non-drug medical entity in the target medical data;
[0177] The second identification unit 503 is used to input the expected data into the second target model to obtain the category of non-pharmaceutical medical entities in the expected data output by the second target model;
[0178] The categories of non-pharmacological medical procedures include surgery, manipulation, and radiotherapy.
[0179] In the aforementioned scheme, the second identification unit 502 is used for:
[0180] The desired data includes the full-word data of non-pharmaceutical medical entities in the target medical data; the full-word data and the word segmentation data of the full-word data are input into the second target model to obtain the category of the non-pharmaceutical medical entity output by the second target model.
[0181] In the aforementioned scheme, the desired data is obtained by the first target model identifying non-pharmaceutical medical entities in the target medical data based on the first reference information and the second reference information;
[0182] The first reference information is the fields that constitute non-pharmaceutical medical entities in medical data and their field types; the second reference information is used to characterize the sorting of fields of different field types.
[0183] In the aforementioned scheme, the second target model is obtained by training a preset model using target training data;
[0184] The target training data includes at least full-word training text for non-pharmaceutical medical entities and word-segmented training text for the full-word training text.
[0185] In the aforementioned scheme, the whole-word training text and the word-segmented training text are obtained by the first target model based on the identification of non-pharmaceutical medical entities in the predetermined medical text.
[0186] In the aforementioned scheme, such as Figure 6 As shown, the device further includes:
[0187] The second acquisition unit 504 is used to acquire the medical data to be identified;
[0188] The preprocessing unit 505 is used to preprocess the medical data to be identified to obtain the target medical data.
[0189] It should be noted that the medical entity recognition device in this application embodiment solves the problem in a similar way to the aforementioned medical entity recognition method. Therefore, the implementation process and implementation principle of the medical entity recognition device can be found in the description of the implementation process and implementation principle of the aforementioned method, and the repeated parts will not be repeated.
[0190] According to embodiments of this disclosure, this disclosure also provides an electronic device and a readable storage medium.
[0191] Figure 7A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0192] like Figure 7 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0193] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0194] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as medical entity recognition methods. For example, in some embodiments, the medical entity recognition method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the medical entity recognition method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the medical entity recognition method by any other suitable means (e.g., by means of firmware).
[0195] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0196] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0197] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0198] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0199] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0200] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0201] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0202] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.
[0203] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A method for medical entity recognition, characterized in that, The method includes: Acquire target medical data; The target medical data is input into a first target model to obtain the expected data output by the first target model; the expected data is the non-pharmaceutical medical entity in the target medical data; the first target model is an AC automaton. The expected data is obtained by the first target model identifying non-pharmaceutical medical entities in the target medical data based on first reference information and second reference information; wherein, the first reference information is the fields that can describe the non-pharmaceutical medical entities in the medical data and their field types. The fields and their types in the first reference information include: core words, anatomical words, lesion words, and descriptive words; the core words are words in the medical text that describe surgical methods or procedures, the anatomical words are words in the medical text that describe the body parts being treated, the lesion words are the targets of treatment, and the descriptive words are words that modify the surgery or procedure. The second reference information is represented as a state transition relationship in the AC automaton model, and the state transition relationship includes: the state of anatomical words transitioning to the state of lesion words or core words, the state of lesion words transitioning to the state of core words or descriptive words, and the state of descriptive words transitioning to the state of anatomical words. For each character or word in the target medical data, the AC automaton identifies them one by one in the order of their appearance. When a character or word and its field type are identified, if the state transition relationship records a transition relationship where the word can be transferred to a subsequent character or word, then the character or word is identified as transferable to a subsequent character or word in the target medical data. If the state transition relationship does not record a transition relationship between the character or word and a subsequent character or word, or if a transition relationship between the character or word and a subsequent character or word is not allowed, then the character or word is identified as not transferable to a subsequent character or word in the target medical data, and identification stops. Starting from the character or word where identification stopped, a second round of identification is performed on the remaining characters or words in the target medical data, using that character or word as the starting word. If a core word, or a core word and an anatomical word, or a core word and a lesion word, or a core word, an anatomical word, and a lesion word exist among the identified characters or words, it is believed that the fields constituting non-drug medical behavior can be selected to obtain the desired data. The desired data is input into the second target model to obtain the categories of non-pharmaceutical medical entities in the desired data output by the second target model; The categories of non-pharmaceutical medical entities include surgery, procedures, and radiotherapy.
2. The method according to claim 1, characterized in that, The expected data includes the full-word data of non-pharmaceutical medical entities in the target medical data; The desired data is input into the second target model to obtain the categories of non-pharmaceutical medical entities in the desired data output by the second target model, including: The full-word data and the segmented data of the full-word data are input into the second target model to obtain the category of the non-drug medical entity output by the second target model.
3. The method according to claim 1, characterized in that, The second target model is obtained by training a preset model using target training data; The target training data includes at least full-word training text for non-pharmaceutical medical entities and word-segmented training text for the full-word training text.
4. The method according to claim 3, characterized in that, The full-word training text and the word-segmented training text are obtained by the first target model based on the identification of non-pharmaceutical medical entities in a predetermined medical text.
5. The method according to any one of claims 1 to 4, characterized in that, The acquisition of the target medical data includes: Acquire medical data to be identified; The medical data to be identified is preprocessed to obtain the target medical data.
6. A medical entity recognition device, characterized in that, The device includes: The first acquisition unit is used to acquire target medical data; The first identification unit is used to input the target medical data into the first target model to obtain the expected data output by the first target model; the expected data is a non-drug medical entity in the target medical data; the first target model is an AC automaton. The expected data is obtained by the first target model identifying non-pharmaceutical medical entities in the target medical data based on first reference information and second reference information; wherein, the first reference information is the fields that can describe the non-pharmaceutical medical entities in the medical data and their field types. The fields and their types in the first reference information include: core words, anatomical words, lesion words, and descriptive words. The core words are words in the medical text that describe surgical methods or procedures. The anatomical words are words in the medical text that describe the body parts being treated. The lesion words are the targets of treatment. The descriptive words are words that modify the surgery or procedure. The second reference information is represented as a state transition relationship in the AC automaton model. The state transition relationship includes: the state of anatomical words transitioning to the state of lesion words or core words, the state of lesion words transitioning to the state of core words or descriptive words, and the state of descriptive words transitioning to the state of anatomical words. For each character or word in the target medical data, the AC automaton identifies them one by one in the order of their appearance. When a character or word and its field type are identified, if the state transition relationship records a transition relationship where the word can be transferred to a subsequent character or word, then the character or word is identified as transferable to a subsequent character or word in the target medical data. If the state transition relationship does not record a transition relationship between the character or word and a subsequent character or word, or if a transition relationship between the character or word and a subsequent character or word is not allowed, then the character or word is identified as not transferable to a subsequent character or word in the target medical data, and identification stops. Starting from the character or word where identification stopped, a second round of identification is performed on the remaining characters or words in the target medical data, using that character or word as the starting word. If a core word, or a core word and an anatomical word, or a core word and a lesion word, or a core word, an anatomical word, and a lesion word exist among the identified characters or words, it is believed that the fields constituting non-drug medical behavior can be selected to obtain the desired data. The second identification unit is used to input the expected data into the second target model to obtain the category of non-pharmaceutical medical entities in the expected data output by the second target model; The categories of non-pharmaceutical medical entities include surgery, procedures, and radiotherapy.
7. The apparatus according to claim 6, characterized in that, The expected data is obtained by the first target model identifying non-pharmaceutical medical entities in the target medical data based on the first reference information and the second reference information; The first reference information is the fields that constitute non-pharmaceutical medical entities in medical data and their field types; the second reference information is used to characterize the sorting of fields of different field types.
8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.
Citation Information
Patent Citations
Electronic medical record text named entity recognition method based on pre-trained language model
CN110705293A