A method, device, equipment and storage medium for pre-annotating medical corpus
By determining corpus attributes, deduplication and knowledge triplets for medical corpus data, and pre-labeling using historical annotation data, the problem of inefficient manual annotation of medical corpus data is solved, and a more efficient annotation process is achieved.
Patent Information
- Application Number
- CN202311045368.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-18
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-08-18
AI Technical Summary
The manual annotation work of medical corpus data is inefficient due to the large number of entities and relationships and the high repetition of types, and the prior art has failed to make specific optimizations for this.
By obtaining the medical corpus data to be marked, determining its corpus properties, and deduplication based on these properties, building a knowledge triple, determining historical annotation data with the same attributes from the preset annotation library for pre-labeling, reducing the workload of subsequent manual annotation.
Pre-marking reduces the workload of manual labeling, improves the efficiency of medical corpus data labeling, and reduces the complexity of labor-intensive work.
Smart Images

Figure CN117056520B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language, and particularly to a method, device, equipment and storage medium for pre-annotating medical corpus. Background Art
[0002] The prerequisite for artificial intelligence to exert its effectiveness is to construct good training data. In the field of natural language, most tasks are supervised learning problems that require a large number of samples. Tasks such as entity recognition and entity relationship extraction all require a large amount of good pre-annotated data for model training. Constructing a training data set through the method of manually annotating text data is a complex labor-intensive task. In the manual annotation of medical corpus data, due to the large number of corpus entities and relationships involved, and the high repetition rate of entity and relationship types, the conventional annotation work is more complex and repetitive.
[0003] In the prior art, tags are defined according to the annotation object and actual requirements, the entity types to be extracted are determined, and the relationship types between entities are defined according to the determined entity types. Then, after manually annotating the whole document, a directed graph structure is constructed with entities as vertices and entity relationships as edges. Finally, the directed graph structure forms a related graph structure according to the logical relationship of the article context. However, there are a large number of entities and relationships in medical corpus data, and the repetition rate of entity and relationship types within a single corpus is high. There is no technology to specifically optimize the manual annotation operation process for medical corpus data to improve the manual annotation efficiency. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method, device, equipment and storage medium for pre-annotating medical corpus, which can pre-annotate through historical annotation data after determining historical annotation data with the same attributes as the medical corpus data to be annotated, so as to reduce the workload of subsequent more accurate manual annotation. The specific scheme is as follows:
[0005] In the first aspect, the present application discloses a method for pre-annotating medical corpus, which is applied to a corpus annotation system and includes:
[0006] Obtain the medical corpus data to be annotated, and determine the corpus attribute of the medical corpus data to be annotated as the target corpus attribute based on the received corpus attribute determination instruction;
[0007] Perform data deduplication on the medical corpus data to be annotated based on the target corpus attribute to obtain deduplicated corpus data;
[0008] Determine the entity types and entity relationships of the deduplicated corpus data according to the target corpus attribute, and construct knowledge triples according to the entity types and the entity relationships;
[0009] Based on the knowledge triples, historical annotation data having attributes identical to those of the target corpus is determined from a preset annotation library, and the deduplicated corpus data is pre-annotated based on the historical annotation data.
[0010] Optionally, the acquiring of the medical corpus data to be annotated, and determining the corpus attribute of the medical corpus data to be annotated as the target corpus attribute based on the received corpus attribute determination instruction, includes:
[0011] Receiving a corpus data file, and determining a file format of the corpus data file;
[0012] If the file format of the corpus data file is an electronic document, extracting the medical corpus data to be annotated in the electronic document; if the file format of the corpus data file is a scanned file, performing an optical character recognition operation on the scanned file to determine the medical corpus data to be annotated;
[0013] A corpus attribute determination instruction sent by a user terminal is received, and a corpus attribute of the medical corpus data to be annotated and a sub-corpus attribute of the corpus attribute are determined based on the corpus attribute determination instruction, so as to determine the corpus attribute and the sub-corpus attribute as the target corpus attribute of the medical corpus data to be annotated.
[0014] Optionally, the deduplication of the medical corpus data to be annotated based on the target corpus attribute to obtain deduplication corpus data includes:
[0015] Determining a semicolon and a period in the medical corpus data to be annotated, and determining the determined semicolon and the period as separators;
[0016] Segmenting the medical corpus data to be annotated based on the separation symbol to obtain segmented corpus data;
[0017] Based on the target corpus attribute, duplicate corpus data in the segmented corpus data is determined, and the duplicate corpus data is removed to obtain deduplicated corpus data.
[0018] Optionally, determining the entity type and entity relationship of the deduplicated corpus data according to the target corpus attributes, and constructing knowledge triples according to the entity type and the entity relationship, includes:
[0019] Determine the entity type of the deduplicated corpus data based on the corpus attribute, and determine the sub-entity type of the deduplicated corpus data based on the sub-corpus attribute;
[0020] Determine the entity relationship between the entity type and the sub-entity type, and construct a knowledge triple of the deduplicated corpus data based on the entity type, the sub-entity type, and the entity relationship.
[0021] Optionally, the medical corpus pre-annotation method further includes:
[0022] If an error annotation modification instruction is received, delete the determined entity type in the deduplicated corpus data based on the error annotation modification instruction, and modify the entity type of the deduplicated corpus data based on the error annotation modification instruction.
[0023] Optionally, the step of determining historical annotation data with the same attributes as the target corpus attributes from a preset annotation library based on the knowledge triple and pre-annotating the deduplicated corpus data based on the historical annotation data includes:
[0024] Determine whether the pre-annotation rule is an entity exact match rule or an entity similarity match rule according to the received pre-annotation instruction;
[0025] If the received pre-annotation instruction indicates that the pre-annotation rule is the entity exact match rule, determine historical annotation data that exactly matches the entity type in the knowledge triple from the preset annotation library based on the target corpus attributes, and pre-annotate the deduplicated corpus data based on the historical annotation data;
[0026] If the received pre-annotation instruction indicates that the pre-annotation rule is the entity similarity match rule, determine historical similar annotation data with a character position distance difference not greater than a preset character position distance difference threshold from the preset annotation library based on the target corpus attributes, and pre-annotate the deduplicated corpus data based on the historical similar annotation data.
[0027] Optionally, the medical corpus pre-annotation method further includes:
[0028] If a annotation copy instruction is received, split the knowledge triple based on the annotation copy instruction to determine the entity type and the entity relationship in the knowledge triple;
[0029] Split the data to be copied based on the entity type and the entity relationship to determine first annotation data corresponding to the entity type and second annotation data corresponding to the entity relationship, so that the client can copy the first annotation data and the second annotation data.
[0030] In a second aspect, the present application discloses a medical corpus pre-annotation device, which is applied to a corpus annotation system and includes:
[0031] An attribute determination module, configured to obtain medical corpus data to be labeled, and determine the corpus attribute of the medical corpus data to be labeled as a target corpus attribute based on a received corpus attribute determination instruction;
[0032] A corpus deduplication module, configured to perform data deduplication on the medical corpus data to be labeled based on the target corpus attribute to obtain deduplicated corpus data;
[0033] A triple construction module, configured to determine the entity type and entity relationship of the deduplicated corpus data according to the target corpus attribute, and construct a knowledge triple according to the entity type and the entity relationship;
[0034] A corpus pre-annotation module, configured to determine historical annotation data with the same attribute as the target corpus attribute from a preset annotation library based on the knowledge triple, and perform pre-annotation on the deduplicated corpus data based on the historical annotation data.
[0035] In a third aspect, the present application discloses an electronic device, including:
[0036] A memory, configured to store a computer program;
[0037] A processor, configured to execute the computer program to implement the foregoing medical corpus pre-annotation method.
[0038] In a fourth aspect, the present application discloses a computer-readable storage medium, configured to store a computer program, and the computer program, when executed by a processor, implements the foregoing medical corpus pre-annotation method.
[0039] In this application, first, the medical corpus data to be annotated is obtained, and based on the received corpus attributes, an instruction is determined to set the corpus attributes of the medical corpus data to be annotated as target corpus attributes. Then, data deduplication is performed on the medical corpus data to be annotated based on the target corpus attributes to obtain the deduplicated corpus data. Next, the entity types and entity relationships of the deduplicated corpus data are determined according to the target corpus attributes, and knowledge triples are constructed based on the entity types and entity relationships. Finally, historical annotation data with attributes identical to the target corpus attributes is determined from a preset annotation library based on the knowledge triples, and the deduplicated corpus data is pre-annotated based on the historical annotation data. It can be seen that through the medical corpus pre-annotation method of this application, after successfully obtaining the medical corpus data to be annotated, the corpus attributes of the medical data to be annotated can be determined, data deduplication can be performed on the medical corpus data to be annotated according to the determined target corpus attributes, the entity types and entity relationships in the deduplicated corpus data can be determined, knowledge triples can be constructed based on the entity types and entity relationships, historical annotation data can be determined based on the constructed knowledge triples, and the deduplicated corpus data can be pre-annotated with the historical annotation data. In this way, after determining the historical annotation data with attributes identical to the medical corpus data to be annotated, pre-annotation can be performed using the historical annotation data to reduce the workload of subsequent more accurate manual annotation. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings according to the provided drawings without creative efforts.
[0041] Figure 1 It is a flowchart of a medical corpus pre-annotation method provided by this application;
[0042] Figure 2 It is a flowchart of a specific medical corpus pre-annotation method provided by this application;
[0043] Figure 3 It is a flowchart of another specific medical corpus pre-annotation method provided by this application;
[0044] Figure 4 It is a schematic diagram of a corpus annotation system provided by this application;
[0045] Figure 5 It is a schematic diagram of the structure of a medical corpus pre-annotation device provided by this application;
[0046] Figure 6A structural diagram of an electronic device provided by this application. Specific implementation manners
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] In the prior art, tags are defined according to the annotation objects and actual requirements, the entity types to be extracted are determined, and the relationship types between entities are defined according to the determined entity types. Then, after manually annotating the entire document, a directed graph structure is constructed with entities as vertices and entity relationships as edges. Finally, the directed graph structure forms a related graph structure according to the logical relationship of the article context. However, there are a large number of entities and relationship quantities in medical corpus data, and the type repetition degree of entities and relationships within a single corpus is high. There is no technology to specifically optimize the manual annotation operation process for medical corpus data to improve the manual annotation efficiency.
[0049] To solve the above technical problems, the purpose of the present invention is to provide a medical corpus pre-annotation method, device, equipment, and storage medium, which can perform pre-annotation through historical annotation data after determining historical annotation data with the same attributes as the medical corpus data to be annotated, so as to reduce the workload of subsequent more accurate manual annotation.
[0050] See Figure 1 As shown, an embodiment of the present invention discloses a medical corpus pre-annotation method, which is applied to a corpus annotation system and includes:
[0051] Step S11, obtain the medical corpus data to be annotated, and determine the corpus attribute of the medical corpus data to be annotated as the target corpus attribute based on the received corpus attribute determination instruction.
[0052] In this embodiment, the medical corpus data to be annotated is obtained, and the corpus attributes of the medical corpus data to be annotated are determined as the target corpus attributes based on the received corpus attribute determination instruction. That is, if the medical corpus data is to be annotated, the medical corpus data needs to be obtained first, and the way of receiving the medical corpus data can be to receive the uploaded corpus data file through the data acquisition module in the corpus annotation system. After receiving the uploaded corpus data file, the medical corpus data to be annotated in the corpus data file needs to be extracted. After determining the medical corpus data to be annotated, the data attribute customization module in the corpus annotation system can provide the medical corpus data to be annotated with an attribute customization function. It should be noted that each medical corpus data to be annotated can customize multiple data attributes, and each customized data attribute can have multiple sub-attribute levels. The internal corpus to be annotated can also be divided into different sub-modules to define data attributes respectively. In this way, the corpus attributes of the medical corpus data to be annotated can be customized by receiving the corpus attribute determination instruction sent by the user end, so as to define the corpus attributes of the medical corpus data to be annotated as the target corpus attributes corresponding to the received corpus attribute determination instruction.
[0053] Step S12: deduplicating the medical corpus data to be annotated based on the target corpus attributes to obtain deduplicated corpus data.
[0054] In this embodiment, the medical corpus data to be annotated is deduplicated based on the target corpus attributes to obtain deduplicated corpus data. That is, due to the professional characteristics of the medical corpus, there are a lot of identical contents within the corpus to be annotated and between corpora of the same type to be annotated. Therefore, the semicolon and period in the corpus to be annotated can be determined by the data deduplication module in the corpus annotation system, and the determined semicolon and period are determined as separators to segment the medical corpus data to obtain segmented corpus data. After determining the segmented corpus data, it is necessary to remove the duplicate data in the segmented corpus data based on the determined target corpus attributes, and only retain the corpus data that appears for the first time. In this way, the processing efficiency of the medical corpus pre-annotation method described in this application can be effectively improved by data deduplication.
[0055] Step S13: determining the entity type and entity relationship of the deduplicated corpus data according to the target corpus attributes, and constructing knowledge triples according to the entity type and the entity relationship.
[0056] In this embodiment, the entity type and entity relationship of the deduplicated corpus data are determined according to the target corpus attribute, and a knowledge triple is constructed based on the entity type and the entity relationship. That is, after performing deduplication on the to-be-annotated medical corpus data to obtain the deduplicated corpus data, it is necessary to customize the entity type and entity relationship type through the construction of an entity type and entity relationship type customization module, so as to realize the customization of the entity type and entity relationship type based on the data attributes of the to-be-annotated medical corpus and the actual annotation requirements, so as to adapt to the medical corpus annotation tasks for different types of corpus data, different clinical annotation tasks, and different model function training requirements. The entity type and entity relationship type can be divided into multiple levels. In manual annotation, the entity annotation task can perform entity annotation by selecting the corresponding content in the to-be-annotated corpus after selecting the entity type, and the relationship annotation task can perform directed relationship annotation by successively clicking on two related entities after selecting the relationship type.
[0057] Step S14, determine historical annotation data with attributes the same as those of the target corpus from the preset annotation library based on the knowledge triple, and pre-annotate the deduplicated corpus data based on the historical annotation data.
[0058] In this embodiment, historical annotation data with attributes the same as those of the target corpus is determined from the preset annotation library based on the knowledge triple, and the deduplicated corpus data is pre-annotated based on the historical annotation data. That is, there are a large number of repeated clinical professional terms of entity types and corresponding entity relationships among the to-be-annotated medical corpora with the same attributes. Therefore, the pre-annotation rule can be determined through the received pre-annotation instruction, and based on the pre-annotation rule and the determined knowledge triple, historical annotation data for annotating the to-be-annotated medical corpus data is determined from the preset annotation library established in advance for storing historical annotation data. It should be noted that the pre-annotation rule can be an entity full match rule or an entity similarity match rule. If it is an entity full match rule, the determined historical annotation data needs to be exactly the same as the entity type in the knowledge triple; if it is an entity similarity match rule, the determined historical annotation data needs to have a character position distance difference from the entity type in the knowledge triple that is not greater than the preset character position distance difference threshold, and the distance difference threshold can be set by the user according to the requirements of the user terminal.
[0059] It can be seen that in this embodiment, the medical corpus data to be annotated is first obtained, and the corpus attribute of the medical corpus data to be annotated is determined as the target corpus attribute based on the received corpus attribute determination instruction. The medical corpus data to be annotated is de-duplicated based on the target corpus attribute to obtain the de-duplicated corpus data. Then, the entity type and entity relationship of the de-duplicated corpus data are determined according to the target corpus attribute, and a knowledge triple is constructed based on the entity type and the entity relationship. Finally, the historical annotation data with the same attribute as the target corpus attribute is determined from the preset annotation library based on the knowledge triple, and the de-duplicated corpus data is pre-annotated based on the historical annotation data. It can be seen that through the medical corpus pre-annotation method described in this application, after successfully obtaining the medical corpus data to be annotated, the corpus attribute of the medical data to be annotated can be determined, the medical corpus data to be annotated can be de-duplicated according to the determined target corpus attribute, the entity type and entity relationship in the de-duplicated corpus data can be determined, a knowledge triple can be constructed based on the entity type and entity relationship, the historical annotation data can be determined based on the constructed knowledge triple, and the de-duplicated corpus data can be pre-annotated with the historical annotation data. In this way, after determining the historical annotation data with the same attribute as the medical corpus data to be annotated, pre-annotation can be performed through the historical annotation data to reduce the workload of subsequent more accurate manual annotation.
[0060] Based on the foregoing embodiments, it can be known that in this application, after receiving the corpus data file, the medical corpus data to be annotated in the file needs to be extracted, and the extracted medical corpus data to be annotated needs to be preprocessed and de-duplicated. For this reason, this embodiment details how to extract the medical corpus data to be annotated from the corpus file and how to de-duplicate the medical corpus data to be annotated. See Figure 2 As shown, an embodiment of the present invention discloses a medical corpus pre-annotation method, including:
[0061] Step S21, receive a corpus data file, and determine the file format of the corpus data file.
[0062] In this embodiment, a corpus data file is received, and the file format of the corpus data file is determined. That is to say, the uploaded corpus data file can be received through the data acquisition module in the corpus annotation system. It should be noted that the uploaded corpus data file can have different formats, which can be in the word document format or a scanned file recognized by OCR (optical character recognition). And according to the different file formats obtained, different processing methods will be used to extract the medical corpus data to be annotated in the file.
[0063] Step S22: If the file format of the corpus data file is an electronic document, extract the medical corpus data to be annotated from the electronic document; if the file format of the corpus data file is a scanned file, perform optical character recognition on the scanned file to determine the medical corpus data to be annotated.
[0064] In this embodiment, if the file format of the corpus data file is an electronic document, extract the medical corpus data to be annotated from the electronic document; if the file format of the corpus data file is a scanned file, perform optical character recognition on the scanned file to determine the medical corpus data to be annotated. That is, if it is determined that the file format of the received corpus data file is an electronic document such as a word electronic document or a hospital medical record document, the electronic document can be transmitted through the corresponding interface in the corpus annotation system, and the medical corpus data to be annotated in the electronic document can be extracted; if it is determined that the file format of the received corpus data file is a scanned file, it is necessary to determine the character data in the scanned file through OCR recognition to determine the medical corpus data to be annotated in the scanned file.
[0065] Step S23: Receive the corpus attribute determination instruction sent by the user terminal, and determine the corpus attribute of the medical corpus data to be annotated and the sub-corpus attributes of the corpus attribute based on the corpus attribute determination instruction, so as to determine the corpus attribute and the sub-corpus attribute as the target corpus attribute of the medical corpus data to be annotated.
[0066] In this embodiment, a corpus attribute determination instruction sent by the user terminal is received, and based on the corpus attribute determination instruction, the corpus attributes of the to-be-annotated medical corpus data and the sub-corpus attributes of the corpus attributes are determined, so as to determine the corpus attributes and the sub-corpus attributes as the target corpus attributes of the to-be-annotated medical corpus data. That is, as described in the foregoing embodiment, the to-be-annotated medical corpus can customize multiple data attributes, and each customized data attribute can have multiple sub-attribute levels. The to-be-annotated corpus can also be divided into different sub-modules to define data attributes respectively. And defining data attributes requires the user terminal to define, that is, by receiving the corpus attribute determination instruction of the user terminal, the corpus attributes of the to-be-annotated medical corpus data are defined, so as to define the corpus attributes of the to-be-annotated medical corpus data as the target corpus attributes corresponding to the received corpus attribute determination instruction. For example, the to-be-annotated medical corpus data can customize data attributes such as medical record data, drug instruction data, literature data, medical policy data, etc. based on the data type, and can also define data attributes based on the medical institutions where the data comes from, disease types, etc.; under the attributes of medical institutions and disease types, sub-attribute levels such as specific departments and disease subtypes can be customized respectively; further, due to the professional characteristics of medical corpus data, the content recorded between different paragraphs of the to-be-annotated medical corpus data may have different clinical meanings and semantic description characteristics. Therefore, the to-be-annotated corpus needs to define sub-module attributes to distinguish different annotation contents, so as to reduce the annotation difficulty and provide training corpus for training artificial intelligence models applicable to different sub-module contents. Taking the medical record data as an example, the attributes can be defined as admission record, first course record, the Nth (N≥2) course record, discharge record, inspection record, operation record, etc.
[0067] Step S24: Determine the semicolons and full stops in the to-be-annotated medical corpus data, and determine the determined semicolons and full stops as delimiter symbols.
[0068] In this embodiment, the semicolons and full stops in the to-be-annotated medical corpus data are determined, and the determined semicolons and full stops are determined as delimiter symbols. That is, after determining the data attributes of the to-be-annotated medical corpus data, it is necessary to perform deduplication processing on the to-be-annotated medical corpus data to eliminate the repeatedly occurring corpus data in the to-be-annotated medical corpus data and improve the processing efficiency. Thus, first, it is necessary to determine the semicolons and full stops in the to-be-annotated medical corpus data, and determine the determined semicolons and full stops as delimiter symbols.
[0069] Step S25: Split the to-be-annotated medical corpus data based on the delimiter symbols to obtain the split corpus data.
[0070] In this embodiment, the medical corpus data to be annotated is segmented based on the separator to obtain segmented corpus data. That is, the medical corpus data to be annotated is segmented by the determined semicolons and periods to deduplicate the segmented corpus data.
[0071] Step S26: determining duplicate corpus data in the segmented corpus data based on the target corpus attribute, and removing the duplicate corpus data to obtain deduplicated corpus data.
[0072] In this embodiment, the repeated corpus data in the segmented corpus data is determined based on the target corpus attribute, and the repeated corpus data is removed to obtain the deduplicated corpus data. That is, taking the segmented corpus data whose attribute is medical record data as an example, since there are often repeated contents in sentences and paragraphs in multiple medical records, the segmented corpus data can be deduplicated using semicolons and periods as segmentation symbols, and only the first-appearing data is retained as the annotation data.
[0073] Step S27: determining the entity type and entity relationship of the deduplicated corpus data according to the target corpus attributes, and constructing knowledge triples according to the entity type and the entity relationship.
[0074] Step S28: determining historical annotation data having the same attributes as the target corpus from a preset annotation library based on the knowledge triples, and pre-annotating the deduplicated corpus data based on the historical annotation data.
[0075] It should be noted that for a more detailed description of step S27 and step S28, reference can be made to the aforementioned embodiment, which will not be repeated here.
[0076] It can be seen that in this embodiment, the corpus data file is first received, and the file format of the corpus data file is judged. If the file format of the corpus data file is an electronic document, the medical corpus data to be annotated in the electronic document is extracted. If the file format of the corpus data file is a scanned file, an optical character recognition operation is performed on the scanned file to determine the medical corpus data to be annotated. Then, the corpus attribute determination instruction sent by the client is received, and the corpus attribute of the medical corpus data to be annotated and the sub-corpus attribute of the corpus attribute are determined based on the corpus attribute determination instruction, so as to determine the corpus attribute and the sub-corpus attribute as the target corpus attribute of the medical corpus data to be annotated. Finally, the semicolons and full stops in the medical corpus data to be annotated are determined, and the determined semicolons and full stops are determined as delimiter symbols. The medical corpus data to be annotated is segmented based on the delimiter symbols to obtain segmented corpus data. The duplicate corpus data in the segmented corpus data is determined based on the target corpus attribute, and the duplicate corpus data is removed to obtain deduplicated corpus data. In this way, the efficiency of pre-annotation can be effectively improved through the preprocessing of the medical corpus data to be annotated.
[0077] Based on the foregoing embodiments, it can be known that in this application, after preprocessing and deduplicating the medical corpus data to be annotated, it is necessary to construct a knowledge triple by determining the entity type and entity relationship in the corpus data, and determine the historical annotation data through the knowledge triple, so as to perform pre-annotation through the historical annotation data. For this reason, this embodiment describes in detail how to determine the knowledge triple and how to determine the historical annotation data through the knowledge triple. See Figure 3 As shown, an embodiment of the present invention discloses a medical corpus pre-annotation method, including:
[0078] Step S31: Obtain the medical corpus data to be annotated, and determine the corpus attribute of the medical corpus data to be annotated as the target corpus attribute based on the received corpus attribute determination instruction.
[0079] Step S32: Perform data deduplication on the medical corpus data to be annotated based on the target corpus attribute to obtain deduplicated corpus data.
[0080] Step S33: Determine the entity type of the deduplicated corpus data based on the corpus attribute, and determine the sub-entity type of the deduplicated corpus data based on the sub-corpus attribute.
[0081] In this embodiment, the entity type of the deduplicated corpus data is determined based on the corpus attribute, and the sub-entity type of the deduplicated corpus data is determined based on the sub-corpus attribute. That is, the entity type and entity relationship type can be customized through the entity type and entity relationship type customization module to achieve the customization of the entity type and entity relationship type based on the data attributes of the medical corpus to be annotated and the actual annotation requirements. Taking the corpus attribute as the medical record data as an example, in the corpus to be annotated with medical record data attributes, the entity types can be defined as "clinical manifestations" and "treatment plans", and the sub-entity types can be defined respectively under the entity types. Among them, the sub-entity types of the entity type "clinical manifestations" can include "symptoms, signs, adverse reactions", etc., and the sub-entity types of the entity type "treatment plans" can include "plan name, plan cycle, implementation frequency, treatment method name, administration method, dosage, treatment time", etc.
[0082] Step S34: Determine the entity relationship between the entity type and the sub-entity type, and construct a knowledge triple of the deduplicated corpus data based on the entity type, the sub-entity type, and the entity relationship.
[0083] In this embodiment, the entity relationship between the entity type and the sub-entity type is determined, and a knowledge triple of the deduplicated corpus data is constructed based on the entity type, the sub-entity type, and the entity relationship. That is, after determining the entity relationship of the deduplicated corpus data, the entity type can be determined. Taking the target attribute as the treatment plan as an example, when the entity type is determined to be "treatment plan" and the sub-entity type is "dosage", the entity relationship between the entity type and the sub-entity type can be determined as the entity relationship corresponding to the entity relationship determination instruction sent by the user terminal. For example, when the received entity relationship determination instruction indicates that the entity relationship is determined to be an "inclusion relationship", an entity-relationship-entity triple can be constructed based on the determined entity type, sub-entity type, and entity relationship, that is, the knowledge triple is "treatment plan - inclusion relationship - dosage".
[0084] It should be noted that during the process of annotating medical corpus data, there are frequently scenarios where the same relationship exists between one entity type and multiple other entity types. The existing annotation systems need to establish relationship types separately for these entity types, which involves a huge amount of work. Therefore, a model for batch establishing relationships needs to be constructed to solve this problem. After selecting the relationship type, the batch relationship establishing module is called to select the entity type and the sub-entity type, and multiple entity-relationship-entity triples with the same relationship type and direction are established, that is, knowledge triples. For example, when the entity type is determined to be "treatment plan" and the entity relationship is "inclusion relationship", the sub-entity types of "implementation frequency", "treatment time", and "dose" can be selected one by one to batch generate knowledge triples of "treatment plan - inclusion relationship - implementation frequency", "treatment plan - inclusion relationship - treatment time", and "treatment plan - inclusion relationship - dose". In this way, the generation efficiency of knowledge triples can be effectively improved, and the usage experience of medical staff can be enhanced.
[0085] Step S35: Determine the pre-annotation rule as the entity exact match rule or the entity similarity match rule according to the received pre-annotation instruction.
[0086] In this embodiment, the pre-annotation rule is determined as the entity exact match rule or the entity similarity match rule according to the received pre-annotation instruction. That is, among the medical corpus data to be annotated with the same attributes, there are a large number of repetitive clinical professional term entity types and entity relationships. A pre-annotation library can be constructed based on historical annotation data, and entity exact match or similarity match rules can be written. Based on the pre-annotation module in the corpus annotation system, pre-annotation before manual annotation of the medical corpus data to be annotated can be performed through the historical annotation corpus with the same attributes that have been annotated, reducing the workload of manual corpus annotation.
[0087] Step S36: If the received pre-annotation instruction indicates that the pre-annotation rule is the entity exact match rule, determine the historical annotation data that exactly matches the entity type in the knowledge triple from the preset annotation library based on the target corpus attribute, and perform pre-annotation on the deduplicated corpus data based on the historical annotation data.
[0088] In this embodiment, if the received pre-annotation instruction indicates that the pre-annotation rule is the entity exact matching rule, historical annotation data that exactly matches the entity type in the knowledge triple is determined from the preset annotation library based on the target corpus attribute, and the deduplicated corpus data is pre-annotated based on the historical annotation data. That is, if it is determined that the pre-annotation rule is the entity exact matching rule, taking the target attribute as the treatment plan as an example, and the knowledge triple is "treatment plan - inclusion relationship - implementation frequency", historical annotation data with the entity type of "treatment plan" and the sub-entity type of "implementation frequency" in the knowledge triple needs to be determined from the preset annotation library, and the deduplicated corpus data is pre-annotated through the determined historical annotation data.
[0089] Step S37: If the received pre-annotation instruction indicates that the pre-annotation rule is the entity similarity matching rule, historical similar annotation data with a character position distance difference not greater than the preset character position distance difference threshold from the entity type in the knowledge triple is determined from the preset annotation library based on the target corpus attribute, and the deduplicated corpus data is pre-annotated based on the historical similar annotation data.
[0090] In this embodiment, if the received pre-annotation instruction indicates that the pre-annotation rule is the entity similarity matching rule, historical similar annotation data with a character position distance difference not greater than the preset character position distance difference threshold from the entity type in the knowledge triple is determined from the preset annotation library based on the target corpus attribute, and the deduplicated corpus data is pre-annotated based on the historical similar annotation data. That is, if it is determined that the pre-annotation rule is the similarity matching rule, taking the target attribute as the treatment plan as an example, and the knowledge triple is "treatment plan - inclusion relationship - implementation frequency", historical similar annotation data with a character position distance difference not greater than the preset character position distance difference threshold between the entity type and "treatment plan" and the sub-entity type and "implementation frequency" in the knowledge triple needs to be determined from the preset annotation library, and the deduplicated corpus data is pre-annotated based on the historical similar annotation data. And the preset character position distance difference threshold can be set according to the requirements of the user terminal.
[0091] It should be noted that the medical corpus pre-annotation method described in this application further includes: if a wrong annotation modification instruction is received, based on the wrong annotation modification instruction, the determined entity type in the deduplicated corpus data is deleted, and the entity type of the deduplicated corpus data is modified based on the wrong annotation modification instruction. That is to say, when the existing annotation system has annotation errors, it needs to be processed by the operation method of deletion - re-annotation. When a large amount of content needs to be modified, the repeated workload is huge. In this embodiment, the quick modification module in the corpus annotation system can be called to quickly modify the pre-annotation errors of the data to be annotated caused by mistakes in the manual annotation process and the quality problems of the content in the pre-annotation library. Specifically, by selecting the entity type, the entity type of the already annotated entity can be redefined, or the annotation range of the entity type can be selected, but the already annotated entity relationship is retained and the entity relationship is not modified.
[0092] It should be further noted that the medical corpus pre-annotation method described in this application further includes: if a annotation copy instruction is received, based on the annotation copy instruction, the knowledge triple is split to determine the entity type and the entity relationship in the knowledge triple; based on the entity type and the entity relationship, the data to be copied is split to determine the first annotation data corresponding to the entity type and the second annotation data corresponding to the entity relationship, so that the client can copy the first annotation data and the second annotation data. That is to say, in the annotation process of medical corpus data, there will frequently be multiple juxtaposed knowledge triple contents with the same entity type and entity relationship type. The existing annotation system needs to annotate these juxtaposed contents separately, and the workload is huge. In this embodiment, after completing the annotation of one knowledge triple content, the annotation copy module of the corpus annotation system can be called. Using the already annotated knowledge triple as a template and selecting the data to be annotated that needs to be copied, the entity type and entity relationship annotation data of the already annotated template can be copied in the data to be annotated. The entity annotation range in the data to be annotated and the positioning of the two entities in the knowledge triple can be realized by identifying the splitting symbols between entities, such as punctuation marks, as well as the differences between Chinese entities and English or digital entities, and the differences between Chinese / English / digital entities and symbol entities, etc., based on the built-in rules of the annotation copy module.
[0093] It should be noted that for the more detailed content of step S31 and step S32 in this embodiment, reference can be made to the foregoing embodiment, and details will not be repeated here.
[0094] It can be seen that in this embodiment, first, the entity type of the deduplicated corpus data is determined based on the corpus attribute, and the sub-entity type of the deduplicated corpus data is determined based on the sub-corpus attribute. The entity relationship between the entity type and the sub-entity type is determined, and a knowledge triple of the deduplicated corpus data is constructed based on the entity type, the sub-entity type, and the entity relationship. Then, according to the received pre-annotation instruction, the pre-annotation rule is determined to be the entity exact match rule or the entity similarity match rule. If the received pre-annotation instruction indicates that the pre-annotation rule is the entity exact match rule, historical annotation data that exactly matches the entity type in the knowledge triple is determined from the preset annotation library based on the target corpus attribute, and the deduplicated corpus data is pre-annotated based on the historical annotation data; if the received pre-annotation instruction indicates that the pre-annotation rule is the entity similarity match rule, historical similar annotation data with a character position distance difference not greater than the preset character position distance difference threshold from the entity type in the knowledge triple is determined from the preset annotation library based on the target corpus attribute, and the deduplicated corpus data is pre-annotated based on the historical similar annotation data. And the annotation errors can be quickly modified by calling the quick modification module in the corpus annotation system, and the content to be copied can be batch-copied by calling the annotation copy module in the corpus annotation system. In this way, on the one hand, the historical annotation data in the preset annotation library can be determined through the constructed knowledge triple, effectively improving the annotation efficiency; on the other hand, the pre-annotation errors of the data to be annotated caused by mistakes in the manual annotation process and the quality problems of the content in the preset annotation library can be quickly modified by calling the quick modification module in the corpus annotation system; on the third hand, the content to be copied can be batch-copied based on the entity relationship and the entity type by calling the annotation copy module in the corpus annotation system, improving the user experience.
[0095] See Figure 4 As shown, an embodiment of the present invention discloses a method for pre-annotating medical corpus, including:
[0096] Receive the uploaded corpus data file through the data acquisition module in the corpus annotation system, and extract the medical corpus data to be annotated in the file. After the medical corpus data to be annotated is extracted, define the corpus attribute of the medical corpus data to be annotated by receiving the corpus attribute determination instruction from the client, so as to define the corpus attribute of the medical corpus data to be annotated as the target corpus attribute corresponding to the received corpus attribute determination instruction. After determining the data attribute of the medical corpus data to be annotated, it is necessary to perform deduplication processing on the medical corpus data to be annotated to eliminate the repeatedly occurring corpus data in the medical corpus data to be annotated. The medical corpus data to be annotated can be segmented by the determined semicolons and periods, so as to perform data deduplication on the segmented corpus data. And through the entity type and entity relationship type customization module, the customization of entity type and entity relationship type can be realized based on the data attribute of the medical corpus data to be annotated and the actual annotation requirements; the relationship batch establishment module can be called to select the entity type and sub-entity type, and establish multiple knowledge triples with the same relationship type and direction. Based on the pre-annotation module in the corpus annotation system, through the historical annotation corpus with the same attributes that has been annotated, pre-annotation can be performed on the medical corpus data to be annotated before manual annotation, reducing the workload of manual corpus annotation; the quick modification module can be called to quickly modify the pre-annotation error of the data to be annotated caused by mistakes in the manual annotation process and the quality problem of the content in the pre-annotation library; the annotation copying module can be called to use the annotated knowledge triple as a template, select the content to be annotated that needs to be copied, and copy the entity type and entity relationship annotation data of the annotated template in the content to be annotated.
[0097] See Figure 5 As shown, an embodiment of the present invention discloses a medical corpus pre-annotation device, which is applied to a corpus annotation system and includes:
[0098] An attribute determination module 11, configured to obtain the medical corpus data to be annotated, and determine the corpus attribute of the medical corpus data to be annotated as the target corpus attribute based on the received corpus attribute determination instruction;
[0099] A corpus deduplication module 12, configured to perform data deduplication on the medical corpus data to be annotated based on the target corpus attribute to obtain deduplicated corpus data;
[0100] A triple construction module 13, configured to determine the entity type and entity relationship of the deduplicated corpus data according to the target corpus attribute, and construct a knowledge triple according to the entity type and the entity relationship;
[0101] A corpus pre-annotation module 14, configured to determine historical annotation data with the same attribute as the target corpus attribute from a preset annotation library based on the knowledge triple, and perform pre-annotation on the deduplicated corpus data based on the historical annotation data.
[0102] It can be seen that in this embodiment, first, the medical corpus data to be annotated is obtained, and based on the received corpus attribute determination instruction, the corpus attribute of the medical corpus data to be annotated is determined as the target corpus attribute. Then, data deduplication is performed on the medical corpus data to be annotated based on the target corpus attribute to obtain the deduplicated corpus data. Next, the entity type and entity relationship of the deduplicated corpus data are determined according to the target corpus attribute, and a knowledge triple is constructed based on the entity type and the entity relationship. Finally, historical annotation data with the same attribute as the target corpus attribute is determined from the preset annotation library based on the knowledge triple, and the deduplicated corpus data is pre-annotated based on the historical annotation data. It can be seen that through the medical corpus pre-annotation method of the present application, after successfully obtaining the medical corpus data to be annotated, the corpus attribute of the medical data to be annotated can be determined, data deduplication can be performed on the medical corpus data to be annotated according to the determined target corpus attribute, the entity type and entity relationship in the deduplicated corpus data can be determined, a knowledge triple can be constructed based on the entity type and entity relationship, historical annotation data can be determined based on the constructed knowledge triple, and the deduplicated corpus data can be pre-annotated with the historical annotation data. In this way, after determining the historical annotation data with the same attribute as the medical corpus data to be annotated, pre-annotation can be performed through the historical annotation data to reduce the workload of subsequent more accurate manual annotation.
[0103] In some embodiments, the attribute determination module 11 may specifically include:
[0104] A file format determination unit, configured to receive a corpus data file and determine the file format of the corpus data file;
[0105] A corpus data determination unit, configured to extract the medical corpus data to be annotated in the electronic document if the file format of the corpus data file is an electronic document, and perform optical character recognition on the scanned file to determine the medical corpus data to be annotated if the file format of the corpus data file is a scanned file;
[0106] A corpus attribute determination unit, configured to receive a corpus attribute determination instruction sent by the user terminal, and determine the corpus attribute of the medical corpus data to be annotated and the sub-corpus attribute of the corpus attribute based on the corpus attribute determination instruction, so as to determine the corpus attribute and the sub-corpus attribute as the target corpus attribute of the medical corpus data to be annotated.
[0107] In some embodiments, the corpus deduplication module 12 may specifically include:
[0108] A delimiter determination unit for determining semicolons and full stops in the medical corpus data to be annotated, and determining the determined semicolons and full stops as delimiters;
[0109] A corpus segmentation unit for segmenting the medical corpus data to be annotated based on the delimiters to obtain segmented corpus data;
[0110] A corpus duplicate removal unit for determining duplicate corpus data in the segmented corpus data based on the target corpus attribute, and removing the duplicate corpus data to obtain deduplicated corpus data.
[0111] In some embodiments, the triple construction module 13 may specifically include:
[0112] An entity type determination unit for determining the entity type of the deduplicated corpus data based on the corpus attribute, and determining the sub-entity type of the deduplicated corpus data based on the sub-corpus attribute;
[0113] A triple construction unit for determining the entity relationship between the entity type and the sub-entity type, and constructing a knowledge triple of the deduplicated corpus data based on the entity type, the sub-entity type, and the entity relationship.
[0114] In some embodiments, the medical corpus pre-annotation device may further include:
[0115] An entity type modification unit for, if an error annotation modification instruction is received, deleting the determined entity type in the deduplicated corpus data based on the error annotation modification instruction, and modifying the entity type of the deduplicated corpus data based on the error annotation modification instruction.
[0116] In some embodiments, the corpus pre-annotation module 14 may specifically include:
[0117] A matching rule determination unit for determining a pre-annotation rule as an entity exact match rule or an entity similarity match rule according to a received pre-annotation instruction;
[0118] A first pre-annotation unit for, if the received pre-annotation instruction indicates that the pre-annotation rule is the entity exact match rule, determining historical annotation data that exactly matches the entity type in the knowledge triple from a preset annotation library based on the target corpus attribute, and pre-annotating the deduplicated corpus data based on the historical annotation data;
[0119] A second pre-annotation unit, configured to, if the received pre-annotation instruction indicates that the pre-annotation rule is the entity similarity matching rule, determine, based on the target corpus attribute, historical similar annotation data in the preset annotation library whose character position distance difference from the character position of the entity type in the knowledge triple is not greater than a preset character position distance difference threshold, and pre-annotate the deduplicated corpus data based on the historical similar annotation data.
[0120] In some embodiments, the medical corpus pre-annotation device may further include:
[0121] A triple splitting unit, configured to, if a annotation copying instruction is received, split the knowledge triple based on the annotation copying instruction to determine the entity type and the entity relationship in the knowledge triple;
[0122] An annotation data copying unit, configured to split the data to be copied based on the entity type and the entity relationship to determine first annotation data corresponding to the entity type and second annotation data corresponding to the entity relationship, so that a client can copy the first annotation data and the second annotation data.
[0123] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 6 which is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be considered as any limitation to the scope of use of the present application.
[0124] Figure 6 This is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the medical corpus pre-annotation method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0125] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and specific limitations are not imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitations are imposed here.
[0126] In addition, as a carrier for storing resources, the memory 22 can be a read-only memory, a random access memory, a magnetic disk, an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc. The storage method can be temporary storage or permanent storage.
[0127] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the medical corpus pre-annotation method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.
[0128] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the medical corpus pre-annotation method disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0129] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0130] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0131] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0132] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0133] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. At the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for pre-annotating medical corpora, characterized in that, Applied to a corpus annotation system, including: Obtain the medical corpus data to be annotated, and determine the corpus attributes of the medical corpus data to be annotated as target corpus attributes based on the received corpus attribute determination instruction; Deduplicate the medical corpus data to be annotated based on the target corpus attributes to obtain the deduplicated corpus data; Determine the entity types and entity relationships of the deduplicated corpus data according to the target corpus attributes, and construct knowledge triples according to the entity types and the entity relationships; Determine the historical annotation data with attributes the same as the target corpus attributes from a preset annotation library based on the knowledge triples, and pre-annotate the deduplicated corpus data based on the historical annotation data; Among them, the target corpus attributes determine the entity types and entity relationships of the deduplicated corpus data, and construct knowledge triples according to the entity types and the entity relationships, including: Determine the entity types of the deduplicated corpus data based on the corpus attributes, and determine the sub-entity types of the deduplicated corpus data based on the sub-corpus attributes; Determine the entity relationships between the entity types and the sub-entity types, and construct the knowledge triples of the deduplicated corpus data based on the entity types, the sub-entity types, and the entity relationships; Among them, the determining the historical annotation data with attributes the same as the target corpus attributes from a preset annotation library based on the knowledge triples, and pre-annotating the deduplicated corpus data based on the historical annotation data, includes: Determine the pre-annotation rule as the entity full match rule or the entity similarity match rule according to the received pre-annotation instruction; If the received pre-annotation instruction indicates that the pre-annotation rule is the entity full match rule, then determine the historical annotation data that exactly matches the entity type in the knowledge triples from the preset annotation library based on the target corpus attributes, and pre-annotate the deduplicated corpus data based on the historical annotation data; If the received pre-annotation instruction indicates that the pre-annotation rule is the entity similarity match rule, then determine the historical similar annotation data with a character position distance difference not greater than the preset character position distance difference threshold from the preset annotation library based on the target corpus attributes, and pre-annotate the deduplicated corpus data based on the historical similar annotation data.
2. The method for pre-annotating medical corpora according to claim 1, characterized in that, The obtaining the medical corpus data to be annotated, and determining the corpus attributes of the medical corpus data to be annotated as target corpus attributes based on the received corpus attribute determination instruction, includes: Receive the corpus data file, and judge the file format of the corpus data file; If the file format of the corpus data file is an electronic document, extract the medical corpus data to be annotated from the electronic document. If the file format of the corpus data file is a scanned file, perform an optical character recognition operation on the scanned file to determine the medical corpus data to be annotated; A corpus attribute determination instruction sent by a user terminal is received, and a corpus attribute of the medical corpus data to be annotated and a sub-corpus attribute of the corpus attribute are determined based on the corpus attribute determination instruction, so as to determine the corpus attribute and the sub-corpus attribute as the target corpus attribute of the medical corpus data to be annotated.
3. The method for pre-annotating medical corpora according to claim 1, characterized in that, The step of performing data deduplication on the medical corpus data to be annotated based on the target corpus attribute to obtain deduplicated corpus data includes: Determining a semicolon and a period in the medical corpus data to be annotated, and determining the determined semicolon and the period as separators; Segmenting the medical corpus data to be annotated based on the separation symbol to obtain segmented corpus data; Based on the target corpus attribute, duplicate corpus data in the segmented corpus data is determined, and the duplicate corpus data is removed to obtain deduplicated corpus data.
4. The method for pre-annotating medical corpora according to claim 1, characterized in that, Also includes: If an error marking modification instruction is received, the entity type determined in the deduplicated corpus data is deleted based on the error marking modification instruction, and the entity type of the deduplicated corpus data is modified based on the error marking modification instruction.
5. The method for pre-annotating medical corpora according to claim 1, characterized in that, Also includes: If a label copy instruction is received, segmenting the knowledge triple based on the label copy instruction to determine the entity type and the entity relationship in the knowledge triple; The data to be copied is segmented based on the entity type and the entity relationship to determine first labeled data corresponding to the entity type and second labeled data corresponding to the entity relationship, so that the client can copy the first labeled data and the second labeled data.
6. A device for pre-annotating medical corpora, characterized in that, Applied to corpus annotation systems, including: An attribute determination module, used for acquiring medical corpus data to be annotated, and determining the corpus attribute of the medical corpus data to be annotated as the target corpus attribute based on the received corpus attribute determination instruction; A corpus deduplication module, used for deduplicating the medical corpus data to be annotated based on the target corpus attributes to obtain deduplicated corpus data; A triple construction module, used to determine the entity type and entity relationship of the deduplicated corpus data according to the target corpus attributes, and to construct knowledge triples according to the entity type and the entity relationship; A corpus pre-annotation module, configured to determine historical annotation data having the same attributes as the target corpus from a preset annotation library based on the knowledge triples, and pre-annotate the deduplicated corpus data based on the historical annotation data; Wherein, the triple building module includes: An entity type determination unit, configured to determine the entity type of the deduplicated corpus data based on the corpus attribute, and to determine the sub-entity type of the deduplicated corpus data based on the sub-corpus attribute; A triple construction unit, used to determine the entity type and the entity relationship between the sub-entity type, and to construct a knowledge triple of the deduplicated corpus data based on the entity type, the sub-entity type and the entity relationship; In some embodiments, the medical corpus pre-annotation device may further include: An entity type modification unit, configured to, if an error annotation modification instruction is received, delete the determined entity type in the deduplicated corpus data based on the error annotation modification instruction, and modify the entity type of the deduplicated corpus data based on the error annotation modification instruction; Wherein, the corpus pre-annotation module includes: A matching rule determination unit, configured to determine a pre-annotation rule as an entity exact matching rule or an entity similarity matching rule according to a received pre-annotation instruction; A first pre-annotation unit, configured to, if the received pre-annotation instruction indicates that the pre-annotation rule is the entity exact matching rule, determine historical annotation data that exactly matches the entity type in the knowledge triple from a preset annotation library based on the target corpus attribute, and perform pre-annotation on the deduplicated corpus data based on the historical annotation data; A second pre-annotation unit, configured to, if the received pre-annotation instruction indicates that the pre-annotation rule is the entity similarity matching rule, determine historical similar annotation data whose character position distance difference from the entity type in the knowledge triple is not greater than a preset character position distance difference threshold from the preset annotation library based on the target corpus attribute, and perform pre-annotation on the deduplicated corpus data based on the historical similar annotation data.
7. An electronic device, characterized in that, Including: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the medical corpus pre-annotation method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, For storing a computer program, which when executed by a processor implements the medical corpus pre-annotation method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Entity relationship automatic labeling method applied to medical text
CN111291568A
Conversation intention intelligent identification model construction method, and device, equipment
CN112131890A