Information Extraction Method, Device, Electronic Device, and Storage Medium

By vectorized processing and model recognition of the statements in the surgical record, and automatic identification and extraction of surgical methods and related information, the problems of low efficiency and error-prone manual audit in the existing technology are solved, and a more efficient and accurate audit process is achieved.

CN114496138BActive Publication Date: 2025-06-13SHENYANG NEUSOFT INTELLIGENT MEDICAL TECH RES INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111640706.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-06-13
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

In the prior art, the review of surgical procedures and related information in surgical records usually relies on manual labor, resulting in low extraction efficiency and prone to errors.

Method used

By vectorizing the statements in the surgical record, and using the translation model and sequence labeling model, the surgical method is automatically identified and information extracted, and verified with preset standard surgical method information.

Benefits of technology

It significantly improves the efficiency of surgical recognition and information extraction, reduces manual errors, and improves the accuracy of audits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114496138B_ABST
    Figure CN114496138B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an information extraction method, apparatus, electronic device, and storage medium, and relates to the field of text processing. The method includes: performing vectorization processing on at least one statement in a surgical record to be extracted to obtain at least one corresponding vectorized statement; using a translation model to translate the at least one vectorized statement to obtain the surgical procedures corresponding to each vectorized statement, where the surgical procedures corresponding to each vectorized statement are one or more; and using a sequence labeling model to perform information extraction on the surgical procedures corresponding to each vectorized statement to obtain the target information corresponding to the surgical procedures corresponding to each vectorized statement. It can improve the efficiency and accuracy of surgical procedure recognition and information extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of text processing, and in particular, to an information extraction method, apparatus, electronic device, and storage medium. Background Art

[0002] As the scope of compensation covered by medical insurance becomes more comprehensive and the number of recipients increases, the payment pressure on medical care has also become greater. The implementation of the DRGs (Diagnosis Related Groups) and DIP (Big Data Diagnosis-Intervention Packet) systems has brought about great changes in the medical insurance payment review method. Especially after the hospital encodes the medical records and uploads them to the medical insurance bureau for encoding verification, the payment process has become the most critical link in the medical insurance payment review system. In this process, for patients who need surgery, since the surgery itself involves complex surgical procedures and the consumables involved during the surgery are very critical content in the medical insurance payment review process, the surgical procedures (operation methods) and other key information during the surgery need to be reviewed. Currently, the review of these contents is usually carried out manually. Due to the complexity of the content involved in the surgery and the large amount of information, there are problems of low extraction efficiency and easy errors in extracting the surgical procedures and related information from the surgical records through manual review. Summary of the Invention

[0003] The purpose of the present disclosure is to provide an information extraction method, apparatus, electronic device, and storage medium to solve the problems of low extraction efficiency and easy errors in the existing method.

[0004] To achieve the above purpose, in the first aspect of the present disclosure, an information extraction method is provided, including:

[0005] Performing vectorization processing on at least one statement in the surgical record to be extracted to obtain at least one corresponding vectorized statement;

[0006] Using a translation model to translate the at least one vectorized statement to obtain the surgical procedures corresponding to each vectorized statement, and the surgical procedures corresponding to each vectorized statement are one or more;

[0007] Using a sequence annotation model to extract information from the surgical procedures corresponding to each vectorized statement to obtain the target information corresponding to the surgical procedures corresponding to each vectorized statement.

[0008] Optionally, the target information corresponding to any surgical procedure includes at least one of the consumable information, drug information, device information, and ICD coding information corresponding to the surgical procedure, and the method further includes:

[0009] Verify the target information corresponding to the surgical procedure corresponding to each vectorized statement by using the preset standard surgical procedure information; wherein, the standard surgical procedure information includes at least one of standard consumable information, standard drug information, standard instrument information, and standard ICD coding information corresponding to multiple surgical procedures;

[0010] When the target information corresponding to the surgical procedure corresponding to each vectorized statement passes the verification, it is determined that the surgical record passes the review;

[0011] When the target information corresponding to any surgical procedure corresponding to any vectorized statement fails to pass the verification, output the surgical procedure that fails to pass the verification and the target information that fails to pass the verification.

[0012] Optionally, the vectorizing at least one statement in the surgical record to be extracted to obtain the corresponding at least one vectorized statement includes:

[0013] Perform holistic vectorization on the first statement in the surgical record to obtain a first vector;

[0014] Perform word segmentation on the first statement;

[0015] After performing vectorization on each word obtained by word segmentation of the first statement, merge them to obtain a second vector;

[0016] Concatenate the first vector and the second vector to obtain the vectorized statement corresponding to the first statement.

[0017] Optionally, the merging after performing vectorization on each vector obtained by word segmentation of the first statement to obtain a second vector includes:

[0018] Perform vectorization on each word respectively to obtain the word vectors corresponding to each word;

[0019] Multiply each word vector by the corresponding coefficient, and sum the products of each word vector and the corresponding coefficient to obtain the second vector.

[0020] Optionally, the translation model is an attention model.

[0021] Optionally, the training method of the translation model includes:

[0022] Obtain multiple surgical record samples, each surgical record sample includes at least one sample statement, and each sample statement has at least one labeled surgical procedure;

[0023] For each surgical record sample among multiple surgical record samples, perform vectorization processing on at least one sample statement in the surgical record sample to obtain at least one corresponding vectorized statement;

[0024] Use the vectorized statements corresponding to each sample statement in the multiple surgical record samples as input data, and use the at least one surgical procedure already marked for each sample statement as verification data to train the initial translation model to obtain the translation model.

[0025] Optionally, the training method of the sequence annotation model includes:

[0026] Obtain multiple surgical record samples, each surgical record sample includes at least one sample statement, each sample statement has one or more pre-marked surgical procedures, and a sequence annotation corresponding to the one or more surgical procedures, where the sequence annotation is used to characterize the attribute information of each word in the sample statement;

[0027] Use the multiple surgical record samples as input data for the initial sequence annotation model, and use the sequence annotations corresponding to the one or more surgical procedures corresponding to each sample statement in the multiple surgical record samples as verification data to train the initial sequence annotation model to obtain the sequence annotation model.

[0028] In a second aspect of the present disclosure, there is provided an information extraction device, including:

[0029] A processing module, configured to perform vectorization processing on at least one statement in the surgical record to be extracted to obtain at least one corresponding vectorized statement;

[0030] A translation module, configured to use the translation model to translate the at least one vectorized statement to obtain the surgical procedures corresponding to each vectorized statement, where the surgical procedures corresponding to each vectorized statement are one or more;

[0031] An information extraction module, configured to use the sequence annotation model to perform information extraction on the surgical procedures corresponding to each vectorized statement to obtain the target information corresponding to the surgical procedures corresponding to each vectorized statement.

[0032] Optionally, the target information corresponding to any surgical procedure includes at least one of the consumable information, drug information, instrument information, and ICD coding information corresponding to the surgical procedure, and the device further includes:

[0033] A verification module, configured to use the preset standard surgical procedure information to verify the target information corresponding to the surgical procedures corresponding to each vectorized statement; where the standard surgical procedure information includes at least one of the standard consumable information, standard drug information, standard instrument information, and standard ICD coding information corresponding to multiple surgical procedures;

[0034] A determination module, configured to determine that the surgical record passes the review when the target information corresponding to the surgical method corresponding to each vectorized statement passes the verification;

[0035] An output module, configured to output the surgical method that fails to pass the verification and the target information that fails to pass the verification when the target information corresponding to any surgical method corresponding to any of the vectorized statements fails to pass the verification.

[0036] Optionally, the processing module includes:

[0037] A vectorization processing sub-module, configured to perform an overall vectorization processing on the first statement in the surgical record to obtain a first vector;

[0038] A word segmentation sub-module, configured to perform word segmentation processing on the first statement;

[0039] A merging sub-module, configured to perform vectorization processing on each word obtained after word segmentation of the first statement and then merge them to obtain a second vector;

[0040] A splicing sub-module, configured to splice the first vector and the second vector to obtain the vectorized statement corresponding to the first statement.

[0041] Optionally, the merging sub-module includes:

[0042] A word vector processing sub-module, configured to perform vectorization processing on each of the words to obtain word vectors corresponding to the words;

[0043] A calculation sub-module, configured to multiply each word vector by a corresponding coefficient and sum the products of each word vector and the corresponding coefficient to obtain the second vector.

[0044] Optionally, the translation model is an attention model.

[0045] Optionally, the apparatus further includes: a first training module, and the first training module is configured to:

[0046] Obtain a plurality of surgical record samples, where each surgical record sample includes at least one sample statement, and each sample statement has at least one labeled surgical method;

[0047] For each surgical record sample in the plurality of surgical record samples, perform vectorization processing on at least one sample statement in the surgical record sample to obtain at least one corresponding vectorized statement;

[0048] Using the vectorized statements corresponding to each sample statement in the multiple surgical record samples as input data, and using at least one surgical procedure already labeled for each sample statement as verification data, train the initial translation model to obtain the translation model.

[0049] Optionally, the apparatus further includes: a second training module, and the second training module is configured to:

[0050] Obtain multiple surgical record samples, where each surgical record sample includes at least one sample statement, each sample statement has one or more pre-labeled surgical procedures, and sequence annotations corresponding to the one or more surgical procedures, and the sequence annotations are used to characterize the attribute information of each word in the sample statement;

[0051] Using the multiple surgical record samples as input data for the initial sequence annotation model, and using the sequence annotations corresponding to one or more surgical procedures corresponding to each sample statement in the multiple surgical record samples as verification data, train the initial sequence annotation model to obtain the sequence annotation model.

[0052] In a third aspect of the present disclosure, there is provided an electronic device, including:

[0053] A memory, on which a computer program is stored;

[0054] A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of the first aspect.

[0055] In a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method according to any one of the first aspect.

[0056] In the above technical solution, by performing vectorization processing on at least one statement in the surgical record to be extracted, obtaining at least one corresponding vectorized statement, and using a translation model to translate the at least one vectorized statement to obtain the surgical procedures corresponding to each vectorized statement, where each vectorized statement corresponds to one or more surgical procedures, and then using a sequence annotation model to extract information from the surgical procedures corresponding to each vectorized statement to obtain the target information corresponding to the surgical procedures corresponding to each vectorized statement. Through the above technical solution, it is possible to automatically identify the surgical procedures in the surgical record and extract the target information corresponding to each surgical procedure. Compared with the existing method of manual extraction, it can significantly improve the efficiency of surgical procedure identification and information extraction, and can avoid the problem of easy errors in the manual method, improving the accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The accompanying drawings are used to provide a further understanding of the present disclosure and form a part of the specification. Together with the following detailed description, they serve to explain the present disclosure, but do not limit the present disclosure. In the accompanying drawings:

[0058] Figure 1 is a flowchart of an information extraction method shown according to an exemplary embodiment of the present disclosure.

[0059] Figure 2 is a flowchart of another information extraction method shown according to an exemplary embodiment of the present disclosure.

[0060] Figure 3 is a flowchart of a vectorization processing method shown according to an exemplary embodiment of the present disclosure.

[0061] Figure 4 is a block diagram of an information extraction device shown according to an exemplary embodiment.

[0062] Figure 5 is a block diagram of an electronic device shown according to an exemplary embodiment.

[0063] Figure 6 is a block diagram of another electronic device shown according to an exemplary embodiment. Detailed Description of the Embodiment

[0064] The following provides a detailed description of the specific embodiments of the present disclosure with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining and understanding the present disclosure, and are not used to limit the present disclosure.

[0065] Figure 1 is a flowchart of an information extraction method shown according to an exemplary embodiment of the present disclosure. As Figure 1 shown, the information extraction method may include the following steps.

[0066] In step S101, at least one statement in the surgical record to be extracted is vectorized to obtain at least one corresponding vectorized statement.

[0067] Among them, the surgical record usually contains relevant information recorded during one or more surgeries of a certain patient, which may include the surgical procedures used during the surgery, that is, the surgical methods, such as XX resection. The surgical record usually also includes relevant information related to each surgical procedure, such as consumable information (such as consumables like bandages, adhesive tapes, cotton swabs, scalpels, syringes, medical masks, gloves, etc.), drug information (such as anesthetics, disinfection drugs, etc.), device information (such as ventilators, oxygen cylinders, etc.), and at least one of the ICD coding information corresponding to the surgical procedure (such as ICD-9 coding). It can be understood that the classification methods of the above consumable information, drug information, and device information are exemplary and may belong to different classifications in different classification criteria. For example, in one possible classification method, consumable information and drug information are classified into one category. Specifically, reference can be made to the classification criteria in the actual situation, and the present disclosure does not make any limitations.

[0068] In step S102, the translation model is used to translate the at least one vectorized statement to obtain the surgical procedures corresponding to each vectorized statement, and the surgical procedures corresponding to each vectorized statement are one or more.

[0069] It can be understood that the surgical record will record one or more surgical procedures and the relevant information of each surgical procedure, such as the name of each surgical procedure, the detailed process, the consumables used, the drugs, etc. In some complex descriptions, multiple surgical procedures may be described simultaneously in one statement, and there may also be some content unrelated to the surgical procedure mixed in the surgical record, such as exploration, intraoperative pathology of tumor surgery, etc. Therefore, it is necessary to determine which information is the surgical procedure in each statement of the surgical record. For the above reasons, the relationship between the statements of the surgical record and the surgical procedures may have multiple mapping relationships, such as one-to-one, one-to-many, many-to-one, and many-to-many. Among them, one-to-one means that one statement corresponds to one surgical procedure, one-to-many means that one statement corresponds to multiple surgical procedures, many-to-one means that multiple statements correspond to the same surgical procedure, and many-to-many means that multiple statements correspond to multiple surgical procedures. Therefore, the mapping relationship from the surgical record to the surgical procedure can be regarded as a translation mapping relationship, and the surgical procedure is regarded as the label of the statement, so that the surgical procedure in the statement can be identified through the translation model, where the translation model is a trained translation model.

[0070] In step S103, the sequence labeling model is used to extract information from the surgical procedures corresponding to each vectorized statement to obtain the target information corresponding to the surgical procedures corresponding to each vectorized statement.

[0071] Among them, sequence tagging is a simple NLP (Natural Language Processing) task. The scope of sequence tagging is relatively wide and can be used to solve a series of problems of classifying characters, such as word segmentation, part-of-speech tagging, named entity recognition, relation extraction, and so on. For the sequence tagging model, first introduce the observation sequence. The observation sequence is the input of the sequence tagging model. The observation sequence is generally a sentence, so it can also be called a word sequence. The state sequence or tag sequence is the output of the sequence tagging model. In the embodiments of the present disclosure, the sequence tagging model can be any trained sequence tagging model, and the present disclosure does not limit the model structure of the sequence tagging model.

[0072] In the embodiments of the present disclosure, using the sequence tagging model to extract information from the surgical method corresponding to each vectorized sentence can be regarded as a task of named entity recognition for the elements in the sentence. When the surgical method has been determined, each sentence with the determined surgical method is used as the input of the sequence tagging model. Through the sequence tagging model, the elements (words, phrases, and / or short sentences) in each sentence are located and classified, and the attribute information (or called tags) of each element is recognized, such as person names, organization names, locations, times, consumables, drugs, and so on, so as to extract the target information related to the surgical method. Among them, the target information is the key information corresponding to the surgical method, such as the above-mentioned consumable information, drug information, device information, ICD coding information, etc.

[0073] For example, in the sentences of the surgical record, for each surgical method, there is a record of the use of consumables, including the name, quantity, etc. of the consumables. Therefore, when a certain sentence is input into the sequence tagging model, the sequence tagging model can label each element in the sentence based on the determined surgical method. For example, if the input sentence is "XXX resection, using two scalpels and 5 packs of gauze", then "scalpel" and "gauze" in the sentence are labeled as the names of consumables, "two" is labeled as the quantity of the consumable "scalpel", and "5 packs" is labeled as the quantity of the consumable "gauze". Among them, the consumable name and the consumable quantity are both used as the above-mentioned consumable information.

[0074] It is worth mentioning that when extracting the target information, for the attribute information of each sentence extracted by the above sequence tagging model, all the attribute information can be selected as the target information, or some information groups of interest can be selected as the target information. Which information in the attribute information extracted from the sentence is actually determined as the target information corresponding to the surgical method can refer to the requirements in the actual medical insurance review standard. For example, the ICD coding information can refer to "ICD-9-CM3 Medical Insurance Version", and the present disclosure does not make a limit.

[0075] Through the above technical solution, the surgical procedures in the operation records can be automatically identified, and the target information corresponding to each surgical procedure can be extracted. Compared with the existing method of manual extraction, the efficiency of surgical procedure identification and information extraction can be significantly improved, and the problem of easy errors in the manual method can be avoided, thereby improving the accuracy rate.

[0076] Optionally, Figure 2 is a flowchart of another information extraction method shown according to an exemplary embodiment of the present disclosure. As Figure 2 shown, the information extraction method may further include the following steps.

[0077] In step S104, the target information corresponding to the surgical procedure corresponding to each vectorized statement is verified by using the preset standard surgical procedure information.

[0078] It can be understood that the surgical procedure corresponding to each vectorized statement is the surgical procedure corresponding to each statement in the above-mentioned surgical record to be reviewed. Therefore, the target information corresponding to the surgical procedure is the target information corresponding to the surgical procedure corresponding to each statement in this surgical record. Therefore, by verifying the target information, the review of this surgical record can be realized, and the reimbursement of the expenses corresponding to this surgical record can be realized after the review is passed.

[0079] Among them, the above standard surgical procedure information includes at least one of standard consumable information, standard drug information, standard instrument information, and standard ICD coding information corresponding to various surgical procedures. That is, when verifying the target information corresponding to each surgical procedure, the standard information corresponding to different types of information can be referred to. For example, the consumable information can be verified by the standard consumable information, and the same applies to other types of information. The standard consumable information, standard drug information, standard instrument information, and standard ICD coding information are industry standards pre-established by authoritative departments. In practice, the medical insurance review standards formulated by relevant departments can be referred to. For example, the ICD coding information can refer to "ICD-9-CM3 Medical Insurance Edition", and the present disclosure does not make any limitations.

[0080] In step S105, when the target information corresponding to the surgical procedure corresponding to each vectorized statement all passes the verification, it is determined that the surgical record passes the review.

[0081] In step S106, when the target information corresponding to any surgical procedure corresponding to any vectorized statement fails to pass the verification, the surgical procedure that fails to pass the verification and the target information that fails to pass the verification are output.

[0082] Exemplarily, in the target information corresponding to each surgical procedure that has been extracted, it includes the ICD-9 code corresponding to the surgical procedure provided by the hospital side. Through the above method, the standard ICD-9 code corresponding to the surgical procedure can be queried, so as to verify whether the ICD-9 code provided by the hospital side is consistent with the standard ICD-9 code. If they are inconsistent, it is determined that the surgical procedure fails. At this time, a prompt message can be output to inform the surgical procedure that fails the verification, and to inform that the surgical procedure fails the verification because of the inconsistent ICD-9 code. For the consumable information or other types of target information, the same principle applies and will not be listed one by one. In the case where any target information of any surgical procedure fails the verification, a prompt message indicating that the surgical record fails the review can be output, or a prompt message indicating that the target information of the surgical procedure fails the review can be output, and other surgical procedures that pass the verification are allowed to be reimbursed. Actually, it can be executed according to the reimbursement regulations of the relevant part, and the present disclosure does not make any restrictions.

[0083] Optionally, Figure 3 is a flowchart of a vectorization processing method shown according to an exemplary embodiment of the present disclosure. As Figure 3 shown, the information extraction method may further include the following steps. Step S101 includes the following steps:

[0084] In step S1011, the first statement in the surgical record is subjected to overall vectorization processing to obtain a first vector.

[0085] Exemplarily, in one way, the overall vectorization processing of the statement can be performed through Google's BERT (Bidirectional Encoder Representations from Transformers) model or Sentence Transformer. Taking BERT as an example, the network architecture of the BERT model uses a multi-layer Transformer structure. Transformer is an encoder-decoder structure, and encoder-decoder is a common model framework in deep learning, which is usually formed by stacking several encoders and decoders. The essence of BERT is to learn a suitable feature representation for words by running self-supervised learning methods with a large amount of corpus, and it can fuse the context information of deep bidirectional language representations. Therefore, in the present disclosure, the statements of the surgical procedure can be vectorized through BERT. Through the above processing, the features of the statement can be subjected to embedding processing in machine learning. Embedding processing can convert a large sparse vector into a low-dimensional space that preserves semantic relationships, so as to project the features in the statement into a low-dimensional vector space.

[0086] In step S1012, perform word segmentation on the first statement.

[0087] In step S1013, after performing vectorization on each word obtained by word segmentation of the first statement, combine them to obtain a second vector.

[0088] Among them, in an optional manner, step S1013 may include the following steps:

[0089] Perform vectorization on each of these words to obtain word vectors corresponding to each of these words.

[0090] Multiply each word vector by a corresponding coefficient, and sum each word vector and its corresponding product to obtain a second vector.

[0091] Exemplarily, for each word obtained by word segmentation, the Word2Vec word vector model can be used for vectorization. Word2vec is one of the Word Embedding methods and belongs to the field of NLP. It can transform words into computable and structured vectors. For example, the word sequence obtained by word segmentation of a surgical record statement is "A, B, A, C, B, F, G". It is desired that each different word in this sequence obtains a corresponding vector representation. Then, the Word2Vec model can be used to perform vectorization on the sequence "A, B, A, C, B, F, G". Suppose the vector corresponding to A is [1 2 -0.5], and the vector corresponding to B is [-0.2 9 7]. Similarly, words C - F can also be represented as vectors in the above form, so that the words can be vectorially represented.

[0092] The value range of the coefficient corresponding to each word vector is [0, 1], which can be understood as the weight corresponding to this word. By using different coefficients, words with important semantic meanings can obtain greater weights. The coefficient corresponding to each word vector can be determined according to its IDF (Inverse Document Frequency) value.

[0093] In step S1014, concatenate the first vector and the second vector to obtain the vectorized statement corresponding to the first statement.

[0094] Optionally, the translation model involved in each embodiment of the present disclosure is an attention model.

[0095] Exemplarily, the attention model, namely the Attention model or the Attention translation model. For the surgical records in the embodiments of the present disclosure, since the entire surgical record contains a very large number of sentences, all sentences can be regarded as a long text of multiple words (each sentence is regarded as a word), so that the Attention translation model can be used for translation to determine the surgical procedures corresponding to each sentence. Optionally, a Seq2Seq model based on Attention can be adopted. Seq2Seq is a network with an Encoder–Decoder structure. For example, when inputting a surgical record of a brain tumor operation, the Attention translation model outputs four surgical procedures: brain tumor resection operation, internal decompression operation, external decompression operation, and cerebrospinal fluid shunt operation.

[0096] Furthermore, the training method of the translation model used in the above method will be introduced below. This method may include the following steps:

[0097] (a) Obtain a plurality of surgical record samples. Each surgical record sample includes at least one sample sentence, and each sample sentence has at least one labeled surgical procedure. It can be understood that, as the surgical record samples are used as training samples, each sentence has been labeled with the surgical procedure(s) it contains.

[0098] (b) For each surgical record sample among the plurality of surgical record samples, perform vectorization processing on at least one sample sentence in the surgical record sample to obtain at least one corresponding vectorized sentence. Among them, the method of vectorizing the sentence is the same as the method described in step S101 and will not be elaborated here.

[0099] (c) Use the vectorized sentences corresponding to each sample sentence in the plurality of surgical record samples as input data, and use at least one labeled surgical procedure of each sample sentence as verification data to train the initial translation model to obtain the translation model. Exemplarily, this training method may include:

[0100] (c1) For each sample sentence, input the vectorized sentence corresponding to the sample sentence into the initial translation model to obtain at least one surgical procedure corresponding to the sample sentence output by the initial translation model; where the initial translation model may be the above-mentioned Attention translation model, and the initial parameters when establishing this model can be determined according to prior knowledge, which is not limited in the present disclosure.

[0101] (c2) Update the initial translation model according to at least one pre-labeled surgical procedure of the sample sentence and at least one surgical procedure corresponding to the sample sentence output by the initial translation model.

[0102] Execute the steps from (c1) to (c2) again using the next sample statement until the initial translation model meets the first preset condition, and then use it as the translation model that has completed training.

[0103] The first preset condition can be achieved by examining the difference between at least one surgical procedure corresponding to the sample statement output by the translation model and at least one pre-marked surgical procedure of the sample statement. For example, a loss function can be established, and the value of this loss function can be used to represent the magnitude of the difference. When the value of the loss function is less than a certain threshold, it is determined that the first preset condition is met.

[0104] Next, the training method of the sequence labeling model used in the above method will be introduced. This method may include the following steps:

[0105] (d) Obtain multiple surgical record samples. Each surgical record sample includes at least one sample statement, and each sample statement has one or more pre-marked surgical procedures, as well as sequence labels corresponding to the one or more surgical procedures. The sequence labels are used to represent the attribute information of each word in the sample statement. Among them, the attribute information can be information such as consumables, drugs, instruments, doctor information, ICD codes, etc. of the surgical procedure. All or part of the attribute information in the attribute information is used as the above-mentioned target information.

[0106] (e) Use the multiple surgical record samples as the input data of the initial sequence labeling model, and use the sequence labels corresponding to one or more surgical procedures of each sample statement in the multiple surgical record samples as the verification data to train the initial sequence labeling model to obtain the sequence labeling model. Exemplarily, the training method may include:

[0107] (e1) For each sample statement, input the vectorized statement corresponding to the sample statement and the recognized surgical procedures corresponding to the sample statement into the initial sequence labeling model, and obtain the attribute information corresponding to the surgical procedures of the sample statement output by the initial sequence labeling model.

[0108] (e2) Update the initial sequence labeling model according to the one or more pre-marked surgical procedures of each sample statement, the sequence labels corresponding to the pre-marked surgical procedures, and the attribute information corresponding to the surgical procedures of the sample statement output by the initial sequence labeling model.

[0109] Execute the steps from (e1) to (e2) again using the next sample statement until the initial sequence labeling model meets the second preset condition, and then use it as the sequence labeling model that has completed training. The second preset condition is the same as the principle of the above first preset condition, and the above loss function method can also be used to determine whether the training is completed.

[0110] Through the above technical solution, it is possible to automatically identify the surgical procedures in the operation record and extract the target information corresponding to each surgical procedure. Compared with the existing method of manual extraction, it can significantly improve the efficiency of surgical procedure identification and information extraction, and can avoid the problem of easy errors in the manual method, thereby improving the accuracy rate.

[0111] Figure 4 It is a block diagram of an information extraction device shown according to an exemplary embodiment. As Figure 4 shown, the information extraction device 400 includes:

[0112] A processing module 401, configured to perform vectorization processing on at least one statement in the operation record to be extracted, so as to obtain at least one vectorized statement corresponding thereto;

[0113] A translation module 402, configured to use a translation model to translate at least one vectorized statement, so as to obtain the surgical procedures corresponding to each vectorized statement, and the surgical procedures corresponding to each vectorized statement are one or more;

[0114] An information extraction module 403, configured to use a sequence labeling model to extract information from the surgical procedures corresponding to each vectorized statement, so as to obtain the target information corresponding to the surgical procedures corresponding to each vectorized statement.

[0115] Optionally, the target information corresponding to any surgical procedure includes at least one of the consumable information, drug information, instrument information, and ICD coding information corresponding to the surgical procedure. The information extraction device 400 further includes:

[0116] A verification module, configured to use the preset standard surgical procedure information to verify the target information corresponding to the surgical procedures corresponding to each vectorized statement; wherein, the standard surgical procedure information includes at least one of the standard consumable information, standard drug information, standard instrument information, and standard ICD coding information corresponding to multiple surgical procedures;

[0117] A determination module, configured to determine that the operation record passes the review when the target information corresponding to the surgical procedures corresponding to each vectorized statement all passes the verification;

[0118] An output module, configured to output the surgical procedures that fail to pass the verification and the target information that fails to pass the verification when the target information corresponding to any surgical procedure corresponding to any vectorized statement fails to pass the verification.

[0119] Optionally, the processing module 401 includes:

[0120] A vectorization processing sub-module, configured to perform holistic vectorization processing on the first statement in the operation record to obtain a first vector;

[0121] A word segmentation sub-module, configured to perform word segmentation processing on the first statement;

[0122] A merging sub-module, configured to perform vectorization processing on each word obtained after segmenting the first statement and then merge them to obtain a second vector;

[0123] A splicing sub-module, configured to splice the first vector and the second vector to obtain a vectorized statement corresponding to the first statement.

[0124] Optionally, the merging sub-module includes:

[0125] A word vector processing sub-module, configured to perform vectorization processing on each word respectively to obtain a word vector corresponding to each word;

[0126] A calculation sub-module, configured to multiply each word vector by a corresponding coefficient and sum the products of each word vector and the corresponding coefficient to obtain a second vector.

[0127] Optionally, the translation model is an attention model.

[0128] Optionally, the apparatus 400 further includes: a first training module, and the first training module is configured to:

[0129] Obtain a plurality of surgical record samples, each surgical record sample includes at least one sample statement, and each sample statement has at least one tagged surgical procedure;

[0130] For each surgical record sample among the plurality of surgical record samples, perform vectorization processing on at least one sample statement in the surgical record sample to obtain at least one corresponding vectorized statement;

[0131] Use the vectorized statements corresponding to each sample statement in the plurality of surgical record samples as input data, and use at least one tagged surgical procedure of each sample statement as verification data to train an initial translation model to obtain a translation model.

[0132] Optionally, the apparatus 400 further includes: a second training module, and the second training module is configured to:

[0133] Obtain a plurality of surgical record samples, each surgical record sample includes at least one sample statement, each sample statement has one or more pre-tagged surgical procedures, and a sequence annotation corresponding to the one or more surgical procedures, and the sequence annotation is used to characterize the attribute information of each word in the sample statement;

[0134] Use the plurality of surgical record samples as input data of an initial sequence annotation model, and use the sequence annotation corresponding to one or more surgical procedures of each sample statement in the plurality of surgical record samples as verification data to train the initial sequence annotation model to obtain a sequence annotation model.

[0135] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0136] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of an information extraction method provided by the present disclosure are implemented.

[0137] Figure 5 is a block diagram of an electronic device 500 shown according to an exemplary embodiment. As Figure 5 shown, the electronic device 500 may include: a processor 501, a memory 502. The electronic device 500 may also include one or more of a multimedia component 503, an input / output (I / O) interface 504, and a communication component 505.

[0138] Among them, the processor 501 is used to control the overall operation of the electronic device 500 to complete all or part of the steps in the above information extraction method. The memory 502 is used to store various types of data to support the operation of the electronic device 500. These data may include, for example, instructions for any application or method operating on the electronic device 500, as well as application-related data, such as contact data, received and sent messages, pictures, audio, video, and so on. The memory 502 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The multimedia component 503 may include a screen and an audio component. Among them, the screen may be a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, and the microphone is used to receive external audio signals. The received audio signals may be further stored in the memory 502 or sent through the communication component 505. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 504 provides an interface between the processor 501 and other interface modules. The above other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 505 is used for wired or wireless communication between the electronic device 500 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more of them is not limited here. Therefore, the corresponding communication component 505 may include: a Wi-Fi module, a Bluetooth module, an NFC module, and so on.

[0139] In an exemplary embodiment, the electronic device 500 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the above information extraction method.

[0140] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the above information extraction method are implemented. For example, the computer-readable storage medium can be the above-mentioned memory 502 including program instructions, and the above program instructions can be executed by the processor 501 of the electronic device 500 to complete the above information extraction method.

[0141] Figure 6 is a block diagram of another electronic device 600 shown according to an exemplary embodiment. For example, the electronic device 600 can be provided as a server. Referring to Figure 6 , the electronic device 600 includes a processor 622, the number of which can be one or more, and a memory 632 for storing computer programs executable by the processor 622. The computer programs stored in the memory 632 can include one or more modules each corresponding to a set of instructions. In addition, the processor 622 can be used to execute the computer program to execute the above information extraction method.

[0142] In addition, the electronic device 600 can further include a power supply component 626 and a communication component 650. The power supply component 626 can be used to perform power management of the electronic device 600, and the communication component 650 can be used to implement communication of the electronic device 600, for example, wired or wireless communication. In addition, the electronic device 600 can further include an input / output (I / O) interface 658. The electronic device 600 can operate based on an operating system stored in the memory 632, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, and so on.

[0143] In another exemplary embodiment, there is also provided a computer-readable storage medium including program instructions, which, when executed by a processor, implement the steps of the above information extraction method. For example, the non-transitory computer-readable storage medium may be the above-mentioned memory 632 including program instructions, and the above program instructions may be executed by the processor 622 of the electronic device 600 to complete the above information extraction method.

[0144] In another exemplary embodiment, there is also provided a computer program product, which includes a computer program capable of being executed by a programmable device, and the computer program has a code portion for executing the above information extraction method when executed by the programmable device.

[0145] The preferred embodiments of the present disclosure have been described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.

[0146] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the present disclosure will not separately describe various possible combination manners.

[0147] In addition, any combination can be made between various different embodiments of the present disclosure as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.

Claims

1. An information extraction method, characterized in that, comprising: Performing vectorization processing on at least one statement in the surgical record to be extracted to obtain at least one corresponding vectorized statement; The performing vectorization processing on at least one statement in the surgical record to be extracted to obtain at least one corresponding vectorized statement includes: performing holistic vectorization processing on the first statement in the surgical record to obtain a first vector; performing word segmentation processing on the first statement; performing vectorization processing on each word obtained after word segmentation of the first statement and then merging them to obtain a second vector; splicing the first vector and the second vector to obtain the vectorized statement corresponding to the first statement; Using a translation model to translate the at least one vectorized statement to obtain the surgical procedures corresponding to each vectorized statement, and each surgical procedure corresponding to a vectorized statement is one or more; Using a sequence labeling model to extract information from the surgical procedures corresponding to each vectorized statement to obtain the target information corresponding to the surgical procedures corresponding to each vectorized statement.

2. The information extraction method according to claim 1, characterized in that, The target information corresponding to any surgical procedure includes at least one of the consumable information, drug information, instrument information, and ICD coding information corresponding to the surgical procedure, and the method further includes: Using the preset standard surgical procedure information to verify the target information corresponding to the surgical procedures corresponding to each vectorized statement; wherein, the standard surgical procedure information includes at least one of the standard consumable information, standard drug information, standard instrument information, and standard ICD coding information corresponding to multiple surgical procedures; When the target information corresponding to the surgical procedures corresponding to each vectorized statement all passes the verification, determining that the surgical record passes the review; When the target information corresponding to any surgical procedure corresponding to each vectorized statement fails to pass the verification, outputting the surgical procedure that fails to pass the verification and the target information that fails to pass the verification.

3. The information extraction method according to claim 1, characterized in that, The performing vectorization processing on each vector obtained after word segmentation of the first statement and then merging them to obtain a second vector includes: Performing vectorization processing on each of the words to obtain the word vectors corresponding to the words; Multiplying each word vector by a corresponding coefficient and summing the products of each word vector and the corresponding coefficient to obtain the second vector.

4. The information extraction method according to claim 1, characterized in that, The translation model is an attention model.

5. The method according to any one of claims 1-4, characterized in that, The training method of the translation model includes: Obtaining a plurality of surgical record samples, each surgical record sample includes at least one sample statement, and each sample statement has at least one labeled surgical procedure; For each surgical record sample in the plurality of surgical record samples, performing vectorization processing on at least one sample statement in the surgical record sample to obtain at least one corresponding vectorized statement; Using the vectorized statements corresponding to each sample statement in the multiple surgical record samples as input data, and using at least one surgical procedure already marked for each sample statement as verification data, train the initial translation model to obtain the translation model.

6. The method according to any one of claims 1-4, wherein, the training method of the sequence labeling model includes: Obtain multiple surgical record samples, each surgical record sample includes at least one sample statement, each sample statement has one or more pre-marked surgical procedures, and a sequence label corresponding to the one or more surgical procedures, the sequence label is used to characterize the attribute information of each word in the sample statement; Using the multiple surgical record samples as input data for the initial sequence labeling model, and using the sequence labels corresponding to the one or more surgical procedures corresponding to each sample statement in the multiple surgical record samples as verification data, train the initial sequence labeling model to obtain the sequence labeling model.

7. An information extraction device, wherein, comprising: A processing module, configured to vectorize at least one statement in the surgical record to be extracted to obtain at least one corresponding vectorized statement; The processing module is further configured to: perform an overall vectorization process on the first statement in the surgical record to obtain a first vector; perform a word segmentation process on the first statement; perform vectorization on each word obtained after word segmentation of the first statement and then merge them to obtain a second vector; splice the first vector and the second vector to obtain the vectorized statement corresponding to the first statement; A translation module, configured to use a translation model to translate the at least one vectorized statement to obtain the surgical procedures corresponding to each vectorized statement, and each vectorized statement corresponds to one or more surgical procedures; An information extraction module, configured to use a sequence labeling model to extract information from the surgical procedures corresponding to each vectorized statement to obtain the target information corresponding to the surgical procedures corresponding to each vectorized statement.

8. An electronic device, wherein, comprising: A memory, on which a computer program is stored; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1-6.

9. A non-transitory computer-readable storage medium, on which a computer program is stored, wherein, when the program is executed by a processor, it implements the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Form recognition method and device based on self-attention mechanism, and storage medium

    CN113569840A