Dual verification method for electronic medical record document

Through the N-gram model and classification model combined with the BM25 algorithm and BGE algorithm, the shortcomings of electronic medical record integrity verification are solved, efficient and reliable medical record integrity verification is achieved, and the risk of misclassification when updating and changing medical order dictionary is reduced.

CN120356595APending Publication Date: 2025-07-22BEIJING TIANTAN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510411208.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing technology lacks the integrity verification of the types and contents of electronic medical records, resulting in errors or tampering during the archiving of medical records, affecting medical quality and doctor-patient relationship.

Method used

The N-gram model and classification model are used to combine the BM25 algorithm and the BGE algorithm with dual detection method, and through matching, prediction and manual review, the accuracy of medical order items classification is ensured.

Benefits of technology

It improves the efficiency and accuracy of the integrity verification of medical record documents, reduces the risk of misclassification when updating and changing medical order dictionaries, and ensures the reliability of electronic medical record documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356595A_ABST
    Figure CN120356595A_ABST
Patent Text Reader

Abstract

The invention provides a dual verification method for an electronic medical record document, and belongs to the technical field of medical record verification, and the method comprises the steps: 1, obtaining medical advice items inputted based on the electronic medical record document, and matching the medical advice items with existing items in a medical advice dictionary one by one; 2, if the matching is unsuccessful, predicting a first category of the doctor's advice items based on an N-gram model and a classification model, and determining a second category of the doctor's advice items by using a BM25 algorithm and a BGE algorithm; 3, judging whether the first category is consistent with the second category or not, and if yes, judging that the medical advice item passes verification; otherwise, turning to manual judgment of the first category and the second category. Through the dual mechanism, the integrity verification of the medical record document caused by the change / addition of the hospital medical advice dictionary is more efficient, reliable and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical record verification, and particularly to a dual verification method for electronic medical record documents. Background Art

[0002] As a medium for the paperless medical record, the medical record filing system needs to integrate all the documents of electronic medical records according to the business system. The main interface modes between the medical record system and other systems are mainly two interaction methods: Web-Service and view. For those with relatively high requirements for the real-time performance of data information interaction, the WebService solution is usually adopted; while for those with a large amount of interactive data and high requirements for data security, the view mode is usually used. Currently, no matter which of the above methods, there is a lack of integrity verification for the types and contents of medical record documents.

[0003] The integrity verification of medical record documents is to prevent errors or tampering during the medical record filing process to ensure the accuracy and reliability of electronic medical record documents. The integrity verification of medical record documents is related to aspects such as the diagnosis and treatment effects of patients, medical quality, and doctor-patient relationships, and is crucial for medical institutions. With the promotion and implementation of the paperless electronic medical records in hospitals and the frequent occurrence of medical disputes caused by incomplete medical records and defective medical record records, it is urgent for hospitals to apply technologies such as big data and artificial intelligence to develop automated tools to perform integrity verification in aspects such as document structure, information integrity, and consistency of key data points, so as to help hospitals better manage digital medical records, improve data integrity, and reduce the risk of potential medical record disputes.

[0004] Therefore, the present invention proposes a dual verification method for electronic medical record documents. Summary of the Invention

[0005] The present invention provides a dual verification method for electronic medical record documents, which, when dealing with the change or addition of hospital doctor's order dictionary items, combines the N-gram + classification model and the BM25 + BGE similarity calculation dual detection method, can not only improve the classification accuracy of new doctor's order dictionary items, but also reduce the risk of misclassification when the dictionary is updated and changed. When the prediction results of the two methods are consistent, it is automatically used as the final classification result; if they are inconsistent, manual review is passed to ensure the accuracy of the final classification. Through such a dual mechanism, the integrity verification of medical record documents caused by the change / addition of hospital doctor's order dictionary is more efficient, reliable, and accurate.

[0006] The present invention provides a dual verification method for electronic medical record documents, including:

[0007] Step 1: Obtain the doctor's order items input based on the electronic medical record documents, and match each of the doctor's order items with the existing items in the doctor's order dictionary one by one;

[0008] Step 2: If the match is unsuccessful, predict the first category of the medical advice item based on the N-gram model and the classification model. Meanwhile, use the BM25 algorithm and the BGE algorithm to determine the second category of the medical advice item;

[0009] Step 3: Determine whether the first category is consistent with the second category. If they are consistent, it is determined that the verification of the medical advice item passes;

[0010] Otherwise, transfer to manual determination of the first category and the second category.

[0011] Preferably, after matching the medical advice item with the existing items in the medical advice dictionary one by one, it further includes:

[0012] If the match is successful, return the label in the medical advice dictionary and query the corresponding report document in the corresponding business system.

[0013] Preferably, predicting the first category of the medical advice item based on the N-gram model and the classification model includes:

[0014] Divide the medical advice item into consecutive words of length N, and use a sliding window to convert the consecutive words into N-gram subsequences;

[0015] Collect the pre-labeled medical advice dictionary item samples, perform preprocessing and model training to obtain a classification model, and input the N-gram subsequence corresponding to the medical advice item into the classification model to predict the first category of the medical advice item.

[0016] Preferably, using the BM25 algorithm and the BGE algorithm to determine the second category of the medical advice item includes:

[0017] Determine the matching degree between the medical advice item and each existing item in the medical advice dictionary respectively, sort the matching degrees in descending order, and select the top N required dictionary entries. Meanwhile, use the BGE algorithm to determine the semantic similarity between the medical advice item and the dictionary entries in the medical advice dictionary to capture deep semantic information;

[0018] Calculate the similarity between the required dictionary entries and the deep semantic information to obtain the second category of the medical advice item.

[0019] Preferably, collecting the pre-labeled medical advice dictionary item samples and performing preprocessing and model training to obtain a classification model includes:

[0020] Based on the labeling results, determine the sample category and sample subsequence of each item sample to obtain a sample vector, and input it into a neural network model for training;

[0021] Classify all labeled item samples according to the sample category, and determine the number of samples in each category respectively;

[0022] Determine the clinical probability distribution and the interrogation probability distribution formed by each item sample in its original state;

[0023] Perform clustering analysis on all samples in the same category to obtain the clustering center, the first distance between each item sample in the same category and the clustering center, and the second distance between each item sample and each of the other samples, and obtain the center coefficient;

[0024]

[0025] Among them, Z represents the center coefficient of the corresponding item sample; L j1 represents the first distance between the corresponding item sample and the clustering center; ln represents the logarithmic function symbol; n1 represents the number of item samples involved in the same category; L max 、L min respectively represent the maximum value and the minimum value in all L j1 ; L j1,j2 represents the second distance between the corresponding item sample and the j2th item sample; represents the clustering density of the samples in the corresponding circle drawn with the corresponding item sample as the center and L min as the radius; ρ0 represents the set density;

[0026] Draw a curve for all center coefficients in ascending order of the first distance, and determine the sample balance value for the corresponding category;

[0027]

[0028] Among them, qx represents the linear coefficient of the fitting curve after fitting the curve formed by the center coefficients; a1 represents the linear threshold; Z1, Z n1 、Z max respectively represent the 1st, the n1th, and the largest center coefficients in the drawn curve; represents the variance of all |Z j1 -Y j1 |; Y j1 represents the j1th center coefficient based on the fitting curve; Z j1 represents the j1th center coefficient based on the drawn curve; sumZj represents the sum of the center coefficients of all intersections between the drawn curve and the fitting curve; Z J,sum represents the analysis function based on the intersection points;

[0029] Align the clinical probability distribution with the inquiry probability distribution, lock the probability pairs of the corresponding item samples in the same original state and in the same category, and calculate the reliability coefficient of the corresponding item samples in combination with the sample quantity and sample balance value involved;

[0030] Determine the average coefficient and coefficient variance of all reliability coefficients involved in the same category, determine the specified quantity of random sampling from the coefficient-variance-quantity comparison table, and perform random sampling of the specified quantity from each category, and then input them into the trained model respectively to generate verification results;

[0031] If the verification results are consistent with the annotation results, regard the trained model as a classification model;

[0032] Otherwise, obtain the result difference sets under each classification respectively, and continue to optimize the trained model to obtain a classification model.

[0033] Preferably, obtaining the result difference sets under each classification respectively and continuing to optimize the trained model includes:

[0034] Calculate the difference coefficient of each random sample in the same classification based on the result difference set, where the result difference set contains the differences of each subsequence involved in the verification result and the annotation result of each random sample in the same classification;

[0035] Sort all the random samples in the same classification in descending order according to the difference coefficient, and divide the sorting result in combination with multiple preset difference thresholds to obtain the sample increment under the clinical probability distribution corresponding to the category;

[0036] Amplify the corresponding random samples according to the sample increment to obtain new samples, and continue to train the trained model.

[0037] Preferably, determining the clinical probability distribution and inquiry probability distribution formed by each item sample in the corresponding original state includes:

[0038] Extract the examination order issued for the inquiry record related to the item sample in the original state, the examination result based on the issued examination order, and the re-issued order based on the examination result from the historical examination database respectively, where the re-issued order based on the examination result does not include the two situations of existence and non-existence;

[0039] When it does not exist, determine that the inquiry probability of the corresponding item sample is consistent with the clinical probability and regard it as 1;

[0040] When present, it is determined that the inquiry probability and the clinical probability of the corresponding item sample are inconsistent, and the number of new prescriptions for re-prescription based on the corresponding examination results is counted. The corresponding inquiry probability is regarded as 1 and the corresponding clinical probability is regarded as where N1 represents the corresponding number of new prescriptions;

[0041] Based on the inquiry probability and the clinical probability of each item sample in its original state, the clinical probability distribution and the inquiry probability distribution formed in the corresponding original state are obtained.

[0042] Preferably, calculating the reliability coefficient of the corresponding item sample includes:

[0043]

[0044] where K represents the reliability coefficient of the corresponding item sample; p01 and p02 respectively represent the inquiry probability and the clinical probability in the pair of occurrence probabilities of the corresponding item sample; p1 i1 , p2 i1 respectively represent the inquiry probability and the clinical probability in the i1-th pair of occurrence probabilities of the same item sample in the corresponding same category; m2 represents the number of pairs of occurrence probabilities of the same item sample in the same category.

[0045] Compared with the prior art, the beneficial effects of the present application are as follows:

[0046] When dealing with the change or addition of hospital order dictionary items, combining the N-gram + classification model and the BM25 + BGE similarity calculation dual detection method can not only improve the classification accuracy of new order dictionary items, but also reduce the risk of misclassification when the dictionary is updated and changed. When the prediction results of the two methods are consistent, it is automatically used as the final classification result; if they are inconsistent, manual review is passed to ensure the accuracy of the final classification. Through such a dual mechanism, the integrity verification of medical record documents caused by the change / addition of the hospital order dictionary is more efficient, reliable and accurate.

[0047] Other features and advantages of the present invention will be described in the subsequent description, and, in part, will be obvious from the description, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in the written description and the drawings.

[0048] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings

[0049] The drawings are used to provide a further understanding of the present invention, and constitute a part of the description. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0050] Figure 1 This is a flowchart of a double verification method for an electronic medical record document in an embodiment of the present invention. Detailed implementation manners

[0051] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for illustrating and explaining the present invention, and are not used to limit the present invention.

[0052] The present invention provides a double verification method for an electronic medical record document, as Figure 1 shown, including:

[0053] Step 1: Obtain the medical order items input based on the electronic medical record document, and match each of the medical order items with the existing items in the medical order dictionary one by one;

[0054] Step 2: If the matching is unsuccessful, predict the first category of the medical order item based on the N-gram model and the classification model. At the same time, use the BM25 algorithm and the BGE algorithm to determine the second category of the medical order item;

[0055] Step 3: Determine whether the first category and the second category are consistent. If they are consistent, determine that the verification of the medical order item passes;

[0056] Otherwise, transfer to manual determination of the first category and the second category.

[0057] Preferably, after matching each of the medical order items with the existing items in the medical order dictionary one by one, it further includes:

[0058] If the matching is successful, return the label in the medical order dictionary, and query the corresponding report document in the corresponding business system.

[0059] Preferably, predicting the first category of the medical order item based on the N-gram model and the classification model includes:

[0060] Divide the medical order item into consecutive words with a length of N, and use a sliding window to convert the consecutive words into N-gram subsequences;

[0061] Collect the pre-labeled medical order dictionary item samples, perform preprocessing and model training to obtain a classification model, and input the N-gram subsequence corresponding to the medical order item into the classification model to predict the first category of the medical order item.

[0062] Preferably, using the BM25 algorithm and the BGE algorithm to determine the second category of the medical order item includes:

[0063] Determine the matching degree between the medical order items and each existing item in the medical order dictionary respectively, sort the matching degrees in descending order, and then screen the top N required dictionary entries. At the same time, use the BGE algorithm to determine the semantic similarity between the medical order items and the dictionary entries in the medical order dictionary, and capture deep semantic information;

[0064] Calculate the similarity between the required dictionary entries and the deep semantic information to obtain the second category of the medical order items.

[0065] In this embodiment, based on the label classification of the medical record document system of the medical order form, then search for and match the corresponding report documents for each business system to which the medical order item belongs, and use the card control details to accurately prompt the document missing situation for the query results, so as to standardize the work process of medical record filing and improve the work efficiency of medical record quality control personnel.

[0066] In this embodiment, for the existing medical order dictionary items, it is necessary to cooperate with relevant departments such as medical records and medical technology to compare the medical order dictionary with the business system. For example, the medical order items "47 biochemical items", "low-field strength fast nuclear magnetic resonance MR", and "16 lymph node pathology items" correspond to the inspection system, examination system, and pathology system respectively.

[0067] In this embodiment, the N-gram model helps to capture the context information of the input text and can process the words that do not exist in the medical order dictionary (such as new words, spelling changes, etc.).

[0068] p(w1w2w3…w n ) = p(w1) * p(w2|w1) * p(w3|w1w2) * … * p(w1w2w3…w n-1 )

[0069] Among them, p(w1w2w3…w n ) represents the probability of the entire new medical order item, and w1w2w3…w n represents the word-based representation.

[0070] In this embodiment, BM25 is used to estimate the correlation between the medical order dictionary D and the newly added or changed medical order item Q. The core idea of BM25 is based on the term frequency (TF) and inverse document frequency (IDF), and at the same time, the length information of the document is introduced to calculate the correlation between the document D and the query Q. The basic formula of the BM25 algorithm is:

[0071]

[0072] Among them, Score(D,Q) is the correlation score between the entire medical order dictionary D and the newly added medical order item Q. q is the i-th word in the query. f(q i ,D) is the word q iThe frequency in the rectified doctor's advice dictionary D, IDF(q i ) is the inverse document frequency of the word q i . |D| is the length of the entire doctor's advice dictionary D. avgdl is the average length of all doctor's advice items. k1 and b are adjustable parameters. Usually, k1 is between 1.2 and 2, and b is usually set to 0.75.

[0073] In this embodiment, BGE converts the input doctor's advice text into a vector representation through the BERT model, then calculates the similarity with the BERT embeddings of the top N doctor's advice dictionary entries obtained by rough ranking, and finally obtains the most similar top 1.

[0074] In this embodiment, for each doctor's advice item, the corresponding report document is searched and matched in the affiliated business system. Since the names of the report documents corresponding to each business system are fixed. For example, in the electronic medical record system, each in-patient includes informed documents (agreement on "no giving or receiving red envelopes" between doctors and patients, authorization and signature form for medical activities of in-patients, and general informed consent form), front and back of the medical record front page, progress notes, admission records, discharge records, discharge diagnosis certificates. For surgical patients, surgical records are also required. The corresponding file title of the inspection system is the inspection report form, and the corresponding inspection report form of the examination system. At this time, only according to the business system tags matched / predicted above and the established rules of the documents required by each business system for different types of patients, the integrity of each type of document can be judged from two aspects: quantity and document type.

[0075] The beneficial effects of the above technical solution are: when dealing with the change or addition of doctor's advice dictionary items in the hospital, combining the dual detection methods of N-gram + classification model and BM25 + BGE similarity calculation can not only improve the classification accuracy of new doctor's advice dictionary items, but also reduce the risk of misclassification when the dictionary is updated and changed. When the prediction results of the two methods are consistent, it is automatically used as the final classification result; if they are inconsistent, manual review is passed to ensure the accuracy of the final classification. Through such a dual mechanism, the verification of the integrity of medical record documents caused by the change / addition of the hospital doctor's advice dictionary is more efficient, reliable and accurate.

[0076] The present invention provides a dual verification method for electronic medical record documents, collecting labeled doctor's advice dictionary item samples, and performing preprocessing and model training to obtain a classification model, including:

[0077] Based on the labeling results, determine the sample category and sample subsequence of each item sample, obtain a sample vector, and input it into a neural network model for training;

[0078] Classify all labeled item samples according to the sample category, and respectively determine the number of samples in each category;

[0079] Determine the clinical probability distribution and the inquiry probability distribution formed by each project sample in its original state;

[0080] Perform clustering analysis on all samples in the same category to obtain the clustering center, the first distance between each project sample in the same category and the clustering center, and the second distance between each project sample and each of the other samples, and obtain the center coefficient;

[0081]

[0082] Among them, Z represents the center coefficient of the corresponding project sample; L j1 represents the first distance between the corresponding project sample and the clustering center; ln represents the logarithmic function symbol; n1 represents the number of project samples involved in the same category; L max 、L min respectively represent the maximum value and the minimum value among all L j1 ; L j1,j2 represents the second distance between the corresponding project sample and the j2th project sample; represents the clustering density of the samples in the corresponding circle drawn with the corresponding project sample as the center and Lm in为 as the radius; ρ0 represents the set density;

[0083] Draw curves for all center coefficients in ascending order of the first distance to determine the sample balance value for the corresponding category;

[0084]

[0085] Among them, qx represents the linear coefficient of the fitting curve after fitting the curve formed by the center coefficients; a1 represents the linear threshold; Z1, Z n1 、Z max respectively represent the first, the n1th, and the largest center coefficients in the drawn curve; represents the variance of all |Z j1 -Y j1 |; Y j1 represents the j1th center coefficient based on the fitting curve; Z j1 represents the j1th center coefficient based on the drawn curve; sumZj represents the sum of the center coefficients of all intersections between the drawn curve and the fitting curve; Z J,sum represents the analysis function based on the intersection points;

[0086] Align the clinical probability distribution and the inquiry probability distribution, lock the occurrence probability pairs that are in the same original state as the corresponding project sample and in the same category, and combine the number of samples involved and the sample balance value to calculate the reliability coefficient of the corresponding project sample;

[0087] Determine the average coefficient and coefficient variance of all reliability coefficients involved in the same category, determine the specified number of random samples from the coefficient-variance-quantity comparison table, and perform a specified number of random samples from each category, and then input them into the trained model to generate verification results;

[0088] If the verification result is consistent with the labeling result, the trained model is regarded as a classification model;

[0089] Otherwise, obtain the result difference set under each category respectively, continue to optimize the trained model, and obtain the classification model.

[0090] Preferably, calculating the reliability coefficient of the corresponding project sample includes:

[0091]

[0092] Among them, K represents the reliability coefficient of the corresponding item sample; p01 and p02 represent the inquiry probability and clinical probability of the occurrence probability pair of the corresponding item sample respectively; p1 i1 、p2 i1 They respectively represent the consultation probability and clinical probability of the same item sample in the i1th probability pair corresponding to the same category; m2 represents the number of probability pairs of the same item sample under the same category.

[0093] In this embodiment, the labeling results are labeled in advance by the doctor, and the sample category and sample subsequence can be directly obtained, and the sample vector = {sample subsequence sample category}.

[0094] In this embodiment, the sample categories include, for example, medical records, admission records, discharge records, etc. Therefore, directly classifying the existing relevant samples can obtain the relevant quantity, and the quantity under each category is greater than 100 or above.

[0095] In this embodiment, the original state corresponding to the project sample refers to a series of operations including a patient registering to see a doctor and a doctor issuing a checkup form.

[0096] In this embodiment, the consultation probability is 1.

[0097] In this embodiment, the clinical probability is

[0098] In this embodiment, clinical probability distribution: clinical probability of the item sample in the original state;

[0099] Probability distribution of consultation: the probability of consultation of the sample of items in the original state.

[0100] It should be noted that a series of steps such as hospitalization, surgery, and discharge after the patient's diagnosis are all based on the project samples in the original state, and the original states corresponding to the project samples involved in subsequent hospitalization, surgery, and discharge are all the states corresponding to the projects at the time of the first diagnosis.

[0101] In this embodiment, the clustering analysis is implemented by using the K-nearest neighbor algorithm. Then, the clustering center can be directly obtained after the clustering analysis to directly obtain the required first distance and second distance.

[0102] In this embodiment, each sample corresponds to a first distance and a center coefficient. Therefore, the center coefficients can be sorted according to the size of the distance, and a relevant curve can be plotted.

[0103] In this embodiment, the alignment process refers to aligning according to the project samples.

[0104] In this embodiment, the result of the alignment process is:

[0105] The interrogation probability of sample 01 The interrogation probability of sample 02 The interrogation probability of sample 03

[0106] The clinical probability of sample 01 The clinical probability of sample 02 The clinical probability of sample 03

[0107] In this embodiment, for example, project sample 01 and project sample 02 under category 1 are in the same original state, and thus there are 2 occurrence probability pairs.

[0108] In this embodiment, the coefficient-variance-quantity comparison table contains the average coefficient, coefficient variance, and the corresponding specified quantity under different categories, which are all preset to ensure the reliable verification of the trained model as much as possible.

[0109] In this embodiment, the verification result includes: output subsequence, output category.

[0110] The beneficial effects of the above technical solution are: clustering the samples under the same category to draw a curve based on the center coefficient, obtaining the sample balance coefficient under this category to ensure the unevenness of the samples, and then combining the clinical probability distribution and the interrogation probability distribution to lock the occurrence probability pairs, calculate the reliable coefficient, effectively screen a specified number of sample pairs for model verification, thereby ensuring the continuous optimization accuracy of the model and improving the verification accuracy.

[0111] The present invention provides a double verification method for electronic medical record documents, which respectively obtains the result difference sets under each classification and further optimizes the trained model, including:

[0112] Calculate the difference coefficient of each random sample under the same classification based on the result difference set, where the result difference set contains the differences of each subsequence involved in the verification result and the annotation result of each random sample under the same classification;

[0113] Sort all the random samples under the same classification in descending order according to the difference coefficient, and divide the sorting result in combination with multiple preset difference thresholds to obtain the sample increment under the clinical probability distribution corresponding to the category;

[0114] Amplify the corresponding random samples according to the sample increment to obtain new samples, and continue to train the trained model.

[0115] In this embodiment, the value of the difference of each subsequence is respectively obtained from the sequence difference - value comparison table, and all the values are accumulated to obtain the difference coefficient, where the comparison table contains the differences of different subsequences and the values matched with the differences.

[0116] In this embodiment, the difference coefficient is divided into difference levels according to the preset difference threshold, and the number of difference coefficients under each difference level is respectively obtained. Combining with the set sample size under the single difference quantity of the difference level, the required sample size under the corresponding difference level is obtained. Then, adding up all the required sample sizes under the same category can obtain the sample increment.

[0117] In this embodiment, amplification means that the random sample quantity is replicated and amplified according to the sample increment to obtain new samples.

[0118] The beneficial effects of the above technical solutions are: By calculating the difference coefficient under each classification and dividing the sorting result by the preset difference threshold to obtain the sample increment, the training accuracy is guaranteed.

[0119] The present invention provides a double - verification method for electronic medical record documents, which determines the clinical probability distribution and the interrogation probability distribution formed by each item sample in its original state, including:

[0120] Extract from the historical examination database the issued examination list of the interrogation record related to the item sample in the original state, the examination result based on the issued examination list, and the re - issued list based on the examination result, where the re - issued list based on the examination result does not include the two situations of existence and non - existence;

[0121] When it does not exist, it is determined that the interrogation probability of the corresponding item sample is consistent with the clinical probability and is regarded as 1;

[0122] When it exists, it is determined that the inquiry probability of the corresponding item sample is inconsistent with the clinical probability, and the number of new issuances of the re-issued order based on the corresponding examination results is counted, and the corresponding inquiry probability is regarded as 1 and the corresponding clinical probability is regarded as wherein, N1 represents the corresponding number of new issuances;

[0123] Based on the inquiry probability and the clinical probability of each item sample in the original state, the clinical probability distribution and the inquiry probability distribution formed in the corresponding original state are obtained.

[0124] In this embodiment, the inquiry record is the basis for determining the original state.

[0125] In this embodiment, the historical examination database includes different inquiry records and corresponding item samples, examination results, issued examination orders, etc.

[0126] In this embodiment, for example, the doctor issues examination order A1 according to the face-to-face consultation description of the patient. However, after the patient undergoes the examination according to examination order A1, there is no relevant disease information.

[0127] In this embodiment, the value of the number of new issuances is greater than or equal to 1.

[0128] The beneficial effect of the above technical solution is: starting from the original state to obtain the examination results and the re-issued order, which is convenient for reasonably determining the inquiry probability and the clinical probability.

[0129] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A double verification method for electronic medical record documents, characterized in that, Including: Step 1: Obtain the medical order items input based on the electronic medical record documents, and match each of the medical order items with the existing items in the medical order dictionary one by one; Step 2: If the match is unsuccessful, predict the first category of the medical order item based on the N-gram model and the classification model. At the same time, use the BM25 algorithm and the BGE algorithm to determine the second category of the medical order item; Step 3: Determine whether the first category and the second category are consistent. If they are consistent, it is determined that the verification of the medical order item passes; Otherwise, transfer it to the manual for determining the first category and the second category.

2. The double verification method for the electronic medical record text according to claim 1, characterized in that After matching each of the medical order items with the existing items in the medical order dictionary, it further includes: If the match is successful, return the label in the medical order dictionary and query the corresponding report document in the corresponding business system.

3. The double verification method for electronic medical record texts according to claim 1, characterized in that, Predicting the first category of the medical order item based on the N-gram model and the classification model includes: Divide the medical order item into consecutive words of length N, and use a sliding window to convert the consecutive words into N-gram subsequences; Collect the pre-labeled medical order dictionary item samples, perform preprocessing and model training to obtain a classification model, and input the N-gram subsequence corresponding to the medical order item into the classification model to predict the first category of the medical order item.

4. The double verification method for electronic medical record text according to claim 1, characterized in that Using the BM25 algorithm and the BGE algorithm to determine the second category of the medical order item includes: Respectively determine the matching degree between the medical order item and each existing item in the medical order dictionary, sort the matching degrees in descending order, and select the top N required dictionary entries. At the same time, use the BGE algorithm to determine the semantic similarity between the medical order item and the dictionary entries in the medical order dictionary to capture deep semantic information; Calculate the similarity between the required dictionary entries and the deep semantic information to obtain the second category of the medical order item.

5. The double-checking method for electronic medical record texts according to claim 3, wherein Collect the pre-labeled medical order dictionary item samples, perform preprocessing and model training to obtain a classification model, including: Based on the annotation results, determine the sample category and sample subsequence of each item sample to obtain a sample vector, and input it into a neural network model for training; Classify all the labeled item samples according to the sample category, and respectively determine the number of samples in each category; Determine the clinical probability distribution and the interrogation probability distribution formed by each item sample in its original state; Perform clustering analysis on all samples in the same category to obtain the cluster center, the first distance between each item sample in the same category and the cluster center, and the second distance between each item sample and each other sample, to obtain the center coefficient; Among them, Z represents the central coefficient of the corresponding project sample; L j1 represents the first distance between the corresponding project sample and the clustering center; ln represents the logarithmic function symbol; n1 represents the number of project samples involved under the same category; L max and L min respectively represent the maximum value and the minimum value among all L j1 ; L j1,j2 represents the second distance between the corresponding project sample and the j2-th project sample; represents the clustering density of the samples in the corresponding circle drawn with the corresponding project sample as the center and L min as the radius; ρ0 represents the set density; Draw a curve of all the center coefficients in ascending order of the first distance to determine the sample balance value for the corresponding category; Among them, qx represents the linear coefficient of the fitting curve after fitting the drawn curve composed of central coefficients; a1 represents the linear threshold; Z1, Z n1 , Z max respectively represent the 1st, n1th, and largest central coefficients in the drawn curve; represents the variance of all |Z j1 -Y j1 |; Y j1 represents the j1th central coefficient based on the fitting curve; Zj1 represents the j1th central coefficient based on the drawn curve; sumZj represents the sum of the central coefficients of all intersections between the drawn curve and the fitting curve; Z J,sum represents the analysis function based on the intersection points; Align the clinical probability distribution and the interrogation probability distribution, lock the pair of the occurrence probabilities of the same original state and in the same category corresponding to the item sample, and combine the number of samples involved and the sample balance value to calculate the reliability coefficient of the corresponding item sample; Determine the average coefficient and coefficient variance of all reliability coefficients involved in the same category, determine the specified number of random samplings from the coefficient-variance-quantity comparison table, and perform the specified number of random samplings from each category, and then input them into the trained model respectively to generate verification results; If the verification results are consistent with the annotation results, regard the trained model as a classification model; Otherwise, obtain the result difference sets under each classification respectively, and continue to optimize the trained model to obtain a classification model.

6. The double verification method for the electronic medical record text according to claim 5, characterized in that, Obtain the result difference sets under each classification respectively, and continue to optimize the trained model, including: Calculate the difference coefficients of each random sample in the same classification based on the result difference sets, where the result difference sets contain the differences of each subsequence involved in the verification results and annotation results of each random sample in the same classification; Sort all random samples in the same classification in descending order according to the difference coefficients, and divide the sorting results in combination with multiple preset difference thresholds to obtain the sample increment under the clinical probability distribution corresponding to the category; Amplify the corresponding random samples according to the sample increment to obtain new samples, and continue to train the trained model.

7. The double-checking method for electronic medical record texts according to claim 5, characterized in that, Determine the clinical probability distribution and consultation probability distribution formed by each project sample in its original state, including: Extract from the historical examination database the examination order for the consultation record related to the project sample in its original state, the examination results based on the examination order, and the re-issuance order based on the examination results, where the re-issuance order based on the examination results does not include the two situations of existence and non-existence; When it does not exist, determine that the consultation probability of the corresponding project sample is consistent with the clinical probability and regard it as 1; When it exists, it is determined that the inquiry probability of the corresponding item sample is inconsistent with the clinical probability, and the number of new prescriptions for re-prescription based on the corresponding examination results is counted, and the corresponding inquiry probability is regarded as 1 and the corresponding clinical probability is regarded as where N1 represents the corresponding number of new prescriptions; Based on the consultation probability and clinical probability of each project sample in its original state, obtain the clinical probability distribution and consultation probability distribution formed by the corresponding original state.

8. The double verification method for electronic medical record text according to claim 5, characterized in that, Calculate the reliability coefficient of the corresponding project sample, including: Among them, K represents the reliability coefficient of the corresponding project sample; p01 and p02 respectively represent the interrogation probability and the clinical probability in the pair of occurrence probabilities of the corresponding project sample; p1 i1 , p2 i1 respectively represent the interrogation probability and the clinical probability in the i1-th pair of occurrence probabilities of the same project sample in the corresponding same category; m2 represents the number of pairs of occurrence probabilities of the same project sample under the same category.