Medical data quality control rule automatic generation method and system, terminal and medium

By generating medical data quality control rules through the BERT model and large language model, combined with dynamic updates of clustering algorithms, the problem that traditional manual rules are difficult to adapt to the diversity and changes of medical data is solved, and automated quality control and improved data accuracy are achieved.

CN120809027APending Publication Date: 2025-10-17NORTH CHINA DIGITAL HEALTH TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510671113.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional medical data quality control methods rely on manually designed rules, which are difficult to adapt to the diversity and complexity of medical data. They also lack the ability to track and analyze dynamic changes in data. Rule updates rely on manual feedback and are difficult to adapt to the ever-changing medical data environment.

Method used

The BERT model is used to extract feature vectors of medical data, and quality control rules are generated through fine-tuning of the large language model. The effectiveness of the rules is evaluated based on correct and incorrect data, and the clustering algorithm is used to analyze error patterns and dynamically update the quality control rules.

Benefits of technology

It realizes automated quality control of medical data, improves data quality and accuracy, dynamically adapts to data changes, and enhances medical decision-making support capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120809027A_ABST
    Figure CN120809027A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical data processing, and particularly provides a medical data quality control rule automatic generation method and system, a terminal and a medium, and the method comprises the steps: obtaining medical data, marking correct data and wrong data in the medical data, and carrying out the feature extraction of different types of medical data, and obtaining corresponding feature vectors; standard medical information of the data is obtained, and the standard medical information is converted into natural language cue words; after fine tuning training is carried out on the large language model, a corresponding quality control rule is output, and the validity of the quality control rule is evaluated based on correct data and wrong data; the method comprises the steps of screening out error data in medical data based on an existing quality control rule, obtaining an error mode of the error data through a clustering algorithm, matching the error mode with the existing quality control rule, and judging whether a quality control rule needs to be newly added or not based on a matching result. A comprehensive and effective medical data quality control system is constructed, and the medical data quality and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of medical data processing, and particularly relates to a medical data quality control rule automatic generation method, system, terminal and medium. BACKGROUND

[0002] Under the current digital wave, medical and health data is growing explosively. From detailed medical records of daily outpatient and inpatient, to images and test reports generated by various advanced medical testing equipment, to data collected by emerging telemedicine and personal health monitoring devices, the sources are extensive and the quantity is huge. Therefore, the quality control of medical data becomes extremely critical.

[0003] The traditional medical data quality control method relies on manual design of rules. According to the actual needs and specifications of medical business, the staff determines the field information that must be included in the data record, manually formulates the corresponding judgment standard for different types of data to ensure its accuracy, and designs the logical relationship rules between data according to medical knowledge and clinical experience.

[0004] Because the sources of medical data are extensive, the rules designed manually are often based on certain experience and limited conditions, which are difficult to adapt to the diversity and complexity of medical data; because medical data is in dynamic change, the rules designed manually lack effective tracking and analysis capabilities for dynamic changes of data; and in the current manual design rule method, the update of rules mainly depends on manual feedback, and the rule system is difficult to adapt to the changing medical data environment. SUMMARY

[0005] In view of the above shortcomings of the prior art, the present application provides a medical data quality control rule automatic generation method, system, terminal and medium to solve the above technical problems.

[0006] In a first aspect, the present application provides a medical data quality control rule automatic generation method, comprising: Obtaining medical data, labeling correct data and incorrect data in the medical data, and dividing the medical data into structured data, semi-structured data and unstructured data; Based on the BERT model, the features of different types of medical data are extracted to obtain corresponding feature vectors; Respectively obtaining the standard medical information of structured data, semi-structured data and unstructured data, and converting the standard medical information into natural language prompt words; The feature vectors and corresponding natural language prompt words are input into a large language model, the large language model is fine-tuned and trained, and the corresponding quality control rules are output, and the effectiveness of the quality control rules is evaluated based on the correct data and incorrect data; Obtaining new medical data, screening error data in the medical data based on existing quality control rules, obtaining error patterns of the error data through a clustering algorithm, matching the error patterns with the existing quality control rules, and determining whether to add new quality control rules based on the matching result.

[0007] In an optional embodiment, the structured data has a clear data model and a fixed format, including patient basic information, diagnosis information, test and examination results, treatment information, surgery information, and electronic medical record text. Semi-structured data contains partially structured elements, including hospital clinical diagnosis records, patient interview records, medical impact reports, and XML or JSON format medical data. Unstructured data has no fixed structure, including doctor's handwritten medical records, patient's self-description of illness, medical literature and research reports, and medical conference records and discussions.

[0008] In an optional embodiment, the BERT model is used to extract features from different types of medical data to obtain corresponding feature vectors: Different types of medical data are preprocessed respectively; The preprocessed medical data sequence is converted into corresponding word embedding, segment embedding and position embedding, and the three embeddings are added to obtain an input embedding vector; The input embedding vector is input into the pre-trained BERT model, which includes multiple layers of Transformer encoders. In each layer of Transformer, self-attention calculation and feedforward neural network calculation are performed on the input; After processing by multiple layers of Transformer, the BERT model outputs a hidden state vector corresponding to each word, and the feature vector is obtained by combining the hidden state vectors corresponding to each word.

[0009] In an optional embodiment, converting the standardized medical information into natural language prompts specifically includes: The structured data is parsed to obtain medical terms, which are mapped to a standard term set. The integrity of the structured data is checked and missing information is supplemented to obtain standardized structured data. The semi-structured data is parsed to extract text, and the extracted text is expressed in a unified medical term. The medical terms in the text are mapped to a standard term set. The integrity of the structured data is checked and missing information is supplemented to obtain standardized semi-structured data. The unstructured data is preprocessed to obtain text, medical entities in the text are recognized based on natural language processing technology, the recognized medical entities are matched and standardized with a standard medical terminology set, and the entity names are converted into standard terminology expressions; relationship extraction technology is used to obtain the relationships between the medical entities in the text, the standardized medical entities and the relationships therebetween are integrated to form complete medical information descriptions, important information missing in the text is supplemented, and standardized unstructured data is obtained. Different types of standardized data are used to generate respective corresponding natural language prompt words, and the natural language prompt words include a target information positioning part, a key information description part, an operation or judgment indication part, and a supplementary explanation or special situation prompt part.

[0010] In an optional implementation, the fine-tuning training of the large language model specifically includes: The feature vectors and the corresponding natural language prompt words are used as a data set, the data set is divided into a training set, a validation set and a test set, and the weights of the large language model pre-trained on a large-scale corpus are loaded into the model. The large language model is trained based on the training set, in each round of training, the feature vectors and the natural language prompt words are input, the model encodes the feature vectors, encodes the natural language prompt words based on the word embedding technology, fuses the encoding of the natural language prompt words with the representation of the encoded feature vectors, and continuously updates and optimizes the understanding of the input information through multiple Transformer layers. After being processed by the multiple Transformer layers, a feature representation is obtained, and the output layer of the model converts the feature representation into a theoretical quality control rule.

[0011] Based on the comparison between the theoretical quality control rule and the actual quality control rule, a loss value is calculated, an optimizer calculates a gradient based on the loss value, and model parameters are updated through a backpropagation algorithm. After each round of training is completed, the model is evaluated using the validation set, the loss value and evaluation indicators of the model on the validation set are calculated, and the hyperparameters are adjusted according to the evaluation results. After the training and validation are completed, the fine-tuned model is comprehensively tested using the test set, the performance of the model on the test set is evaluated, and the generalization ability of the model is determined.

[0012] In an optional implementation, the evaluation of the effectiveness of the quality control rule based on the correct data and the error data specifically includes: A set of pre-labeled correct data and error data is constructed, the quality control rule output by the model is applied to the set, for each piece of data, it is judged whether it is error data according to the quality control rule, and the judgment result of each piece of data is recorded. counting the number of correctly recognized correct data, the number of correctly recognized error data, the number of error-recognized correct data, and the number of error-recognized error data based on the judgment result of each piece of data; calculating the accuracy, recall rate, precision, F1 value, and false positive rate of the number of each type of data respectively; Comparing the accuracy of the three types of medical data, if the accuracy of one type of medical data is lower than the other two, it is determined that there is a problem in processing this type of data by the quality control rule; When the recall rate of the medical data is lower than the recall threshold, it is determined that there is error data that has not been identified; When the precision is lower than the precision threshold, it is determined that there is a misjudgment; When the F1 value is lower than the F1 threshold, it is determined that the overall performance of the quality control rule on the medical data is poor; When the false positive rate is higher than the false positive threshold, it is determined that there is a correct data that is misjudged as error data; When the above determination conditions exist, the quality control rule is partially effective or invalid, and the quality control rule is adjusted based on the specific problems that occur.

[0013] In an optional implementation, determining whether to add a quality control rule specifically includes: Applying each existing quality control rule to each record in the medical data set in turn, marking the error data, and constructing an error data set from all the error data; Clustering the error data based on a clustering algorithm and calculating the cosine similarity between the data points, dividing the error data set into different clusters based on the cosine similarity between the data points, and each cluster representing a potential error mode; Extracting the key features of each error mode, comparing the key features of the error mode with the key information of the quality control rule, and calculating the matching degree; When the matching degree of the error mode and a certain existing quality control rule exceeds the matching threshold, it is determined to be matched; otherwise, it is determined to be not matched; when it is determined to be not matched, a corresponding quality control rule is added.

[0014] In a second aspect, the present application provides a medical data quality control rule automatic generation system, which implements the medical data quality control rule automatic generation method described above, and the system includes: A data acquisition module acquires medical data, labels correct data and error data in the medical data, and divides the medical data into structured data, semi-structured data, and unstructured data; A feature extraction module extracts features based on a BERT model to obtain corresponding feature vectors for different types of medical data; The prompt construction module respectively acquires the standardized medical information of the structured data, the semi-structured data and the unstructured data, and converts the standardized medical information into natural language prompt words; The rule generation module inputs the feature vector and the corresponding natural language prompt word into a large language model, outputs corresponding quality control rules after fine-tuning training of the large language model, and evaluates the effectiveness of the quality control rules based on correct data and error data; The rule updating module acquires new medical data, screens error data in the medical data based on existing quality control rules, acquires error patterns of the error data through a clustering algorithm, matches the error patterns with the existing quality control rules, and judges whether new quality control rules need to be added based on the matching result.

[0015] In a third aspect, a terminal is provided, comprising: a processor and a memory, the memory is configured to store a computer program, the processor is configured to call and run the computer program from the memory, so that the terminal executes the method of the terminal described above.

[0016] In a fourth aspect, a computer readable storage medium is provided, which stores instructions when running on a computer, so that the computer executes the method described in the above aspects.

[0017] The medical data quality control rule automatic generation method, system, terminal and medium provided by the application have the beneficial effects that the medical data is acquired and labeled as correct or incorrect, the data types are divided, the feature vector is extracted based on the BERT model, the standardized medical information is acquired and converted into prompt words, the prompt words are input into the large language model for fine-tuning training to output the quality control rules, the rule effectiveness can be evaluated according to the correct and error data, the error data can be screened from the new medical data, the error patterns can be clustered and analyzed, and the existing rules can be matched to judge whether new rules need to be added, so that a comprehensive and effective medical data quality control system is constructed, the medical data quality and accuracy are improved, and strong support is provided for medical decision-making.

[0018] In addition, the design principle of the application is reliable, the structure is simple, and the application prospect is very wide. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0020] Figure 1is a schematic flow chart of a medical data quality control rule automatic generation method of an embodiment of the present application.

[0021] Figure 2 is a schematic block diagram of a medical data quality control rule automatic generation system of an embodiment of the present application.

[0022] Figure 3 is a structural schematic diagram of a terminal provided by an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order for those skilled in the art to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of the present application.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terminology used in the description of the present application herein only for the purpose of describing specific embodiments, and is not intended to limit the present application.

[0025] The medical data quality control rule automatic generation method provided by the embodiments of the present application is executed by a computer device, and accordingly, the medical data quality control rule automatic generation system runs in the computer device.

[0026] Figure 1 is a schematic flow chart of a medical data quality control rule automatic generation method of an embodiment of the present application. In which, Figure 1 The execution subject can be a medical data quality control rule automatic generation system. According to different needs, the order of steps in the flow chart can be changed, and some can be omitted.

[0027] As Figure 1 shown, the method comprises: Step S1, obtaining medical data, labeling correct data and error data in the medical data, and dividing the medical data into structured data, semi-structured data and unstructured data; Collect medical data from multiple channels such as hospital information systems and medical databases, organize professional medical personnel to label correct and error data according to medical standards and experience, and then divide the data into three categories according to whether there is a fixed format and structure characteristics, i.e. structured (such as patient basic information table), semi-structured (such as XML format medical record) and unstructured (such as doctor's handwritten medical record); Clearing medical data lays the foundation for subsequent in-depth analysis and processing. Different types of data correspond to different processing methods, which improves data processing efficiency and pertinence. Accurate labeling of true and false data helps to find data problems and improve data quality.

[0028] Step S2, feature extraction of different types of medical data based on BERT model to obtain corresponding feature vectors; First, preprocess different types of medical data, normalize structured data values, clean text data, etc., and then input the processed data into the pre-trained BERT model. Through the internal Transformer layer calculation of the model, the vector representation that can represent the data semantics and features is obtained.

[0029] Complex medical data is converted into a vector form that is easy for computers to process, key features are extracted, and important data information is retained, facilitating subsequent model learning and analysis, and improving the ability to understand and process medical data.

[0030] Step S3, respectively obtaining structured data, semi-structured data and unstructured data of standard medical information, and converting the standard medical information into natural language prompts; For structured data, the field content is standardized according to medical terminology standards; after parsing semi-structured data, text specifications are extracted; and for unstructured data, natural language processing technology is used to identify and standardize entities. Then, according to the standard information, generate simple and easy-to-understand natural language prompts, such as "Please verify whether the patient's diagnosis meets the standard".

[0031] Make medical information more standardized and unified, easy to understand and communicate, and natural language prompts reduce the threshold of data processing, allowing medical personnel and related systems to quickly understand the key issues and processing direction of the data.

[0032] Step S4, input the feature vector and corresponding natural language prompt into the large language model, and output the corresponding quality control rules after fine-tuning the large language model, and evaluate the effectiveness of the quality control rules based on correct data and error data; Integrate the feature vector and natural language prompt into training samples, input them into the large language model, adjust the model parameters using cross-entropy loss function and optimizer, output the quality control rules after training, and calculate the accuracy, recall rate and other indicators based on the labeled correct and error data to evaluate the effectiveness of the rules.

[0033] Generate quality control rules that fit the characteristics of medical data to improve data quality control capabilities. Through evaluation, the advantages and disadvantages of the rules are clear, providing a basis for optimizing the rules and ensuring the accuracy and reliability of medical data.

[0034] Step S5, obtain new medical data, filter out error data in the medical data based on existing quality control rules, obtain error mode of error data through clustering algorithm, match error mode with existing quality control rules, and judge whether new quality control rules need to be added based on matching result.

[0035] Continuously collect new medical data, filter error data with existing quality control rules, analyze error data with K-Means clustering algorithm, classify similar errors into a class to form error mode, and compare and match with existing rules. According to the matching degree, it is decided whether to add new rules.

[0036] Dynamic monitoring of new medical data quality, timely discovery of errors, mining of error rules, continuous improvement of quality control rule system, improvement of rule coverage and adaptability, and better protection of medical data quality.

[0037] Optionally, as an embodiment of the present application, structured data has a clear data model and fixed format, including: Electronic medical record text; Patient basic information: including name, gender, age, date of birth, nationality, ID number, contact information, family address and other clear data fields. These information is usually stored in the form of table in the database, which is convenient for query and analysis. For example, in the patient management system of the hospital, the basic information of each patient is recorded in the corresponding field.

[0038] Diagnosis information: disease diagnosis name and corresponding international disease classification code (such as ICD-10 code), diagnosis time, diagnosis doctor's name, etc. For example, the patient is diagnosed as "pneumonia", and the corresponding ICD-10 code is J18.9. These information is recorded in a structured way, which is convenient for statistical analysis of the distribution of diseases.

[0039] Test results: the name, result value, unit, reference range, test time of various test items. For example, the results of white blood cell count, red blood cell count, hemoglobin content in blood routine test, and the detection results of blood sugar, blood lipid, liver function and other indexes in biochemical test are presented in structured data.

[0040] Treatment information: the name of treatment measures (such as drug treatment, surgical treatment, etc.), treatment start time, end time, drug name, dose, administration route, etc. For example, the patient received "amoxicillin" drug treatment, the dose was 0.5g each time, 3 times a day, and intravenous infusion. These treatment information are accurately recorded.

[0041] Surgery information: surgery name, surgery time, lead surgeon, anesthesia method, surgery site, etc. For example, record the relevant information of an "appendectomy", including the specific time of the start and end of the surgery, the name of the lead surgeon, etc.

[0042] Semi-structured data contains partially structured elements, including: Hospital clinical diagnosis records, patient interview records; Medical image reports: image examination (such as X-ray, CT, MRI, etc.) reports usually contain descriptions of findings and diagnostic opinions, which are presented in text form, but the description of the location, size, shape, etc. of the lesion has certain structured characteristics. For example, "chest CT shows a 2.0cm x 1.5cm nodule in the left upper lobe, with clear boundaries", the size and location of the nodule are clear.

[0043] XML or JSON format medical data: Some medical information systems use XML or JSON format to store and transmit data, which has a certain structure, but not as strict as traditional database tables. For example, a patient health record stored in JSON format, which contains the patient's basic information, medical history, allergy history, etc. The data is organized in the form of key-value pairs.

[0044] Unstructured data has no fixed structure, including: Doctor's handwritten medical records: In some medical institutions, there are still some doctors' handwritten medical records, which record patients' illness, symptoms, diagnosis process, etc. in the form of text, without fixed format and structure, completely in the form of free text. For example, the doctor describes the patient's onset process, medical history, family history, etc. in detail, and these information needs to be extracted by natural language processing technology.

[0045] Patient's self-description of illness: When patients tell their doctors about their illness, their description is usually free expression without fixed pattern. For example, the patient may say "I have been feeling a headache for the past few days, and sometimes I also feel dizzy, especially when I stand up", which needs to be analyzed and sorted out by the doctor.

[0046] Medical literature and research reports: A large number of medical literature, research reports, academic papers, etc. in the medical field contain rich medical knowledge and research results, but these contents are in the form of unstructured text. For example, a research paper on a new treatment method for a certain disease, the experimental method, result discussion, etc. are described in detail through words.

[0047] Medical conference records and discussions: records of case discussion conferences within hospitals, academic exchange conferences, etc., usually record the content of the conference, the opinions and suggestions of experts, etc. in the form of text, and these records are also unstructured. For example, the discussion process of a difficult case and the diagnosis and treatment suggestions of experts in the conference.

[0048] Optionally, as an embodiment of the present application, step S2 specifically comprises: Different types of medical data are preprocessed respectively; the numerical data in the structured medical data is normalized, the corresponding analysis tool (such as an XML parser, a JSON parser, etc.) is used for the semi-structured data, and the unstructured data is segmented into individual words or tokens through word segmentation processing.

[0049] The preprocessed medical data sequence is converted into corresponding word embeddings, segment embeddings and position embeddings, and the three embeddings are added to obtain an input embedding vector, wherein: Word embedding: each word is mapped to a fixed-dimensional vector, which represents the semantic information of the word. In the BERT model, each word has a corresponding pre-trained word embedding vector.

[0050] Segment embedding: if the input is multiple text segments (for example, when processing text containing multiple sentences), it is used to distinguish different text segments, and 0 and 1 are usually used to represent different segments.

[0051] Position embedding: since the BERT model is based on the Transformer architecture, it does not have the ability to perceive the position information of the word. Position embedding is used to encode the position information of the word in the text sequence. Different positions of the word have different position embedding vectors.

[0052] The input embedding vector is input into the pre-trained BERT model, and the model includes multiple layers of Transformer encoders. In each layer of Transformer, self-attention calculation and feedforward neural network calculation are performed on the input. After processing by multiple layers of Transformer, the BERT model outputs a hidden state vector corresponding to each word, and a feature vector is obtained by integrating the hidden state vector corresponding to each word.

[0053] Optionally, as an embodiment of the present application, step S4 specifically comprises: The structured data is parsed to obtain medical terms, and the medical terms are mapped to a standard term set; the integrity of the structured data is checked and missing information is supplemented to obtain standardized structured data. The semi-structured data is parsed to extract text, the extracted text is expressed in uniform medical terminology, and the medical terminology in the text is mapped to a standard terminology set; the integrity of the structured data is checked and missing information is supplemented to obtain standardized semi-structured data; The unstructured data is preprocessed to obtain text, medical entities in the text are identified based on natural language processing technology, the identified medical entities are matched and standardized with a standard medical terminology set, and entity names are converted into standard terminology expressions; relationship extraction technology is used to obtain the relationships between medical entities in the text, and the standardized medical entities and their relationships are integrated to form a complete medical information description, supplementing important information missing in the text to obtain standardized unstructured data; According to the standardized data of different types, respective natural language prompt words corresponding thereto are generated, and the natural language prompt words include a target information positioning part, a key information description part, an operation or judgment indication part, and a supplementary explanation or special situation prompt part.

[0054] Optionally, as an embodiment of the present application, some examples of the natural language prompt words include: a. Patient basic information: Target information positioning part: Basic information of patient [name], with hospitalization number [X].

[0055] Key information description part: Contains name, gender, age, contact information, home address, etc., where the age needs to be accurate to a specific value, and the contact information needs to be specific in the form of telephone or email, etc.

[0056] Operation or judgment indication part: Please verify whether these information is accurate, whether the name has any misspelling, whether the gender is consistent with other medical records, whether the age is consistent with the actual situation of the patient, and whether the contact information is valid and contactable.

[0057] Supplementary explanation or special situation prompt part: If the patient is a minor, special attention should be paid to whether the guardian information is complete and accurate; if the contact information is an emergency contact, the relationship between the emergency contact and the patient should be confirmed.

[0058] b. Diagnosis information: Target information positioning part: Diagnosis information of patient [name], with hospitalization number [X].

[0059] Key information description part: Covers disease diagnosis name, corresponding ICD-10 code, diagnosis time, and diagnosis doctor name. The diagnosis name needs to use standard medical terminology, the ICD-10 code needs to be accurate to a specific code, and the diagnosis time needs to be accurate to minutes.

[0060] Operation or judgment instruction part: Check if the diagnosis name matches the ICD-10 code, if the diagnosis time is logically consistent with the patient's onset and treatment time, and if the diagnosis doctor's signature is clear and identifiable and meets the hospital's specifications.

[0061] Supplementary explanation or special situation prompt part: For difficult diagnosis, check if there is expert consultation opinion as the basis for diagnosis; if the diagnosis changes, confirm the change reason and whether the related records are complete.

[0062] c. Medical imaging report: Target information positioning part: Medical imaging report of patient [name], hospital number [X].

[0063] Key information description part: The report contains the findings of the image examination (such as lesion location, size, shape), diagnosis opinion. The lesion location needs to be accurate to the anatomical site, the size has specific numerical value, the shape has accurate description, and the diagnosis opinion needs to be clear.

[0064] Operation or judgment instruction part: Determine if the image examination findings description is accurate, if the diagnosis opinion is reasonably derived based on the image performance, and if the lesion-related information is confirmed by other examination results.

[0065] Supplementary explanation or special situation prompt part: For complex lesion imaging, check if there is analysis of previous imaging data; if the image has artifacts or interference factors, confirm if there is related description in the report.

[0066] d. Doctor's handwritten medical record: Target information positioning part: Handwritten medical record of patient [name] written by [doctor's name].

[0067] Key information description part: The medical record contains disease description, diagnosis thought, treatment suggestion, etc. It is recorded in the form of text without fixed format.

[0068] Operation or judgment instruction part: Check if the handwriting is clear and identifiable, if the disease description is complete, if the diagnosis thought is logically coherent, and if the treatment suggestion is reasonable and operable.

[0069] Supplementary explanation or special situation prompt part: If there are traces of erasing in the medical record, confirm if there is a doctor's signature or a note on the modification reason; for foreign language writing part, confirm if there is accurate translation.

[0070] e. Patient's self-description of disease: Target information positioning part: Patient [name]'s self-description of their own disease.

[0071] Key information description part: including the onset time, symptom manifestation, symptom change, past medical history, etc., mainly based on the patient's free expression.

[0072] Operation or judgment indication part: comb the patient's self-description, judge whether the onset time is clear, whether the symptom description is specific, whether the symptom change is consistent with the time line, and whether the past medical history is consistent with other materials.

[0073] Supplementary explanation or special situation prompt part: if the patient's expression is unclear or logically confused, further inquiry is needed; for the special living habits or environmental factors mentioned by the patient, attention should be paid to whether they are related to the disease.

[0074] Optionally, as an embodiment of the present application, in step S4, the fine-tuning training of the large language model specifically includes: The feature vector and the corresponding natural language prompt word are taken as a data set, and the data set is divided into a training set, a validation set and a test set; the weights of the large language model pre-trained on a large-scale corpus are loaded into the model; The training of the large language model is performed based on the training set. In each round of training, the feature vector and the natural language prompt word are input, the model encodes the feature vector, encodes the natural language prompt word based on the word embedding technology, fuses the encoding of the natural language prompt word with the representation of the encoded feature vector, and continuously updates and optimizes the understanding of the input information through multiple Transformer layers. After processing through multiple Transformer layers, a feature representation is obtained, and the output layer of the model converts the feature representation into theoretical quality control rules.

[0075] Based on the comparison between the theoretical quality control rules and the actual quality control rules, the loss value is calculated, the optimizer calculates the gradient based on the loss value, and the model parameters are updated through the backpropagation algorithm; After each round of training is completed, the model is evaluated using the validation set, the loss value and evaluation indicators of the model on the validation set are calculated, and the hyperparameters are adjusted according to the evaluation results; After training and verification are completed, the fine-tuned model is comprehensively tested using the test set, the performance of the model on the test set is evaluated, and the generalization ability of the model is determined.

[0076] Optionally, as an embodiment of the present application, in step S4, the evaluation of the effectiveness of the quality control rules based on correct data and error data specifically includes: A set of pre-labeled correct data and error data is constructed, the quality control rules output by the model are applied to the set, for each data, whether it is error data is judged according to the quality control rules, and the judgment result of each data is recorded; counting the number of correctly recognized correct data, the number of correctly recognized error data, the number of error-recognized correct data, and the number of error-recognized error data based on the judgment result of each piece of data; calculating the accuracy, recall, precision, F1 value, and false positive rate of the number of each type of data respectively; comparing the accuracy of the three types of medical data, if the accuracy of one type of medical data is lower than the other two, it is determined that there is a problem in processing this type of data by the quality control rule; When the recall rate of the medical data is lower than the recall threshold, it is determined that there is error data that has not been identified; When the precision is lower than the precision threshold, it is determined that there is a misjudgment; When the F1 value is lower than the F1 threshold, it is determined that the overall performance of the quality control rule on the medical data is poor; When the false positive rate is higher than the false positive threshold, it is determined that there is a correct data that is misjudged as error data; When the above determination conditions exist, the quality control rule is partially effective or invalid, and the quality control rule is adjusted based on the specific problems that occur.

[0077] Optionally, as an embodiment of the present application, step S5 specifically comprises: extracting features from the screened error data, wherein Structured data: For structured data in electronic medical records, such as patient basic information, test results, etc., features such as age, gender, test index value, diagnosis code, etc. can be extracted. For example, when analyzing error data of blood routine test results, specific index values such as white blood cell count, red blood cell count, and platelet count are extracted as features.

[0078] Semi-structured data: For semi-structured data such as clinical diagnosis records, features such as diagnosis name, key information in symptom description, and related examination items can be extracted. For example, from the diagnosis record, extract symptom keywords such as "cough" and "fever", and examination item names such as "chest X-ray" and "CT examination".

[0079] Unstructured data: For unstructured data such as doctor's handwritten medical records and patient's self-description, features such as keywords, word frequency, and text sentiment tendency can be extracted with the help of natural language processing technology. For example, use the TF-IDF algorithm to extract important keywords in medical record text, or judge the emotional tendency of patients when describing their illness through sentiment analysis.

[0080] Select a suitable clustering algorithm (such as K-Means) to cluster the error data set after extracting features: First, the number of clusters K needs to be determined. The elbow method, silhouette coefficient method, etc. can be used to determine the appropriate K value. Then, the algorithm will divide the error data into K clusters according to the cosine similarity between data points. The data points in each cluster have high similarity, while the data points between different clusters have large differences. For example, when clustering error diagnosis coding data, the K - Means algorithm may divide data with similar coding error types into the same cluster.

[0081] Observe the common characteristics of error data in each cluster and summarize the error pattern represented by the cluster. For example, error data in a certain cluster are concentrated in a certain age group of patients, and the error type is the abnormal value of a certain test indicator. Therefore, it can be concluded that the error pattern is related to the test indicator of patients in a certain age group.

[0082] Compare the feature differences between different clusters to determine whether there are significantly different error patterns. If the feature differences between different clusters are large, it means that there are multiple types of error patterns; if the differences are small, further adjustment of the clustering parameters or re-extraction of the features may be needed.

[0083] Extract the key features of each error pattern, and compare the key features of the error pattern with the key information of the quality control rule to calculate the matching degree. When the matching degree of the error pattern and a certain existing quality control rule exceeds the matching threshold, it is determined to be matched; otherwise, it is determined to be not matched. When it is determined to be matched, if the error pattern represented by a certain cluster can be completely covered by an existing quality control rule, it means that the error pattern is already within the scope of the rule and no new rule is needed. For example, the existing rule states that "blood glucose values should be between 3.9 - 6.1 mmol / L", and the error data in a certain cluster are all cases where the blood glucose value exceeds this range. Therefore, the error pattern corresponding to the cluster matches the existing rule.

[0084] When it is determined to be matched, if the error pattern of a certain cluster can only be partially covered by the existing rule, or is related to the existing rule but not exactly the same, further evaluation is needed to determine whether the existing rule needs to be refined or supplemented. For example, the existing rule states that "the blood pressure value of a general patient should be within [normal range]", and the error data in a certain cluster are abnormal blood pressure values of patients with a certain disease. In this case, a quality control rule for blood pressure values of patients with a certain disease may need to be supplemented.

[0085] If the error mode of a cluster does not match any existing quality control rule, it means that a new error mode has appeared, and a new quality control rule needs to be added. For example, a new format error of a test report is found in the new medical data, and the existing rule cannot cover this error situation, so a new quality control rule for this format error needs to be added. According to the characteristics and commonalities of the error data in the cluster, the specific content of the quality control rule is formulated. For example, for the newly discovered test report format error, the rule can stipulate that “the test report should contain [specific mandatory items], and the format should comply with [specific format requirements]”.

[0086] In some embodiments, the medical data quality control rule automatic generation system can include a plurality of functional modules composed of computer program segments. The computer programs of each program segment in the medical data quality control rule automatic generation system can be stored in the memory of the computer device and executed by at least one processor to perform the functions of (see Figure 1 described) medical data quality control rule automatic generation.

[0087] In this embodiment, the medical data quality control rule automatic generation system can be divided into a plurality of functional modules according to the functions it performs, as shown in Figure 2 The functional modules of the system can include a data acquisition module, a feature extraction module, a prompt construction module, a rule generation module, and a rule update module. The module referred to by the present application refers to a series of computer program segments that can be executed by at least one processor and can complete a fixed function, which is stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments. The system includes: The data acquisition module acquires medical data, labels correct data and error data in the medical data, and divides the medical data into structured data, semi-structured data, and unstructured data; The feature extraction module extracts features based on the BERT model to obtain corresponding feature vectors for different types of medical data; The prompt construction module acquires the standard medical information of structured data, semi-structured data, and unstructured data respectively, and converts the standard medical information into natural language prompts; The rule generation module inputs the feature vectors and corresponding natural language prompts into a large language model, fine-tunes the large language model, and outputs corresponding quality control rules, and evaluates the effectiveness of the quality control rules based on correct data and error data; The rule update module acquires new medical data, filters error data in the medical data based on existing quality control rules, obtains error modes of the error data through a clustering algorithm, matches the error modes with existing quality control rules, and determines whether new quality control rules need to be added based on the matching results.

[0088] Figure 3 A structural diagram of a terminal is provided for an embodiment of the present application, which can be used to execute the method for automatically generating medical data quality control rules provided by the embodiment of the present application.

[0089] The terminal can include a processor, a memory and a communication unit. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation on the present application. It can be a bus structure or a star structure. It can also include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0090] The memory can be used to store the execution instructions of the processor. The memory can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk. When the execution instructions in the memory are executed by the processor, the terminal can execute some or all of the steps in the following method embodiments.

[0091] The processor is the control center of the storage terminal. It connects all parts of the electronic terminal through various interfaces and lines, executes software programs and / or modules stored in the memory, and calls data stored in the memory, to perform various functions of the electronic terminal and / or process data. The processor can be composed of integrated circuits (ICs), such as a single packaged IC or multiple packaged ICs with the same or different functions. For example, the processor can only include a central processing unit (CPU). In the present embodiment, the CPU can be a single operation core or can include multiple operation cores.

[0092] The communication unit is used to establish a communication channel, so that the storage terminal can communicate with other terminals. It receives user data sent by other terminals or sends user data to other terminals.

[0093] The application further provides a computer storage medium, wherein the computer storage medium can store a program, and the program can include some or all steps in the embodiments provided by the application when executed. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM) or a random access memory (RAM) or the like.

[0094] Those skilled in the art can clearly understand that the technologies in the embodiments of the application can be realized by means of software and necessary general hardware platforms. Based on such understanding, the technical solutions in the embodiments of the application can be embodied in the form of a software product, which is stored in a storage medium such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disc or an optical disc, and includes a plurality of instructions for causing a computer terminal (which can be a personal computer, a server or a second terminal, a network terminal or the like) to execute all or part of the steps of the method described in the embodiments of the application.

[0095] In the specification, the same or similar parts among various embodiments can be referred to each other. Especially, for the terminal embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiments.

[0096] In the several embodiments provided by the application, it should be understood that the disclosed system and method can be implemented in other ways. For example, the division of the system embodiments described above is merely a logical function division, and other division manners can be adopted during actual implementation, for example, a plurality of modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the modules shown or discussed can be indirect coupling or communication connection through some interfaces, and can be electrical, mechanical or in other forms.

[0097] The modules illustrated as separated components can or can not be physically separated, and the components illustrated as modules can or can not be physical modules, i.e., can be located in one place, or can be distributed on a plurality of network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments.

[0098] In addition, each function module in each embodiment of the present application can be integrated in one processing module, or each module can exist physically separately, or two or more modules can be integrated in one module.

[0099] Although the present application has been described in detail with reference to the preferred embodiments, the present application is not limited to the preferred embodiments. Various equivalent modifications or replacements can be made to the embodiments of the present application by those skilled in the art without departing from the spirit and essence of the present application, and these modifications or replacements shall be within the scope of the present application. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and these changes or replacements shall be within the protection scope of the present application.

Claims

1. A method for automatically generating medical data quality control rules, characterized in that: The following steps are involved: Acquire medical data, mark correct and incorrect data in the medical data, and divide the medical data into structured data, semi-structured data, and unstructured data; Based on the BERT model, feature extraction is performed on different types of medical data to obtain corresponding feature vectors; Obtain standardized medical information of structured data, semi-structured data, and unstructured data respectively, and convert the standardized medical information into natural language prompt words; The feature vector and the corresponding natural language prompt word are input into the large language model. After fine-tuning the large language model, the corresponding quality control rules are output and the effectiveness of the quality control rules is evaluated based on the correct data and the incorrect data. Acquire new medical data, filter out erroneous data in the medical data based on existing quality control rules, obtain the error pattern of the erroneous data through clustering algorithms, match the error pattern with the existing quality control rules, and determine whether new quality control rules are needed based on the matching results.

2. The method for automatically generating medical data quality control rules according to claim 1, characterized in that: Structured data has a clear data model and fixed format, including basic patient information, diagnosis information, test results, treatment information, surgical information, and electronic medical record text; Semi-structured data contains some structured elements, including hospital clinical diagnosis records, patient consultation records, medical impact reports, and medical data in XML or JSON format; Unstructured data has no fixed structure and includes doctors’ handwritten medical records, patients’ self-reports of their conditions, medical literature and research reports, medical meeting minutes and discussions.

3. The method for automatically generating medical data quality control rules according to claim 1, characterized in that: Steps for extracting features from different types of medical data based on the BERT model to obtain corresponding feature vectors: Preprocess different types of medical data separately; Convert the preprocessed medical data sequence into corresponding word embedding, segment embedding, and position embedding, and add the three embeddings to obtain the input embedding vector; The input embedding vector is fed into the pre-trained BERT model, which consists of multiple layers of Transformer encoders. In each layer of Transformer, self-attention calculations and feed-forward neural network calculations are performed on the input. After multi-layer Transformer processing, the BERT model outputs the hidden state vector corresponding to each word, and the hidden state vector corresponding to each word is combined to obtain the feature vector.

4. The method for automatically generating medical data quality control rules according to claim 1, characterized in that: Converting standardized medical information into natural language prompts specifically includes: Parse structured data to obtain medical terms and map them to standard terminology sets; check the integrity of structured data and supplement missing information to obtain standardized structured data; Parse semi-structured data to extract text, unify the extracted text into medical terms, and then map the medical terms in the text to a standard terminology set; check the integrity of the structured data and supplement missing information to obtain standardized semi-structured data; Unstructured data is preprocessed to obtain text. Medical entities in the text are identified using natural language processing technology. The identified medical entities are matched and normalized with standard medical terminology, and entity names are converted into standard terminology. Relationship extraction technology is used to obtain the relationships between medical entities in the text. The normalized medical entities and their relationships are integrated to form a complete description of medical information, supplementing the important information missing in the text and obtaining normalized unstructured data. According to the normalized different types of data, corresponding natural language prompt words are generated respectively. The natural language prompt words include target information positioning part, key information description part, operation or judgment instruction part, supplementary explanation or special situation prompt part.

5. The method for automatically generating medical data quality control rules according to claim 1, characterized in that: Fine-tuning the large language model specifically includes: The feature vectors and the corresponding natural language prompt words are used as a dataset, which is divided into a training set, a validation set, and a test set. The weights of the large language model pre-trained on a large corpus are loaded into the model. The large language model is trained based on the training set. In each round of training, feature vectors and natural language prompts are input. The model encodes the feature vectors and then encodes the natural language prompts using word embedding technology. The encoding of the natural language prompts is then fused with the encoded representation of the feature vectors. Through multiple layers of Transformer layers, the model continuously updates and optimizes its understanding of the input information. After processing through these multiple layers of Transformer layers, a feature representation is obtained. The model's output layer converts this feature representation into theoretical quality control rules. Based on the comparison between theoretical quality control rules and actual quality control rules, the loss value is calculated, the optimizer is used to calculate the gradient based on the loss value, and the model parameters are updated through the back propagation algorithm; After each round of training, the model is evaluated using the validation set, the loss value and evaluation index of the model on the validation set are calculated, and the hyperparameters are adjusted based on the evaluation results; After training and validation, the fine-tuned model is fully tested using the test set to evaluate the model's performance on the test set and determine the model's generalization ability.

6. The method for automatically generating medical data quality control rules according to claim 1, characterized in that: The effectiveness of quality control rules based on correct and incorrect data is evaluated as follows: Construct a pre-labeled set of correct and incorrect data, apply the quality control rules output by the model to the set, determine whether each piece of data is incorrect based on the quality control rules, and record the judgment results for each piece of data; Based on the judgment result of each data, the number of correct data correctly identified, the number of incorrect data correctly identified, the number of correct data incorrectly identified, and the number of incorrect data incorrectly identified are counted; Calculate the accuracy, recall, precision, F1 value and false positive rate for each type of data; Compare the accuracy of three types of medical data. If the accuracy of one type of medical data is lower than that of the other two, it is determined that there is a problem with the quality control rule when processing this type of data; When the recall rate of medical data is lower than the recall threshold, it is determined that erroneous data has not been identified; When the accuracy rate is lower than the accuracy threshold, it is determined that there is a misjudgment; When the F1 value is lower than the F1 threshold, it is determined that the overall performance of the quality control rule on the medical data is poor; When the false positive rate is higher than the false positive threshold, it is determined that correct data is misjudged as incorrect data; When the above judgment situation occurs, the quality control rule is partially valid or invalid, and the quality control rule is adjusted based on the specific problems that arise.

7. The method for automatically generating medical data quality control rules according to claim 1, characterized in that: Determining whether new quality control rules are needed includes: Apply each existing quality control rule to each record in the medical data set in turn, mark the erroneous data, and construct an erroneous data set from all the erroneous data; Cluster the error data based on the clustering algorithm and calculate the cosine similarity between data points. Then, divide the error data set into different clusters based on the cosine similarity between data points. Each cluster represents a potential error pattern. Extract the key features of each error pattern and calculate the matching degree by comparing the key features of the error pattern with the key information of the quality control rules; When the degree of matching between the error pattern and an existing quality control rule exceeds the matching threshold, it is determined to be a match; otherwise, it is determined to be a mismatch; when it is determined to be a mismatch, a corresponding quality control rule is added.

8. A medical data quality control rule automatic generation system, characterized by: When implemented, the system executes the method for automatically generating medical data quality control rules according to any one of claims 1 to 7, and the system includes: The data acquisition module acquires medical data, marks correct and incorrect data in the medical data, and divides the medical data into structured data, semi-structured data, and unstructured data; The feature extraction module extracts features from different types of medical data based on the BERT model to obtain corresponding feature vectors; A prompt construction module obtains standardized medical information of structured data, semi-structured data, and unstructured data, and converts the standardized medical information into natural language prompt words; The rule generation module inputs the feature vector and the corresponding natural language prompt word into the large language model, fine-tunes the large language model, outputs the corresponding quality control rules, and evaluates the effectiveness of the quality control rules based on correct and incorrect data; The rule update module obtains new medical data, filters out erroneous data in the medical data based on existing quality control rules, obtains the error pattern of the erroneous data through clustering algorithms, matches the error pattern with the existing quality control rules, and determines whether new quality control rules are needed based on the matching results.

9. A terminal, characterized in that: include: A memory for storing a program for automatically generating medical data quality control rules; A processor is configured to implement the steps of the method for automatically generating medical data quality control rules as described in any one of claims 1 to 7 when executing the program for automatically generating medical data quality control rules.

10. A computer-readable storage medium, characterized in that The readable storage medium stores a medical data quality control rule automatic generation program, which, when executed by a processor, implements the steps of the medical data quality control rule automatic generation method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Executable clinical pathway generation method and device, electronic equipment and storage medium

    CN120995987A