Special disease library medical document intelligent identification method and system, terminal and medium
By constructing a Prompt template and automatically extracting medical labeled data using a large language model, the problem of low efficiency of Chinese medicine documents in specialized disease databases is solved, and efficient and accurate extraction of medical information is achieved.
Patent Information
- Application Number
- CN202510283302.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, medical documents identification in the construction of specialized disease databases rely on manual analysis, resulting in low efficiency, experience-dependent and inconsistent results, which cannot meet the needs of efficient processing of large-scale data.
Based on the type definition of key medical features, a Prompt template is built, combined with natural language processing and large language model (LLM), and automatically extracts medical annotation data, and realizes standardized medical annotation through training and optimization of the model.
It improves the efficiency and accuracy of medical document processing, reduces manual intervention, provides timely and reliable information support, and adapts to the needs of large-scale data processing.
Smart Images

Figure CN120337912A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical documents, and particularly relates to an intelligent recognition method, system, terminal and medium for medical documents in a specialized disease database. Background Art
[0002] The construction of a specialized disease database is an important task in the medical field, aiming to collect, integrate and manage relevant clinical data, patient information, research results, etc. for specific diseases, providing strong support for the diagnosis, treatment, research and management of diseases.
[0003] Currently, the recognition process of medical documents in the construction of specialized disease databases mostly relies on manual analysis of medical record documents and imaging data. The recognition and analysis of medical documents in traditional disease database construction are usually carried out manually by professional medical personnel or imaging experts, relying on personal knowledge, experience and skills.
[0004] This method has certain limitations. Manual analysis is cumbersome and time-consuming, with low speed and efficiency when dealing with a large number of cases, unable to meet the needs of efficient medical treatment. Moreover, the analysis accuracy depends on experience, and there are large differences in subjective judgments among different personnel, which easily leads to inconsistent recognition results and cannot cope with the high-efficiency processing requirements of large-scale data brought by the increasing amount of medical imaging and case document data. Summary of the Invention
[0005] In view of the above deficiencies of the prior art, the present invention provides an intelligent recognition method, system, terminal and medium for medical documents in a specialized disease database to solve the above technical problems.
[0006] In a first aspect, the present invention provides an intelligent recognition method for medical documents in a specialized disease database, including: S1, determining key medical features based on the recognition requirements of various medical documents and defining the types of key medical features; S2, constructing a Prompt template based on the recognition instruction of the medical document, the input medical document data, the relevant background information of the medical document and the output specification requirements after recognition; S3, extracting preliminary medical annotation data from unannotated medical documents based on natural language processing methods combined with the Prompt template; S4, training an LLM model based on the medical document and the medical annotation data related to the medical document combined with the Prompt template; S5, inputting the medical document to be recognized and the corresponding Prompt template into the LLM model to obtain medical annotation data that meets the output specification requirements.
[0007] In an optional implementation manner, the construction of the Prompt template in step S2 specifically includes: Recognition instructions for medical documents, specifying the professional perspective that the model needs to be in during the recognition of medical documents; The input medical document data uses placeholders to determine how to embed the medical document to be recognized and the annotation requirements into the Prompt template; Background information related to medical documents, summarizing the important content of medical documents and giving the key medical features to be extracted; The output specification after recognition stipulates the output structure and the data type of each field in the output in JSON format.
[0008] In an optional implementation manner, step S3 specifically includes: Preprocess the medical document based on natural language processing tools; Based on medical knowledge and the standard expression methods in medical documents, construct a key medical feature extraction rule library, and extract specific medical features from the medical text based on the key medical feature extraction rule library; Based on the recognition instructions of the medical document in the Prompt template, look up the vocabulary related to the recognition instructions in the medical dictionary, and obtain the information in the medical document that matches the vocabulary as a specific medical feature; Organize all the extracted specific medical features into preliminary annotation data based on the format specified by the Prompt template; Verify the preliminary annotation data based on medical logic and medical common sense, and optimize the rule library and medical dictionary based on the verification results.
[0009] In an optional implementation manner, step S4 specifically includes: Organize the medical document and the corresponding preliminary medical annotation data into an input format that conforms to the Prompt template and use it as input data, and divide the input data into a training set, a validation set, and a test set; Set the hyperparameters required during training, train the LLM model based on the training set, the model inputs the medical document and the Prompt template, and the model outputs medical annotation data; Calculate the loss value between the output medical annotation data and the true medical annotation data based on the cross-entropy loss function. When the loss value is greater than the preset threshold, use the backpropagation algorithm to calculate the model gradient, and update the model's parameters based on the Adam optimizer; After each training round, evaluate the performance of the model using the validation set, adjust the hyperparameters based on the evaluation results of the validation set, and retrain; After the model training is completed, evaluate the performance of the model using the test set, and verify the stability and reliability of the model based on the fluctuations in performance.
[0010] In an alternative embodiment, in step S5, when the number of words in the medical document is greater than the threshold, the medical document is split into multiple window texts based on the sliding window method. The multiple texts and the corresponding Prompt templates are input into the LLM model to obtain the medical annotation data for each window text. The medical annotation data for each window text is associated and synthesized to obtain the medical annotation data for the medical document.
[0011] In an alternative embodiment, splitting the medical document into multiple window texts based on the sliding window method specifically includes: Determine the window size and sliding step based on the input limit of the LLM model, the characteristics of the medical document, and the Prompt template; Starting from the starting position of the medical document, intercept the window text based on the window size, and slide the window backward in sequence according to the set sliding step, and intercept the new window text each time until the window moves to the end of the medical document.
[0012] In an alternative embodiment, associating and synthesizing the medical annotation data for each window text specifically includes: For those with consistent annotations for the same key medical feature in different window texts, use this content as the medical annotation data for the current key medical feature; When the same key medical feature has different annotation values in different window texts, for numerical features, merge them into a numerical range, and for text features, merge them into a complete description; Verify whether there are inconsistent situations in the synthesized medical annotation data, and adjust the medical annotation data based on the verification results.
[0013] In a second aspect, the present invention provides a medical document intelligent recognition system for a specialized disease database. When the system is implemented, it executes the above-mentioned medical document intelligent recognition method for the specialized disease database. The system includes: A feature determination module that determines key medical features based on the recognition requirements of various medical documents and defines the types of key medical features; A template construction module that constructs a Prompt template based on the recognition instructions of the medical document, the input medical document data, the relevant background information of the medical document, and the output specification requirements after recognition; A data annotation module that extracts preliminary medical annotation data from unannotated medical documents based on natural language processing methods combined with the Prompt template; A model training module that trains the LLM model based on the medical document, the medical annotation data related to the medical document, and the Prompt template; An output recognition module that inputs the medical document to be recognized and the corresponding Prompt template into the LLM model to obtain medical annotation data that meets the output specification requirements.
[0014] In a third aspect, a terminal is provided, including: a processor and a memory, where the memory is used to store a computer program, the processor is used to call and run the computer program from the memory, so that the terminal executes the method of the above terminal.
[0015] In a fourth aspect, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, the computer is made to execute the methods described in the above aspects.
[0016] The beneficial effects of the present invention are as follows. The intelligent recognition method, system, terminal and medium of medical documents in a specialized disease database provided by the present invention determine key medical features and define types according to the recognition requirements of medical documents, construct a Prompt template including recognition instructions, input data, background information and output specifications, extract preliminary annotation data from unannotated medical documents by means of natural language processing methods combined with the template, then train the LLM model with medical documents and annotation data combined with the template, and finally input the document to be recognized and the corresponding template into the model to obtain standardized medical annotation data, efficiently and accurately extracting key information from medical documents, without relying on the subjective judgment of personnel, and improving the processing efficiency and quality of medical documents.
[0017] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very wide application prospect. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the present invention, the drawings required to be used in the description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 is a schematic flowchart of the intelligent recognition method of medical documents in a specialized disease database according to an embodiment of the present invention.
[0020] Figure 2 is a schematic block diagram of the intelligent recognition system of medical documents in a specialized disease database according to an embodiment of the present invention.
[0021] Figure 3 is a schematic structural diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.
[0024] Natural Language Processing (NLP) is an important branch in the field of artificial intelligence, aiming to enable computers to understand, process, and generate human language.
[0025] LLM, i.e., Large Language Model, is a type of artificial intelligence model in the field of natural language processing with powerful language understanding and generation capabilities.
[0026] The intelligent recognition method for medical documents in a disease-specific database provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the intelligent recognition system for medical documents in a disease-specific database runs in the computer device.
[0027] Figure 1 It is a schematic flowchart of the intelligent recognition method for medical documents in a disease-specific database according to an embodiment of the present invention. Among them, Figure 1 The execution subject can be an intelligent recognition system for medical documents in a disease-specific database. According to different requirements, the order of steps in this flowchart can be changed, and some can be omitted.
[0028] As Figure 1 shown, the method includes: Step S1, determining key medical features based on the recognition requirements of various medical documents and defining the types of key medical features; Organize medical experts and natural language processing experts to jointly discuss various medical documents, such as medical records, inspection reports, diagnosis certificates, etc. According to clinical needs and research purposes, key medical features such as patient basic information (name, age, gender), symptom manifestations (fever, cough), disease diagnosis (pneumonia, diabetes), treatment plans (drug names, surgical methods), etc. are screened out. Subsequently, these features are classified according to data types. For example, patient names and disease names are defined as text types, age is defined as a numerical type, and whether there is an allergy history is defined as a boolean type.
[0029] Identifying key medical features and their types provides clear goals and specifications for subsequent medical document processing and analysis, which helps improve the accuracy and consistency of information extraction.
[0030] Step S2: Construct a Prompt template based on the recognition instructions of medical documents, the input medical document data, the background information related to medical documents, and the requirements of the recognized output specifications. Write clear and specific recognition instructions, such as "Please accurately extract the patient's basic information, symptoms, diagnosis results, and treatment plans from the medical document." Use the medical document to be processed in a specific format (such as "Medical document content: {specific text}") as the input data. Provide background information related to medical documents, such as the explanations of common medical terms and the typical symptoms of diseases. Finally, specify the output format, for example, require it to be output in JSON format, and clearly define the key names and data types of each feature in the JSON structure.
[0031] It provides clear task guidance and input / output specifications for the model, enabling the model to better understand the task requirements, improving the efficiency and accuracy of information extraction. The unified output format facilitates subsequent data processing and analysis, and can effectively reduce errors and biases caused by inconsistent understanding.
[0032] Step S3: Extract preliminary medical annotation data from unannotated medical documents based on natural language processing methods combined with the Prompt template. Apply a pre-trained natural language processing model (such as BioBERT) combined with rule-based methods. First, combine the medical document and the Prompt template into an input text and input it into the pre-trained model. Utilize the language understanding ability of the model to identify key information in the text, and at the same time screen and correct the recognition results through pre-set rules, such as judging the reasonableness of the tumor size based on medical common sense. Finally, extract the preliminary medical annotation data.
[0033] With the help of natural language processing technology and the Prompt template, it is possible to quickly and automatically extract preliminary annotation data from a large number of unannotated medical documents, greatly saving the time and cost of manual annotation.
[0034] Step S4: Train the LLM model based on medical documents, the medical annotation data related to medical documents, and the Prompt template. Medical documents, corresponding medical annotation data, and Prompt templates are combined into a training dataset. Select a suitable large language model (such as GPT-3, ChatGLM, etc.) and perform fine-tuning training on it. During the training process, continuously adjust the model's parameters to minimize the error between the prediction results and the annotation data. Methods such as cross-validation can be used to evaluate the performance of the model to ensure the stability and accuracy of the model on different datasets.
[0035] Step S5, input the medical document to be recognized and the corresponding Prompt template into the LLM model to obtain medical annotation data that meets the output specification requirements.
[0036] Using the trained LLM model and Prompt template, it is possible to efficiently and accurately extract medical annotation data that meets the specifications from the medical document to be recognized. This automated processing method improves the efficiency of medical document processing, reduces manual intervention, and provides timely and reliable information support for medical research, clinical diagnosis, and treatment.
[0037] Optionally, as an embodiment of the present invention, a specific example of step S1 is as follows: In a CT examination record, there are usually multiple information points about tumors. According to the content of the CT examination report and the imaging data, the following key features can be extracted: Whether there is a solid tumor: This feature is a boolean value indicating whether there is a suspicious solid tumor in the examination result. Solid tumors usually refer to those tumors that can be directly observed on CT images through imaging techniques, usually appearing in the form of nodules, masses, etc. If the examination result shows the presence of a tumor, the value of this feature is "True", otherwise it is "False".
[0038] Tumor location: This is a text-based feature indicating the anatomical location where the tumor is located, such as the lung, liver, kidney, etc. The CT report usually clearly indicates the anatomical location of the tumor.
[0039] Number of tumors: This is a numerical feature indicating the number of tumors found in the CT examination. The number of tumors is usually presented in numerical form in the report.
[0040] Optionally, as an embodiment of the present invention, the construction of the Prompt template in step S2 specifically includes: Recognition instructions for medical documents, clarifying the professional perspective that the model needs to be in during the recognition process of medical documents; Input medical document data, using placeholders to determine how to embed the medical document to be recognized and the annotation requirements into the Prompt template; Background information related to medical documents, an overview of the important content of medical documents, and key medical features to be extracted are given; The recognized output specification stipulates the output structure and the data type of each field in the output in JSON format.
[0041] For example, disease information extraction: "Instruction": "You are an experienced medical information analyst. Please accurately extract key disease-related information from the following medical document and strictly output the result in the specified JSON format.", "Background information": "Medical documents may contain content such as patients' symptoms, disease names, disease stages, and whether there are complications. This extraction focuses on several key information: disease name, symptom manifestations, disease stage, and complications. Symptom manifestations refer to abnormal conditions that occur in the patient's body; disease stage is used to describe the development stage of the disease; complications refer to other diseases triggered during the process of suffering from a certain disease.", "Input format": { "The following is the content of the medical document: {medical document text}", "The following are the annotation requirements: {annotation requirements}" } "Output format": { "Disease name": "string", "Symptom manifestations": "string", "Disease stage": "string", "Complications": "string" } If the medical document text is "The patient reported headache and fever in the past week. After examination, it was diagnosed as influenza. Currently, it is in the early stage and there are no complications for the time being.", The model output should be: "Disease name": "Influenza", "Symptom manifestations": "Headache, fever", "Disease stage": "Early", "Complications": "None for the time being.
[0042] Optionally, as an embodiment of the present invention, step S3 specifically includes: Data preprocessing: Use natural language processing tools (such as NLTK, spaCy) to clean medical documents, removing special characters, HTML tags (if any), and irrelevant formatting information. At the same time, convert the text to a unified case form for subsequent processing. For example, use regular expressions to remove useless content such as "[Picture]" in the text. After that, perform word segmentation on the cleaned text, splitting the continuous text into individual words or phrases. A word segmentation tool combined with a medical professional dictionary can be used to improve the accuracy of word segmentation. For example, for "myocardial infarction", an ordinary word segmentation tool may mis-segment it into "myocardium" and "infarction", but combining with a medical dictionary can correctly identify it as a professional term. Finally, perform part-of-speech tagging on the word segmentation results, tagging the part of speech of each word (such as noun, verb, adjective, etc.) to facilitate subsequent analysis of the text structure and semantics based on the part of speech.
[0043] Construct and apply a rule base: Based on medical knowledge and common medical document expressions, construct a set of rule bases. For tumor-related features, set rules such as "If the text appears 'found a space-occupying lesion' and is followed by a word indicating a body part, such as 'lung', 'liver', etc., then it is judged that there may be a tumor, and the tumor location is that body part". When judging the size of a tumor, set the rule "If the text appears keywords such as 'diameter','size approximately', etc., and is followed by a number and a length unit (such as 'cm','mm'), then extract the number and unit as the tumor size". According to these rules, match and judge the text of the preprocessed medical document, and extract key medical features. Traverse the text, and when a text segment that meets the tumor location judgment rule is found, extract the corresponding information as the tumor location annotation data.
[0044] Combine Prompt template with dictionary matching: Associate the feature requirements in the Prompt template with the medical dictionary. For example, the Prompt template requires extracting features such as "whether there is a tumor" and "tumor location". Look up words related to tumors (such as "tumor", "mass", "tumor body", etc.) and words indicating body parts (such as "lung", "liver", "kidney", etc.) in the medical dictionary. Search for these dictionary words in the medical document text. If a word related to a tumor is found, judge whether there is a tumor in combination with the context; if a word indicating a body part is found and there is a certain semantic association with the word related to the tumor (such as in the same sentence), then it is determined as the tumor location. When the text appears "A mass is visible in the lung", through dictionary matching, it can be determined that "whether there is a tumor" is "yes" and "tumor location" is "lung".
[0045] Generate preliminary labeled data: Organize the extracted medical features into preliminary labeled data according to the format specified in the Prompt template. If the Prompt template requires output in JSON format, construct the corresponding JSON structure and fill in the fields with features such as "Whether there is a tumor", "Tumor location", "Tumor size", etc. and their extraction results. If the relevant information of a certain feature is not extracted from the medical document, fill it in according to the regulations (such as filling in "not mentioned" or null). Example of the organized labeled data: {"Whether there is a tumor": "Yes", "Tumor location": "Lung", "Tumor size": "2cm", "Tumor morphology": "not mentioned"}.
[0046] Result verification and optimization: Verify the generated labeled data according to some basic medical logics and common senses. For example, the tumor size cannot be negative, and the tumor location must be a real part of the human body, etc. If it is found that the labeled data does not conform to the logic, return to the previous steps for rechecking and extraction. Continuously optimize the rule library and dictionary through statistical analysis of the labeled results of a large number of medical documents. For example, if it is found that the expressions of certain medical terms vary in different documents, supplement these variants to the medical dictionary to improve the accuracy and integrity of the labeling.
[0047] Optionally, as an embodiment of the present invention, step S4 specifically includes: Organize the medical document and the corresponding preliminary medical labeled data into an input format that conforms to the Prompt template and use it as input data, and divide the input data into a training set, a validation set, and a test set; Set the hyperparameters required during training, and train the LLM model based on the training set. The model takes the medical document and the Prompt template as input, and the model outputs medical labeled data; The hyperparameter settings include the learning rate (learning rate) that controls the step size of model parameter updates, the batch size that represents the number of samples input to the model each time during training, and the number of epochs that represents the number of times the model trains on the entire training dataset.
[0048] In each training batch, combine the medical document and the Prompt template into an input text and input it into the LLM model. The model performs forward propagation based on the input information and outputs medical labeled data.
[0049] Calculate the loss value between the output medical annotation data and the true medical annotation data based on the cross-entropy loss function. When the loss value is greater than the preset threshold, it indicates that there is a large deviation between the prediction result of the model and the true label, and the parameters of the model need to be adjusted. Use the backpropagation algorithm to calculate the model gradient. The gradient represents the rate of change of the loss function with respect to the parameters, and update the parameters of the model based on the Adam optimizer. The Adam optimizer combines the advantages of AdaGrad and RMSProp and can adaptively adjust the learning rate of each parameter to improve the training efficiency and stability.
[0050] After each training epoch, use the validation set to evaluate the performance of the model. Select appropriate evaluation metrics, such as accuracy, precision, recall, F1-score, etc., to measure the performance of the model on the validation set. Adjust the hyperparameters based on the evaluation results of the validation set and retrain the model; for example, if the accuracy does not improve with the increase in the number of training epochs, the learning rate may need to be decreased; if the model shows overfitting, try reducing the batch size or early stopping the training.
[0051] After the model training is completed, use the test set to evaluate the performance of the model. Also use evaluation metrics such as accuracy, precision, recall, F1-score, etc. to measure the performance of the model on the test set. The evaluation results of the test set can reflect the generalization ability of the model on unseen data, and verify the stability and reliability of the model based on the performance fluctuations.
[0052] Optionally, as an embodiment of the present invention, in step S5, when the number of words in the medical document is greater than the threshold, split the medical document into multiple window texts based on the sliding window method, input the multiple texts and the corresponding Prompt templates into the LLM model to obtain the medical annotation data of each window text, and correlate and synthesize the medical annotation data of each window text to obtain the medical annotation data of the medical document.
[0053] Optionally, as an embodiment of the present invention, splitting the medical document into multiple window texts based on the sliding window method specifically includes: Determine the window size and sliding step based on the input limit of the LLM model, the characteristics of the medical document, and the Prompt template; for example, some models support a maximum of 512 tokens as input. When determining the window size and sliding step, ensure that the total length of the window text plus the Prompt template does not exceed the input limit of the model; the language characteristics of the medical document, the density of professional terms, and the length of the Prompt template will all affect the selection of the window size and sliding step. If there are many professional terms and complex sentence structures in the medical document, the window size can be appropriately reduced to ensure that the model can process accurately; when the Prompt template is long, the window size also needs to be adjusted accordingly.
[0054] Starting from the beginning position of the medical document, intercept the window text based on the window size, and slide the window backward in sequence according to the set sliding step length. Each time, intercept the new window text until the window moves to the end of the medical document; Add the corresponding Prompt template to each window text to form the complete input data.
[0055] Input each input data into the LLM model in sequence to obtain the medical annotation data of each window text.
[0056] Optionally, as an embodiment of the present invention, the correlation and synthesis of the medical annotation data of each window text specifically include: For those with the same key medical feature annotated and the same content in different window texts, use this content as the medical annotation data of the current key medical feature; When the same key medical feature has different annotation values in different window texts, for numerical features, merge them into a numerical range, and for text features, merge them into a complete description; Verify whether there are inconsistent situations in the synthesized medical annotation data, and adjust the medical annotation data based on the verification results.
[0057] Optionally, as an embodiment of the present invention, the following is a specific example of medical document extraction: There is a medical document with the content: "Patient Li Si, male, 65 years old, has had cough and expectoration symptoms in the past two weeks, accompanied by low fever. After chest CT examination, a nodule about 2 cm × 3 cm in size was found in the lower lobe of the right lung, with clear boundaries. The preliminary diagnosis is a nodule in the lower lobe of the right lung, and the nature is to be determined. It is recommended to further examine. Cough-relieving and expectorant drugs were given, such as ambroxol tablets, 30 mg each time, 3 times a day." Determination of key medical features and type definition: According to the identification requirements of the medical document, determine the key medical features. For example, the patient's basic information (name, gender, age) is text type and numerical type, the symptom manifestations (cough, expectoration, low fever) are text type, the disease diagnosis (nodule in the lower lobe of the right lung, nature to be determined) is text type, and the treatment plan (drug name, dosage, frequency of administration) are text type and numerical type respectively.
[0058] Construct a Prompt template: { "Identification instruction": "You are a professional medical information extraction expert. Please accurately extract relevant information from the medical document.", "Input data": { "Placeholder for medical document content": "Patient Li Si, male, 65 years old, has had cough and expectoration symptoms in the past two weeks, accompanied by low fever. After chest CT examination, a nodule about 2 cm × 3 cm was found in the lower lobe of the right lung, with clear boundaries. The initial diagnosis is a nodule in the lower lobe of the right lung, and the nature is to be determined. Further examination is recommended. The patient was treated with cough-relieving and expectoration-reducing drugs, such as ambroxol tablets, 30 mg each time, 3 times a day.", "Annotation requirements placeholder": "Please extract the patient's basic information (name, gender, age), symptoms, disease diagnosis, and treatment plan (drug name, dosage, frequency).", }, "Background information": "Medical documents include the patient's basic situation, symptoms, disease diagnosis, and treatment plan, etc. The patient's basic information includes name, gender, and age; symptoms are the manifestations of the patient's physical discomfort; disease diagnosis is the doctor's preliminary judgment result; the treatment plan involves drugs or other treatment means.", "Output specification": { "Patient name": "string", "Gender": "string", "Age": "number", "Symptom manifestations": "string", "Disease diagnosis": "string", "Drug name": "string", "Dosage": "string", "Frequency": "string" } } Extract preliminary medical annotation data Data preprocessing: Use natural language processing tools (such as NLTK) to clean the text, remove irrelevant information, convert to unified case, tokenize and tag parts of speech.
[0059] Build and apply a rule base: Build a rule base based on medical knowledge, such as rules like following the word "patient" with the name, using "male" or "female" to determine gender, and using numbers to determine age, to extract relevant information.
[0060] Combine Prompt template with dictionary matching: According to the requirements of the Prompt template, look up relevant words in the medical dictionary and determine the feature information in combination with the context, such as looking up symptom words like "cough" and "expectoration", and drug words like "ambroxol".
[0061] Generate preliminary labeled data: Organize the extracted features into preliminary labeled data according to the Prompt template format, such as {"Patient Name": "Li Si", "Gender": "Male", "Age": 65, "Symptom Manifestation": "Cough, expectoration, low fever", "Disease Diagnosis": "Nodule in the lower lobe of the right lung, nature to be determined", "Drug Name": "Ambroxol Tablets", "Dosage": "30mg", "Frequency of Medication": "3 times a day"}.
[0062] Result verification and optimization: Verify the labeled data according to medical logic and common sense, such as whether the age is reasonable, whether the dosage is in line with the regulations, etc. If there are problems, return to re-check the extraction and continuously optimize the rule library and dictionary.
[0063] Train the LLM model: Organize the medical documents and preliminary medical labeled data into an input format that conforms to the Prompt template, and divide them into a training set, a validation set, and a test set. Set hyperparameters, such as learning rate, batch size, number of training epochs, select a suitable LLM model (such as ChatGLM) for training. Use the cross-entropy loss function to calculate the loss value, and update the model parameters through the backpropagation algorithm and the Adam optimizer. During the training process, use the validation set to evaluate the model performance, adjust the hyperparameters according to the evaluation results and retrain. After training is completed, use the test set to evaluate the model performance and verify the stability and reliability of the model.
[0064] In some embodiments, the intelligent identification system for medical documents in a specialized disease library may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the intelligent identification system for medical documents in a specialized disease library can be stored in the memory of a computer device and executed by at least one processor to perform (see Figure 1 description) the functions of intelligent identification of medical documents in a specialized disease library.
[0065] In this embodiment, the intelligent identification system for medical documents in a specialized disease library can be divided into multiple functional modules according to the functions it performs, such as Figure 2 shown. The functional modules of the system may include: a feature determination module, a template construction module, a data annotation module, a model training module, and an output recognition module. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments. The system includes: A feature determination module, which determines key medical features based on the identification requirements of various medical documents and defines the types of key medical features; A template construction module constructs a Prompt template based on the recognition instructions of medical documents, the input medical document data, the background information related to medical documents, and the requirements of the recognized output specifications; A data annotation module extracts preliminary medical annotation data from unannotated medical documents based on natural language processing methods combined with the Prompt template; A model training module trains an LLM model based on medical documents, medical annotation data related to medical documents, and the Prompt template; An output recognition module inputs the medical document to be recognized and the corresponding Prompt template into the LLM model to obtain medical annotation data that meets the output specification requirements.
[0066] The feature determination module clarifies the key medical features and their types, providing a basis for subsequent processing; the template construction module generates a standardized Prompt template to guide the model's work; the data annotation module efficiently obtains preliminary annotation data with the help of natural language processing and the template; the model training module uses relevant data and the template to train the LLM model to improve its performance; the output recognition module uses the trained model to obtain accurate medical annotation data. Each module works together to efficiently and accurately extract key information from medical documents, reduce manual intervention, improve the efficiency and quality of medical document processing, provide strong support for medical research, diagnosis, and treatment, and promote the intelligent development of the medical field. Figure 3 FIG. is a schematic structural diagram of a terminal provided by an embodiment of the present invention, and the terminal can be used to execute the method for intelligent recognition of medical documents in a specialized disease database provided by an embodiment of the present invention.
[0067] Among them, the terminal may include: a processor, a memory, and a communication unit. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation on the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0068] Among them, the memory can be used to store the execution instructions of the processor. The memory can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. When the execution instructions in the memory are executed by the processor, the terminal can execute some or all of the steps in the above method embodiments.
[0069] The processor is the control center of the storage terminal, connecting various parts of the entire electronic terminal through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory, and by calling data stored in the memory, it performs various functions of the electronic terminal and / or processes data. The processor can be composed of an integrated circuit (IC for short), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor can include only a central processing unit (CPU for short). In the embodiments of the present invention, the CPU can be a single arithmetic core or can include multiple arithmetic cores.
[0070] A communication unit, which is used to establish a communication channel so that the storage terminal can communicate with other terminals. It receives user data sent by other terminals or sends user data to other terminals.
[0071] The present invention also provides a computer-readable storage medium. Among them, this computer storage medium can store a program, and when the program is executed, it can include some or all of the steps in the embodiments provided by the present invention. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM for short), a random access memory (RAM for short), etc.
[0072] Therefore, the technical effects that can be achieved by this embodiment can be referred to the description above, and will not be elaborated here.
[0073] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disk, etc., which can store program codes, and includes several instructions to enable a computer terminal (which can be a personal computer, a server, or a second terminal, a network terminal, etc.) to execute all or part of the steps of the methods described in the embodiments of the present invention.
[0074] In each embodiment described in this specification, the same or similar parts among the embodiments can be referred to each other. In particular, for the terminal embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, refer to the descriptions in the method embodiments.
[0075] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of systems or modules can be in electrical, mechanical or other forms.
[0076] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0077] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0078] Although the present invention has been described in detail by referring to the drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention. / Any person familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, and they should all be covered within the protection scope of the present invention.
Claims
1. An intelligent recognition method for medical documents in a disease-specific database, characterized in that, It includes the following steps: S1. Determine the key medical features based on the recognition requirements of various medical documents, and define the types of the key medical features; S2. Construct a Prompt template based on the recognition instructions of medical documents, the input medical document data, the relevant background information of medical documents, and the requirements of the recognition output specification; S3. Extract preliminary medical annotation data from unannotated medical documents based on the natural language processing method combined with the Prompt template; S4. Train the LLM model based on medical documents, the relevant medical annotation data of medical documents, and the Prompt template; S5. Input the medical document to be recognized and the corresponding Prompt template into the LLM model to obtain medical annotation data that meets the requirements of the output specification.
2. The intelligent recognition method for medical documents in a specialized disease database according to claim 1, wherein, The construction of the Prompt template in step S2 specifically includes: The recognition instructions of medical documents, which clarify the professional perspective that the model needs to be in during the recognition process of medical documents; The input medical document data, which uses placeholders to determine how to embed the medical document to be recognized and the annotation requirements into the Prompt template; The relevant background information of medical documents, which summarizes the important content of medical documents and gives the key medical features to be extracted; The output specification after recognition, which stipulates the structure of the output and the data type of each field of the output in JSON format.
3. The intelligent recognition method for medical documents in a specialized disease database according to claim 1, wherein Step S3 specifically includes: Preprocess the medical documents based on natural language processing tools; Construct a key medical feature extraction rule library based on medical knowledge and the standard expression methods in medical documents, and extract specific medical features from medical texts based on the key medical feature extraction rule library; Look up the vocabulary related to the recognition instructions in the medical dictionary based on the recognition instructions of medical documents in the Prompt template, and obtain the information matching the vocabulary in the medical document as specific medical features; Organize all the extracted specific medical features into preliminary annotation data based on the format specified by the Prompt template; Verify the preliminary annotation data based on medical logic and medical common sense, and optimize the rule library and medical dictionary based on the verification results.
4. The intelligent recognition method of medical documents in a specialized disease database according to claim 1, wherein Step S4 specifically includes: Organize medical documents and the corresponding preliminary medical annotation data into an input format that conforms to the Prompt template and use it as input data, and divide the input data into a training set, a validation set, and a test set; Set the hyperparameters required during training, train the LLM model based on the training set, input the medical document and the Prompt template into the model, and the model outputs medical annotation data; Calculate the loss value between the output medical annotation data and the real medical annotation data based on the cross-entropy loss function. When the loss value is greater than the preset threshold, calculate the model gradient using the backpropagation algorithm and update the parameters of the model based on the Adam optimizer; After each training round, evaluate the performance of the model using the validation set, adjust the hyperparameters based on the evaluation results of the validation set, and retrain; After the model training is completed, evaluate the performance of the model using the test set, and verify the stability and reliability of the model based on the fluctuations of the performance.
5. The intelligent recognition method of medical documents in a specialized disease database according to claim 1, wherein, In step S5, when the number of words in the medical document is greater than the threshold, the medical document is split into multiple window texts based on the sliding window method. The multiple texts and the corresponding Prompt templates are input into the LLM model to obtain the medical annotation data for each window text. The medical annotation data for each window text is associated and integrated to obtain the medical annotation data for the medical document.
6. The intelligent recognition method of medical documents in a specialized disease database according to claim 5, wherein Specifically, splitting the medical document into multiple window texts based on the sliding window method includes: Determining the window size and sliding step based on the input limit of the LLM model, the characteristics of the medical document, and the Prompt template; Starting from the starting position of the medical document, intercepting the window text based on the window size, and sliding the window backward in sequence according to the set sliding step, and intercepting a new window text each time until the window moves to the end of the medical document.
7. The intelligent recognition method of medical documents in a specialized disease database according to claim 5, characterized in that, Specifically, associating and integrating the medical annotation data for each window text includes: For the same key medical feature that is annotated and has the same content in different window texts, using this content as the medical annotation data for the current key medical feature; When there are different annotation values for the same key medical feature in different window texts, merging numerical features into a numerical range and merging text features into a complete description; Verifying whether there are inconsistent situations within the integrated medical annotation data, and adjusting the medical annotation data based on the verification results.
8. An intelligent recognition system for medical documents in a specialized disease database, characterized in that, When the system is implemented, it executes the intelligent recognition method for medical documents in a specialized disease database as described in any one of claims 1-7. The system includes: A feature determination module that determines key medical features based on the recognition requirements of various medical documents and defines the types of key medical features; A template construction module that constructs a Prompt template based on the recognition instructions of the medical document, the input medical document data, the relevant background information of the medical document, and the output specification requirements after recognition; A data annotation module that extracts preliminary medical annotation data from unannotated medical documents based on natural language processing methods in combination with the Prompt template; A model training module that trains the LLM model based on the medical document, the medical annotation data related to the medical document, and the Prompt template; An output recognition module that inputs the medical document to be recognized and the corresponding Prompt template into the LLM model to obtain medical annotation data that meets the output specification requirements.
9. A terminal, characterized in that, Including: A memory for storing the intelligent recognition program for medical documents in a specialized disease database; A processor that, when executing the intelligent recognition program for medical documents in a specialized disease database, implements the steps of the intelligent recognition method for medical documents in a specialized disease database as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The intelligent recognition program for medical documents in a specialized disease database is stored on the readable storage medium, and when the intelligent recognition program for medical documents in a specialized disease database is executed by the processor, it implements the steps of the intelligent recognition method for medical documents in a specialized disease database as described in any one of claims 1-7.