Medical text structured processing method based on Prompt technology
Through a comprehensive process based on Prompt technology, problems such as difficult non-structured text processing and strong dependence on labeled data in medical text processing are solved, and efficient structured medical text processing and privacy protection are achieved, improving the understanding ability and generation quality of the model.
Patent Information
- Application Number
- CN202510091839.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
Existing medical text processing technology is difficult to process unstructured text, it has strong dependence on labeled data, insufficient generalization ability, lacks dynamic optimization mechanisms and low system integration.
Using the structured medical text processing method based on Prompt technology, the efficient structured medical text processing is achieved through the comprehensive process of data desensitization and privacy protection, multiple candidate generation and evaluation, Embedding layer similarity calculation, dynamic example update and February-shot learning optimization.
It improves the accuracy and consistency of medical text processing, realizes privacy protection and data desensitization, enhances the model's understanding of medical text, and ensures the high quality of generated data and the consistency of standards in the medical field through automated testing and evaluation mechanisms.
Smart Images

Figure CN120015344A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical information processing, and in particular to a medical text structured processing method based on Prompt technology. Background Art
[0002] With the rapid development of medical information technology, medical texts (such as pathology reports, examination reports, diagnosis records, surgical records, discharge summaries, etc.) play a vital role in the medical industry. These texts record the patient's medical history, examination results, diagnosis conclusions and treatment plans, and are an important basis for clinicians to diagnose diseases, formulate treatment plans and conduct medical research. However, medical texts usually exist in an unstructured form, with complex content and diverse expressions, and lack a unified standard format, resulting in low efficiency in information extraction and analysis, which seriously restricts the value of medical data.
[0003] At present, the main technical routes for medical text processing include the following:
[0004] Rule-based text processing methods,Traditional medical text processing methods rely on predefined rules and templates, and extract information from text through regular expressions, keyword matching and other technologies.,This method is simple to implement, but has limited applicability, is difficult to deal with complex text structures and diverse expressions, and has high rule maintenance costs.
[0005] Natural language processing (NLP) methods based on machine learning. With the development of machine learning technology, NLP methods based on statistical learning have gradually been applied to medical text processing tasks, such as named entity recognition (NER), relationship extraction, and text classification. These methods rely on a large amount of labeled data for model training, which can improve the accuracy of information extraction to a certain extent. However, traditional machine learning methods have limited semantic understanding capabilities and are difficult to handle the complex semantic relationships implicit in medical texts.
[0006] NLP methods based on deep learning. In recent years, deep learning technology has made significant progress in the field of NLP, especially neural network-based models (such as LSTM, Transformer, etc.) have shown strong performance in medical text processing tasks. For example, pre-trained language models such as BERT significantly improve the semantic understanding ability of medical text through contextual semantic modeling. However, the training and optimization of these models rely on a large amount of annotated data, and their performance is limited when processing small sample scenarios.
[0007] Prompt technology based on large language models is an emerging NLP method that guides large language models to complete specific tasks by designing prompts. Large language models (such as the GPT series, T5, etc.) have powerful semantic understanding and generation capabilities, and can complete complex text processing tasks with few or even zero samples. The core of Prompt technology is to convert tasks into generation tasks that the model is good at by reasonably designing prompts, thereby giving full play to the capabilities of large language models. However, the application of Prompt technology in the medical field is still in the exploratory stage, especially in the structured processing of medical texts, there is still a lack of systematic research.
[0008] Although medical text processing technology has made significant progress in recent years, existing technologies still face many challenges and limitations in practical applications. These limitations are mainly reflected in the following aspects: difficulty in processing unstructured text, strong dependence on annotated data, insufficient generalization ability, lack of dynamic optimization mechanism, and low system integration. Summary of the invention
[0009] In order to overcome the shortcomings and deficiencies of the prior art, the present invention provides a medical text structured processing method based on Prompt technology.
[0010] A medical text structured processing method based on Prompt technology, the method comprising:
[0011] Step S1: Data desensitization and privacy protection, including removal of personal information, data encryption, hierarchical access control, compliance review, and text segmentation;
[0012] Step S2, multi-candidate generation and evaluation, including multi-candidate generation, self-evaluation and ranking, automated structured template generation, user feedback-driven optimization, automated testing and evaluation, and multi-round interactive optimization;
[0013] Step S3: Embedding layer similarity calculation, including semantic representation, example selection, text embedding extraction, multiple example selection, and example weight assignment;
[0014] Step S4, dynamic example update, including dynamic update in the medical field, task diversity expansion, user feedback driven, update based on generation results, update based on user feedback, automated example expansion, example classification and grouping, and example version control;
[0015] Step S5, comprehensive process of few-shot learning optimization, including text embedding extraction, similarity calculation and example selection, generation results and user feedback, dynamic example update, and loop iterative optimization;
[0016] Step S6: Self-optimization of the large model, including feedback-driven improvement, generation quality evaluation, multi-round evaluation and feedback optimization, multi-dimensional evaluation fusion, anomaly detection and confidence analysis.
[0017] Furthermore, the step S1 comprises:
[0018] Remove personal information: Delete sensitive information such as patient name, ID number, address, and contact information in the text;
[0019] Data encryption: Encrypt the stored and transmitted medical text data to ensure that the data is not leaked during transmission;
[0020] Hierarchical access control: Set data access levels based on user permissions to ensure that only authorized personnel can view sensitive information;
[0021] Compliance review: used to ensure that the data processing process complies with relevant laws, regulations and industry standards;
[0022] Text segmentation: Segment medical text into paragraphs or lines to extract key diagnostic information.
[0023] Furthermore, the step S2 comprises:
[0024] Multi-candidate generation: For the same medical text, Prompt is used to generate multiple candidate results to cover different possibilities and expressions;
[0025] Self-evaluation and ranking: The self-evaluation capability of the large language model is used to evaluate and rank the quality of the generated candidate results;
[0026] Automatic structured template generation: Prompt guides the large language model to extract information from medical texts and automatically generate structured templates for subsequent data storage and analysis;
[0027] User feedback-driven optimization: Collect user feedback on generated results, analyze the causes of errors, and adjust the prompt template accordingly;
[0028] Automated testing and evaluation: Build test sets to automatically evaluate the accuracy and consistency of prompt generated results;
[0029] Multi-round interaction optimization: The prompt design is gradually optimized through multiple rounds of interaction.
[0030] Furthermore, the step S3 comprises:
[0031] Semantic representation: The Embedding layer maps text to a high-dimensional vector space to capture the semantic information of the text. Similar texts are closer in the Embedding space.
[0032] Example selection: By calculating the embedding similarity between medical text and example data, select the few-shot examples most relevant to the current task from the example library to avoid using irrelevant or low-relevance examples;
[0033] Text Embedding Extraction: Use the Embedding layer of a large language model to extract vector representations of medical text and example data;
[0034] Multiple example selection: select multiple highly relevant examples based on similarity scores to construct contextual input;
[0035] Example weight assignment: Assign weights to selected examples based on their similarity scores. Examples with higher weights have a greater impact on model generation.
[0036] Furthermore, the step S4 comprises:
[0037] Dynamic updates in the medical field: The sample library is dynamically updated to keep pace with the latest medical knowledge;
[0038] Task diversity expansion: Based on different examples of different types of medical texts, the example library is dynamically expanded to cover more task scenarios;
[0039] User feedback-driven: User feedback on model generation results helps identify deficiencies in the example library and is used to guide the optimization and update of examples;
[0040] Updates based on generated results: Analyze the quality of model generation results, identify scenarios with generation errors or inaccuracies, and supplement or replace examples for the scenarios;
[0041] Updates based on user feedback: Collect user evaluations of generated results and adjust the sample library based on feedback;
[0042] Automatic example expansion: Generate new examples using large language models and select high-quality examples through manual review or automatic evaluation;
[0043] Example classification and grouping: Classify examples according to task types to ensure that each task has targeted example support;
[0044] Example version control: Perform version management on the example library, record the content and reason of each update, and use it for backtracking and optimization.
[0045] Furthermore, the step S5 comprises:
[0046] Text Embedding Extraction: Extract Embedding vectors of medical text and sample libraries;
[0047] Similarity calculation and example selection: Calculate similarity and select the most relevant few-shot examples as context input for Prompt;
[0048] Generate results and user feedback: Generate results using selected examples and collect user feedback;
[0049] Dynamic example update: Dynamically update the example library based on generation results and user feedback to optimize example quality and coverage;
[0050] Iterative optimization: Repeat the cycle of generating results, self-evaluating the results by the large model, correcting the results, and self-evaluating to continuously improve the effect of few-shot learning.
[0051] Furthermore, the step S6 comprises:
[0052] Feedback-driven improvement: Using an iterative feedback loop, a large language model generates initial prompt words, self-evaluates the prompt word effects, corrects the results, self-evaluates, criticizes and improves its own prompt templates and generated examples. This continuous improvement mechanism is used to ensure that each iteration is better than the previous one, thereby producing efficient prompts and examples.
[0053] Generation quality evaluation: Use the self-evaluation ability of the large language model to score the generation results and screen out high-quality results;
[0054] Multi-round evaluation and feedback optimization: Through multiple rounds of generation and evaluation, the quality of candidate results is gradually optimized;
[0055] Multi-dimensional evaluation fusion: Combine the multi-dimensional evaluation criteria of accuracy, completeness, logic and consistency to weightedly score the candidate results and select the best result;
[0056] Anomaly detection and confidence analysis: Detect anomalies in candidate results through rules or models, and eliminate low-quality results to ensure the reliability of output results.
[0057] Beneficial effects:
[0058] The present invention proposes a medical text structured processing method based on Prompt technology: the method utilizes the Prompt technology of dynamic generation and self-optimization to break the bottleneck of traditional manual design of Prompt, automatically generates relevant few-shot examples through the Embedding mechanism to improve the accuracy and consistency of text processing, and automatically optimizes the structured data generation process through self-evaluation and expert feedback through the automated structured template generation and feedback mechanism, thereby improving the model's ability to understand medical text. The method has privacy protection and data desensitization functions, which are used to ensure that medical data processing complies with relevant regulatory requirements, and protects patient privacy through automatic desensitization and encryption. At the same time, the automated testing and evaluation mechanism enables the system to continuously optimize and verify its generation effect, which is used to ensure the high quality of generated data and the consistency of medical field standards. The method can efficiently process medical texts, automatically generate structured data, and can be continuously optimized in actual use, ultimately providing strong support for data analysis, storage and application in the medical industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a flow chart of the overall steps of the present invention. DETAILED DESCRIPTION
[0060] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application may be combined with each other. The present application is further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0061] like Figure 1 As shown, a medical text structured processing method based on Prompt technology, the method includes:
[0062] Step S1: Data desensitization and privacy protection, including removal of personal information, data encryption, hierarchical access control, compliance review, and text segmentation;
[0063] In the process of data desensitization and privacy protection, sensitive information in medical texts is first removed, such as patient names, ID numbers, addresses, and contact information. This personal information must be deleted from the original data to ensure that patient privacy is protected. Next, in order to prevent data leakage, all stored and transmitted medical data need to be encrypted to ensure that the data is not tampered with or leaked during transmission. At the same time, hierarchical access control is adopted to ensure that only authorized personnel can access relevant sensitive information. This step also requires compliance review to ensure that data processing complies with relevant laws, regulations, and industry standards. Finally, through text segmentation technology, medical texts are segmented by paragraphs or lines to facilitate the extraction of important diagnostic information and the processing of subsequent analysis tasks.
[0064] Medical texts usually contain patients' personal information (such as name, ID number, address, etc.). Patient privacy must be strictly protected during data processing to ensure compliance with relevant laws and regulations (such as the "Personal Information Protection Law of the People's Republic of China"). Removing personal information: Delete sensitive information such as the patient's name, ID number, address, contact information, etc. in the text. For example: Original text: Patient Zhang San, male, 45 years old, ID number: 123456789012345678. After desensitization: Patient [name has been desensitized], male, 45 years old, ID number: [has been desensitized]. Data anonymization: Use anonymous identifiers to replace patient information. For example, replace "Zhang San" with "Patient A001".
[0065] Data encryption: Encrypt the stored and transmitted medical text data to ensure that the data is not leaked during transmission. Hierarchical access control: Set data access levels based on user permissions to ensure that only authorized personnel can view sensitive information.
[0066] Compliance review: used to ensure that the data processing process complies with relevant laws, regulations and industry standards. Text segmentation: medical text is segmented into paragraphs or lines to extract key diagnostic information.
[0067] Step S2, multi-candidate generation and evaluation, including multi-candidate generation, self-evaluation and ranking, automated structured template generation, user feedback-driven optimization, automated testing and evaluation, and multi-round interactive optimization;
[0068] In the step of multi-candidate generation and evaluation, firstly, multiple candidate results are generated for the same medical text through prompt design. These candidate results can show different expressions and diverse solutions, thus covering different possibilities. Then, the self-evaluation ability of the large language model is used to evaluate the quality of these candidate results and sort them according to the evaluation results. At the same time, the system automatically generates structured templates, extracts key information from the medical text and organizes it into a structured format to facilitate subsequent data storage and analysis. In this process, the collection of user feedback is also very important. By analyzing the user's feedback on the generated results, errors can be identified and the prompt template can be optimized. In order to ensure the quality and consistency of the results, the system also builds an automated testing and evaluation mechanism, and further adjusts and improves the prompt design through multiple rounds of interactive optimization.
[0069] There may be multiple possibilities for the results generated by Prompt. In order to ensure the accuracy and reliability of the output results, the final results can be optimized through a multi-candidate generation and evaluation mechanism.
[0070] Multi-candidate generation: For the same medical text, Prompt is used to generate multiple candidate results to ensure that different possibilities and expressions are covered.
[0071] For example:
[0072] Candidate 1:
[0073] {
[0074] "Lesion type":"Nodule",
[0075] "Lesion location":"Right lower lobe",
[0076] "Diagnosis conclusion": "Highly likely to be malignant"
[0077] }
[0078] Candidate 2:
[0079] {
[0080] "Lesion type":"mass",
[0081] "Lesion location":"Right lower lobe",
[0082] "Diagnosis conclusion":"malignant"
[0083] }
[0084] Self-evaluation and ranking: Use the self-evaluation ability of the large language model to evaluate and rank the quality of the generated candidate results. The evaluation criteria may include:
[0085] Semantic accuracy: whether the results accurately reflect the key information in the medical text.
[0086] Logical consistency: whether the results are consistent with the logic and common sense in the medical field.
[0087] Format compliance: whether the result conforms to the expected structured format (such as JSON).
[0088] Example:
[0089] Please evaluate the following candidate results and select the best one:
[0090] Candidate result 1: {Candidate result 1}
[0091] Candidate result 2: {Candidate result 2}
[0092] Candidate result 3: {Candidate result 3}
[0093] Model output:
[0094] Best result: Candidate 1
[0095] Reason: The lesion type, location and diagnosis conclusion of candidate result 1 are completely consistent with the medical text and the format is standard.
[0096] Automated structured template generation, through prompts to guide large language models, can not only extract key information from medical texts, but also automatically generate structured templates (such as JSON, XML, etc.) for subsequent data storage and analysis.
[0097] Example of a structured template:
[0098] Generate standardized JSON format for medical text:
[0099] Enter medical text:
[0100] A nodular shadow was found in the patient's right lower lobe, which was considered to be highly likely to be malignant.
[0101] Output JSON template:
[0102] {
[0103] "Lesion type":"Nodule",
[0104] "Lesion location":"Right lower lobe",
[0105] "Diagnosis conclusion": "Highly likely to be malignant"
[0106] }
[0107] Dynamically adjust the template structure according to different task requirements. For example:
[0108] Pathology report: extract information such as lesion type, lesion location, and pathological grade.
[0109] Examination records: Extract information such as examination type, examination results, doctor's advice, etc.
[0110] Discharge summary: extract information such as diagnosis conclusion, treatment plan, follow-up recommendations, etc.
[0111] Optimize the iterative process of prompt design. Prompt design is a process that needs to be continuously optimized. The following methods can further improve the effect of prompt:
[0112] User feedback-driven optimization: Collect user feedback on generated results, analyze the causes of errors, and adjust the prompt template accordingly. For example, if users report that some diagnostic conclusions are not correctly extracted, you can add a clear prompt for the diagnostic conclusion in the prompt.
[0113] Sentence-BERT (SBERT) is a variant based on BERT, which is specially designed to generate sentence embeddings, mainly used for tasks such as calculating semantic similarity, text clustering, and information retrieval. The following is the model structure and workflow of SBERT:
[0114] Input layer: Input a single sentence or sentence pair. For sentence pairs, the format is usually [CLS]sentence1[SEP]sentence2[SEP].
[0115] Encoding layer: Use pre-trained BERT or its variants (such as RoBERTa, DistilBERT) to independently encode each sentence and generate contextual representations (token embeddings) of each word.
[0116] Pooling layer: The word vector is converted into a fixed-length representation of the sentence through pooling. The commonly used pooling method is mean-pooling, which takes the average of all word embeddings as the global representation of the sentence. You can also use maximum pooling or directly use the embedding of the [CLS] token.
[0117] Model training: (1) Data preparation: For the original text of each case (e.g., doctor's diagnosis record, patient's medical history, etc.), we input it into the SBERT model. SBERT will generate a fixed-length sentence embedding for each medical sentence. This embedding represents the semantic information contained in the sentence and can be used for subsequent similarity calculations. Embedding generation for structured data: For structured data (such as patient's diagnosis, laboratory test results, drug prescriptions, etc.), we can use numerical embedding (e.g., converting numerical values into standardized vectors) or category embedding (e.g., converting category data into vector representations through an embedding layer).
[0118] (2) Similarity calculation: Once you have the embedding of the sentence (text embedding) and the embedding of the structured data, you can calculate the similarity between them. Similarity can be calculated using metrics such as cosine similarity, Euclidean distance, or Manhattan distance.
[0119]
[0120] (3) Loss design based on target f1:
[0121] Weighted Contrastive Loss: In order to make the model pay more attention to key sentence pairs in the medical field (such as diagnosis results, drug treatment, etc.), a weighting factor can be introduced into the contrast loss function to make the loss of certain samples more significant. The specific steps are as follows:
[0122]
[0123] The loss formula based on f1 weighted optimization, where: y i represents the positive and negative sample labels (1 means similar, 0 means dissimilar). i') is the distance between the two sentence embeddings. m is the margin value, ensuring that the distance of negative samples is greater than that of positive samples i, w i is a weighting factor, which indicates the importance of different sample pairs.
[0124] (4) F1 value as loss weighting factor:
[0125] In order to further improve the accuracy of medical text information extraction and structured data matching, the F1 value is used as a weighting factor to optimize the loss function of the model.
[0126]
[0127] w i The weight is calculated based on the F1 value:
[0128] During the training process of the model, medical text and structured data are first taken as input to generate corresponding text embeddings and structured data embeddings, respectively. According to the task requirements, the F1 value of each sample is calculated and used as a weighting factor for the calculation of weighted loss. Then, the similarity between the text embedding and the structured data embedding is calculated. Commonly used metrics include cosine similarity or Euclidean distance. The loss function uses weighted contrast loss and triplet loss, in which the importance of each sample is weighted by the F1 value to ensure that the model pays more attention to high-quality prediction results during training. At each iteration, the model dynamically adjusts the sample weights according to the currently calculated F1 value, thereby optimizing the loss function and finally updating the model parameters through the back-propagation algorithm. By continuously optimizing the weighted contrast loss and triplet loss, the model can better understand the relationship between medical text and structured data and improve its performance in medical tasks.
[0129] Whenever new medical text is input, the system selects multiple historical examples that are most relevant to the input from the database based on similarity calculations, and inputs these examples into Prompt as context. If there are not enough relevant examples in the example library, the system will automatically generate relevant examples through a generative model (such as GPT-4) and add them to the example library.
[0130] Automated testing and evaluation:
[0131] Build a test set to automatically evaluate the accuracy and consistency of Prompt's generated results. For example, compare the generated results with standard answers in the medical field and calculate the precision, recall, and F1 score.
[0132] Multi-round interaction optimization: The prompt design is gradually optimized through multiple rounds of interaction. For example, the results generated by the initial prompt may be incomplete, and additional prompts (such as "Please supplement the diagnosis conclusion") can be used to guide the model to generate more complete results.
[0133] Few-shot learning is an efficient method to use a small amount of example data (Few-shot examples) to guide large language models (LLMs) to complete specific tasks. In the medical text processing scenario, Few-shot learning can significantly improve the model's understanding and generation capabilities for complex tasks. However, how to select the most relevant Few-shot examples and dynamically optimize the example library is the key to the effectiveness of Few-shot learning. The following is a detailed method and extended content for optimizing Few-shot learning through Embedding layer similarity calculation and dynamic example update.
[0134] Step S3: Embedding layer similarity calculation, including semantic representation, example selection, text embedding extraction, multiple example selection, and example weight assignment;
[0135] In the step of similarity calculation of the Embedding layer, the medical text is first mapped to a high-dimensional vector space through the Embedding layer to capture the semantic information of the text. In this space, semantically similar texts are mapped to positions that are close to each other. Then, by calculating the Embedding similarity between the current medical text and the example data, the few-shot examples that are most relevant to the current task are selected. This step can effectively avoid using irrelevant or low-relevance examples. Next, the Embedding vectors of the medical text and the example data are extracted, and the similarity scores between them are calculated. Based on the scores, multiple highly relevant examples are selected to construct the input context, and weights are assigned to the selected examples based on the similarity. Examples with higher weights have a greater impact on the generated results, thereby optimizing the generation process.
[0136] The Embedding layer is a core component used to represent text semantics in large language models. By calculating the similarity between medical text and sample data in the Embedding space, the most relevant few-shot examples can be automatically selected as the context input of the Prompt, thereby improving the generation quality of the model.
[0137] The role of the Embedding layer: semantic representation: The Embedding layer maps text to a high-dimensional vector space to capture the semantic information of the text. Similar texts are closer in the Embedding space. Example selection: By calculating the Embedding similarity between medical text and example data, the few-shot examples most relevant to the current task can be selected from the example library to avoid using irrelevant or low-relevance examples.
[0138] Similarity calculation method, text embedding extraction: Use the Embedding layer of a large language model (such as GPT, BERT, etc.) to extract vector representations of medical text and example data.
[0139] Example:
[0140] Medical text: A nodule was found in the patient's right lower lobe of the lung, which is likely to be malignant.
[0141] Example 1: A mass was found in the left upper lobe of the patient's lung and was diagnosed as benign.
[0142] Example 2: A lesion was found in the right lower lobe of the patient's lung, which was considered to be an infectious lesion.
[0143] Extract the embedding vector:
[0144] Medical text → vector A;
[0145] Example 1 → Vector B;
[0146] Example 2 → Vector C;
[0147] Similarity calculation formula:
[0148] Use cosine similarity to calculate the similarity between texts:
[0149]
[0150] Among them, A and B are the embedding vectors of two texts, ||A|| and ||B|| are the moduli of the vectors.
[0151] Example calculation:
[0152] Similarity between medical text and example 1 = 0.65;
[0153] The similarity between medical text and example 2 = 0.92.
[0154] Example selection:
[0155] According to the similarity score, the few-shot examples with the highest similarity are selected from the example library as the context input of Prompt.
[0156] Example selection result: Example 2 (similarity 0.92) was selected as the few-shot example.
[0157] Optimization strategy,Multi-example selection: Instead of just selecting one example, multiple highly relevant examples are selected based on similarity scores to build richer contextual input.
[0158] Example:
[0159] Few-shot example 1: A lesion was found in the right lower lobe of the patient's lung, which was considered to be an infectious lesion.
[0160] Few-shot example 2: A nodule was found in the patient's right middle lobe and was diagnosed as malignant.
[0161] Example weight assignment:
[0162] The selected examples are assigned weights based on their similarity scores, and examples with higher weights have a greater impact on model generation.
[0163] Example:
[0164] Example 1 weight: 0.6;
[0165] Example 2 Weight: 0.4.
[0166] Step S4, dynamic example update, including dynamic update in the medical field, task diversity expansion, user feedback driven, update based on generation results, update based on user feedback, automated example expansion, example classification and grouping, and example version control;
[0167] In the step of dynamic example update, the example library will be updated in real time according to the latest medical knowledge to ensure that the model always uses the most cutting-edge medical data. The system dynamically expands the example library to cover more scenarios based on the diversity of tasks and the needs of different types of medical texts. User feedback is an important driving force for the optimization of the example library. By collecting and analyzing user feedback on the generated results, the deficiencies in the example library can be discovered and optimized in a timely manner. On this basis, the system will analyze the generated results, identify erroneous or inaccurate scenarios, and supplement or replace examples based on these scenarios. In addition, the system automatically generates new examples through a large language model, and selects high-quality examples through manual review or automatic evaluation and expands them. In order to manage the example library, the system will also classify and group examples according to the task type, and version control the example library, recording the content and reasons for each update for subsequent backtracking and optimization.
[0168] The quality and coverage of the example library directly affect the effect of few-shot learning. By dynamically updating the example library, the example library can be continuously optimized based on the model generation results and user feedback, thereby improving the long-term effect of few-shot learning.
[0169] The necessity of dynamic updates, the dynamic nature of the medical field: Medical knowledge and practice are constantly updated, such as the discovery of new diseases, the application of new treatments, etc. The sample library needs to be updated in a timely manner to keep pace with the latest medical knowledge.
[0170] Task diversity: Different types of medical texts (such as pathology reports, imaging test results, and discharge summaries) may require different examples, and the example library needs to be dynamically expanded to cover more task scenarios.
[0171] User feedback driven: User feedback on model generation results can help identify deficiencies in the example library, thereby guiding the optimization and updating of examples.
[0172] Dynamic update method, based on the update of generation results: analyze the quality of the model generation results, identify scenarios with generation errors or inaccuracies, and supplement or replace examples for these scenarios.
[0173] Example:
[0174] Generates the following result:
[0175] {
[0176] "Lesion type":"Nodule",
[0177] "Lesion location":"Right middle lobe",
[0178] "Diagnosis conclusion":"benign"
[0179] }
[0180] User feedback: The lesion location should be "right lower lobe".
[0181] Update examples: Supplement or replace examples related to "right lower lobe of the lung".
[0182] Updates based on user feedback: Collect user evaluations of generated results (such as accuracy, completeness, and logic) and adjust the example library based on feedback.
[0183] Example:
[0184] User feedback: The model failed to correctly extract treatment options.
[0185] Updated examples: Added examples that include treatment options, such as:
[0186] Example: The patient underwent surgery to remove a nodule in the right lower lobe of the lung and recovered well after the surgery.
[0187] Automatic example expansion: Generate new examples using a large language model and select high-quality examples through manual review or automatic evaluation.
[0188] Example:
[0189] Input: Please generate a medical example containing lesion type, lesion location, and diagnosis conclusion.
[0190] Output: A mass was found in the patient's left upper lung lobe, which is considered to be highly likely to be malignant.
[0191] Example library management, example classification and grouping: Classify examples according to task types (such as diagnosis extraction, treatment plan extraction) to ensure that each task has targeted example support.
[0192] Example grouping:
[0193] Diagnostic Extraction Sample Group;
[0194] Treatment plan extraction sample group;
[0195] Medical history record extraction sample group;
[0196] Example version control:
[0197] Perform version management on the sample library and record the content and reason of each update to facilitate backtracking and optimization.
[0198] Example:
[0199] Version 1.0: Initial example library, containing 50 examples.
[0200] Version 1.1: Added 10 new examples related to lung lesions.
[0201] Version 1.2: Replaced 5 low-quality examples and optimized the diagnostic extraction example set.
[0202] Step S5, comprehensive process of few-shot learning optimization, including text embedding extraction, similarity calculation and example selection, generation results and user feedback, dynamic example update, and loop iterative optimization;
[0203] In the comprehensive process of few-shot learning optimization, the embedding vectors of medical text and example library are first extracted, and the similarity between them is calculated. According to the similarity score, the most relevant few-shot examples are selected as the context input of Prompt to generate more accurate results. The generated results will be combined with user feedback to further optimize the example library and improve the quality and coverage of the examples. This optimization process is iterative. After each result is generated, the system will self-evaluate and correct, thereby gradually improving the effect of the model. The specific cycle includes generating results-the large model self-evaluates the results-corrects the generated results-self-evaluates again until the optimization effect is finally achieved. Through this continuous optimization process, the effect of few-shot learning is continuously improved.
[0204] Text Embedding Extraction: Extract Embedding vectors of medical text and sample libraries. Similarity Calculation and Sample Selection: Calculate similarity and select the most relevant few-shot samples as contextual input for Prompt.
[0205] Generate results and user feedback: Generate results using selected examples and collect user feedback.
[0206] Dynamic example update: Dynamically update the example library based on generation results and user feedback to optimize example quality and coverage.
[0207] Iterative optimization: Repeat the above steps to continuously improve the effect of few-shot learning.
[0208] Step S6: Self-optimization of the large model, including feedback-driven improvement, generation quality evaluation, multi-round evaluation and feedback optimization, multi-dimensional evaluation fusion, anomaly detection and confidence analysis.
[0209] In the self-optimization process of the large model, first through the feedback mechanism of cyclic iteration, the model will become more efficient each time it generates, criticizes and improves its own prompt templates and examples. Specifically, the model will first generate the initial prompt words, then self-evaluate the effect of its generation, modify the prompt content according to the evaluation results, and finally self-evaluate and modify it. Through this continuous improvement mechanism, the model can ensure that each iteration is more accurate and efficient than the previous one. The evaluation of generation quality is equally important. The model will score the generated results and filter out higher quality results. At the same time, the model will perform multiple rounds of generation and evaluation to gradually optimize the quality of candidate results. Throughout the process, multi-dimensional evaluation criteria such as accuracy, completeness, logic and consistency will be combined and comprehensively scored to select the best results. In addition, an anomaly detection mechanism will be introduced to detect anomalies in candidate results and eliminate low-quality results to ensure the reliability of the output results.
[0210] Feedback-driven improvement: At the heart of the system design is the use of an iterative feedback loop in which LLM generates, critiques, and improves its own prompt templates and generated examples. This continuous improvement mechanism is used to ensure that each iteration is better than the last, resulting in highly effective prompts and examples.
[0211] Generation quality evaluation: Utilize the self-evaluation ability of the large language model to score the generation results and screen out high-quality results.
[0212] Multi-round evaluation and feedback optimization: Through multiple rounds of generation and evaluation, the quality of candidate results is gradually optimized.
[0213] Multi-dimensional evaluation fusion: Combine multi-dimensional evaluation criteria such as accuracy, completeness, logic and consistency to perform weighted scoring on candidate results and select the best result.
[0214] Anomaly detection and confidence analysis: Detect anomalies in candidate results through rules or models, and eliminate low-quality results to ensure the reliability of output results.
[0215] Although embodiments of the present invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A medical text structured processing method based on Prompt technology, characterized in that , the method comprises: Step S1: Data desensitization and privacy protection, including removal of personal information, data encryption, hierarchical access control, compliance review, and text segmentation; Step S2, multi-candidate generation and evaluation, including multi-candidate generation, self-evaluation and ranking, automated structured template generation, user feedback-driven optimization, automated testing and evaluation, and multi-round interactive optimization; Step S3: Embedding layer similarity calculation, including semantic representation, example selection, text embedding extraction, multiple example selection, and example weight assignment; Step S4, dynamic example update, including dynamic update in the medical field, task diversity expansion, user feedback driven, update based on generation results, update based on user feedback, automated example expansion, example classification and grouping, and example version control; Step S5, comprehensive process of few-shot learning optimization, including text embedding extraction, similarity calculation and example selection, generation results and user feedback, dynamic example update, and loop iterative optimization; Step S6: Self-optimization of the large model, including feedback-driven improvement, generation quality evaluation, multi-round evaluation and feedback optimization, multi-dimensional evaluation fusion, anomaly detection and confidence analysis.
2. A medical text structured processing method based on Prompt technology as claimed in claim 1, characterized in that: The step S1 comprises: Remove personal information: Delete sensitive information such as patient name, ID number, address, and contact information in the text; Data encryption: Encrypt the stored and transmitted medical text data to ensure that the data is not leaked during transmission; Hierarchical access control: Set data access levels based on user permissions to ensure that only authorized personnel can view sensitive information; Compliance review: used to ensure that the data processing process complies with relevant laws, regulations and industry standards; Text segmentation: Segment medical text into paragraphs or lines to extract key diagnostic information.
3. The method for structuring medical text based on Prompt technology as claimed in claim 1, characterized in that: The step S2 comprises: Multi-candidate generation: For the same medical text, Prompt is used to generate multiple candidate results to cover different possibilities and expressions; Self-evaluation and ranking: The self-evaluation capability of the large language model is used to evaluate and rank the quality of the generated candidate results; Automatic structured template generation: Prompt guides the large language model to extract information from medical texts and automatically generate structured templates for subsequent data storage and analysis; User feedback-driven optimization: Collect user feedback on generated results, analyze the causes of errors, and adjust the prompt template accordingly; Automated testing and evaluation: Build test sets to automatically evaluate the accuracy and consistency of prompt generated results; Multi-round interaction optimization: The prompt design is gradually optimized through multiple rounds of interaction.
4. The method for structuring medical text based on Prompt technology as claimed in claim 1, characterized in that: The step S3 comprises: Semantic representation: The Embedding layer maps text to a high-dimensional vector space to capture the semantic information of the text. Similar texts are closer in the Embedding space. Example selection: By calculating the embedding similarity between medical text and example data, select the few-shot examples most relevant to the current task from the example library to avoid using irrelevant or low-relevance examples; Text Embedding Extraction: Use the Embedding layer of a large language model to extract vector representations of medical text and example data; Multiple example selection: select multiple highly relevant examples based on similarity scores to construct contextual input; Example weight assignment: Assign weights to selected examples based on their similarity scores. Examples with higher weights have a greater impact on model generation.
5. The method for structuring medical text based on Prompt technology as claimed in claim 1, characterized in that: The step S4 comprises: Dynamic updates in the medical field: The sample library is dynamically updated to keep pace with the latest medical knowledge; Task diversity expansion: Based on different examples of different types of medical texts, the example library is dynamically expanded to cover more task scenarios; User feedback-driven: User feedback on model generation results helps identify deficiencies in the example library and is used to guide the optimization and update of examples; Updates based on generated results: Analyze the quality of model generation results, identify scenarios with generation errors or inaccuracies, and supplement or replace examples for the scenarios; Updates based on user feedback: Collect user evaluations of generated results and adjust the sample library based on feedback; Automatic example expansion: Generate new examples using large language models and select high-quality examples through manual review or automatic evaluation; Example classification and grouping: Classify examples according to task types to ensure that each task has targeted example support; Example version control: Perform version management on the example library, record the content and reason of each update, and use it for backtracking and optimization.
6. The method for structuring medical text based on Prompt technology as claimed in claim 1, characterized in that: The step S5 comprises: Text Embedding Extraction: Extract Embedding vectors of medical text and sample libraries; Similarity calculation and example selection: Calculate similarity and select the most relevant few-shot examples as context input for Prompt; Generate results and user feedback: Generate results using selected examples and collect user feedback; Dynamic example update: Dynamically update the example library based on generation results and user feedback to optimize example quality and coverage; Iterative optimization: Repeat the cycle of generating results, self-evaluating the results by the large model, correcting the results, and self-evaluating to continuously improve the effect of few-shot learning.
7. The method for structuring medical text based on Prompt technology as claimed in claim 1, characterized in that: The step S6 comprises: Feedback-driven improvements: Using an iterative feedback loop, a large language model generates initial prompt words, self-evaluates the prompt word effects, corrects the results, self-evaluates, criticizes and improves its own prompt templates and generated examples; Generation quality evaluation: Use the self-evaluation ability of the large language model to score the generation results and screen out high-quality results; Multi-round evaluation and feedback optimization: Through multiple rounds of generation and evaluation, the quality of candidate results is gradually optimized; Multi-dimensional evaluation fusion: Combine the multi-dimensional evaluation criteria of accuracy, completeness, logic and consistency to weightedly score the candidate results and select the best result; Anomaly detection and confidence analysis: Detect anomalies in candidate results through rules or models, and eliminate low-quality results to ensure the reliability of output results.
Citation Information
Patent Citations
Voice question answering method and system based on short text matching
CN114328881A
Class case retrieval system and method based on retrieval enhancement generation technology
CN118260391A
Medical literature intelligent question answering system and method based on RAG and LLM technologies
CN118364088A
Large-model-based medical industry-oriented medical record fine adjustment data set generation and evaluation method and system for inquiry dialogue writing
CN119049626A
Method for generating diversified instruction data in medical field based on large language model
CN119069138A