Large medical model training method, medium, computer equipment and program product
By screening out target records containing maximized keywords from medical records, training medical big models is solved, and the problem of high difficulty in obtaining sample data and long training time is achieved, efficient model training and performance approximation are achieved.
Patent Information
- Application Number
- CN202510491438.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-29
AI Technical Summary
During the training of existing medical big model, the acquisition of sample data is difficult and the training takes time, resulting in inefficient model training.
The target medical records are selected from the medical records through multiple iterations, provided that the medical keywords included in the currently screened records and the total number of keywords included in the historical screened records is maximized, and the medical big model is trained using the filtered target medical records.
The number of sample data is reduced, the training efficiency is improved, the model performance is approaching the training effect of full data, and the data acquisition difficulty and training time are reduced.
Smart Images

Figure CN120387518A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of large language models, and particularly to a training method, medium, computer device, and program product for a medical large model. Background Art
[0002] With the rapid development of technology, the application of artificial intelligence (AI) in the medical field is gradually changing the traditional medical service model, bringing profound changes to the medical industry. Currently, a solution has been proposed to provide medical Q&A services for users through a medical large model. To improve the accuracy and reliability of the medical large model, a large amount of sample data is usually used for model training. On the one hand, this increases the difficulty of obtaining sample data, and on the other hand, it also increases the time-consuming of the model training process. Summary of the Invention
[0003] In a first aspect, an embodiment of this application provides a training method for a medical large model, the method including:
[0004] Obtain multiple medical records and multiple medical keywords;
[0005] Iteratively screen out multiple target medical records from the multiple medical records; in each iteration, based on the following condition, screen out target medical records from the medical records that have not been selected among the multiple medical records: the total number of medical keywords included in the currently screened target medical records and the medical keywords included in the historically screened target medical records is maximized;
[0006] Train a first large language model based on the multiple target medical records to obtain a medical large model, and the medical large model is used to provide Q&A services related to the medical field for users.
[0007] In a second aspect, an embodiment of this application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any embodiment of this application is implemented.
[0008] In a third aspect, an embodiment of this application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the method described in any embodiment of this application is implemented.
[0009] In a fourth aspect, an embodiment of this application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any embodiment of this application is implemented.
[0010] In the embodiments of the present application, multiple rounds of iteration are used to screen out multiple target medical records from the obtained multiple medical records, and then the screened target medical records are used to train a large medical model. Among them, in each iteration, target medical records are screened from the medical records that have not been selected among the multiple medical records based on the following condition: the total number of medical keywords included in the currently screened target medical records and the medical keywords included in the historically screened target medical records is maximized. On the one hand, the target medical records screened in the above manner can contain as many medical keywords as possible, that is, the screened target medical records are representative data in the full set of medical records. In this way, the performance of the large medical model trained based on these target medical records can approach the performance of the large medical model trained based on the full set of medical records, so that it is not necessary to use a large amount of sample data to train the large medical model; on the other hand, since the obtained target medical records are only a subset of the full set of medical records, the number of sample data is effectively reduced. Therefore, the training time can be reduced and the training efficiency can be improved.
[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings herein are incorporated into the specification and constitute a part of this application. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to explain the technical solutions of the present application.
[0013] Figure 1 is a flowchart of the method for training a large medical model according to an embodiment of the present application.
[0014] Figure 2 is a schematic diagram of the process of screening target medical records according to an embodiment of the present application.
[0015] Figure 3 is a schematic diagram of the pruning process according to an embodiment of the present application.
[0016] Figure 4 is a general flowchart of an embodiment of the present application.
[0017] Figure 5 is a flowchart of the training device for a large medical model according to an embodiment of the present application.
[0018] Figure 6 is a schematic diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0020] The terms used in the present application are for the purpose of describing particular embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. Additionally, the term "at least one" as used herein represents any one of a plurality or any combination of at least two of a plurality.
[0021] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0022] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the embodiments of the present application and to make the above-mentioned objects, features, and advantages of the embodiments of the present application more apparent and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0023] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select authorization or rejection.
[0024] The medical large model is a natural language interaction system based on medical data processing technology. In some embodiments, the medical large model can provide medical knowledge retrieval and analysis services for users by integrating medical literature databases, clinical guidelines, and public health data. The system uses deep learning algorithms to semantically parse the input user queries and generate response information. Specifically, the functions that the medical large model can perform include but are not limited to:
[0025] Medical knowledge association analysis: Match the symptom characteristics described by the user with the disease characteristics in the medical knowledge base, and output the disease characteristic descriptions recorded in relevant literature and possible relevance analysis;
[0026] Literature retrieval of treatment plans: Extract academic content such as treatment principle explanations and drug action mechanisms from public literature according to the medical concepts input by the user;
[0027] Push of health management knowledge: Based on the health parameters provided by the user and combined with the general health management principles in medical guidelines, generate popular science content of health knowledge.
[0028] The response information includes but is not limited to medical literature abstracts, disease characteristic descriptions, treatment principle explanations, and health knowledge items. These response information can be further provided to professional medical personnel so that professional medical personnel can make comprehensive judgments in combination with clinical practice.
[0029] In order to improve the accuracy and reliability of the medical large model, a large amount of sample data is usually used for model training, which brings the following problems:
[0030] (1) The difficulty of obtaining sample data is high. On the one hand, in order to train a medical large model with better performance, a large amount of data often needs to be collected, and the data collection difficulty is high; on the other hand, there are often a large amount of noises in the collected data, and a large amount of data needs to be cleaned, resulting in a high data cleaning difficulty.
[0031] (2) Since the quantity of the obtained sample data is huge, the training process needs to process a large amount of data, resulting in a long time-consuming for the model training process.
[0032] Based on this, the embodiments of the present application propose a training method for a medical large model. Through multiple rounds of iteration, multiple target medical records are screened out from the obtained multiple medical records, and then the screened target medical records are used to train the medical large model. In each iteration, the target medical records that can maximize the total number of medical keywords included in each target medical record are screened out. In this way, it is possible to screen out as much representative data as possible from the full amount of medical records, so that the performance of the medical large model trained using these target medical records can approach the performance of the medical large model trained based on the full amount of medical records, without the need to use a large amount of sample data to train the medical large model, thereby reducing the difficulty of obtaining sample data and improving the model training efficiency. The following is an example of the specific implementation details of the embodiments of the present application.
[0033] See Figure 1 , the training method of the medical large model of the present application includes:
[0034] Step S12: Obtain multiple medical records and multiple medical keywords;
[0035] Step S14: Iteratively screen out multiple target medical records from the multiple medical records; in each iteration, based on the following condition, screen out target medical records from the medical records that have not been selected from the multiple medical records: the total number of medical keywords included in the currently screened target medical records and the medical keywords included in the historically screened target medical records is maximized;
[0036] Step S16: Train a first large language model based on the multiple target medical records to obtain a medical large model, and the medical large model is used to provide question-and-answer services related to the medical field for users.
[0037] In step S12, multiple medical records can be obtained. Among them, medical records are records such as texts, charts, and imaging materials formed during medical treatment activities (such as cases, medical images, etc.), which reflect the patient's condition, treatment process, etc. The medical records may include, but are not limited to, the following information: the patient's basic information (such as name, gender, age, contact information, etc.), medical history, symptoms, diagnosis results, treatment process, examination results, test results, medication records, and / or effect evaluations. In some embodiments, the medical records may be question-and-answer data. For example, when a user visits a hospital, the doctor and the patient will communicate about the patient's symptoms, diagnosis results, treatment plan, etc., and the medical records may include question-and-answer data generated based on the communication records between the doctor and the patient. Another example is that when a user interacts with a medical large model, multiple rounds of question-and-answer dialogues will be carried out with the medical large model, and the medical records may include the question-and-answer data in the above multiple rounds of question-and-answer dialogues.
[0038] In addition, multiple medical keywords can be obtained. Medical keywords refer to words used to describe information in medical activities, including but not limited to disease names, symptom descriptions, treatment methods, drug names, examination items, etc.
[0039] In some embodiments, multiple medical keywords can be extracted from medical records. The medical records used here to obtain medical keywords can include some or all of the multiple medical records obtained in the foregoing embodiments, or can also include other medical records. For any obtained medical record, the medical record can be segmented to obtain multiple segments, and then multiple keywords related to the medical field can be screened out from the multiple segments. Among them, the segmentation process can be implemented using a segmentation tool.
[0040] Further, in the case where the medical record is non-text data (such as image data or audio data, etc.), before segmenting the medical record, the medical record can also be converted into text data. For example, voice information can be extracted from the image data or audio data, and the extracted voice information can be converted into text information through voice recognition technology. Another example is that text information can be extracted from the image data through optical character recognition (OCR) technology.
[0041] Further, before segmenting the medical record, the medical record can also be text-cleaned to remove irrelevant characters (such as special symbols, numbers, etc.) in the medical record and retain only the pure text content.
[0042] Further, before screening out multiple keywords related to the medical field from the multiple segments, stop words can also be filtered out from the multiple segments, such as words that appear frequently but do not carry actual meaning (such as "de", "shi", etc.) to reduce noise interference.
[0043] A specific implementation method for screening out multiple keywords related to the medical field from the multiple segments will be illustrated by way of example below.
[0044] In some embodiments, multiple word segments obtained can be optimized through a pre-established medical domain dictionary to obtain multiple keywords. Here, the medical domain dictionary refers to a term library established for the medical domain, which contains frequently occurring professional vocabulary in the medical domain and may also include corresponding phrases and / or abbreviations, aiming to improve the accuracy and efficiency of text processing. In the medical domain, such dictionaries usually contain core terms such as disease names, symptoms, drugs, and diagnosis and treatment techniques, such as "myocardial infarction", "CT scan", "penicillin", etc. The word segmentation in the foregoing embodiments is based on general domain vocabulary and grammar rules. However, in the medical domain, due to its professionalism and the complexity of terms, general word segmentation methods often cannot achieve ideal results. In this embodiment, using the medical domain dictionary to optimize the multiple word segments obtained can improve the accuracy of word segmentation.
[0045] In other embodiments, a large language model can be used to screen out multiple keywords related to the medical domain from the multiple word segments obtained.
[0046] For example, multiple word segment weights can be obtained through a large language model, and multiple keywords related to the medical domain can be screened out from the multiple word segments based on the weights of the multiple word segments. The large language model can perform semantic understanding and context analysis on each word segment based on a large amount of text data learned during its pre-training process, thereby calculating the weight of the word segment. The weight of the word segment reflects the relevance of the word segment to the medical domain (i.e., the probability that the word segment is a keyword related to the medical domain). The higher the weight, the more likely the word segment is to be a keyword in the medical domain. For example, word segments such as "hypertension", "diabetes", and "myocardial infarction" have relatively high weights in the medical domain. Specifically, word segments with weights greater than a preset threshold can be determined as keywords related to the medical domain. Or, the top K word segments with the largest weights can be determined as keywords related to the medical domain.
[0047] Alternatively, the large language model can also filter out word segments unrelated to the medical domain among the multiple word segments, determine the weights of the filtered word segments based on the term frequency-inverse document frequency (TF-IDF) scores of the filtered word segments, and screen out multiple keywords related to the medical domain from the filtered candidate keywords based on the weights of the filtered word segments.
[0048] Specifically, the number of times each token appears in the medical records can be divided by the total number of tokens in the medical records to obtain the term frequency (TF) of the corresponding token. It is also possible to calculate the ratio between the number of medical records containing the token and the total number of medical records, take the logarithm, and perform smoothing to obtain the inverse document frequency. Multiply the term frequency by the inverse document frequency to obtain the TF-IDF score of the token. The higher the TF-IDF score of the token, the higher the importance of the token in the medical records, and thus the stronger the relevance of the token to the medical field.
[0049] In some other embodiments, the multiple tokens obtained can be optimized first through a pre-established medical domain dictionary to obtain multiple candidate keywords, and then the large language model can be used to screen out multiple keywords related to the medical field from the multiple candidate keywords. In this embodiment, the medical domain dictionary and the large language model are combined to screen keywords. The keywords screened out by the medical domain dictionary are used as preliminary keywords (i.e., candidate keywords), and the large language model is used to further process the subsequent keywords. The specific implementation of screening out multiple keywords related to the medical field from the multiple candidate keywords by the large language model can refer to the method of screening out multiple keywords related to the medical field from the multiple tokens obtained in the foregoing embodiments, which will not be elaborated here.
[0050] In addition to extracting multiple keywords from medical records, the above-mentioned multiple medical keywords can also be directly output by the large language model. For example, prompt information in the following form can be input to the large language model: "Please give the medical keywords commonly used in the medical field", so that the large language model outputs multiple medical keywords based on the above prompt information.
[0051] In other examples, the multiple medical keywords can also be obtained by other means, and this application does not limit this.
[0052] In step S14, multiple target medical records can be iteratively screened out from multiple medical records. By performing data screening, it helps to remove noise interference and retain valuable data for further analysis and processing. The medical field has accumulated a vast amount of data. How to efficiently and accurately screen out valuable information from the vast and complex data to support clinical decision-making and scientific research is an urgent problem to be solved currently. In the related art, there are generally two ways to screen medical records:
[0053] One is a statistically based data screening solution, which mainly relies on statistical analysis techniques such as t-test, chi-square test, regression analysis, etc. This method determines which variables (such as the patient's age, gender, medical history, etc.) have a significant impact on a specific disease by setting a certain statistical threshold or significance level. In practical applications, researchers usually first preset some hypotheses based on domain knowledge, and then use statistical software to perform data analysis to verify whether these hypotheses are valid. In addition, dimensionality reduction techniques such as cluster analysis and principal component analysis can be used to process high-dimensional data sets to identify the most important features. This type of solution has the following disadvantages:
[0054] (1) Large data volume and high data quality requirements: Traditional statistical methods often require a large number of high-quality samples to ensure the validity and reliability of the results.
[0055] (2) When faced with complex nonlinear relationships, traditional statistical models may not be able to accurately capture the associations between variables, and thus have poor processing capabilities for complex patterns.
[0056] (3) Lack of flexibility: The pre-set assumptions limit the ability to explore unknown patterns and can easily overlook potentially important information.
[0057] (4) Over-reliance on the knowledge of domain experts: From variable selection to model construction, strong professional background support is required, which increases the cost and time expenditure of the project.
[0058] Another approach is data screening based on deep learning. This approach automatically learns feature representations from input data by building multi-layer neural networks, using these representations to perform tasks such as classification and prediction. Common architectures include convolutional neural networks (CNNs) for image recognition, recurrent neural networks (RNNs) for sequence data analysis, and autoencoders for anomaly detection. Compared to traditional statistical methods, deep learning can directly extract high-level, abstract features from raw data without the need for manual feature engineering. However, this approach has the following drawbacks:
[0059] (1) Poor interpretability: Although deep learning models can achieve high accuracy, their internal working mechanism is like a “black box”, making it difficult to intuitively understand the logic behind each decision.
[0060] (2) Long training time: Building effective deep learning models usually requires a lot of computing resources and time, especially when dealing with large-scale datasets.
[0061] (3) High risk of overfitting: If the model is too complex and the training samples are insufficient, overfitting is likely to occur, resulting in a decrease in generalization ability.
[0062] (4) Data privacy issues: To train a deep learning model, a large amount of personal health information often needs to be collected and stored, which raises concerns about data security and personal privacy protection.
[0063] For this reason, the present application designs target constraint conditions to iteratively screen target medical records. The target medical records screened out each time satisfy the following target constraint conditions:
[0064] (1) The currently screened out target medical records are the medical records that have not been selected among multiple medical records;
[0065] (2) The total number of medical keywords included in the currently screened out target medical records and the medical keywords included in the historically screened out target medical records is maximized. Among them, the total number of the above-mentioned medical keywords refers to the number of non-repeated medical keywords in each of the screened out target medical records, and the repeatedly appearing medical keywords are only counted once. For example, assume that one medical record includes medical keywords "hypertension" and "heart disease", and another medical record includes medical keywords "hypertension" and "coronary heart disease", then the total number of medical keywords included in these two medical records is 3.
[0066] The target medical records screened out each time can be one or more. For the sake of description, the following takes screening out one target medical record each time as an example for illustration. Refer to Figure 2 , assume that the initially obtained medical records include medical record A, medical record B, medical record C, medical record D, and medical record E. During the first iteration screening, the medical record with the largest total number of medical keywords included (assume it is medical record A) can be screened out from multiple medical records as the target medical record. During the second iteration screening, screening can be performed from the other medical records (including medical record B, medical record C, medical record D, and medical record E) except medical record A among multiple medical records. When screening, it can be determined that the total number n of medical keywords included in medical record B and the medical keywords included in medical record A A+B , the total number n of medical keywords included in medical record C and the medical keywords included in medical record A A+C , the total number n of medical keywords included in medical record D and the medical keywords included in medical record A A+D , and the total number n of medical keywords included in medical record E and the medical keywords included in medical record A A+E , and take n A+B , n A+C , n A+D , and n A+EThe medical record corresponding to the largest value among them (assumed to be medical record B) is used as the target medical record in the second iteration. When screening in the third iteration, screening can be performed from other medical records (including medical record C, medical record D, and medical record E) among the multiple medical records except medical record A and medical record B. When screening, it can be determined that the total number of medical keywords included in medical record C, the total number of medical keywords included in medical record A, and the total number of keywords included in medical record B is n A+B+C , the total number of medical keywords included in medical record D, the total number of medical keywords included in medical record A, and the total number of keywords included in medical record B is n A+B+D , and the total number of medical keywords included in medical record E, the total number of medical keywords included in medical record A, and the total number of keywords included in medical record B is n A+B+E , and take the medical record corresponding to the largest value among n A+B+C , n A+B+D and n A+B+E (assumed to be medical record C) as the target medical record in the third iteration. And so on.
[0067] In some embodiments, after each iteration, the total number of the screened target medical records can be determined. If the total number reaches a preset quantity, the iteration is stopped; otherwise, return to the step of iteratively screening multiple target medical records from multiple medical records. In theory, when the total number of initially obtained medical records is large enough, as the number of iterations increases, the number of medical keywords included in the medical records will gradually approach or cover all medical keywords. At this time, the contribution rate of newly added target medical records to keyword coverage will decrease significantly. By setting the preset quantity as the termination condition, the embodiments of the present application can effectively avoid waste of computing resources caused by blind iteration and achieve a balance between efficiency and the number of covered medical keywords while ensuring the sufficiency of target medical record coverage.
[0068] It can be understood that the total number of the screened target medical records in the above embodiments is only an optional iteration stop condition. In other examples, other conditions can also be used as iteration stop conditions. For example, other conditions can include but are not limited to that the current iteration number is greater than a preset number threshold, etc.
[0069] In some embodiments, after each iteration, the medical keywords included in the screened target medical records can be deleted from the multiple medical keywords to obtain several remaining medical keywords. Each time when iterating, screen out the keywords that meet the target constraint conditions, that is, from the medical records that have not been selected among the multiple medical records, screen out the medical record with the largest number of remaining medical keywords included as the target medical record. See Figure 3, assume that the multiple medical keywords obtained in step S12 include {w1, w2, w3, w4, w5, w6, ……}, and the target medical records selected in the first iteration include medical keywords w1 and w2. Then, after the first iteration, medical keywords w1 and w2 can be deleted, and from the unselected medical records among the multiple medical records, the medical records with the largest number of the remaining medical keywords {w3, w4, w5, w6, ……} included are selected as the target medical records. Assume that the target medical records selected in the second iteration include medical keywords w3 and w4. Then, medical keywords w3 and w4 can be deleted again, and from the unselected medical records among the multiple medical records, the medical records with the largest number of the remaining medical keywords {w5, w6, ……} included are selected as the target medical records. And so on.
[0070] In this embodiment, through the screening mechanism (greedy selection strategy) that maximizes the coverage rate of the remaining medical keywords, the currently optimal representative target medical records can be selected in each iteration; moreover, in each iteration, the covered medical keywords in the selected target medical records are deleted (dynamic pruning strategy), making the set of keywords to be matched show a shrinking trend. In summary, through the collaborative optimization of the dynamic pruning mechanism and the greedy selection strategy in the embodiments of the present application, a significant improvement in the efficiency of medical record screening and the coverage effect of medical keywords is achieved.
[0071] In some embodiments, the multiple medical records obtained in step S12 can be added to a queue. Each time an iteration is performed, the target medical records selected during this iteration are moved to the front of the queue. In the next iteration, the target medical record at the front of the queue is first removed from the queue, and then a new round of target medical records is selected. Repeat the above process until the iteration stop condition is met.
[0072] In step S16, a medical large model can be trained based on the selected multiple target medical records. The trained medical large model can provide question-and-answer services related to the medical field for users. Specifically, it can be used to implement functions such as medical knowledge association analysis, treatment plan literature retrieval, and health management knowledge push.
[0073] After training the medical large model, the user query can be input into the medical large model as a prompt message, so that the medical large model outputs a response message based on the user query. Further, the context information related to the user query can be retrieved, a prompt message is generated based on the user query and the context information, and the prompt message is input into the medical large model.
[0074] It should be noted that the large language models mentioned in the embodiments of the present application, such as the large language model used to screen out multiple keywords related to the medical field from multiple word segments, the large language model used to filter out word segments unrelated to the medical field from multiple word segments, and the first large language model trained based on the selected multiple target medical records, can be the same large language model or different large language models.
[0075] In addition, although the data used to train the medical large model in the present application may include disease treatment and diagnosis data, the medical large model in the present application is not directly used to implement disease instructions and diagnoses. Instead, it realizes the purpose of assisting users in understanding their own health status and assisting medical personnel in making clinical judgments by outputting response information such as medical literature abstracts, disease symptom descriptions, treatment principle explanations, and health knowledge entries.
[0076] Figure 4 The overall process of an embodiment of the present application is shown, mainly including the following steps:
[0077] (1) Data preprocessing
[0078] First, text cleaning is performed to remove irrelevant characters (such as special symbols, numbers, etc.) in the initially obtained medical records and retain the pure text content. Then, Chinese word segmentation is performed. A Chinese word segmentation tool (such as jieba) can be used to segment each document to obtain multiple word segments, such as hepatitis B vaccine, melatonin, weight loss, cold, etc. Then, stop words are filtered out from the multiple word segments, that is, words that appear frequently but do not carry actual meaning (such as "of", "is", etc.) are removed to reduce noise interference. According to the characteristics of the medical field, a medical field dictionary can also be constructed and used to further optimize the word segmentation results to obtain multiple candidate keywords to ensure that important terms are accurately captured.
[0079] (2) Medical keyword extraction
[0080] A pre-trained large language model (such as BERT or other large-scale pre-trained models customized for the medical field) is used to analyze the obtained multiple candidate keywords to identify those candidate keywords with higher weights in the current context. The weights of the candidate keywords can be obtained by the large language model. After the large language model obtains the weights of each candidate keyword, the TOP K candidate keywords with weights from large to small are determined as multiple keywords related to the medical field. Among them, the weights of the candidate keywords are used to reflect their importance and uniqueness in multiple medical records. Or, the large language model can further filter out candidate keywords unrelated to the medical field from the multiple candidate keywords. For the remaining candidate keywords, the TOP K optimal algorithm (such as an algorithm for calculating the TF-IDF score of the candidate keywords) is used to screen out multiple keywords related to the medical field.
[0081] (3) Modeling of the maximum coverage problem
[0082] Regard all the keywords obtained in the previous step as the target elements to be covered. Each medical record is regarded as one of the potential options, and the "coverage range" it can provide depends on the number of keywords it contains that are not covered by other records. Among them, a keyword is covered by a medical record if the keyword is included in the medical record. The objective constraint is to find N medical records in all possible combinations of choices such that these N medical records together contain as many different keywords as possible.
[0083] (4) Solving process
[0084] Apply the greedy strategy to solve the above maximum coverage problem: At each iteration, select the medical record that can increase the total coverage quantity the most from the remaining current options and add it to the finally selected set until the predetermined quantity N is reached. In the actual operation process, the computational efficiency issue also needs to be considered. Unnecessary comparisons can be avoided through appropriate pruning techniques to improve the running speed of the algorithm.
[0085] (5) Output result
[0086] The finally selected N medical records will be used to train a higher-level medical large model.
[0087] The embodiment of this application combines the maximum coverage algorithm with the large language model, and uses the maximum coverage algorithm to screen out target medical records from multiple medical records and train the large language model, which has the following advantages:
[0088] (1) By using text cleaning, word segmentation, and stop word filtering techniques specifically for Chinese and the medical field, and introducing a professional term dictionary to optimize the word segmentation results, it is ensured that the extracted keywords more accurately reflect the core information of the medical records. This method can capture key concepts in medical literature better than the general-purpose preprocessing process.
[0089] (2) Use the large language model to deeply understand the text and determine weights for each candidate keyword accordingly. This not only considers the importance of the word segmentation itself, but also comprehensively considers its distribution in the entire corpus (TF-IDF), so as to be able to more scientifically and reasonably quantify the value of each vocabulary.
[0090] (3) Through the maximum coverage algorithm, this application can select the most representative samples from a limited dataset. This method, by effectively integrating and utilizing existing resources, does not require a large amount of high-quality data to ensure the effectiveness of the results. Because it optimizes the selection process based on the existing keyword coverage quantity, it can ensure that as much important information as possible is obtained under limited resource conditions, thus alleviating the demand for a large amount of data to a certain extent and improving resource utilization rate.
[0091] (4) By adopting a method based on the greedy strategy to approximately solve the maximum coverage problem, while ensuring the quality of the selection results, the computational complexity is greatly reduced, making the whole process more feasible and efficient.
[0092] (5) In addition to the core algorithm design, this application also performs pruning techniques during the data screening process, which not only ensures the rapid execution of the screening process but also does not affect the quality of the finally selected dataset.
[0093] (6) It provides a better basic dataset for subsequent large models. By selecting medical records with rich information and obvious features as training materials, the dataset can better reflect the diversity in the real world, which helps to build a more robust and generalization-capable prediction model and can indirectly help improve the model's ability to handle complex patterns.
[0094] (7) Compared with traditional statistical methods, the method combined with large models has stronger learning ability and adaptability. The pre-trained large model can automatically extract features from the original text without the need for artificial assumptions or feature engineering, greatly reducing the cost of human intervention. Therefore, it is more flexible and can explore more unknown patterns.
[0095] (8) Although there may be interpretability issues with deep learning models themselves, this application focuses on the data screening stage and does not completely rely on deep neural networks to make the final decision. Instead, it screens target medical records based on target constraint conditions, improving the interpretability of the solution. In addition, interpretable artificial intelligence techniques (such as LIME, SHAP, etc.) can be adopted in practical applications to enhance model transparency.
[0096] (9) By carefully selecting training samples, this application reduces the input of unnecessary redundant information to the large language model, which helps to shorten the training time and reduce the risk of overfitting. At the same time, high-quality data is also beneficial to improving model performance.
[0097] (10) Although this solution still needs to collect a certain amount of data for analysis, since its goal is to select a small-scale subset that can best represent the overall characteristics rather than all data, it can reduce the risk of personal health information leakage to a certain extent. In addition, measures such as encrypted storage can be taken to protect user privacy during the implementation process.
[0098] (11) This application combines medical expertise with computer science, providing a new perspective and technical means to solve specific problems in clinical practice.
[0099] See Figure 5 , an embodiment of this application also provides a training device for a medical large model, and the device includes:
[0100] An acquisition module 202, configured to acquire multiple medical records and multiple medical keywords;
[0101] A screening module 204, configured to iteratively screen out multiple target medical records from the multiple medical records; in each iteration, based on the following condition, screen out target medical records from the medical records that have not been selected among the multiple medical records: the total number of medical keywords included in the currently screened target medical records and the medical keywords included in the historically screened target medical records is maximized;
[0102] A training module 206, configured to train a first large language model based on the multiple target medical records to obtain a medical large model, and the medical large model is used to provide question-and-answer services related to the medical field for users.
[0103] For the specific implementation details of the above device embodiment, please refer to the foregoing method embodiment, and details will not be described herein again.
[0104] An embodiment of this application also provides a computer device, which at least includes a memory, a processor, and a computer program stored on the memory and executable on the processor. Among them, when the processor executes the program, it implements the method described in any one of the foregoing embodiments.
[0105] Figure 6 FIG. shows a more specific schematic diagram of the hardware structure of a computer device provided by an embodiment of this application. The device may include: a processor 302, a memory 304, an input / output interface 306, a communication interface 308, and a bus 310. Among them, the processor 302, the memory 304, the input / output interface 306, and the communication interface 308 are communicatively connected to each other inside the device through the bus 310.
[0106] The processor 302 can be implemented in the form of a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application. The processor 302 may further include a graphics card, and the graphics card may be an Nvidia titan X graphics card or a 1080Ti graphics card, etc.
[0107] The memory 304 can be implemented in the form of a read-only memory (ROM), a random access memory (RAM), a static storage device, a dynamic storage device, etc. The memory 304 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of the present application through software or firmware, the relevant program codes are stored in the memory 304 and are called and executed by the processor 302.
[0108] The input / output interface 306 is used to connect to an input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0109] The communication interface 308 is used to connect to a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through a wired method (such as USB, network cable, etc.) or can achieve communication through a wireless method (such as a mobile network, Wi-Fi, Bluetooth, etc.).
[0110] The bus 310 includes a path for transmitting information between various components of the device (such as the processor 302, the memory 304, the input / output interface 306, and the communication interface 308).
[0111] It should be noted that although the above device only shows the processor 302, the memory 304, the input / output interface 306, the communication interface 308, and the bus 310, in the specific implementation process, the device may further include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solutions of the embodiments of the present application, and do not have to include all the components shown in the figure.
[0112] An embodiment of the present application provides a computer program product, including a computer program, which implements the method described in any embodiment of the present application when executed by a processor.
[0113] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and the program implements the method described in any of the foregoing embodiments when executed by a processor.
[0114] Computer-readable media includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computer device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0115] Each embodiment in the present application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the description of the method embodiment. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. When implementing the solution of the embodiment of the present application, the functions of the modules can be implemented in the same or multiple software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0116] The above is only the specific implementation manner of the embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the embodiments of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of the present application.
Claims
1. A training method for a medical large model, the method comprising: Obtaining multiple medical records and multiple medical keywords; Iteratively screening out multiple target medical records from the multiple medical records; In each iteration, screening out target medical records from the medical records that have not been selected among the multiple medical records based on the following condition: the total number of medical keywords included in the currently screened target medical records and the medical keywords included in the historically screened target medical records is maximized; Training a first large language model based on the multiple target medical records to obtain a medical large model, where the medical large model is used to provide question-and-answer services related to the medical field for users.
2. The method according to claim 1, where obtaining multiple medical records and multiple medical keywords includes: Obtaining several medical records from the multiple medical records; Performing word segmentation on the obtained several medical records to obtain multiple word segments; Screening out multiple keywords related to the medical field from the multiple word segments.
3. The method according to claim 2, where screening out multiple keywords related to the medical field from the multiple word segments includes: Optimizing the multiple word segments through a pre-established medical field dictionary to obtain multiple candidate keywords; Screening out multiple keywords related to the medical field from the multiple candidate keywords through a second large language model.
4. The method according to claim 3, where screening out multiple keywords related to the medical field from the multiple candidate keywords through a second large language model includes: Obtaining the weights of the multiple candidate keywords through the second large language model, where the weight of a candidate keyword is used to represent the probability that the candidate keyword is a keyword related to the medical field; Screening out multiple keywords related to the medical field from the multiple candidate keywords based on the weights of the multiple candidate keywords.
5. The method according to claim 3, where screening out multiple keywords related to the medical field from the multiple candidate keywords through a second large language model includes: Filtering out the candidate keywords unrelated to the medical field among the multiple candidate keywords through the second large language model; Determining the weights of the filtered candidate keywords based on the TF-IDF scores of the filtered candidate keywords; Screening out multiple keywords related to the medical field from the filtered candidate keywords based on the weights of the filtered candidate keywords.
6. The method according to claim 1, the method further comprising: After each iteration, determining the total number of the already screened target medical records; If the total number reaches a preset quantity, stop the iteration; Otherwise, return to the step of iteratively screening out multiple target medical records from the multiple medical records.
7. The method according to claim 6, the method further comprising: After each iteration, deleting the medical keywords included in the already screened target medical records from the multiple medical keywords to obtain several remaining medical keywords; The iteratively screening out multiple target medical records from the multiple medical records includes: In each iteration, from the medical records that have not been selected among the multiple medical records, the medical record with the largest number of remaining medical keywords included is screened out as the target medical record.
8. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
9. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer program product, including a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.