Large language model optimization method and system based on clinical test
By fine-tuning training and multimodal data fusion based on target drug properties, the large language model is optimized, which solves the problem of lack of professional knowledge in clinical trials and achieves accurate adaptation and efficient application of the model in clinical trial scenarios.
Patent Information
- Application Number
- CN202510773095.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-12
AI Technical Summary
Large language models lack professional knowledge coverage in the field of clinical trials, resulting in factual errors and deviations in the understanding of professional terminology when processing related tasks, affecting the accuracy and reliability of the application.
By determining the properties of the target drug, matching domain data for fine-tuning training, integrating multimodal clinical trial data, performing reinforcement learning optimization, and establishing a multi-dimensional evaluation indicator system, the model's adaptability in clinical trial scenarios can be gradually improved.
It significantly improves the accuracy and reliability of large language models in clinical trials, providing efficient and professional intelligent support for clinical trials.
Smart Images

Figure CN120632459A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of model optimization technology, and in particular to a large language model optimization method and system based on clinical trials. Background Art
[0002] With the rapid development of artificial intelligence (AI), large language models (LLMs), with their powerful natural language processing capabilities, have shown tremendous potential in healthcare and are widely used in scenarios such as drug development and trials. In clinical trials, large language models can assist researchers with protocol design and data management, significantly improving research efficiency and quality.
[0003] However, the training data for general-purpose large language models often comes from internet text, which has varying quality. The model lacks in-depth coverage of clinical trial expertise, industry standards, and the latest research findings. This makes it prone to factual errors and misinterpretation of professional terminology when handling clinical trial-related tasks. For example, when interpreting inclusion and exclusion criteria in clinical trial protocols, it may provide inaccurate results due to misunderstandings of medical terminology and study conditions.
[0004] In order to give full play to the application value of large language models in clinical trials, it is urgent to propose a large language model optimization method based on clinical trials to promote the safe, reliable and efficient application of large language models in the field of clinical trials. Summary of the Invention
[0005] To this end, the present invention provides a large language model optimization method, system, electronic device, computer storage medium and computer program product based on clinical trials to solve at least one of the above technical problems.
[0006] In a first aspect, the present invention provides a large language model optimization method based on clinical trials, comprising the following method steps: determining a first attribute of a target drug targeted by the clinical trial, obtaining domain data based on matching of the first attribute; generating a small sample data set based on the domain data, and using the small sample data set to fine-tune the large language model; obtaining a plurality of second attributes based on matching of the first attribute, obtaining clinical trial data based on the first attribute and each of the second attributes, performing multimodal data fusion on the clinical trial data, and obtaining a reinforced training data set; wherein the second attribute is an attribute of a related drug that is correlated with the target drug; using the reinforced training data set to perform reinforcement learning optimization on the fine-tuned large language model, and performing model evaluation and iteration on the large language model until the optimization end condition is met.
[0007] In a second aspect, the present invention provides a large language model optimization system based on clinical trials, the system including a controller and a storage medium, wherein the storage medium stores a computer program, and the controller calls and executes the computer program to implement: determining a first attribute of a target drug targeted by the clinical trial, and obtaining domain data based on matching of the first attribute; generating a small sample data set based on the domain data, and using the small sample data set to fine-tune the large language model; obtaining several second attributes based on matching of the first attribute, obtaining clinical trial data based on the first attribute and each of the second attributes, performing multimodal data fusion on the clinical trial data, and obtaining an enhanced training data set; wherein the second attribute is an attribute of a related drug that is correlated with the target drug; using the enhanced training data set to perform reinforcement learning optimization on the fine-tuned large language model, and performing model evaluation and iteration on the large language model until the optimization end condition is met.
[0008] In a third aspect of the present invention, an electronic device is provided, comprising: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the computer program implements any of the methods described above when executed by the processor.
[0009] According to a fourth aspect of the present invention, a computer storage medium is provided, wherein the computer storage medium stores a computer program executable by a processor to implement any of the methods described above.
[0010] According to a fifth aspect of the present invention, a computer program product is provided, which comprises a computer program executable by a processor to implement any of the methods described above.
[0011] The present invention extracts the first attribute matching field data of the target drug for fine-tuning, so that the model can master professional knowledge; the present invention also fully explores the value of clinical trial data of related drugs, integrates multi-source and multi-modal clinical trial data, and enhances the generalization ability of the model; through reinforcement learning and iterative optimization, it achieves precise adaptation to the drug clinical trial scenario, effectively improving the accuracy and reliability of the model in tasks such as program design and data interpretation, and providing efficient and professional intelligent support for clinical trials. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0013] Figure 1 This is a flow chart of a large language model optimization method based on clinical trials disclosed in an embodiment of the present invention.
[0014] Figure 2 It is a schematic diagram of feature representation obtained by processing using a neural network model disclosed in an embodiment of the present invention.
[0015] Figure 3 This is a structural diagram of a large language model optimization system based on clinical trials disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0016] The following specific embodiments illustrate the implementation of this application. Those familiar with the art can easily understand the other advantages and functions of this application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of this application, but not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0017] In addition, the technical features involved in the different embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.
[0018] like Figure 1 As shown, an embodiment of the present invention discloses a large language model optimization method based on clinical trials, including the following method steps: S10, determining the first attribute of the target drug targeted by the clinical trial, and obtaining domain data based on the first attribute matching; generating a small sample data set based on the domain data, and using the small sample data set to fine-tune the large language model.
[0019] Because general-purpose large language models lack specific knowledge about clinical trials of specific drugs, they are prone to bias when handling related tasks. Therefore, we first need to determine the primary attributes of the target drug being tested in the clinical trial, including the target drug's category (such as antibiotics, anticancer drugs, etc.), mechanism of action (such as inhibiting the activity of a certain enzyme), and the type of disease being treated (such as cardiovascular disease, diabetes, etc.).
[0020] Based on the primary attribute, a large amount of clinical trial-related professional literature, research reports, standard operating procedures, and other materials are accurately matched from professional medical databases, academic literature, clinical trial reports, and other channels. This data constitutes the domain data. The selected domain data is cleaned, annotated, and structured to form a small sample dataset. The small sample dataset contains multiple training data related to the primary attribute of the target drug that have been screened, organized, and annotated (e.g., less than 200 pieces). This annotation includes interpreting medical terms in the data and classifying key information in the data.
[0021] This small sample dataset is used to fine-tune the large language model. Through this process, the large language model can initially learn the professional knowledge and language patterns of the target drug clinical trials, significantly reducing errors caused by lack of knowledge when handling related tasks and providing the foundational professional capabilities for subsequent optimization.
[0022] S20, deriving several second attributes based on the matching of the first attribute, obtaining clinical trial data based on the first attribute and each of the second attributes, performing multimodal data fusion on the clinical trial data, and obtaining a reinforced training data set; wherein the second attribute is an attribute of a related drug that is correlated with the target drug.
[0023] Considering that clinical trial data for a single drug may be insufficient in quantity and have a single information dimension, it is difficult to meet the needs of a large language model to fully learn complex clinical trial scenarios. Therefore, based on the determined primary attribute, further matching is performed to identify several other drugs that are related to the target drug, and their attributes are used as secondary attributes. Relevance can be determined based on dimensions such as similar mechanisms of action, treatment of similar diseases, and similar chemical structures.
[0024] Based on the primary and secondary attributes, we extensively acquire multimodal clinical trial data, including text (e.g., trial protocols, medical records), images (e.g., medical imaging, pathological slides), and numerical values (e.g., test indicators, efficacy data). Through multimodal data fusion technology, we transform these different types of data into a unified feature representation and integrate them to form a reinforced training dataset.
[0025] This step integrates clinical trial experience of similar drugs and multi-dimensional information, providing the model with richer and more comprehensive learning materials. This helps the model gain a deeper understanding of the complexity and diversity of clinical trial scenarios, and enhances the model’s generalization and comprehensive analysis capabilities.
[0026] S30: Use the reinforcement training dataset to perform reinforcement learning optimization on the fine-tuned large language model, and perform model evaluation and iteration on the large language model until the optimization end conditions are met.
[0027] The fine-tuned large language model is optimized through reinforcement learning using a reinforced training dataset. By designing a reasonable reward function and using the actual performance of the model in handling clinical trial tasks as feedback signals, the model is guided to adjust its parameters and decision-making strategies, enabling it to make better decisions when faced with various clinical trial tasks.
[0028] At the same time, an evaluation index system covering multiple dimensions such as accuracy, reliability, security, and interpretability is established to conduct a comprehensive and systematic evaluation of the model. If the model does not meet the preset optimization termination conditions, such as the prediction accuracy is lower than the standard, the decision-making process cannot be reasonably traced, etc., the training strategy is adjusted according to the evaluation results, such as optimizing the data screening method, improving the multimodal fusion algorithm, adjusting the reward function, etc., and then a new round of iterative training is carried out. Through continuous optimization and iteration, the performance of the large language model in the field of clinical trials is gradually improved, so that it can accurately adapt to a variety of clinical trial scenarios, provide accurate, reliable, safe and interpretable auxiliary support for clinical trials, and ultimately realize the efficient application of large language models in the field of clinical trials.
[0029] The present invention extracts the first attribute matching field data of the target drug for fine-tuning, so that the model can master professional knowledge; the present invention also fully explores the value of clinical trial data of related drugs, integrates multi-source and multi-modal clinical trial data, and enhances the generalization ability of the model; through reinforcement learning and iterative optimization, it achieves precise adaptation to the drug clinical trial scenario, effectively improving the accuracy and reliability of the model in tasks such as program design and data interpretation, and providing efficient and professional intelligent support for clinical trials.
[0030] As an example, the matching based on the first attribute derives several second attributes, including: using a graph neural network to construct a drug knowledge graph, taking the target drug as the central node, calculating the topological distance with other drug nodes through node embedding technology, and screening other drugs within the distance threshold; using a clustering algorithm to cluster the multidimensional features of the target drug, and selecting other drugs at the cluster center and boundary; the multidimensional features include chemical structure fingerprint maps, pharmacological activity parameters, and clinical trial result vectors; using an association rule mining algorithm to analyze the implicit association between the first attribute and the second attribute, and screening other drugs with a confidence level higher than a confidence threshold; taking the intersection of the other drugs obtained by the above three methods to obtain several target other drugs and their corresponding second attributes.
[0031] During drug clinical trials, clinical trial data from other drugs related to the target drug can be used to inform R&D personnel about clinical trial design and assist in analyzing drug pharmacology. However, the correlations between the target drug and related drugs are complex. To accurately determine the secondary attributes of these other drugs, this paper designs three correlation analysis methods and combines the results of these three methods to obtain more accurate secondary attributes.
[0032] (1) Knowledge graphs can intuitively represent the complex relationships between drugs in the form of graphs, such as the synergistic effects and interaction mechanisms of drugs. Graph neural networks (GNNs) can effectively process data with such graph structures and learn the association information between nodes (drugs). Node embedding technology can be used to map drug nodes into a low-dimensional vector space. Calculating the topological distance in this space can quantify the similarity between drugs and help screen out other drugs that are similar to the target drug in the knowledge graph structure. The specific steps are: 1. Construct a drug knowledge graph: Use a graph neural network to process drug information collected from various data sources (such as medical literature, clinical trial databases, etc.) and construct a knowledge graph containing drug nodes and the relationship edges between them.
[0033] 2. Determine the central node: Use the target drug as the central node of the knowledge graph.
[0034] 3. Node embedding and distance calculation: Using node embedding technology, each drug node is represented as a vector, and the topological distance between the target drug node vector and other drug node vectors is calculated, such as Euclidean distance, cosine similarity, etc.
[0035] 4. Drug Screening: Set a distance threshold to screen out other drugs whose topological distance from the target drug node is within the threshold. These drugs are similar to the target drug in the knowledge graph structure and have certain correlations in terms of mechanism of action and therapeutic effect.
[0036] (2) Clustering algorithms can group drugs with similar characteristics together. By selecting drugs at the center and edge of the cluster, other drugs with similar characteristics to the target drug but with a certain degree of diversity can be found. The specific steps are: 1. Feature extraction: Collect multi-dimensional features of the target drug and other drugs, including chemical structure fingerprints (used to describe the chemical structure characteristics of the drug), pharmacological activity parameters (such as the affinity of the drug for a specific target, etc.), and clinical trial result vectors (such as vectorized data of data such as the efficacy indicators of the drug in different trials).
[0037] 2. Clustering operation: Use clustering algorithms (such as DBSCAN) to perform cluster analysis on these multi-dimensional features and divide similar drugs into the same cluster.
[0038] 3. Drug selection: Select the cluster center of each cluster and other drugs on the edge. The drugs at the cluster center represent the typical characteristics of the cluster, while drugs on the edge may have some unique characteristics and are related to the cluster where the target drug is located.
[0039] (3) The association rule mining algorithm can discover some implicit associations between the first attribute and the second attribute from a large amount of data. By setting a confidence threshold, other drugs corresponding to the second attribute that have a strong correlation with the first attribute can be screened out. This ensures that the selected drugs have a high correlation with the target drug in terms of attributes, improving the data quality of subsequent model training. The specific steps are: 1. Data preparation: Organize data related to the first attribute of the target drug and the second attributes of other drugs.
[0040] 2. Association rule mining: Use association rule mining algorithms (such as Apriori) to mine this data and find association rules between the first attribute and the second attribute. For example, it is found that when a drug has a certain first attribute, there is a high probability that it also has a certain second attribute.
[0041] 3. Drug Screening: Set another confidence threshold to screen out other drugs whose association rule confidence is higher than the confidence threshold. These drugs have stronger attribute associations with the target drug and are more likely to provide valuable information for training the large language model.
[0042] The three aforementioned screening methods screen other drugs from different perspectives (knowledge graph structure, multi-dimensional features, and attribute associations). Each method has its own unique advantages, but may also have certain limitations. Intersection takes into account the results of all three methods, removing drugs that were identified by a single method but are otherwise unrelated to the target drug. This results in a set of other drugs that are highly correlated with the target drug across multiple dimensions. This ensures that the ultimately selected drugs and their corresponding secondary attributes provide more accurate and valuable information for optimizing the large language model.
[0043] As an example, the multimodal data fusion of clinical trial data to obtain an enhanced training data set includes: preprocessing the text data, image data, and numerical data in the clinical trial data respectively; performing a multidimensional analysis on the first attribute of the target drug to obtain typical characteristics of the target drug, wherein the typical characteristics include mechanism of action, applicable diseases, and clinical decision-making process; using a fuzzy logic model to infer the typical characteristics to obtain three fusion proportions, and using a deep learning model to perform multimodal data fusion on the preprocessed text data, image data, and numerical data based on the corresponding fusion proportions to obtain a piece of enhanced training data, thereby constructing an enhanced training data set; wherein the sum of the three fusion proportions is 1.
[0044] Drug clinical trial data includes text, image, and numerical data. This paper fuses these multimodal data to generate a single piece of enhanced training data, which is then used for reinforcement learning optimization of the fine-tuned large language model. Furthermore, the present invention specifically determines the fusion ratio of each modality based on the characteristics of the target drug to improve the training efficiency of the enhanced training data.
[0045] First, the text data is cleaned, segmented, and stop words are removed; the image data is preprocessed by noise reduction, normalization, and size adjustment; and the numerical data is filled with missing values, outliers are processed, and normalized.
[0046] Then, the typical characteristics of the target drug are analyzed, including mechanism of action, applicable diseases, and clinical decision-making process.
[0047] The mechanism of action refers to the specific way and principle by which a drug exerts its therapeutic effect in the body. It typically involves the interaction of the drug with a target within the organism, such as a cell surface receptor, enzyme, or ion channel. For example, aspirin's mechanism of action is to inhibit cyclooxygenase (COX) activity, reducing prostaglandin synthesis and thereby exerting antipyretic, analgesic, and anti-inflammatory effects. Statins, on the other hand, inhibit hydroxymethylglutaryl coenzyme A (HMG-CoA) reductase, reducing cholesterol synthesis and thereby lowering blood lipids and preventing cardiovascular disease.
[0048] Indicated diseases refer to the specific diseases or conditions a drug is approved for treating, preventing, or alleviating. A single drug may be effective for multiple diseases, or it may be indicated for only a single condition. For example, penicillin is primarily indicated for infections caused by susceptible bacteria, such as pneumonia, meningitis, and sepsis; whereas metformin is primarily indicated for type 2 diabetes, managing the condition by improving insulin resistance and lowering blood sugar.
[0049] The clinical decision-making process refers to the process by which doctors, in clinical practice, develop the best treatment plan for their patients based on their specific circumstances, including symptoms, signs, test results, medical history, genetic factors, and other factors, combining their professional knowledge and clinical experience with the latest clinical guidelines and research evidence. For example, for a patient with hypertension, a doctor needs to consider factors such as the patient's age, blood pressure level, and the presence of other comorbidities to decide whether to use a single drug or a combination of drugs, which type of antihypertensive drug is most suitable for the patient, and determine the dosage and duration of medication.
[0050] This paper uses a fuzzy logic model to analyze the different characteristics of the target drug (mechanism of action, applicable diseases, and clinical decision-making process). The performance of the target drug on these three factors is used as the input of the fuzzy logic model. Through the three main steps of fuzzification, rule reasoning, and defuzzification, the fusion ratio of each modality data (text, image, and numerical value) is finally output. The specific process is as follows: 1. Define the input variables: mechanism of action factor ( ): Measures the degree to which the drug's mechanism of action depends on molecular-level interactions, with a value range of 0-1, where 0 indicates no dependence and 1 indicates complete dependence.
[0051] Applicable disease factors ( ): Measures the degree to which drugs treat diseases with typical imaging features, with a value range of 0-1, where 0 means not involved at all and 1 means mainly targeting such diseases.
[0052] Clinical decision factors ( ): Measures the extent to which drug use involves a complex clinical decision-making process, with a value range of 0-1, where 0 indicates a simple decision and 1 indicates a very complex decision.
[0053] 2. Define fuzzy sets: For each input variable, define three fuzzy sets: low (L), medium (M), high (H), and determine the corresponding membership functions. For example, for the mechanism of action factor , the following membership function can be defined: Low (L):
[0054] Medium (M):
[0055] High (H):
[0056] For applicable disease factors and clinical decision factors Similar fuzzy set and membership function definitions are also performed.
[0057] 3. Formulate fuzzy rules: Based on the three typical characteristics given above, formulate the following fuzzy rules: Rule 1: If is high and is low and If it is low, then the fusion ratio of numerical data is high, the fusion ratio of image data is medium, and the fusion ratio of text data is low.
[0058] Rule 2: If is low and is high and If it is low, then the fusion ratio of image data is high, the fusion ratio of numerical data is medium, and the fusion ratio of text data is low.
[0059] Rule 3: If is low and is low and If it is high, then the fusion ratio of text data is high, the fusion ratio of numerical data is medium, and the fusion ratio of image data is low.
[0060] A total of 27 rules are formulated to cover all possible input combinations.
[0061] 4. Rule matching and reasoning: For a given input value , calculate the membership of each input variable in each fuzzy set, and then make inferences based on the fuzzy rules. For example, for rule 1, calculate ,in represents the minimum value), and the activation degree of the rule is obtained. This calculation is performed on all the above rules to obtain the activation degree of each rule.
[0062] 5. Defuzzification: Define output variables: numerical data ratio ( ): The value range is 0-100%; the image data ratio ( ): The value range is 0-100%; the proportion of text data ( ): The value range is 0-100%, and meets .
[0063] For each output variable, define three fuzzy sets: low (L), medium (M), and high (H), and determine the corresponding membership functions. Aggregate the outputs of all rules, for example, using the maximum-minimum synthesis method. Use defuzzification methods such as the centroid method to convert the aggregated fuzzy outputs into specific numerical values, determining the fused proportion of each modal data.
[0064] After using the fuzzy logic model to infer the typical characteristics and obtain three fusion ratios, the deep learning model is used to perform multimodal data fusion. Specifically: Figure 2 As shown, a multi-input neural network model is constructed, feeding preprocessed text, image, and numerical data into different sub-networks. For text data, features are extracted using a pretrained language model (such as BERT); for image data, features are extracted using a convolutional neural network (such as ResNet); and for numerical data, features are extracted using fully connected layers. The normalized features extracted by each sub-network are weighted and fused according to the corresponding fusion ratios described above, for example, by weighted summation to obtain a unified feature representation.
[0065] The fused feature representations are combined with corresponding labels (such as clinical trial results and efficacy evaluations) to form a reinforcement training dataset. This reinforcement training dataset is divided into a training set, a validation set, and a test set. This is used for subsequent reinforcement learning optimization of the fine-tuned large language model. The details are not detailed here.
[0066] As an example, before using the reinforcement training dataset to perform reinforcement learning optimization on the fine-tuned large language model, the method also includes: if it is determined that the amount of reinforcement training data in the reinforcement training dataset is lower than a first threshold, using a generative adversarial network to perform data expansion on a number of first-type reinforcement training data, second-type reinforcement training data, and third-type reinforcement training data using different expansion strategies to obtain a new reinforcement training dataset with an amount of reinforcement training data higher than a second threshold; wherein the first threshold is lower than the second threshold.
[0067] Because the amount of data in the reinforcement training dataset directly affects the optimization results of the large language model, insufficient data can lead to inadequate model training and poor generalization. To address this issue, before using the reinforcement training dataset to optimize the model for reinforcement learning, if the amount of training data in the reinforcement training dataset is determined to be below a first threshold (e.g., the minimum data amount required for model training), differentiated generative adversarial network (GAN) strategies are used to augment the data corresponding to the three types of data in the dataset, namely, the first type of reinforcement training data, the second type of reinforcement training data, and the third type of reinforcement training data.
[0068] Data augmentation strategies can vary depending on the type of GAN used. For example, for text data, a text-generating adversarial network (e.g., TextGAN) is used. A discriminator determines the differences between the generated text and real clinical text, guiding the generator to mimic the language style and professional logic of real clinical trial protocols and medical records, generating more standardized medical text. For image data, due to its high dimensionality, complex spatial structure, and the extremely high requirements for authenticity and detail in medical images, a conditional generative adversarial network (e.g., cGAN) is used with conditional inputs (e.g., disease type, imaging modality, etc.) to generate medical images for specific scenarios (e.g., lung CT, tumor MRI). A discriminator is used to ensure that the newly added images conform to the texture and structural characteristics of real medical images. For numerical data, a standard generative adversarial network is used to generate reasonable numerical samples based on the distribution patterns of real clinical indicators (e.g., test values, efficacy data).
[0069] As an example, the generative adversarial network is used to perform data expansion on a number of first-type reinforcement training data, second-type reinforcement training data, and third-type reinforcement training data using different expansion strategies, including: adding a multi-head attention mechanism to the generative adversarial network, and the multi-head attention mechanism includes a scaling factor and a multi-head attention weight; for the first-type reinforcement training data, the word frequency and text length of the professional terms in the corresponding text data are obtained, and the scaling factor is adjusted according to the word frequency, and the multi-head attention weight is adjusted according to the text length; for the second-type reinforcement training data, the significance mask value and image resolution of the lesion area in the corresponding image data are obtained, and the scaling factor is adjusted according to the image resolution, and the multi-head attention weight is adjusted according to the significance mask value; for the third-type reinforcement training data, the coefficient of variation and numerical dimension of the corresponding numerical data are obtained, and the scaling factor is adjusted according to the number of variations, and the multi-head attention weight is adjusted according to the numerical dimension.
[0070] This paper introduces a multi-head attention mechanism into the Generative Adversarial Network (GAN). The multi-head attention mechanism can help the GAN focus on the important parts of the data. Moreover, the following multi-head attention mechanisms can be added to the three different types of GANs mentioned above: ;in, (query vector), (key vector), (value vector) is the input, 、 is a learnable parameter, is the key vector dimension. The expansion strategy of the present invention is to adjust the scaling factor and multi-head attention weights To achieve this, specifically: 1) For the first type of enhanced training data, it is mainly based on text data (the fusion of text data accounts for the highest proportion), with strong semantic relevance and high reliance on professional terminology, and attention should be paid to the consistency of context logic and medical knowledge.
[0071] The scaling factor The frequency of professional terms in the corresponding text data is associated with the higher the frequency of the text data, the higher the scaling weight is, so as to enhance the attention of the generative adversarial network to key information; at the same time, the multi-head attention weight is dynamically adjusted according to the length of the text. ,For long texts, we increase the number of heads to capture multi-scale semantics.
[0072] For example, using the TF-IDF algorithm, the frequency of professional terms (several professional terms are pre-stored in the database) and the length of the text are extracted from the corresponding text data. The weight is obtained based on the frequency comparison. , then the scaling factor ,in is a hyperparameter.
[0073] If the text length exceeds the threshold , then increase the number of long positions to , To the text length, and reinitialize To adapt to the new multi-head structure.
[0074] 2) For the second type of enhanced training data, it is mainly derived from image data (image data fusion accounts for the highest proportion), its spatial structure is complex, the lesion area needs to be focused on, and it requires precise capture of local features.
[0075] Multi-head attention weights The saliency association with the lesion area in the image is used to locate the lesion and generate a saliency mask through the target detection model (such as YOLO) , the area with higher mask value has greater corresponding attention weight; at the same time, the scaling factor is dynamically adjusted according to the image resolution , high-resolution images increase the scaling factor to balance the computational complexity.
[0076] For example, using a pre-trained medical image detection model to generate a saliency mask for an input image ,in is the number of pixels in the vertical direction of the image, is the number of pixels in the horizontal direction of the image; the mask and multi-head attention weights Element-wise multiplication, that is , strengthen the attention of the lesion area; according to the image resolution Adjust the zoom factor for , is the base resolution.
[0077] 3) For the third type of reinforcement training data, it is mainly based on numerical data (the fusion of numerical data accounts for the highest proportion), with sensitive statistical distribution and strong correlation between indicators. It is necessary to highlight the impact of outliers and key indicators.
[0078] The scaling factor Coefficient of variation of the value Correlation, the indicator with a larger coefficient of variation corresponds to a lower scaling factor to enhance the attention to fluctuating data; at the same time, the multi-head attention weight is dynamically adjusted according to the numerical dimension , high-dimensional data increases the number of multiple heads to separate complex relationships.
[0079] For example, the coefficient of variation of each indicator is calculated for the input numerical data. ( is the standard deviation, is the mean); adjust the scaling factor based on the coefficient of variation for ,in is a hyperparameter. If the numerical dimension Exceeding the threshold , then increase the number of long positions to , and redistribute To adapt to the new dimension.
[0080] In addition, the corresponding scaling factors can be determined in advance for the three types of reinforcement training data mentioned above and multi-head attention weights The multi-head attention mechanism parameters are jointly trained with the GAN generator and discriminator parameters, and all parameters are updated through backpropagation. After each training batch of N, data characteristics (such as text term weights, image saliency, and numerical coefficient of variation) are recalculated and the attention mechanism parameters are dynamically updated. The quality of the generated data is evaluated on the validation set. If the generation effect of a certain type of data is poor, the attention parameters associated with this data are adjusted through hyperparameter search (such as Bayesian optimization). .
[0081] like Figure 3 As shown, an embodiment of the present invention further provides a large language model optimization system based on clinical trials, the system comprising a controller (101) and a storage medium (201), wherein the storage medium (201) stores a computer program, and the controller (101) calls and executes the computer program to implement: determining a first attribute of a target drug targeted by the clinical trial, obtaining domain data based on the first attribute matching; generating a small sample data set based on the domain data, and fine-tuning the large language model using the small sample data set; obtaining a plurality of second attributes based on the first attribute matching, obtaining clinical trial data based on the first attribute and each of the second attributes, performing multimodal data fusion on the clinical trial data, and obtaining a reinforced training data set; wherein the second attribute is an attribute of a related drug that is correlated with the target drug; performing reinforcement learning optimization on the fine-tuned large language model using the reinforced training data set, and performing model evaluation and iteration on the large language model until the optimization end condition is met.
[0082] An embodiment of the present invention further provides an electronic device comprising: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the computer program implements any of the aforementioned methods when executed by the processor.
[0083] An embodiment of the present invention further provides a computer storage medium storing a computer program that can be executed by a processor to implement any of the methods described above.
[0084] An embodiment of the present invention further provides a computer program product, which includes a computer program that can be executed by a processor to implement any of the methods described above.
[0085] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0086] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A large language model optimization method based on clinical trials, characterized by: The method comprises the following steps: determining a first attribute of a target drug for a clinical trial, obtaining domain data based on matching of the first attribute; generating a small sample data set based on the domain data, and fine-tuning a large language model using the small sample data set; deriving a plurality of second attributes based on the matching of the first attributes, acquiring clinical trial data based on the first attributes and each of the second attributes, and performing multimodal data fusion on the clinical trial data to obtain a reinforced training data set; wherein the second attributes are attributes of related drugs that are correlated with the target drug; Use the reinforced training dataset to perform reinforcement learning optimization on the fine-tuned large language model, and perform model evaluation and iteration on the large language model until the optimization end conditions are met.
2. The method for optimizing a large language model based on clinical trials according to claim 1, characterized in that: The method of obtaining several second attributes based on the matching of the first attribute includes: using a graph neural network to construct a drug knowledge graph, taking the target drug as the central node, calculating the topological distance with other drug nodes through node embedding technology, and screening other drugs within a distance threshold; using a clustering algorithm to cluster the multidimensional features of the target drug, and selecting other drugs at the cluster center and boundaries; the multidimensional features include chemical structure fingerprint maps, pharmacological activity parameters, and clinical trial result vectors; using an association rule mining algorithm to analyze the implicit association between the first attribute and the second attribute, and screening other drugs with a confidence level higher than a confidence threshold; taking the intersection of the other drugs obtained by the above three methods to obtain several target other drugs and their corresponding second attributes.
3. The method for optimizing a large language model based on clinical trials according to claim 2, characterized in that: The multimodal data fusion of the clinical trial data to obtain an enhanced training data set includes: preprocessing the text data, image data, and numerical data in the clinical trial data respectively; performing a multidimensional analysis on the first attribute of the target drug to obtain typical characteristics of the target drug, wherein the typical characteristics include mechanism of action, applicable diseases, and clinical decision-making process; using a fuzzy logic model to infer the typical characteristics to obtain three fusion proportions, and using a deep learning model to perform multimodal data fusion on the preprocessed text data, image data, and numerical data based on the corresponding fusion proportions to obtain a piece of enhanced training data, thereby constructing a reinforced training data set; wherein the sum of the three fusion proportions is 1.
4. The method for optimizing a large language model based on clinical trials according to claim 3, characterized in that: Before using the reinforcement training dataset to perform reinforcement learning optimization on the fine-tuned large language model, the method further includes: if it is determined that the amount of reinforcement training data in the reinforcement training dataset is lower than a first threshold, using a generative adversarial network to respectively perform data expansion on a plurality of first-type reinforcement training data, second-type reinforcement training data, and third-type reinforcement training data using different expansion strategies to obtain a new reinforcement training dataset with an amount of reinforcement training data higher than a second threshold; wherein the first threshold is lower than the second threshold.
5. The method for optimizing a large language model based on clinical trials according to claim 4, characterized in that: The method uses a generative adversarial network to perform data expansion on a number of first-type reinforcement training data, second-type reinforcement training data, and third-type reinforcement training data using different expansion strategies, including: adding a multi-head attention mechanism to the generative adversarial network, wherein the multi-head attention mechanism includes a scaling factor and a multi-head attention weight; for the first-type reinforcement training data, obtaining the word frequency and text length of professional terms in the corresponding text data, adjusting the scaling factor according to the word frequency, and adjusting the multi-head attention weight according to the text length; for the second-type reinforcement training data, obtaining the significance mask value and image resolution of the lesion area in the corresponding image data, adjusting the scaling factor according to the image resolution, and adjusting the multi-head attention weight according to the significance mask value; for the third-type reinforcement training data, obtaining the coefficient of variation and numerical dimension of the corresponding numerical data, adjusting the scaling factor according to the number of variations, and adjusting the multi-head attention weight according to the numerical dimension.
6. A large language model optimization system based on clinical trials, characterized by: The system includes a controller and a storage medium, wherein the storage medium stores a computer program, and the controller calls and executes the computer program to: determine a first attribute of a target drug for a clinical trial, obtain domain data based on the first attribute matching; generate a small sample data set based on the domain data, and use the small sample data set to fine-tune a large language model; deriving a plurality of second attributes based on the matching of the first attributes, acquiring clinical trial data based on the first attributes and each of the second attributes, and performing multimodal data fusion on the clinical trial data to obtain a reinforced training data set; wherein the second attributes are attributes of related drugs that are correlated with the target drug; Use the reinforced training dataset to perform reinforcement learning optimization on the fine-tuned large language model, and perform model evaluation and iteration on the large language model until the optimization end conditions are met.
7. The large language model optimization system based on clinical trials according to claim 6, characterized in that: The multimodal data fusion of the clinical trial data to obtain an enhanced training data set includes: preprocessing the text data, image data, and numerical data in the clinical trial data respectively; performing a multidimensional analysis on the first attribute of the target drug to obtain typical characteristics of the target drug, wherein the typical characteristics include mechanism of action, applicable diseases, and clinical decision-making process; using a fuzzy logic model to infer the typical characteristics to obtain three fusion proportions, and using a deep learning model to perform multimodal data fusion on the preprocessed text data, image data, and numerical data based on the corresponding fusion proportions to obtain a piece of enhanced training data, thereby constructing a reinforced training data set; wherein the sum of the three fusion proportions is 1.
8. An electronic device, characterized in that: The electronic device includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the computer program implements the method according to any one of claims 1 to 5 when executed by the processor.
9. A computer storage medium, characterized in that: The computer storage medium stores a computer program that can be executed by a processor to implement the method according to any one of claims 1 to 5.
10. A computer program product, characterized in that: The computer program product comprises a computer program that can be executed by a processor to implement the method according to any one of claims 1 to 5.
Citation Information
Cited By
Multi-role configuration and effective judgment method based on semantic arbitration
CN120809299A