Medical question and answer model distillation construction method for data feature analysis
By introducing MOE layers and generative adversarial networks into the Q&A big model, combining medical Q&A dataset optimization and pruning processing, a lightweight medical Q&A model is built, solving the problems of large amount of parameters and high deployment costs, and achieving efficient application in resource-constrained scenarios.
Patent Information
- Application Number
- CN202510441389.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
AI Technical Summary
The existing Q&A big model has huge parameters, high deployment costs, high time delays and is difficult to effectively apply in resource-constrained scenarios, especially in the medical field, real-time and accuracy requirements are difficult to meet.
Using the data feature analysis method, by obtaining the MOE layer in the basic large model, using the medical Q&A data set for supervision and fine-tuning tests and parameter restoration, combining the generation of adversarial network to generate sample data, optimize the Q&A data set, and perform pruning processing of the expert network to build a distilled medical Q&A model.
The miniaturization and efficiency of the Q&A model is realized, and deployment efficiency and performance are improved, meeting the real-time and accuracy needs of the medical field.
Smart Images

Figure CN120371957A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of model optimization, and particularly to a method for constructing a distilled medical question-answering model for data feature analysis. Background Art
[0002] With the rapid development of artificial intelligence technology, question-answering models are increasingly widely used in the medical field. However, current large question-answering models generally face problems such as huge parameters, high local deployment costs, and high latency, which greatly limit their practical applications in the medical field. Specifically, existing large question-answering models, such as GPT, DeepSeek, etc., usually have more than hundreds of billions of parameters. This not only makes the deployment and inference costs of the model extremely high, but also makes it difficult to be effectively applied in resource-constrained scenarios, such as mobile terminals and edge computing environments. Especially in the medical field, there are extremely high requirements for the real-time performance and accuracy of the model, and the huge model parameters often lead to slow inference speeds, unable to meet the needs of scenarios such as clinical decision-making assistance and intelligent medical consultations.
[0003] In addition, although the general large model can be fine-tuned to meet the needs of the medical field, this method often has difficulty in efficiently retaining in-depth knowledge in the vertical field, and the fine-tuned model still has redundant parameters, resulting in limited improvement in model performance. At the same time, although traditional model distillation methods can reduce the number of parameters and computational costs by training small models through knowledge transfer, these methods often cannot directly optimize the characteristics of new model architectures such as the mixture of experts architecture, and it is difficult to fully retain the capabilities of experts in specific fields. Summary of the Invention
[0004] In view of the technical problems in the prior art that large question-answering models have huge numbers of parameters, high deployment costs, high latency, and are difficult to be effectively applied in resource-constrained scenarios, the present invention provides a method for constructing a distilled medical question-answering model for data feature analysis to solve these problems.
[0005] The technical solution of the present invention to solve the above technical problems is as follows:
[0006] The present invention provides a method for distilling and constructing a medical Q&A model for data feature analysis. The method includes: obtaining a basic large model, where the basic large model includes a Mixture of Experts (MOE) layer, and the MOE layer includes a gating network and multiple expert networks; obtaining a medical Q&A data set, performing supervised fine-tuning tests on the basic large model, obtaining multiple test scores of the multiple expert networks, and restoring the parameters of the basic large model after the tests are completed; based on the multiple test scores, optimizing the generation of sample data according to the medical Q&A data set to obtain an optimized Q&A data set. Among them, the basic large model is subjected to supervised fine-tuning tests using the generated Q&A data set to obtain the generated test scores of the multiple expert networks, calculating the similarity with the multiple test scores, and performing optimization; using the optimized Q&A data set and the medical Q&A data set, performing supervised fine-tuning tests on the basic large model, obtaining multiple fusion test scores of the multiple expert networks, and pruning the multiple expert networks according to the multiple fusion test scores to obtain a distilled medical Q&A model.
[0007] The beneficial effects of the present invention are as follows: By obtaining a basic large model and embedding an MOE layer, using a medical Q&A data set for supervised fine-tuning tests and parameter restoration, optimizing sample data based on test scores, and then using the optimized data set and the original data set for supervised fine-tuning tests and pruning, a distilled medical Q&A model is finally obtained, effectively solving the technical problems of the large number of parameters, high deployment cost, and difficulty in effectively applying in resource-constrained scenarios of the Q&A large model, realizing the miniaturization and high efficiency of the model, and improving the deployment efficiency and performance of the medical Q&A model. Description of the Drawings
[0008] Figure 1 It is a schematic flow chart of a method for distilling and constructing a medical Q&A model for data feature analysis provided by the present invention.
[0009] Figure 2 It is a schematic flow chart of the medical Q&A data set test in the method for distilling and constructing a medical Q&A model for data feature analysis provided by the present invention. Detailed Embodiments
[0010] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.
[0011] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.
[0012] In the description of the present invention, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present invention is not necessarily construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope that conforms to the principles and features disclosed in the present invention.
[0013] Embodiment:
[0014] As Figure 1 shown, an embodiment of the present invention provides a method for distilling and constructing a medical Q&A model for data feature analysis, the method comprising:
[0015] S10: Obtain a base large model, wherein the base large model includes a MOE layer, and the MOE layer includes a gating network and a plurality of expert networks.
[0016] S20: Obtain a medical Q&A data set, perform supervised fine-tuning tests on the base large model, obtain multiple test scores of the plurality of expert networks, and restore the parameters of the base large model after the tests are completed.
[0017] S30: Based on the multiple test scores, optimize the generation of sample data according to the medical Q&A data set to obtain an optimized Q&A data set. Among them, use the generated Q&A data set to perform supervised fine-tuning tests on the base large model, obtain the generated test scores of the plurality of expert networks, calculate the similarity with the multiple test scores, and perform optimization.
[0018] S40: Use the optimized Q&A data set and the medical Q&A data set to perform supervised fine-tuning tests on the base large model, obtain multiple fusion test scores of the plurality of expert networks, and perform pruning processing on the plurality of expert networks according to the multiple fusion test scores to obtain a distilled medical Q&A model.
[0019] Exemplarily, a method for distillation construction of a medical Q&A model for data feature analysis includes steps such as introducing an MOE layer, performing supervised fine-tuning using a medical Q&A dataset, optimizing sample data generation, model distillation and pruning, aiming to significantly compress the model parameters while maintaining the Q&A performance in the medical field, so as to meet the requirements of the medical field for low-cost, low-latency, and high-precision Q&A models. The medical Q&A model is an automatic Q&A system based on artificial intelligence technology, aiming to answer medical-related questions raised by users, which can cover multiple aspects such as disease symptoms, diagnostic methods, treatment plans, disease prevention, drug use, etc. Distillation means extracting knowledge from a large and complex model, i.e., the base large model, and compressing it into a smaller and more efficient model, aiming to retain the most important information or characteristics while removing redundant or unnecessary parts.
[0020] Specifically, in the distillation process of constructing a medical Q&A model, first, a base large model needs to be obtained. Here, the base large model can be an open-source and widely verified large language model, such as DeepSeek, or a huge language model trained by oneself with a large amount of data. The base large model is internally designed with an MOE (Mixture of Experts) layer, which is a key component in the model structure. The MOE layer consists of two main parts: a gating network and multiple expert networks. The role of the gating network is similar to an intelligent allocator, which dynamically determines which data each expert network should process according to the characteristics of the input data; while the expert networks are a series of smaller models dedicated to processing different subtasks or data characteristics, each with its own unique weights and parameters, capable of processing data in parallel, thereby improving the efficiency and performance of the overall model. After obtaining such a base large model, the distillation construction process can begin. The goal of distillation is to extract the most valuable knowledge from this complex and large model and compress it into a smaller and more efficient model. This usually involves means such as knowledge distillation, model pruning, parameter sharing, etc., to minimize the complexity and computational requirements of the model while ensuring model performance. In summary, obtaining the base large model is the first step in distilling and constructing a medical Q&A model, and this base model has powerful processing capabilities and flexibility due to the MOE layer it contains, providing a solid foundation for subsequent knowledge distillation and model optimization.
[0021] Furthermore, a large amount of high-quality Q&A data is collected from the medical field, and this data will be used as supervision information for fine-tuning the base large model. When conducting the supervised fine-tuning test, the base large model is trained using the medical Q&A dataset. At this time, the gating network in the model will selectively activate the corresponding expert networks according to the characteristics of the input data. Since the dataset focuses on the medical field, the expert networks related to medicine will be frequently activated, while the expert networks dealing with other fields will remain silent. This is because the activation function determines which neurons should be activated to handle the current task based on the characteristics of the input data and the properties of the dataset. During the supervised fine-tuning test, record the activation status of each expert network, especially the proportion of activated neurons is statistically counted. This step helps to understand which expert networks are the most active in handling medical Q&A tasks, so as to give more attention in the subsequent distillation process. At the same time, multiple test scores of multiple expert networks are also obtained, and these scores reflect their performance in handling medical Q&A tasks. It should be noted that the parameters of the base large model need to be restored after the supervised fine-tuning test. This is because it is desired to keep the original state of the base large model unchanged during the distillation process, so that it can be used for the supervised fine-tuning test multiple times and obtain the test scores of different expert networks. Through this process, the contribution and value of each expert network in handling medical Q&A tasks can be evaluated more accurately, providing strong support for subsequent knowledge distillation and model optimization.
[0022] Next, in the process of distillation construction of the medical Q&A model, when facing the challenge of insufficient data volume in the medical Q&A dataset, an optimization strategy for sample data generation based on Generative Adversarial Networks (GANs) can be adopted to enhance the dataset. The core of this strategy lies in leveraging the generation ability of GANs to randomly generate Q&A datasets related to medical Q&A, thereby expanding the diversity of training samples and improving the generalization ability of the model. First, train a generative adversarial network based on the existing medical Q&A dataset. This network consists of two parts: a generator and a discriminator. The task of the generator is to generate realistic medical Q&A pairs, while the task of the discriminator is to distinguish the generated Q&A pairs from those in the real dataset. Through continuous confrontation between the two, the generator gradually learns to generate Q&A pairs that are increasingly close to the real data. After that, use the generated Q&A dataset to conduct supervised fine-tuning tests on the basic large model. In this process, multiple expert networks in the basic large model will make predictions based on the generated data and give corresponding generation test scores, which reflect the performance of the expert networks when processing the generated data. To evaluate the quality of the generated data, calculate the similarity between the generation test scores and multiple test scores obtained previously based on the real dataset. The higher the similarity, the closer the generated data is to the real data, and the greater the training value for the model. Based on the calculation results of the similarity, the generated data can be further optimized, which includes adjusting the parameters of the generative adversarial network, changing the strategy of generating data, or adding additional regularization terms, etc., to improve the quality and diversity of the generated data. By iterating this process multiple times, an optimized Q&A dataset can be gradually obtained. This dataset not only contains the original real data but also a large amount of high-quality generated data. Using this optimized Q&A dataset to train the basic large model can further improve the performance of the model, especially the accuracy and generalization ability when dealing with medical Q&A tasks. In summary, through the optimization strategy of sample data generation based on generative adversarial networks, the problem of insufficient data volume in the medical Q&A dataset can be effectively solved, providing richer and more diverse training samples for the subsequent distillation construction process.
[0023] Finally, make full use of the optimized Q&A dataset and the original medical Q&A dataset to further improve the performance of the model. The core of this step lies in conducting more in-depth supervised fine-tuning tests on the basic large model and pruning multiple expert networks in the model based on the test results. Combine the optimized Q&A dataset with the original medical Q&A dataset to form a more abundant and diverse training sample set. This sample set not only contains real medical Q&A data but also data related to medical Q&A generated through generative adversarial networks. Such a combination helps improve the generalization ability and accuracy of the model when dealing with medical Q&A tasks. Then, use this combined dataset to conduct supervised fine-tuning tests on the basic large model. During this process, multiple expert networks in the model will make predictions based on the input data and give corresponding fusion test scores. These scores comprehensively reflect the performance of the expert networks when dealing with the mixed dataset, including their accuracy, generalization ability, and stability, etc. To further improve the efficiency and performance of the model, prune multiple expert networks according to the fusion test scores. The goal of pruning is to remove those expert networks that contribute less to the model performance and have a low activation rate, such as expert networks focusing on dealing with other fields. These networks are often in a silent state when dealing with medical Q&A tasks and have little impact on the prediction results of the model, so they can be safely removed. Through pruning, the number of parameters and computational complexity of the model can be significantly reduced, thus improving the efficiency and performance of the model during local deployment. At the same time, due to the removal of unnecessary expert networks, the speed and accuracy of the model when dealing with medical Q&A tasks will also be improved. In summary, conducting supervised fine-tuning tests using the optimized Q&A dataset and the original medical Q&A dataset and pruning the expert networks according to the fusion test scores is an important step in the process of distilling and constructing a medical Q&A model. This step helps improve the performance and efficiency of the model and provides strong support for subsequent deployment and applications.
[0024] In a preferred embodiment, obtaining the basic large model includes: obtaining the basic large model; embedding an MOE layer in the basic large model, where the MOE layer includes a gating network and multiple expert networks.
[0025] Optionally, in the early stage of building a medical question-answering model, you first need to obtain a basic large model, which is usually a pre-trained large language model with strong natural language processing capabilities and generalization performance. Obtaining a basic large model is the basis of the entire process and provides a necessary starting point for subsequent steps. Next, a MOE layer (Mixture of Experts) is embedded inside the basic large model. The MOE layer is designed to enhance the professionalism and efficiency of the model when dealing with question-answering tasks in specific fields (such as medical). The MOE layer mainly consists of two parts: a gating network and multiple expert networks. The gating network is one of the core components of the MOE layer. It is responsible for dynamically selecting and activating the corresponding expert network based on the characteristics of the input data. The gating network determines which expert network should be activated to handle the current task by calculating the degree of match between the input data and each expert network. This mechanism ensures that the model can flexibly call the most appropriate expert network when processing different tasks, thereby improving the performance and accuracy of the model. Multiple expert networks constitute another important component of the MOE layer. Each expert network focuses on processing data of a specific type or field and has corresponding professional knowledge and skills. For example, in a medical question-answering model, there may be an expert network that specializes in processing diagnostic information about a disease, while another expert network focuses on processing instructions for the use of a drug. These expert networks process input data in parallel and output their own professional opinions, providing a rich source of information for the selection of the gating network. In summary, obtaining a basic large model and embedding the MOE layer in it is one of the key steps in building a medical question-answering model. This step provides the model with powerful natural language processing capabilities, and the design of the MOE layer enhances the model's professionalism and efficiency in handling question-answering tasks in specific fields.
[0026] In a preferred embodiment, Figure 2 As shown, a medical question and answer dataset is obtained, a supervised fine-tuning test is performed on the basic large model, a plurality of test scores of the plurality of expert networks are obtained, and the parameters of the basic large model are restored after the test is completed, including: obtaining a medical question and answer dataset; inputting each question and answer data in the medical question and answer dataset into the basic large model, performing a supervised fine-tuning test, obtaining the activation rate of each node in the expert network, and obtaining a plurality of activation rate sets; based on the plurality of activation rate sets, a plurality of test scores of the plurality of expert networks are calculated; and the parameters of the basic large model are restored after the test is completed.
[0027] Furthermore, in the process of constructing and optimizing the medical Q&A model, a medical Q&A dataset is obtained. This dataset contains a large number of Q&A pairs related to the medical field, providing data support for the training of the model. Next, the basic large model is subjected to supervised fine-tuning tests using this dataset to further improve the model's performance in handling medical Q&A tasks. Specifically, each Q&A data in the medical Q&A dataset is sequentially input into the basic large model. After receiving the input data, the model activates its internal MOE layer, which contains multiple expert networks. These expert networks will perform parallel processing based on the characteristics of the input data and output their respective professional opinions. During this process, the activation rates of the nodes within each expert network are recorded. These activation rates reflect the activity level of the expert network when processing specific input data. By collecting the activation rates of multiple Q&A data, multiple sets of activation rates are obtained, which provide an important basis for the subsequent calculation of test scores. Next, the test scores of multiple expert networks are calculated based on these sets of activation rates. The test scores comprehensively consider factors such as the activation rates of the expert networks, prediction accuracy, and the collaborative effect with other expert networks, and can objectively reflect the performance of the expert networks in handling medical Q&A tasks. By comparing the test scores of different expert networks, the performance differences between them can be understood, providing a basis for subsequent model optimization. After completing the supervised fine-tuning test, the parameters of the basic large model are restored in a timely manner. This step is to ensure that the model can maintain its original state during subsequent training or deployment processes, avoiding performance degradation or instability caused by parameter changes. By restoring the parameters, the stability and consistency of the model can be ensured, providing a strong guarantee for subsequent model applications. In summary, the basic large model is subjected to supervised fine-tuning tests using the medical Q&A dataset. The test scores of multiple expert networks are calculated by recording and analyzing the activation rates of the expert networks, and the model parameters are restored in a timely manner after the test. This process provides an effective method and steps for constructing and optimizing the medical Q&A model.
[0028] In a preferred embodiment, according to the multiple sets of activation rates, multiple test scores of multiple expert networks are calculated, including: calculating multiple total activation rates according to the multiple sets of activation rates; calculating the ratio of each total activation rate to the sum of the multiple total activation rates as the test score, obtaining multiple test scores of multiple expert networks.
[0029] Specifically, the total activation rate of each expert network is calculated based on multiple collected activation rate sets. The total activation rate refers to the total number or proportion of times the internal nodes of the expert network are activated when processing the entire medical Q&A dataset. This metric can intuitively reflect the activity level of the expert network when processing a specific task. Next, the ratio of each total activation rate to the sum of multiple total activation rates is calculated as the test score of the expert network. This ratio reflects the relative contribution or importance of the expert network among all expert networks. By comparing the test scores of different expert networks, the performance differences and advantages and disadvantages between them can be intuitively understood. Specifically, if the test score of a certain expert network is relatively high, it indicates that it has a high activity level and contribution when processing medical Q&A tasks, and may possess stronger professionalism and accuracy. On the contrary, if the test score of a certain expert network is relatively low, it may mean that it performs poorly when processing a specific task, or its professionalism does not match the current task well. Through this scoring mechanism, the performance of multiple expert networks in medical Q&A tasks can be evaluated more objectively, providing a strong basis for subsequent model optimization and selection. At the same time, this mechanism also provides a method for quantifying the contribution of expert networks.
[0030] In a preferred embodiment, based on the multiple test scores, sample data generation optimization is performed according to the medical Q&A dataset to obtain an optimized Q&A dataset, including: collecting a sample original Q&A dataset according to historical medical Q&A data, and rewriting, supplementing, and annotating each sample original Q&A data to obtain a sample generated Q&A dataset; performing multiple random partitions with replacement on the sample original Q&A dataset and the sample generated Q&A dataset to obtain M groups of generated training data; using a generative adversarial network to construct multiple data generation paths, where each data generation path includes a generator and a discriminator; using the M groups of generated training data to perform supervised training on the multiple data generation paths until convergence; randomly extracting some medical Q&A data from the medical Q&A dataset and randomly inputting them into the data generation paths to generate a first generated Q&A dataset; using the first generated Q&A dataset to perform supervised fine-tuning tests on the basic large model, obtaining multiple first generated test scores of the multiple expert networks, and restoring the parameters of the basic large model after the test; calculating the first generated fitness of the first generated Q&A dataset according to the multiple test scores and the multiple first generated test scores; continuing to randomly extract medical Q&A data to generate a generated Q&A dataset and processing and calculating the generated fitness, performing sample data generation optimization until the optimization converges, and outputting the generated Q&A dataset with the maximum generated fitness to obtain the optimized Q&A dataset.
[0031] Preferably, in order to further improve the performance of the medical Q&A model, the sample data generation of the medical Q&A dataset is optimized based on the test scores of multiple expert networks to obtain a higher-quality and more diverse training dataset. First, a sample original Q&A dataset is collected according to historical medical Q&A data, and each sample of the original Q&A data is rewritten, supplemented, and annotated to obtain a sample-generated Q&A dataset. This step aims to enrich the diversity and accuracy of the dataset and provide strong support for subsequent data generation and model training. Then, the sample original Q&A dataset and the sample-generated Q&A dataset are randomly divided multiple times with replacement to obtain M groups of generated training data. This division method helps to ensure the independence and diversity of each group of training data, thereby improving the stability and accuracy of data generation and model training. Then, the generative adversarial network (GAN) technology is used to construct multiple data generation paths. Each data generation path includes a generator and a discriminator, which continuously generate more realistic and diverse medical Q&A data through mutual competition and cooperation. The M groups of generated training data are used to supervise the training of multiple data generation paths until they converge. This step aims to improve the performance of the generator and the discriminator so that they can generate data that is more in line with real medical Q&A scenarios. After the training is completed, some medical Q&A data is randomly selected from the medical Q&A dataset and randomly input into the data generation path to generate the first generated Q&A dataset. Then, the first generated Q&A dataset is used to conduct a supervised fine-tuning test on the basic large model, and multiple first generated test scores of multiple expert networks are obtained. After the test is completed, the parameters of the basic large model are restored in a timely manner to ensure its stability and consistency. Next, the first generation fitness of the first generated Q&A dataset is calculated based on multiple test scores and multiple first generated test scores. This metric reflects the matching degree and quality of the generated Q&A dataset with the real medical Q&A scenario. To further optimize the quality of the generated Q&A dataset, medical Q&A data is continuously randomly selected, and the above-mentioned generation, testing, and evaluation processes are repeated to calculate the generation fitness and optimize the sample data generation. This optimization process is continuously iterated until the convergence condition is reached, that is, the generation fitness no longer increases significantly. Finally, the generated Q&A dataset with the maximum generation fitness is output as the optimized medical Q&A dataset. Through this optimization process, a higher-quality and more diverse medical Q&A dataset is successfully obtained, providing strong support for subsequent model training and deployment.
[0032] In a preferred embodiment, the first generation fitness of the first generation Q&A dataset is calculated based on the multiple test scores and the multiple first generation test scores, including: obtaining the medical branch field to which each medical Q&A data in the medical Q&A dataset belongs, and statistically obtaining the multiple branch field proportions of the multiple medical branch fields; obtaining the medical branch field to which each first generation Q&A data in the first generation Q&A dataset belongs, and statistically obtaining the multiple first branch field proportions of the multiple medical branch fields; and calculating the first generation fitness of the first generation Q&A dataset based on the multiple test scores, the multiple first generation test scores, the multiple branch field proportions, and the multiple first branch field proportions.
[0033] Specifically, to evaluate the quality of the first generation Q&A dataset, the first generation fitness of the first generation Q&A dataset is calculated based on the multiple test scores and the multiple first generation test scores, in combination with the medical branch field information to which the medical Q&A data in the medical Q&A dataset and the first generation Q&A dataset belong. First, a detailed analysis of the medical Q&A dataset is performed to obtain the medical branch field to which each medical Q&A data belongs, and the multiple branch field proportions of the multiple medical branch fields are statistically obtained. This step aims to understand the distribution of real medical Q&A data and provide a benchmark for subsequent calculations. Next, the same process is performed on the first generation Q&A dataset to obtain the medical branch field to which each first generation Q&A data belongs, and the multiple first branch field proportions of the multiple medical branch fields are statistically obtained. This step aims to understand the distribution of the generated Q&A data for comparison with the real data. Then, in combination with the multiple test scores (i.e., the performance of the expert network on the real medical Q&A dataset) and the multiple first generation test scores (i.e., the performance of the expert network on the generated Q&A dataset), as well as the multiple branch field proportions and the multiple first branch field proportions, a comprehensive calculation is performed. This calculation process takes into account the performance of the expert network, the distribution of the dataset, and the matching degree between them, thereby obtaining the first generation fitness of the first generation Q&A dataset. The first generation fitness is a comprehensive indicator that reflects the matching degree and quality between the generated Q&A dataset and the real medical Q&A scenario. By comparing the first generation fitness of different generated Q&A datasets, an optimized Q&A dataset with higher quality and more in line with actual requirements can be selected, providing strong support for subsequent model training and deployment.
[0034] In a preferred embodiment, the first generation fitness of the first generation Q&A dataset is calculated based on the multiple test scores, the multiple first generation test scores, the multiple branch field proportions, and the multiple first branch field proportions, as follows: where GEN fit is the first generation fitness, α and β are weights, M is the number of multiple expert networks, K iis the test score of the i-th expert network, is the first generated test score of the i-th expert network, σ is a small real number, N is the number of multiple medical branch fields, Z j is the proportion of the j-th medical branch field, is the first proportion of the j-th medical branch field.
[0035] Furthermore, when calculating the first generated fitness of the first generated Q&A dataset, it is necessary to comprehensively consider the test scores and the first generated test scores of multiple expert networks, as well as the proportion of each medical branch field in the medical Q&A dataset and the first generated Q&A dataset. This process aims to comprehensively evaluate the quality of the generated Q&A dataset to ensure that it not only conforms to the performance of the expert network but also accurately reflects the distribution of real medical Q&A data. Specifically, first, the first generated fitness GEN fit this comprehensive index is used to quantify the quality of the generated Q&A dataset. In the calculation process, weights α and β are introduced, which respectively represent the relative importance of the test score and the first generated test score in calculating the fitness. By adjusting these two weights, the performance of the expert network on real data and generated data can be flexibly balanced to meet different application requirements. The specific calculation formula is: where GEN fit is the first generated fitness, α and β are weights, M is the number of multiple expert networks, K i is the test score of the i-th expert network, is the first generated test score of the i-th expert network, σ is a small real number, N is the number of multiple medical branch fields, Z j is the proportion of the j-th medical branch field, It is the proportion of the first branch field in the j-th medical branch field. Then, all expert networks are traversed to obtain their test scores and the first generated test scores. These scores reflect the performance of the expert networks when processing real medical Q&A data and generating Q&A data, and are important bases for evaluating the quality of the generated dataset. At the same time, the proportion of each medical branch field in the medical Q&A dataset and the first generated Q&A dataset is also counted. These proportion data reflect the distribution of the dataset and help to understand the coverage and balance of the generated dataset in each field. Finally, all the above information is substituted into the fitness calculation formula to obtain the first generated fitness of the first generated Q&A dataset. During the calculation process, a small real number σ is introduced to avoid the situation of the denominator being zero, ensuring the stability and accuracy of the calculation. To sum up, by comprehensively considering the test scores of the expert networks, the first generated test scores, and the proportion of the medical branch fields, the first generated fitness of the first generated Q&A dataset is calculated. This index provides strong support for evaluating and optimizing the quality of the generated dataset, and helps to select an optimized Q&A dataset that better meets the actual needs.
[0036] In a preferred embodiment, the optimized Q&A dataset and the medical Q&A dataset are used to perform supervised fine-tuning tests on the basic large model to obtain multiple fusion test scores of the multiple expert networks, and pruning processing is performed on the multiple expert networks according to the multiple fusion test scores to obtain a distilled medical Q&A model, including: using the optimized Q&A dataset and the medical Q&A dataset as the fusion Q&A dataset; using the fusion Q&A dataset to perform supervised fine-tuning tests on the basic large model to obtain multiple fusion test scores of the multiple expert networks; and performing pruning processing on a preset proportion of the expert networks with the smallest fusion test scores according to the multiple fusion test scores to obtain a distilled medical Q&A model.
[0037] Specifically, in order to further improve the performance and efficiency of the medical Q&A model, an optimized Q&A dataset and the original medical Q&A dataset were used together to form a fused Q&A dataset, and supervised fine-tuning tests were conducted on the basic large model. In this process, the high-quality characteristics of the optimized dataset and the authenticity of the original dataset were fully utilized to obtain a more accurate and efficient medical Q&A model. First, the optimized Q&A dataset and the medical Q&A dataset were merged to form a fused Q&A dataset. This dataset not only contains high-quality and diverse generated Q&A data but also retains the real and reliable original Q&A data, providing rich data support for subsequent model training. Then, the basic large model was subjected to supervised fine-tuning tests using the fused Q&A dataset. In this process, multiple expert networks were launched and trained and tested on the fused Q&A dataset simultaneously. Multiple fused test scores were obtained by comparing the performance of each expert network on the test set, and these scores reflect the performance level of the expert network when processing the fused Q&A dataset. Then, pruning was performed on the expert network according to multiple fused test scores. Specifically, a preset ratio was set, and a part of the expert networks with the smallest fused test scores was selected for pruning. This step aims to remove expert networks with poor performance and low contribution, thereby simplifying the model structure and improving computational efficiency. After pruning, a distilled medical Q&A model was obtained. This model not only retains the powerful performance of the basic large model but also further improves computational efficiency and real-time response ability through pruning. At the same time, since the optimized Q&A dataset and the original medical Q&A dataset were used for training and testing, the accuracy and reliability of the distilled medical Q&A model are also fully guaranteed. In summary, by using the optimized Q&A dataset and the medical Q&A dataset for fused training and combining pruning based on the fused test scores of multiple expert networks, a distilled medical Q&A model was successfully obtained. This model performs excellently in terms of performance and efficiency, providing strong support for the practical application of medical Q&A systems.
[0038] A method for distilling and constructing a medical Q&A model for data feature analysis provided by an embodiment of the present invention has at least the following technical effects:
[0039] 1. By introducing a basic large model containing a gating network and multiple expert networks, refined processing of different medical Q&A data is achieved. Each expert network can focus on specific data features or domains, and the gating network is responsible for selecting the appropriate expert network for prediction according to the characteristics of the input data. This design not only improves the generalization ability of the model but also enhances computational efficiency through parallel processing of multiple expert networks. During the training process, the test scores of the expert networks are obtained through supervised fine-tuning tests and optimized accordingly, further improving the accuracy and stability of the model.
[0040] 2. Multiple data generation paths are constructed using generative adversarial networks, and high-quality sample Q&A data is generated through supervised training. This technology not only enriches the content of the medical Q&A dataset but also improves the diversity and authenticity of the data. By continuously iteratively optimizing the first generation fitness of the generated Q&A dataset, the optimal data distribution can be gradually approximated, thus significantly enhancing the model's performance in real medical Q&A scenarios. In addition, the proportion of medical sub-domains is also considered to ensure that the generated data can comprehensively cover all medical fields, further enhancing the model's domain adaptability.
[0041] 3. After obtaining the optimized Q&A dataset, it is combined with the original medical Q&A dataset for fusion training, and the fusion test scores of multiple expert networks are obtained. By pruning a preset proportion of expert networks with the smallest fusion test scores, redundant and poorly performing network nodes can be removed, thereby simplifying the model structure and reducing the computational complexity. This pruning and model distillation strategy not only retains the powerful performance of the basic large model but also achieves model lightweighting, improving the model's real-time response ability and deployment efficiency. At the same time, since the pruning process is based on the fusion test scores, it can ensure that the pruned model has better generalization ability and robustness while maintaining high performance.
[0042] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for distilling and constructing a medical Q&A model for data feature analysis, characterized in that The method includes: Obtain a basic large model, where the basic large model includes a Mixture of Experts (MOE) layer, and the MOE layer includes a gating network and multiple expert networks; Obtain a medical Q&A dataset, perform supervised fine-tuning testing on the basic large model, obtain multiple test scores of the multiple expert networks, and restore the parameters of the basic large model after the testing is completed; Based on the multiple test scores, optimize the generation of sample data according to the medical Q&A dataset to obtain an optimized Q&A dataset. Among them, use the generated Q&A dataset to perform supervised fine-tuning testing on the basic large model, obtain the generated test scores of the multiple expert networks, calculate the similarity with the multiple test scores, and perform optimization; Use the optimized Q&A dataset and the medical Q&A dataset to perform supervised fine-tuning testing on the basic large model, obtain multiple fusion test scores of the multiple expert networks, and perform pruning on the multiple expert networks according to the multiple fusion test scores to obtain a distilled medical Q&A model.
2. The method for constructing a distilled medical Q&A model for data feature analysis according to claim 1, characterized in that Obtain a basic large model, including: Obtain a basic large model; Embed an MOE layer in the basic large model, where the MOE layer includes a gating network and multiple expert networks.
3. The method for constructing a medical Q&A model distillation for data feature analysis according to claim 1, characterized in that Obtain a medical Q&A dataset, perform supervised fine-tuning testing on the basic large model, obtain multiple test scores of the multiple expert networks, and restore the parameters of the basic large model after the testing is completed, including: Obtain a medical Q&A dataset; Input each Q&A data in the medical Q&A dataset into the basic large model for supervised fine-tuning testing, obtain the activation rate of each node in each expert network, and obtain multiple sets of activation rates; Calculate multiple test scores of the multiple expert networks based on the multiple sets of activation rates; Restore the parameters of the basic large model after the testing is completed.
4. The method for constructing a distilled medical Q&A model for data feature analysis according to claim 3, wherein Calculate multiple test scores of the multiple expert networks based on the multiple sets of activation rates, including: Calculate multiple total activation rates based on the multiple sets of activation rates; Calculate the ratio of each total activation rate to the sum of the multiple total activation rates as the test score to obtain multiple test scores of the multiple expert networks.
5. The method for constructing a distilled medical Q&A model for data feature analysis according to claim 1, wherein Based on the multiple test scores, optimize the generation of sample data according to the medical Q&A dataset to obtain an optimized Q&A dataset, including: Collect a sample original Q&A dataset according to historical medical Q&A data, rewrite, supplement, and annotate each sample original Q&A data to obtain a sample generated Q&A dataset; Perform multiple random partitions with replacement on the sample original Q&A dataset and the sample generated Q&A dataset to obtain M groups of generated training data; Use a generative adversarial network to construct multiple data generation paths, where each data generation path includes a generator and a discriminator; Use the M groups of generated training data to perform supervised training on the multiple data generation paths until convergence; Randomly extract some medical Q&A data from the medical Q&A dataset and randomly input them into the data generation paths to generate a first generated Q&A dataset; Using the first generated Q&A dataset, perform supervised fine-tuning testing on the basic large model, obtain multiple first generated test scores of the multiple expert networks, and restore the parameters of the basic large model after the testing is completed; Calculate the first generation fitness of the first generated Q&A dataset based on the multiple test scores and the multiple first generated test scores; Continue to randomly extract medical Q&A data to generate a generated Q&A dataset, process and calculate the generated fitness, and perform sample data generation optimization until the optimization converges, output the generated Q&A dataset with the maximum generated fitness, and obtain the optimized Q&A dataset.
6. The method for constructing a distilled medical Q&A model for data feature analysis according to claim 5, characterized in that, Calculating the first generation fitness of the first generated Q&A dataset based on the multiple test scores and the multiple first generated test scores includes: Obtain the medical branch field to which each medical Q&A data in the medical Q&A dataset belongs, and statistically obtain the proportion of multiple branch fields in multiple medical branch fields; Obtain the medical branch field to which each first generated Q&A data in the first generated Q&A dataset belongs, and statistically obtain the proportion of multiple first branch fields in multiple medical branch fields; Calculate the first generation fitness of the first generated Q&A dataset based on the multiple test scores, the multiple first generated test scores, the multiple branch field proportions, and the multiple first branch field proportions.
7. The method for constructing a distilled medical Q&A model for data feature analysis according to claim 6, characterized in that Calculate the first generation fitness of the first generated Q&A dataset based on the multiple test scores, the multiple first generated test scores, the multiple branch field proportions, and the multiple first branch field proportions, as shown in the following formula: Among them, GEN fit is the first generation fitness, α and β are weights, M is the number of multiple expert networks, and K i is the test score of the i-th expert network, is the first generation test score of the i-th expert network, σ is a small real number, N is the number of multiple medical branch fields, and Z j is the proportion of the j-th medical branch field, is the first proportion of the j-th medical branch field.
8. The method for constructing a distilled medical Q&A model for data feature analysis according to claim 1, characterized in that Using the optimized Q&A dataset and the medical Q&A dataset, perform supervised fine-tuning testing on the basic large model, obtain multiple fusion test scores of the multiple expert networks, and perform pruning processing on the multiple expert networks according to the multiple fusion test scores to obtain a distilled medical Q&A model, including: Use the optimized Q&A dataset and the medical Q&A dataset as the fusion Q&A dataset; Using the fusion Q&A dataset, perform supervised fine-tuning testing on the basic large model to obtain multiple fusion test scores of the multiple expert networks; According to the multiple fusion test scores, perform pruning processing on a preset proportion of expert networks with the smallest fusion test scores to obtain a distilled medical Q&A model.