Human body metabolism multi-task analysis method based on large language model
By constructing a multi-task analysis method for human metabolism based on a large language model, the problem of human metabolism analysis in existing technologies has been solved. It enables efficient and accurate processing of compound description, molar mass calculation, enzyme classification, metabolic reaction type identification, and product prediction, thereby improving the efficiency and analytical capabilities of biomedical research.
Patent Information
- Application Number
- CN202511264156.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies in human metabolic analysis suffer from several drawbacks, including lagging knowledge updates, strong reliance on feature engineering, insufficient ability to process unstructured data, difficulty in generalizing to novel compounds and unknown metabolic pathways, lack of a unified framework for multi-task collaborative processing, inability to effectively understand compound descriptions and enzyme functions, and low efficiency in reaction type identification and product prediction.
A multi-task analysis method for human metabolism based on a large language model is constructed. Data is obtained from the KEGG database, cleaned and converted into a text question-and-answer format, and fine-tuned using the Qwen2.5-7B model. The method supports compound description, molar mass calculation, enzyme classification, metabolic reaction type identification and product prediction. Supervised training and optimization are performed using a GPU cluster, and the method is deployed as an API service for multi-task parallel processing.
It enables rapid and accurate processing of various human metabolic tasks, improves the efficiency and analytical accuracy of biomedical research, reduces development costs, and provides a unified framework for multi-task collaborative processing and cross-task generalized prediction.
Smart Images

Figure SMS_15 
Figure SMS_16 
Figure SMS_22
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence and bioinformatics, in particular to a human metabolism multi-task analysis method based on a large language model. BACKGROUND
[0002] In the fields of systems biology research and drug development, it is crucial to accurately understand and analyze human metabolism. The human metabolism system is composed of a complex network of chemical reactions catalyzed by thousands of enzymes, and these reactions involve a large number of small molecule compounds, enzymes, cofactors, and their spatiotemporal dynamic regulation. Traditional metabolic analysis methods rely on manual annotation or rule-based computational models, which have achieved a certain degree of success in metabolite recognition and reaction type classification, but face problems such as knowledge update lag, strong dependence on feature engineering, and insufficient ability to process unstructured data.
[0003] In addition, existing tools have obvious bottlenecks in tasks such as SMILES structure analysis, metabolic reaction semantic understanding, enzyme function classification, and reaction product prediction. For example, traditional methods can only handle known molecules and reactions, and cannot be generalized to new compounds and unknown metabolic pathways. They lack semantic understanding ability in compound description and enzyme function analysis, making it difficult to associate molecular structure with biological activity and the internal logic of enzyme catalytic mechanism. Reaction type recognition and product prediction tasks rely on independent models, lack a unified framework, and have low efficiency in multi-task collaboration, which limits their application potential in human metabolism data analysis.
[0004] In recent years, artificial intelligence technology, especially large language models, has shown unprecedented generalization ability in natural language understanding, information extraction, and knowledge reasoning tasks. LLM can automatically learn language, structure, and semantic rules from massive heterogeneous texts in unsupervised learning, providing the potential for cross-modal information integration and complex task generalization, and providing a technical basis for building a general metabolic intelligent system with chemical knowledge understanding and reasoning ability. However, there is still a lack of high-performance, large-scale, and large-coverage large language models specifically designed for human metabolism systems, which cannot systematically handle multi-task goals such as SMILES string to molecular weight mapping, compound structure semantic generation, metabolic reaction and enzyme function identification, reaction type inference, and metabolic product prediction. Therefore, it is urgent to develop a metabolic large model system that integrates molecular structure analysis, biochemical semantic understanding, and metabolic knowledge reasoning capabilities to fully realize the application potential of LLM in bioinformatics and promote the evolution of metabolomics towards intelligence and automation. SUMMARY
[0005] The present application provides a human metabolism multi-task analysis method based on a large language model, which aims to solve the problems raised in the background art.
[0006] A human metabolism multi-task analysis method based on a large language model, comprising the following steps: (1) Constructing a multi-task metabolism dataset: obtaining human metabolism-related data from the KEGG database, including but not limited to metabolic reaction pathways, enzyme information, and compound structures, after cleaning the obtained data, the multiple human metabolism-related tasks of compound description, molar mass calculation, compound functional description, enzyme EC classification, metabolic reaction type identification, and metabolic product prediction are uniformly converted into a text question and answer format containing "question-thought-answer", in the uniform format, the thinking process and the answer are both marked with a specific label; (2) Setting fine-tuning parameters: selecting a large language model with generation capability as the base model, configuring fine-tuning parameters, including learning rate setting, training round number, gradient accumulation step number, precision training, scheduler strategy, and distributed training framework; (3) Performing supervised fine-tuning training: based on the dataset constructed in step (1), using the full-parameter fine-tuning method, performing supervised training on the GPU cluster, saving model checkpoints every 500 steps during training, and evaluating multi-task performance on the validation set, while monitoring the loss curve and task accuracy changes, and adjusting the training parameters according to the characteristics of the dataset; (4) Model evaluation and optimization: evaluate the performance of the model on each metabolic sub-task in the validation set, calculate the BLEU-4 score for the compound description task, calculate the integer accuracy, average relative error, and maximum relative error for the molar mass calculation task, calculate the accuracy for the enzyme classification task, reaction type identification task, and metabolic product prediction task, and optimize the model according to the indicators through sample screening, difficult example mining, task reweighting, and knowledge injection; (5) Model deployment and application: deploy the fine-tuned model as an API service, supporting molecular analysis tasks in text input form, input is a joint text containing chemical structure representation and natural language questions, output is a natural language reasoning process and final result marked with specific labels, and is applied to drug metabolism pathway prediction, enzyme function analysis, and metabolic product toxicity evaluation scenarios, and supports multi-task parallel processing and cross-task generalization prediction.
[0007] Preferably, the data cleaning process in step (1) includes identifying and deleting duplicate data, correcting incorrect data according to KEGG system standards and related literature, and filling in missing data based on the internal logic of the human metabolism network.
[0008] Preferably, the "question-thought-answer" format in step (1) is as follows: the question part contains task description and input parameters, the thinking part contains a specific label marked reasoning process, and the answer part contains a specific label marked final result.
[0009] Preferably, the GPU cluster in step (3) adopts H100 model GPU.
[0010] Preferably, the average relative error in step (4) is calculated according to the formula: , The maximum relative error is calculated according to the formula: ; wherein, is the number of samples in the test set, represents the true value of the th sample, represents the predicted value of the th sample.
[0011] Preferably, the specific label in step (5) is a label containing <think>The thinking process of the tag is in accordance with <answer>The answer set of the label.
[0012] A human metabolism large model based on a large language model is trained by using any one of the human metabolism multi-task analysis methods based on a large language model in claims 1-6, and can realize the functions of compound description, molar mass calculation, enzyme EC category classification, metabolic reaction type identification and metabolic product prediction.
[0013] Preferably, the model parameter scale is not less than 7 billion, and the support long context input length is not less than 4K tokens.
[0014] Due to the above-mentioned scheme, the beneficial effects of the present application are: (1) The human metabolism large model can quickly and accurately complete various human metabolism related tasks, greatly improving the processing efficiency and analysis accuracy of human metabolism data in biomedical research.
[0015] (2) Data collection and cleaning based on the KEGG system ensures the authority and comprehensiveness of the data, providing a solid foundation for the accuracy of the model.
[0016] (3) The generation of question-thinking-answer pairs and the formulation of targeted evaluation indicators make the model training and evaluation more scientific and effective, which helps to improve the performance of the model.
[0017] (4) The SFT fine-tuning using the Qwen2.5-7B model fully utilizes the advantages of existing advanced models, and can be optimized according to the characteristics of human metabolism tasks, reducing the development cost and difficulty of the model. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical scheme and advantages of the present application clearer and more apparent, the following embodiments are used to further illustrate the present application. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0019] A human metabolism multi-task analysis method based on a large language model, comprising the following steps: (1) Constructing a multi-task metabolism dataset: obtaining human metabolism related data from the KEGG database, including but not limited to metabolic reaction pathways, enzyme information and compound structures, after cleaning the obtained data, the multi-human metabolism related tasks of compound description, molar mass calculation, compound functional description, enzyme EC category classification, metabolic reaction type identification and metabolic product prediction are uniformly converted into a text question and answer format containing "question-thinking-answer", in the uniform format, the thinking process and the answer are marked with specific labels; (2) Set fine-tuning parameters: Select a large language model with generative capabilities as the base model and configure fine-tuning parameters, including learning rate setting, number of training rounds, gradient accumulation steps, precision training, scheduler strategy and distributed training framework. (3) Perform supervised fine-tuning training: Based on the dataset constructed in step (1), use the full parameter fine-tuning method to perform supervised training on the GPU cluster. Save the model checkpoints every 500 steps during the training process, evaluate the multi-task performance on the validation set, monitor the loss curve and task accuracy changes, and fine-tune the training parameters according to the characteristics of the dataset. (4) Model evaluation and optimization: The performance of the model on each metabolic subtask was evaluated in the validation set. BLEU-4 score was calculated for the compound description task, integer accuracy, average relative error and maximum relative error were calculated for the molar mass calculation task, and accuracy was calculated for the enzyme classification task, reaction type identification task and metabolite prediction task. Based on each indicator, the model was fine-tuned and optimized by means of sample screening, difficult case mining, task reweighting and knowledge injection. (5) Model deployment and application: The fine-tuned model is deployed as an API service, which supports molecular analysis tasks in the form of text input. The input is a joint text containing chemical structure representation and natural language questions, and the output is a natural language reasoning process and final result with specific tags. It can be applied to drug metabolism pathway prediction, enzyme function analysis, and metabolite toxicity assessment scenarios, and supports multi-task parallel processing and cross-task generalization prediction.
[0020] The data cleaning process in step (1) includes identifying and deleting duplicate data, correcting erroneous data according to the KEGG system standard and relevant literature, and filling in missing data based on the inherent logic of the human metabolic network.
[0021] In step (1), the "question-thinking-answer" format is as follows: the question part contains the task description and input parameters, the thinking part contains the reasoning process marked with specific tags, and the answer part contains the final result marked with specific tags.
[0022] In step (3), the GPU cluster uses H100 model GPUs.
[0023] In step (4), the average relative error is calculated using the formula: calculate, The maximum relative error is calculated using the formula: calculate; in, That is the number of samples in the test set. Indicates the first The true value of each sample Indicates the first Predicted values for each sample.
[0024] The specific tag in step (5) is a tag containing <think>The thinking process of the tag is in accordance with <answer>The answer set of the tag.
[0025] A large language model-based human metabolism large model, which can realize the functions of compound description, molar mass calculation, enzyme EC category classification, metabolic reaction type identification and metabolic product prediction; the model parameter scale is not less than 7 billion, and the long context input length is not less than 4Ktoken. Embodiments
[0026] Data collection and cleaning steps (1) Use network crawler technology and API interface provided by KEGG system to comprehensively collect data related to human metabolism in KEGG database, including metabolic pathway information, enzyme information, compound structure and property information, etc.
[0027] (2) Use data cleaning algorithm to process the collected data as follows: (i) For duplicate data, identify and delete them through data comparison algorithm.
[0028] (ii) For error data, such as incorrect compound structure representation, incorrect enzyme classification information, etc., correct them according to KEGG system standards and related literature. Implementation of generating question-thought-answer pairs For the smiles to molar mass task: (i) Extract a large number of different smiles expressions from the cleaned data as question input.
[0029] (ii) Use chemical calculation tools and algorithms to write the thinking process, including how to parse smiles expressions and the steps of calculating molar mass according to the parsed structure.
[0030] (iii) Use chemical calculation software to calculate the accurate molar mass as the answer.
[0031] Example: Q: Calculate the molar mass of C12H22O11 A: <think>1. Count the number of each element: C12 (12 carbons), H22 (22 hydrogens), O11 (11 oxygens)\n2. Find the molar mass of each element (in g / mol):\n-C: 12.0107\n-H: 1.00794\n-O: 15.9994\n3. Calculate the total mass of each element:\n-Carbon: 12 × 12.0107 = 144.1284\n-Hydrogen: 22 × 1.00794 = 22.17468\n-Oxygen: 11 × 15.9994 = 175.9934\n4. Add the masses of each part together: 144.1284 + 22.17468 + 175.9934 = 342.29648< / think> <answer> 342.29< / answer> (2) Compound description task: (i) Select various representative compounds as question input.
[0032] (ii) Combine chemical knowledge and related databases to write the thinking process, analyze chemical properties from the aspects of element composition, chemical bond type, functional group of the compound, and analyze biological activity from the aspects of the mechanism of action of the compound in the organism, the binding ability with the target.
[0033] (iii) Write a detailed and accurate compound description as the answer by referring to authoritative chemical and biomedical databases.
[0034] Example: Q: Describe '[O-]C(=O)CCC(=O)C([O-])=O' A: <think>This SMILES string describes a molecule with the following atomic composition: 5 carbons (C), 6 oxygens (O), and 8 hydrogens (H) inferred from the bonds. The bond types include single bonds (e.g., C-C) and double bonds (e.g., C=O). There are two carboxylate (COO-) groups, one ketone group (C=O), and a central carbon atom connecting two carboxylate groups and a propyl chain. There are no obvious chiral centers in the molecule.< / think> <answer>A dicarboxylic acid is a compound in which the hydrogens on two methyl groups of pyruvic acid are replaced by carboxyl groups.< / answer> (3) Classification task of enzymes involved in metabolic reactions: (i) Collect enzymes involved in different metabolic reactions as the input of the question.
[0035] (ii) Write the thinking process based on enzyme classification standards and related literature, such as classifying according to the reaction type catalyzed by the enzyme (oxidoreductases, transferases, hydrolases, etc.) and the structural characteristics of the enzyme (such as whether it contains a specific cofactor binding site, etc.).
[0036] (iii) Give the accurate enzyme classification result as the answer according to the standard enzyme classification system.
[0037] Example: Q: What is 'D-iditol 2-dehydrogenase'? A: <think>D-iditol 2-dehydrogenase is an oxidoreductase that acts on the CH-OH group of donors with NAD+ or NADP+ as acceptor. This enzyme is usually involved in sugar metabolic pathways, catalyzing the oxidation-reduction of specific substrates.< / think> <answer>{'cls':'Oxidoreductases','ligand':'CH-OH group of donors','receptor':'NAD+ or NADP+'}< / answer> (4) Reaction type identification: (i) Organize a large number of different metabolic reactions and enzymes involved as the input of the question.
[0038] (ii) Analyze the changes of substrates and products in metabolic reactions, and write the thinking process by combining the catalytic mechanism of enzymes to determine the reaction type (such as addition reaction, decomposition reaction, substitution reaction, etc.).
[0039] (iii) Give the accurate reaction type as the answer according to the definition of reaction type in chemistry and biochemistry.
[0040] Example: Q: What is the reaction type of D-Mannitol 1-phosphate + NAD+ <=> beta-D-Fructose 6-phosphate + NADH + H+? A: <think>This question involves a chemical reaction where D-Mannitol 1-phosphate and NAD+ are reactants, producing beta-D-Fructose 6-phosphate, NADH, and H+. The conversion of NAD+ to NADH indicates that this is an oxidation-reduction reaction. Additionally, the conversion of D-Mannitol 1-phosphate to beta-D-Fructose 6-phosphate involves the oxidation of a hydroxyl group, further supporting the classification as an oxidation-reduction reaction.< / think> <answer>Oxidoreductase reaction< / answer> (5) Metabolic product prediction task: (i) Select different metabolic reactions and initial substrates as the input of the question.
[0041] (ii) Refer to known metabolic reaction rules and literature of similar reactions to write the thinking process, consider reaction conditions, enzyme action, etc. to predict possible metabolic products.
[0042] (iii) The list of predicted metabolites is given as an answer by validating the experimental data or referring to an authoritative database.
[0043] Example: Q: Under the catalysis of oxidoreductase 4-hydroxybutanoate: NAD+ oxidoreductase, the reactants [H] C(=O) CCC(O)=O and NC(=O) C1=CN(C=CC1)[C@@H]1O[C@H](COP(O)(=O)OP(O)(=O)OC[C@H]2O[C@H]([C@H](O)[C@@H]2O)n2cnc3c(N)ncnc23)[C@@H](O)[C@H]1O and [H+] participate in the reaction. What is the product? A: <think>1. Reactant 1 [H]C(=O)CCC(O)=O is 4-hydroxybutyric acid, containing a carboxyl group (-COOH) and an aldehyde group (-CHO); Reactant 2 is NAD+, whose C1' position of the nicotinamide ring forms a covalent bond with the substrate.\n2. In the redox reaction, NAD+ acts as an electron acceptor, and the aldehyde group (-CHO) is oxidized to a carboxyl group (-COOH), while NAD+ is reduced to NADH.\n3. Chemical bond changes: the aldehyde group C=O double bond breaks to form the carboxyl group O=C-OH, and the nicotinamide ring of NAD+ obtains a hydrogen ion and an electron to form NADH.\n4. Product 1 should be the carboxylic acid form of 4-hydroxybutyric acid OC(=O)CCCC(O)=O; Product 2 is reduced NADH NC(=O)c1ccc[n+](c1)[C@@H]1O[C@H](COP(O)(=O)OP(O)(=O)OC[C@H]2O[C@H]([C@H](O)[C@@H]2O)n2cnc3c(N)ncnc23)[C@@H](O)[C@H]1O< / think> <answer>OC(=O)CCCC(O)=O, NC(=O)c1ccc[n+](c1)[C@@H]1O[C@H](COP(O)(=O)OP(O)(=O)OC[C@H]2O[C@H]([C@H](O)[C@@H]2O)n2cnc3c(N)ncnc23)[C@@H](O)[C@H]1O< / answer> Model fine-tuning implementation (1) Construct multi-task fine-tuning data: organize the question and answer pairs in multiple metabolic tasks (including molecular formula prediction, molecular weight calculation, enzyme classification, product prediction, etc.) into a unified data format. Each sample combines the question and the thinking part into the prompt input, and the answer part as the answer label, forming <think> + <answer>The format of the supervised fine-tuning sample is unified as the Qwen2.5-7B template structure organization.
[0044] (2) Set the fine-tuning parameters: select Qwen2.5-7B-Instruct as the base model, and use full parameter tuning (full tuning) mode. Configure the training parameters as follows: learning rate is 1.5e-5, training rounds is 4, gradient accumulation step is 16, use cosine learning rate scheduling strategy, use bf16 precision training, and enable FlashAttention2 to improve training efficiency. During training and verification, save the model and perform evaluation every 500 steps, and the evaluation indicators cover multi-task accuracy, BLEU[1], ROUGE[2], etc.
[0045] (3) Start SFT fine-tuning training: perform supervised fine-tuning (SFT) on the constructed training set on the GPU cluster, and the model periodically learns semantic mapping ability and knowledge generation ability for each task. In the training process, real-time monitoring of the performance of the validation set (such as loss curve, task accuracy change), and according to the performance, adjust the parameters or interrupt and resume the training.
[0046] (4) Model evaluation and optimization: after training, use the multi-task validation set (including molecular formula prediction, molecular weight calculation, enzyme classification, product prediction, etc.) to evaluate the fine-tuned model as a whole, collect the performance indicators of each sub-task, and perform data playback or hard example mining for tasks that perform poorly, and fine-tune again if necessary to improve the generalization ability. At the same time, test the original Qwen2.5-7B and deepseek-v3 on five tasks according to the same evaluation indicators and compare the indicators. The evaluation indicators are calculated as follows: (i) For the smiles to molar mass task error rate calculation: (i) In the test data set, calculate whether the molar mass obtained by the model and the true molar mass integer are the same size, divide the number of correct calculations by the total number of test samples, and calculate the coarse-grained accuracy.
[0047] (ii) Calculate the average relative error and the maximum relative error for all samples as evaluation indicators.
[0048] The average relative error is calculated as follows:
[0049] The maximum relative error is calculated as follows:
[0050] where, is the number of samples in the test set, represents the true value of the th sample, represents the a sample predicted value.
[0051] (ii) Compound description task matching degree and completeness calculation: (i) Matching degree calculation: text comparison of the compound description generated by the model with the description in the authoritative database and literature, calculation of the proportion of the same information as the matching degree.
[0052] (ii) Completeness calculation: according to the pre-defined key information points that the compound description should contain, check the number of information points contained in the model-generated description, divide by the total number of information points to obtain the completeness index.
[0053] (iii) Classification task accuracy rate of enzymes participating in metabolic reactions: In the test data set, the number of enzymes classified correctly by the model is counted, and the total number of enzymes is divided to obtain the classification accuracy.
[0054] (iv) Given metabolic reaction and enzyme recognition reaction type task accuracy rate calculation: In the test data set, the number of correctly identified reaction types by the model is counted, and the total number of reactions is divided to obtain the identification accuracy.
[0055] (v) Metabolic product prediction task coincidence degree calculation: In the test data set, the number of metabolic products predicted by the model that are the same as the actual detected metabolic products is calculated, and the number of actual detected metabolic products is divided to obtain the coincidence degree index.
[0056] The results of the comparison test of the five tasks are as follows: Table 1
[0057] In this embodiment, the SMILES to molar mass task: precision and error control advantage The model of the present application is superior to the two comparison models, especially in error control: Integer accuracy: the model of the present application reaches 78.81%, which is 2.65 percentage points higher than Deepseekv3 (76.16%) and nearly 2.7 times higher than the original Qwen2.5-7B (29.14%), which means that the prediction accuracy of the integer bits of the molar mass of the compound is greatly improved; Mean relative error (MRE): the model of the present application is as low as 0.0215, lower than Deepseekv3 (0.0232) and the original Qwen2.5-7B (0.0791), which means that the overall prediction deviation of the model for molar mass is smaller and more stable; MaxRE: The MaxRE of the model of the present application is 0.3813, which is much lower than that of the original Qwen2.5-7B (5.0098) and is also better than that of Deepseekv3 (0.4821), avoiding the interference of extreme errors on the metabolic analysis results.
[0058] In the present embodiment, in the compound description task, the text matching degree and the integrity are crushed and lead, The task evaluates the ability of the model to generate related indicators to describe the structure and properties of the compound. The model of the present application has obvious advantages: BLEU-4 score: The BLEU-4 score of the model of the present application is 55.2167, which is nearly 30 times that of Deepseekv3 (1.8282) and 2.35 times that of the original Qwen2.5-7B (23.4705), indicating that the semantic matching degree of the compound description generated by the model is higher than that of the authoritative literature / database description; ROUGE series indicators: The ROUGE-1 (70.4138), ROUGE-2 (51.1334), and ROUGE-L (61.4744) of the model of the present application are all much higher than those of the two comparison models. Among them, the ROUGE-2 (measuring phrase-level matching) is 2.34 times that of the original Qwen2.5-7B (21.8322) and 9.56 times that of Deepseekv3 (5.3484), proving that the model can more completely and accurately cover the key information of the compound description (such as functional groups and biological activity).
[0059] In the present embodiment, the classification task of enzymes participating in metabolic reactions, the accuracy rate has achieved a qualitative leap, The task evaluates the classification ability of the model to the enzyme EC category. The model of the present application performs much better than the comparison models: The accuracy rate of the model of the present application is 82.33%, which is 5.55 times that of Deepseekv3 (14.83%) and 249.48 times that of the original Qwen2.5-7B (0.33%), completely solving the problem of low precision of traditional models in enzyme function classification and providing reliable enzyme function labeling for metabolic reaction mechanism analysis.
[0060] In the present embodiment, the reaction type recognition task. Under the condition of small sample test, the accuracy rate is 100%, In the test set of 200 samples, the model of the present application shows perfect reaction type recognition ability: The accuracy rate of the model of the present application is 100%, while that of the original Qwen2.5-7B is 61.62% and that of Deepseekv3 is only 57.00%, indicating that the model can accurately capture the structural changes of substrates / products, the conversion of coenzyme states (such as NAD+→NADH), and other key features in metabolic reactions, and accurately determine the reaction type (such as redox reaction).
[0061] In the present embodiment, the metabolite prediction task, the prediction ability is far beyond the existing model, The task is directly related to the drug metabolism path analysis and other core applications, and the model of the present application has obvious advantages: The accuracy of the model of the present application is 72.29%, which is 14.4 times higher than that of Deepseekv3 (5.02%) and 180.7 times higher than that of the original Qwen2.5-7B (0.40%), and can effectively predict the types and structures of metabolites based on enzyme catalytic mechanism and substrate structure, providing key support for drug toxicity evaluation and metabolic pathway analysis.
[0062] In the five core tasks of human metabolism, the model of the present application (based on Qwen2.5-7B-Instruct fine-tuning) is significantly better than the original Qwen2.5-7B and Deepseekv3 model in accuracy, error control, and text generation quality. Especially in traditional difficult tasks such as enzyme classification and metabolite prediction, the performance has achieved a magnitude breakthrough, fully verifying the effectiveness and advancement of the technical solution.
[0063] The above description of the embodiments is to facilitate the understanding and use of the present application by those skilled in the art. Those skilled in the art can easily make various modifications to these embodiments, and apply the general principles described herein to other embodiments without having to go through creative labor. Therefore, the present application is not limited to the above embodiments. Those skilled in the art can make improvements and modifications to the present application without departing from the scope of the present application. The above description is only a preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.< / answer> < / think> < / answer> < / think> < / answer> < / think>
Claims
1. A method for human metabolism multi-task analysis based on a large language model, characterized in that, Comprise the following steps: (1) Constructing multi-task metabolic dataset: obtaining human metabolic related data from KEGG database, including but not limited to metabolic reaction path, enzyme information and compound structure, after cleaning the obtained data, the multiple human metabolic related tasks of compound description, molar mass calculation, compound functional description, enzyme EC category classification, metabolic reaction type identification and metabolic product prediction are uniformly transformed into the text question and answer format containing "question-thinking-answer", in the uniform format, the thinking process and the answer are marked with specific tags; (2) Setting fine-tuning parameters: selecting a large language model with generation capability as the base model, configuring fine-tuning parameters, including learning rate setting, training round number, gradient accumulation step number, precision control, scheduler strategy and distributed training framework; (3) Performing supervised fine-tuning training: based on the dataset constructed in step (1), using full parameter fine-tuning method, performing supervised training on GPU cluster, saving model checkpoints every 500 steps during training, and evaluating multi-task performance on validation set, monitoring loss curve and task accuracy change, and adjusting training parameters according to dataset characteristics; (4) Model evaluation and optimization: evaluate the performance of the model on each metabolic sub-task in the validation set, calculate the BLEU-4 score for the compound description task, calculate the integer accuracy, average relative error and maximum relative error for the molar mass calculation task, calculate the accuracy for the enzyme classification task, reaction type identification task and metabolic product prediction task, and optimize the model through sample screening, difficult example mining, task reweighting and knowledge injection according to the indicators; (5) Model deployment and application: deploy the fine-tuned model as an API service, support molecular analysis tasks in text input form, input is joint text containing chemical structure representation and natural language question, output is natural language reasoning process and final result marked with specific tags, applied to drug metabolic pathway prediction, enzyme function analysis, metabolic product toxicity evaluation scene, and support multi-task parallel processing and cross-task generalization prediction.
2. The human metabolism multi-task analysis method based on a large language model according to claim 1, characterized in that, The data cleaning process in step (1) includes identifying and deleting duplicate data, correcting incorrect data according to KEGG system standards and related literature, and filling in missing data based on the internal logic of human metabolic network.
3. The human metabolism multi-task analysis method based on a large language model according to claim 1, characterized in that, The "question-thinking-answer" format in step (1) is as follows: the question part contains task description and input parameters, the thinking part contains specific tag marked reasoning process, and the answer part contains specific tag marked final result.
4. The human metabolism multi-task analysis method based on a large language model according to claim 1, characterized in that, The GPU cluster in step (3) uses H100 model GPU.
5. The human metabolism multi-task analysis method based on a large language model according to claim 1, characterized in that, The average relative error in step (4) is calculated according to the formula: calculated, The maximum relative error is calculated according to the formula: ; wherein, is the number of samples of the test set, denotes the true value of the th sample, denotes the predicted value of the th sample.
6. The human metabolism multi-task analysis method based on a large language model according to claim 1, characterized in that, The specific tag in step (5) comprises <think>The thinking process of the tag is in accordance with <answer>The answer of the label is composed of.< / answer> < / think> 7. A large language model-based human metabolism large model, characterized by, The method of claim 1-6 can realize compound description, molar mass calculation, enzyme EC category classification, metabolic reaction type identification and metabolic product prediction functions.
8. The human metabolism large model based on a large language model according to claim 7, characterized in that, The model parameter size is not less than 7 billion, and the long context input length is not less than 4K tokens.
Citation Information
Cited By
Method and computing device for generating training data for chemical reaction
CN122245480A
Method of generating training data for chemical reactions, computing device
CN122245480B