Construction method of type 2 diabetes mouse model combining ai and multi-target knockout

By constructing an interaction map encompassing pancreatic β-cell function, insulin signaling, and hepatic lipid metabolism pathways, and combining artificial intelligence prediction with ethical compliance conditions, a type 2 diabetes mouse model was optimized. This solved the problems of model complexity and inefficiency in existing technologies, achieving efficient and accurate model construction and validation.

CN121122378BActive Publication Date: 2026-04-17NANJING LINGNUO BIOMEDICAL TECH RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING LINGNUO BIOMEDICAL TECH RES INST CO LTD
Filing Date
2025-08-28
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, single-gene knockout models are difficult to fully reflect the complex multifactorial pathological mechanisms of type 2 diabetes, and multi-gene knockout lacks systematic prediction methods, resulting in low efficiency and poor reproducibility in screening target combinations. Animal model construction lacks dynamic feedback mechanisms, and resource consumption is high and research cycles are long.

Method used

This study combines AI with a multi-target knockout method to construct a type 2 diabetes mouse model. By acquiring and standardizing multi-omics data and phenotypic baseline data of mice, an interaction map of pancreatic β-cell function, insulin signaling, and hepatic lipid metabolism pathways is constructed. Artificial intelligence is used to predict and generate candidate target combinations and phenotypic response results. Target target combinations and phenotypic prediction results are screened in accordance with ethical compliance conditions, and the model is optimized in real time.

Benefits of technology

This approach enables mouse models to more closely resemble the pathological characteristics of clinical type 2 diabetes, improves research efficiency, reduces the number of invalid experiments, shortens the model establishment cycle, ensures that the model phenotype is highly consistent with expectations, meets regulatory requirements, and forms a traceable research resource.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121122378B_ABST
    Figure CN121122378B_ABST
Patent Text Reader

Abstract

This invention discloses a method for constructing a type 2 diabetes mouse model combining artificial intelligence prediction and multi-target knockout. The method involves acquiring and standardizing multi-omics data and baseline phenotypic data of mice to form a standardized modeling dataset; constructing an interaction map covering pancreatic β-cell function based on this data; selecting target target combinations and target phenotypic prediction results; generating construction instructions and submitting them to a controlled facility for multi-gene knockout modification to obtain modified mice and modification receipts; obtaining measured phenotypic data of the modified mice based on the modification receipts; comparing this data with the target phenotypic prediction results to generate phenotypic deviation metrics and phenotypic determination results; and registering the target target combinations, prediction results, measured phenotypic data, and modification receipts when the phenotypic determination results are satisfactory to form a type 2 diabetes composite mouse model. This invention achieves closed-loop optimization of AI-driven multi-target genetic modification, improving the efficiency and accuracy of constructing animal models of complex metabolic diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mouse model technology, and in particular to a method for constructing a type 2 diabetic mouse model that combines AI with multi-target knockout. Background Technology

[0002] In current technologies, animal model research for type 2 diabetes has been extensively conducted. Researchers typically use genetic engineering techniques such as single-gene knockout, point mutation, or transgenic overexpression to construct mouse models that exhibit some characteristics of diabetes. Furthermore, some models incorporate dietary interventions or environmental factors to enhance metabolic disorder phenotypes, thereby mimicking insulin resistance, impaired glucose tolerance, or obesity-related pathologies. In recent years, with the development of high-throughput sequencing and multi-omics technologies, researchers have been able to conduct more systematic explorations of disease-related mechanisms at the gene, transcription, protein, and metabolic levels, which has, to some extent, driven improvements in diabetes models.

[0003] However, several problems exist in existing technologies. First, single-gene knockout models cannot fully reflect the complex multifactorial pathological mechanisms of type 2 diabetes, and their phenotypes often differ from those of clinical patients. Second, although studies on multi-gene knockout have emerged, they mostly rely on empirical judgment and lack systematic predictive methods, resulting in low efficiency and poor reproducibility in screening target combinations. Third, the construction of animal models often lacks dynamic feedback mechanisms, and model validation results are difficult to promptly inform model design, leading to high resource consumption and long research cycles.

[0004] Given the above, there is an urgent need for a new approach that combines data-driven and experimental feedback to develop more accurate and efficient combined mouse models of type 2 diabetes. Summary of the Invention

[0005] This application provides a method for constructing a type 2 diabetic mouse model that combines AI and multi-target knockout, in order to improve the efficiency and accuracy of constructing animal models of complex metabolic diseases.

[0006] This application provides a method for constructing a type 2 diabetes mouse model that combines AI and multi-target knockout, including:

[0007] Acquire and standardize mouse multi-omics data and phenotypic baseline data to form a standardized modeling dataset;

[0008] Based on the standardized modeling dataset, an interaction map containing pancreatic β-cell function, insulin signaling, and hepatic lipid metabolism pathways was constructed, and artificial intelligence was used to predict and generate candidate target combinations and candidate phenotypic response results.

[0009] Based on the candidate target combination and the candidate phenotype response results, and in conjunction with ethical compliance conditions, the target target combination and target phenotype prediction results are obtained through screening.

[0010] Based on the target combination and the target phenotype prediction results, a construction instruction is generated, and the construction instruction is submitted to a controlled facility to perform genetic modification operations, resulting in modified mice and modification receipts;

[0011] Based on the modification receipt, the measured phenotypic data of the modified mice are obtained, and the measured phenotypic data are compared with the target phenotypic prediction results to generate phenotypic deviation measurement and phenotypic determination results.

[0012] When the phenotypic determination result is not met, the interaction map is updated using the phenotypic deviation metric to generate a corrected candidate target combination and a corrected candidate phenotypic response result, and the process is returned to the screening step; when the phenotypic determination result is met, the target target combination, the target phenotypic prediction result, the measured phenotypic data, and the modification receipt are registered to form a traceable type 2 diabetes composite mouse model.

[0013] The beneficial effects of the technical solution provided in this application include:

[0014] (1) Through artificial intelligence prediction based on standardized multi-omics data and phenotypic baseline data, it is possible to systematically identify and screen multi-gene target combinations related to pancreatic β-cell function, insulin signaling, and liver lipid metabolism pathways, thereby making the constructed mouse model closer to the complex pathological characteristics of clinical type 2 diabetes. (2) Using AI to predict candidate target combinations and combining them with experimental results for iterative optimization avoids the traditional approach of relying on experience and random trials, significantly reduces the number of invalid experiments, shortens the model establishment cycle, and improves the overall research efficiency. (3) By comparing the measured phenotypic data of modified mice with the predicted results of target phenotypes, the interaction map can be corrected in real time, thereby continuously optimizing the target combination, forming an adaptive improvement mechanism, and ensuring that the final model phenotype is highly consistent with the expectations. (4) Introducing ethical compliance conditions and result registration mechanisms in the process of generating construction instructions and registering modification receipts makes the entire model development process comply with regulatory requirements, and by registering and saving target target combinations, prediction results, and measured data, traceable research resources are formed, which are convenient for subsequent verification and application. Attached Figure Description

[0015] Figure 1 This is a flowchart of a method for constructing a type 2 diabetic mouse model that combines AI and multi-target knockout, provided in the first embodiment of this application. Detailed Implementation

[0016] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.

[0017] The first embodiment of this application provides a method for constructing a type 2 diabetic mouse model combining AI and multi-target knockout. Please refer to... Figure 1 This figure is a schematic diagram of the first embodiment of this application. The following is in conjunction with... Figure 1 The first embodiment of this application provides a detailed description of a method for constructing a type 2 diabetic mouse model that combines AI and multi-target knockout.

[0018] Step S101: Acquire and standardize mouse multi-omics data and phenotypic baseline data to form a standardized modeling dataset.

[0019] In step S101, acquiring and standardizing mouse multi-omics data and phenotypic baseline data is a fundamental step in the entire construction method. This step first requires clarifying the source and type of multi-omics data. Multi-omics data typically includes genomics data, transcriptomics data, proteomics data, and metabolomics data. When acquiring genomics data, whole-genome sequencing or targeted region capture sequencing can be used to obtain differential information on gene sequences from different mouse strains or experimental conditions. When acquiring transcriptomics data, high-throughput RNA sequencing technology can be used to measure gene expression levels in different tissues or under different conditions. When acquiring proteomics data, liquid chromatography-mass spectrometry (LC-MS / MS) is often used to quantify protein abundance in tissues or serum. When acquiring metabolomics data, nuclear magnetic resonance (NMR) or gas chromatography / liquid chromatography-mass spectrometry can be used to detect small molecule metabolites related to energy metabolism, glucose metabolism, and lipid metabolism. These data must be acquired following a unified experimental protocol and quality control standards to ensure comparability and consistency between data from different sources.

[0020] While collecting multi-omics data, it is also necessary to obtain phenotypic baseline data related to the mouse population. Phenotypic baseline data should cover key physiological indicators that characterize susceptibility to type 2 diabetes, such as body weight, blood glucose levels, insulin sensitivity, serum lipid levels, and energy metabolism rate. In addition, variables that may affect phenotypic performance, such as mouse housing conditions, age, sex, and environmental factors, should also be included. The process of obtaining phenotypic data must follow standardized measurement methods. For example, fasting blood glucose levels can be measured by collecting blood from the tail vein and using a standardized blood glucose meter, and insulin sensitivity can be obtained through a glucose tolerance test or an insulin tolerance test. Only by using standardized measurement methods and recording formats during these data collection processes can the scientific comparability of data be ensured in subsequent modeling.

[0021] After data acquisition, multi-omics data and phenotypic baseline data must be standardized to form a unified modeling dataset. The core of standardization is to eliminate data bias caused by differences in measurement platforms, experimental batches, or sample sources. For example, gene expression data needs to be normalized using transcript length and sequencing depth, employing methods such as TPM or FPKM to unify gene expression levels across different samples to the same dimension. Proteomics and metabolomics data require peak area normalization and internal control calibration to reduce the impact of instrument sensitivity fluctuations. Phenotypic baseline data requires Z-score or min-max normalization methods to map different indicators to a uniform 0-1 range, ensuring comparability of overall results during modeling without bias due to differences in numerical dimensions.

[0022] Furthermore, the standardization process requires handling missing and outlier values. For data with partially failed or missing measurements, multiple imputation or extrapolation based on similar samples can be used to complete the data. For outlier data that significantly deviates from a reasonable range, statistical tests combined with physiological significance should be used to remove or correct them to prevent erroneous values ​​from affecting the accuracy of the model. After data cleaning and standardization, multi-omics data and phenotypic data need to be integrated. Specifically, the genomic, transcriptomic, proteomic, and metabolomic data of the same mouse can be associated with its phenotypic data using unique sample identifiers to form a complete multidimensional description of an individual's characteristics. The resulting standardized modeling dataset not only contains multi-level information at the genetic background and molecular level but also incorporates actual phenotypic performance, providing comprehensive and accurate input for subsequent artificial intelligence prediction and interaction mapping.

[0023] During data integration, the stability and traceability of data storage and retrieval should also be considered. Standardized modeling datasets typically need to be stored in a structured database to ensure that each data point has clear annotation information, including data source, experimental conditions, processing methods, and timestamps, thereby achieving traceability management. This traceability not only helps ensure the compliance of experiments but also provides a solid foundation for verification and correction during subsequent model training. Ultimately, the standardized modeling dataset obtained through step S101 is the core input throughout the entire method; its quality and completeness directly determine the reliability of the artificial intelligence prediction results and the scientific validity and feasibility of subsequent mouse model construction.

[0024] Furthermore, the acquisition and standardization of mouse multi-omics data and phenotypic baseline data to form a standardized modeling dataset includes:

[0025] Collect genomic, transcriptomic, proteomic, and metabolomic data of mice, and simultaneously collect baseline phenotypic data of mice;

[0026] Data cleaning was performed on mouse multi-omics data and phenotypic baseline data, including batch effect correction, missing value imputation, and outlier removal.

[0027] Normalization was performed on the cleaned mouse multi-omics data and phenotypic baseline data to unify different types of data into a comparable dimensional range.

[0028] Based on individual mouse identifiers, normalized mouse multi-omics data and phenotypic baseline data are integrated to form a standardized modeling dataset.

[0029] First, during the data acquisition phase, whole-genome sequencing data, transcriptome sequencing data, quantitative proteome data, and metabolome analysis data should be obtained for each mouse. Whole-genome sequencing data can be obtained using second- or third-generation sequencing platforms. Transcriptome sequencing requires RNA samples from pancreatic islets, liver, and adipose tissue at the tissue level. Proteomics data can be obtained using mass spectrometry to detect protein abundance, while metabolomics data should be obtained using chromatography-mass spectrometry to detect metabolite levels in serum and liver tissue. Simultaneously with the acquisition of omics data, baseline phenotypic data should be acquired. This baseline phenotypic data includes fasting blood glucose levels, insulin resistance test results, body weight, serum lipid profile, and liver histological indicators to ensure a one-to-one correspondence between multi-omics data and phenotypic data at the individual level.

[0030] After collecting the raw data, rigorous data cleaning should be performed on the mouse multi-omics data and phenotypic baseline data. Batch effect correction is an important step in handling systematic biases caused by different sequencing batches and experimental batches, and can be achieved using the ComBat algorithm or adjustment methods based on mixed linear models. Missing value imputation can be achieved through nearest neighbor imputation or multiple imputation methods based on data distribution to avoid the impact of missing data on downstream analysis. Outlier removal requires identifying data points that significantly deviate from the normal range based on statistical distribution or principal component analysis to ensure the reliability of the data for the next step.

[0031] After data cleaning, multi-omics data and phenotypic baseline data need to be normalized to eliminate differences between different measurement methods and units of measurement. Genomic and transcriptomic data can be normalized to uniform expression levels using per million sequences or the TPM method; proteomic data can be normalized to ensure comparability using total normalization or quantile normalization; and metabolomic data can be adjusted to a uniform scale using logarithmic transformation and Z-score standardization. Phenotypic baseline data should also be standardized, for example, by converting indicators such as blood glucose, insulin levels, and blood lipid concentrations into standard deviations, so that different types of phenotypic variables can be directly compared in the same numerical space.

[0032] After completing the above normalization process, data from different sources need to be integrated based on individual mouse identifiers to ensure that the baseline data of genomics, transcriptomics, proteomics, metabolomics, and phenotype for each mouse can form a complete multidimensional data vector. Specifically, a unique individual identifier should be established for each mouse, and all normalized data should be integrated into the same database or data matrix according to this identifier, thereby forming a standardized modeling dataset. This dataset contains both high-dimensional molecular features at multiple omics levels and numerical indicators directly related to disease phenotypes, providing complete and consistent input for the subsequent construction of interaction maps.

[0033] Step S102: Based on the standardized modeling dataset, construct an interaction map including pancreatic β-cell function, insulin signaling, and hepatic lipid metabolism pathways, and use artificial intelligence to predict and generate candidate target combinations and candidate phenotypic response results.

[0034] In step S102, constructing an interaction map encompassing pancreatic β-cell function, insulin signaling, and hepatic lipid metabolism pathways based on a standardized modeling dataset, and utilizing artificial intelligence to predict and generate candidate target combinations and candidate phenotypic responses, is a crucial step in achieving the intelligence and systematization of the method of this invention. In this process, the standardized modeling dataset obtained in step S101 first needs to be input into a structured data processing framework. This framework should be able to organically integrate different types of data according to four dimensions: genes, transcription, proteins, and metabolites. By establishing a multi-layered biological network, the impact of gene sequence variations on transcriptional levels, the regulation of protein expression by changes in transcription products, and how differences in protein abundance further affect the generation and consumption of metabolites can be demonstrated within the same system. In this way, the entire interaction map can comprehensively cover the causal relationships from genetic information to phenotypic expression, rather than simply a single-dimensional list of data.

[0035] In constructing an interaction map, it is crucial to emphasize the integration of three pathways: pancreatic β-cell function, insulin signaling, and hepatic lipid metabolism. These three pathways are central to the pathogenesis of type 2 diabetes. Pancreatic β-cell function involves the regulation of insulin synthesis and secretion, typically including insulin gene transcription levels, expression of insulin processing-related enzymes, calcium ion signaling, and the state of secretory granule transport. The insulin signaling pathway encompasses the interactions of a series of downstream signaling molecules, including the insulin receptor, insulin receptor substrate, PI3K / AKT pathway, and MAPK pathway, directly affecting the uptake and utilization of glucose by peripheral tissues. The hepatic lipid metabolism pathway mainly involves fatty acid synthesis, β-oxidation, triglyceride synthesis, and VLDL secretion. Disorders of these metabolic activities can lead to phenotypes such as insulin resistance and fatty liver. By establishing the nodes and edges of these three pathways and their intersections in the map, the functional relationships and signal transduction pathways between different biomolecules can be systematically reflected.

[0036] After obtaining a preliminary interaction map, mathematical modeling and topological analysis are required. Graph theory methods can be used to calculate node importance, such as degree centrality, betweenness centrality, and eigenvector centrality, to identify genes or proteins that play key roles in the network. For feedback loops and redundant branches in the pathway, dynamic modeling methods are also needed to simulate system behavior under different perturbations to reveal key node combinations that may cause pathological changes. The results of this analysis will generate a series of candidate potential regulatory points, and joint interventions of these regulatory points are expected to more closely approximate the true phenotype of type 2 diabetes.

[0037] Building upon this foundation, artificial intelligence (AI) algorithms are introduced into the prediction process. AI systems can employ deep learning networks, graph neural networks, or Bayesian inference models to train and infer on standardized datasets and interaction maps. Specifically, by learning from existing experimental data and data features from known disease models, AI can predict the phenotype of unknown target combinations. The system outputs a series of candidate target combinations and provides a corresponding phenotypic response for each combination. These results typically manifest as predicted values ​​for specific phenotypic indicators such as changes in blood glucose levels, changes in insulin sensitivity, and lipid metabolism status. AI here goes beyond simply screening high-frequency nodes; it comprehensively considers the complex nonlinear relationships between multi-omics data, pathway coupling effects, and phenotypic results, generating a more comprehensive and reliable candidate target scheme than human experience-based judgment.

[0038] To ensure the reliability of the prediction results, the AI ​​model should employ cross-validation and independent test datasets during training, and output confidence intervals or credibility scores during the prediction phase, allowing researchers to intuitively assess the reliability of each candidate target combination. Furthermore, the candidate phenotypic response results should not only include single numerical indicators but also provide joint predictions of multiple related indicators, such as the overall trend of the glucose tolerance curve, the dynamic response of signaling pathways under insulin stimulation, and the likelihood of hepatic lipid accumulation. This multi-dimensional prediction ensures that the generated candidate target combinations are more likely to exhibit a state consistent with the type 2 diabetes complex phenotype in subsequent experiments.

[0039] In summary, step S102, by constructing an interaction map encompassing pancreatic β-cell function, insulin signaling, and hepatic lipid metabolism pathways, and utilizing artificial intelligence to perform predictive analysis on a standardized modeling dataset, can output clear combinations of candidate targets and candidate phenotypic responses. This process not only provides a systematic scientific basis for subsequent target screening but also ensures the reproducibility and traceability of the model construction process, thus laying a solid technical foundation for the accurate construction of animal models of complex diseases.

[0040] Furthermore, the construction of an interaction map encompassing pancreatic β-cell function, insulin signaling, and hepatic lipid metabolism pathways based on the standardized modeling dataset, and the generation of candidate target combinations and candidate phenotypic response results using artificial intelligence prediction, includes:

[0041] The standardized modeling dataset was used to extract features from genes, proteins and metabolites related to pancreatic β-cell function, insulin signaling and liver lipid metabolism pathways, and to obtain multidimensional features corresponding to pathway nodes.

[0042] The multidimensional features are connected in a network based on intermolecular regulatory relationships and pathway mechanisms to form an interaction map that includes pancreatic β-cell function, insulin signaling and hepatic lipid metabolism pathways;

[0043] The interaction map is input into an artificial intelligence model to infer and simulate the effects of perturbations at different nodes and joint perturbations at multiple nodes, generating candidate target combinations and their corresponding dynamic response curves.

[0044] Based on the dynamic response curve, the predicted results of candidate target combinations in terms of glucose homeostasis, insulin sensitivity, and lipid metabolism are calculated, forming candidate target combinations and candidate phenotypic response results.

[0045] Specifically, genes, proteins, and metabolites directly related to the aforementioned pathways should be identified from standardized modeling datasets. These include ion channel genes regulating insulin secretion, key signaling proteins in the insulin receptor signaling chain, and metabolites involved in hepatic triglyceride metabolism. For each molecular node, multi-dimensional features such as expression levels, activity changes, and interaction strengths should be extracted to accurately represent its functional performance under different states. These features must employ a unified numerical encoding format to ensure direct comparison and combination with features from other nodes during subsequent network construction.

[0046] After obtaining multidimensional features, network connections need to be established based on known intermolecular regulatory relationships and pathway mechanisms. For example, directed edge relationships are established between glucose-sensing genes in pancreatic β-cells and insulin secretion-related proteins; functional connections are established between insulin signaling receptors and downstream PI3K / AKT pathway nodes; and bidirectional relationships are established between key enzymes in liver lipid metabolism and their metabolites. Networks constructed in this way can reflect dynamic cross-level interactions and form a complete interaction map. This map not only encompasses first-order interactions between molecular nodes but also covers multiple pathway intersections and feedback loops, ensuring the model possesses high biological realism.

[0047] The constructed interaction map needs to be fed into an artificial intelligence (AI) model for systematic reasoning and simulation. The AI ​​model can employ graph neural networks, dynamic Bayesian networks, or causal inference models to calculate perturbations on single nodes and multiple nodes jointly. During the simulation, knockout or overexpression states of certain nodes can be artificially set, and their propagation effects on downstream nodes in the network can be observed. The AI ​​model generates candidate target combinations based on large-scale iterative calculations and provides dynamic response curves for each combination under different perturbation conditions. These dynamic response curves can display trends over time, such as the rate of blood glucose homeostasis recovery or the dynamic changes in lipid accumulation levels after simulating the knockout of a certain group of targets, thus providing quantitative evidence for the effectiveness of candidate combinations.

[0048] After obtaining the dynamic response curves, further calculations and quantifications of the effects of candidate target combinations are needed. Evaluation indicators should be established for three aspects: glucose homeostasis, insulin sensitivity, and lipid metabolism, converting key parameters in the dynamic response curves into comparable numerical results. For example, the reduction in blood glucose fluctuations after knocking out a particular combination can be calculated, or the efficiency of insulin signaling enhancement and the degree of reduction in hepatic lipid deposition can be assessed. These indicators allow for a direct comparison of the predictive effects of different candidate target combinations, ultimately forming candidate target combinations and candidate phenotypic response results. These results include multiple potentially feasible gene intervention combinations screened through artificial intelligence reasoning and their predicted metabolic responses, providing a solid scientific basis for subsequent screening and experimental validation.

[0049] Thus, the entire process begins with feature extraction from multi-omics data, proceeds through the construction of pathway-level networks, enters AI-driven perturbation simulation, and finally generates candidate target combinations and candidate phenotypic response results.

[0050] In a specific implementation, the artificial intelligence model can employ a composite architecture combining graph neural networks and causal inference mechanisms to fully utilize the multi-level information of the standardized modeling dataset and the interaction graph. The overall input to the artificial intelligence model is a set of multi-dimensional feature nodes extracted from the standardized modeling dataset and an interaction graph constructed based on pathway mechanisms. The overall output is the candidate target combination and its corresponding candidate phenotypic response results. The model can be divided into four parts: a feature embedding layer, a dynamic graph convolutional layer, a causal inference layer, and a phenotypic prediction layer.

[0051] In the feature embedding layer, the input is the multidimensional feature data of each node, including gene expression levels, protein abundance, and metabolite concentrations. The embedding network transforms these multidimensional features into a unified low-dimensional vector representation, enabling each node to participate in subsequent calculations in the same feature space. The output is the feature embedding vector of each node.

[0052] In the dynamic graph convolutional layer, the input consists of feature embedding vectors and node connectivity information from the interaction graph. This layer simulates the signal propagation process between molecules through graph convolution operations and updates the node state representation in each convolution iteration, thereby reflecting the diffusion effect of perturbations in the pathway. Its output is the node state matrix after multiple rounds of propagation, which contains the temporal response features of the nodes under dynamic perturbations.

[0053] In the causal inference layer, the input consists of a node state matrix and predefined perturbation conditions. This layer uses causal structural equations to infer single-node knockout and multi-node joint knockout, eliminating spurious connections derived solely from correlations and ensuring the causal interpretability of the output. Its output comprises candidate target combinations and the causal effect weights of these combinations under perturbation conditions, reflecting the direct driving force of specific combinations on phenotypic changes.

[0054] In the phenotypic prediction layer, the input is the candidate target combination and its causal effect weights output from the causal inference layer. Combined with the phenotypic baseline data in the standardized modeling dataset, dynamic response curves are generated using regression prediction and time-series modeling methods. The output of this layer is the candidate phenotypic response results, specifically manifested as a blood glucose level regulation curve, an insulin sensitivity change curve, and a liver lipid metabolism curve.

[0055] In this way, the artificial intelligence model can not only reveal the interactions between targets at the molecular level, but also provide quantitative predictive responses at the phenotypic level. Its innovation lies in the first-ever organic combination of graph neural networks and causal reasoning for predicting multi-target combined intervention strategies. This approach retains the global characteristics of complex network modeling while introducing the interpretability of causal effects, making the generated candidate target combinations and candidate phenotypic response results highly reliable and traceable. Those skilled in the art can directly construct and train this artificial intelligence model based on the explicit descriptions of the above inputs, outputs, and functional modules at each layer, thereby reproducing the prediction process in practice.

[0056] In a specific implementation, the artificial intelligence model can employ a composite architecture combining graph neural networks and causal inference mechanisms to fully utilize the multi-level information of the standardized modeling dataset and the interaction graph. The overall input to the artificial intelligence model is a set of multi-dimensional feature nodes extracted from the standardized modeling dataset and an interaction graph constructed based on pathway mechanisms. The overall output is the candidate target combination and its corresponding candidate phenotypic response results. The model can be divided into four parts: a feature embedding layer, a dynamic graph convolutional layer, a causal inference layer, and a phenotypic prediction layer.

[0057] In the feature embedding layer, the input is the multidimensional feature data of each node, including gene expression levels, protein abundance, and metabolite concentrations. The embedding network transforms these multidimensional features into a unified low-dimensional vector representation, enabling each node to participate in subsequent calculations in the same feature space. The output is the feature embedding vector of each node.

[0058] In the dynamic graph convolutional layer, the input consists of feature embedding vectors and node connectivity information from the interaction graph. This layer simulates the signal propagation process between molecules through graph convolution operations and updates the node state representation in each convolution iteration, thereby reflecting the diffusion effect of perturbations in the pathway. Its output is the node state matrix after multiple rounds of propagation, which contains the temporal response features of the nodes under dynamic perturbations.

[0059] In the causal inference layer, the input consists of a node state matrix and predefined perturbation conditions. This layer uses causal structural equations to infer single-node knockout and multi-node joint knockout, eliminating spurious connections derived solely from correlations and ensuring the causal interpretability of the output. Its output comprises candidate target combinations and the causal effect weights of these combinations under perturbation conditions, reflecting the direct driving force of specific combinations on phenotypic changes.

[0060] In the phenotypic prediction layer, the input is the candidate target combination and its causal effect weights output from the causal inference layer. Combined with the phenotypic baseline data in the standardized modeling dataset, dynamic response curves are generated using regression prediction and time-series modeling methods. The output of this layer is the candidate phenotypic response results, specifically manifested as a blood glucose level regulation curve, an insulin sensitivity change curve, and a liver lipid metabolism curve.

[0061] In this way, the artificial intelligence model can not only reveal the interactions between targets at the molecular level, but also provide quantitative predictive response results at the phenotypic level. Its innovation lies in the first organic combination of graph neural networks and causal reasoning for the prediction of multi-target combination intervention strategies. It retains the global characteristics of complex network modeling while introducing the interpretability of causal effects, making the generated candidate target combinations and candidate phenotypic response results highly reliable and traceable.

[0062] The model takes a standardized modeling dataset and an interaction graph as unified input and outputs candidate target combinations and candidate phenotypic response results. Internally, the model consists of a feature embedding layer, a dynamic graph convolutional layer, a causal inference layer, and a phenotypic prediction layer. Targeted loss and data ledger are introduced during the training and calibration phases to ensure that the results are traceable, reproducible, and can be directly used for the generation of construction instructions and subsequent registration.

[0063] At the feature embedding layer, the multi-omics features of each pathway node in the standardized modeling dataset need to be mapped to the same representation space. Each node should contain at least three types of numerical vectors: gene expression level, protein abundance, and metabolite concentration. All inputs have been dimensionally aligned in the standardized modeling dataset. To avoid the superposition of local noise from different omics, it is recommended to use gated linear projection during vector mapping to automatically suppress dimensions weakly correlated with pancreatic β-cell function, insulin signaling, and hepatic lipid metabolism pathways, making the output a node initial representation more consistent with the pathway mechanism. To reduce the loss of functional roles due to node homogenization, type priors should be introduced for the three types of nodes: "gene," "protein," and "metabolite." Trainable type biases are directly superimposed on the initial node representation, enabling subsequent propagation on the interaction map to explicitly distinguish molecular levels. The output is a node representation matrix with a consistent shape, used for subsequent time-spread propagation.

[0064] In the dynamic graph convolutional layer, perturbations need to propagate temporally across the interaction graph, and intervention awareness needs to be implemented for nodes that have been "knocked out" or "jointly knocked out." The interaction graph can be derived from an authoritative pathway library or manually reviewed pathway files, serving as fixed or weighted adjacency inputs. At each time step, adjacencies are first stabilized and normalized, and then trainable gating weights are applied to edges connected to the intervened node, ensuring that incoming and outgoing edges from the intervened node are treated differently during propagation, thus explicitly injecting the effects of "joint knockout" into the graph propagation operator. At each time step, message aggregation, nonlinear transformation, and residual normalization are performed to obtain new node states; the state of the intervened node is pulled to a low-energy or zero state according to the intervention intensity before each propagation step, simulating the direct effects of "do-intervention." The number of time steps should be consistent with the pathway biology timescale, typically ranging from several to more than ten steps, and the output is a sequence of node states, characterizing the diffusion trajectory of the perturbation from local to global. The essential difference between this layer and traditional static graph convolution lies in the introduction of intervention-aware gating and temporal unfolding mechanisms, thereby transforming "joint knockout" into learnable diffusion dynamics.

[0065] At the causal inference layer, the node states of the time series need to be transformed into interpretable causal effects, thereby generating candidate target combinations and causal strength weights. It is recommended to adopt a counterfactual advancement framework centered on structural equation modeling, establishing causal evolution relationships between node states at adjacent times through learnable mappings. A baseline trajectory is obtained before intervention. When a "do-intervention" is applied to a set of nodes, the states of these nodes are replaced and the process is rolled forward to obtain the counterfactual trajectory. To map the node time series to phenotypic latent variables, stable aggregation readouts can be performed in the node and time dimensions to obtain three-dimensional phenotypic vectors corresponding to blood glucose levels, insulin sensitivity, and lipid metabolism. By comparing the distances between the baseline trajectory and the counterfactual trajectory in these three-dimensional vectors, the phenotypic closeness of the intervention combination can be obtained. The causal effect weight is defined as the unintervention distance minus the intervention distance; a larger weight indicates a closer alignment with the target phenotypic direction. To generate candidate target combinations from a large number of nodes, a small-scale greedy bundle search strategy can be adopted, gradually expanding the combination size, and sorting and pruning based on causal effect weights. This strategy is naturally adapted to the combination space of multiple targets and can limit the candidate pool to the vicinity of key nodes in the pathway to improve search efficiency. The output of this layer is the counterfactual trajectory or endpoint statistics required for the candidate target combinations and the corresponding candidate phenotypic response results, directly serving the screening step.

[0066] In the phenotypic prediction layer, the counterfactual trajectory needs to be transformed into directly comparable and registerable candidate phenotypic response results, including blood glucose level regulation curves, insulin sensitivity change curves, and liver lipid metabolism curves. Attention pooling can be applied to the node states at each time point, with weights automatically learned based on the node's contribution to the target phenotype. The resulting global state sequence is then input into a stable temporal modeler (which can be a gated recurrent unit or other equivalent structure), finally outputting three curves in a multi-head regression format. At the zeros and scales of the curves, the phenotypic baseline should be introduced as a bias term to align the output with the individual baseline of the modified mice, facilitating item-by-item comparison with the target phenotypic prediction results. This layer should also provide robust transformations from curves to endpoint indicators to support scenarios requiring single-value threshold judgments.

[0067] In terms of training and calibration, the model needs to be jointly constrained by multiple task objectives. First, self-supervised or reconstructed objectives should be constructed using interaction maps or cross-omics consistency to stabilize the representation quality of feature embeddings and graph propagation. Second, phenotypic curves or endpoint indicators should be fitted using historical or small-scale gold standard data to align the phenotypic prediction layer and the causal inference layer under supervised signaling. Third, intervention direction consistency regularization should be introduced, using a small number of known single-target or dual-target intervention controls as exogenous constraints to ensure that counterfactual advancements align with existing biological common sense in terms of direction. Data partitioning should maintain cross-validation between strains and batches, and early cessation and weight decay can use generally accepted and stable configurations. Those skilled in the art can then complete end-to-end training and evaluate the model on the validation set using three types of indicators: curve correlation, endpoint error, and causal direction consistency.

[0068] In terms of output management and compliance, candidate target combinations, causal effect weights, and candidate phenotypic response results need to be solidified in a unified data structure, and associated with the version number of the standardized modeling dataset, the version number of the interaction graph, and the training checkpoint identifier to ensure complete traceability. When connecting with the screening step, candidate target combinations and candidate phenotypic response results are directly used as input, and convergence screening is performed in conjunction with ethical and compliance conditions. When connecting with subsequent construction instructions, the top-ranked candidate target combinations are transformed into clear modification targets and expected phenotypic descriptions to support the generation of construction instructions and the verification of modification receipts. When connecting with comparison and iteration, the measured phenotypic data is compared with the candidate phenotypic response results, and the phenotypic deviation metric is calculated and then fed back into the weights and gating of the interaction graph to form a closed-loop optimization.

[0069] Step S103: Based on the candidate target combination and the candidate phenotype response results, and in conjunction with ethical compliance conditions, select the target target combination and target phenotype prediction results.

[0070] In step S103, further screening is required based on the candidate target combinations and candidate phenotypic response results generated by artificial intelligence to ensure that the target schemes that finally enter the experimental stage are both scientifically sound and meet ethical and compliance requirements. First, a systematic evaluation of the candidate target combinations output by artificial intelligence should be conducted. Candidate combinations typically contain multiple genes or regulatory factors, which may involve key regulatory sites of pancreatic β-cell function, receptors or downstream molecules in the insulin signaling pathway, and key enzymes in liver lipid metabolism. The first step in the evaluation is to examine whether these targets have a clear association with known pathogenic mechanisms of diabetes. This process requires referencing disease gene annotation information in publicly available databases, existing animal experimental literature, and internationally recognized lists of genes related to metabolic diseases. By comparing these knowledge bases, it can be preliminarily confirmed whether the candidate target combinations could theoretically induce a complex phenotype associated with type 2 diabetes.

[0071] After confirming the scientific rationale behind candidate target combinations, it is necessary to verify the consistency and interpretability of their phenotypic response results. The phenotypic response results predicted by artificial intelligence should not be limited to a single indicator but should encompass multiple dimensions such as blood glucose levels, insulin tolerance, lipid accumulation, and energy metabolism efficiency. This step requires cross-comparison of the phenotypic changes predicted by artificial intelligence using simulated experimental data or historical experimental results. For example, if a combination predicts an increase in fasting blood glucose, accompanied by a weakened insulin response and elevated liver triglyceride levels, then this phenotypic pattern should be consistent with previously reported diabetic mouse phenotypes to demonstrate the rationality of the prediction results. If the prediction results generated by certain candidate combinations differ significantly from the type 2 diabetes phenotype or contradict existing biological laws, they should be excluded.

[0072] After confirming scientific rationality and phenotypic consistency, this step further introduces ethical and compliance constraints. Multi-target gene knockout experiments often involve the simultaneous modification of multiple genes in an animal, which may severely impact the survival of mice. Therefore, it is essential to conduct a compliance review of candidate target combinations, in accordance with the guidelines of the laboratory animal ethics committee and the regulations of the region. This process includes assessing whether the modified mice can maintain a basic quality of life, whether there is an unacceptably high risk of mortality, and whether it involves severe non-specific damage to non-target systems. Target combinations that predict the possibility of excessive pathological damage or significant deviation from controllable limits should be eliminated during the screening stage. Simultaneously, it is necessary to ensure that data recording and experimental design comply with the regulatory requirements for the filing and approval of transgenic animal experiments to avoid legal or ethical obstacles in subsequent implementation.

[0073] After completing scientific evaluation, phenotypic consistency verification, and ethical compliance review, the candidate target combinations can be progressively narrowed down to select the most representative and research-aligned combination. The corresponding target phenotypic prediction results will then serve as the standard input for subsequent experimental execution. This target combination should theoretically reproduce the main clinical features of type 2 diabetes and exhibit multidimensional phenotypic abnormalities consistent with clinicopathology in the prediction results. Furthermore, it should meet operability requirements, meaning it should enable stable and reproducible modification using existing gene-editing techniques, without being impractical due to excessive technical difficulty.

[0074] Through this screening process, step S103 not only ensures that the target combinations selected from the AI ​​prediction results are scientifically sound and experimentally feasible, but also, by incorporating ethical compliance requirements, allows the entire research methodology to be conducted in accordance with legal and social norms. The result of this process is that the originally numerous and diverse candidate target combinations are converged into a few clearly defined target combinations, providing solid scientific and compliance support for their corresponding target phenotype prediction results. This lays a reliable foundation for subsequent generation of construction instructions and actual genetic modification operations.

[0075] Furthermore, the step of selecting target target combinations and target phenotype prediction results based on the candidate target combination and the candidate phenotype response results, and in conjunction with ethical compliance conditions, includes:

[0076] The pathogenic mechanism correlation analysis was performed on the candidate target combinations. Based on the interaction between pancreatic β-cell function, insulin signaling and hepatic lipid metabolism pathways, candidate target combinations that can cause multidimensional metabolic abnormalities were screened.

[0077] The candidate target combinations obtained through correlation analysis are compared with the candidate phenotypic response results. Candidate target combinations whose phenotypic response results do not match the core phenotype of type 2 diabetes are eliminated, and candidate target combinations with multi-phenotypic consistency and their corresponding candidate phenotypic response results are obtained.

[0078] Candidate target combinations with multi-phenotypic consistency and their corresponding candidate phenotypic response results are compared with ethical compliance conditions to exclude candidate target combinations that may cause serious irreversible damage or do not meet the ethical requirements of animal experiments, thereby obtaining compliant candidate target combinations and their corresponding candidate phenotypic response results.

[0079] Priority evaluation is performed based on the compliant candidate target combination and its corresponding candidate phenotypic response results. The ranking is determined according to the comprehensive weight of the predicted blood glucose regulation ability, the degree of change in insulin sensitivity, and the effect of liver lipid deposition, and the target target combination and the target phenotypic prediction results are selected.

[0080] First, a pathogenic mechanism correlation analysis is required for candidate target combinations. This process necessitates examining each candidate target combination individually based on the interactions between pancreatic β-cell function, insulin signaling, and hepatic lipid metabolism pathways. By locating the node positions of candidate targets in the aforementioned interaction map and analyzing their pathways within the signal transduction network, it can be determined whether these targets may synergistically cause multidimensional metabolic abnormalities such as glucose metabolism, insulin secretion, or lipid deposition. Combinations that produce isolated effects on only a single pathway are eliminated, while candidate target combinations that simultaneously demonstrate significant effects across multiple pathways are retained as the basis for further screening.

[0081] After obtaining candidate target combinations that pass correlation analysis, these combinations need to be compared with the candidate phenotypic response results for consistency. This comparison requires verifying whether the candidate phenotypic response results are consistent with the core phenotypes of type 2 diabetes, including persistent hyperglycemia, insulin resistance, and abnormal hepatic lipid metabolism. If the phenotypic response results of a candidate target combination deviate from the core characteristics, such as showing no significant increase in blood glucose levels or no change in insulin sensitivity, the combination will be excluded. In this way, candidate target combinations with multi-phenotypic consistency and their corresponding candidate phenotypic response results can be obtained, ensuring that the selected combinations are not only reasonable at the molecular level but also able to reproduce the key features of the disease at the phenotypic level.

[0082] For candidate target combinations demonstrating multiphenotypic consistency, ethical compliance must also be considered. This step requires a risk assessment of the potential impact of candidate target combinations, excluding mouse model construction schemes that could lead to severe irreversible damage. For example, if the phenotypic response results of certain target combinations show that they cause rapid death in animals in the early stages of the experiment, or induce irreversible systemic damage far exceeding the needs of disease research, then such combinations do not comply with animal experimental ethical norms and must be excluded. Only candidate target combinations that are within ethical compliance, stably reproduce the target disease phenotype, and do not exceed ethical standards can proceed to the final selection stage.

[0083] After compliance screening, a priority evaluation is required for compliant candidate target combinations and their corresponding candidate phenotypic response results. This evaluation should employ a comprehensive weighting method, using predicted glycemic regulation capacity, changes in insulin sensitivity, and hepatic lipid deposition effects as core indicators. By weighting and ranking these three types of indicators, the overall value of each candidate target combination can be quantified. Combinations that demonstrate high consistency and significant effects in glycemic regulation, insulin sensitivity, and lipid metabolism abnormalities will be ranked higher. The final selected combinations will serve as the target target combinations, and their corresponding phenotypic prediction results will be the target phenotypic prediction results. This approach ensures that the target target combinations and target phenotypic prediction results are consistent not only at the mechanistic and phenotypic levels but also undergo ethical and priority-based filtering, guaranteeing the scientific validity, feasibility, and compliance of the final results. This provides a clear and reliable basis for subsequent genetic modification and model construction.

[0084] Step S104: Generate a construction instruction based on the target combination and the target phenotype prediction results, and submit the construction instruction to the controlled facility to perform genetic modification operations to obtain modified mice and modification receipts.

[0085] In step S104, the selected target combination and target phenotype prediction results need to be transformed into construction instructions that can be actually executed in the experimental facility, thereby guiding the implementation of genetic modification operations. The generation of construction instructions must be a systematic, precise, and standardized process to ensure that different researchers or controlled facilities strictly adhere to the same specifications during execution, thus obtaining consistent and reproducible results. First, at the starting point of construction instruction generation, each gene locus in the target combination needs to be clearly defined. This process typically includes the full name of the gene, standard gene symbol, gene location in the genome, reference sequence number, and the specific type of target mutation or knockout, such as whether it is a whole-genome knockout, exon deletion, point mutation, or conditional knockout. This information must be recorded in a precise format to avoid any ambiguity or vague interpretation.

[0086] After defining the target, it's necessary to further clarify the direction and expected effects of the modification by combining the target phenotype prediction results. Target phenotype prediction results typically involve indicators such as trends in blood glucose levels, decreased insulin sensitivity, and the degree of liver lipid accumulation. Therefore, the construction instructions must reflect the correspondence between these predicted phenotypes and the target. In this way, the experimental team can understand how the modification of each target serves the final overall disease model phenotype, rather than simply performing isolated gene manipulations. This ensures that the experimental team has a comprehensive grasp of the design approach and also facilitates retrospective analysis and interpretation during the subsequent phenotypic validation phase.

[0087] In the specific process of generating construction instructions, a standardized set of operational descriptions needs to be developed. These descriptions do not directly provide experimental details, but rather present key operational points and a logical sequence. For example, they specify the modification strategy required for each target (such as a double-cut deletion strategy using the CRISPR / Cas9 system, the Cre-LoxP system setup required for conditional knockout, or reducing specific gene expression through RNA interference), and clearly label the correspondence between these strategies and target targets within the instructions. Simultaneously, the instructions must include instructions on the modification order, because in multi-target modification, the knockout of different genes may have sequential dependencies. For example, certain basal metabolic genes must be modified first; otherwise, subsequent target modification will not proceed correctly. In such cases, the construction instructions must explicitly specify the execution order to avoid conflicts or adverse results.

[0088] In addition to operational requirements, the construction instructions should also include explanations of key validation points. Since every genetic modification requires molecular-level testing to confirm its success, the instructions must specify necessary validation steps. For example, whether PCR amplification is used to detect target fragment deletion, whether sequencing verifies that the mutation occurs in the correct location, and whether protein expression analysis confirms functional loss. These validation points do not provide specific experimental protocols but are included in the construction instructions as mandatory verification steps to ensure that the modification operation conforms to the intended design.

[0089] Once the construct instructions are generated, they must be submitted to a legally qualified, controlled facility for actual genetic modification. Controlled facilities typically refer to experimental centers or biotechnology companies that meet compliance requirements for animal testing and transgenic research. These facilities are capable of performing multi-gene knockout or editing under legal and ethical conditions. The submission process should be completed in the form of a digital document containing the entire construct instructions, along with data traceability information, including the target combination upon which the instructions are based, the predicted target phenotypes, and screening records from previous steps. This traceability information ensures that the controlled facility understands the context of the instructions during execution, preventing experimental bias due to missing information.

[0090] After the genetic modification operation is performed in a controlled facility, a formal confirmation of the modification results will be returned. This confirmation typically includes information confirming the modification process, such as the batch number of the modified mice, the specific target modifications completed, whether any targets were unsuccessful, and preliminary survival and health status. To ensure data traceability, the confirmation should also include a correspondence with the construction instructions, such as whether the modification of each target was successful and whether the results of the validation phase met the requirements. This confirmation will serve as the core input for subsequent steps of phenotypic validation and data comparison, ensuring the integrity and reliability of the entire closed-loop process.

[0091] Therefore, step S104 is not only an intermediate step in translating theoretical predictions into experimental practice, but also a crucial step in ensuring that the multi-target knockout process is executable, traceable, and verifiable. By standardizing the generation of instructions, ensuring the standardized execution of controlled facilities, and returning complete modification feedback, a solid foundation can be provided for subsequent phenotypic data collection and bias measurement, ensuring that the entire method for constructing a type 2 diabetes composite mouse model is scientifically sound and feasible.

[0092] Furthermore, the step of generating a construction instruction based on the target target combination and the target phenotype prediction results, and submitting the construction instruction to a controlled facility to perform genetic modification operations, to obtain modified mice and modification receipts, includes:

[0093] The target combination is transformed into a specific genetic modification scheme, and the modification method, modification location and required editing tool type for each target are determined to form a target modification specification.

[0094] The target modification specifications are compared with the target phenotype prediction results to generate a modification priority sequence, and the modification order and corresponding phenotype expectations are marked in the construction instructions.

[0095] The execution conditions for genetic modification are integrated into the construction instructions, including mouse strain selection, controlled facility operation requirements, and necessary verification steps, to form a complete construction instruction.

[0096] The construction instructions are submitted to the controlled facility to perform genetic modification operations, and the modified mouse information and modification receipt returned by the controlled facility are received and recorded.

[0097] In this implementation, the process of generating construction instructions based on the target target combination and the target phenotype prediction results, and submitting these instructions to a controlled facility for genetic modification to obtain modified mice and modification receipts, requires starting with the specific genetic modification design of the target points and gradually transforming it into a complete operational plan that can be directly executed in the facility. First, the target target combination needs to be transformed into an operable genetic modification plan one by one. For each target target, the modification method must be clearly defined, such as knockout, knock-in, point mutation, or conditional regulation, while precisely determining its modification location in the mouse genome. This location is usually based on a reference genome sequence and annotation database, combined with the specific molecular identifier of the target in the interaction map. Simultaneously, the type of gene editing tool required also needs to be determined; for example, when using the CRISPR-Cas9 system, the sequence of its guide RNA must be specified, or in other cases, tools such as TALEN or ZFN may be selected. Through this process, a target modification specification for the target target combination is formed, which clearly describes the modification requirements and implementation details for each point.

[0098] After generating the target modification specifications, they need to be compared with the predicted target phenotypes to ensure that the genetic modification plan aligns with the predicted phenotypic direction. This comparison identifies the contribution of different targets to the final phenotypic results and generates a modification priority sequence accordingly. For example, if some targets show significant regulatory effects on blood glucose levels in the predicted results, while others primarily affect lipid deposition, the priority modification order can be determined by weighted evaluation of different phenotypic indicators. During this process, the modification order of each target needs to be clearly indicated in the construction instructions, along with the corresponding phenotypic expectations, so that subsequent validation can compare the actual phenotype with the predicted results.

[0099] Furthermore, the implementation conditions for genetic modification must be integrated into the construction instructions. This includes the selection of experimental mouse strains, such as determining the background of commonly used models like C57BL / 6J or db / db, to ensure that the modified mice are genetically suitable for research on the type 2 diabetes phenotype. The operational requirements for controlled facilities also need to be specified, including the biosafety level of the gene editing laboratory, specific conditions for embryo microinjection or in vitro fertilization, and standards for animal husbandry and phenotypic monitoring. Simultaneously, the construction instructions should clearly define necessary validation steps, such as genotype sequencing confirmation, transcriptional level detection, and preliminary phenotypic testing, to ensure the accuracy and controllability of the modification process.

[0100] After completing the above steps, the construction instructions can be submitted to the controlled facility as a complete operational document. After the facility performs the genetic modification operation, it will return information about the modified mice and a modification receipt. The modified mouse information includes its genetic background, successfully modified sites, and the number of surviving mice, while the modification receipt is a formal record of the execution of the submitted construction instructions. These results need to be systematically recorded as an important basis for subsequent phenotypic data acquisition and comparison.

[0101] Step S105: Based on the modification receipt, obtain the measured phenotypic data of the modified mice, and compare the measured phenotypic data with the target phenotypic prediction results to generate phenotypic deviation measurement and phenotypic determination results.

[0102] In step S105, comprehensive phenotypic data needs to be collected from the genetically modified mice in the controlled facility. A data correspondence is established based on the modification information provided in the modification receipt to ensure that all experimental results can be accurately compared with the target combination and the predicted target phenotype. First, a tracking file must be established for each mouse according to the batch number, individual identifier, and actual target modification status recorded in the modification receipt. This ensures that there is no information confusion between mice in different experimental or control groups during phenotypic testing and facilitates subsequent statistical analysis and verification of experimental results.

[0103] The collection of experimental phenotypic data should cover multiple core physiological indicators related to type 2 diabetes, including basal metabolic rate, glucose metabolism capacity, insulin sensitivity, and lipid metabolism status. Routine procedures require fasting blood glucose measurement in modified mice, typically performed by collecting blood via the tail vein after a certain period of fasting using a standardized blood glucose meter. Simultaneously, a glucose tolerance test should be conducted, involving intraperitoneal injection of glucose solution and blood glucose level measurement at multiple time points to obtain a glucose clearance curve. In addition, an insulin tolerance test should be performed, evaluating peripheral tissue sensitivity to insulin by injecting exogenous insulin and monitoring the rate of glucose decline. For liver lipid metabolism detection, triglyceride and cholesterol levels can be measured by collecting serum samples; if necessary, liver tissue sections can be stained to assess lipid droplet accumulation and the degree of steatosis. These tests not only directly reflect the core pathology of diabetes but also provide multi-dimensional references for comparing phenotypic prediction results.

[0104] After data collection, these measured phenotypic data need to be systematically compared with the predicted results of the target phenotypic traits. The comparison process involves not only examining whether a single indicator closely approximates the predicted value, but also comprehensively analyzing the relative relationships between multiple indicators. For example, if the predicted results show that mice should exhibit elevated fasting blood glucose, decreased insulin sensitivity, and increased hepatic lipid accumulation under the target combination, then the measured data must simultaneously show a consistent trend in these three directions for the prediction to be considered validated. If some indicators deviate, such as blood glucose elevation exceeding the predicted value or insufficient lipid accumulation, a comprehensive phenotypic deviation measure needs to be calculated by comparing the magnitude and direction of the deviation. This deviation measure can be a numerical score, such as calculating the difference distribution between measured and predicted values ​​using statistical methods, or a graphical visualization result to more intuitively display the deviation.

[0105] After measuring phenotypic deviation, phenotypic determination results must be generated. These results should adhere to a pre-defined standard threshold system, typically set by researchers during the experimental design phase based on key indicators of the clinical diabetes phenotype. For example, if a mouse model meets the criteria for a composite diabetes phenotype when measured fasting blood glucose exceeds twice the normal value, the glucose tolerance test curve shows a significantly delayed recovery, and the insulin tolerance test indicates no peripheral tissue response to insulin, then the model is considered to have met the criteria. Conversely, if only some indicators meet the requirements, but the overall trend does not match the predicted phenotype, the model should be deemed unmet. The process of generating phenotypic determination results must be clearly documented, including reference thresholds, comparison methods, specific values ​​for deviation measurements, and the final conclusion, to ensure traceability and reproducibility throughout the process.

[0106] In summary, step S105 rigorously compares the measured phenotypic data of the modified mice with the predicted results of the target phenotypic, and generates phenotypic deviation measures and phenotypic determination results based on this comparison, thus achieving closed-loop verification between experimental data and theoretical predictions in model construction. This step not only verifies whether the modification operation has achieved the expected results, but also provides a solid basis for subsequent interaction map correction and candidate target combination optimization.

[0107] Furthermore, the step of obtaining measured phenotypic data of the modified mice based on the modification receipt, and comparing the measured phenotypic data with the target phenotypic prediction results to generate phenotypic deviation measures and phenotypic determination results includes:

[0108] Based on the modification receipt, the modified mice are grouped and identified, and an individual tracking relationship corresponding to the combination of target points is established.

[0109] The experimental phenotypic data of the modified mice were collected, including blood glucose levels, insulin resistance, and liver lipid metabolism indicators.

[0110] The measured phenotypic data are compared with the target phenotypic prediction results item by item to obtain the difference results of each phenotypic index.

[0111] Based on the difference results, a phenotypic deviation metric is calculated, and a phenotypic determination result is generated according to a preset determination criterion.

[0112] In this embodiment, the process of obtaining measured phenotypic data of the modified mice based on the modification receipt, and comparing the measured phenotypic data with the target phenotypic prediction results to generate phenotypic deviation measures and phenotypic determination results requires full consideration of the correspondence between the modified mice and the target target combination, as well as the measurement and calculation logic of multidimensional phenotypes. First, the modified mice must be clearly grouped and identified according to the modification receipt to ensure that each modified mouse can establish a unique tracking relationship with its corresponding target target combination. This process not only requires recording the modified mouse's number and modification site in the receipt information, but also establishing a data structure to correspond these numbers one-to-one with the target target combination, thereby ensuring that the genetic background and modification information of each group of mice can be accurately traced in subsequent phenotypic analysis.

[0113] After establishing the tracking relationship, experimental phenotypic data needs to be collected from the modified mice. This data should include at least fasting blood glucose levels, oral glucose tolerance test or insulin tolerance test results, and key indicators of hepatic lipid metabolism, such as hepatic triglyceride content and lipid droplet accumulation. These data should be obtained using standardized experimental methods. For example, blood glucose levels should be measured using a tail vein blood sample and a glucometer; insulin tolerance should be assessed by injecting insulin at set time points and monitoring the blood glucose decline curve; and hepatic lipid metabolism should be obtained through histological staining, metabolomics analysis, or quantitative mass spectrometry. Throughout the process, the controllability and reproducibility of experimental conditions should be ensured to guarantee that the experimental data accurately reflect the metabolic status of the modified mice.

[0114] After collecting the measured phenotypic data, each of these data needs to be compared with the predicted results of the target phenotypic. The comparison method involves using the predicted value as a reference standard and mapping the measured values ​​to the predicted values ​​item by item to obtain the differences in each phenotypic indicator. For example, regarding blood glucose levels, the difference between the measured blood glucose curve and the predicted blood glucose regulation curve can be compared; regarding insulin tolerance, the fit between the measured rate of blood glucose decline and the predicted curve can be compared; and regarding lipid metabolism, the difference between liver lipid content and the predicted level can be compared. These item-by-item comparison results will directly serve as the basis for calculating the phenotypic deviation metric.

[0115] Finally, a phenotypic deviation metric needs to be calculated based on the difference results, and the phenotypic determination result is generated according to a preset judgment criterion. The phenotypic deviation metric is typically implemented using a weighted error function, where different phenotypic indicators are assigned different weights based on their contribution to the pathology of type 2 diabetes. For example, the deviation weight of blood glucose level can be higher than that of lipid deposition indicators to reflect its central role. By calculating the comprehensive deviation metric, the closeness between the modified mouse phenotype and the predicted results can be quantitatively assessed. Subsequently, the phenotypic deviation metric is compared with a threshold according to the preset judgment criterion to generate the phenotypic determination result. If the deviation metric is within the threshold range, the determination result is satisfactory; if it exceeds the threshold range, the determination result is unsatisfactory.

[0116] Step S106: When the phenotypic determination result is not satisfied, update the interaction map using the phenotypic deviation metric, generate a corrected candidate target combination and a corrected candidate phenotypic response result, and return to the screening step; when the phenotypic determination result is satisfied, register the target target combination, the target phenotypic prediction result, the measured phenotypic data and the modification receipt to form a traceable type 2 diabetes composite mouse model.

[0117] In step S106, the core task is to handle the unmet and met cases based on the different phenotypic determination results, thereby ensuring that the entire method can form a dynamic iterative closed loop. First, when the phenotypic determination result shows that the expectation is not met, it means that there is a significant deviation between the measured phenotypic data and the predicted result of the target phenotypic. In this case, the deviation metric needs to be input as a feedback signal into the interaction map to readjust the node weights and edge connections to reflect the actual biological differences revealed by the modification results. Specifically, if the knockout of certain key genes does not cause the expected insulin resistance or hyperglycemia, the causal weight of the gene in the map needs to be reduced; conversely, if the knockout of a minor gene leads to significant metabolic disorders, the influence of the gene in the map should be increased. In this way, the map is gradually corrected after each round of experiments, gradually approaching the true pathological state.

[0118] The revised interaction map will be used as a new predictive model to run AI inference again, generating revised candidate target combinations and corresponding revised phenotypic responses. These revised candidate combinations not only compensate for the reasons for the previous failures but also incorporate new causal relationships, making them more likely to yield the expected disease phenotype in the next experiment. These new candidate results will return to the screening stage, undergoing further screening according to established scientific rationality verification, phenotypic consistency comparison, and ethical compliance constraints, thus entering a new round of construction instruction generation and experimental verification. Through this iterative process, the method itself possesses self-learning and self-correcting capabilities, continuously narrowing the gap between prediction and actual results, enabling the final model to stably reproduce the core phenotype of clinical type 2 diabetes.

[0119] When the phenotypic determination results show that the expected outcomes have been met, it means that the measured phenotypic data and the predicted target phenotypic results are highly consistent with each other in key indicators, and the modified mice exhibit characteristics consistent with the type 2 diabetes complex phenotype. In this case, all information related to the model construction needs to be registered and solidified. The registration content must be comprehensive and detailed, including a complete record of the target target combination, the target phenotypic prediction results output by artificial intelligence during the screening stage, the genetic modification implementation information provided by the modification receipt, and the finally collected measured phenotypic data. This content should be uniformly stored in a traceable database or archive system and protected by digital signatures or timestamp technology to ensure that the data is tamper-proof. This registration mechanism ensures that subsequent researchers can obtain complete reference information when they need to reproduce the model, and also provides a transparent basis for compliance and supervision.

[0120] After registration, the type 2 diabetes composite mouse model was officially confirmed as effective and incorporated into the experimental animal model system as a standardized resource. This model can be used not only for basic research but also in multiple scenarios such as drug development, mechanism exploration, and preclinical validation. Through the closed-loop control of step S106, the entire method achieves the dual function of iterative correction in case of failure and registration and solidification in case of success, thereby ensuring the reliability, reproducibility, and traceability of the method.

[0121] Furthermore, when the phenotypic determination result is not satisfied, the interaction map is updated using the phenotypic bias metric to generate a corrected candidate target combination and a corrected candidate phenotypic response result, and the process is returned to the screening step; when the phenotypic determination result is satisfied, the target target combination, the target phenotypic prediction result, the measured phenotypic data, and the modification receipt are registered to form a traceable type 2 diabetes composite mouse model, including;

[0122] The nodes and connections of the interaction map are weighted and adjusted based on the phenotypic bias metric to correct the dynamic interactions reflecting pancreatic β-cell function, insulin signaling, and hepatic lipid metabolism pathways.

[0123] Using the revised interaction map, a revised candidate target combination and revised candidate phenotypic response results are regenerated through an artificial intelligence model;

[0124] When the phenotypic determination result is not satisfied, the modified candidate target combination and the modified candidate phenotypic response result are returned to the screening step to enter a new round of iteration.

[0125] When the phenotypic determination result is satisfied, the target target combination, the target phenotypic prediction result, the measured phenotypic data and the modification receipt are uniformly archived, and a traceability relationship is established through digital identification to form the traceable type 2 diabetes composite mouse model.

[0126] In this embodiment, when the phenotypic determination result is not met, the interaction map is updated using the phenotypic deviation metric to generate a corrected candidate target combination and a corrected candidate phenotypic response result, and the process returns to the screening step. When the phenotypic determination result is met, the target target combination, the target phenotypic prediction result, the measured phenotypic data, and the modification receipt are registered to form a traceable type 2 diabetes composite mouse model. This process involves closed-loop correction of the difference between model prediction and measurement, as well as the registration and tracing of the final model. First, the interaction map needs to be weighted and adjusted based on the phenotypic deviation metric. Specifically, the phenotypic deviation metric obtained in the previous stage not only indicates the overall difference between the predicted and measured results, but also allows the identification of pathway nodes and connections related to the deviation through the distribution of differences. For example, if there is a significant deviation in measured blood glucose regulation, abnormalities in pancreatic β-cell functional nodes can be traced; if there is a large error in lipid metabolism, key nodes in the liver lipid metabolism pathway can be located. By converting the deviation values ​​into weight correction coefficients, the parameters of relevant nodes and connecting edges are updated, thereby making the corrected interaction graph more realistically reflect the dynamic interaction relationship under multi-target perturbation.

[0127] After refining the interaction map, it needs to be re-input into the AI ​​model for inference. Based on the new weighted map, the AI ​​model recalculates the node perturbation effects, generating revised candidate target combinations and revised candidate phenotypic responses. This process ensures that the model can iteratively refine its predictions from the original data, gradually approximating the true phenotypic response. The revised candidate target combinations and phenotypic responses are logically consistent with the previous screening process, allowing for a direct return to the screening step, where they can be combined with ethical compliance criteria for a new round of evaluation. This closed-loop iterative mechanism ensures that the entire method continuously optimizes prediction accuracy and dynamically adapts to different experimental feedback.

[0128] If the phenotypic determination results are satisfactory, the iteration process ceases, and the relevant data is uniformly archived. Specifically, a systematic registration and management system is required for target combination, target phenotypic prediction results, measured phenotypic data, and modification receipts. This process should employ a digital identification system to establish a unique correspondence between each set of data and the modified mouse, ensuring rapid tracking in subsequent research or source tracing analysis. Registration includes not only simple data storage but also the creation of indexes and source tracing chains in the database, enabling researchers to clearly query the association between a specific phenotypic result and a specific genetic modification operation. In this way, the resulting type 2 diabetes composite mouse model possesses traceability and scientific credibility, not only meeting the requirements of experimental reproducibility but also laying a data foundation for subsequent research and clinical translation.

[0129] This process introduces for the first time a dynamic map correction driven by phenotypic bias and an iterative mechanism for artificial intelligence models, and forms a complete closed-loop process of prediction, measurement and registration, so that the model tends to be stable and realistic through continuous correction, thereby ensuring that the constructed composite mouse model is not only reliable, but also fully traceable.

[0130] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.

Claims

1. A method for constructing a type 2 diabetes mouse model combining AI and multi-target knockout, characterized in that, include: Acquire and standardize mouse multi-omics data and phenotypic baseline data to form a standardized modeling dataset; Based on the standardized modeling dataset, an interaction map containing pancreatic β-cell function, insulin signaling, and hepatic lipid metabolism pathways was constructed, and artificial intelligence was used to predict and generate candidate target combinations and candidate phenotypic response results. Based on the candidate target combination and the candidate phenotype response results, and in conjunction with ethical compliance conditions, the target target combination and target phenotype prediction results are obtained through screening. Based on the target combination and the target phenotype prediction results, a construction instruction is generated, and the construction instruction is submitted to a controlled facility to perform genetic modification operations, resulting in modified mice and modification receipts; Based on the modification receipt, the measured phenotypic data of the modified mice are obtained, and the measured phenotypic data are compared with the target phenotypic prediction results to generate phenotypic deviation measurement and phenotypic determination results. When the phenotypic determination result is not satisfied, the interaction map is updated using the phenotypic bias metric to generate a corrected candidate target combination and a corrected candidate phenotypic response result, and the process is returned to the screening step. When the phenotypic determination result is satisfied, the target target combination, the target phenotypic prediction result, the measured phenotypic data and the modification receipt are registered to form a traceable type 2 diabetes composite mouse model.

2. The method for constructing a type 2 diabetic mouse model combining AI and multi-target knockout according to claim 1, characterized in that, The acquisition and standardization of mouse multi-omics data and phenotypic baseline data to form a standardized modeling dataset includes: Collect genomic, transcriptomic, proteomic, and metabolomic data of mice, and simultaneously collect baseline phenotypic data of mice; Data cleaning was performed on mouse multi-omics data and phenotypic baseline data, including batch effect correction, missing value imputation, and outlier removal. Normalization was performed on the cleaned mouse multi-omics data and phenotypic baseline data to unify different types of data into a comparable dimensional range. Based on individual mouse identifiers, normalized mouse multi-omics data and phenotypic baseline data are integrated to form a standardized modeling dataset.

3. The method for constructing a type 2 diabetic mouse model combining AI and multi-target knockout according to claim 1, characterized in that, The process involves constructing an interaction map based on the standardized modeling dataset, encompassing pancreatic β-cell function, insulin signaling, and hepatic lipid metabolism pathways. Artificial intelligence is then used to predict and generate candidate target combinations and candidate phenotypic response results, including: The standardized modeling dataset was used to extract features from genes, proteins and metabolites related to pancreatic β-cell function, insulin signaling and liver lipid metabolism pathways, and to obtain multidimensional features corresponding to pathway nodes. The multidimensional features are connected in a network based on intermolecular regulatory relationships and pathway mechanisms to form an interaction map that includes pancreatic β-cell function, insulin signaling and hepatic lipid metabolism pathways; The interaction map is input into an artificial intelligence model to infer and simulate the effects of perturbations at different nodes and joint perturbations at multiple nodes, generating candidate target combinations and their corresponding dynamic response curves. Based on the dynamic response curve, the predicted results of candidate target combinations in terms of glucose homeostasis, insulin sensitivity, and lipid metabolism are calculated, forming candidate target combinations and candidate phenotypic response results.

4. The method for constructing a type 2 diabetic mouse model combining AI and multi-target knockout according to claim 1, characterized in that, The step of selecting target target combinations and target phenotype prediction results based on the candidate target combination and the candidate phenotype response results, and in conjunction with ethical compliance conditions, includes: The pathogenic mechanism correlation analysis was performed on the candidate target combinations. Based on the interaction between pancreatic β-cell function, insulin signaling and hepatic lipid metabolism pathways, candidate target combinations that can cause multidimensional metabolic abnormalities were screened. The candidate target combinations obtained through correlation analysis are compared with the candidate phenotypic response results. Candidate target combinations whose phenotypic response results do not match the core phenotype of type 2 diabetes are eliminated, and candidate target combinations with multi-phenotypic consistency and their corresponding candidate phenotypic response results are obtained. Candidate target combinations with multi-phenotypic consistency and their corresponding candidate phenotypic response results are compared with ethical compliance conditions to exclude candidate target combinations that may cause serious irreversible damage or do not meet the ethical requirements of animal experiments, thereby obtaining compliant candidate target combinations and their corresponding candidate phenotypic response results. Priority evaluation is performed based on the compliant candidate target combination and its corresponding candidate phenotypic response results. The ranking is determined according to the comprehensive weight of the predicted blood glucose regulation ability, the degree of change in insulin sensitivity, and the effect of liver lipid deposition, and the target target combination and the target phenotypic prediction results are selected.

5. The method for constructing a type 2 diabetic mouse model combining AI and multi-target knockout according to claim 1, characterized in that, The process of generating a construction instruction based on the target target combination and the target phenotype prediction results, and submitting the construction instruction to a controlled facility to perform genetic modification operations, resulting in modified mice and modification receipts, includes: The target combination is transformed into a specific genetic modification scheme, and the modification method, modification location and required editing tool type for each target are determined to form a target modification specification. The target modification specifications are compared with the target phenotype prediction results to generate a modification priority sequence, and the modification order and corresponding phenotype expectations are marked in the construction instructions. The genetic modification execution conditions, including mouse strain selection, controlled facility operation requirements, and necessary verification steps, are integrated into the construction instructions to form a complete construction instruction. The construction instructions are submitted to the controlled facility to perform genetic modification operations, and the modified mouse information and modification receipt returned by the controlled facility are received and recorded.

6. The method for constructing a type 2 diabetic mouse model combining AI and multi-target knockout according to claim 1, characterized in that, The process of obtaining measured phenotypic data of the modified mice based on the modification receipt, comparing the measured phenotypic data with the target phenotypic prediction results, and generating phenotypic deviation measures and phenotypic determination results includes: Based on the modification receipt, the modified mice are grouped and identified, and an individual tracking relationship corresponding to the combination of target points is established. The experimental phenotypic data of the modified mice were collected, including blood glucose levels, insulin resistance, and liver lipid metabolism indicators. The measured phenotypic data are compared with the target phenotypic prediction results item by item to obtain the difference results of each phenotypic index. Based on the difference results, a phenotypic deviation metric is calculated, and a phenotypic determination result is generated according to a preset determination criterion.

7. The method for constructing a type 2 diabetic mouse model combining AI and multi-target knockout according to claim 1, characterized in that, When the phenotypic determination result is not satisfied, the interaction map is updated using the phenotypic bias metric to generate a corrected candidate target combination and a corrected candidate phenotypic response result, and the process is returned to the screening step. When the phenotypic determination result is satisfied, the target target combination, the target phenotypic prediction result, the measured phenotypic data, and the modification receipt are registered to form a traceable type 2 diabetes composite mouse model, including: The nodes and connections of the interaction map are weighted and adjusted based on the phenotypic bias metric to correct the dynamic interactions reflecting pancreatic β-cell function, insulin signaling, and hepatic lipid metabolism pathways. Using the revised interaction map, a revised candidate target combination and revised candidate phenotypic response results are regenerated through an artificial intelligence model; When the phenotypic determination result is not satisfied, the modified candidate target combination and the modified candidate phenotypic response result are returned to the screening step to enter a new round of iteration. When the phenotypic determination result is satisfied, the target target combination, the target phenotypic prediction result, the measured phenotypic data and the modification receipt are uniformly archived, and a traceability relationship is established through digital identification to form the traceable type 2 diabetes composite mouse model.

Citation Information

Patent Citations

  • Construction method and application of pancreatic beta cell Pik3r3 gene conditional knockout mouse model

    CN117025686A

  • Application of serpinb3 / b4 as target in drugs for treating inflammatory skin diseases such as rosacea

    WO2023092768A1