Genome selection method, device and equipment based on machine learning, medium and computer program product
By integrating multiple machine learning models and deep learning frameworks, and combining them with automated parameter optimization and parallel computing technologies, we have solved the problems of existing genome selection tools, such as model uniformity, data scale limitations, insufficient flexibility, lack of parameter optimization functions, and lack of human-computer interaction friendliness, and achieved an efficient and accurate genome selection method.
Patent Information
- Application Number
- CN202510738200.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Existing genomic selection tools have problems such as model uniformity, data scale limitations, insufficient flexibility, lack of parameter optimization functions, single molecular marker types, and lack of human-computer interaction friendliness. They are difficult to adapt to the efficient processing of different genetic structures and large-scale genomic data.
It adopts a machine learning-based approach, integrates multiple machine learning models and deep learning frameworks, combines automated parameter optimization with parallel computing technology, provides a graphical interface and command line interaction, supports the processing of multi-omics molecular marker data, parallel computing and adaptive modeling, and achieves efficient prediction.
It significantly improves the prediction accuracy and computational efficiency of genomic selection, supports large-scale breeding data analysis, simplifies the operation process, and improves user-friendliness and model generalization ability.
Smart Images

Figure CN120613016A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular to a genome selection method, apparatus, device, medium and computer program product. Background Art
[0002] With the advancement of high-throughput SNP typing technology, genomic selection, a statistical analysis method, has been widely used not only in animal and plant breeding to assess the genetic potential of complex traits due to its powerful predictive capabilities, but is also increasingly used in human genetics research. However, existing genomic selection tools for plants and animals often rely on single statistical models, such as linear regression or mixed linear models, which have the following limitations:
[0003] 1. Model singleness: A single model is difficult to adapt to the prediction of complex traits with different genetic structures, which easily leads to insufficient generalization ability.
[0004] 2. Data scale limitations: Existing tools cannot be effectively applied when processing massive genomic data, such as SNP data of tens of thousands of samples, due to low computational efficiency or insufficient memory.
[0005] 3. Lack of flexibility: Existing tools often use the same model to calculate data from different species, populations, and batches, lack the ability to differentiate data from different sources, and do not support user-defined model structures.
[0006] 4. Lack of parameter optimization capabilities: Most existing tools do not integrate automated hyperparameter optimization functions and rely on manual experience adjustment, which affects model performance.
[0007] 5. Single type of molecular marker: Current methods mainly focus on genomic SNP markers and cannot support transcriptome, metabolome, proteome and other omics data. A single marker type makes it difficult to achieve higher-precision predictions.
[0008] 6. Lack of user-friendly human-computer interaction: Existing tools are mostly command-line based and lack graphical user interfaces (GUIs) and visualization analysis modules, which creates a high barrier to entry for users with a weak bioinformatics background. Furthermore, data processing, model training, and result analysis are not streamlined, requiring users to manually connect the various calculation steps, significantly increasing operational complexity.
[0009] Based on the above technical bottlenecks, there is an urgent need to develop a genome selection method, device, equipment, medium and computer program product that can support multi-omics molecular marker data, support parallel computing, have adaptive modeling capabilities, be human-computer interactive and simple to operate. Summary of the Invention
[0010] In response to the above-mentioned deficiencies in the existing technology, the present invention provides a genome selection method, device, equipment, medium and computer program product based on machine learning. By integrating multiple machine learning models and deep learning frameworks, combined with automated parameter optimization and parallel computing technology, it significantly improves prediction accuracy and computing efficiency, and meets the needs of large-scale breeding data analysis.
[0011] The present invention is achieved in that:
[0012] The present invention provides a genome selection method based on machine learning, which is applied to a cloud platform and is used by users to calculate breeding data with limited scale online. Data uploading, model selection, covariate correction, real-time log monitoring and result downloading are achieved through a graphical interface. The operation does not require programming knowledge and is completed through web page interaction. Users with a certain coding foundation can calculate large-scale breeding data on a personal computer or server, configure parameters such as genotype files, phenotype files, model types through command line parameters, support multiple machine learning algorithms, and output result files and model files.
[0013] Preferably, the machine learning-based genomic selection method is written in Python language as IASML software, which uses the phenotype and genotype of experimental animal individuals as input information, and uses individuals with both phenotypes and genotypes as a reference group to build a model. The model is used to predict the phenotypic values of individuals with unknown phenotypes, and the prediction results are output as the breeding values of the individuals.
[0014] The genome selection method based on machine learning of the present invention comprises the following steps:
[0015] S1: Data preprocessing, processing genotype and phenotype files separately:
[0016] S11 Genotype data processing: Convert the input PLINK format file to a standardized NumPy matrix, process missing values in parallel blocks, fill in the mean, and output a `.npy` file for subsequent calculations; or digitize the genotypes, with the codes for genotypes AA, AB, and BB being 0, 1, and 2, respectively, and input the digitized results as a TXT file;
[0017] S12 Phenotype data processing: extract the target trait column specified by `--phe-pos` from the input `.txt` phenotype file, use the linear model to correct the covariates, and generate the cleaned training set data;
[0018] S2: Model configuration and training, select the model to be used for calculation and the configuration method of the model parameters, including two methods: pre-training based on input data and direct specification:
[0019] If the `--model` parameter is specified, the input reference group data is used for pre-training. The scikit-learn model uses random search `n_iter` iterations to optimize parameters, and then uses three-fold cross-validation to build the model on the reference group data. The best performing parameter combination is ultimately selected as the final model hyperparameters. The deep learning model uses a fixed network structure, continuously updates parameters through gradient descent during training, and ultimately saves the optimal weights to a `.keras` file.
[0020] If `--model-params` is specified, the hyperparameters specified by the user are used for training directly, and the parameter search step is skipped;
[0021] S3: prediction and result output, including,
[0022] S31. Phenotype prediction: Use the trained model to predict the phenotype of an individual with unknown phenotype and output the predicted value. Save the result to `{out}_predict.txt`.
[0023] S32. Model saving: saves the scikit-learn model parameter file (`{out}_model.txt`) and the Keras model file (`.keras`), supporting direct input for subsequent calculations;
[0024] S4: Extended function, perform K-fold cross validation of phenotypic data through `--split-seed`, generate
[0025] `Ref{fold}.txt` training set and `Val{fold}.txt` validation set are used to evaluate the model used for prediction.
[0026] The present invention also provides a machine learning-based genome selection device, including an IASML breeding software module and an IASML breeding cloud platform module;
[0027] The IASML breeding software module includes:
[0028] R1 supports reading and preprocessing of genotype data in PLINK binary format (bed / bim / fam) and TXT format, including missing value filling, genotype encoding conversion, and parallel block processing;
[0029] R2 integrates a variety of machine learning models, including support vector machines (SVM), linear regression, Ridge regression, Lasso regression, Elastic net regression, PLS regression, decision trees, random forests, gradient boosting machines, LightGBM, XGBoost, convolutional neural networks (CNN), and multi-layer perceptrons (MLP), and provides random search hyperparameter optimization capabilities;
[0030] R3 supports two training modes: hyperparameter search and training with specified parameters. The neural network model can save all parameters and use them directly for inference.
[0031] R4 supports cross-validation segmentation of phenotypic data and generates training set and validation set files;
[0032] The IASML breeding cloud platform module includes:
[0033] Y1 provides a responsive graphical interface that allows users to upload PLINK or TXT format genotype files and TXT format phenotype files by dragging and dropping or selecting, and verifies file size and format in real time;
[0034] Y2 dynamically generates phenotypic column selectors and interactive selection components for factor / numeric covariates, supporting multi-column collaborative filtering and missing value prompts;
[0035] Y3 provides the function of uploading preset models, custom hyperparameter files, and pre-trained Keras models, and triggers the calculation process through a visual button;
[0036] Y4 displays and analyzes log streams in real time, supports one-click downloading of prediction files and model files, and integrates error handling mechanisms and user notification systems;
[0037] Y5 uses temporary working directories to isolate computing tasks of different users, automatically retains result files, and cleans up intermediate data to ensure system efficiency.
[0038] Preferably, the IASML breeding cloud platform module of the machine learning-based genomic selection device includes a homepage, an IASML software download and online usage page, and its user interface and deployment architecture include:
[0039] The interactive web-based interface includes dynamic navigation, mathematical formula rendering in KaTeX, and real-time log monitoring. It uses a front-end and back-end separation architecture, with the front-end implementing a responsive layout using HTML / CSS, and the back-end processing computing tasks based on a Python asynchronous framework.
[0040] Supports one-click cloud deployment and localized private deployment.
[0041] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the aforementioned genome selection methods when executing the computer program.
[0042] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program implements any of the aforementioned genome selection methods when executed by a processor.
[0043] The present invention also provides a computer program product, comprising a computer program, characterized in that when the computer program is executed by a processor, any of the aforementioned genome selection methods is implemented.
[0044] The machine learning-based genome selection method, apparatus, device, medium, and computer program product provided by the present invention have the following advantages:
[0045] 1. It integrates 13 machine learning algorithms, including SVM, Ridge, Lasso, Elastic net, Decision_tree, Random_forest, LightGBM, XGBoost, Linear, PLS, GBM, CNN, and MLP, for calculation. It is used to support the calculation of breeding values of individuals of different species, populations, and batches, including four major types: linear model, tree model, support vector machine, and neural network.
[0046] 2. Efficient data preprocessing capabilities. Compatible with PLINK format input for genotype data, it fills missing values through parallel block processing and outputs a standardized matrix. For phenotypic data, it automatically corrects for factorial and numerical covariates, filters missing samples, and ultimately generates training and prediction sets.
[0047] 3. It has an automated parameter optimization function, using the randomized search (CV) method to optimize hyperparameters for the scikit-learn model through cross-validation (the default is 3-fold). The optimization target is the negative mean square error (NMSE), calculated as follows:
[0048]
[0049] In addition, deep learning model optimization introduces an integrated early stopping mechanism (Early Stopping), dynamic learning rate adjustment (Reduce LR On Plateau) and model checkpoint (Model Checkpoint) to effectively prevent overfitting and improve the generalization ability and stability of the model.
[0050] 4. Implemented parallel computing and resource optimization management. The Joblib framework supports multi-threaded parallel computing, significantly improving the efficiency of data preprocessing and model training. Furthermore, combined with the dynamic memory recycling mechanism GC module, it effectively optimizes resource usage and ensures efficient management of memory usage.
[0051] 5. Achieve full process traceability, with a comprehensive logging system that records key operations and error messages, ensuring transparency and auditability at every stage. Furthermore, the system automatically generates model parameter files (.txt) and prediction result files (.txt), facilitating the reproduction of results and subsequent in-depth analysis, enhancing workflow reproducibility and data integrity.
[0052] This invention significantly improves breeding accuracy by integrating multiple machine learning algorithms, efficiently processing large-scale molecular marker data, and combining parallel computing and parameter optimization technology, providing an intelligent solution for phenotypic prediction, genetic evaluation and selection optimization in biological breeding. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 This is the usage diagram of the IASML breeding cloud platform module of the present invention Figure 1
[0054] Figure 2 This is the usage diagram of the IASML breeding cloud platform module of the present invention Figure 2 DETAILED DESCRIPTION
[0055] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0056] The genome selection method, apparatus, device, medium, and computer program product of the present invention are further described below with reference to the accompanying drawings.
[0057] The device of the present invention includes two parts: the IASML breeding software module and the IASML breeding cloud platform module. The specific implementation of the IASML breeding software module is as follows:
[0058] 1. Data preprocessing process
[0059] The main data preprocessing process includes:
[0060] 1.1 Molecular marker data preprocessing:
[0061] The input format supports PLINK binary format (.bed / .bim / .fam) or TXT text format (.txt). "joblib.Parallel" is used to process large-scale labeled data in parallel (block size is 10,000 markers), fill missing sites according to the sample mean, and finally generate a standardized NumPy array file (.npy) for subsequent calculations.
[0062] 1.2 Phenotypic data preprocessing
[0063] 2. Machine Learning Model Implementation Details
[0064] 2.1 Model Principle:
[0065] 2.1.1 Linear Model
[0066] Linear regression uses mean square error (MSE) as the objective function. By minimizing the objective function, the regression coefficients (i.e., model parameters) in the model are optimized, and the optimal parameter combination is finally fitted, thereby minimizing the model's fitting error to the data. The model is as follows:
[0067] y=β0+β1x1+β2x2+…+β i x i +∈
[0068] Where y is the phenotypic value; β0 is the intercept term, which represents the predicted value of y when all independent variables are 0;
[0069] β1, β2, ..., β i is the regression coefficient, also called weight or parameter, which represents the influence of each variable on y; x1, x2, ..., x i is a molecular marker, also known as an input feature; ∈ is an error term, which is usually assumed to be normally distributed with a mean of 0. During the training process, the linear regression model gradually adjusts the regression coefficients by minimizing the mean squared error (MSE), so that the model can more accurately predict the phenotype. MSE is defined as follows:
[0070]
[0071] Where n is the number of samples, y i is the actual phenotypic value, is the phenotypic value predicted by the model.
[0072] Lasso regression, Ridge regression, and Elastic net regression are three machine learning methods based on linear regression models, but they use different objective functions for parameter optimization. They respectively incorporate L1 regularization, L2 regularization, and a combination of L1 and L2 regularization into the objective function to improve the ability to adjust the regression coefficients and effectively address overfitting and multicollinearity issues. Their objective functions are as follows:
[0073]
[0074] Among them, y i is the true phenotypic value of the i-th sample; is the phenotypic value predicted by the model; is the regression coefficient for the jth feature; λ is the regularization parameter used to control the regularization weight. λ1 is the weight of the L1 regularization term, and λ2 is the weight of the L2 regularization term. p is the number of features. Introducing regularization into the objective function can compress the regression coefficients, thereby reducing feature participation. L1 and L2 regularization achieve varying degrees of feature compression, which can both improve multicollinearity in the model and prevent overfitting. The specific value of λ will be determined through hyperparameter search.
[0075] Partial least squares regression (PLS) can be viewed as a linear model. Its goal is to model the relationship between the independent variable (molecular marker) and the dependent variable (phenotype) y by extracting a set of orthogonal factors (latent variables). This also provides a new strategy for explaining the relationship between molecular markers and phenotypes. These latent variables maximize their covariance with the response variable y. This method assumes that both X and y can be decomposed into a bilinear form:
[0076]
[0077] Among them, T A is an n×A score matrix (latent variable), P A is a p×A loading matrix, q A is a 1×A load vector, E A is an n×p residual matrix, f A is an n×1 residual vector. The goal is to find the maximum score t j The weight w of the covariance between (latent variable) and the response variable y j This goal is achieved through the following optimization problem:
[0078]
[0079] The number of latent variables will exist as a tunable parameter, and the optimal value will be determined through hyperparameter search.
[0080] 2.1.2 Tree Model
[0081] The tree model constructs prediction rules by recursively partitioning the feature space, which can capture nonlinear relationships and high-order interactions. It is often used to deal with complex genetic effects and environmental interactions in breeding. Its core idea is to gradually divide the samples into different subsets according to the values of molecular markers (features), so that the phenotypic values in each subset are as homogeneous as possible. The tree models used in the present invention are decision trees, random forests, gradient boosting machines with regression trees as weak learners, LightGBM and XGBoost.
[0082] The decision tree (decision_tree) is constructed based on a feature splitting criterion (Gini index or information gain), recursively partitioning the data by maximizing the purity of the child nodes. For regression tasks, the splitting features and split points are usually selected with the goal of minimizing the mean squared error (MSE). The model can be expressed as:
[0083]
[0084] Among them, x is the feature vector composed of molecular markers, is the phenotypic value; R m is a region of the feature space (corresponding to the leaf node of the tree), c m is the mean of the sample phenotypes in the region, and I(·) is the indicator function. The decision tree selects the split feature J and the split point s by the following rules:
[0085]
[0086] Each splitting node only evaluates a portion of randomly selected features, and the decision tree will continue to split according to the above rules until any hyperparameter reaches the upper limit or the features in the node cannot be split again.
[0087] Random Forest improves generalization ability by integrating multiple decision trees. The samples used by each tree are obtained by bootstrap sampling, and the phenotypic prediction value is the mean of all trees:
[0088]
[0089] Among them, T k (x) is the predicted value of the kth tree, where K is the total number of trees. Random forests identify key molecular markers by evaluating feature importance. For example, for the jth feature, the sum of the MSE reductions when splitting nodes in all trees is calculated and normalized to obtain the importance value:
[0090]
[0091] where ΔMSEj,t is the MSE reduction of the jth feature when it is split at node t, ΔMSE j,t is the mean square error MSE of the parent node parent and the mean square error MSE of the left child node after splitting left Mean square error (MSE) with the right child node right The difference of the sum of , where:
[0092]
[0093] ΔMSE j,t =MSE parent -(MSE left +MSE right )
[0094] The Gradient Boosting Machine (GBM) uses an additive model to gradually fit the residuals and optimizes the loss function through gradient descent. Its prediction value is the weighted sum of the outputs of multiple decision trees:
[0095]
[0096] Among them, f k is the K-th tree, η is the learning rate, is the set of all trees. Its core is to iteratively train multiple decision trees, each tree correcting the pseudo residual of the previous model. In the kth iteration, the pseudo residual is calculated as:
[0097]
[0098] That is, the pseudo residual is the loss function L with respect to the current model prediction value The negative gradient of . A new decision tree f will be constructed k , the goal is to fit these pseudo residuals r ik , that is, minimize:
[0099]
[0100] Iterations will continue until the number of decision trees built reaches the preset hyperparameter upper limit.
[0101] XGBoost further optimizes the objective function based on the gradient boosting machine and introduces a regularization term to control the complexity of the model:
[0102]
[0103] Among them, N k is the number of leaf nodes in the kth tree, w k is the leaf weight, γ and λ penalize the tree structure and weight size respectively. When splitting, the features and split points that maximize the gain are selected:
[0104]
[0105] G and H are the first and second order derivatives of the loss function, respectively.
[0106] LightGBM uses a histogram algorithm to accelerate the search for split points and discretize continuous features into histogram bins. For example, when processing genotype data, it will be directly binned according to genotype:
[0107]
[0108] Directly count the total gradient G of samples of this category in each bin b b and the sum of the second derivatives H b :
[0109]
[0110] in is the first-order gradient, is a second-order gradient. The split point only needs to be selected between b bins, not all samples, for the candidate split point s
[0111] (i.e., the boundary between adjacent bins), calculate the gain formula:
[0112]
[0113] Where λ is the L2 regularization coefficient and γ is the leaf splitting gain threshold. The amount of training data is reduced by using gradient unilateral sampling (GOSS). The gradient statistics after sampling are:
[0114]
[0115] Where A is a set of large gradient samples, and B is a set of randomly sampled small gradient samples. The objective function of LightGBM is similar to that of XGBoost, but it is optimized by the above technology:
[0116]
[0117] Among them, T k is the number of leaf nodes in the kth tree, w k For leaf weights, γ and λ control the complexity.
[0118] 2.1.3 Support Vector Machine
[0119] Support vector machines, with their unique fitting strategy, can effectively reduce the impact of outliers on the model. Their core concept is not to fit the data with an optimal straight line, but rather to find an optimal regression hyperplane for phenotypic prediction. The goal of SVR is to maximize the margin between the regression hyperplane and the training data points while minimizing the model's prediction error. SVR solves the following optimization problem:
[0120]
[0121] Constraints:
[0122]
[0123] Where w is the weight vector, b is the bias term, ξ is the slack vector, and ∈ is the tolerance interval that defines the acceptable range of prediction error. C is the regularization parameter that balances the interval width and the prediction error penalty. i ) is a kernel function that maps input features to a high-dimensional space to handle nonlinearity. Kernel functions include the following types:
[0124] Linear kernel: The mapping function is an identity transformation, which directly calculates the inner product of the original feature space:
[0125]
[0126] Applicable to scenarios where there is a dominant linear relationship between traits (additive genetic effects).
[0127] Polynomial kernel: Extends nonlinear expression capabilities by the order of the polynomial:
[0128]
[0129] Among them, c is the constant term and d is the polynomial order, which is suitable for capturing moderately complex nonlinear relationships such as gene-environment interactions.
[0130] Radial basis kernel: Mapping to infinite dimensional space through Gaussian function, fitting highly nonlinear patterns: K(x i , x j )=exp(-γ||x i -x j || 2 )
[0131] where ||x i -x j || 2 is the square of the Euclidean distance between samples, γ is the Gaussian kernel width parameter, and the larger the value, the more sensitive the model is to local structure, and it is suitable for complex epistatic effects or non-additive genetic structures.
[0132] 2.1.4 Neural Networks
[0133] The MLP model is designed with two hidden layers and one input layer. The first hidden layer contains 96 neurons and uses the sigmoid activation function. The second hidden layer contains 64 neurons and uses the GELU activation function. The model is as follows:
[0134] y=f(W (3) g(W (2) h(X)+b (2) )+b (3) )
[0135] Where X represents the pruned valid marker information, y represents the phenotypic information, and W (2) and b (2) represents the weight and bias of the first hidden layer, W (3) and b (3) represents the weights and biases of the second hidden layer. h(X), g(X), and f(X) represent the activation functions of each layer starting from the input layer, and are expressed as follows:
[0136] h(x)=x
[0137]
[0138]
[0139] The loss function of MLP is as follows:
[0140]
[0141] where y i represents the true phenotypic value. represents the predicted value of the phenotype.
[0142] The CNN model uses a one-dimensional convolutional neural network (Conv1D) model. The model's input layer receives preprocessed SNP marker data in the form of a one-dimensional feature vector. Next, the input data passes through the first one-dimensional convolutional layer (Conv1D). This convolutional layer uses 36 convolutional kernels (filters), each with a kernel size of 3, and applies the GELU activation function:
[0143]
[0144] Among them, h t is the convolution output of the t-th SNP position, ω i is the weight of the convolution kernel, b is the bias term, x trepresents the value of the t-th SNP in the input. After the convolutional layer, a dropout layer is introduced with a dropout rate set to 0.5. During training, the dropout layer randomly sets the output of the neuron to zero with a probability of 50% to reduce overfitting. The dropout operation can be expressed as:
[0145]
[0146] Where f is the final output of the convolutional layer, and p is the dropout rate, which is set to 0.5 here. After the Dropout layer, the output of the convolutional layer will be flattened by the Flatten layer so that the data can be input into the fully connected layer. The Flatten operation flattens the multi-dimensional feature map into a one-dimensional vector:
[0147] f′=Flatten(f)
[0148] Where f′ is the flattened vector. The flattened vector is fed into the fully connected layer, which will be used to integrate the features extracted from the convolutional layer and finally perform regression prediction:
[0149] y=Wf′+b
[0150] Where y is the phenotypic value vector, W is the weight matrix of the fully connected layer, and b is the bias matrix of the fully connected layer. This allows for phenotype prediction.
[0151] 2.2 Hyperparameter Search
[0152] The linear model, tree model, and support vector machine used in the present invention all require specific hyperparameters. Each method sets a corresponding parameter space. The training set data is searched for the optimal parameter combination in the parameter space according to three-fold cross-validation. The optimal parameter combination is found after a certain number of iterative searches. The parameters and parameter spaces of each method are shown in Table 1.
[0153] Table 1
[0154]
[0155]
[0156]
[0157]
[0158]
[0159] 2.3 Neural Network Optimization Strategies
[0160] The main optimization strategies of the neural network in this invention are as follows:
[0161] 2.3.1 Basic Optimizer Configuration
[0162] 1. Use the Adam optimizer with an initial learning rate of 0.001.
[0163] 2. The loss function uses mean square error (MSE).
[0164] 3. The batch size (batch_size) is fixed to 32.
[0165] 4. The maximum number of training epochs is set to 50.
[0166] 2.3.2 Dynamic Learning Rate Adjustment Strategy
[0167] 1. Monitoring indicator: training loss.
[0168] 2. Adjustment strategy: When the loss stagnates (no improvement after 5 epochs), the learning rate is decayed to 50% of the current value.
[0169] 3. Lower limit: The minimum learning rate is not less than 1e-7.
[0170] 4. Design purpose: Fine-tune parameters in the later stages of training.
[0171] 2.3.3 Training Early Stopping Strategy
[0172] 1. Primary early stopping: Monitor the loss and tolerate no improvement for 50 epochs.
[0173] 2. Assisted early stopping: Custom callback function: relative improvement threshold (min_delta) = 0.1; tolerance (patience) = 3 times; automatically save the best model (best_model.keras).
[0174] 2.3.4 Training Process Control
[0175] 1. Double early stopping strategy: combining absolute threshold and relative improvement threshold.
[0176] 2. Model checkpoint: Save both the intermediate best model and the final model.
[0177] 3. Random seed fixation: ensure reproducibility through set_random_seed(42).
[0178] The IASML breeding cloud platform module utilizes a B / S architecture, consisting of a front-end interactive interface, a back-end analysis engine, and a distributed computing module. The front-end is built on a responsive web framework, supporting multi-terminal access. The back-end utilizes a microservices architecture, integrating a machine learning algorithm library with high-throughput data analysis modules. The computing layer implements asynchronous parallel processing through task queues, ensuring efficient computation of large-scale molecular marker data.
[0179] The main functional modules are as follows:
[0180] 1. Data upload module
[0181] 1.1 Molecular marker processing unit: supports uploading molecular phenotype files in Plink binary format (bed / bim / fam) and TXT text format, automatically verifies file integrity (MD5 checksum comparison) and size limit (≤100MB).
[0182] 1.2 Phenotypic data parser: uses streaming reading technology to parse TXT format phenotypic files, dynamically identifies the trait name and individual ID columns in the file header, and constructs a sparse matrix storage structure.
[0183] 1.3 Covariate Selector: Provides a dual selection interface for categorical variables (such as variety and environment) and continuous variables (such as temperature and humidity), and supports regular expression matching.
[0184] 2. Model configuration module
[0185] 2.1 Pre-built model library: integrates 13 machine learning models (including linear models, tree models, support vector machines and neural network models), and has built-in parameter optimization strategies (consistent with the strategies of each method in the breeding method).
[0186] 2.2 Custom model interface: supports uploading Keras model files (.keras) and parsing hyperparameter configuration files (TXT format).
[0187] 3. Task Execution Engine
[0188] 3.1 Create an independent temporary working directory.
[0189] 3.2 Call the core algorithm module (IASML.py) through a subprocess and capture the standard output stream in real time.
[0190] 3.3 Error recovery mechanism: When the process terminates abnormally, the breakpoint status is automatically saved and a SIGTERM signal is sent.
[0191] 4. Result output module
[0192] 4.1 Dynamic log system: Uses a ring buffer to store the latest 100 log lines and pushes them in real time via WebSocket.
[0193] 4.2 A split result output architecture is adopted, allowing users to download result files and model files separately as needed.
[0194] Example 1
[0195] The IASML breeding software was used to perform cross-generational genomic prediction of the three important slaughter traits of white-feathered meat duck populations, namely breast muscle weight, leg weight and carcass weight, using a support vector machine model.
[0196] Step 1: Obtain the required files: The three files "genotype.bim, genotype.bed, and genotype.fam" include the genotypes of the previous and next generations of white-feathered meat ducks; the file "phenotype.txt" includes the breast muscle weight, leg weight, and carcass weight phenotypes of the previous generation of white-feathered meat ducks.
[0197] Step 2: Use the following command in the command line (taking chest muscle weight as an example):
[0198] "python IASML.py --bfile genotype --phe phenotype.txt --phe-pos 15 --model svm --out result" uses "--bfile" and "--phe" to specify the genotype and phenotype files used in the calculation, "--phe-pos" determines that the breast muscle weight trait is located in the 15th column, uses "--model" to specify the model type, and finally uses "--out" to specify the output file name to start the calculation.
[0199] Step 3: The log information shows that the "genotype.npy" file was generated. After three-fold cross-validation of 20 parameter combinations and 60 rounds of model training and validation, the optimal parameter combination {'kernel':'rbf','gamma':'scale','epsilon':0.2,'C':1} was found. Finally, the support vector machine model built using this optimal parameter calculated the next generation of white-feathered broiler duck breast muscle weight breeding value and saved it in the "result_predict.txt" file. The model parameter information was saved in the "result_model.txt" file.
[0200] Step 4: Calculate the Pearson correlation between the calculated breeding value and the true value. The final results are shown in the following table:
[0201]
[0202]
[0203] The prediction accuracy of the three traits was higher than that of GBLUP.
[0204] Example 2:
[0205] The IASML breeding cloud platform was used to perform cross-generational genomic prediction of two important feed efficiency traits, feed intake and feed conversion rate, in white-feathered broiler duck populations using the Ridge regression model.
[0206] Step 1: Obtain the required files: The three files "genotype.bim, genotype.bed, and genotype.fam" include the genotypes of the previous and next generations of white-feathered broiler ducks; the file "phenotype.txt" includes the phenotypes of feed intake and feed conversion rate of the previous generation of white-feathered broiler ducks.
[0207] Step 2: Log in to the cloud platform https: / / iasbreeding.cn / IASML. Upload the prepared genotype and phenotype files, select the phenotype, select the Rdige model, and click the Start Calculation button. Figure 1 .
[0208] Step 3: Click the button to download the model parameters and prediction results to obtain the "result_model.txt" file and the "result_predict.txt" file.
[0209] Step 4:
[0210] The calculated breeding value and the true value are subjected to Pearson correlation, and the final results are shown in the following table:
[0211]
[0212] The prediction accuracy of both traits was higher than that of GBLUP.
[0213] Example 3: Cross-generational prediction of three important economic traits of white-feathered meat ducks: head weight, neck weight, and wing weight. First, use IASML software to perform a large number of iterations to find the optimal parameter combination, and then upload the model parameter file on the IASML cloud platform for fast calculation.
[0214] Step 1: Obtain the required files: The three files "genotype.bim, genotype.bed, and genotype.fam" include the genotypes of the previous and next generations of white-feathered meat ducks; the file "phenotype.txt" includes the head weight, neck weight, and wing weight phenotypes of the previous generation of white-feathered meat ducks.
[0215] Step 2: Use the following command in the command line (taking head weight as an example):
[0216] The "python IASML.py --bfile genotype --phe phenotype.txt --phe-pos 17 --model svm --n-iter 100 --out result" command adds the "--n-iter" option to control the parameter search intensity. This parameter search intensity can be adjusted based on the computing environment and actual needs. The default parameter search combination is 8. Finally, save the generated "result_model.txt" file. In actual production, this step can be performed before the new generation is generated. After the new generation is generated, the saved model parameter file can be directly used to significantly reduce computing time. Therefore, if cost allows, the parameter search intensity can be increased to find the optimal parameters as much as possible.
[0217] Step 3: Log in to the cloud platform https: / / iasbreeding.cn / IASML. Upload the prepared genotype and phenotype files, select the phenotype, upload the model parameter file, and click the Start Calculation button. Figure 2 .
[0218] Step 4: Click the button to download the model parameters and prediction results to obtain the "result_predict.txt" file.
[0219] Step 5: Calculate the Pearson correlation between the calculated breeding value and the true value. The final results are shown in the following table:
[0220]
[0221]
[0222] The prediction accuracy of the three traits was higher than that using GBLUP.
Claims
1. A genome selection method based on machine learning, characterized in that: It is applied to the cloud platform for users to calculate breeding data with limited scale online. Data upload, model selection, covariate correction, real-time log monitoring and result download can be achieved through a graphical interface. No programming knowledge is required for operation, and it is completed through web page interaction. Users with a certain coding foundation can calculate large-scale breeding data on personal computers or servers, configure parameters such as genotype files, phenotype files, model types through command line parameters, support multiple machine learning algorithms, and output result files and model files.
2. The genome selection method based on machine learning according to claim 1, characterized in that The IASML software was written in Python. It takes the phenotype and genotype of experimental animal individuals as input information, and uses individuals with both phenotype and genotype as a reference group to build a model. The model is used to predict the phenotypic values of individuals with unknown phenotypes, and the prediction results are output as the breeding values of the individuals.
3. The genome selection method based on machine learning according to claim 1 or 2, characterized in that The following steps are involved: S1: Data preprocessing, processing genotype and phenotype files separately: S11 Genotype data processing: Convert the input PLINK format file to a standardized NumPy matrix, process missing values in parallel blocks, fill in the mean, and output a `.npy` file for subsequent calculations; or digitize the genotypes, with the codes for genotypes AA, AB, and BB being 0, 1, and 2, respectively, and input the digitized results as a TXT file; S12 Phenotype data processing: extract the target trait column specified by `--phe-pos` from the input `.txt` phenotype file, use the linear model to correct the covariates, and generate the cleaned training set data; S2: Model configuration and training, select the model to be used for calculation and the configuration method of the model parameters, including two methods: pre-training based on input data and direct specification: If the `--model` parameter is specified, the input reference group data is used for pre-training. The scikit-learn model uses random search `n_iter` iterations to optimize parameters, and then uses three-fold cross-validation to build the model on the reference group data. The best performing parameter combination is ultimately selected as the final model hyperparameters. The deep learning model uses a fixed network structure, continuously updates parameters through gradient descent during training, and ultimately saves the optimal weights to a `.keras` file. If `--model-params` is specified, the hyperparameters specified by the user are used for training directly, and the parameter search step is skipped; S3: prediction and result output, including, S31. Phenotype prediction: Use the trained model to predict the phenotype of an individual with unknown phenotype and output the predicted value. Save the result to `{out}_predict.txt`. S32. Model saving: saves the scikit-learn model parameter file (`{out}_model.txt`) and the Keras model file (`.keras`), supporting direct input for subsequent calculations; S4: Extended functionality, using `--split-seed` to perform K-fold cross-validation on phenotypic data, generating `Ref{fold}.txt` training set and `Val{fold}.txt` validation set for evaluating the prediction model.
4. A genome selection device based on machine learning, comprising: IASML breeding software module and IASML breeding cloud platform module are characterized by: The IASML breeding software module includes: R1 supports reading and preprocessing of genotype data in PLINK binary format (bed / bim / fam) and TXT format, including missing value filling, genotype encoding conversion, and parallel block processing; R2 integrates a variety of machine learning models, including support vector machines (SVM), linear regression, Ridge regression, Lasso regression, Elastic net regression, PLS regression, decision trees, random forests, gradient boosting machines, LightGBM, XGBoost, convolutional neural networks (CNN), and multi-layer perceptrons (MLP), and provides random search hyperparameter optimization capabilities; R3 supports two training modes: hyperparameter search and training with specified parameters. The neural network model can save all parameters and use them directly for inference. R4 supports cross-validation segmentation of phenotypic data and generates training set and validation set files; The IASML breeding cloud platform module includes: Y1 provides a responsive graphical interface that allows users to upload PLINK or TXT format genotype files and TXT format phenotype files by dragging and dropping or selecting, and verifies file size and format in real time; Y2 dynamically generates phenotypic column selectors and interactive selection components for factor / numeric covariates, supporting multi-column collaborative filtering and missing value prompts; Y3 provides the function of uploading preset models, custom hyperparameter files, and pre-trained Keras models, and triggers the calculation process through a visual button; Y4 displays and analyzes log streams in real time, supports one-click downloading of prediction files and model files, and integrates error handling mechanisms and user notification systems; Y5 uses temporary working directories to isolate computing tasks of different users, automatically retains result files, and cleans up intermediate data to ensure system efficiency.
5. The genome selection device based on machine learning according to claim 4, characterized in that The IASML breeding cloud platform module includes a homepage, IASML software download and online usage pages, and its user interface and deployment architecture include: The interactive web-based interface includes dynamic navigation, mathematical formula rendering in KaTeX, and real-time log monitoring. It uses a front-end and back-end separation architecture, with the front-end implementing a responsive layout using HTML / CSS, and the back-end processing computing tasks based on a Python asynchronous framework. Supports one-click cloud deployment and localized private deployment.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the genome selection method according to any one of claims 1, 2, and 3 is implemented.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the genome selection method according to any one of claims 1, 2, and 3 is implemented.
8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the genome selection method according to any one of claims 1, 2, and 3 is implemented.
Citation Information
Patent Citations
Genome selective breeding method and system based on federal learning
CN118538294A
Cited By
Intelligent self-adaptive rasterization genome association analysis and molecular phenotype QTL data visualization method and visualization system thereof
CN122050502A