An intelligent selection and matching method, system, device and medium for sugarcane hybrid combinations
By predicting the phenotype of sugarcane hybrid combinations and evaluating their value, the problem of low selection efficiency of sugarcane hybrid combinations in the prior art is solved, and rapid and efficient hybrid combination selection is achieved.
Patent Information
- Application Number
- CN202411321893.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-09-23
AI Technical Summary
In the prior art, the selection of hybrid combinations during sugarcane hybrid breeding is usually dependent on the performance of offspring, resulting in a long time, high cost and low efficiency.
By acquiring multiple sugarcane parents, predicting the phenotype of each hybrid combination, evaluating the value of each hybrid combination based on the phenotype and preset value assessment rules, thus selecting the target hybrid combination.
This method can quickly select high-value sugarcane hybrid combinations without waiting for offspring to cultivate, reduce workload, save breeding years, and improve selection efficiency.
Smart Images

Figure CN119296642B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of bioinformatics and neural networks. Specifically, the present invention relates to a method, system, device and medium for intelligent selection and matching of sugarcane hybrid combinations. Background Art
[0002] In the prior art, for the selection of hybrid combinations in the process of cross-breeding, it is usually based on the performance of the offspring cultivated after cross-breeding to reversely evaluate the corresponding hybrid combinations. This method not only takes a long time, has high costs, but also has low efficiency. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a method, system, device and medium for intelligent selection and matching of sugarcane hybrid combinations, aiming to solve at least one of the above technical problems.
[0004] In a first aspect, the technical solution of the present invention to solve the above technical problem is as follows: A method for intelligent selection and matching of sugarcane hybrid combinations, the method comprising:
[0005] Obtaining a plurality of sugarcane parents;
[0006] Taking any two of the plurality of sugarcane parents as a hybrid combination, and predicting the phenotype corresponding to each hybrid combination according to the phenotypic traits of each sugarcane parent;
[0007] Evaluating the value of each hybrid combination according to the phenotype corresponding to each hybrid combination and a preset value evaluation rule;
[0008] Determining a target hybrid combination according to the values of each hybrid combination.
[0009] The beneficial effect of the present invention is that in the process of selecting and matching sugarcane hybrid combinations, first, in combination with the phenotypic traits of the hybrid combinations, the phenotypes of each possible hybrid combination are predicted, and then based on the phenotypes corresponding to each hybrid combination and a preset value evaluation rule, the value of each hybrid combination is evaluated, and then based on the values of each hybrid combination, the target hybrid combination is selected and matched, without waiting until the offspring of each hybrid combination are actually cultivated to select and match the target hybrid combination, thereby reducing the workload, saving the breeding period, and improving the selection and matching efficiency.
[0010] On the basis of the above technical solution, the present invention can also be improved as follows.
[0011] Further, for each hybrid combination, if the parents in the hybrid combination are homozygous parents, the value of the hybrid combination is higher than that of the hybrid combination composed of heterozygous parents.
[0012] Furthermore, the performance traits of each of the sugarcane parents are expressed by SNP gene data, and any two of the plurality of sugarcane parents are used as a hybrid combination, including:
[0013] Performing phenotypic identification on the SNP gene data of each of the sugarcane parents according to different phenotypic types;
[0014] According to the phenotypic identifiers corresponding to the sugarcane parents, any two sugarcane parents corresponding to the same phenotypic type are regarded as a hybrid combination, and any two sugarcane parents corresponding to different phenotypic types are regarded as a hybrid combination.
[0015] Furthermore, the prediction of the phenotype corresponding to each of the hybrid combinations according to the performance traits of each of the sugarcane parents is determined based on a phenotype prediction model, and the phenotype prediction model is trained based on the following method:
[0016] S1, obtaining SNP gene data and phenotypic data of multiple sugarcane sample parents, and preprocessing each of the SNP gene data and phenotypic data;
[0017] S2, for each of the SNP gene data, performing numerical mapping on the preprocessed SNP gene data to obtain a gene sequence after numerical mapping;
[0018] S3, for each of the numerically mapped gene sequences, using a fixed window to perform discrete Fourier transform on the numerically mapped gene sequence, calculating a power spectrum curve of each window sequence after the transform, determining whether each window is a protein coding region, and performing feature enhancement on the protein coding region according to the determination result to obtain a feature-enhanced gene sequence;
[0019] S4, for each gene sequence after feature enhancement, the gene sequence after feature enhancement is subjected to wavelet transform to obtain high-frequency features and low-frequency features, the high-frequency features are denoised, and the low-frequency features are subjected to inverse wavelet transform. The initial network is trained using the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels to obtain a phenotypic prediction model.
[0020] Furthermore, the size of the gene sequence after feature enhancement is 4*N. For each gene sequence after feature enhancement, the high-frequency features are denoised, the low-frequency features are inversely transformed by wavelet transform, and the initial network is trained using the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels to obtain a phenotypic prediction model, including:
[0021] Perform wavelet transform on the N-dimensional sequence of each channel to obtain high-frequency features and low-frequency features, and denoise the high-frequency features;
[0022] Perform inverse wavelet transform on the low-frequency features after transformation for each channel;
[0023] Input the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels into the initial network for phenotypic prediction training, and optimize the network model parameters through the loss functions Focal_L1 and Triplet_Contrast to obtain the phenotypic prediction model.
[0024] Furthermore, the method further includes:
[0025] Obtain the transcriptomes and metabolomes of multiple sugarcane sample parents;
[0026] For each of the sugarcane sample parents, use the SNP gene data, transcriptome, and metabolome of the sugarcane sample parent as training data;
[0027] Take the same phenotypes in all the preprocessed phenotypic data as the target phenotypes, calculate the first similarity between each piece of training data corresponding to each sugarcane sample parent corresponding to the target phenotype, and construct the graph structure data of each sugarcane sample parent using the first similarity;
[0028] The step of training the initial network with the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels to obtain the phenotypic prediction model includes:
[0029] Use the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels to perform the first training on the initial network to obtain the first predicted phenotypes;
[0030] Use the graph structure data of each sugarcane sample parent and the preprocessed phenotypic data to perform the second training on the initial network to obtain the second predicted phenotypes;
[0031] Calculate the first loss function value of the initial network based on each of the first predicted phenotypes and each of the preprocessed phenotypic data;
[0032] Calculate the second loss function value of the initial network based on each of the second predicted phenotypes and each of the preprocessed phenotypic data;
[0033] Train the initial network according to the first loss function value and the second loss function value to obtain the phenotypic prediction model.
[0034] Further, evaluating the value of each of the hybridization combinations according to the phenotypes corresponding to the respective hybridization combinations and a preset value evaluation rule includes:
[0035] For each of the hybridization combinations, evaluating the value of the hybridization combination according to the phenotype corresponding to the hybridization combination and a first correspondence relationship corresponding to the value evaluation rule, where the first correspondence relationship is the correspondence relationship between each different hybridization combination and the value of each of the hybridization combinations.
[0036] In a second aspect, the present invention also provides a smart selection and matching system for sugarcane hybridization combinations to solve the above technical problems. The system includes:
[0037] An acquisition module, configured to acquire a plurality of sugarcane parents;
[0038] A phenotype prediction module, configured to use any two of the plurality of sugarcane parents as a hybridization combination, and predict the phenotype corresponding to each hybridization combination according to the phenotypic traits of each sugarcane parent;
[0039] A value evaluation module, configured to evaluate the value of each hybridization combination according to the phenotypes corresponding to the respective hybridization combinations and a preset value evaluation rule;
[0040] A hybridization combination determination module, configured to determine a target hybridization combination according to the values of the respective hybridization combinations.
[0041] In a third aspect, the present invention also provides an electronic device to solve the above technical problems. The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the smart selection and matching method for sugarcane hybridization combinations of the present application is implemented.
[0042] In a fourth aspect, the present invention also provides a computer-readable storage medium to solve the above technical problems. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the smart selection and matching method for sugarcane hybridization combinations of the present application is implemented.
[0043] Additional aspects and advantages of the present application will be given in part in the following description, and these will become obvious from the following description, or can be understood through the practice of the present application. Description of the Drawings
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention.
[0045] Figure 1 It is a schematic flowchart of a smart selection and matching method for sugarcane hybridization combinations provided by an embodiment of the present invention;
[0046] Figure 2 Schematic diagram of the structure of an intelligent selection and matching system for sugarcane hybrid combinations provided by an embodiment of the present invention;
[0047] Figure 3 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Specific embodiments
[0048] The principles and features of the present invention are described below. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0049] The technical solutions of the present invention and how the technical solutions of the present invention solve the above technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other. For the same or similar concepts or processes, they may not be repeated in some embodiments. The embodiments of the present invention will be described below with reference to the drawings.
[0050] The solution provided by the embodiment of the present invention can be applied to any application scenario that requires intelligent selection and matching of sugarcane hybrid combinations. The solution provided by the embodiment of the present invention can be executed by any electronic device. For example, it can be the user's terminal device, including at least one of the following: smart phone, tablet computer, notebook computer, desktop computer, smart speaker, smart watch, smart TV, smart vehicle-mounted device.
[0051] The embodiment of the present invention provides a possible implementation manner, as Figure 1 shown, a flowchart of a method for intelligent selection and matching of sugarcane hybrid combinations is provided. This solution can be executed by any electronic device. For example, it can be a terminal device, or jointly executed by a terminal device and a server. For the convenience of description, the method provided by the embodiment of the present invention will be described below taking the terminal device as the execution subject as an example. As Figure 1 shown in the flowchart, the method may include the following steps:
[0052] Step S10, obtaining a plurality of sugarcane parents;
[0053] Step S20, taking any two of the plurality of sugarcane parents as a hybrid combination, and predicting the phenotype corresponding to each hybrid combination according to the phenotypic traits of each sugarcane parent;
[0054] Step S30, evaluating the value of each hybrid combination according to the phenotypes corresponding to each hybrid combination and a preset value evaluation rule;
[0055] Step S40, determining a target hybrid combination according to the values of each hybrid combination.
[0056] Through the method of the present invention, in the process of selecting and matching sugarcane hybrid combinations, first, in combination with the phenotypic traits of the hybrid combinations, predict the phenotypes of each possible hybrid combination, and then, based on the phenotypes corresponding to each of the hybrid combinations and a preset value evaluation rule, evaluate the value of each of the hybrid combinations. Furthermore, based on the values of the hybrid combinations, select and match the target hybrid combination, without waiting until the offspring of each hybrid combination are actually cultivated to select and match the target hybrid combination, thereby reducing the workload, saving the breeding years, and improving the selection and matching efficiency.
[0057] The following further illustrates the solution of the present invention in combination with the following specific embodiments. In this embodiment, the intelligent selection and matching method for sugarcane hybrid combinations may include the following steps:
[0058] Step S10, obtain multiple sugarcane parents;
[0059] Among them, multiple sugarcane parents refer to the data basis for selecting the target hybrid combination, that is, select the sugarcane parents that can form the target hybrid combination from multiple sugarcane parents. Multiple sugarcane parents may include homozygous parents and heterozygous parents. Multiple sugarcane parents may include parents with the same phenotypic traits, or may also include parents with different phenotypic traits. The phenotypic traits of each sugarcane parent represent the characteristics of the corresponding sugarcane parent and can be used as a reference for the subsequent target hybrid combination.
[0060] Among them, multiple sugarcane parents can be provided by the MS ACCESS platform.
[0061] Step S20, take any two of the multiple sugarcane parents as a hybrid combination, and according to the phenotypic traits of each of the sugarcane parents, predict the phenotype corresponding to each of the hybrid combinations;
[0062] Among them, for each hybrid combination, the hybrid combination may include parents with the same phenotypic traits or may also include parents with different phenotypic traits. In this solution, cover as many hybrid combinations formed by various phenotypic traits as possible.
[0063] Among them, for each of the hybrid combinations, if the parents in the hybrid combination are homozygous parents, the value of the hybrid combination is higher than that of the hybrid combination composed of heterozygous parents.
[0064] Among them, if the phenotypic traits of each of the sugarcane parents are represented by SNP gene data, then in the above S20, taking any two of the multiple sugarcane parents as a hybrid combination includes:
[0065] S201, label the SNP gene data of each of the sugarcane parents according to different phenotypic types;
[0066] Among them, the SNP genotype data can be identified according to different phenotypic types. As an example, for instance, the SNP gene data of different phenotypic types can be identified based on 0, 1, and 2. Among them, 0 represents the SNP gene data with a high homozygote frequency, 1 represents the SNP gene data corresponding to the heterozygote, and 2 represents the SNP gene data with a low homozygote frequency.
[0067] Among them, the SNP gene data can be determined based on genomic-BLUP analysis
[0068] S202. According to the phenotypic identifiers corresponding to each sugarcane parent, any two sugarcane parents corresponding to the same phenotypic type are used as a hybridization combination, and any two sugarcane parents corresponding to different phenotypic types are used as a hybridization combination.
[0069] Optionally, in the above S20, according to the performance traits of each sugarcane parent, predicting the phenotype corresponding to each hybridization combination is determined based on a phenotype prediction model, and the phenotype prediction model is trained based on the following method:
[0070] S1. Obtain the SNP gene data and phenotype data of multiple sugarcane sample parents, and preprocess each SNP gene data and phenotype data;
[0071] Among them, since there are quality problems in the gene sequences of each sugarcane sample parent due to sequencing technology, the SNP gene data and phenotype data are preprocessed: filter gene loci according to the gene locus deletion rate and the minor allele frequency; perform unknown filtering on the phenotype data to remove data with unknown phenotype values.
[0072] S2. For each SNP gene data, perform numerical mapping on the preprocessed SNP gene data to obtain the gene sequence after numerical mapping;
[0073] Among them, the gene length in the SNP gene data of each sugarcane sample parent is Mbp, and numerical mapping is performed on the SNP gene data of each sugarcane sample parent. Let I = {A, T, C, G}, and the gene sequence corresponding to the SNP gene data of each sugarcane sample parent can be expressed as: S = {S(n)|S(n) ∈ I, n = 0, 1, 2…N - 1}, and numerical mapping is performed on each gene sequence in the following way:
[0074]
[0075] Among them, b are respectively A, G, C, T; S(n) represents the genotype at the gene sequence position n, and t b [n] is the result of numerical mapping of the genotype at the gene sequence position n of b. Thus, through the above encoding method, each sugarcane sample parent will generate 4 numerical mapping sequences.
[0076] S3. For each of the numerically mapped gene sequences, perform discrete Fourier transform on the numerically mapped gene sequence using a fixed window, calculate the power spectral curve of each window sequence after the transform, determine whether each window is a protein-coding region, and enhance the features of the protein-coding region according to the determination result to obtain the gene sequence with enhanced features.
[0077] One realizable way of the above S3 is as follows:
[0078] Perform discrete Fourier transform on the numerically mapped gene sequence encoding using a sliding window with window size W and step size L, and calculate the power spectral curve of each window sequence after the transform:
[0079] For the power spectral curve of each window, if the spectral peak at W / 3 and the spectral mean of the entire window are higher than the threshold, then consider this window as a protein-coding region, otherwise it is a non-coding region, and record the starting position of the protein-coding region.
[0080] Enhance the features of the obtained protein-coding region in the following way:
[0081]
[0082] Among them, P(W / 3) is the spectral value of the coding region at W / 3. Use a sliding window with window size W and step size L to merge the coding of each window after feature enhancement, that is, for the same gene position, use the weighted mean of the coding at this position as the final coding at this position. Among them, the non-coding region still maintains the numerical mapping of S2.
[0083] Among them, the value range of W is from 200bp to 1000bp, and the value range of L is from 100bp to 300bp. Find the optimal values in the two intervals.
[0084] S4. For each of the gene sequences with enhanced features, perform wavelet transform on the gene sequence with enhanced features to obtain high-frequency features and low-frequency features, denoise the high-frequency features, perform inverse wavelet transform on the low-frequency features, and use the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels to train the initial network to obtain a phenotypic prediction model.
[0085] Optionally, the initial network can be a three-stream network. Each branch of the three-stream network includes three convolutional layers, two LSTM layers, and one fully connected layer. After each convolutional layer, it is processed by a batch normalization layer and a ReLU layer.
[0086] Optionally, the size of the gene sequence after the above feature enhancement is 4*N. For each of the gene sequences after the feature enhancement, in S4 above, the high-frequency features are denoised, the low-frequency features are subjected to inverse wavelet transform, and the initial network is trained using the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels to obtain a phenotypic prediction model, including:
[0087] S41, perform wavelet transform on the N-dimensional sequence of each channel to obtain high-frequency features and low-frequency features F1, and denoise the high-frequency features to obtain denoised high-frequency features F2;
[0088] S42, perform inverse wavelet transform on the transformed low-frequency features of each channel to obtain low-frequency features F3 after inverse wavelet transform;
[0089] S43, input the low-frequency features F1, the denoised high-frequency features F2, the low-frequency features F3 after inverse wavelet transform, and the preprocessed phenotypic data as labels into the initial network for phenotypic prediction training, and optimize the network model parameters through the loss functions Focal_L1 and Triplet_Contrast to obtain a phenotypic prediction model.
[0090] Optionally, the method further includes:
[0091] Obtain the transcriptomes and metabolomes of multiple sugarcane sample parents;
[0092] For each of the sugarcane sample parents, use the SNP gene data, transcriptome, and metabolome of the sugarcane sample parent as training data, and each training data includes SNP gene data, transcriptome, and metabolome;
[0093] Take the same phenotypes in all the preprocessed phenotypic data as the target phenotypes, calculate the first similarity between each training data corresponding to each sugarcane sample parent corresponding to the target phenotypes, and construct graph structure data for each sugarcane sample parent using the first similarity;
[0094] Among them, in the graph structure data of each sugarcane sample parent, the nodes of the graph represent different omics (SNP gene data, transcriptome or metabolome), and the edges of the graph represent the first similarity between different omics.
[0095] It should be understood that the weight of the edge reflects the degree of association between different omics, that is, the greater the first similarity, the higher the degree of association between different omics.
[0096] In S4 above, the initial network is trained using the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels to obtain a phenotypic prediction model, including:
[0097] Using the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels, the initial network is trained for the first time to obtain the first predicted phenotype;
[0098] Using the graph structure data of each of the sugarcane sample parents and the preprocessed phenotypic data, the initial network is trained for the second time to obtain the second predicted phenotype;
[0099] According to each of the first predicted phenotypes and each of the preprocessed phenotypic data, the first loss function value of the initial network is calculated;
[0100] According to each of the second predicted phenotypes and each of the preprocessed phenotypic data, the second loss function value of the initial network is calculated;
[0101] According to the first loss function value and the second loss function value, the initial network is trained to obtain the phenotypic prediction model.
[0102] Step S30: According to the phenotypes corresponding to each of the hybridization combinations and a preset value evaluation rule, evaluate the value of each of the hybridization combinations;
[0103] Among them, the above value evaluation rule can be pre-configured based on actual needs, and the value can be represented numerically. The larger the value, the larger the corresponding numerical value.
[0104] Optionally, in the above S30, evaluating the value of each of the hybridization combinations according to the phenotypes corresponding to each of the hybridization combinations and a preset value evaluation rule includes:
[0105] For each of the hybridization combinations, evaluate the value of the hybridization combination according to the phenotype corresponding to the hybridization combination and the first correspondence corresponding to the value evaluation rule, where the first correspondence is the correspondence between each different hybridization combination and the value of each of the hybridization combinations.
[0106] Among them, the above first correspondence is preset, which describes the correspondence between different hybridization combinations and the values of different hybridization combinations. That is, hybridization combinations composed of different parents may correspond to different values. The above first correspondence can be determined based on historical hybridization data, that is, from historical hybridization combinations, determine the values corresponding to different hybridization combinations.
[0107] Step S40: Determine the target hybridization combination according to the values of each of the hybridization combinations.
[0108] Among them, one implementable manner of the above S40 is:
[0109] From the values of each hybridization combination, select the hybridization combination with the highest value as the target hybridization combination.
[0110] Optionally, in the solution of this application, if the initial network is a three-stream network, the three-stream network includes three branches, each of the branches includes three convolutional layers, two LSTM layers and one fully connected layer, and each convolutional layer is processed by a batch normalization layer and a ReLU layer.
[0111] Optionally, a feasible implementation of using the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels to perform the first training on the initial network to obtain the first predicted phenotype is as follows:
[0112] Input the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels into the trained three-stream network according to a set mode, and output the first predicted phenotype.
[0113] Among them, different modes may include:
[0114] Mode 1: Only select one of the low-frequency features, the denoised high-frequency features, and the low-frequency features after inverse wavelet transform and input it into the corresponding branch network to obtain the first predicted phenotype;
[0115] Mode 2: Input the low-frequency features and the denoised high-frequency features into the corresponding branch networks respectively, take the average of the two obtained phenotypic values to obtain the first predicted phenotype;
[0116] Mode 3: Input the denoised high-frequency features and the low-frequency features after inverse wavelet transform into the corresponding branch networks respectively, take the average of the two obtained phenotypic values to obtain the first predicted phenotype;
[0117] Mode 4: Input all of the low-frequency features, the denoised high-frequency features, and the low-frequency features after inverse wavelet transform into the corresponding branch networks, take the average of the three predicted phenotypic values to obtain the first predicted phenotype.
[0118] Optionally, the method further includes:
[0119] Divide multiple sugarcane sample parents into test samples and training samples;
[0120] Select multiple training models, and the training models at least include the GBLUP model and the RRBLUP model; the input of the training models is the training data of multiple sugarcane sample parents, and the output is the actual phenotypic value;
[0121] Train the training models with the training data of all the training samples and the corresponding actual phenotypic values to obtain multiple optimized models;
[0122] Input the training data of all the test samples into each of the optimized models respectively to obtain a third predicted phenotype corresponding to each optimized model and corresponding to the first characteristic parameter, where the first characteristic parameter is the characteristic parameter corresponding to one sugarcane sample parent among the same sugarcane sample parents of all sugarcane sample parents;
[0123] Calculate the correlation values corresponding to each optimized model based on the third predicted phenotypes of all the test samples corresponding to each optimized model and the actual phenotype values;
[0124] Determine the third loss function value according to each of the correlation values;
[0125] The training of the initial network according to the first loss function value and the second loss function value to obtain the phenotype prediction model includes:
[0126] Train the initial network according to the first loss function value, the second loss function value and the third loss function value to obtain the phenotype prediction model.
[0127] Among them, during the training process, the aim is to minimize the correlation value corresponding to the third loss function value.
[0128] In this solution, genetic methods such as G-BLUP analysis, genomic-BLUP analysis and MiXBLUP are used to conduct phenotype prediction and value evaluation on common sugarcane breeding parents and combinations, so as to realize the prediction of the field performance of various unconfigured hybrid combinations.
[0129] In this solution, the true phenotypes of the parents can also be encoded on the barcodes, that is, one barcode corresponds to one parent.
[0130] In this solution, based on the phenotype prediction and value evaluation of common sugarcane breeding parents and hybrid combinations, a database management system for sugarcane hybrid combination selection and matching is established. The computer is enabled to automatically match high-value combinations according to the hybridization plan.
[0131] In this solution, based on the database for sugarcane hybrid combination selection and matching, a suitable machine learning model (phenotype prediction model) is constructed, and a fully automatic phenotype prediction and value evaluation process for parents and hybrid offspring is developed using the python programming language.
[0132] For the combinations configured in the intelligent sugarcane hybrid combination selection and matching software, combined with the previous phenotype data, use the mixed linear model to obtain the phenotype prediction values of the traits. For the combinations with good prediction effects, directly use the phenotype prediction to judge the performance of the offspring, so as to reduce the workload and save the breeding years.
[0133] Based on and Figure 1Based on the same principle as the method shown in [reference], an embodiment of the present invention also provides a smart selection and matching system 20 for sugarcane hybrid combinations, as shown in Figure 2 As shown in [reference], the smart selection and matching system 20 for sugarcane hybrid combinations may include an acquisition module 210, a phenotype prediction module 220, a value evaluation module 230, and a hybrid combination determination module 240, where:
[0134] The acquisition module 210 is used to acquire a plurality of sugarcane parents;
[0135] The phenotype prediction module 220 is used to take any two of the plurality of sugarcane parents as a hybrid combination, and predict the phenotype corresponding to each hybrid combination according to the phenotypic traits of each sugarcane parent;
[0136] The value evaluation module 230 is used to evaluate the value of each hybrid combination according to the phenotype corresponding to each hybrid combination and a preset value evaluation rule;
[0137] The hybrid combination determination module 240 is used to determine a target hybrid combination according to the value of each hybrid combination.
[0138] Optionally, for each hybrid combination, if the parents in the hybrid combination are homozygous parents, the value of the hybrid combination is higher than that of the hybrid combination composed of heterozygous parents.
[0139] Optionally, the phenotypic traits of each sugarcane parent are represented by SNP gene data. When the phenotype prediction module 220 takes any two of the plurality of sugarcane parents as a hybrid combination, it is specifically used for:
[0140] Label the SNP gene data of each sugarcane parent according to different phenotypic types;
[0141] According to the phenotypic labels corresponding to each sugarcane parent, take any two sugarcane parents corresponding to the same phenotypic type as a hybrid combination, and take any two sugarcane parents corresponding to different phenotypic types as a hybrid combination.
[0142] Optionally, the prediction of the phenotype corresponding to each hybrid combination based on the phenotypic traits of each sugarcane parent is determined based on a phenotype prediction model, and the phenotype prediction model is trained by the following model training module. The model training module is used for:
[0143] Acquire the SNP gene data and phenotypic data of a plurality of sugarcane sample parents, and preprocess each SNP gene data and phenotypic data;
[0144] For each of the SNP gene data, perform numerical mapping on the preprocessed SNP gene data to obtain a gene sequence after numerical mapping;
[0145] For each of the gene sequences after numerical mapping, perform discrete Fourier transform on the gene sequences after numerical mapping using a fixed window, calculate the power spectrum curve of each window sequence after transformation, determine whether each window is a protein coding region, and perform feature enhancement on the protein coding region according to the determination result to obtain a gene sequence after feature enhancement;
[0146] For each of the gene sequences after feature enhancement, perform wavelet transform on the gene sequences after feature enhancement to obtain high-frequency features and low-frequency features, denoise the high-frequency features, perform inverse wavelet transform on the low-frequency features, and use the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels to train an initial network to obtain a phenotypic prediction model.
[0147] Optionally, the size of the gene sequence after feature enhancement is 4*N. For each of the gene sequences after feature enhancement, when the model training module denoises the high-frequency features, performs inverse wavelet transform on the low-frequency features, and uses the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels to train an initial network to obtain a phenotypic prediction model, it is specifically used for:
[0148] Perform wavelet transform on the N-dimensional sequence of each channel to obtain high-frequency features and low-frequency features, and denoise the high-frequency features;
[0149] Perform inverse wavelet transform on the low-frequency features after transformation of each channel;
[0150] Input the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels into the initial network for phenotypic prediction training, and optimize the network model parameters through the loss functions Focal_L1 and Triplet_Contrast to obtain a phenotypic prediction model.
[0151] Optionally, the system further includes:
[0152] A graph structure data determination module, which is used for:
[0153] Obtain the transcriptomes and metabolomes of multiple sugarcane sample parents;
[0154] For each of the sugarcane sample parents, use the SNP gene data, transcriptome, and metabolome of the sugarcane sample parent as training data;
[0155] Take the same phenotypes in all pre - processed phenotypic data as the target phenotypes, calculate the first similarity between each piece of training data corresponding to the parents of each sugarcane sample for the target phenotypes, and construct graph structure data for each sugarcane sample parent using the first similarity;
[0156] When the above - mentioned model training module uses low - frequency features, denoised high - frequency features, low - frequency features after wavelet inverse transform, and pre - processed phenotypic data as labels to train the initial network to obtain a phenotypic prediction model, it is specifically used for:
[0157] Use low - frequency features, denoised high - frequency features, low - frequency features after wavelet inverse transform, and pre - processed phenotypic data as labels to perform the first training on the initial network to obtain the first predicted phenotypes;
[0158] Use the graph structure data of each sugarcane sample parent and the pre - processed phenotypic data to perform the second training on the initial network to obtain the second predicted phenotypes;
[0159] Calculate the first loss function value of the initial network according to each of the first predicted phenotypes and each of the pre - processed phenotypic data;
[0160] Calculate the second loss function value of the initial network according to each of the second predicted phenotypes and each of the pre - processed phenotypic data;
[0161] Train the initial network according to the first loss function value and the second loss function value to obtain the phenotypic prediction model.
[0162] Optionally, when the value evaluation module 230 evaluates the value of each hybridization combination according to the phenotypes corresponding to each hybridization combination and a preset value evaluation rule, it is specifically used for:
[0163] For each hybridization combination, evaluate the value of the hybridization combination according to the phenotype corresponding to the hybridization combination and the first correspondence relationship corresponding to the value evaluation rule, where the first correspondence relationship is the correspondence relationship between each different hybridization combination and the value of each hybridization combination.
[0164] The sugarcane hybridization combination intelligent selection and matching system of the embodiments of the present invention can execute the sugarcane hybridization combination intelligent selection and matching method provided by the embodiments of the present invention, and its implementation principle is similar. The actions performed by each module and unit in the sugarcane hybridization combination intelligent selection and matching system in each embodiment of the present invention correspond to the steps in the sugarcane hybridization combination intelligent selection and matching method in each embodiment of the present invention. For the detailed function descriptions of each module of the sugarcane hybridization combination intelligent selection and matching system, reference can specifically be made to the descriptions in the corresponding sugarcane hybridization combination intelligent selection and matching method shown above, and details are not described herein again.
[0165] Among them, the above-mentioned intelligent selection and matching system for sugarcane hybrid combinations can be a computer program (including program code) running on a computer device. For example, the intelligent selection and matching system for sugarcane hybrid combinations is an application software; this system can be used to execute the corresponding steps in the method provided by the embodiments of the present invention.
[0166] In some embodiments, the intelligent selection and matching system for sugarcane hybrid combinations provided by the embodiments of the present invention can be implemented in a combination of software and hardware. As an example, the intelligent selection and matching system for sugarcane hybrid combinations provided by the embodiments of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the intelligent selection and matching method for sugarcane hybrid combinations provided by the embodiments of the present invention. For example, the processor in the form of a hardware decoding processor can employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0167] In other embodiments, the intelligent selection and matching system for sugarcane hybrid combinations provided by the embodiments of the present invention can be implemented in a software manner. Figure 2 The intelligent selection and matching system for sugarcane hybrid combinations stored in the memory is shown. It can be software in the form of programs and plugins, etc., and includes a series of modules, including an acquisition module 210, a phenotype prediction module 220, a value evaluation module 230, and a hybrid combination determination module 240, which are used to implement the intelligent selection and matching method for sugarcane hybrid combinations provided by the embodiments of the present invention.
[0168] The modules involved in the embodiments described in the present invention can be implemented in a software manner or in a hardware manner. Among them, the name of the module does not constitute a limitation to the module itself in some cases.
[0169] Based on the same principle as the method shown in the embodiments of the present invention, an electronic device is also provided in the embodiments of the present invention. The electronic device may include, but is not limited to: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the method shown in any embodiment of the present invention by calling the computer program.
[0170] In an alternative embodiment, an electronic device is provided, as Figure 3 shown. Figure 3The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as being connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between this electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present invention.
[0171] The processor 4001 can be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in connection with the disclosure of the present invention. The processor 4001 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0172] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 can be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard structure) bus, etc. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 3 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0173] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0174] The memory 4003 is used to store the application program code (computer program) for implementing the solution of the present invention and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the application program code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.
[0175] Among them, the electronic device can also be a terminal device. Figure 3 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0176] The embodiments of the present invention provide a computer-readable storage medium, on which a computer program is stored. When the computer program runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments.
[0177] According to another aspect of the present invention, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above various implementation manners of the embodiments.
[0178] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0179] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the methods and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0180] The computer-readable storage medium provided by the embodiments of the present invention may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0181] The above computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to execute the method shown in the above embodiments.
[0182] The above description is only a preferred embodiment of the present invention and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of disclosure involved in the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present invention.
Claims
1. A method for intelligent selection of sugarcane hybrid combinations, characterized in that: The following steps are involved: Obtain multiple sugarcane parents; Taking any two sugarcane parents from the plurality of sugarcane parents as a hybrid combination, and predicting the phenotype corresponding to each of the hybrid combinations according to the performance traits of each of the sugarcane parents; Evaluate the value of each of the hybrid combinations according to the phenotypes corresponding to each of the hybrid combinations and preset value evaluation rules; Determining a target hybrid combination according to the value of each of the hybrid combinations; The prediction of the phenotype corresponding to each of the hybrid combinations according to the performance traits of each of the sugarcane parents is determined based on a phenotype prediction model, and the phenotype prediction model is trained based on the following method: S1, obtaining SNP gene data and phenotypic data of multiple sugarcane sample parents, and preprocessing each of the SNP gene data and phenotypic data; S2, for each of the SNP gene data, performing numerical mapping on the preprocessed SNP gene data to obtain a gene sequence after numerical mapping; S3, for each of the numerically mapped gene sequences, using a fixed window to perform discrete Fourier transform on the numerically mapped gene sequence, calculating a power spectrum curve of each window sequence after the transform, determining whether each window is a protein coding region, and performing feature enhancement on the protein coding region according to the determination result to obtain a feature-enhanced gene sequence; S4, for each of the gene sequences after feature enhancement, performing wavelet transform on the gene sequence after feature enhancement to obtain high-frequency features and low-frequency features, denoising the high-frequency features, performing inverse wavelet transform on the low-frequency features, and using the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the preprocessed phenotypic data as labels to perform a first training on the initial network to obtain a first predicted phenotype; S5, obtain transcriptomes and metabolomes of multiple sugarcane sample parents; S6, for each of the sugarcane sample parents, using the SNP gene data, transcriptome and metabolome of the sugarcane sample parent as training data; S7, taking the same phenotype in all the preprocessed phenotypic data as the target phenotype, calculating the first similarity between the training data corresponding to each sugarcane sample parent corresponding to the target phenotype, and constructing the graph structure data of each sugarcane sample parent using the first similarity; S8, using the graph structure data and pre-processed phenotypic data of each of the sugarcane sample parents, training the initial network for a second time to obtain a second predicted phenotype; S9, calculating a first loss function value of the initial network according to each of the first predicted phenotypes and each of the preprocessed phenotype data; S10, calculating a second loss function value of the initial network according to each of the second predicted phenotypes and each of the preprocessed phenotype data; S11, dividing multiple sugarcane sample parents into test samples and training samples; S12, selecting a plurality of training models, wherein the training models at least include a GBLUP model and an RRBLUP model; the input of the training model is the training data of a plurality of sugarcane sample parents, and the output is the actual phenotypic value; S13, training the training model using the training data of all the training samples and the corresponding actual phenotypic values to obtain multiple optimization models; S14, inputting the training data of all the test samples into each of the optimization models respectively, and obtaining a third predicted phenotype corresponding to each of the optimization models and corresponding to the first characteristic parameter, wherein the first characteristic parameter is a characteristic parameter corresponding to one of the same sugarcane sample parents among all the sugarcane sample parents; S15, calculating the correlation value corresponding to each optimization model according to the third predicted phenotype of all the test samples corresponding to each optimization model and the actual phenotype value; determining the third loss function value according to each correlation value; S16, training the initial network according to the first loss function value, the second loss function value and the third loss function value to obtain the phenotype prediction model.
2. The method according to claim 1, characterized in that For each of the hybrid combinations, if the parents in the hybrid combination are homozygous parents, the value of the hybrid combination is higher than the value of the hybrid combination whose parents are heterozygous parents.
3. The method according to claim 1, characterized in that The performance traits of each sugarcane parent are expressed by SNP gene data, and any two sugarcane parents among the plurality of sugarcane parents are used as a hybrid combination, including: Performing phenotypic identification on the SNP gene data of each of the sugarcane parents according to different phenotypic types; According to the phenotypic identifiers corresponding to the sugarcane parents, any two sugarcane parents corresponding to the same phenotypic type are regarded as a hybrid combination, and any two sugarcane parents corresponding to different phenotypic types are regarded as a hybrid combination.
4. The method according to claim 1, characterized in that: The size of the gene sequence after feature enhancement is 4*N. For each gene sequence after feature enhancement, the high-frequency features are denoised, the low-frequency features are inversely transformed by wavelet transform, and the initial network is trained using the low-frequency features, the denoised high-frequency features, the low-frequency features after inverse wavelet transform, and the pre-processed phenotypic data as labels to obtain a phenotypic prediction model, including: Perform wavelet transform on the N-dimensional sequence of each channel to obtain high-frequency features and low-frequency features, and denoise the high-frequency features; Perform inverse wavelet transform on the transformed low-frequency features of each channel; The low-frequency features, high-frequency features after denoising, low-frequency features after inverse wavelet transform, and preprocessed phenotypic data as labels were input into the initial network for phenotypic prediction training, and the network model parameters were optimized by the loss functions Focal_L1 and Triplet_Contrast to obtain the phenotypic prediction model.
5. The method according to any one of claims 1 to 3, characterized in that The step of evaluating the value of each hybrid combination according to the phenotype corresponding to each hybrid combination and a preset value evaluation rule includes: For each of the hybrid combinations, the value of the hybrid combination is evaluated according to a first correspondence between a phenotype corresponding to the hybrid combination and the value evaluation rule, wherein the first correspondence is a correspondence between different hybrid combinations and the values of each of the hybrid combinations.
6. A sugarcane hybrid combination intelligent selection system, characterized in that: The method for intelligent selection of sugarcane hybrid combinations according to claim 1 comprises: An acquisition module, used for acquiring multiple sugarcane parents; A phenotype prediction module, for taking any two sugarcane parents from the plurality of sugarcane parents as a hybrid combination, and predicting the phenotype corresponding to each of the hybrid combinations according to the performance traits of each of the sugarcane parents; A value assessment module, used to assess the value of each of the hybrid combinations according to the phenotypes corresponding to each of the hybrid combinations and preset value assessment rules; The hybrid combination determination module is used to determine the target hybrid combination according to the value of each hybrid combination.
7. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Hybrid seed prediction method based on Bayesian model integrating parent phenotypes
CN113053459A
Phenotype prediction method and device based on frequency domain transform enhancement
CN117174161A