Corn backcross breeding optimization design method and system, application and related products
By simulating the genotypes of maize backcross breeding and optimizing the breeding route, and utilizing Poisson distribution and bin genotype identification, the problem of balancing breeding cycle and cost was solved, achieving rapid and efficient screening of breeding target genes and high background recovery rate, thus improving breeding efficiency.
Patent Information
- Application Number
- CN202511339556.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2026-01-02
AI Technical Summary
In the process of maize backcross breeding, how to balance the breeding cycle and cost, especially how to optimize the number of backcross generations and the number of samples to quickly and economically screen for target genes and achieve a high background recovery rate under the needs of different breeders.
By simulating the genotypes of backcross progeny and using the Poisson distribution hypothesis regarding recombination times, the population size under different success probabilities and background reversion rates is calculated. Combined with the bin genotype identification method, the breeding route is optimized to provide the optimal number of backcross generations and population size.
This approach enables rapid screening of target genes with high background recovery rates while taking into account both cost and time, thus saving resources, improving breeding efficiency, and shortening the breeding cycle.
Smart Images

Figure CN121241906A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of maize backcross breeding, in particular to a maize backcross breeding optimization design method and system, application and related products. BACKGROUND
[0002] Maize (Zea mays L.) as the world's major food crops, plays a pivotal role in the world's agricultural development and food security. In 2012, the maize yield has exceeded the rice yield, and has jumped to the first food crop in China. However, there is still a big gap between China's maize yield and the United States. Therefore, in a short time, breeding high-yield maize varieties is the goal pursued by breeders. Backcross breeding is a key link in the process of maize breeding.
[0003] Backcross breeding is a method used by breeders to improve individual traits of varieties by backcrossing, which belongs to the category of hybrid breeding. Backcross breeding refers to the method of introducing other traits of excellent parents lacking specific excellent traits into another parent with specific excellent traits through backcrossing. The recipient of the target trait is called the recurrent parent or the recipient parent, and the provider of the target trait is called the non-recurrent parent or the donor parent. Backcross breeding is mainly used to improve the undesirable traits of excellent varieties, overcome the linkage of favorable genes and unfavorable genes, aggregate the favorable traits of both parents, and provide materials for genetic research.
[0004] In the process of multiple generations of backcrossing, it is a time-consuming and laborious process to screen materials that meet the purpose of gene introduction and have high background recovery rate from a large number of materials. However, the needs of different breeders are different, for example; some breeders want to shorten the breeding cycle, and some breeders want to save costs, that is, to achieve the purpose of breeding with fewer samples, the sample detection cost generated by fewer samples is relatively less, and the cost can be saved. The trend is to achieve the same breeding purpose, the more the number of backcrossing generations, the fewer the number of samples needed, and the lower the cost; the fewer the number of backcrossing generations, the larger the sample size needed, resulting in higher cost. How to balance the breeding cycle and cost is a problem that breeders need to consider.
[0005] Therefore, it is urgent to design a maize backcross breeding optimization design scheme, which can give an optimal breeding route according to the actual needs of breeders. SUMMARY
[0006] The purpose of the present application is to provide a maize backcross breeding optimization design method and system, application and related products, which can simulate the genotype of backcross offspring, judge the number of backcross generations and the corresponding sample size of different routes needed to achieve the breeding purpose of breeders (introducing target genes and achieving the desired value of background recovery rate), and provide reference and suggestion for the breeding selection of breeders.
[0007] The application is achieved by the following technical solutions:
[0008] A corn backcross breeding optimization design method, comprising the following steps:
[0009] S1, obtaining genotype data of a receptor parent and a donor parent;
[0010] S2, according to the genotype data of the parents and the genetic map, assuming that the recombination number obeys Poisson distribution, the genotype of different individuals of the BC1F1 generation is simulated;
[0011] S3, using the simulation method of step S2, based on N times of simulation of different BC1F1 generations, obtaining the probability value of the BC1F1 generation that meets the target gene introduction and the background recovery rate is greater than a certain level According to the size of the population scale of N times of simulated different BC1F1 generations, a success probability P of the target material is obtained S , the probability value And the relationship of the BC1F1 generation, wherein N is a positive integer greater than or equal to 20;
[0012] S4, the relationship formula obtained in step S3 is used to calculate different success probabilities P S The population scale n of the BC1F1 generation that meets the target gene introduction and the background recovery rate is greater than a certain level appears;
[0013] S5, based on the calculation result of step S4, under the condition of different success probabilities P S , find out how large the population scale n of the BC1F1 generation is required under the condition that the number of materials m meets the conditions;
[0014] S6, based on the genotype data of the backcross offspring and the receptor parent, referring to steps S2-S5, the population scale n of the next generation of backcross offspring is calculated;
[0015] S7, under certain cost conditions, when the target gene introduction and the background recovery rate reach a certain value, the required backcross generation number and the optimal route of the corresponding population scale n are given.
[0016] Firstly, according to the genotype data of the parents and the genetic map, based on the assumption that the recombination number of chromosomes obeys Poisson distribution, the genotype of different individuals of the BC1F1 generation can be simulated, so that the genotype data of the backcross offspring can be obtained without actual breeding process, which provides a prerequisite for providing different breeding routes for corn backcross breeding before breeding for breeders to refer to.
[0017] Secondly, the application is based on N different real simulations (real means that the data used is derived from a certain specific backcross improved population data; and simulation means that based on the recombination exchange rule of whole genome, the possible situation is calculated by using statistical deduction, because the real recombination exchange only occurs once), to obtain a success probability P of the target material S , the probability value and BC1F1 generation relationship formula, that is, the calculation formula can calculate the sample number of BC1F1 generation required for different success probability and different background recovery rate, and then the population size required for finding the number of materials meeting the conditions under different success probability can be calculated.
[0018] Based on the above calculation results, according to the different backcross generations and the population size n corresponding to different routes required by the actual needs of the applicant (the specific number of materials meeting the conditions m, when the target gene introduction and background recovery rate reach a certain value), the cost of genotype detection in the actual breeding process can be calculated based on the population size n. Breeders can choose breeding routes according to the cost and the corresponding time of backcross generations.
[0019] In a preferred mode, the genotype of the backcross population is corrected by using the bin genotype identification method; the bin genotype identification method comprises the following steps:
[0020] Step 1: The genotype of each chromosome of each sample of the backcross population is coded and the sample is screened, the coded genotype value is divided into different fragments, and the length of each fragment is calculated;
[0021] Step 2: Starting from the longest fragment, traverse upwards and downwards, and smooth the small fragments;
[0022] Step 3: Three fragments are judged together, the middle fragment is determined by the fixed determination of the two fragments, so as to realize the bin genotype identification of 10 pairs of chromosomes of the whole genome.
[0023] The bin genotype identification method designed in the application can correct and optimize the genotype of the backcross population, and can ensure higher accuracy of the population size n calculation.
[0024] Specifically, in step 1, the coding rule is to distinguish gametes from the donor parent, gametes from the recipient parent and genotype defects by using different numbers, and the genotype defects include errors or missing genotypes.
[0025] Specifically, in step 1, the sample screening process is:
[0026] Remove samples with single chromosome deletion rate greater than or equal to 0.5; remove samples with whole genome genotype deletion rate greater than or equal to 0.5; and correct the remaining samples to the coding of the locus from the gamete of the donor parent or the gamete of the recipient parent.
[0027] Specifically, the specific process of encoding and sample screening is as follows:
[0028] Step 1): Encode the genotype of the backcross population, with 1 indicating that the gamete comes from the non-recurrent parent, 2 indicating that the gamete comes from the recurrent parent, 3 indicating that the judgment is wrong or the genotype is missing, hereinafter referred to as missing;
[0029] Step 2): Remove samples with single chromosome deletion rate greater than or equal to 0.5; remove samples with whole genome (10 pairs of chromosomes) genotype deletion rate greater than or equal to 0.5;
[0030] Step 3): Correct the genotype value of the remaining samples to 3 to 1 or 2, and perform bin genotype identification at the same time;
[0031] Specifically, in step S2, the simulation process is as follows:
[0032] Using the information of the genetic map, assuming that the number of recombination times obeys the Poisson distribution; simulate the number of recombination times of 10 pairs of chromosomes, the parent source of different chromosomes and the specific position of recombination of different chromosomes; according to the parent source of different chromosomes and the position of recombination of different chromosomes, splice the genotype, and finally simulate the genotype of different individuals.
[0033] Specifically, in step S3, the relationship is as follows:
[0034]
[0035] Wherein, represents the average probability value of appearing materials that meet the target gene introduction and the background return rate is greater than a certain level (0.85, 0.90, etc.) under multiple population sizes, randomly repeated M times, M is greater than or equal to 10; P S is the success probability; n is the number of backcross offspring individuals required.
[0036] Specifically, in step S5, find out how large the population size n of the BC1F1 generation is required under the number m of materials that meet the conditions, then there is at least a P S probability of finding m samples that need to meet:
[0037]
[0038] Wherein, n is the population size, m is the number of materials that meet the conditions, P SThe probability of success; This represents the average probability of finding materials that meet the target gene introduction criteria and have a background recovery rate greater than a certain level (0.85, 0.90, etc.) after 20 random repetitions across multiple population sizes.
[0039] A system for optimizing the design of maize backcross breeding includes:
[0040] The first data storage module is used to store the genotype data of the recipient parent and the donor parent, as well as the genetic map of maize;
[0041] The offspring genotype simulation module is used to simulate the genotypes of different individuals in the backcross offspring based on the genotype data and genetic map of the parents.
[0042] The second data storage module is used to store the genotype data of each sample in N simulations;
[0043] The average probability value calculation module is used to calculate the average probability value of materials that meet the target gene import requirements and have a background recovery rate greater than a certain level, based on the genotype data obtained from N simulations.
[0044] The first-generation individual count calculation module is used to calculate the number of offspring individuals based on the average probability value. The population size n of backcross offspring required for different success probabilities and different background response rates was calculated.
[0045] The second-generation offspring count calculation module is used to calculate the number of individuals under different success probabilities P. S Given the number of materials m that meet the conditions, what is the required population size n of backcross offspring?
[0046] In a preferred embodiment, the system further includes:
[0047] The correction module is used to correct the genotypes obtained by the offspring genotype simulation module. This correction module uses the identification of bin genotypes for correction.
[0048] The above system is used to simulate the BC1F1 generation genotype in maize backcross breeding.
[0049] The above system is used in calculating the population size n of the BC1F1 generation during maize backcross breeding. Its key feature is that it includes the population size n of the BC1F1 generation required for different success probabilities and different background reversion rates, as well as different success probabilities P. S Given the number of materials m that meet the conditions, what is the required population size n for generation BC1F1?
[0050] A computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0051] A computer readable storage medium stores a computer program, and the computer program implements the steps of the above method when executed by a processor.
[0052] Compared with the prior art, the present application has the following advantages and beneficial effects:
[0053] According to the actual needs of the applicant (the specific number of the number of eligible materials m, the different backcross generations required when the target gene is introduced and the background recovery rate reaches a certain value), and the population size n corresponding to different routes, the cost of genotype detection in the actual breeding process can be calculated based on the population size n. Breeders can choose a breeding route by considering the cost and the time corresponding to the backcross generations. That is, the present application provides customers with the optimal and most suitable strategy (how many generations of backcrossing and how large a population size to set for each generation) under certain cost conditions when the target gene is introduced and the background recovery rate reaches a certain value. Thus, resource waste is avoided, and breeders can better and faster screen target materials during multiple generations of backcrossing, saving a large amount of time and cost, improving breeding efficiency, and accelerating the cultivation of new varieties. BRIEF DESCRIPTION OF DRAWINGS
[0054] The accompanying drawings, which are included to provide a further understanding of the embodiments of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and, together with the description, serve to explain the principles of the application. In the drawings:
[0055] Figure 1 It is a different backcross scheme route map in embodiment 1 of the present application;
[0056] Figure 2 It is a schematic diagram of the possible genotypes of the BC1F1 population simulated in embodiment 1 of the present application;
[0057] Figure 3 It is a schematic diagram of the individual genotype heat map distribution of the 20th repetition of 1000 BC1F1 individuals simulated in embodiment 1 of the present application, wherein A in the legend represents the genotype from parent A, and B represents the genotype from parent B;
[0058] Figure 4 It is a method for bin genotype identification developed in embodiment 1 of the present application. After bin genotype identification of real data, the distribution diagram of the number of recombination on each chromosome is obtained.
[0059] Figure 5After the bin genotype is identified in the embodiment 1 of the present application, the average recombination frequency map of 4 chromosomes (chromosomes 1, 2, 6 and 7) is calculated in the form of sliding window (window size is 3M, step is 1.5M);
[0060] Figure 6 After the bin genotype is identified in the embodiment 1 of the present application, the average recombination frequency map of 4 chromosomes (chromosomes 3, 4, 8 and 9) is calculated in the form of sliding window (window size is 3M, step is 1.5M);
[0061] Figure 7 After the bin genotype is identified in the embodiment 1 of the present application, the average recombination frequency map of 2 chromosomes (chromosomes 5 and 10) is calculated in the form of sliding window (window size is 3M, step is 1.5M);
[0062] Figure 8 The genotype consistency rate distribution map of 1776 samples before and after the bin genotype is identified in the embodiment 1 of the present application;
[0063] Figure 9 The relationship distribution map of the physical position and the genetic position of chromosomes 1, 2, 6 and 7 after the bin genotype is identified in the embodiment 1 of the present application;
[0064] Figure 10 The relationship distribution map of the physical position and the genetic position of chromosomes 3, 4, 8 and 9 after the bin genotype is identified in the embodiment 1 of the present application;
[0065] Figure 11 The relationship distribution map of the physical position and the genetic position of chromosomes 5 and 10 after the bin genotype is identified in the embodiment 1 of the present application. DETAILED DESCRIPTION
[0066] In order to make the objectives, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application with reference to the embodiments and the drawings, the illustrative embodiments and the description thereof are only used to explain the present application, and do not limit the present application. The embodiments described below are part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0067] In the following description, numerous specific details are set forth to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without the specific details. In other instances, well-known structures, materials, or processes have not been described in detail in order to avoid obscuring the present application. The materials, instruments, and reagents used in the following examples, as well as other examples, are obtained from commercial vendors, unless otherwise specified. The techniques used in the examples, unless otherwise specified, are routine techniques commonly used in the art.
[0068] In addition, the terms "first", "second", etc. are used herein only to describe different instances, and do not imply or suggest relative importance or a number of the technical features indicated. Therefore, the features defined as "first", "second", etc. can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "a plurality" is two or more, unless otherwise specifically limited.
[0069] Embodiment 1
[0070] A method for optimizing the design of a maize backcross breeding population, comprising the following steps:
[0071] S1, obtaining genotype data of the recipient parent and the donor parent.
[0072] S2, according to the genotype data of the parents and the genetic map, assuming that the number of recombination events obeys Poisson distribution, the genotype of different individuals of BC1F1 generation is simulated; when simulating, the information of the genetic map is used, and it is assumed that the number of recombination events obeys Poisson distribution Wherein, λ is the expected value of Poisson distribution, which is also the number of recombination events of 10 pairs of chromosomes, the parent source of different chromosomes and the specific position of recombination of different chromosomes; according to the parent source of different chromosomes and the position of recombination of different chromosomes, the genotype is spliced, and finally the genotype of different individuals is simulated. The specific possible genotype of BC1F1 is shown in Figure 2 .
[0073] The model used in this embodiment is the existing model, such as hidden Markov model. Those skilled in the art can use hidden Markov model according to actual needs, and train through historical sample data. Using this model needs to input the genotype data of the parents, the genetic map information, and then the genotype data equivalent to the number of offspring individuals will be generated.
[0074] In order to verify the reliability of the backcross breeding population BC1F1 genotype simulation technology designed in this embodiment, the following comparison of real results and simulation results is used for verification:
[0075] By comparing the real data and the simulation data, from all the sample numbers (all_sam_nums), the sample numbers meeting the target conditions (meet_obj_sam), and the proportion (meet_sam_per) (target gene introduction and background recovery rate greater than 0.9), the average background recovery rate (ave_bck_rec), the minimum background recovery rate (min_bck_rec), the maximum background recovery rate (max_bck_rec), and the like, it is shown that the simulation method is reliable, and the specific results are shown in Table 1:
[0076] Table 1
[0077] all_sam_nums meet_obj_sams meet_sam_per ave_bck_rec min_bck_rec max_bck_rec Real 1695 850 0.5015 0.7521 0.5682 0.9202 Simulation 2000 999 0.500 0.745 0.5902 0.8965
[0078] From the data in Table 1, it can be seen that:
[0079] The difference between the real data and the simulation data is very small, indicating that the simulation method designed in this embodiment is reliable.
[0080] Note: The process of genotype simulation needs to input the genotype data of parents, genetic map, and the genotype results of the offspring after backcrossing.
[0081] S3, using the simulation method of step S2, based on N times of simulation of different BC1F1 generations, obtaining the probability value of the occurrence of BC1F1 generations meeting the target gene introduction and the background recovery rate greater than a certain level According to the size of the population scale of N times of simulation of different BC1F1 generations, referring to the statistical principle, obtaining a success probability P of the target material S , the probability value and the relationship of the BC1F1 generation, wherein N is a positive integer greater than or equal to 20;
[0082] The relationship is shown in formula (1):
[0083]
[0084] wherein, represents the average probability value of the occurrence of materials meeting the target gene introduction and the background recovery rate greater than a certain level (0.85, 0.90, etc.) under multiple population scales, randomly repeated 20 times; P S is the success probability; n is the number of backcross offspring individuals required.
[0085] In this embodiment, N is 20, 20 times of repetition are set in simulation, and 5 population scales are set. The gradient of the 5 population scales is 1000, 2000, 3000, 4000, and 5000 (the data here refers to the number of simulated backcross offspring BC1F1 individuals). The individual genotype heat map distribution diagram of the 20th repetition of 1000 BC1F1 individuals is as follows:Figure 3 as shown.
[0086] In the above five population sizes, the average probability value of the material meeting the target gene introduction and the background recovery rate greater than a certain level (0.85, 0.90, etc.) is calculated by repeating 20 times, and the average probability is calculated by calculating the probability of the material meeting the target gene introduction and the background recovery rate greater than a certain level in each repetition, and the total is divided by 20.
[0087] According to the simulation of different BC1F1 population sizes, the calculation formula of the success probability Ps greater than or equal to 99% of finding a target material is shown as formula (2):
[0088]
[0089] Based on the above formula (2), the calculation formula of the population size n of the BC1F1 generation is shown as formula (3):
[0090]
[0091] The base of formula (3) is 10.
[0092] S4, calculating different success probabilities P S The population size n of the BC1F1 generation meeting the target gene introduction and the background recovery rate greater than a certain level appears; the population size n of the BC1F1 generation corresponding to different success probabilities P S The population size n of the BC1F1 generation corresponding to different background recovery rates is shown in Table 2:
[0093] Table 2
[0094] bck_rec(0.9) bck_rec(0.85) bck_rec(0.80) P S (0.90)]]> 8194 260 32 P S (0.80)]]> 5727 181 22 P S (0.70)]]> 4284 136 17 P S (0.60)]]> 4284 136 17 P S (0.50)]]> 3261 104 13
[0095] From the data in Table 2, it can be seen that:
[0096] When the customer's requirement is that the background recovery rate (bck_rec) reaches 0.9, the success probability P S and reaches 0.90, the required population size of the BC1F1 generation is 8194 samples; when the customer's requirement is that the background recovery rate reaches 0.9, the success probability P S and reaches 0.80, the required population size of the BC1F1 generation is 5727 samples, etc.
[0097] Note: In formula (3), P is calculated based on a certain background recovery rate, that is, different background recovery rates is different, wherein the background recurrence rate is based on the genotype data, and the number of loci in which the genotype of the offspring is consistent with the genotype of the recurrent parent is calculated, and the specific formula is BCrate = [A + (1 / 2)H] / [A + H + B], wherein the homozygous genotype of the recurrent parent is denoted as A, the homozygous genotype of the donor is denoted as B, and the heterozygous genotype is denoted as H. As can be seen from formula 3, the population size n is related to Ps and is related to the background recurrence rate, so n is related to Ps and the background recurrence rate. is related to the background recurrence rate, so n is related to Ps and the background recurrence rate.
[0098] S5, based on the calculation result of step S4, the number of materials m that meet the conditions is found, and the population size n of the BC1F1 generation is required under different success probabilities P S ; then at least P S of the probability of finding m samples needs to meet the formula as shown in formula (4):
[0099]
[0100] wherein i is the iteration number, n is the population size, m is the number of materials that meet the conditions, P S is the success probability; represents that under multiple population sizes, the average probability value of materials that meet the target gene introduction and have a background recurrence rate greater than a certain level (0.85, 0.90, etc.) is obtained by repeating 20 times, and the more the number of repetitions, the better.
[0101] Step S4 of the embodiment is to find at least one material, and the population size n of the BC1F1 generation corresponding to different success probabilities P S and different background recurrence rates; step S5 is to set the number of target materials to m, and find the population size n of the BC1F1 generation under different success probabilities P S .
[0102] S6, based on the genotype data of the backcross offspring and the recipient parent, the population size n of the next generation of backcross offspring is calculated by referring to steps S2-S5;
[0103] For example: the genotype of the backcross second-generation individual is obtained by selecting the genotype of the BC1F1 individual and the genotype of the recurrent parent, and is obtained by a backcross individual genotype simulation program; then the principle of calculating the population size n is the same as above.
[0104] S7, under certain cost conditions, when the target gene introduction and the background recurrence rate reach a certain value, the required backcross generation number and the optimal route of the corresponding population size n are given. In a specific case, for example Figure 1As shown, if the breeder wants to backcross one generation, the individual that meets the target gene introduction and the background recovery rate is greater than 90%, and in the case of funds, the breeder can create 13814 backcross offspring, that is, there is a 90% success rate to find individuals that meet the requirements; similarly, if the success rate is reduced, the number of backcross offspring individuals required is also reduced, and the number of individuals required under different success probabilities is shown in the reference route. In addition, if the breeder needs to backcross two generations to find individuals that meet the target gene introduction and the background recovery rate is greater than 90%, the first backcross generation can screen materials with a background recovery rate of ≥85%, and on this basis, 42% of the individuals in the second backcross generation meet the target gene introduction and the background recovery rate is greater than 90%. In summary, the breeder can choose the optimal route suitable for himself according to his own funds and time, so as to screen the target individual.
[0105] Example 2:
[0106] This example is based on Example 1, and the difference from Example 1 is:
[0107] In this example, the genotype of the backcross population is corrected by the bin genotype identification method; the bin genotype identification method includes the following steps:
[0108] Step 1: Encode the genotype of each chromosome of each sample of the backcross population and sample screening, the encoded genotype value is divided into different fragments, and the length of each fragment is calculated;
[0109] The process of encoding and sample screening is as follows:
[0110] Step 11): Encode the genotype of the backcross population, 1 indicates that the gamete comes from the non-recurrent parent, 2 indicates that the gamete comes from the recurrent parent, 3 indicates that the judgment is wrong or the genotype is missing, hereinafter referred to as missing;
[0111] Step 12): Remove samples with a single chromosome missing rate greater than or equal to 0.5; remove samples with a whole genome (10 pairs of chromosomes) genotype missing rate greater than or equal to 0.5;
[0112] Step 13): Correct the genotype value of the remaining samples to 1 or 2 at the site encoded as 3, and perform bin genotype identification at the same time.
[0113] Step 2: Start from the longest fragment and go up and down, and perform smoothing processing on the fragment that is too small;
[0114] Step 3: Judge the middle fragment with the two side fragments fixed, so as to realize the bin genotype identification of the whole genome 10 pairs of chromosomes.
[0115] The reliability of the bin genotype identification in the embodiment is verified by the following method:
[0116] 1) The method of bin genotype identification developed in the embodiment is used to identify the bin genotype of real data (genotype data of the backcross offspring obtained by detection), from which the bin genotype of each sample is obtained. Figure 4 It can be seen that the recombination times of each chromosome after bin identification are about 2, which is lower than that before bin genotype identification, and meets the theoretical value;
[0117] 2) After bin genotype identification, the average recombination times of 10 chromosomes in the whole genome are calculated in the form of sliding window (window size is 3M, step is 1.5M). It can be seen from Figures 5-7 that the recombination rules of different chromosomes meet the expectation, that is, the recombination times of the two ends of the chromosome are high, and the recombination times of the centromere region in the middle are low;
[0118] 3) The genotype consistency rate of all 1776 samples is between 0.9 and 1.0 before and after bin genotype identification, the genotype consistency rate is high, which indicates that the method of bin genotype identification is reliable, and the specific genotype consistency rate is shown in Figure 8 ;
[0119] 4) After bin genotype identification, by analyzing and comparing the relationship between the physical position and the genetic position, the graph is drawn (see Figures 9-10 ), it can be found that the recombination rate of the middle position of the chromosome is low, and the recombination rate of the two ends is high, which meets the biological law.
[0120] Therefore, steps 1) to 4) show that the method of bin genotype identification developed in the embodiment is reliable and available.
[0121] The method of bin genotype identification developed in the embodiment can effectively identify the genotype of the backcross offspring, and improve the accuracy of the population size n calculation of the backcross offspring.
[0122] Embodiment 3:
[0123] The system for corn backcross breeding optimization design method of embodiment 1 or embodiment 2 comprises:
[0124] A first data storage module is configured to store genotype data of the recipient parent, the donor parent and the genetic map of the corn;
[0125] A progeny genotype simulation module is configured to simulate the genotype of different individuals of the backcross offspring based on the genotype data of the parents and the genetic map;
[0126] A second data storage module is configured to store the genotype data of each sample simulated for N times;
[0127] an average probability value calculation module for calculating an average probability value of N times of simulation based on the genotypic data obtained in the N times of simulation, the material meeting the target gene introduction and the background recovery rate being greater than a certain level
[0128] a first offspring individual number calculation module for calculating the population size n of the backcross offspring required by different success probabilities and different background recovery rates based on the average probability value the population size n of the backcross offspring required by different success probabilities and different background recovery rates is calculated
[0129] a second offspring individual number calculation module for calculating how large the population size n of the backcross offspring is required under the condition that the number m of materials meeting the conditions is found under different success probabilities P S
[0130] Preferably, the system further comprises:
[0131] a correction module for correcting the genotypes obtained by the offspring genotype simulation module, the correction module using the identification of the bin genotype for correction, and the method for identifying the bin genotype described in the embodiments being stored in the correction module.
[0132] The system of the embodiments is applied in the simulation of the genotypes of the BC1F1 generation in the process of corn backcross breeding.
[0133] The system of the embodiments is applied in the calculation of the population size n of the BC1F1 generation in the process of corn backcross breeding, and the population size n of the BC1F1 generation required by different success probabilities and different background recovery rates and how large the population size n of the BC1F1 generation is required under the condition that the number m of materials meeting the conditions is found under different success probabilities P S
[0134] Embodiment 4:
[0135] A computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the steps of the method described in Embodiment 1 or Embodiment 2 when executing the computer program.
[0136] Embodiment 5:
[0137] A computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the method described in Embodiment 1 or Embodiment 2.
[0138] The above detailed description of the specific embodiments of the present application has been given to understand the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for optimizing maize backcross breeding, characterized in that, Includes the following steps: S1. Obtain genotype data of the recipient parent and the donor parent; S2. Based on the genotype data and genetic map of the parents, assuming that the number of recombinations follows a Poisson distribution, the genotypes of different individuals in the BC1F1 generation are simulated. S3. Using the simulation method in step S2, based on N simulations of different BC1F1 generations, obtain the probability value of a BC1F1 generation that satisfies the target gene introduction requirement and has a background recovery rate greater than a certain level. Based on the size of different BC1F1 generation populations in N simulations, a success probability P for the target material is obtained. S probability value The relationship between N and BC1F1 is given by the expression, where N is a positive integer greater than or equal to 20. S4. Calculate the different success probabilities P using the relationship obtained in step S3. S The population size n of the BC1F1 generation that meets the target gene introduction requirements and has a background recovery rate greater than a certain level; S5. Based on the calculation results of step S4, calculate the success probabilities P under different conditions. S Given the number of materials m that meet the conditions, what is the required population size n for generation BC1F1? S6. Based on the genotype data of the backcross progeny and the recipient parent, refer to steps S2-S5 to calculate the population size n of the next generation of backcross progeny. S7. Under certain cost conditions, when the target gene is successfully introduced and the background recovery rate reaches a certain value, provide the optimal route for the required number of backcross generations and the corresponding population size n.
2. The maize backcross breeding optimization design method according to claim 1, characterized in that, The genotypes of the backcross population were corrected using the bin genotyping method; the bin genotyping method includes the following steps: Step 1: Encode the genotype of each chromosome of each sample in the backcross population and screen the samples. Divide the encoded genotype values into different segments and calculate the length of each segment. Step 2: Starting from the longest segment, traverse upwards and downwards, smoothing out the more fragmented segments; Step 3: Judge the three segments together, fix the two outer segments and judge the middle segment, so as to realize the bin genotype identification of all 10 pairs of chromosomes in the whole genome.
3. The maize backcross breeding optimization design method according to claim 2, characterized in that, In step 1, the coding rule is to use different numbers to distinguish gametes from donor parents, gametes from recipient parents, and genotype defects, which include errors or missing genotypes.
4. The maize backcross breeding optimization design method according to claim 2, characterized in that, In step 1, the sample screening process is as follows: Remove samples with a single chromosome deletion ratio greater than or equal to 0.5; remove samples with a whole-genome genotype deletion ratio greater than or equal to 0.
5. The remaining samples were then shown to have their genotype defect coding sites corrected to codes from gametes from either the donor or recipient parent.
5. The maize backcross breeding optimization design method according to claim 1, characterized in that, In step S2, the simulation process is as follows: Using information from genetic maps, we assume that the number of recombinations follows a Poisson distribution; we simulate the number of recombinations of 10 pairs of chromosomes, the parental origins of different chromosomes, and the specific locations where recombinations occur on different chromosomes; based on the parental origins of different chromosomes and the locations where recombinations occur on different chromosomes, we splice the genotypes and finally simulate the genotypes of different individuals.
6. A method for optimizing maize backcross breeding according to any one of claims 1-5, characterized in that, In step S3, the relationship is as follows: in, P represents the average probability of finding materials that meet the target gene introduction criteria and have a background recovery rate greater than a certain level (0.85, 0.90, etc.) after M random repetitions (M ≥ 10) across multiple population sizes; S denoted as , where is the probability of success; and n is the number of backcross offspring individuals required.
7. A maize backcross breeding optimization design method according to any one of claims 1-5, characterized in that, In step S5, given the number of materials m that meet the conditions, what is the required population size n for generation BC1F1? Then, at least P... S To find m samples with a probability of [missing information], the following conditions must be met: Where n is the group size, m is the number of materials that satisfy the conditions, and P S The probability of success; This represents the average probability of finding materials that meet the target gene introduction criteria and have a background recovery rate greater than a certain level (0.85, 0.90, etc.) after 20 random repetitions across multiple population sizes.
8. A system for the maize backcross breeding optimization design method as described in any one of claims 1-7, characterized in that, include: The first data storage module is used to store the genotype data of the recipient parent and the donor parent, as well as the genetic map of maize; The offspring genotype simulation module is used to simulate the genotypes of different individuals in the backcross offspring based on the genotype data and genetic map of the parents. The second data storage module is used to store the genotype data of each sample in N simulations; The average probability value calculation module is used to calculate the average probability value of materials that meet the target gene import requirements and have a background recovery rate greater than a certain level, based on the genotype data obtained from N simulations. The first-generation individual count calculation module is used to calculate the number of offspring individuals based on the average probability value. The population size n of backcross offspring required for different success probabilities and different background response rates was calculated. The second-generation offspring count calculation module is used to calculate the number of individuals under different success probabilities P. S Given the number of materials m that meet the conditions, what is the required population size n of backcross offspring? 9. The system according to claim 8, characterized in that, Also includes: The correction module is used to correct the genotypes obtained by the offspring genotype simulation module. This correction module uses the identification of bin genotypes for correction.
10. The application of the system as described in claim 8 or 9 in simulating the BC1F1 generation genotype during maize backcross breeding.
11. The application of the system as described in claim 8 or 9 in calculating the population size n of the BC1F1 generation during maize backcross breeding, characterized in that, This includes the population size n required for BC1F1 generations and the different success probabilities P required for different success probabilities and different background response rates. S Given the number of materials m that meet the conditions, what is the required population size n for generation BC1F1? 12. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.