A method, device, equipment and storage medium for selecting high-quality breeding populations

By using graph theory dynamic weighted alignment algorithm and deep learning model to analyze genomic data on cloud computing platform, the efficiency and accuracy problems of genomic data processing and association analysis in existing technologies are solved, and the efficiency and accuracy of the breeding process are achieved.

CN119694406BActive Publication Date: 2025-09-09JIHUA LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510199086.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-09-09
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

Existing technologies are inefficient and lack accuracy in genomic data processing, genetic variation detection, and genomic association analysis, making it difficult to meet the efficiency and precision requirements of modern breeding.

Method used

By uploading genome resequencing data to a cloud computing platform, a dynamic weighted alignment algorithm based on graph theory is used for comparison and analysis. Combined with a deep learning multi-layer convolutional neural network and attention mechanism, genetic variation detection and genome association analysis are performed to screen out breeding individuals with excellent traits.

Benefits of technology

It achieves efficient comparison and precise analysis of genomic data, can accurately identify breeding individuals with excellent traits, improves the efficiency and accuracy of the breeding process, and reduces human intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119694406B_ABST
    Figure CN119694406B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of breeding technology and discloses a method, apparatus, device, and storage medium for selecting high-quality breeding populations. The method comprises the following steps: uploading genome resequencing data to a cloud computing platform, fully utilizing the high-performance computing resources and distributed storage technology of cloud computing to achieve automated processing and efficient analysis of genomic data; employing a dynamic weighted alignment algorithm based on graph theory to perform segmented alignment and weight optimization on the uploaded genomic data to generate a genomic sequence map, greatly improving alignment efficiency and accuracy; furthermore, through genomic association analysis technology and deep learning models, breeding individuals with superior traits can be accurately identified and screened; ultimately, individuals with the highest comprehensive scores are screened through comprehensive scoring to form a high-quality breeding population. The present invention not only improves the accuracy of individual selection but also significantly reduces manual intervention, making the breeding process more efficient and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of breeding technology, and in particular to a method, device, equipment and storage medium for selecting a high-quality breeding population. Background Art

[0002] The rapid development of genome resequencing technology has provided unprecedented opportunities for agricultural breeding. Traditional breeding methods, such as phenotypic selection and crossbreeding, while somewhat successful, still suffer from numerous challenges, including low efficiency, long cycles, and uncertain results. Through genome resequencing, scientists can gain a comprehensive understanding of crop genomes and identify loci associated with desirable traits, thus providing a scientific basis for breeding decisions. However, with the rapid accumulation of genomic data, efficiently processing and analyzing this massive amount of data has become a significant challenge.

[0003] In traditional breeding methods, the processing and analysis of genomic data usually rely on local computing resources, which not only requires high-performance hardware support, but also requires a lot of time for data processing and analysis. For example, the genome resequencing data of a single breeding individual is usually as high as hundreds of GB. Faced with the data processing task of hundreds of breeding individuals, traditional methods often seem to be unable to cope with it. In addition, when processing genomic data, existing technologies often require manual steps such as quality control, sequence splicing, and genetic variation detection. These steps are cumbersome and error-prone, further limiting breeding efficiency. Another shortcoming of existing technologies is that it is difficult to accurately identify breeding individuals with excellent traits while processing large-scale data. Traditional genomic association analysis methods mostly rely on linear models or simple statistical methods. These methods are often unable to cope with complex traits and it is difficult to effectively identify gene loci that are significantly associated with the target traits. In addition, traditional methods usually ignore the effects of multi-gene interactions and environmental factors on traits, resulting in low accuracy and reliability of screening results.

[0004] The rise of cloud computing offers a new approach to addressing these challenges. By uploading genome resequencing data to cloud platforms, we can fully leverage cloud computing's high-performance computing resources and distributed storage technologies to rapidly process and analyze large amounts of data. However, achieving automated processing and efficient analysis of genomic data on cloud computing platforms still faces numerous challenges. For example, achieving efficient alignment and variant detection of genomic data on cloud platforms, leveraging deep learning and artificial intelligence technologies for genomic association analysis, and accurately screening individuals for superior traits based on association analysis results are all current research hotspots and challenges.

[0005] Traditional genome alignment algorithms are inefficient when processing large-scale data and are difficult to meet the needs of practical applications. Existing alignment algorithms mostly use fixed weights or simple dynamic programming methods, which make it difficult to fully utilize the parallel computing capabilities of cloud computing. In addition, traditional genetic variation detection methods usually require manual adjustment of parameters and thresholds, which are difficult to adapt to the characteristics of different data sets, affecting the accuracy of the test results. In terms of genome association analysis, existing technologies mainly rely on single statistical methods or simple machine learning models, which perform poorly when dealing with complex traits and multi-gene interactions. Although deep learning technology has made significant progress in fields such as image recognition and natural language processing, its application in genome association analysis is still in the exploratory stage. How to design an effective deep learning model and combine genomic data and phenotypic data for association analysis remains an urgent problem to be solved.

[0006] In summary, existing technologies have many shortcomings in genomic data processing, genetic variation detection and genomic association analysis, and are unable to meet the efficiency and precision requirements of modern breeding.

[0007] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0008] In order to overcome the shortcomings of the existing technology, the purpose of the present invention is to provide a method, device, equipment and storage medium for selecting high-quality breeding populations, aiming to solve the problem that the existing technology has many shortcomings in genome data processing, genetic variation detection and genome association analysis, and is difficult to meet the efficiency and accuracy requirements of modern breeding.

[0009] The first aspect of the present invention provides a method for selecting high-quality breeding populations, comprising: performing multiple high-throughput sequencing on each breeding individual in the breeding population to obtain genome sequence data; performing preliminary processing on the genome sequence data to obtain a preprocessed genome sequence; uploading the preprocessed genome sequence to a cloud platform, and using a dynamic weighted alignment algorithm based on graph theory to compare and analyze the uploaded genome sequence data to generate an individual genome sequence map; performing genetic variation detection on the genome sequence map to identify genetic variation data, wherein the genetic variation data includes single nucleotide polymorphisms, insertions and deletions, and structural variations; performing genome association analysis based on the genetic variation data to identify key gene sites related to target traits; using a multi-layer convolutional neural network based on deep learning, combined with an attention mechanism, to construct an association model that predicts the association between target traits and key gene sites; and screening out breeding individuals with excellent traits based on the results of the genome association analysis and the output value of the association model.

[0010] Optionally, in a first implementation of the first aspect of the present invention, the collected genome sequence data Perform preliminary quality control, use quality control software to evaluate the quality of each sequence, remove low-quality sequence data with a quality score lower than Q20, and obtain a high-quality sequence set , where Q(d ij ) represents the sequence d ij The quality score of D i QC It is a set of high-quality sequences retained after quality control, among which, , d in represents the nth DNA sequencing of the i-th individual sample; the software is used to remove the adapter sequence of the high-quality sequence set to generate trimmed sequence data , where m≤n; use splicing software to assemble the trimmed sequence data Splice to generate the preliminary genome splicing sequence G for each individual i ; The initial assembled genome sequence G i Remove repetitive sequences, identify and remove redundant and repetitive sequence fragments, and generate pre-processed genome sequences after removing repetitive sequences , where p≤k.

[0011] Optionally, in a second implementation of the first aspect of the present invention, the pre-processed genome sequence is uploaded to a cloud platform, and a dynamic weighted alignment algorithm based on graph theory is used to compare and analyze the uploaded genome sequence data to generate an individual genome sequence map, including the steps of: constructing a sequence fragment graph structure G(V, E) based on graph theory on the cloud platform, where V represents a node set, and each node corresponds to a sequence fragment s ijk , E represents the edge set, each edge corresponds to the alignment relationship between two sequence fragments, and the nodes and edges are constructed as follows: ; ; Where N represents the number of samples, m i represents the number of fragments of sample i, v ij represents the fragment j,w of sample i ijk represents the comparison weight between segments j and k; design an alignment algorithm that combines breeding objectives, where each sequence segment Divided into several sub-segments s ijk , each sub-segment corresponds to a child node v in the graph structure ijk : ,in, Representation fragment The sub-segment between positions a and b, Representation fragment The length of the breeding target is determined by introducing a scoring function for high-quality traits specific to the breeding goal, and determining the weights w between the child nodes. ijkl, and construct the initial weight matrix W, where w ijkl Represents the child node v ijk and v ijl The association weights of high-quality traits between ; ; ,in, and are the balance coefficients of similarity weight and quality trait weight, is the Kronecker delta function, is the scoring function for high-quality traits, is the weight of gene locus y, Represents sub-segment s ijk and s ijl The specific breeding target similarity at the gene locus y, T is the gene locus set associated with high-quality traits; the characteristic information of high-quality breeding individuals and the specific breeding target are introduced, and the weight matrix W is dynamically adjusted to obtain the adjusted weight matrix W', where ; ;in, represents the adjusted weight, is the weight adjustment coefficient, Represents sub-segment s ijk and s ijl The matching function between the gene loci T associated with a specific breeding target is used; the optimized dynamic programming algorithm is used to solve the optimal alignment path for the adjusted weight matrix W' to generate the individual genome sequence map M i , , where π represents the alignment path, (u, v) represents the edge in the path, Represents the adjusted weight.

[0012] Optionally, in a third implementation of the first aspect of the present invention, genetic variation detection is performed on the genome sequence map to identify genetic variation data, wherein the genetic variation data includes single nucleotide polymorphisms, insertions and deletions, and structural variations, including the steps of: using a cloud platform tool to perform genetic variation detection on the genome sequence map to obtain genetic variation data, wherein the detection method for single nucleotide polymorphisms is: , where p represents the genomic position, b represents the base type, q represents the quality score of the variant, and Q min is the quality score threshold; the detection method for insertion and deletion is: ,in, Indicates the length of insertion or deletion; the detection method for structural variation is: , where t represents the mutation type, and DUP, DEL, INV, and TRA represent duplication, deletion, inversion, and translocation, respectively.

[0013] Optionally, in a fourth implementation of the first aspect of the present invention, performing genomic association analysis based on the genetic variation data to identify key gene loci associated with the target trait comprises the steps of:

[0014] Annotate the detected single nucleotide polymorphisms, insertions and deletions, and structural variations and integrate them into a comprehensive annotation result A i , including the location, type, frequency, functional impact and related annotation information of each variant; the comprehensive annotation results A i Perform statistical analysis and generate statistical reports on genetic variation, including the number, distribution, and functional classification of variants: ,in, represents the total number of single nucleotide polymorphisms detected in the genome of individual i, Indicates the sum of the number of single nucleotide polymorphisms detected in each sample, m i is the number of sequenced fragments in individual i, SNP ij represents the set of SNPs detected in fragment j, represents the size of the set, i.e. the number of SNPs in fragment j, represents the total number of indels detected in the genome of individual i, Indicates the sum of the number of insertions and deletions detected in each sample. represents the set of indels detected in fragment j, represents the size of the set, i.e. the number of indels in segment j, represents the total number of structural variations detected in the genome of individual i, Indicates the sum of the number of structural variations detected in each sample, SV i represents the set of structural variations detected in fragment j, represents the functional classification of all genetic variants detected in the genome of individual i, This means that all detected genetic variations will be classified by functional category; the statistical reports and annotation results of genetic variations will be associated with breeding goals to identify key gene loci related to target traits: ,in, Indicates whether the variation v is related to the breeding target T.

[0015] Optionally, in a fifth implementation of the first aspect of the present invention, a multi-layer convolutional neural network based on deep learning is used in combination with an attention mechanism to construct an association model for predicting the association between the target trait and the key gene loci, comprising the steps of: based on the detected comprehensive annotation result A i, perform preliminary screening on each variant and select a set of potential candidate gene loci C associated with the target trait i , , where v represents the variation, T is the breeding target, and P is the characteristic vector of the variation. is a function that evaluates the correlation between the variation v and the breeding target T, is a set threshold; a multi-layer convolutional neural network based on deep learning is used to further analyze the candidate gene sites and construct an initial association model for predicting the association between the target trait and the key gene sites. The multi-layer convolutional neural network structure includes multiple convolutional layers, pooling layers and fully connected layers. An attention mechanism is introduced into the convolutional neural network to assign different weights to the characteristics of different gene sites and generate an attention weight matrix A; the predicted value of the association between the key gene site and the target trait is output through the fully connected layer; the classification loss function is used to train the initial association model, optimize the network parameters, and obtain the association model.

[0016] Optionally, in a sixth implementation of the first aspect of the present invention, breeding individuals with excellent traits are screened out based on the results of the genome association analysis and the output value of the association model, including the steps of: obtaining a set of key gene loci for each breeding individual based on the results of the genome association analysis; , where v represents the variation, C i is the set of candidate gene loci, T k is the kth target trait, Represents the relationship between the variation v and the target trait T k The correlation, p k Represents the target trait T k The weight of The comprehensive score of each breeding individual is calculated based on the weight of each key gene locus and the predicted value output by the association model. , where w v represents the weight of the key gene site v, Represents the output value of the association model for the gene locus v, Indicates the similarity between breeding individuals i and j; according to the comprehensive score S i Sort the breeding individuals and select the individuals with the highest comprehensive scores from high to low to build a high-quality breeding group: ,in, is the ratio of the set comprehensive score threshold, and N represents the total number of breeding individuals.

[0017] The second aspect of the present invention provides a high-quality breeding population selection device, comprising: a sequencing module for performing multiple high-throughput sequencing on each breeding individual in the breeding population to obtain genome sequence data; a preprocessing module for performing preliminary processing on the genome sequence data to obtain a preprocessed genome sequence; a comparison and analysis module for uploading the preprocessed genome sequence to a cloud platform, and using a dynamic weighted comparison algorithm based on graph theory to compare and analyze the uploaded genome sequence data to generate an individual genome sequence map; a variation detection module for performing genetic variation detection on the genome sequence map to identify genetic variation data, wherein the genetic variation data includes single nucleotide polymorphisms, insertions and deletions, and structural variations; an association analysis module for performing genome association analysis based on the genetic variation data to identify key gene sites related to target traits; a model construction module for using a multi-layer convolutional neural network based on deep learning, combined with an attention mechanism, to construct an association model for predicting the association between target traits and key gene sites; a screening module for screening breeding individuals with excellent traits based on the results of the genome association analysis and the output value of the association model.

[0018] The third aspect of the present invention provides a high-quality breeding population selection device, comprising: a memory and at least one processor, wherein the memory stores computer-readable instructions, and the memory and the at least one processor are interconnected through lines; the at least one processor calls the computer-readable instructions in the memory to enable the high-quality breeding population selection device to perform each step of the high-quality breeding population selection method as described above.

[0019] A fourth aspect of the present invention provides a computer-readable storage medium having computer-readable instructions stored therein, which, when executed on a computer, enables the computer to execute the various steps of the method for selecting a high-quality breeding population as described above.

[0020] Beneficial effects: The present invention uploads genome resequencing data to a cloud computing platform, making full use of the high-performance computing resources and distributed storage technology of cloud computing to achieve automated processing and efficient analysis of genome data. Through distributed storage technology, a large amount of genome data can be efficiently uploaded, stored and backed up to ensure the security and availability of the data. A dynamic weighted alignment algorithm based on graph theory is used to perform segmented alignment and weight optimization on the uploaded genome data to generate a genome sequence map, greatly improving the alignment efficiency and accuracy. In addition, through genome association analysis technology and deep learning models, breeding individuals with excellent traits can be accurately identified and screened. A multi-layer convolutional neural network structure is used to extract features through multiple convolutional layers, pooling layers and fully connected layers, and an attention mechanism is introduced to assign different weights to the features of different gene sites, thereby improving the accuracy of association analysis. Finally, through comprehensive scoring and further verification, the individuals with the highest comprehensive scores are screened out to form a high-quality breeding population. This process not only improves the accuracy of individual selection, but also significantly reduces manual intervention, making the breeding process more efficient and reliable. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 Flowchart of the method for selecting high-quality breeding populations provided in an embodiment of the present invention.

[0022] Figure 2 This is a schematic structural diagram of the high-quality breeding population selection device provided by the present invention.

[0023] Figure 3 This is a schematic diagram of the structure of the high-quality breeding population selection equipment provided by the present invention. DETAILED DESCRIPTION

[0024] The embodiments of the present invention provide a method, device, equipment and storage medium for selecting a high-quality breeding population. The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0025] See also Figure 1 , Figure 1The flow chart of a method for selecting a high-quality breeding population provided by the present invention is shown in the figure, which includes the following steps:

[0026] S10, performing high-throughput sequencing on each breeding individual in the breeding population multiple times to obtain genome sequence data;

[0027] Specifically, this embodiment collects samples from each breeding individual in the breeding population. Individual samples include plant tissues of leaves, stems, and roots. Ensure that each sample is large enough and healthy to avoid collecting damaged or diseased parts. Each sample is uniquely marked and relevant information (such as individual number, collection date, location, etc.) is recorded. The samples are stored in an appropriate preservation medium (such as liquid nitrogen or a refrigerator at a specific temperature) to maintain the integrity of their genetic information. DNA is extracted from the collected individual samples using a sequencing instrument; the extracted DNA is sequenced multiple times using sequencing technology, and each individual sample is sequenced at least three times independently to generate genome sequence data. ,In this embodiment, each individual sample is sequenced at least three times independently to reduce the impact of sequencing errors and biases;

[0028] During the sequencing process in this example, the Illumina platform was used to generate short-read or long-read sequence data. Read length is defined as the length of a single sequence read generated by the sequencing instrument, ranging from 100 to 300 base pairs. The Illumina platform offers high throughput, high accuracy, and low cost, making it suitable for large-scale genome sequencing. By setting appropriate sequencing parameters, short-read or long-read sequence data can be generated to meet research needs.

[0029] S20, performing preliminary processing on the genome sequence data to obtain a preprocessed genome sequence;

[0030] In this embodiment, the collected genome data D i Perform preliminary quality control, use quality control software to evaluate the quality of each sequence, remove low-quality sequence data with a quality score lower than Q20, and obtain a high-quality sequence set , where Q(d ij ) represents the sequence d ij The quality score of D i QC It is a set of high-quality sequences retained after quality control, among which, , d in represents the nth DNA sequencing of the i-th individual sample; preliminary quality control can remove low-quality sequence data and improve the accuracy and reliability of subsequent analysis; setting the quality threshold Q20 can ensure that the retained sequence data has a high quality level.

[0031] Then use the software to remove the adapter sequence from the high-quality sequence collection to generate trimmed sequence data , where m≤n, removing the linker sequence can reduce the interference and error of subsequent analysis; using splicing software to assemble the trimmed sequence data Splice to generate the preliminary genome splicing sequence G for each individual i , the spliced ​​sequence data can splice multiple short sequences into long sequences, providing complete genome sequence information for subsequent genome analysis; finally, the preliminary spliced ​​genome sequence Remove repetitive sequences, identify and remove redundant and repetitive sequence fragments, and generate genome sequences after removing repetitive sequences , where p≤k, removing duplicate sequences can reduce the complexity and computational effort of subsequent analysis.

[0032] S30, uploading the pre-processed genome sequence to a cloud platform, performing comparison and analysis on the uploaded genome sequence data using a dynamic weighted alignment algorithm based on graph theory, and generating a genome sequence map of the individual;

[0033] In this embodiment, the pre-processed genome sequence is first uploaded to the cloud platform, and then a sequence segment graph structure G(V, E) based on graph theory is constructed on the cloud platform, where V represents a node set and each node corresponds to a sequence segment s. ijk , E represents the edge set, each edge corresponds to the alignment relationship between two sequence fragments, and the nodes and edges are constructed as follows: ; ; Where N represents the number of samples, m i represents the number of fragments of sample i, v ij represents the fragment j,w of sample i ijk represents the alignment weight between fragments j and k;

[0034] Design an alignment algorithm that combines breeding objectives, where each sequence fragment Divided into several sub-segments s ijk , each sub-segment corresponds to a child node v in the graph structure ijk : ,in, Representation fragment The sub-segment between positions a and b, Representation fragment length;

[0035] By introducing a scoring function for high-quality traits specific to breeding objectives, the weights w between child nodes are determined. ijkl , and construct the initial weight matrix W, where w ijkl Represents the child node v ijk and v ijlThe association weights of high-quality traits between ; ; ,in, and are the balance coefficients of similarity weight and quality trait weight, is the Kronecker delta function, is the scoring function for high-quality traits, is the weight of gene locus y, Represents sub-segment s ijk and s ijl The similarity of a specific breeding target at the gene locus y, T is the set of gene loci associated with high-quality traits;

[0036] Introducing the characteristic information of high-quality breeding individuals and specific breeding goals, dynamically adjusting the weight matrix W, and obtaining the adjusted weight matrix W', where ; ;in, represents the adjusted weight, is the weight adjustment coefficient, Represents sub-segment s ijk and s ijl The matching function between the gene loci T associated with a specific breeding goal;

[0037] The optimized dynamic programming algorithm is used to solve the optimal alignment path for the adjusted weight matrix W' to generate the individual genome sequence map M i , , where π represents the alignment path, (u, v) represents the edge in the path, Indicates the adjusted weight. i Finally, the map is verified and corrected, and the accuracy and stability of high-quality traits for specific breeding targets are verified.

[0038] This embodiment, by combining a dynamic weighted alignment algorithm for specific breeding targets, not only achieves efficient alignment of genome sequences, but also introduces high-quality trait information of specific breeding targets into the alignment process, ensuring that the generated genome sequence map more accurately reflects the specific high-quality traits of breeding individuals, providing more accurate data support for subsequent genetic variation detection and genome association analysis.

[0039] S40, performing genetic variation detection on the genome sequence map to identify genetic variation data, wherein the genetic variation data includes single nucleotide polymorphisms, insertions and deletions, and structural variations;

[0040] In this embodiment, a cloud platform tool is used to perform genetic variation detection on the genome sequence map to obtain genetic variation data, wherein the detection method for single nucleotide polymorphism is: , where p represents the genomic position, b represents the base type, q represents the quality score of the variant, and Q min is the quality score threshold; the detection method for insertion and deletion is: ,in, Indicates the length of the insertion or deletion;

[0041] Structural variation detection methods are: , where t represents the variant type, and DUP, DEL, INV, and TRA represent duplication, deletion, inversion, and translocation, respectively. This example uses tools on the cloud platform to detect genetic variations, accurately identifying the location and type of variant sites, providing a reliable basis for in-depth research. The cloud platform also provides visualization tools and a wealth of parameter settings, allowing users to flexibly adjust and optimize according to research needs.

[0042] S50, performing genome association analysis based on the genetic variation data to identify key gene loci associated with the target trait;

[0043] In this embodiment, the detected single nucleotide polymorphisms are annotated, and the location, function and potential impact of each single nucleotide polymorphism are annotated with reference to the genome and annotation database, and annotation results are generated; the detected indels are annotated, the location and length of each indel are analyzed, and the potential functional impact of the indel is annotated using tools on the cloud platform, and annotation results are generated; the detected structural variations are counted and annotated, and structural variations including chromosomal rearrangements, copy number variations and large indels are identified. Each structural variation is annotated in detail using tools on the cloud platform, and annotation results are generated; all detected single nucleotide polymorphisms, indels and structural variations are integrated into a comprehensive annotation result A i , including the location, type, frequency, functional impact and related annotation information of each variant; the comprehensive annotation results A i Perform statistical analysis and generate statistical reports on genetic variation, including the number, distribution, and functional classification of variants: ,in, represents the total number of single nucleotide polymorphisms detected in the genome of individual i, Indicates the sum of the number of single nucleotide polymorphisms detected in each sample, m i is the number of sequenced fragments in individual i, SNP ij represents the set of SNPs detected in fragment j, represents the size of the set, i.e. the number of SNPs in fragment j, represents the total number of indels detected in the genome of individual i, Indicates the sum of the number of insertions and deletions detected in each sample. represents the set of indels detected in fragment j, represents the size of the set, i.e. the number of indels in segment j, represents the total number of structural variations detected in the genome of individual i, Indicates the sum of the number of structural variations detected in each sample, SV i represents the set of structural variations detected in fragment j, represents the functional classification of all genetic variants detected in the genome of individual i, This means that all detected genetic variations are classified by functional category. The statistical report and annotation results of genetic variations are analyzed in association with breeding goals to identify key gene loci related to target traits: ,in, Indicates whether the variation v is related to the breeding target T.

[0044] S60. Using a multi-layer convolutional neural network based on deep learning and combined with an attention mechanism, we construct an association model to predict the association between target traits and key gene loci.

[0045] In this embodiment, based on the detected comprehensive annotation result A i , perform preliminary screening on each variant and select a set of potential candidate gene loci C associated with the target trait i , , where v represents the variation, T is the breeding target, and P is the characteristic vector of the variation. is a function that evaluates the correlation between the variation v and the breeding target T, is a set threshold; a multi-layer convolutional neural network based on deep learning is used to further analyze the candidate gene sites and construct an initial association model for predicting the association between the target trait and the key gene sites. The multi-layer convolutional neural network structure includes multiple convolutional layers, pooling layers and fully connected layers. An attention mechanism is introduced into the convolutional neural network to assign different weights to the characteristics of different gene sites and generate an attention weight matrix A; the predicted value of the association between the key gene site and the target trait is output through the fully connected layer; the classification loss function is used to train the initial association model, optimize the network parameters, and obtain the association model.

[0046] Specifically, the main purpose of this embodiment is to analyze the specific degree of association between candidate loci and target traits and construct a high-precision association model between traits and genetic variation. The convolutional layer in deep learning automatically extracts key features of candidate loci (such as genetic variation patterns and functional associations). Each candidate locus is modeled as a feature vector P, which can be a property of the genetic variation (such as location, variation frequency, annotation function, etc.). Convolution operations are used to extract potential relationships between candidate loci and target traits, specifically analyzing associations between candidate loci and specific attributes of the target trait (such as gene function and impact pathways). The multi-layer convolutional neural network architecture comprises multiple convolutional layers, pooling layers, and fully connected layers. The convolutional layers extract higher-level features. During network training, an attention mechanism is incorporated to assign different weights to different loci, emphasizing those more relevant to the target trait. Finally, the fully connected layer generates predictions, quantifying the contribution of each candidate locus to the target trait. A classification loss function is used to train the network and optimize the neural network parameters to ensure higher accuracy in the analysis results.

[0047] In this embodiment, the feature vector p of the training data includes: genetic variation information (such as variation type and annotation function) of the candidate gene loci and target traits, where the target traits refer to traits that need to be optimized or improved in breeding goals, such as disease resistance and yield, and the values ​​of these traits are obtained through experiments or field trials; the association information between the candidate gene loci and the target traits is used as a label.

[0048] The final output of the model is the predicted value of the association between the key gene site and the target trait. The predicted value is not a simple probability or percentage, but a quantitative "association score", which is a scalar (such as a floating point number) that represents the degree of contribution of a specific gene site to the target trait.

[0049] S70. Screen out breeding individuals with excellent traits based on the results of the genome association analysis and the output value of the association model.

[0050] Based on the results of genome association analysis, obtain the key gene loci set for each breeding individual , where v represents the variation, C i is the set of candidate gene loci, T k is the kth target trait, Represents the relationship between the variation v and the target trait T k The correlation, p k Represents the target trait T k The weight of is the set threshold;

[0051] The comprehensive score of each breeding individual is calculated based on the weight of each key gene locus and the predicted value output by the association model , where w v represents the weight of the key gene site v, Represents the output value of the association model for the gene locus v, represents the similarity between breeding individuals i and j;

[0052] According to the comprehensive score S i Sort the breeding individuals and select the individuals with the highest comprehensive scores from high to low to build a high-quality breeding group: ,in, is the ratio of the set comprehensive score threshold, and N represents the total number of breeding individuals.

[0053] In the present invention, the selected high-quality breeding individuals can be further verified, including field trials and gene expression analysis, to confirm their breeding potential and the stability of excellent traits; based on the verification results, the high-quality breeding individuals are finally determined and used for breeding.

[0054] The present invention uploads genome resequencing data to a cloud computing platform, making full use of cloud computing's high-performance computing resources and distributed storage technology to achieve automated processing and efficient analysis of genome data. Through distributed storage technology, a large amount of genome data can be efficiently uploaded, stored, and backed up to ensure data security and availability. A dynamic weighted alignment algorithm based on graph theory is used to segment and optimize the weight of the uploaded genome data to generate a genome sequence map, greatly improving alignment efficiency and accuracy. In addition, through genome association analysis technology and deep learning models, breeding individuals with excellent traits can be accurately identified and screened. A multi-layer convolutional neural network structure is used to extract features through multiple convolutional layers, pooling layers, and fully connected layers, and an attention mechanism is introduced to give different weights to the features of different gene sites, thereby improving the accuracy of association analysis. Ultimately, through comprehensive scoring and further verification, the individuals with the highest comprehensive scores are screened out to form a high-quality breeding population. This process not only improves the accuracy of individual selection, but also significantly reduces manual intervention, making the breeding process more efficient and reliable.

[0055] The above describes the method for selecting a high-quality breeding population in the embodiment of the present invention. The following describes the device for selecting a high-quality breeding population in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, a high-quality breeding population selection device includes:

[0056] The sequencing module 10 is used to perform multiple high-throughput sequencing on each breeding individual in the breeding population to obtain genome sequence data;

[0057] A preprocessing module 20 is used to perform preliminary processing on the genome sequence data to obtain a preprocessed genome sequence;

[0058] The comparison and analysis module 30 is used to upload the pre-processed genome sequence to the cloud platform, and use a dynamic weighted comparison algorithm based on graph theory to compare and analyze the uploaded genome sequence data to generate an individual genome sequence map;

[0059] a variation detection module 40 for performing genetic variation detection on the genome sequence map to identify genetic variation data, wherein the genetic variation data includes single nucleotide polymorphisms, insertions and deletions, and structural variations;

[0060] an association analysis module 50 for performing genomic association analysis based on the genetic variation data to identify key gene loci associated with the target trait;

[0061] A model building module 60 is used to construct an association model for predicting the association between the target trait and the key gene loci using a multi-layer convolutional neural network based on deep learning and combined with an attention mechanism;

[0062] The screening module 70 is used to screen out breeding individuals with excellent traits based on the results of the genome association analysis and the output value of the association model.

[0063] Based on the same idea as the method in the above embodiment, the device provided in the present application can implement the method in the above embodiment. For the convenience of explanation, the structural diagram of the device embodiment only shows the parts related to the embodiment of the present application. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and may include more or fewer modules than shown in the figure, or a combination of certain modules, or different module arrangements.

[0064] Figure 2 The high-quality breeding population selection device in the embodiment of the present invention is described in detail from the perspective of modular functional entities, and the high-quality breeding population selection device in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0065] Figure 3Figure 1 is a schematic diagram of the structure of a high-quality breeding population selection device provided by an embodiment of the present invention. The high-quality breeding population selection device 100 may vary significantly depending on configuration or performance. It may include one or more central processing units (CPUs) 11 (e.g., one or more processors), memory 12, and one or more storage media 13 (e.g., one or more mass storage devices) storing application programs 133 or data 132. The memory 12 and storage medium 13 may be either transient or persistent storage. The program stored in the storage medium 13 may include one or more modules (not shown), each of which may include a series of instructions for operating on the high-quality breeding population selection device 100. Furthermore, the processor 11 may be configured to communicate with the storage medium 13 to execute the series of instructions stored in the storage medium 13 on the high-quality breeding population selection device 100.

[0066] The high-quality breeding population selection device 100 may further include one or more power supplies 14, one or more wired or wireless network interfaces 15, one or more input and output interfaces 16, and / or one or more operating systems 131, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be understood by those skilled in the art that Figure 3 The device structure shown does not constitute a limitation on the high-quality breeding population selection device 100, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0067] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of the method for selecting a high-quality breeding population.

[0068] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0069] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0070] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for selecting a high-quality breeding population, characterized in that: Including steps: Performing multiple high-throughput sequencing on each breeding individual in the breeding population to obtain genome sequence data; After preliminary processing of the genome sequence data, a preprocessed genome sequence is obtained; The pre-processed genome sequence is uploaded to a cloud platform, and a sequence segment graph structure G(V, E) based on graph theory is constructed on the cloud platform, where V represents a set of nodes, each node corresponds to a sequence segment, and E represents a set of edges, each edge corresponds to the alignment relationship between two sequence segments. The nodes and edges are constructed as follows: ; ; Where N represents the number of samples, m i represents the number of fragments of sample i, v ij represents the fragment j,w of sample i ijk represents the alignment weight between fragments j and k; Design an alignment algorithm that combines breeding objectives, where each sequence fragment is divided into several sub-segments s ijk , each sub-segment corresponds to a child node v in the graph structure ijk : ,in, Representation fragment The sub-segment between positions a and b, Representation fragment length; By introducing a scoring function for high-quality traits specific to breeding objectives, the weights w between child nodes are determined. ijkl , and construct the initial weight matrix W, where w ijkl Represents the child node v ijk and v ijl The association weights of high-quality traits between ; ; ,in, and are the balance coefficients of similarity weight and quality trait weight, is the Kronecker delta function, is the scoring function for high-quality traits, is the weight of gene locus y, Represents sub-segment s ijk and s ijl The similarity of a specific breeding target at the gene locus y, T is the set of gene loci associated with high-quality traits; Introducing the characteristic information of high-quality breeding individuals and specific breeding goals, dynamically adjusting the weight matrix W, and obtaining the adjusted weight matrix W', where ; ;in, represents the adjusted weight, is the weight adjustment coefficient, Represents sub-segment s ijk and s ijl The matching function between the set of gene loci T associated with high-quality traits; The optimized dynamic programming algorithm is used to solve the optimal alignment path for the adjusted weight matrix W' to generate the individual genome sequence map M i ; Performing genetic variation detection on the genome sequence map to identify genetic variation data, wherein the genetic variation data includes single nucleotide polymorphisms, insertions and deletions, and structural variations; Performing genomic association analysis based on the genetic variation data to identify key gene loci associated with the target trait; Using a multi-layer convolutional neural network based on deep learning and combined with an attention mechanism, we constructed an association model to predict the association between target traits and key gene loci. Breeding individuals with excellent traits are screened out based on the results of genome association analysis and the output values ​​of the association model.

2. The method for selecting a high-quality breeding population according to claim 1, wherein: The genome sequence data is preliminarily processed, comprising the steps of: The collected genome sequence data D i Perform preliminary quality control, use quality control software to evaluate the quality of each sequence, remove low-quality sequence data with a quality score lower than Q20, and obtain a high-quality sequence set , where Q(d ij ) represents the sequence d ij The quality score of D i QC It is a set of high-quality sequences retained after quality control, among which, , d in represents the nth DNA sequencing of the i-th individual sample; Use software to remove the adapter sequence from the high-quality sequence set to generate trimmed sequence data , where m≤n; Use splicing software to assemble the trimmed sequence data Splice to generate the preliminary genome splicing sequence G for each individual i ; The preliminary genome assembly sequence G i Remove repetitive sequences, identify and remove redundant and repetitive sequence fragments, and generate pre-processed genome sequences after removing repetitive sequences .

3. The method for selecting a high-quality breeding population according to claim 1, wherein: Performing genetic variation detection on the genome sequence map to identify genetic variation data, wherein the genetic variation data includes single nucleotide polymorphisms, insertions and deletions, and structural variations, comprises the steps of: The cloud platform tool is used to perform genetic variation detection on the genome sequence map to obtain genetic variation data, wherein the detection method for single nucleotide polymorphism is: , where p represents the genomic position, b represents the base type, q represents the quality score of the variant, and Q min is the quality score threshold; The detection method for insertion and deletion is: ,in, Indicates the length of the insertion or deletion; Structural variation detection methods are: , where t represents the mutation type, and DUP, DEL, INV, and TRA represent duplication, deletion, inversion, and translocation, respectively.

4. The method for selecting a high-quality breeding population according to claim 1, wherein: Performing genomic association analysis based on the genetic variation data to identify key gene loci associated with the target trait includes the following steps: Annotate the detected single nucleotide polymorphisms, insertions and deletions, and structural variations and integrate them into a comprehensive annotation result A i , including the location, type, frequency, functional impact and related annotation information of each variant; Comprehensive annotation results A i Perform statistical analysis and generate statistical reports on genetic variation, including the number, distribution, and functional classification of variants: ,in, represents the total number of single nucleotide polymorphisms detected in the genome of individual i, Indicates the sum of the number of single nucleotide polymorphisms detected in each sample, m i is the number of sequenced fragments in individual i, SNP ij represents the set of SNPs detected in fragment j, represents the size of the set, i.e. the number of SNPs in fragment j, represents the total number of indels detected in the genome of individual i, Indicates the sum of the number of insertions and deletions detected in each sample. represents the set of indels detected in fragment j, represents the size of the set, i.e. the number of indels in segment j, represents the total number of structural variations detected in the genome of individual i, Indicates the sum of the number of structural variations detected in each sample, SV i represents the set of structural variations detected in fragment j, represents the functional classification of all genetic variants detected in the genome of individual i, It means that all detected genetic variants are classified into functional categories; Correlate the statistical reports and annotations of genetic variation with breeding goals to identify key gene loci associated with target traits: ,in, Indicates whether the variation v is related to the breeding target T.

5. The method for selecting a high-quality breeding population according to claim 1, wherein: Using a multi-layer convolutional neural network based on deep learning and combined with an attention mechanism, we constructed an association model to predict the association between target traits and key gene loci, including the following steps: Based on the comprehensive annotation results detected A i , perform preliminary screening on each variant and select a set of potential candidate gene loci C associated with the target trait i , , where v represents the variation, T is the breeding target, and P is the characteristic vector of the variation. is a function that evaluates the correlation between the variation v and the breeding target T, is the set threshold; A multi-layer convolutional neural network based on deep learning is used to further analyze candidate gene loci and construct an initial association model for predicting the association between target traits and key gene loci. The multi-layer convolutional neural network structure includes multiple convolutional layers, pooling layers, and fully connected layers. An attention mechanism is introduced into the convolutional neural network to assign different weights to the characteristics of different gene loci, generating an attention weight matrix A. Output the predicted value of the association between key gene loci and target traits through the fully connected layer; Use the classification loss function to train the initial association model, optimize the network parameters, and obtain the association model.

6. The method for selecting a high-quality breeding population according to claim 1, wherein: Screening out breeding individuals with excellent traits based on the results of genome association analysis and the output value of the association model includes the following steps: Based on the results of genome association analysis, obtain the key gene loci set for each breeding individual , where v represents the variation, C i is the set of candidate gene loci, T k is the kth target trait, Represents the relationship between the variation v and the target trait T k The correlation, p k Represents the target trait T k The weight of is the set threshold; The comprehensive score of each breeding individual is calculated based on the weight of each key gene locus and the predicted value output by the association model , where w v represents the weight of the key gene site v, Represents the output value of the association model for the gene locus v, represents the similarity between breeding individuals i and j; According to the comprehensive score S i Sort the breeding individuals and select the individuals with the highest comprehensive scores from high to low to build a high-quality breeding group: ,in, is the ratio of the set comprehensive score threshold, and N represents the total number of breeding individuals.

7. A high-quality breeding population selection device, characterized in that: include: A sequencing module is used to perform multiple high-throughput sequencing on each breeding individual in the breeding population to obtain genome sequence data; A preprocessing module, configured to perform preliminary processing on the genome sequence data to obtain a preprocessed genome sequence; The comparison analysis module is used to upload the pre-processed genome sequence to the cloud platform and construct a sequence segment graph structure G(V, E) based on graph theory on the cloud platform, where V represents a set of nodes, each node corresponds to a sequence segment, and E represents a set of edges, each edge corresponds to the alignment relationship between two sequence segments. The nodes and edges are constructed as follows: ; ; Where N represents the number of samples, m i represents the number of fragments of sample i, v ij represents the fragment j,w of sample i ijk represents the comparison weight between segments j and k; design an alignment algorithm that combines breeding objectives, where each sequence segment is divided into several sub-segments s ijk , each sub-segment corresponds to a child node v in the graph structure ijk : ,in, Representation fragment The sub-segment between positions a and b, Representation fragment The length of the breeding target is determined by introducing a scoring function for high-quality traits specific to the breeding goal, and determining the weights w between the child nodes. ijkl , and construct the initial weight matrix W, where w ijkl Represents the child node v ijk and v ijl The association weights of high-quality traits between ; ; ,in, and are the balance coefficients of similarity weight and quality trait weight, is the Kronecker delta function, is the scoring function for high-quality traits, is the weight of gene locus y, Represents sub-segment s ijk and s ijl The specific breeding target similarity at the gene locus y, T is the gene locus set associated with high-quality traits; the characteristic information of high-quality breeding individuals and the specific breeding target are introduced, and the weight matrix W is dynamically adjusted to obtain the adjusted weight matrix W', where ; ;in, represents the adjusted weight, is the weight adjustment coefficient, Represents sub-segment s ijk and s ijl The matching function between the gene loci T associated with a specific breeding target is used; the optimized dynamic programming algorithm is used to solve the optimal alignment path for the adjusted weight matrix W' to generate the individual genome sequence map M i ; a variation detection module, configured to perform genetic variation detection on the genome sequence map and identify genetic variation data, wherein the genetic variation data includes single nucleotide polymorphisms, insertions and deletions, and structural variations; An association analysis module, configured to perform genome association analysis based on the genetic variation data to identify key gene loci associated with the target trait; A model building module is used to build an association model that predicts the association between target traits and key gene loci using a multi-layer convolutional neural network based on deep learning and combined with an attention mechanism; The screening module is used to screen out breeding individuals with excellent traits based on the results of genome association analysis and the output values ​​of the association model.

8. A high-quality breeding population selection device, characterized in that: comprising a memory and at least one processor, wherein the memory has computer-readable instructions stored therein; The at least one processor calls the computer-readable instructions in the memory to execute each step of the method for selecting a high-quality breeding population as described in any one of claims 1 to 6.

9. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the various steps of the method for selecting a high-quality breeding population as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Gene phenotype training and predicting method and device based on graph neural network

    CN115331732A

  • Kandelia candel germplasm resource analysis and improved variety breeding method

    CN118773365A