Wheat genome structure variation detection method based on deep learning

Through the wheat genome structural variation detection method based on deep learning, genomic data is transformed into images and a deep learning prediction model is constructed, solving the problem of limited detection accuracy and efficiency in the wheat genome complexity and high-throughput data in the existing technology, and achieving efficient and accurate structural variation prediction.

CN120220804APending Publication Date: 2025-06-27HENAN AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510375779.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing genome structural variation detection methods are limited under the complexity of the wheat genome and the huge scale of high-throughput sequencing data, making it difficult to effectively identify structural variation in wheat.

Method used

The wheat genome structural variation detection method based on deep learning is adopted to transform genomic data into image form through the genomic structural variation image generation algorithm, and a deep learning-based gene structural variation prediction model is constructed to automatically extract and analyze the variation characteristics in the image to achieve efficient and accurate structural variation prediction.

Benefits of technology

It improves the accuracy and comprehensiveness of detection of wheat genome structural variants, overcomes the shortcomings of traditional methods, and provides a new idea and tool for the study of wheat genome structural variants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220804A_ABST
    Figure CN120220804A_ABST
Patent Text Reader

Abstract

The invention discloses a wheat genome structure variation detection method based on deep learning, and develops a genome structure variation visualization modeling method integrating RD, DRP and SR multi-dimensional features in order to solve the problem that information features are difficult to extract due to the fact that wheat genome data is large in volume and complex in structure. According to the method, multi-dimensional feature analysis is carried out on original sequencing data in an FASTQ format, accurate positioning of candidate variation intervals is achieved, a structural variation discrimination model based on RD signal strength, DRP spatial distribution and SR breakpoint features is established, a structural variation visualization map with multi-dimensional feature fusion is constructed, and the structural variation discrimination model is used for discriminating the structural variation of the FASTQ format. A deep convolutional neural network is adopted to perform feature learning and pattern recognition on the visual atlas; the method can efficiently identify the structural variation in the wheat genome, has better prediction performance and higher detection precision compared with a traditional method, provides an extensible analysis framework for crop genome structural variation detection, and has important application value for promoting wheat molecular breeding technology innovation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text image conversion, recognition and analysis, and specifically relates to a method for graphical detection of genomic structural variations. Background Art

[0002] In genomics research, the crucial role of structural variations (SVs) in plant genetics and breeding has become increasingly prominent. However, due to the complexity of the wheat genome and the huge scale of high-throughput sequencing data, the accurate detection of structural variations still faces many challenges.

[0003] Hexaploid bread wheat (Triticum aestivum L., AABBDD) is one of the world's important food crops, and its excellent yield and adaptability make it a key guarantee for global food security. The genomic structure of bread wheat is complex, containing two different genomes (A, B, D) from three different ancestral species, forming the genomic constitution of AABBDD. This complex genomic structure endows bread wheat with high genetic diversity and rich adaptability, but also leads to the complexity of genomic rearrangement and gene expression regulation. With the challenges brought about by global population growth and climate change, improving the yield and stress resistance of wheat has become the core goal of agricultural research.

[0004] In this context, structural variations (SVs) in the genome play an important role in the evolution, adaptability and trait formation of wheat. Structural variations refer to variation types such as insertion (INS), deletion (DEL), inversion (INV), duplication (DUP) and translocation (TL) that occur in fragments in the genome with a length greater than 50bp.

[0005] These structural variations have been deeply analyzed through high-throughput genomic sequencing technology. The research reveals that the high-yield traits and stress resistance of wheat are closely related to multiple gene groups and regulatory networks. By comparing the genomes of different varieties, the research has also discovered key genes related to major agronomic traits (such as spike weight, disease resistance, drought tolerance). These findings provide a theoretical basis for wheat molecular breeding and strong support for improving the yield and stress resistance of wheat.

[0006] The wheat genome is large and complex, with a total genome of approximately 17 billion base pairs, and is rich in a large number of repetitive sequences and homologous genes. Traditional methods for detecting genomic structural variations, such as fluorescence in situ hybridization (FISH), chromosome staining techniques, and PCR amplification, although have made some progress in small-scale studies, due to the complexity of the wheat genome, especially when identifying large-scale repetitive sequences and highly polymorphic regions, the accuracy and efficiency of these methods have been significantly limited. With the development of high-throughput sequencing technology (HTS), researchers can detect structural variations in wheat more comprehensively and accurately through whole-genome data. However, limited by problems such as the complexity of data analysis and the variety of structural variations, existing computational methods still have certain limitations. Therefore, developing new methods to improve the detection accuracy and comprehensiveness of wheat genomic structural variations remains an urgent research problem to be solved. In recent years, deep learning technology has made remarkable progress in the field of bioinformatics, especially in the analysis and interpretation of genomic data. Deep learning can automatically extract potential features from a large amount of genomic data by constructing complex neural network models, and then accurately identify structural variations. Compared with traditional machine learning methods, deep learning models have stronger expressive power and higher accuracy when dealing with large-scale and high-dimensional data. In the detection of wheat genomic structural variations, the advantages of deep learning are particularly prominent. By combining advanced learning methods such as deep neural networks (DNN) and convolutional neural networks (CNN), the automatic identification and classification of complex structural variations in genomic data can be achieved. Especially for some complex variations that are difficult to capture by traditional methods (such as large-scale insertions, deletions, inversions, etc.), deep learning is expected to provide a more efficient and accurate solution.

[0007] Given the massive, high-dimensional, and sequential characteristics of high-throughput genomics data, deep learning, as a data-driven algorithm, shows great feasibility and potential in the field of bioinformatics, and is expected to break through the limitations of traditional algorithms through the complex function fitting ability of deep neural networks and improve the task accuracy.

[0008] In view of the above, the present application provides a method for detecting wheat genomic structural variations based on deep learning to solve the above problems. Summary of the Invention

[0009] In view of the above situation, in the research of wheat genome structural variation detection, traditional methods have limitations in the problems of difficult data reading and low prediction accuracy. This study proposes a method for wheat genome structural variation detection based on deep learning, which is used to predict two common and frequently occurring structural variations, namely deletion and tandem repeat. This method includes two core steps: First, a genome structural variation image generation algorithm is adopted to convert genome data into image form, thereby improving the data processing efficiency; Second, a gene structural variation prediction model based on deep learning is constructed, and by automatically extracting and analyzing the variation characteristics in the image, efficient and accurate structural variation prediction is achieved. This method can overcome the deficiencies of traditional means and provide a new idea and tool for the research of wheat genome structural variation.

[0010] A method for wheat genome structural variation detection based on deep learning, characterized by including the following steps:

[0011] Use the FASTQ file generated by sequencing data through a sequencer to align with the reference genome using the BWA software, and obtain a SAM file containing variation information from it. Subsequently, use the samtools tool to convert the SAM file into a binary format BAM file. After preprocessing the BAM file and combining RD (read depth data), DRP (discordant read pair data), and SR (split read data), generate structural variation images, and input these images into a deep learning model for training of variation prediction.

[0012] The beneficial effects of the above technical solutions are as follows:

[0013] (1) This study can efficiently identify various structural variations in the wheat genome, has better prediction performance and higher detection accuracy than traditional methods, provides a novel deep learning framework for the efficient detection of wheat genome structural variations, and provides strong technical support for wheat genetic improvement and breeding research;

[0014] (2) By adopting a genome structural variation image generation algorithm to convert genome data into image form, the relevant information of structural variations can be intuitively presented, and the output images can clearly display the characteristic information of the variation regions, thereby improving the data processing efficiency;

[0015] (3) By constructing a gene structural variation prediction model based on deep learning, automatically extracting and analyzing the variation characteristics in the generated images, achieving efficient and accurate structural variation prediction, which is superior to traditional prediction algorithms in multiple indicators. The image encoding scheme combining RD, DRP, and SR three kinds of data is both scientific and effective, and significantly improves the classification performance of the model. Description of the Drawings

[0016] Figure 1 This is the flowchart of the gene structure variation image generation process of the present invention;

[0017] Figure 2 This is the schematic diagram corresponding to the three alignment strategies in the specific implementation manner of the present invention;

[0018] Figure 3 This is the schematic diagram of the structure variation image splicing strategy in the specific implementation manner of the present invention;

[0019] Figure 4 This is the schematic diagram showing the picture effects of the DEL type and DUP type in the specific implementation manner of the present invention;

[0020] Figure 5 This is the schematic diagram showing the enhanced effects of three pictures in the specific implementation manner of the present invention;

[0021] Figure 6 This is the schematic diagram of the principle of the prior art Swin-Transformer;

[0022] Figure 7 This is the schematic diagram showing the change of the loss rate of the Swin-Transformer model with the number of Epochs in the specific implementation manner of the present invention;

[0023] Figure 8 This is the schematic diagram showing the change of the accuracy rate with the number of Epochs during the training process of the Swin-Transformer model in the specific implementation manner of the present invention;

[0024] Figure 9 This is the schematic diagram of the ROC curve of the test set of the Swin-Transformer in the specific implementation manner of the present invention;

[0025] Figure 10 This is the schematic diagram of the confusion matrix of the test set of the Swin-Transformer in the specific implementation manner of the present invention;

[0026] Figure 11 This is the schematic diagram comparing the accuracy, precision, recall rate, and F1 score in the specific implementation manner of the present invention. Specific implementation manner

[0027] Regarding the foregoing and other technical contents, features, and effects of the present invention, they will be clearly presented in the following detailed description of the embodiments with reference to the accompanying drawings of the present application. The contents mentioned in the following embodiments are all referenced to the drawings of the specification.

[0028] Example 1. In this example, the overall algorithm for image generation in the wheat genome structural variation detection method based on deep learning is designed: The algorithm encodes three types of candidate variant interval data, namely Read Depth (RD), Discount ReadPair (DRP), and Split Read (SR), through the RGB color mode to generate structural variation images. In this way, the relevant information of structural variation is presented intuitively, and the output images can clearly show the characteristic information of the variant regions. The workflow of the algorithm is as Figure 1 shown. The core task is to effectively express the representative features of structural variation in an image of limited size. To achieve this goal, first, it is necessary to analyze the characteristic manifestations of the three data types of RD, DRP, and SR under the conditions of genomic deletion and tandem duplication variations, and then encode them into a three-channel image tensor. The image generation process includes steps such as data analysis and extraction, drawing and splicing of structural variation images. By expressing the gene sequence alignment information in the form of images, not only the complexity of extracting variant information from BAM data is reduced, but also an intuitive genomic input is provided for the deep learning model, thereby improving the accuracy of structural variation prediction, as shown in Table 1:

[0029] Table 1 Data types and principles used in structural variation

[0030]

[0031] Example 2. In this example, the sample data is preprocessed on the basis of Example 1. The sequencing data used in this study is from the "Chinese Spring Genome Sequencing and Assembly" project, and the reference genome selects the latest version of the Chinese Spring reference genome sequence IWGSC RefSeq v2.1 released by the International Wheat Genome Sequencing Consortium (IWGSC). The FASTQ files generated by the sequencer are aligned with this reference genome (iwgsc_refseqv2.1_assembly.fa) using BWA to obtain a SAM file containing variant information. Subsequently, the samtools tool is used to convert the SAM file into a binary format BAM file. After preprocessing the BAM file and combining the Read Depth (RD), Discount ReadPair (DRP), and Split Read (SR) data, structural variation images are generated. These images are then used as input data and fed into the deep learning model for variant prediction.

[0032] BAM files contain sequence alignment positions and quality information, enabling precise mapping to the reference genome. However, BAM files themselves cannot directly reveal variations. Therefore, it is necessary to extract potential variant sites and convert them into VCF format for annotation, classification, and comparison. VCF files record information such as the positions, types, and genotypes of variant sites, facilitating subsequent analysis. However, VCF files lack detailed regional annotations. Therefore, they are converted into BED format to facilitate the classification, annotation, and comparison of genomic regions, supporting structural variant analysis.

[0033] In the algorithm module of this study, the converted VCF files generate BED files containing candidate variant regions. Considering that predicted structural variants are usually larger than 50bp, the algorithm sets a threshold of 50bp to filter out variants smaller than this threshold, thus ensuring that the generated BED files contain only valid structural variant regions.

[0034] Example 3 designs an image generation strategy based on Example 2, including designing an image encoding method, defining the image coverage range, and selecting optimized image stitching rules to prepare the sample images for subsequent deep learning.

[0035] Example 4, as Figure 2 shown, designs the image encoding method based on Example 3. Since gene structural variations include insertions, deletions, inversions, and tandem duplications, and deletions and tandem duplications are the most common and easily detectable variant types by sequencing technology, this study focuses on predicting these two types of variations to improve the accuracy and reliability of prediction.

[0036] The principle of generating structural variant images is to convert the read fragments of three data types around the candidate variant regions into three-dimensional tensor images. Its main purpose is to map the alignment information in the BAM files into images, thereby showing the distribution of the three data types, namely RD, DRP, and SR, in the images. The key challenge lies in how to effectively integrate these three types of data into the images. The prediction target of this study is structural variants larger than 50bp, focusing on large-scale fragment data alignment anomalies. In terms of image generation design, the variant image generation algorithm is based on the RGB color model, encoding the three data types, RD, DRP, and SR, with different colors respectively. The specific image generation algorithm is described as follows:

[0037] (1) The R channel is used to represent the read fragment depth (RD) data, and its channel value is set to a. If the base at this position is covered, the value of the R channel, a, is 1; otherwise, the value is 255. Therefore, by observing the image of the R channel, the overall trend of the read coverage depth in the candidate structural variant site region can be intuitively shown.

[0038] (2) The G channel is used to represent discordant read pair (DRP) data, and its channel value is set to b. If the base at this position is covered and the data type is DRP, the value b of the G channel is 1; otherwise, the value is 255. Therefore, viewing the image of the G channel can clearly reflect the quantity and distribution of DRP data types of qualified abnormal alignments within the image area.

[0039] (3) The B channel is used to represent split read (SR) data, and its channel value is set to c. If the base at this position is covered and the data type is SR, the value c of the B channel is 1; otherwise, the value is 255. Therefore, viewing the image of the B channel can visually display the quantity and distribution of SR data types of qualified abnormal alignments within the image area.

[0040] The mapping of the three data types is as Figure 2 shown. The BAM file is equivalent to the canvas of the image, and the data types of each candidate variant interval are presented in the image with corresponding colors through the BED file.

[0041] The pixel color in the image represents the base coverage of the corresponding coordinate. White (255, 255, 255) represents the background color, cyan (1, 255, 255) represents that the base at this coordinate has been covered by the RD type; blue (1, 1, 255) represents that the base coverage at this position comes from the RD and DRP types; green (1, 255, 1) represents that the base coverage at this position comes from the RD and SR types; black (1, 1, 1) represents that this position meets all characteristics. The background color of the image is defaulted to white (255, 255, 255). When selecting an image, ensure that the size of each picture is consistent.

[0042] Example 5. Based on Example 4, the image coverage range is divided. Since structural variations usually have a relatively large coverage range, which can range from 50 bp to tens of thousands of bp, and even for the same type of variation, there may be significant differences in their coverage ranges. For each candidate variant interval, the specific positions of the left and right breakpoints can be calculated through formulas (1) and (2):

[0043] Left = start - 0.1 × candidate_len (1)

[0044] Right = end + 0.1 × candidate_len (2)

[0045] Among them, Left and Right respectively represent the left break point and the right break point positions of the structural variation image, and the range of the finally generated structural variation image is (Left, Right). During this process, start and end respectively represent the starting positions of the candidate variation intervals in the BED file, and candidate_len is the length of the candidate variation, that is, the difference between end and start. To ensure that the integrity of the variation region is not damaged when intercepting the variation interval, Formulas (1) and (2) stipulate that 0.1 times of candidate_len is added to both the left and right sides, which can ensure that the generated structural variation image has a certain degree of fault tolerance and avoid damaging the structure of the candidate variation interval.

[0046] Example 6. Based on Example 5, the selection and optimization of the image splicing rules are carried out. For the start and end break points of each candidate variation interval, two regional images near the break points are respectively formed, and the break points are included in the figure. The splicing rules for the left and right images can be divided into the following three types:

[0047] (1) Splicing along the channel direction (i.e., 3-channel input tensor);

[0048] (2) Horizontally splicing along the abscissa direction;

[0049] (3) Vertically splicing along the ordinate direction.

[0050] In order to adapt to the input format of the model and avoid the sequence range of image coverage being forced to be compressed, it is decided to adopt the last splicing form. When performing left-right flipping and amplification based on this splicing form, the positions of the upper and lower images need to be exchanged at the same time, and finally a rectangular image of 224×224 pixels is output, as Figure 3 shown.

[0051] As Figure 4 shown, it can be seen that the variation characteristics of DUP are concentrated in the 1st and 3rd quadrants, and the covered areas of cyan, blue, and black are relatively large. While the variation characteristics of DEL are concentrated in the 2nd and 4th quadrants, and the covered areas of cyan, blue, and black are relatively small.

[0052] Example 7. As Figure 5 shown, based on Example 6, the design and optimization of the deep learning-based structural variation prediction model are carried out. Through the operations of Examples 1-6, a total of about 500GB of BAM files are obtained, and after converting them into VCF files, a BED file containing 86084 pieces of variation information is generated. Among them, the number of deletion variations and tandem duplication variations are 79376 and 6708 respectively. For each candidate variation interval, a corresponding type of structural variation image can be generated.

[0053] When the data volumes of the two mutation types vary significantly, performing a binary classification task may lead to poor classification performance for the smaller class. This is because during training, the model tends to predict samples of the larger class, thus ignoring samples of the smaller class. To address this issue, the solution proposed in this study is to enhance the image data, including random rotation, Gaussian blur, and adjustments to brightness, contrast, and saturation. The effects after processing are as Figure 5 shown.

[0054] By rotating, applying Gaussian blur, adjusting brightness and contrast to the training images, the diversity of the dataset has been greatly enhanced. This enables the model to encounter more diverse samples during training, thereby enhancing its generalization ability, especially being able to perform more stably when facing complex scenarios or imperfect data. In addition, for the problem of class imbalance, data augmentation techniques effectively increase the number of samples in the minority classes, improving the model's recognition ability for these classes, and ultimately contributing to improving the overall classification performance.

[0055] Example 8, as Figures 6 - 8 shown, an experiment was carried out on the basis of Examples 1 - 7. 12,500 pictures were used in the experiment, with half of the pictures being of the DEL type and half being of the DUP type. The dataset was randomly divided into an 80% training set, a 10% validation set, and a 10% test set. It should be noted that only the data in the training set underwent data augmentation, while the data in the validation set and the test set were not processed with data augmentation. For the training of the deep learning model, considering factors such as data volume, experimental environment, and computing resources, the batch size was set to 32, the number of training epochs was set to 50, the activation function used was ReLU, the learning rate was set to 0.001, and the Adam optimization algorithm was used for gradient update. Except for differences in the model network parameters and the preprocessing methods of the image input, the training and validation processes were basically the same. The experimental process was divided into:

[0056] (1) Image preprocessing and feature extraction. First, the input images were fed into the Patch Partition module for block processing. In this module, each image was divided into several adjacent 4x4 pixel blocks (Patches), and then flattened in the channel direction. Assuming the input image is an RGB three-channel image, each Patch contains 4x4 = 16 pixels, and each pixel contains three values of R, G, and B. Therefore, after flattening, the feature dimension of each Patch is 16x3 = 48. After passing through the Patch Partition module, the shape of the image changes from the original [H, W, 3] to [H / 4, W / 4, 48], where H and W are the height and width of the image respectively.

[0057] Next, the image undergoes a linear transformation of the channel data for each pixel through the Linear Embedding layer, mapping the 48-dimensional features of each pixel to C dimensions. At this time, the shape of the image changes from [H / 4, W / 4, 48] to [H / 4, W / 4, C], where C is a hyperparameter representing the output feature dimension of each Patch. In fact, in the implementation, Patch Partition and Linear Embedding are completed through a convolutional layer, which is similar to the Embedding layer structure in the traditional Vision Transformer.

[0058] (2) Multi-scale construction of the feature map. During the feature extraction process, feature maps of different sizes are gradually constructed through four Stages. In Stage 1, the image is first processed through a Linear Embedding layer; while in the subsequent three Stages, the image in each stage is first downsampled through the Patch Merging layer (which will be introduced in detail later). Then, each Stage further extracts features by stacking multiple Swin Transformer Blocks.

[0059] It should be noted that when stacking Swin Transformer Blocks, two different structures are used, as shown in (b) of Figure 6 . The main difference between the two structures is that one uses the W-MSA (Window-based Multi-Head SelfAttention) structure, while the other uses the SW-MSA (Shifted Window-based Multi-HeadSelfAttention) structure. Moreover, these two structures are used in pairs, that is, the W-MSA structure is used first, and then the SW-MSA structure is used. Therefore, when constructing the Swin Transformer, the number of stacked Blocks is usually even (because each pair of Blocks includes one W-MSA and one SW-MSA).

[0060] (3) Classification and final output. After completing multiple Transformer Blocks, the network is connected to a LayerNormalization layer, followed by global pooling. Finally, the final classification result is output through a fully connected layer. Although these modules are not specifically shown in the diagram, they are indispensable components in the source code implementation.

[0061] Since the Swin-Transformer introduced a spatial attention mechanism in the image classification task and demonstrated significant advantages, after comparing six other deep learning models that performed well in image classification - including AlexNet, GoogleNet, EfficientNet, ShuffleNet, RegNet, and ResNet - the corresponding experimental results were obtained, as shown in Table 2:

[0062] Table 2 Experimental results of each model on structurally variant images

[0063] model Accuracy(%) Precision(%) Recall(%) F1Score(%) AlexNet 94.32 94.32 95.20 94.37 GoogleNet 96.80 97.14 96.32 96.94 EfficientNet 98.00 98.39 97.60 98.14 ShuffleNet 97.92 98.33 97.60 98.16 RegNet 98.48 98.72 98.24 98.71 ResNet 99.12 98.89 99.36 99.12 Swin-Transformer 99.92 99.84 100 99.84

[0064] The results shown in Table 2 indicate that EfficientNet, ShuffleNet, RegNet, ResNet, and Swin-Transformer performed relatively well in terms of accuracy, all achieving over 98%. Specifically, the accuracies of ResNet and Swin-Transformer were 98.89% and 99.84% respectively, with similar performance. In terms of recall rate and F1 score, ResNet and Swin-Transformer had higher scores, and Swin-Transformer performed best overall. Therefore, considering the three metrics of accuracy, precision, and F1 score, Swin-Transformer was superior to other classification models in the final result. As the number of training epochs increased, the loss change trends of Swin-Transformer on the training set and validation set were as Figure 7 shown, demonstrating the training effect of this deep learning model in the structurally variant image classification task.

[0065] From Figures 6 - 8 as shown, it can be observed from the loss rate curve that the loss rates of the training set and validation set began to converge when the training reached the 8th epoch, and by the 50th epoch, the loss rates had basically stabilized.

[0066] Example 9, based on Example 8 as Figures 9 - 10 shown, analyzed the model training results and selected the performance of the optimal model obtained from the 8th epoch of training on the test set. From Figure 9 it can be seen that the ROC curves and AUC values of all models performed well. The overall ROC curve was close to the ideal coordinate point (FPR = 0, TPR = 1), indicating that the models had very strong ability to distinguish positive and negative samples. At the same time, the AUC values under the ROC curves of each model were also close to 1. As Figure 10The confusion matrix in shows the classification results of the Swin-Transformer model on the test set, where the true positives (TP) are 625, the false negatives (FN) are 0, the false positives (FP) are 1, and the true negatives (TN) are 624. These data indicate that the Swin-Transformer model performs excellently in the classification task, with almost no misclassifications and extremely high accuracy. This further verifies that the Swin-Transformer model outperforms other comparison models in overall performance, especially achieving excellent results in terms of recall rate and precision. Therefore, the Swin-Transformer model is considered the most outstanding model in this study.

[0067] Example 10, as Figure 11 shown, on the basis of Example 9, in order to explore the importance of three data types, namely RD, DRP, and SR, in the classification of structural variation images and whether they are the key factors for improving classification performance, experiments were carried out again by combining three algorithms based on different data types: CNVnator, BreakDancer, and Pindel. In the experiment, a fixed network model - Swin-Transformer was used, and the preprocessing process of the training and validation data sets, network parameter settings, etc. were kept consistent with the previous section. The original data set used in the comparative experiment in this subsection comes from the "Chinese Spring Genome Sequencing and Assembly" project (PRJNA392179).

[0068] As Figure 11 shown, it presents the comparison results between the prediction algorithm combining three types of feature information and the prediction algorithm using only a single type of data. The experimental results show that when using only the DRD data information extracted by the BreakDancer algorithm alone, the prediction effect is significantly worse than that of the CNVnator algorithm using RD data and the Pindel algorithm using SR data. However, after integrating the three types of data, RD, DRP, and SR, the algorithm presented in image format shows the best prediction effect. From the above comparative experiment, it can be seen that the algorithm combining these three types of data achieves the best performance in terms of prediction accuracy, which also verifies the improvement of the strategy proposed in this study in all aspects.

[0069] In summary, the framework proposed in this study based on sequence conversion images and deep learning methods is superior to traditional prediction algorithms in multiple metrics. The image coding scheme combining the three types of data, RD, DRP, and SR, is both scientific and effective, and significantly improves the classification performance of the model.

[0070] The above description is only for the purpose of illustrating the present invention. It should be understood that the present invention is not limited to the above embodiments, and various flexible forms conforming to the idea of the present invention are within the protection scope of the present invention.

Claims

1. A method for detecting wheat genome structural variation based on deep learning, characterized in that: The following steps are involved: The FASTQ file generated by the sequencing data of the sequencer is compared with the reference genome using BWA software to obtain the SAM file containing the variation information. Subsequently, the samtools tool is used to convert the SAM file into a binary BAM file. After preprocessing the BAM file and combining it with RD (read depth data), DRP (discordant read pair data) and SR (split read data), structural variation images are generated. These images are input into the deep learning model for variation prediction training.

2. The method for detecting wheat genome structural variation based on deep learning according to claim 1, characterized in that: The preprocessing of the BAM file specifically includes: converting the BAM file into a VCF format through a variation identification tool to extract candidate variation site information, and then converting the VCF format file into a BED format for genome coordinate positioning. The BED file after format conversion can be combined with a genome annotation database to perform refined region annotation based on genome coordinates, and candidate structural variations with a length of less than 50 bp are filtered out by setting a threshold, so as to finally obtain a standardized BED file containing only valid structural variation regions.

3. The method for detecting wheat genome structural variation based on deep learning according to claim 2, characterized in that: The generation process of the structural variation image includes: designing an image encoding method, determining an image coverage range, and selecting and optimizing image splicing rules.

4. The method for detecting wheat genome structural variation based on deep learning according to claim 3, characterized in that: The design of the image encoding method specifically includes converting the read fragments of the three data types around the candidate variation region in the BAM file into a three-dimensional tensor image, that is, mapping the variation information in the BAM file to the image, and drawing the distribution of the three data types of RD, DRP and SR in the image. The image generation is based on the RGB color mode, and the three data of RD, DRP and SR are encoded with different colors respectively. The image generation algorithm specifically includes: The R channel is used to represent RD, and its channel value is set to a. If the base at this position is covered, the value a of the R channel is 1, otherwise, the value is 255; The G channel is used to represent DRP, and its channel value is set to b. If the base at this position is overwritten and the data type is DRP, the value b of the G channel is 1, otherwise, the value is 255; The B channel is used to represent SR, and its channel value is set to c. If the base at this position is overwritten and the data type is SR, the value c of the B channel is 1, otherwise, the value is 255; The pixel color in the image indicates the base coverage of the corresponding coordinates. The white color value (255,255,255) is the default background color. The cyan color value (1,255,255) indicates that the base at the coordinate has been covered by the RD type; the blue color value (1,1,255) indicates that the base coverage at this position comes from the RD and DRP types; the green color value (1,255,1) indicates that the base coverage at this position comes from the RD and SR types; and the black color (1,1,1) indicates that the position meets all the features.

5. The method for detecting wheat genome structural variation based on deep learning according to claim 3, characterized in that: The determination of the image coverage specifically includes that for each candidate variant region in the BAM file, the specific positions of the left and right breakpoints can be calculated by formulas (1) and (2): Left=start-0.1×candidate_len (1) Right=end+0.1×candidate_len (2) Among them, Left and Right represent the left and right breakpoint positions of the structural variation image respectively. The range of the final generated structural variation image is (Left, Right). In this process, start and end represent the starting position of the candidate variation interval in the BED file respectively, and candidate_len is the length of the candidate variation, that is, the difference between end and start. In order to ensure that the integrity of the variation region is not destroyed when intercepting the variation interval, formulas (1) and (2) stipulate that candidate_len is increased by 0.1 times on the left and right sides respectively. This can ensure that the generated structural variation image has a certain fault tolerance and avoid damaging the structure of the candidate variation interval.

6. The method for detecting wheat genome structural variation based on deep learning according to claim 3, characterized in that: The selection of the image stitching rules includes selecting a longitudinal stitching method along the vertical coordinate direction to adapt to the input format of the subsequent deep learning model and avoid forced compression of the sequence range of image stitching, and when performing left-right flipping and amplification based on the longitudinal stitching method, the positions of the upper and lower images need to be swapped at the same time.

7. The method for detecting wheat genome structural variation based on deep learning according to claim 3, characterized in that: The variation prediction training of the deep learning model includes selecting Swin-Transformer, AlexNet, GoogleNet, EfficientNet, ShuffleNet, RegNet and ResNet as deep learning models, and enhancing the generated structural variation images as training image data by rotating, Gaussian blurring, and adjusting brightness and contrast, thereby diversifying the training samples and enhancing the recognition ability of the deep learning model.