A data analysis method that integrates targeted sequencing and whole-genome sequencing

By combining data analysis methods with targeted sequencing and whole genome sequencing, the problems of high efficiency and cost of variant detection in the prior art are solved, and efficient and accurate variant detection and breeding data utilization are achieved.

CN119541625BActive Publication Date: 2025-08-29CHANGSHA BAIAOYUN DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411524235.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-30
Publication Date
2025-08-29
Estimated Expiration
2044-10-30

AI Technical Summary

Technical Problem

In the prior art, targeted sequencing and whole genome sequencing each have limitations in the breeding process, and cannot detect variants comprehensively and efficiently, resulting in high detection costs and low efficiency.

Method used

Targeted sequencing and whole-genome sequencing are combined, and through data splitting, comparison, quality control and biometric analysis, common variant sites are extracted and consistency tests are performed to eliminate wrong data, and the accuracy and efficiency of variant detection are improved.

Benefits of technology

Efficient and accurate variation detection is achieved, which reduces detection costs and improves data utilization in the breeding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119541625B_ABST
    Figure CN119541625B_ABST
Patent Text Reader

Abstract

The present invention discloses a data analysis method integrating targeted sequencing and whole genome sequencing, comprising: obtaining sequencing data information; judging whether the sequencing data can be split and identified, and if so, splitting the sequencing data into targeted sequencing data and WGS sequencing data; comparing the targeted sequencing data and the WGS sequencing data respectively to obtain a targeted comparison result and a WGS comparison result; merging the targeted comparison result and the WGS comparison result to obtain a first comparison result, and performing variation detection and quality control on the first comparison result to obtain high-quality variation data; if not, sending the sequencing data to a preset second module to obtain a second comparison result, performing variation detection and quality control on the second comparison result to obtain high-quality variation data; annotating the high-quality variation data and sending it to a preset management terminal for downstream breeding applications. The present invention effectively improves the variation detection effect of sequencing data by combining targeting and whole genome sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of biomedicine, agricultural breeding, and more specifically, to a data analysis method that integrates targeted sequencing and whole-genome sequencing. Background Art

[0002] Currently, the most commonly used genotyping methods for plant and animal breeding include targeted sequencing and whole-genome sequencing. Targeted sequencing is more targeted and can quickly screen for promising loci in breeding materials, but it cannot reveal the genetic background across the entire genome. Whole-genome sequencing, on the other hand, provides more comprehensive variant detection and can be used as background for hybridization and backcrossing breeding. Combining these two detection technologies, generating both targeted and whole-genome sequencing reads in a single sequencing run, and leveraging the complementary nature of the two sequencing data, allowing for mutual verification and effective supplementation through data analysis, can significantly improve the efficiency and accuracy of variant detection and reduce the cost of genetic testing.

[0003] Therefore, the existing technology has defects and needs to be improved urgently. Summary of the Invention

[0004] In view of the above problems, the purpose of the present invention is to provide a data analysis method that integrates targeted sequencing and whole genome sequencing, which can effectively improve the variation detection effect of sequencing data.

[0005] The present invention provides a data analysis method that integrates targeted sequencing and whole genome sequencing, comprising:

[0006] Obtain sequencing data information;

[0007] Determine whether the sequencing data can be split and identified. If so, split the sequencing data into targeted sequencing data and WGS sequencing data;

[0008] Compare the targeted sequencing data and the WGS sequencing data respectively to obtain the targeted alignment results and the WGS alignment results;

[0009] The targeted alignment results and the WGS alignment results are merged to obtain the first alignment result, and the first alignment result is subjected to variant detection and quality control to obtain high-quality variant data;

[0010] If not, the sequencing data is sent to a preset second module to obtain a second alignment result, and the second alignment result is subjected to variation detection and quality control to obtain high-quality variation data;

[0011] Annotate high-quality variant data and send them to the preset management terminal for downstream breeding applications.

[0012] In this solution, after obtaining the targeted alignment results and the WGS alignment results, the following steps are also included:

[0013] Based on the preset first software tool, all variant sites in the targeted comparison results and the WGS comparison results are extracted to obtain the targeted detection results and the WGS detection results;

[0014] Intersect the targeted detection results and WGS detection results to obtain the common variant sites;

[0015] Based on the preset second software tool, the common variant sites are tested for consistency at the site level and the preset sample level. If they are consistent, the corresponding WGS test results are accurate;

[0016] If there is any inconsistency, the corresponding WGS test result is wrong, and the inconsistent position is marked as a prompt.

[0017] In this solution, the step of splitting the sequencing data into targeted sequencing data and WGS sequencing data specifically includes:

[0018] Extract the adapter sequence of reads in sequencing data;

[0019] Compare and analyze the adapter sequences and index tags of reads in the sequencing data with the preset targeted sequencing library and WGS sequencing library adapter sequences respectively;

[0020] If the index tag of the adapter sequence of the reads in the sequencing data is the same as the index of the adapter sequence in the targeted sequencing library, the corresponding sequencing reads will be set as the targeted sequencing data;

[0021] If the index tag of the adapter sequence of the reads in the sequencing data is the same as the index of the adapter sequence in the WGS sequencing library, the corresponding sequencing reads are set as WGS sequencing data.

[0022] This plan also includes:

[0023] Extracting a comparison index and a corresponding comparison index value from the first comparison result or the second comparison result;

[0024] Multiply the comparison index value by the preset weight coefficient of the corresponding comparison index to obtain the corresponding comparison index proportion value;

[0025] Accumulate the proportions of all comparison indicators to obtain the accuracy index of the first comparison result or the second comparison result;

[0026] If the accuracy index of the first alignment result or the second alignment result is less than the preset accuracy index threshold, the reads that are not aligned or aligned incorrectly are eliminated, and the accuracy index of the first alignment result or the second alignment result is recalculated until the accuracy index is greater than the preset threshold, thereby obtaining the quality-controlled reads;

[0027] The reads after quality control are used as the final clean data for variant detection.

[0028] This plan also includes:

[0029] Obtain targeted and whole-genome alignment results, and calculate alignment rate, depth, and coverage;

[0030] If the alignment rate, depth, and coverage in the targeted alignment results are less than the preset thresholds, reads that are not aligned or aligned incorrectly in the targeted sequencing will be removed for subsequent variant detection;

[0031] If the alignment rate, depth, and coverage in the whole-genome alignment results are less than the preset thresholds, reads that are not aligned or aligned incorrectly in the whole-genome sequencing will be removed for subsequent variant detection.

[0032] In this solution, the quality control step also includes:

[0033] Extract reads from sequencing data that are not aligned to the reference genome;

[0034] The reads in the sequencing data that are not aligned to the reference genome are locally assembled and Blast aligned with the preset NT database to determine the source of the unaligned reads;

[0035] After aligning the read sources that are not aligned to the reference genome through the Blast NT database and the sequence after local assembly, if non-target species alignment information is obtained, the corresponding sequencing data is contaminated, and an invalid sequencing read warning message is generated. The invalid sequencing warning message is sent to the preset management terminal for prompting;

[0036] If the unaligned reads cannot obtain the homologous sequences of non-target species, the corresponding sequencing reads will be set as high-quality sequencing data.

[0037] In this solution, the quality control step also includes:

[0038] Extract the reads insert length recorded in the targeted alignment result bam;

[0039] If the reads insert fragment length is greater than the preset target probe sequence length, the corresponding reads will be deleted.

[0040] This plan also includes:

[0041] Based on the preset statistical software, the depth and coverage of the target interval and non-target interval are obtained;

[0042] According to the depth and coverage of the target interval and the non-target interval, the depth ratio and coverage ratio of the two intervals are obtained;

[0043] If the depth ratio of the two intervals is greater than the preset first ratio threshold and the coverage ratio is greater than the preset first ratio threshold, then the library is constructed once, that is, the sequencing data of the second module is normal.

[0044] In this solution, after obtaining high-quality variant data, the following steps are also included:

[0045] Extracting biometric values ​​corresponding to high-quality variant data;

[0046] Based on the same biometric feature, the difference between the biometric value of the high-quality variant data and the corresponding standard biometric value is calculated to obtain the corresponding biometric difference;

[0047] If the corresponding biometric feature difference is greater than the corresponding biometric feature difference threshold, the corresponding biometric feature is set as a variant feature;

[0048] The mutation features are annotated on all mutation sites in the corresponding high-quality mutation data.

[0049] In this solution, the breeding application step also includes:

[0050] Based on a preset third-party software tool, genotype phasing is performed on the variant sites to determine the target sample and match the corresponding reference panel based on the target sample.

[0051] Based on the preset genotype filling algorithm and the corresponding reference panel, the missing genotypes in the target sample are predicted to obtain high-density genotypes, and then filled or covered to obtain variant breeding gene data;

[0052] The variant breeding gene data is sent to a preset downstream end for breeding application.

[0053] The present invention discloses a data analysis method that integrates targeted sequencing and whole-genome sequencing. By combining targeted and whole-genome sequencing, the variation detection effect of sequencing data is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 A flow chart of a data analysis method integrating targeted sequencing and whole genome sequencing is shown;

[0055] Figure 2 A flowchart showing the steps of splitting sequencing data into targeted sequencing data and WGS sequencing data according to the present invention is shown. DETAILED DESCRIPTION

[0056] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present application and the features therein can be combined with each other.

[0057] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.

[0058] Figure 1 A flow chart of a data analysis method integrating targeted sequencing and whole genome sequencing is shown.

[0059] like Figure 1 As shown, the present invention discloses a data analysis method integrating targeted sequencing and whole genome sequencing, comprising:

[0060] S101, obtaining sequencing data information;

[0061] S102, determining whether the sequencing data can be split and identified, and if so, splitting the sequencing data into targeted sequencing data and WGS sequencing data;

[0062] S103, aligning the targeted sequencing data and the WGS sequencing data, respectively, to obtain a targeted alignment result and a WGS alignment result;

[0063] S104, merging the targeted alignment result and the WGS alignment result to obtain a first alignment result, and performing variant detection and quality control on the first alignment result to obtain high-quality variant data;

[0064] S105: If not, the sequencing data is sent to a preset second module to obtain a second alignment result, and the second alignment result is subjected to variation detection and quality control to obtain high-quality variation data;

[0065] S106, annotate the high-quality variation data and send it to the preset management terminal for downstream breeding applications.

[0066] According to an embodiment of the present invention, the data to be sequenced are divided into one type of data to be sequenced that can distinguish between two sequencing reads and one type of data to be sequenced that cannot distinguish between two sequencing reads, according to whether they can be identified. The data to be sequenced that can distinguish between two sequencing reads are split into targeted sequencing data and whole genome (WGS) sequencing data; the data to be sequenced that cannot distinguish between two sequencing reads are sent to a preset second module. By combining multiple sequencing methods, the range and effect of variation detection in the data to be sequenced are improved.

[0067] According to an embodiment of the present invention, after obtaining the targeted alignment result and the WGS sequencing alignment result, the method further includes:

[0068] Based on the preset first software tool, all variant sites in the targeted comparison results and the WGS comparison results are extracted to obtain the targeted detection results and the WGS detection results;

[0069] Intersect the targeted detection results and WGS detection results to obtain the common variant sites;

[0070] Based on the preset second software tool, the common variant sites are tested for consistency at the site level and the preset sample level. If they are consistent, the corresponding WGS test results are accurate;

[0071] If there is any inconsistency, the corresponding WGS test result is wrong, and the inconsistent position is marked as a prompt.

[0072] It should be noted that, in this embodiment, the preset first software tool is bcftools, and the preset second software tool is GATK. Furthermore, when recording the marked inconsistent positions in the WGS detection results, the marked number values ​​are recorded. If the marked number value is greater than the preset number threshold, the WGS sequencing data corresponding to the WGS detection results is deleted. If the marked number value is less than or equal to the preset number threshold, the corresponding inconsistent positions will be marked and displayed.

[0073] Figure 2 A flowchart showing the steps of splitting sequencing data into targeted sequencing data and WGS sequencing data according to the present invention is shown.

[0074] like Figure 2 As shown, in this embodiment of the present invention, the step of splitting the sequencing data into targeted sequencing data and WGS sequencing data specifically includes:

[0075] S201, extracting the adapter sequences of reads in the sequencing data;

[0076] S202, comparing and analyzing the adapter sequences and index tags of reads in the sequencing data with the preset targeted sequencing library and WGS sequencing library adapter sequences;

[0077] S203, if the index tag of the adapter sequence of the reads in the sequencing data is the same as the index of the adapter sequence in the targeted sequencing library, the corresponding sequencing reads are set as the targeted sequencing data;

[0078] S204: If the index tag of the adapter sequence of the reads in the sequencing data is the same as the index of the adapter sequence in the WGS sequencing library, the corresponding sequencing reads are set as WGS sequencing data.

[0079] It should be noted that the mixed sequencing library is composed of a mixture of a targeted sequencing library and a whole-genome sequencing library. When constructing the targeted sequencing library and the whole-genome sequencing library, the same sample is subjected to targeted sequencing and whole-genome sequencing respectively. The adapter sequences use two different index tags, and the header line of the index sequence information fastq file is added during the sequencing process. The result of the sequencing data after one sequencing on the machine contains reads of the two library construction results. The command line, such as seqkit grep, is used to separate the reads of the two index tags, thus obtaining two different sequencing data: targeted sequencing data and WGS sequencing data.

[0080] Extracting a comparison index and a corresponding comparison index value from the first comparison result or the second comparison result;

[0081] Multiply the comparison index value by the preset weight coefficient of the corresponding comparison index to obtain the corresponding comparison index proportion value;

[0082] Accumulate the proportions of all comparison indicators to obtain the accuracy index of the first comparison result or the second comparison result;

[0083] If the accuracy index of the first or second alignment result is less than the preset accuracy index threshold, reads that are not aligned or aligned incorrectly are removed, and the operation is repeated until the accuracy index is greater than the preset threshold to obtain quality-controlled reads;

[0084] The reads after quality control are used as the final clean data for variant detection.

[0085] It should be noted that the comparison indicators include alignment rate, sequencing depth, coverage, etc. For example, the comparison variation indicator is sequencing depth, the indicator parameter corresponding to the targeted sequencing reads is the total number of bases in the targeted sequencing reads, and the indicator parameter corresponding to the first genome to be tested is the size of the first genome to be tested. For example, the size of the first genome to be tested is 1Mb (million bases), and the total number of bases in the targeted sequencing reads is 10Mb, then the corresponding sequencing depth is 10X; the weight coefficient of each indicator is set in advance, and the accuracy index is used as the benchmark. When it is less than the preset accuracy index threshold, the data to be sequenced corresponding to the comparison result is set as variation data, and the comparison result includes the first comparison result and the second comparison result.

[0086] According to an embodiment of the present invention, the present invention includes:

[0087] Obtain targeted and whole-genome alignment results, and calculate alignment rate, depth, and coverage;

[0088] If the alignment rate, depth, and coverage in the targeted alignment results are less than the preset thresholds, reads that are not aligned or aligned incorrectly in the targeted sequencing will be removed for subsequent variant detection;

[0089] If the alignment rate, depth, and coverage in the whole-genome alignment results are less than the preset thresholds, reads that are not aligned or aligned incorrectly in the whole-genome sequencing will be removed for subsequent variant detection.

[0090] It should be noted that the sequencing accuracy can be further improved by screening the reads that are not mapped or mis-mapped in the targeted sequencing reads or whole-genome sequencing reads through the depth, coverage and alignment rate of the targeted and whole-genome range.

[0091] According to an embodiment of the present invention, the quality control step further includes:

[0092] Extract reads from sequencing data that are not aligned to the reference genome;

[0093] The reads in the sequencing data that are not aligned to the reference genome are locally assembled and Blast aligned with the preset NT database to determine the source of the unaligned reads;

[0094] After aligning the read sources that are not aligned to the reference genome through the Blast NT database and the sequence after local assembly, the non-target species alignment information is obtained. If the corresponding sequencing data is contaminated, an invalid sequencing read warning message is generated and sent to the preset management terminal for prompting;

[0095] If the unaligned reads cannot obtain the homologous sequences of non-target species, the corresponding sequencing reads will be set as high-quality sequencing data.

[0096] It should be noted that reads that are not mapped to the reference genome (Unmaped reads) are extracted for local assembly and Blast alignment is performed with the NT database to determine the source of these reads and detect whether the sample is contaminated.

[0097] According to an embodiment of the present invention, the quality control step further includes:

[0098] Extract the reads insert length recorded in the targeted alignment result bam;

[0099] If the reads insert fragment length is greater than the preset target probe sequence length, the corresponding reads will be deleted.

[0100] It should be noted that quality control of sequencing data includes quality assessment, removal of adapter sequences, filtering of low-quality sequences, removal of N-containing sequences, removal of low-complexity sequences, removal of repetitive sequences, etc. When the length of the reads inserted is greater than the preset target probe sequence length, the corresponding reads will be deleted. This can eliminate interference caused by sites with high similarity to the target probe sequence, thereby improving the accuracy of the sequencing data.

[0101] According to an embodiment of the present invention, the further embodiment includes:

[0102] Based on the preset statistical software, the depth and coverage of the target interval and non-target interval are obtained;

[0103] According to the depth and coverage of the target interval and the non-target interval, the depth ratio and coverage ratio of the two intervals are obtained;

[0104] If the depth ratio of the two intervals is greater than the preset first ratio threshold and the coverage ratio is greater than the preset first ratio threshold, then the library is constructed once, that is, the sequencing data of the second module is normal.

[0105] It should be noted that the target interval and non-target interval were counted using depth and coverage statistical software, and the depth and coverage ratios of the two intervals were calculated to confirm that the sequencing depth of the target interval was much higher than that of the non-target interval; when the depth ratio was the target interval divided by the depth of the non-target interval, the coverage ratio was the coverage of the target interval divided by the coverage of the non-target interval, and the preset first ratio threshold was greater than 1; when the depth ratio was the non-target interval divided by the depth of the target interval, the coverage ratio was the coverage of the non-target interval divided by the coverage of the target interval, and the preset first ratio threshold was less than 1.

[0106] According to an embodiment of the present invention, after obtaining high-quality variation data, the method further includes:

[0107] Extracting biometric values ​​corresponding to high-quality variant data;

[0108] Based on the same biometric feature, the difference between the biometric value of the high-quality variant data and the corresponding standard biometric value is calculated to obtain the corresponding biometric difference;

[0109] If the corresponding biometric feature difference is greater than the corresponding biometric feature difference threshold, the corresponding biometric feature is set as a variant feature;

[0110] The mutation features are annotated on all mutation sites in the corresponding high-quality mutation data.

[0111] It should be noted that the biological characteristics include the stress resistance and tolerance of the target sample corresponding to the sequencing data, and the corresponding biological characteristic value is the value range of the corresponding stress resistance, such as the adapted temperature range and light intensity range.

[0112] According to an embodiment of the present invention, the breeding application step further includes:

[0113] Based on a preset third-party software tool, genotype phasing is performed on the variant sites to determine the target sample and match the corresponding reference panel based on the target sample.

[0114] Based on the preset genotype filling algorithm and the corresponding reference panel, the missing genotypes in the target sample are predicted to obtain high-density genotypes, and then filled or covered to obtain variant breeding gene data;

[0115] The variant breeding gene data is sent to a preset downstream end for breeding application.

[0116] It should be noted that by marking the variant sites through annotation and then filling or covering the annotated variant sites respectively through the preset genotype filling algorithm, it is possible to carry out foreground and background selection of breeding materials, molecular marker-assisted backcrossing, population structure, genetic diversity and evolution analysis, gene discovery, whole genome selection analysis, etc.

[0117] According to an embodiment of the present invention, the further embodiment includes:

[0118] Based on multiple sets of sequencing data for the same target sample, high-quality variant data of any sequencing data is extracted as reference data, and then compared and analyzed with high-quality variant data of other sequencing data in turn;

[0119] If the variant sites of the reference data and the high-quality variant data of other sequencing data are the same, the annotations of redundant variant features in the high-quality variant data of the reference data or other sequencing data are deleted;

[0120] If the variant sites of the high-quality variant data of the reference data and other sequencing data are different, when the variant features are the same, the different variant sites and the annotations on the corresponding variant sites are deleted; when the variant features are different, the different variant features are extracted and the corresponding different variant features are annotated on the corresponding different variant sites respectively;

[0121] Traverse all the sequencing data sets until there are no annotation changes, obtain the variation features annotated on the variant sites, and superimpose the same variation features to obtain the number of times the same variation feature annotated on the variant sites is present;

[0122] The probability value of the same variant feature annotated on the variant site is obtained by dividing the number of times the same variant feature is annotated on the variant site by the total number of times all variant features are annotated on the corresponding variant site;

[0123] If the probability value of the same variant feature annotated at the corresponding variant site is less than the preset probability threshold, the same variant feature will be deleted at the corresponding variant site;

[0124] If the probability value of the same variant feature annotated at the corresponding variant site is greater than or equal to the preset probability threshold, the probability value of the variant feature is annotated on the corresponding variant site.

[0125] It should be noted that the mutation characteristics of the detected high-quality mutations are annotated to facilitate the subsequent evaluation of their possible pathogenicity or biological significance. The high-quality mutation data of any sequencing data is used as reference data, and is compared and analyzed with the high-quality mutation data of other sequencing data in turn. If the annotation of the mutation characteristics in the mutation site is modified, the modified high-quality mutation data are then compared pairwise for annotation modification until there is no annotation change.

[0126] According to an embodiment of the present invention, after annotating the probability value of the variation feature on the corresponding variation site, the method further includes:

[0127] Extract the mutation features annotated above the mutation site;

[0128] Send the mutation features annotated on the mutation site to the preset evaluation terminal for evaluation, and obtain the evaluation score of the corresponding mutation site;

[0129] The evaluation value index of the corresponding variant site is obtained by multiplying the evaluation score of the corresponding variant site by the probability value of the variant feature annotated on the corresponding variant site;

[0130] The evaluation value indexes of the variant sites are marked in order from large to small, and the corresponding marks are sent to the preset genotype filling algorithm for orderly filling or covering.

[0131] It should be noted that in the same batch of sequencing data, there may be multiple mutation sites and the mutation characteristics are inconsistent. Therefore, it is necessary to select some higher-value mutation sites to prioritize downstream breeding applications.

[0132] The present invention discloses a data analysis method that integrates targeted sequencing and whole-genome sequencing. By combining targeted and whole-genome sequencing, the variation detection effect of sequencing data is effectively improved.

[0133] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0134] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0135] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0136] Those skilled in the art will appreciate that all or part of the steps of the above-mentioned method embodiments may be implemented by hardware associated with program instructions, and the aforementioned program may be stored in a computer-readable storage medium. When the program is executed, the program executes the steps of the above-mentioned method embodiments. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0137] Alternatively, if the integrated units described above are implemented as software modules and sold or used as standalone products, they can also be stored on a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product, stored on a storage medium, includes instructions for enabling a computer device (such as a personal computer, server, or network device) to execute all or part of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as removable storage devices, ROM, RAM, magnetic disks, or optical disks.

Claims

1. A data analysis method integrating targeted sequencing and whole genome sequencing, characterized in that: include: Obtain sequencing data information; Determine whether the sequencing data can be split and identified. If so, split the sequencing data into targeted sequencing data and WGS sequencing data; Compare the targeted sequencing data and the WGS sequencing data respectively to obtain the targeted alignment results and the WGS alignment results; The targeted alignment results and the WGS alignment results are merged to obtain the first alignment result, and the first alignment result is subjected to variant detection and quality control to obtain high-quality variant data; If not, the sequencing data is sent to a preset second module to obtain a second alignment result, and the second alignment result is subjected to variation detection and quality control to obtain high-quality variation data; Annotate high-quality variant data and send them to the preset management terminal for downstream breeding applications; The step of splitting the sequencing data into targeted sequencing data and WGS sequencing data specifically includes: Extract the adapter sequence of reads in sequencing data; Compare and analyze the adapter sequences and index tags of reads in the sequencing data with the preset targeted sequencing library and WGS sequencing library adapter sequences respectively; If the index tag of the adapter sequence of the reads in the sequencing data is the same as the index of the adapter sequence in the targeted sequencing library, the corresponding sequencing reads will be set as targeted sequencing data; If the index tag of the adapter sequence of the reads in the sequencing data is the same as the index of the adapter sequence in the WGS sequencing library, the corresponding sequencing reads will be set as WGS sequencing data; The preset second module is provided with a mixed sequencing library, and the mixed sequencing library is composed of a mixture of a targeted sequencing library and a whole genome sequencing library; The second alignment result includes alignment rate, sequencing depth and coverage.

2. The data analysis method integrating targeted sequencing and whole genome sequencing according to claim 1, characterized in that: After obtaining the targeted alignment results and the WGS alignment results, the method further includes: Based on the preset first software tool, all variant sites in the targeted comparison results and the WGS comparison results are extracted to obtain the targeted detection results and the WGS detection results; Intersect the targeted detection results and WGS detection results to obtain the common variant sites; Based on the preset second software tool, the common variant sites are tested for consistency at the site level and the preset sample level. If they are consistent, the corresponding WGS test results are accurate; If there is any inconsistency, the corresponding WGS test result is wrong, and the inconsistent position is marked as a prompt.

3. The data analysis method integrating targeted sequencing and whole genome sequencing according to claim 1, characterized in that: Also includes: Extracting a comparison index and a corresponding comparison index value from the first comparison result or the second comparison result; Multiply the comparison index value by the preset weight coefficient of the corresponding comparison index to obtain the corresponding comparison index proportion value; Accumulate the proportions of all comparison indicators to obtain the accuracy index of the first comparison result or the second comparison result; If the accuracy index of the first alignment result or the second alignment result is less than the preset accuracy index threshold, the reads that are not aligned or aligned incorrectly are eliminated, and the accuracy index of the first alignment result or the second alignment result is recalculated until the accuracy index is greater than the preset threshold, thereby obtaining the quality-controlled reads; The reads after quality control are used as the final clean data for variant detection.

4. The data analysis method integrating targeted sequencing and whole genome sequencing according to claim 1, characterized in that: Also includes: Obtain targeted and whole-genome alignment results, and calculate alignment rate, depth, and coverage; If the alignment rate, depth, and coverage in the targeted alignment results are less than the preset thresholds, reads that are not aligned or aligned incorrectly in the targeted sequencing will be removed for subsequent variant detection; If the alignment rate, depth, and coverage in the whole-genome alignment results are less than the preset thresholds, reads that are not aligned or aligned incorrectly in the whole-genome sequencing will be removed for subsequent variant detection.

5. The data analysis method integrating targeted sequencing and whole genome sequencing according to claim 1, characterized in that: The quality control step further includes: Extract reads from sequencing data that are not aligned to the reference genome; The reads in the sequencing data that are not aligned to the reference genome are locally assembled and Blast aligned with the preset NT database to determine the source of the unaligned reads; After aligning the read sources that are not aligned to the reference genome through the Blast NT database and the sequence after local assembly, if non-target species alignment information is obtained, the corresponding sequencing data is contaminated, and an invalid sequencing read warning message is generated. The invalid warning message is sent to the preset management terminal for prompting; If the unaligned reads cannot obtain the homologous sequences of non-target species, the corresponding sequencing reads will be set as high-quality sequencing data.

6. The data analysis method integrating targeted sequencing and whole genome sequencing according to claim 1, characterized in that: The quality control step further includes: Extract the reads insert length recorded in the targeted alignment result bam; If the reads insert fragment length is greater than the preset target probe sequence length, the corresponding reads will be deleted.

7. The data analysis method integrating targeted sequencing and whole genome sequencing according to claim 1, characterized in that: Also includes: Based on the preset statistical software, the depth and coverage of the target interval and non-target interval are obtained; According to the depth and coverage of the target interval and the non-target interval, the depth ratio and coverage ratio of the two intervals are obtained; If the depth ratio of the two intervals is greater than the preset first ratio threshold and the coverage ratio is greater than the preset first ratio threshold, then the library is constructed once, that is, the sequencing data of the second module is normal.

8. The data analysis method integrating targeted sequencing and whole genome sequencing according to claim 1, characterized in that: After obtaining high-quality variation data, the following steps are also included: Extracting biometric values ​​corresponding to high-quality variant data; Based on the same biometric feature, the difference between the biometric value of the high-quality variant data and the corresponding standard biometric value is calculated to obtain the corresponding biometric difference; If the corresponding biometric feature difference is greater than the corresponding biometric feature difference threshold, the corresponding biometric feature is set as a variant feature; The mutation features are annotated on all mutation sites in the corresponding high-quality mutation data.

9. The data analysis method integrating targeted sequencing and whole genome sequencing according to claim 1, characterized in that: The breeding application step also includes: Based on a preset third-party software tool, the variant sites are genotyped and phased, the target samples are determined, and the corresponding reference panels are matched according to the target samples; Based on the preset genotype filling algorithm and the corresponding reference panel, the missing genotypes in the target sample are predicted to obtain high-density genotypes, and then filled or covered to obtain variant breeding gene data; The variant breeding gene data is sent to a preset downstream end for breeding application.

Citation Information

Patent Citations

  • Targeted gene next-generation sequencing data somatic mutation detection method, terminal and medium

    CN116312780A

  • LP-WGS and DNA methylation-based lung cancer early screening model construction method and electronic equipment

    CN117275585A