Resequencing data analysis methods, electronic devices, and readable storage media

By generating a sample resequencing program, high-throughput whole-genome resequencing data of batch samples can be automatically analyzed, solving the problem of low analysis efficiency in existing technologies and realizing efficient automated analysis of batch samples.

CN117275584BActive Publication Date: 2026-05-26ZYBIO INC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZYBIO INC
Filing Date
2023-09-28
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

The current technology for analyzing high-throughput whole-genome resequencing data from batches of samples is inefficient, mainly because the analysis processes for different samples run independently, requiring users to set up independent sequencing logic for each sample, which is time-consuming.

Method used

By acquiring the raw sequencing data and workflow configuration data of the batch of samples to be analyzed, a sample resequencing program is generated. The high-throughput whole-genome resequencing data of the batch of samples is automatically analyzed using the resequencing workflow data, integrating the analysis workflow of multiple sequencing platforms and strategies.

Benefits of technology

It enables automated analysis of high-throughput whole-genome resequencing data from batches of samples, improving analysis efficiency and avoiding the tedious process of users setting up independent sequencing logic for each sample.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117275584B_ABST
    Figure CN117275584B_ABST
Patent Text Reader

Abstract

This application discloses a resequencing data analysis method, electronic device, and readable storage medium, applied to a resequencing data analysis platform. The resequencing data analysis method includes: acquiring raw sequencing data and workflow configuration data of a batch of samples to be analyzed, wherein the raw sequencing data is generated by the batch of samples to be analyzed on at least one sequencing platform based on at least one sequencing strategy; generating a sample resequencing program for the batch of samples to be analyzed based on the raw sequencing data and the workflow configuration data, wherein the sample resequencing program includes resequencing workflow data; and performing resequencing data analysis on the batch of samples to be analyzed by running the sample resequencing program based on the resequencing workflow data. This application solves the technical problem of low analysis efficiency for high-throughput whole-genome resequencing data of batch samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of gene sequencing technology, and in particular to a resequencing data analysis method, electronic device, and readable storage medium. Background Technology

[0002] With the rapid development of science and technology, gene sequencing technology is constantly iterating. Among them, whole genome sequencing (WGS) is widely used in fields such as molecular breeding of animals and plants, population evolution, clinical medicine and treatment, population evolution and genetic diseases due to its personalized and precise characteristics. At the same time, with the increasing richness of sequencing platforms and sequencing strategies, the continuous increase in sequencing throughput, the decreasing sequencing costs and the large accumulation of sequencing data, research on high-throughput whole genome resequencing data analysis for batch samples has become the focus of attention for industry professionals.

[0003] Currently, the analysis of high-throughput whole-genome resequencing data from batch samples typically requires setting up multiple sequencing strategies on multiple sequencing platforms. In other words, users need to set up independent sequencing logic for different samples. However, since the analysis processes for different samples run independently, users need to set up specific sequencing logic for all samples, which can easily lead to long analysis times for high-throughput whole-genome resequencing data from batch samples. Therefore, the current analysis efficiency for high-throughput whole-genome resequencing data from batch samples is low. Summary of the Invention

[0004] The main purpose of this application is to provide a resequencing data analysis method, electronic device, and readable storage medium, aiming to solve the technical problem of low analysis efficiency of high-throughput whole genome resequencing data of batch samples in the prior art.

[0005] To achieve the above objectives, this application provides a resequencing data analysis method, applied to a resequencing data analysis platform, the resequencing data analysis method comprising:

[0006] Obtain the raw sequencing data and process configuration data of the batch of samples to be analyzed, wherein the raw sequencing data is generated by the batch of samples to be analyzed on at least one sequencing platform based on at least one sequencing strategy;

[0007] Based on the original sequencing data and the process configuration data, a sample resequencing program for the batch of samples to be analyzed is generated, wherein the sample resequencing program includes resequencing process data.

[0008] Based on the resequencing process data, the batch of samples to be analyzed is resequencing data analysis by running the sample resequencing program.

[0009] To achieve the above objectives, this application also provides a resequencing data analysis device for use in a resequencing data analysis platform, the resequencing data analysis device comprising:

[0010] The acquisition module is used to acquire the raw sequencing data and process configuration data of the batch of samples to be analyzed, wherein the raw sequencing data is generated by the batch of samples to be analyzed on at least one sequencing platform based on at least one sequencing strategy;

[0011] A generation module is used to generate a sample resequencing program for the batch of samples to be analyzed based on the original sequencing data and the process configuration data, wherein the sample resequencing program includes resequencing process data.

[0012] The analysis module is used to perform resequencing data analysis on the batch of samples to be analyzed by running the sample resequencing program based on the resequencing process data.

[0013] This application also provides an electronic device comprising: at least one processor and a memory communicatively connected to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the resequencing data analysis method described above.

[0014] This application also provides a computer-readable storage medium storing a program for implementing a resequencing data analysis method, wherein when the program for the resequencing data analysis method is executed by a processor, it implements the steps of the resequencing data analysis method as described above.

[0015] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the resequencing data analysis method described above.

[0016] This application provides a resequencing data analysis method, electronic device, and readable storage medium. Specifically, it involves acquiring raw sequencing data and process configuration data of a batch of samples to be analyzed, wherein the raw sequencing data is generated by the batch of samples to be analyzed on at least one sequencing platform based on at least one sequencing strategy; generating a sample resequencing program for the batch of samples to be analyzed based on the raw sequencing data and the process configuration data, wherein the sample resequencing program includes resequencing process data; and performing resequencing data analysis on the batch of samples to be analyzed by running the sample resequencing program based on the resequencing process data.

[0017] In the process of analyzing high-throughput whole-genome resequencing data of a batch of samples to be analyzed, this application first obtains the original sequencing data and workflow configuration data of the batch of samples to be analyzed, then generates a sample resequencing program for the batch of samples to be analyzed based on the original sequencing data and workflow configuration data, and finally analyzes the resequencing data of the batch of samples to be analyzed by running the sample resequencing program based on the resequencing workflow data carried in the sample resequencing program. Since the resequencing data analysis platform can directly generate the sample resequencing program, the purpose of high-throughput whole-genome resequencing data analysis of the batch of samples to be analyzed can be achieved by running the sample resequencing program.

[0018] Since the sample resequencing program is generated based on the original sequencing data and the process configuration data, the sample resequencing program carries resequencing process data. Therefore, the process of the batch of samples to be analyzed can be integrated through the resequencing process data. Finally, through the sample resequencing program, the sequencing data of the batch of samples generated on multiple sequencing platforms based on various strategies can be automatically analyzed. That is, the resequencing data analysis platform can achieve the purpose of automatically analyzing the high-throughput whole genome resequencing data of the batch of samples to be analyzed by acquiring the original sequencing data and process configuration data.

[0019] Based on this, this application automatically generates a sample resequencing program for the batch of samples to be analyzed from the original sequencing data and workflow configuration data. Finally, by running the sample resequencing program, the purpose of automatically analyzing the high-throughput whole-genome resequencing data of the batch of samples to be analyzed is achieved. This avoids the need for users to set independent sequencing logic for each sample when analyzing high-throughput whole-genome resequencing data of batches. In other words, it overcomes the technical drawback of requiring users to set specific sequencing logic for all samples due to independent operation of the analysis workflow for different samples, which easily leads to long analysis times for high-throughput whole-genome resequencing data of batches. Therefore, it improves the analysis efficiency of high-throughput whole-genome resequencing data of batches. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1A flowchart illustrating the resequencing data analysis method provided in Embodiment 1 of this application;

[0023] Figure 2 This is a schematic diagram of light transmission during the turbidity test of the resequencing data analysis method provided in Embodiment 1 of this application;

[0024] Figure 3 This is a flowchart illustrating the resequencing data analysis method provided in Embodiment 2 of this application;

[0025] Figure 4 This is a schematic diagram of the resequencing data analysis device provided in Embodiment 3 of this application;

[0026] Figure 5 This is a schematic diagram of the structure of the electronic device provided in Embodiment 4 of this application.

[0027] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0028] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] Example 1

[0030] First, it's important to understand that whole-genome resequencing involves sequencing the genomes of different individuals within a species with known reference sequences. This allows for comparative analysis of structural differences in the genomes of different individuals or populations based on the sequencing data. However, most existing resequencing data analysis methods are only applicable to sequencing data generated by a single species, a single sequencing platform, or a single sequencing strategy. Furthermore, the degree of parameter customization in each analysis step is low, limiting the application scenarios. Additionally, the instability of the system during analysis leads to a lack of verification of the integrity of data input or output, reducing the accuracy of the analysis results. Simultaneously, there is insufficient statistical analysis and visualization of key indicators for each step; that is, users set independent sequencing logic for each sample in a batch, resulting in low automation in the analysis of high-throughput whole-genome resequencing data from batches. This leads to long analysis times for high-throughput whole-genome resequencing data from batches. Therefore, there is an urgent need for a method to improve the efficiency of analyzing high-throughput whole-genome resequencing data from batches.

[0031] This application provides a resequencing data analysis method, applied to a resequencing data analysis platform. In Embodiment 1 of this application's resequencing data analysis method, referring to... Figure 1 The resequencing data analysis method includes:

[0032] Step S10: Obtain the raw sequencing data and process configuration data of the batch of samples to be analyzed, wherein the raw sequencing data is generated by the batch of samples to be analyzed on at least one sequencing platform based on at least one sequencing strategy.

[0033] Step S20: Based on the original sequencing data and the process configuration data, generate a sample resequencing program for the batch of samples to be analyzed, wherein the sample resequencing program includes resequencing process data;

[0034] Step S30: Based on the resequencing process data, the batch of samples to be analyzed is resequencing data analyzed by running the sample resequencing program.

[0035] In this embodiment, it should be noted that, although Figure 1The logical order is shown, but in some cases, the steps shown or described may be performed in a different order than that shown here. Resequencing data analysis methods are applied to resequencing data analysis platforms, which are deployed on resequencing data analysis equipment. Specifically, the resequencing data analysis equipment can be a fully automated sequencer. Through this equipment, high-throughput whole-genome resequencing data analysis of batches of samples can be completed. Specifically, it can perform raw sequencing data quality control, reference genome alignment, redundant sequence marking or deletion, BQSR (Base Quality Score Recalibration), and VQSR (Variant Quality Score Recalibration). Multi-threaded automated analysis includes functions such as recalibration (variable quality value recalibration), variant detection, deletion of intermediate quality control files, integrity verification of data input or output at each step, and statistical analysis of key indicator information. The batch samples to be analyzed refer to batch samples awaiting high-throughput whole-genome resequencing data analysis. Specifically, these can be combinations of samples such as *E. coli*, yeast, rhizobium, or lactobacillus. It is understood that this application does not impose specific limitations on the batch samples or sequencing platform. The batch samples can be any combination of all known reference gene species. The sequencing platform can be any existing sequencing platform. The sequencing strategy can include single-end or paired-end strategies, etc. That is, the type of raw sequencing data can be SE (Single End) sequencing data or PE (Pair End) sequencing data. End (paired) sequencing data: Raw sequencing data is generated from the batch of samples to be analyzed using at least one sequencing platform and at least one sequencing strategy. Specifically, the raw sequencing data can be represented in a list file format. The list file has no limit on the number of lines, and each line contains information from the sequencing FASTQ file. Each column can use tabs as delimiters. For example, in one feasible approach, the raw sequencing data is a sequencing FASTQ file containing six columns of information: sample name, quality assessment system (phred), library name, sequencing chip number and lane number (Flowcell_Lane), sequencing chip number, lane number, and tag number (Flowcell_Lane). The storage path for the wcell_Lane_Barcode and the sequencing FASTQ file is specified. Sample names should ideally begin with a letter, avoiding numbers and special characters as much as possible. Spaces are prohibited. Sample names can be repeated across different lines; repetition indicates the resequencing data analysis platform is analyzing multiple sequencing FASTQ files from the same sample. Typically, a quality assessment system of 33 or 64 is used for subsequent base quality control, removing low-quality sequences and bases. Samples within the same batch to be analyzed should use the same quality assessment system; that is, the information in this column should be identical for all samples in a given list file.The library name is used to characterize the library constructed in the experiment, to identify the library information of the sequences, and for sequence grouping. The library name does not contain spaces. In the sequencing FASTQ file, library names can be repeated between different lines, meaning that multiple sequencing FASTQ files from the same library can be analyzed. The sequencing chip number and lane number are used to identify the sequencing information of the sequences and for sequence grouping. Spaces are prohibited. The sequencing chip number and lane number can be repeated for each line of a sequence segment, meaning that sequencing FASTQ files from different tags in the same chip and lane can be analyzed. The sequencing chip number, lane number, and tag number are used to identify the sequencing information of the sequences and for sequence grouping. Spaces are prohibited. The sequencing chip number, lane number, and tag number in each line of the sequencing FASTQ file must be unique and cannot be repeated between different lines. The storage path of the sequencing FASTQ file refers to the path of the folder containing the sequencing FASTQ files. The raw sequencing data can also be referred to as the "sample list file."

[0036] Additionally, it should be noted that the workflow configuration data is a configuration file used to characterize the sequencing workflow. Specifically, it can be a workflow configuration file. Users can refer to the instructions in the sample workflow configuration file for modification. After modification, the workflow configuration file is obtained. The workflow configuration file includes information and descriptions such as the reference sequence database, software, software analysis options, and running threads. Recommended default parameters are set for different information. Users can also interactively import relevant custom information for personalized analysis. In one feasible approach, the workflow configuration data specifically includes the project name, database path, reference genome, file storage path, analysis software, quality control functional parameters, model species parameters, software analysis parameters, and running threads. The parameters consist of several components. The project name can be a unique, automated analysis number for each project, ensuring that automated analyses from different projects do not interfere with each other. Furthermore, analysis execution permissions on the server can be granted based on the project name, effectively managing server resources. The database path is the storage path for the reference genome data. The default database path is configured with two E. coli strains, ATCC_10798 and E. coli_K12_MG1655, and Homosapiens, i.e., the fastq files and fast search FAI files, alignment index files, and dictionary files of the hg19 and hg38 versions of the reference genome of the model organism, as well as the hg... The system includes two versions of GATK_bundle mutation datasets: version 19 and version hg38. Users can also provide the folder path for reference genome data of other species for personalized analysis. The reference genome is a FAI file; the default reference genome is *E. coli*_K12_MG1655. Users can also provide reference genomes for other species, along with their fast search FAI files, alignment index files, and dictionary files, enabling personalized analysis for those species. The BIN path is the storage path for the analysis software used during the analysis process; users can specify other storage paths. The analysis software refers to the analysis software used during the analysis process; the configured default analysis software... The software package includes not only open-source software for sequencing data quality control, reference genome alignment, redundant sequence markers, BQSR, variant detection, and VQSR, but also self-developed auxiliary software for integrity checks of data input and output in each analysis step and for statistical analysis of data results. It provides visualized statistical results in readable tables and images. Furthermore, users can provide other versions of the default software to meet the personalized analysis needs of different software versions. The quality control function parameters characterize the quality control status of the batch of samples to be analyzed, i.e., they represent whether the sequencing data quality control function is enabled or disabled. It is set to enabled by default, but users can customize to disable the quality control function to meet the personalized analysis needs of quality-controlled data obtained from various channels.The model species parameter is used to determine whether a sample belongs to the model species (Homo sapiens) or a non-model species (Homo sapiens). The sample resequencing program will automatically select the sequencing strategy based on the model species parameter. That is, if the batch of samples to be analyzed does not belong to the model species Homo sapiens, BQSR will be performed based on the sample's own data, and VQSR will be performed by setting the variation index filtering conditions in the software analysis parameters; if the batch of samples to be analyzed belongs ... variation index filtering conditions in the software analysis parameters, and VQSR will be performed based on the variation index filtering conditions in the software analysis parameters. The analysis will perform BQSR based on the GATK_bundle mutation dataset in the database, and VQSR based on the variation indicators in the software analysis parameters. The software analysis parameters refer to the optional parameter settings of the analysis software used in the analysis process, mainly including: 1) Sequencing data quality control parameters: on the one hand, by using the default or providing a custom sequencing adapter FASTA file, sequencing data analysis from different platforms can be achieved; on the other hand, by setting quality control indicators such as sliding window size, quality value, and minimum sequence length, personalized analysis for different needs can be achieved; 2) Reference genome alignment parameters: by using the default or providing a custom alignment seed and alignment quality value, personalized analysis for different needs can be achieved; 3) BQSR parameters: by selecting multiple mutation datasets of the reference genome hg19 or hg38 version GATK_bundle in the database, or based on the variation detection results of the sample's own data, BQSR can be performed on different reference genome versions of the model species Homo sapiens and the non-model species Homo sapiens. 4) VQSR parameters: By selecting multiple mutation sets from the reference genome hg19 or hg38 version GATK_bundle in the database, or by setting filtering conditions such as deep variant confidence, coverage sequence quality, strand preference Fisher test p-value, and strand bias, VQSRs for the model species Homo sapiens and non-model species Homo sapiens with different reference genome versions can be achieved; Run thread parameters refer to the thread parameter settings of the analysis software used during the analysis process, mainly including: sequencing data quality control thread, reference genome alignment thread, and data redundancy marker thread. Users can increase the number of threads according to the available server thread resources, thereby improving analysis speed and shortening analysis time.

[0037] Additionally, it should be noted that after obtaining the prepared sample list file and workflow configuration file, the sample list file, workflow configuration file, and other information are entered into the preset sample resequencing main program. This other information may include sequencing platform information, sequencing strategy information, redundant sequence marking or deletion options, intermediate quality control file deletion options, and result output path. Specifically, the sequencing platform information is used to select the sequencing platform, which may include ILLUMINA, SLX, SOLEXA, SOLID, 454, LS454, COMPLETE, PACBIO, IONTORRENT, CAPILLARY, or HELICOS, etc. Other platforms can be selected as UNKNOWN, with ILLUMINA as the default. The sequencing strategy information can be PE or SE, i.e., paired-end sequencing or single-end sequencing, with SE as the default. The redundant sequence marking or deletion options retain or delete optical and PCR redundant sequences, with deletion as the default. The intermediate quality control file deletion option indicates whether to delete the sequencing data file after quality control of the original sequencing data, i.e., clean. The FASTQ file is deleted by default. The output path is the folder path for storing the output results. If the folder does not exist, the sample resequencing program will automatically create the folder path. In one feasible method, the specific process of generating the sample resequencing program is as follows: 1) The main program reads and checks whether the sample list file, workflow configuration file, and output path are provided. If not provided, the main program terminates and prompts the correct usage instructions; 2) The main program reads and checks the sequencing platform information, sequencing strategy information, redundant sequence marker or deletion options, and intermediate quality control file deletion options. If not provided, the main program uses the corresponding default information. If the provided information is incorrect, the main program terminates and prompts the correct usage instructions; 3) The main program reads the workflow configuration file (workflow configuration data) and checks whether the BIN path, database path, reference genome FASTA file and its fast access index FAI file, sequencing adapter FASTA file, mutation true set file in GATK_bundle and its fast access index TBI file exist.If the file does not exist, the main program will terminate and indicate that the missing file is missing. Additionally, if no information is provided in the configuration file, the default information will be used. If the workflow configuration file does not exist, the main program will terminate. 4) The main program reads the sample list file (raw sequencing data) and checks if the sequencing data file exists. If it does not exist, the main program will terminate and indicate that the sequencing data file to be analyzed does not exist. Based on the information in the sample list, the main program will automatically create an automatic analysis folder and multiple folders named after different samples in the output path. Furthermore, under each sample name folder, it will further create a script storage folder, a sequencing data quality control folder, an alignment folder, and a temporary folder. The main program is used to store scripts, quality control data, alignment and variant detection data, and temporary files output during the analysis process. On the other hand, by combining the read sequencing platform information, sequencing strategy information, redundant sequence marking or deletion options, intermediate quality control file deletion options, and workflow configuration files, the main program will automatically generate sequencing data quality control scripts, reference genome alignment scripts, redundant sequence marking or deletion scripts, variant detection scripts, intermediate quality control file deletion scripts, and script delivery order files in the script storage folder of each sample name folder. Finally, for the entire project, the main program will automatically generate all sample script delivery order files and automatic delivery execution scripts in the project automatic analysis folder. If the sample list file does not exist, the main program will terminate.

[0038] Additionally, it should be noted that the resequencing workflow data is used to characterize the resequencing data analysis process. Specifically, it can be the delivery and execution order of different script files. After the sample resequencing program is generated, the automatic delivery and execution scripts automatically generated in the project automatic analysis folder in the output path will automatically deliver and execute all scripts in the project according to the execution order in the script delivery order file of all samples, that is, run the sample resequencing program to complete the resequencing data analysis of the batch of samples to be analyzed.

[0039] As an example, steps S10 to S30 include: obtaining the raw sequencing data of the batch of samples to be analyzed and a sample file of the workflow configuration file input by the user; and obtaining workflow configuration data after detecting the workflow configuration information input in the sample file of the workflow configuration file, wherein the workflow configuration information may specifically include project name, database name, reference genome, BIN path, analysis software, quality control functional parameters, model species parameters, software analysis parameters, and running thread parameters, etc.; inputting the raw sequencing data and the workflow configuration data into a preset sample resequencing program to generate a sample resequencing program for the batch of samples to be analyzed, wherein the sample resequencing program includes resequencing workflow data; and performing resequencing data analysis on the batch of samples to be analyzed by running the sample resequencing program under the resequencing workflow corresponding to the resequencing workflow data.

[0040] This application first acquires the raw sequencing data and workflow configuration data of the batch of samples to be analyzed, input by the user. Then, it inputs these data into a preset sample resequencing program to obtain the sample resequencing program for the batch of samples to be analyzed. Finally, the sample resequencing program is run under the resequencing workflow corresponding to the resequencing workflow data, enabling resequencing data analysis of the batch of samples. In other words, it achieves the goal of high-throughput whole-genome resequencing data analysis of the batch of samples by running the sample resequencing program. Through simple information interaction between the user and the resequencing data analysis platform, the purpose of interactively analyzing high-throughput whole-genome resequencing data of batch samples can be completed, rather than requiring the user to set independent sequencing logic for each different sample when analyzing high-throughput whole-genome resequencing data of batch samples. This overcomes the technical defect that requires users to set specific sequencing logic for all samples due to the independent operation of analysis workflows for different samples, which easily leads to long analysis times for high-throughput whole-genome resequencing data of batch samples. Therefore, it improves the analysis efficiency of high-throughput whole-genome resequencing data of batch samples.

[0041] The sample resequencing program includes resequencing quality control data, and the step of performing resequencing data analysis on the batch of samples to be analyzed by running the sample resequencing program based on the resequencing process data includes:

[0042] Step A10: Detect the quality control function status of the batch of samples to be analyzed according to the process configuration data.

[0043] Step A20: If the quality control function status is the first quality control function status, then the first data analysis graph is obtained by performing quality control data statistics on the original sequencing data;

[0044] Step A30: If the quality control function status is the second quality control function status, then run the resequencing quality control data to obtain the quality control post-sequencing data of the batch of samples to be analyzed.

[0045] Step A40: Perform quality control data statistics on the original sequencing data and the post-quality control sequencing data to obtain the second data analysis graph of the batch of samples to be analyzed.

[0046] In this embodiment, it should be noted that the resequencing quality control data is used to characterize the original sequencing data used for quality control. When running the main sample resequencing program, the resequencing quality control data runs with six threads by default. Users can add or remove threads based on server resources. To display the quality control analysis results to users, a first data analysis graph or a second data analysis graph can be output. Both the first and second data analysis graphs are visual data analysis graphs, which can specifically be graphics or tables. The quality control function status includes a first quality control function status and a second quality control function status. In the process, the first quality control function status can be the quality control function off state, and the second quality control function status is the quality control function on state. The quality control function status of the batch of samples to be analyzed can be determined by the quality control function parameters carried in the process configuration data. The sequencing data after quality control is used to characterize the raw sequencing data after quality control. Among them, different raw sequencing data can be set to be quality controlled independently without interference, which can improve the data analysis speed. The quality control data may include the total number of sequences, the total number of bases, the average sequence length, the number of N bases, the GC content, the proportion of Q20, the proportion of Q30, and the Clean ratio, etc. Among them, the GC (Gas chromatography) amount, Q20 characterizes the error probability given for the identified bases during the base identification process of sequencing, and Q30 is used to characterize the reliability of base identification.

[0047] As an example, steps A10 to A40 include: detecting the quality control function status of the batch of samples to be analyzed according to the quality control function status carried by the process configuration data; if the quality control function status is in the quality control function off state, then performing quality control data statistics on the original sequencing data to obtain a first data analysis graph; if the quality control function status is in the quality control function on state, then running the resequencing quality control data to obtain the quality control post-sequencing data of the batch of samples to be analyzed; and calculating the total number of sequences, total number of bases, average sequence length, number of N bases, GC content, Q20 proportion, Q30 proportion, and Clean ratio of the original sequencing data and the quality control post-sequencing data to obtain a second data analysis graph of the batch of samples to be analyzed.

[0048] In one feasible approach, the default adapter sequence FASTA file is TruSeq-SE.fa from the ILLUMINA platform. The maximum allowed mismatch is 2 bp, the global alignment threshold is 30, and the local alignment threshold is 10. First, quality and N-containing sequences are removed. The average quality value of the sliding window sequencing is calculated, and bases or N-values ​​below the threshold starting from the beginning and end of the sequence are removed. The default window size is 5 bp, the average quality value is 20, and the threshold is 5. Sequences that are too short are removed, with a default minimum length of 50 bp. Sequence data after quality control is output. Then, the quality control function parameters in the workflow configuration data are checked. If the quality control function parameters are turned off, only the following are checked: The system measures the base quality distribution and base class percentage of each cycle in the raw sequencing data and outputs a visualization. It also calculates the total number of sequences, total number of bases, average sequence length, number of N bases, GC content, Q20 percentage, and Q30 percentage in the raw sequencing data. If the quality control function is enabled, it checks the integrity of the quality-controlled sequencing data, the base quality distribution and base class percentage of each cycle, and outputs a visualization. It also calculates the total number of sequences, total number of bases, average sequence length, number of N bases, GC content, Q20 percentage, Q30 percentage, and Clean ratio in both the raw and quality-controlled sequencing data and outputs a readable table.

[0049] The resequencing data analysis method further includes, after the step of running the resequencing quality control data to obtain the quality control post-sequencing data of the batch of samples to be analyzed, the step of running the resequencing quality control data to obtain the quality control post-sequencing data of the batch of samples to be analyzed.

[0050] Step B10: Obtain sequencing reference data of the batch reference samples of the batch samples to be analyzed;

[0051] Step B20: Compare and sort the quality control sequencing data and the sequencing reference data to obtain at least one sequencing storage data.

[0052] Step B30: Based on the data type of the original sequencing data, perform statistics on each of the sequencing storage data to obtain the third data analysis graph of the batch of samples to be analyzed.

[0053] In this embodiment, it should be noted that during the resequencing data analysis process, in addition to quality control of the raw sequencing data, reference genome alignment can also be performed. The reference genome alignment runs by default with 6 threads, which users can adjust based on server resources. Batch reference samples are used to characterize the reference genome of the batch samples to be analyzed, and sequencing reference data is used to characterize the raw sequencing data of the batch reference samples. The alignment and sorting method can be one-to-one. For example, in one feasible approach, the quality-controlled sequencing data can be aligned with the reference genome, removing unaligned sequences and retaining only aligned sequences. After sorting, a binary format BAM file is output, and a quick access index file (index file) is created for the BAM file to quickly verify its integrity and statistically analyze the input data, namely the total number of sequences, total number of bases, number of sequences aligned with the reference genome, and the number of sequences aligned with the reference genome. The sequence counts of the reference genome, the number of unaligned sequences to the reference genome, the number of aligned bases to the reference genome, the number of mismatched bases, the number of unaligned sequences to the reference genome, the number of supplementary aligned sequences to the reference genome, the alignment rate, the mismatch rate, and the number of bases at unique positions aligned to the reference genome are sorted to generate sequencing storage data. Finally, the sequencing storage data is statistically analyzed based on the data type of the original sequencing data. The data type can be single-end or paired-end. When the original sequencing data is single-end, the above data is statistically analyzed; when the original sequencing data is multi-end, the mean length and standard deviation of the inserted fragments are statistically analyzed. Additionally, if the quality control function parameter in the workflow configuration file is turned off, the original sequencing data is aligned with the reference genome. Sequencing data from different samples and multiple quality control sequences of the same sample are aligned with the reference genome independently, without interference, which improves data analysis speed.

[0054] As an example, steps B10 to B30 include: obtaining sequencing reference data of the batch reference samples of the batch samples to be analyzed; aligning and sorting the quality control sequencing data and the sequencing reference data from high to low to obtain at least one sequencing storage data; determining the statistical data of each sequencing storage data according to the data type of the original sequencing data, and performing statistics on each statistical data to obtain a third data analysis graph of the batch samples to be analyzed. Specifically, the statistical data may include the average length of sequence fragments, the standard deviation of sequence fragments, the total number of sequences and bases in the quality control sequencing data, the number of sequences aligned with the reference genome, the number of sequences aligned with the reference genome, the number of sequences not aligned with the reference genome, the number of bases aligned with the reference genome, the number of mismatched bases, the number of sequences not aligned with the reference genome, the number of sequences supplemented to align with the reference genome, the alignment rate, the mismatch rate, and the number of bases at unique positions aligned with the reference genome.

[0055] The resequencing data analysis method further includes, after the step of comparing and sorting the quality control sequencing data and the sequencing reference data to obtain at least one sequenced storage data, the step of resequencing data analysis including:

[0056] Step C10: According to the sample type of the batch of samples to be analyzed, merge the sequencing storage data to obtain at least one sequencing storage data of the same type.

[0057] Step C20: Redundancy processing is performed on the stored sequencing data of the same type to obtain the target sequencing data of the same type;

[0058] Step C30: Based on the data type of the target similar sequencing data, perform statistics on the target similar sequencing data to obtain the fourth data analysis graph of the batch of samples to be analyzed.

[0059] In this embodiment, it should be noted that since the batch of samples to be analyzed during the resequencing data analysis process may contain samples of the same type, redundant sequence marking can be performed on the data during the resequencing data analysis process. The redundant sequence marking has a default of 6 threads, and users can also add or reduce threads according to server resources. Target similar sequencing data is used to characterize similar sequencing data after redundancy processing. The data types of target similar sequencing data include single-end sequencing data and multi-end sequencing data. The sequencing storage data statistically analyzed for different types of data are not the same. For example, in one feasible approach, when performing redundant sequence marking during the resequencing data analysis process, multiple quality control samples of the same type need to be merged according to the redundant sequence marking or deletion options of the sample resequencing program. Following the alignment of the BAM files, redundant optical and PCR sequences are marked or deleted based on the library information in the BAM files. This generates BAM files named after the sample names with marked or deleted redundancies. The integrity of the BAM files is quickly verified, and a fast access index file for the BAM files is created. The aligned BAM files and their fast access index files after quality control are deleted to reduce storage usage. At the same time, the sequence information, GC bias distribution, average quality value of each cycle, base quality value distribution of the sequences, and reference genome coverage depth of the marked or deleted redundancies BAM files are statistically analyzed to obtain a fourth data analysis graph containing the above data. In addition, for paired-end sequencing data, the distribution of inserted fragments will be statistically analyzed and a visualization image will be output.

[0060] As an example, steps C10 to C30 include: obtaining the sample types of the batch of samples to be analyzed; merging the sequencing storage data according to the sample types to obtain at least one type of sequencing storage data; deleting or marking the sequencing storage data of the same type to obtain target sequencing data of the same type; and statistically analyzing the target sequencing data of the same type according to the data type of the target sequencing data to obtain a fourth data analysis graph of the batch of samples to be analyzed.

[0061] The resequencing data analysis method further includes, after the step of performing redundancy processing on each of the aforementioned similar sequencing storage data to obtain the target similar sequencing data, the method further includes:

[0062] Step D10: Detect the sample type of the batch of samples to be analyzed according to the process configuration data.

[0063] Step D20: Extract sequencing mutation data from the target sequencing data of the same type;

[0064] Step D30: Based on the sample type, the sequencing mutation data, and the preset mutation filtering conditions, perform mutation quality correction on the target sequencing data of the same type to obtain the corrected first sequencing correction data.

[0065] In this embodiment, it should be noted that during resequencing data analysis, base quality recalibration can also be performed through the resequencing data analysis platform. Base quality recalibration is fixed as a single thread. Before base quality recalibration, mutation filtering conditions can be preset. Sequencing mutation data is used to characterize the data that produce mutations in the target similar sequencing data. After the existence of sequencing mutation data, it is necessary to perform type-specific mutation quality correction on the target similar sequencing data of the batch of samples to be analyzed through preset mutation filtering conditions. For example, in one feasible approach, when the model species parameters carried in the workflow configuration file determine that the batch of samples to be analyzed is a non-model species, the mutation information extracted directly from the redundant sequence marker or deleted BAM file is combined with the preset mutation filtering conditions to perform BQSR on the redundant sequence marker or deleted BAM file. The specific process is as follows: First, use GATK to extract the whole genome mutation information in the redundant sequence marker or deleted BAM file, compress the mutation file, and create a fast access index (tbi) file for the compressed file; second, extract the SNP and INDEL information of the compressed mutation file respectively, and further filter according to the filtering conditions. The default filtering condition for SNP is the confidence level of deep variation (Variant). Confidence / Quality by Depth (QD) less than 2.0, RMS Mapping Quality (MQ) less than 40.0, p-value using Fisher's exact test to detect strand bias (FS) greater than 60.0, Symmetric Odds Ratio of 2x2 contingency table to detect strand bias (SOR) greater than 3.0, Z-score from Wilcoxon rank sum test of Alt vs. Ref read mapping qualities (MQRankSum) less than -12.5, or Z-score from Wilcoxon rank sum test of Alt vs. Ref read position bias. If the bias (ReadPosRankSum) is less than -8.0, the default filtering conditions for INDEL are QD less than 2.0, FS greater than 200.0, SOR greater than 10.0, MQRankSum less than -12.5, or ReadPosRankSum less than -8.0; Finally, the filtered SNP and INDEL information is merged and output as the filtered mutation file. This mutation file is used to perform BQSR on the BAM files containing redundant sequence markers or deletions, outputting the BQSR-enhanced BAM file. An MD5 file for this BAM file is created, and the BAM files containing redundant sequence markers or deletions, their fast access indexes, the mutation files extracted from the BAM files containing redundant sequence markers or deletions, the SNP mutation files and their filtered files, and the INDEL mutation files and their filtered files are deleted, reducing storage usage. Furthermore, users can provide custom filtering conditions for personalized analysis. If the model species parameter in the configuration file is true, meaning the sample is the model species *Homosapiens*, then GATK is used to perform BQSR on the BAM files containing redundant sequence markers or deletions based on the DBSNP, MILLS, and 1000G mutation sets of the reference genome version hg19 or hg38 of the GATK_bundle, outputting the BQSR-enhanced BAM file, creating the MD5 file for this BAM file, and deleting the BAM files containing redundant sequence markers or deletions and their fast access indexes, further reducing storage usage.

[0066] As an example, steps D10 to D30 include: detecting the sample type of the batch of samples to be analyzed according to the process configuration data, wherein the sample type can be divided into non-model species samples and model species samples; extracting sequencing mutation data from the target similar sequencing data; and performing mutation quality correction on the target similar sequencing data according to the sample type, the sequencing mutation data and preset mutation filtering conditions to obtain the corrected first sequencing correction data.

[0067] In one feasible approach, the resequencing data analysis platform also includes a variant detection function. The variant detection operation is fixed as a single thread, using GATK to extract mutation information from the BAM file after BQSR, compressing the mutation file, and creating a fast access index tbi file for the compressed file.

[0068] The resequencing data analysis method further includes, after the step of performing data quality correction on the target similar sequencing data based on the sequencing mutation data and preset mutation filtering conditions to obtain the corrected first sequencing correction data:

[0069] Step E10: Extract sequencing variant data from the first sequencing correction data;

[0070] Step E20: Based on the sample type, the sequencing variant data, and the preset variant filtering conditions, perform variant quality correction on the first sequencing correction data to obtain the corrected second sequencing correction data.

[0071] In this embodiment, it should be noted that, in addition to base quality recalibration, the resequencing data analysis platform also has a variant quality recalibration function. Sequencing variant data is used to characterize the variant data generated in the target type of sequencing data. For example, in one feasible approach, the sample type is detected based on the model species parameter carried in the workflow configuration data. If the model species parameter is detected as false, i.e., the sample is not the model species Homo sapiens, then the VQSR is performed directly on the mutation compressed file extracted from the BAM file after BQSR according to the preset variant filtering conditions. The process is as follows: First, GATK is used to extract SNP and INDEL information from the mutation compressed file extracted from the BAM file after BQSR, and the SNP and INDEL are further screened using the same preset mutation filtering conditions as described above. Then, the filtered SNP and INDEL information is merged and the final mutation VQSR file is output, and the SNP and INDEL files extracted from the mutation compressed file and their filtered files are deleted to reduce storage usage.If the model species parameter in the configuration file is true, meaning the sample is the model species *Homosapiens*, then the HAPMAP, OMNI, 1000G, DBSNP, and MILLS mutation datasets from the reference genome hg19 or hg38 version of the GATK_bundle are used. The weights of these five mutation datasets in the VQSR model training are 15.0, 12.0, 10.0, 5.0, and 12.0, respectively. Further, combined with mutation filtering conditions, VQSR is performed on the mutation files extracted from the BQSR-derived BAM file. The process is as follows: First, VQSR is performed on the SNPs in the mutation compressed files extracted from the BQSR-derived BAM file using GATK. The default known variant dataset is DBSNP, the default model training dataset is HAPMAP, OMNI, and 1000G, the default validation model training dataset is HAPMAP, and the default information used for calculation is the filtered sequence coverage (Approximate). The data includes readdepth (DP), QD, FS, SOR, ReadPosRankSum, and MQRankSum. The default Gaussian model calculates a set of recalibrated true mutation quality scores of 100.0, 99.9, 99.0, 95.0, and 90.0 for each mutation site. The default maximum Gaussian clustering is 6, and the default sensitivity threshold for the model's true dataset is 99.0%. The output is a mutation file that performs VQSR on the SNPs in the mutation compressed file. Then, GATK is used to perform VQSR on the INDELs in the mutation file output after the SNP VQSR. The default known mutation dataset, model training dataset, and validation model training dataset are all Mills. The default information used for computation, the default set of recalibrated true mutation quality scores calculated by the Gaussian model for each mutation site, the default maximum Gaussian clustering, and the default sensitivity threshold for the model's true dataset are all consistent with the SNP VQSR. The final mutation VQSR file is output, and the mutation file that performed VQSR on the SNPs in the mutation compressed file extracted from the BAM file after BQSR is deleted to reduce storage usage. In addition, users can customize the information used for calculation, the set of quality scores, the maximum Gaussian clustering, and the sensitivity threshold to achieve personalized analysis.

[0072] As an example, steps E10 to E20 include: extracting sequencing variant data from the first sequencing correction data; and performing variant quality correction on the first sequencing correction data according to the sample type, the sequencing variant data, and preset variant filtering conditions to obtain corrected second sequencing correction data.

[0073] This application provides a resequencing data analysis method, namely, acquiring the raw sequencing data and process configuration data of a batch of samples to be analyzed, wherein the raw sequencing data is generated by the batch of samples to be analyzed on at least one sequencing platform based on at least one sequencing strategy; generating a sample resequencing program for the batch of samples to be analyzed based on the raw sequencing data and the process configuration data, wherein the sample resequencing program includes resequencing process data; and performing resequencing data analysis on the batch of samples to be analyzed by running the sample resequencing program based on the resequencing process data.

[0074] In the process of analyzing high-throughput whole-genome resequencing data of a batch of samples to be analyzed, this embodiment first obtains the original sequencing data and process configuration data of the batch of samples to be analyzed, then generates a sample resequencing program for the batch of samples to be analyzed based on the original sequencing data and process configuration data, and finally analyzes the resequencing data of the batch of samples to be analyzed by running the sample resequencing program based on the resequencing process data carried in the sample resequencing program. Since the resequencing data analysis platform can directly generate the sample resequencing program, the purpose of high-throughput whole-genome resequencing data analysis of the batch of samples to be analyzed can be achieved by running the sample resequencing program.

[0075] Since the sample resequencing program is generated based on the original sequencing data and the process configuration data, the sample resequencing program carries resequencing process data. Therefore, the process of the batch of samples to be analyzed can be integrated through the resequencing process data. Finally, through the sample resequencing program, the sequencing data of the batch of samples generated on multiple sequencing platforms based on various strategies can be automatically analyzed. That is, the resequencing data analysis platform can achieve the purpose of automatically analyzing the high-throughput whole genome resequencing data of the batch of samples to be analyzed by acquiring the original sequencing data and process configuration data.

[0076] Based on this, this application automatically generates a sample resequencing program for the batch of samples to be analyzed from the original sequencing data and workflow configuration data. Finally, by running the sample resequencing program, the purpose of automatically analyzing the high-throughput whole-genome resequencing data of the batch of samples to be analyzed is achieved. This avoids the need for users to set independent sequencing logic for each sample when analyzing high-throughput whole-genome resequencing data of batches. In other words, it overcomes the technical drawback of requiring users to set specific sequencing logic for all samples due to independent operation of the analysis workflow for different samples, which easily leads to long analysis times for high-throughput whole-genome resequencing data of batches. Therefore, it improves the analysis efficiency of high-throughput whole-genome resequencing data of batches.

[0077] Example 2

[0078] Furthermore, referring to Figure 2 In another embodiment of this application, content that is the same as or similar to that in Embodiment 1 above can be referred to the above description and will not be repeated hereafter. Based on this, after the step of performing variant quality correction on the first sequencing correction data according to the sample type, the sequencing variant data, and preset variant filtering conditions to obtain the corrected second sequencing correction data, the resequencing data analysis method further includes:

[0079] Step F10: Detect the deletion function status of the batch of samples to be analyzed according to the sample resequencing procedure.

[0080] Step F20: When the quality control function is in the second quality control function state and the deletion function is enabled, delete the post-quality control sequencing data.

[0081] In this embodiment, it should be noted that, to ensure the analysis speed of the resequencing data analysis platform, the resequencing data analysis platform also has an intermediate quality control file deletion function. The intermediate quality control file deletion function operates in a single thread. For example, in one feasible approach, the execution logic of the intermediate quality control file deletion function is as follows: Based on the quality control function parameters in the process configuration data and the intermediate quality control file deletion option in the sample resequencing program, it is selected whether to generate and execute the sequencing data file generated after the original sequencing quality control. Only when both the quality control function parameters and the intermediate quality control file deletion option are enabled will the sample resequencing program automatically generate the intermediate quality control file deletion script and execute the script after VQSR. That is, when the quality control function status is the second quality control function status and the deletion function status is enabled, the quality control post-sequencing data output in the original sequencing data quality control can be deleted, i.e., the quality control post-sequencing data is deleted.

[0082] As an example, steps F10 to F20 include: detecting the deletion function status of the batch of samples to be analyzed according to the sample resequencing procedure; when the quality control function status is detected to be the second quality control function status and the deletion function status is that the deletion function is enabled, deleting the quality control-enhanced sequencing data.

[0083] The resequencing data analysis method further includes:

[0084] Step G10: Check whether the resequencing data analysis process of the batch of samples to be analyzed is proceeding normally;

[0085] Step G20: If so, the resequencing process file for the batch of samples to be analyzed is obtained by integrating the resequencing data analysis process, and the resequencing process file is stored according to the specified storage path of the sample resequencing program.

[0086] Step G30: If not, after detecting that the reason for the abnormal process meets the preset abnormal conditions, determine the resequencing process breakpoint of the batch of samples to be analyzed, and run the resequencing program by returning to the resequencing process breakpoint to perform resequencing data analysis on the batch of samples to be analyzed.

[0087] In this embodiment, it should be noted that the resequencing data analysis process can be recorded by setting up a storage resequencing workflow file. For example, in one feasible approach, the automatic delivery script will check the integrity of data reading or output at each step of the sample according to the statistical commands in the script delivery sequence file for all samples, and statistically analyze the sequencing strategy, sequencing length, total number of original sequences, total number of sequences after quality control, number of sequences aligned to the reference genome, number of sequences not aligned to the reference genome, number of sequences supplemented to be aligned to the reference genome (statistics by two methods), number of redundant sequences generated by optical and PCR, total number of original bases, total number of bases after quality control, number of bases aligned to the reference genome, number of bases at unique positions aligned to the reference genome, number of mismatched bases, Clean rate, alignment rate, unique alignment rate, mismatch rate, redundancy rate, effective average coverage depth, mean and standard deviation of average whole genome coverage depth, and 1, 5, 10, 15, 20, 25, 30. The whole genome coverage at different depths (40, 50, 60, 70, 80, 90, 100X) is finally output in a text file in the output path. If there are incomplete data reads or outputs in any step, this file will not be output, and a message indicating incomplete data will be displayed. After completing the above verification and statistics, the automatic delivery script will draw a flowchart of the entire project execution and output it in the project automatic analysis folder in the output path. In addition, for multiple scripts with no order relationship between different samples, the automatic delivery script will be delivered and run simultaneously according to the principle that the total number of threads set in the automatic delivery script and the total number of threads of the delivered scripts are less than or equal to the total number of threads of the project. If the task is terminated due to system failure or instability of the analysis platform, the automatic delivery script will start the breakpoint re-delivery function to ensure that all scripts in the project are executed. The preset abnormal conditions can be platform system failure or instability, etc. The resequencing process breakpoint is used to characterize the process point of the breakpoint re-delivery function.

[0088] As an example, steps G10 to G30 include: detecting whether the resequencing data analysis process of the batch of samples to be analyzed is proceeding normally; if the resequencing data analysis process of the batch of samples to be analyzed is detected to be proceeding normally, then integrating the resequencing data analysis process to obtain the resequencing process file of the batch of samples to be analyzed, and storing the resequencing process file in the specified storage path of the sample resequencing program; if the resequencing data analysis process of the batch of samples to be analyzed is detected not proceeding normally, then detecting the reason why the resequencing data analysis process of the batch of samples to be analyzed is not proceeding normally, and after detecting that the reason for not proceeding normally meets the preset abnormal conditions, querying the resequencing process breakpoint of the batch of samples to be analyzed, and running the resequencing program by returning to the resequencing process breakpoint to perform resequencing data analysis on the batch of samples to be analyzed.

[0089] Among them, reference Figure 3 , Figure 3 To illustrate the data analysis scenario flowchart of the resequencing data analysis platform, the raw FASTQ files are processed through quality control steps 1 to 10, which are implemented using the analysis functions of the resequencing data analysis platform.

[0090] This application provides a method for deleting post-quality control sequencing data. Specifically, it detects the deletion function status of the batch of samples to be analyzed according to the sample resequencing program. When the quality control function status is the second quality control function status and the deletion function status is enabled, the post-quality control sequencing data is deleted. This application detects the deletion function status of both the samples to be analyzed and the sample resequencing program during resequencing data analysis. When both are enabled, the post-quality control sequencing data is deleted, thereby clearing the data cache of the resequencing data analysis platform during the data analysis process. Therefore, this lays the foundation for improving the analysis speed of the resequencing data analysis platform.

[0091] Example 3

[0092] This application also provides a resequencing data analysis device, applied to a resequencing data analysis platform, as described above. Figure 4 The resequencing data analysis device includes:

[0093] The acquisition module 101 is used to acquire the raw sequencing data and process configuration data of the batch of samples to be analyzed, wherein the raw sequencing data is generated by the batch of samples to be analyzed on at least one sequencing platform based on at least one sequencing strategy.

[0094] The generation module 102 is used to generate a sample resequencing program for the batch of samples to be analyzed based on the original sequencing data and the process configuration data, wherein the sample resequencing program includes resequencing process data.

[0095] Analysis module 103 is used to perform resequencing data analysis on the batch of samples to be analyzed by running the sample resequencing program based on the resequencing process data.

[0096] Optionally, the sample resequencing procedure includes resequencing quality control data, and the analysis module 103 is further used for:

[0097] Based on the process configuration data, detect the quality control function status of the batch of samples to be analyzed;

[0098] If the quality control function status is the first quality control function status, then the first data analysis graph is obtained by performing quality control data statistics on the original sequencing data;

[0099] If the quality control function status is the second quality control function status, then the resequencing quality control data is run to obtain the quality control post-sequencing data of the batch of samples to be analyzed;

[0100] The original sequencing data and the post-quality control sequencing data are statistically analyzed to obtain the second data analysis graph of the batch of samples to be analyzed.

[0101] Optionally, the analysis module 103 includes a statistical unit, which is used for:

[0102] Obtain sequencing reference data of the batch reference sample of the batch sample to be analyzed;

[0103] The quality control sequencing data and the sequencing reference data are compared and sorted to obtain at least one sequencing storage data:

[0104] Based on the data type of the original sequencing data, statistics are performed on each of the sequencing storage data to obtain the third data analysis graph of the batch of samples to be analyzed.

[0105] Optionally, the analysis module 103 further includes a processing unit, the processing unit being used for:

[0106] Based on the sample type of the batch of samples to be analyzed, the sequencing storage data of each sample are merged to obtain at least one sequencing storage data of the same type.

[0107] Redundancy processing is performed on the aforementioned similar sequencing storage data to obtain the target similar sequencing data;

[0108] Based on the data type of the target-type sequencing data, statistics are performed on the target-type sequencing data to obtain the fourth data analysis graph of the batch of samples to be analyzed.

[0109] Optionally, the analysis module 103 further includes a mutation correction unit, which is used for:

[0110] Based on the process configuration data, detect the sample type of the batch of samples to be analyzed;

[0111] Extract sequencing mutation data from the target-type sequencing data;

[0112] Based on the sample type, the sequencing mutation data, and the preset mutation filtering conditions, mutation quality correction is performed on the target sequencing data of the same type to obtain the corrected first sequencing correction data.

[0113] Optionally, the analysis module 103 further includes a variation correction unit, which is used for:

[0114] Extract sequencing variant data from the first sequencing correction data;

[0115] Based on the sample type, the sequencing variant data, and the preset variant filtering conditions, the first sequencing correction data is subjected to variant quality correction to obtain the corrected second sequencing correction data.

[0116] Optionally, the analysis module 103 further includes a deletion unit, which is used for:

[0117] The deletion function status of the batch of samples to be analyzed is detected according to the sample resequencing procedure.

[0118] When the quality control function is in the second quality control function state and the deletion function is enabled, the post-quality control sequencing data is deleted.

[0119] Optionally, the resequencing analysis device further includes an integration module, the integration module being used for:

[0120] Check whether the resequencing data analysis process of the batch of samples to be analyzed is proceeding normally;

[0121] If so, the resequencing process file for the batch of samples to be analyzed is obtained by integrating the resequencing data analysis process, and the resequencing process file is stored according to the specified storage path of the sample resequencing program.

[0122] If not, after detecting that the cause of the abnormal process meets the preset abnormal conditions, the resequencing process breakpoint of the batch of samples to be analyzed is determined, and the resequencing program is run by returning to the resequencing process breakpoint to perform resequencing data analysis on the batch of samples to be analyzed.

[0123] The resequencing data analysis device provided by this invention, employing the resequencing data analysis method in the above embodiments, solves the technical problem of low analysis efficiency for high-throughput whole-genome resequencing data of batch samples. Compared with the prior art, the beneficial effects of the resequencing data analysis device provided by this invention are the same as those of the resequencing data analysis method provided in the above embodiments, and other technical features in this resequencing data analysis device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0124] Example 4

[0125] This invention provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the resequencing data analysis method in Embodiment 1 above.

[0126] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0127] like Figure 5 As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processor, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus.

[0128] Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication devices allow electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although electronic devices with various systems are shown in the figures, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented alternatively.

[0129] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1009, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of embodiments of this disclosure.

[0130] The electronic device provided by this invention employs the resequencing data analysis method described in the above embodiments, solving the technical problem of low analysis efficiency for high-throughput whole-genome resequencing data of batch samples. Compared with the prior art, the beneficial effects of the electronic device provided by this invention are the same as those of the resequencing data analysis method provided in the above embodiments, and other technical features of this electronic device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0131] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.

[0132] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0133] Example 5

[0134] This embodiment provides a computer-readable storage medium having computer-readable program instructions stored thereon, which are used to execute the resequencing data analysis method in the above embodiment.

[0135] The computer-readable storage medium provided in this embodiment of the invention may be, for example, a USB flash drive, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0136] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.

[0137] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to: acquire raw sequencing data and process configuration data of a batch of samples to be analyzed, wherein the raw sequencing data is generated by the batch of samples to be analyzed on at least one sequencing platform based on at least one sequencing strategy; generate a sample resequencing program for the batch of samples to be analyzed based on the raw sequencing data and the process configuration data, wherein the sample resequencing program includes resequencing process data; and perform resequencing data analysis on the batch of samples to be analyzed by running the sample resequencing program based on the resequencing process data.

[0138] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0140] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0141] The computer-readable storage medium provided by this invention stores computer-readable program instructions for executing the above-described resequencing data analysis method, thus solving the technical problem of low analysis efficiency in analyzing high-throughput whole-genome resequencing data from batch samples. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this invention are the same as those of the resequencing data analysis method provided in the above-described embodiments, and will not be repeated here.

[0142] Example 6

[0143] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the resequencing data analysis method described above.

[0144] The computer program product provided in this application solves the technical problem of low analysis efficiency in analyzing high-throughput whole-genome resequencing data from batch samples. Compared with the prior art, the beneficial effects of the computer program product provided in this embodiment are the same as those of the resequencing data analysis method provided in the above embodiments, and will not be repeated here.

[0145] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.

Claims

1. A method for resequencing data analysis, characterized in that, The resequencing data analysis method, applied to a resequencing data analysis platform, includes: Obtain the raw sequencing data and process configuration data of the batch of samples to be analyzed, wherein the raw sequencing data is generated by the batch of samples to be analyzed on at least one sequencing platform based on at least one sequencing strategy; Based on the original sequencing data and the process configuration data, a sample resequencing program for the batch of samples to be analyzed is generated, wherein the sample resequencing program includes resequencing process data. Based on the resequencing process data, the resequencing data of the batch of samples to be analyzed is performed by running the sample resequencing program. The sample resequencing procedure includes resequencing quality control data. The step of performing resequencing data analysis on the batch of samples to be analyzed by running the sample resequencing program based on the resequencing process data includes: Based on the process configuration data, detect the quality control function status of the batch of samples to be analyzed; If the quality control function status is the first quality control function status, then the first data analysis graph is obtained by performing quality control data statistics on the original sequencing data; If the quality control function status is the second quality control function status, then the resequencing quality control data is run to obtain the quality control post-sequencing data of the batch of samples to be analyzed; The original sequencing data and the post-quality control sequencing data were statistically analyzed to obtain the second data analysis graph of the batch of samples to be analyzed. Wherein, the first quality control function state is the quality control function closed state, and the second quality control function state is the quality control function open state; The quality control data includes total number of sequences, total number of bases, average sequence length, number of N bases, GC content, Q20 percentage, Q30 percentage, and Clean ratio. The GC content and Q20 represent the error probability given for the identified bases during the sequencing base identification process, and Q30 is used to characterize the reliability of base identification.

2. The resequencing data analysis method as described in claim 1, characterized in that, After the step of running the resequencing quality control data to obtain the quality control post-sequencing data of the batch of samples to be analyzed, the resequencing data analysis method further includes: Obtain sequencing reference data of the batch reference sample of the batch sample to be analyzed; The quality control sequencing data and the sequencing reference data are compared and sorted to obtain at least one sequencing storage data: Based on the data type of the original sequencing data, statistics are performed on each of the sequencing storage data to obtain the third data analysis graph of the batch of samples to be analyzed.

3. The resequencing data analysis method as described in claim 2, characterized in that, After the step of comparing and sorting the quality control sequencing data and the sequencing reference data to obtain at least one sequencing storage data, the resequencing data analysis method further includes: Based on the sample type of the batch of samples to be analyzed, the sequencing storage data of each sample are merged to obtain at least one sequencing storage data of the same type. Redundancy processing is performed on the aforementioned similar sequencing storage data to obtain the target similar sequencing data; Based on the data type of the target-type sequencing data, statistics are performed on the target-type sequencing data to obtain the fourth data analysis graph of the batch of samples to be analyzed.

4. The resequencing data analysis method as described in claim 3, characterized in that, After the step of performing redundancy processing on each of the aforementioned similar sequencing storage data to obtain the target similar sequencing data, the resequencing data analysis method further includes: Based on the process configuration data, detect the sample type of the batch of samples to be analyzed; Extract sequencing mutation data from the target-type sequencing data; Based on the sample type, the sequencing mutation data, and the preset mutation filtering conditions, mutation quality correction is performed on the target sequencing data of the same type to obtain the corrected first sequencing correction data.

5. The resequencing data analysis method as described in claim 4, characterized in that, After the step of performing data quality correction on the target similar sequencing data based on the sequencing mutation data and preset mutation filtering conditions to obtain the corrected first sequencing correction data, the resequencing data analysis method further includes: Extract sequencing variant data from the first sequencing correction data; Based on the sample type, the sequencing variant data, and the preset variant filtering conditions, the first sequencing correction data is subjected to variant quality correction to obtain the corrected second sequencing correction data.

6. The resequencing data analysis method as described in claim 5, characterized in that, After the step of performing variant quality correction on the first sequencing correction data according to the sample type, the sequencing variant data, and preset variant filtering conditions to obtain the corrected second sequencing correction data, the resequencing data analysis method further includes: The deletion function status of the batch of samples to be analyzed is detected according to the sample resequencing procedure. When the quality control function is in the second quality control function state and the deletion function is enabled, the post-quality control sequencing data is deleted.

7. The resequencing data analysis method as described in claim 6, characterized in that, The resequencing data analysis method further includes: Check whether the resequencing data analysis process of the batch of samples to be analyzed is proceeding normally; If so, the resequencing process file for the batch of samples to be analyzed is obtained by integrating the resequencing data analysis process, and the resequencing process file is stored according to the specified storage path of the sample resequencing program. If not, after detecting that the cause of the abnormal process meets the preset abnormal conditions, the resequencing process breakpoint of the batch of samples to be analyzed is determined, and the resequencing program is run by returning to the resequencing process breakpoint to perform resequencing data analysis on the batch of samples to be analyzed.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; A memory that is communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the steps of the resequencing data analysis method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program for implementing a resequencing data analysis method, the program for implementing the resequencing data analysis method being executed by a processor to implement the steps of the resequencing data analysis method as described in any one of claims 1 to 7.