A multi-generation sequencing data quality control management method and system

By employing a multi-layered nested quality control analysis and source tracing detection method, the problems of single-dimensional quality control of sequencing data and insufficient identification of deep anomalies were solved, realizing multi-dimensional and precise quality control of sequencing data and improving the accuracy and reliability of the data.

CN120833091BActive Publication Date: 2025-12-16INST OF AQUATIC LIFE ACAD SINICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511332142.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-12-16
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

Existing technologies suffer from limited sequencing data quality control dimensions, poor dynamic adaptability, and insufficient ability to identify deep anomalies, resulting in poor quality and reliability of sequencing data quality control management.

Method used

By employing a multi-layered nested quality control analysis method, including multi-dimensional quality control at the original signal layer, alignment mapping layer, annotation variation layer, and biological logic layer, combined with a source tracing analysis channel, abnormal influences are identified and corrected, generating quality control analysis results.

Benefits of technology

It achieves multi-dimensional and precise quality control, improves the quality and reliability of sequencing data quality control management, and ensures the accuracy and reliability of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833091B_ABST
    Figure CN120833091B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-generation sequencing data quality control management method, system, it is related to data management related technical field, method includes: reading the file header data of sequencing data, load quality control rule template;Sequencing data is imported after conversion and is carried out feature extraction under single read segment granularity, establishes feature extraction result;According to quality control rule template, sequencing data, feature extraction result carries out multi-layer nested quality control analysis, establishes sample quality grade mark;Configuration trigger threshold and carry out sample quality grade mark trigger analysis, execute trace detection, establish abnormal influence path chart;Sample quality grade mark correction and generate quality control analysis result.It solves the technical problems that sequencing data quality control dimension is single, dynamic adaptability is poor, deep abnormal recognition ability is insufficient in the prior art, leads to sequencing data quality control management quality and reliability is poor, achieves multi-dimensional precision quality control, improves sequencing data quality control management quality and reliability technical effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data management, and particularly relates to a multi-generation sequencing data quality control management method and system. BACKGROUND

[0002] Multi-generation sequencing platforms are widely used in the fields of genomics, transcriptomics, epigenetics, etc. However, the quality of sequencing data directly affects the accuracy of bioinformatics analysis, especially in key fields such as precision medicine, cancer genomics and genetic disease diagnosis. Low-quality data may lead to incorrect biological conclusions or clinical misjudgments. Traditional sequencing data quality control is usually based on single-dimensional indicators such as Phred quality score, sequencing depth, GC content, etc., which is difficult to comprehensively evaluate the biological rationality of the data. Moreover, multi-generation sequencing data is highly complex, with multiple noise sources such as base recognition errors, chimeric reads, and amplification bias. Existing quality control strategies cannot accurately and effectively identify deep abnormal data, thereby affecting the reliability of multi-generation sequencing data quality control management.

[0003] Therefore, in the related art, there are technical problems of single dimension of sequencing data quality control, poor dynamic adaptability, insufficient deep abnormality recognition capability, resulting in poor quality and reliability of sequencing data quality control management. SUMMARY

[0004] The present application provides a multi-generation sequencing data quality control management method and system, which solves the technical problems of single dimension of sequencing data quality control, poor dynamic adaptability, insufficient deep abnormality recognition capability, resulting in poor quality and reliability of sequencing data quality control management in the prior art, and achieves the technical effects of multi-dimensional accurate quality control and improved quality and reliability of sequencing data quality control management.

[0005] The present application provides a multi-generation sequencing data quality control management method, which comprises the following steps: reading file header data of sequencing data, loading a quality control rule template according to the file header data; performing feature extraction at a single read granularity after importing and converting the sequencing data, and establishing a feature extraction result; performing multi-layer nested quality control analysis according to the quality control rule template, the sequencing data and the feature extraction result, establishing a sample quality level identifier, wherein the multi-layer nested quality control analysis comprises original signal layer analysis, alignment mapping layer analysis, annotation variation layer analysis and biological logic layer analysis; configuring a trigger threshold, performing trigger analysis of the sample quality level identifier using the trigger threshold, performing traceability detection based on a traceability analysis channel according to the trigger analysis result, and establishing an abnormal influence path diagram; correcting the sample quality level identifier according to the abnormal influence path diagram, and generating a quality control analysis result.

[0006] In a possible implementation, the multi-generation sequencing data quality control management method further performs the following processing: calling a raw signal processing layer, reading base quality sequences and platform standard offset templates in feature extraction results; performing first-order difference on the base quality sequences to establish quality change sequences; calculating local offset trend vectors of the quality change sequences based on a sliding window; comparing the local offset trend vectors with the platform standard offset templates to establish read segment granularity-based offset anomaly scores; and completing raw signal layer analysis according to the offset anomaly scores.

[0007] In a possible implementation, the multi-generation sequencing data quality control management method further performs the following processing: calling a raw signal processing layer, reading sequence end base information at a read segment granularity, and loading an adapter template database by using the quality control rule template; constructing a base-Q value two-dimensional pattern graph according to the sequence end base information; performing similarity evaluation on the base-Q value two-dimensional pattern graph based on the adapter template database to generate adapter artifact anomaly scores; and completing raw signal layer analysis according to the adapter artifact anomaly scores and the offset anomaly scores.

[0008] In a possible implementation, the multi-generation sequencing data quality control management method further performs the following processing: calling an alignment mapping processing layer, reading alignment configuration parameters in the quality control rule template, selecting a reference genome and an alignment tool, performing standardized alignment on the sequencing data to generate an alignment result; extracting alignment positions, alignment scores, and alignment state identifiers at a read segment granularity according to the alignment result, establishing an association of read segment granularity feature extraction results, constructing an alignment mapping feature graph in combination with structural annotation information of the reference genome; performing position fidelity score analysis and mismatch pattern clustering identification analysis by using the alignment mapping feature graph, establishing an alignment quality index set, and completing alignment mapping layer analysis according to the alignment quality index set.

[0009] In a possible implementation, the multi-generation sequencing data quality control management method further performs the following processing: activating an annotated variation processing layer, calling the alignment result to perform variation detection, and establishing a variation candidate set; performing average base quality value calculation based on variation position range search on the variation candidate set to generate a first confidence score; performing annotated variation matching based on the quality control rule template on the variation candidate set, and establishing a second confidence score according to a matching result; and completing annotated variation layer analysis according to the first confidence score and the second confidence score.

[0010] In a possible implementation, the multi-generation sequencing data quality control management method further performs the following processing: activating a biological logic analysis layer, calling a population frequency database according to sample basic information of the sequencing data; performing variation sample distribution proportion analysis in the sequencing data by using the population frequency database, and establishing a proportion deviation value; and completing biological logic layer analysis according to the proportion deviation value.

[0011] In a possible implementation, the multi-generation sequencing data quality control management method further performs the following processing: activating a traceability analysis channel, inputting a trigger analysis result and a sample quality level identifier as input data into the traceability analysis channel; constructing an authentication question chain for each abnormal item, performing evidence query based on the authentication question chain, and establishing an authentication node; and completing traceability detection by using the authentication node, and establishing an abnormal influence path graph.

[0012] The application also provides a multi-generation sequencing data quality control management system, which comprises: a file header data reading module, configured to read file header data of sequencing data, and load a quality control rule template according to the file header data; a feature extraction result establishing module, configured to perform feature extraction on the sequencing data after import conversion, and establish a feature extraction result at a single read segment granularity; a quality control analysis module, configured to perform multi-layer nested quality control analysis according to the quality control rule template, the sequencing data and the feature extraction result, and establish a sample quality level identifier, wherein the multi-layer nested quality control analysis comprises original signal layer analysis, alignment mapping layer analysis, annotation variation layer analysis and biological logic layer analysis; a traceability detection module, configured to configure a trigger threshold, perform trigger analysis on the sample quality level identifier by using the trigger threshold, perform traceability detection based on a traceability analysis channel according to a trigger analysis result, and establish an abnormal influence path graph; and a quality control analysis result generating module, configured to generate a quality control analysis result after correcting the sample quality level identifier according to the abnormal influence path graph.

[0013] The application provides a multi-generation sequencing data quality control management method and system, which read file header data of sequencing data, load a quality control rule template, perform feature extraction on the sequencing data after import conversion at a single read segment granularity, establish a feature extraction result, perform multi-layer nested quality control analysis according to the quality control rule template, the sequencing data and the feature extraction result, establish a sample quality level identifier, configure a trigger threshold, perform trigger analysis on the sample quality level identifier, perform traceability detection, and establish an abnormal influence path graph. The technical problem of poor sequencing data quality control management quality and reliability due to single quality control dimension, poor dynamic adaptability and insufficient deep abnormality recognition capability in the prior art is solved, multi-dimensional accurate quality control is achieved, and the technical effect of improving sequencing data quality control management quality and reliability is achieved. BRIEF DESCRIPTION OF DRAWINGS

[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments of the present disclosure will be briefly introduced below. In the present application, flowcharts are used to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the foregoing or the following operations are not necessarily performed in sequence. On the contrary, various steps can be processed in reverse order or simultaneously according to needs. Meanwhile, other operations can be added to these processes, or one or more steps of operations can be removed from these processes.

[0015] Figure 1 A multi-generation sequencing data quality control management method flowchart is provided for the embodiments of the present application.

[0016] Figure 2 A multi-generation sequencing data quality control management system structure diagram is provided for the embodiments of the present application.

[0017] Legend: file header data reading module 10, feature extraction result establishing module 20, quality control analysis module 30, traceability detection module 40, quality control analysis result generating module 50. DETAILED DESCRIPTION

[0018] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application, the content of the specification can be implemented, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described.

[0019] In order to make the purposes, technical solutions and advantages of the present application more clear, the present application will be further described in detail below with reference to the drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without making creative labor belong to the scope of protection of the present application.

[0020] In the following description, "some embodiments" are referred to, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict, the term "first\second" referred to only distinguishes similar objects, and does not represent a specific order for the objects. The terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or modules not clearly listed or inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application.

[0021] The embodiments of the present application provide a multi-generation sequencing data quality control management method, as shown in the method, which comprises: Figure 1

[0022] Step S100, reading the file header data of the sequencing data, and loading the quality control rule template according to the file header data.

[0023] Preferably, in the process of sequencing data analysis, the sequencing data is usually stored in a standard file format, and the header part of the file contains key metadata information. The file header data of the sequencing data is read and parsed, and a suitable quality control rule template is dynamically selected to ensure that the quality control process can adapt to different sequencing platforms, experimental types or data analysis requirements. The file header data of the sequencing data includes sequencing platform information, sequencing running parameters, sample information, and data generation time and instrument ID. Specifically, different platforms in the sequencing platform information need different quality control strategies; the sequencing running parameters such as read length, sequencing chemistry version and chip type affect the data quality evaluation standard, such as additional detection of chimeric reads for long read data; the sample information such as sample ID and library preparation method needs special quality control rules; the data generation time and instrument ID can be used to identify batch effects or systematic bias. The quality control rule template is loaded and adjusted according to the key fields of the header data, mainly including platform adaptation rules, experimental type rules and dynamic threshold adjustment, for example, Illumina data needs to check the Phred quality value distribution, and Nanopore data needs to pay attention to signal stability; whole genome sequencing pays attention to coverage uniformity, and RNA-seq needs to check strand specificity; tumor low-frequency mutation detection needs a more stringent alignment quality threshold, and population genetics can relax some standards. Further ensure that the multi-layer nested quality control can be based on the correct benchmark, thereby improving the accuracy and reliability of the overall quality control.

[0024] ​Step S200, after importing and converting the sequencing data, feature extraction is performed at single read granularity to establish a feature extraction result.

[0025] Preferably, importing and converting the sequencing data refers to reading the original sequencing data from a storage format and converting it into a standardized data structure, ensuring data integrity and format uniformity. Specifically, the sequencing data is parsed, the file header and main content are read, metadata such as sequencing platform and sample information are extracted, encoding conversion is performed, and the encoding method of base sequence and quality score data is ensured to meet the analysis standard. At the same time, data verification is performed, including checking file integrity and whether there are damaged reads. Then, the read is objectified, i.e. each read is converted into a data object containing uniform fields, including sequence, quality score, read ID, and alignment information, and then converted into a standardized data structure. Then, feature extraction is performed at single read granularity, wherein the sequencing data usually consists of millions or even billions of reads, each representing a DNA / RNA fragment being sequenced. Single read granularity analysis refers to analyzing each read and performing anomaly detection to find local problems, such as local low quality and alignment anomalies of a read. Then, feature extraction is performed for each read, including but not limited to original signal layer features such as base quality distribution, signal intensity anomaly, and low complexity sequence; alignment mapping layer features such as alignment position information, insertions / deletions, and soft clipping proportion; sequence composition features such as GC content deviation and k-mer frequency anomaly; structural anomaly features such as chimeric reads and read directionality. Finally, the extracted features of each read are structured and stored to form a feature extraction result, which may include a feature matrix, each row representing a read and each column representing a feature; anomaly labels, labeling potential problems for each read; and metadata association with file header information such as sequencing platform and sample ID, thereby significantly improving the accuracy and reliability of sequencing data quality control.

[0026] Step S300, according to the quality control rule template, the sequencing data, and the feature extraction result, multi-layer nested quality control analysis is performed to establish a sample quality level identifier, wherein the multi-layer nested quality control analysis includes original signal layer analysis, alignment mapping layer analysis, annotation variation layer analysis, and biological logic layer analysis.

[0027] Preferably, the multi-layer nested quality control analysis is performed according to the quality control rule template, the sequencing data and the feature extraction result, i.e. the data quality is checked step by step from different levels to finally form a quality grade identification of the sample, such as qualified, suspicious and unqualified, so as to ensure the accuracy and reliability of the data, wherein the multi-layer nested quality control analysis includes raw signal layer analysis, alignment mapping layer analysis, annotation variation layer analysis and biological logic layer analysis. Specifically, the raw signal layer analysis refers to evaluating the quality of the raw signal generated by the sequencer, identifying technical noise, including analyzing base quality distribution, checking Phred quality score and removing low-quality reads, for the FASTQ (base sequence and quality score) and PacBio / Nanopore current signal data.

[0028] Preferably, the alignment mapping layer analysis refers to evaluating the accuracy of the read alignment to the reference genome, identifying abnormal alignment, including analyzing alignment rate, the proportion of reads successfully aligned to the reference genome, low alignment rate may indicate contamination or mismatch of the reference genome, alignment quality, high alignment quality value represents high alignment reliability, low alignment quality value may indicate repeat region or alignment error, insert size distribution, checking whether it meets the library building expectation, soft clipping proportion, too much unaligned part at both ends of the read may indicate structural variation or sequencing error, strand specificity, checking whether it meets the expected direction of strand-specific library, and further marking abnormal alignment reads and calculating effective alignment rate.

[0029] Preferably, the annotation variation layer analysis refers to evaluating whether the detected variation is reliable and excluding technical false positives, including analyzing variation quality, filtering low-confidence variations, allele frequency, checking whether it meets the expectation, such as the allele frequency of somatic mutation of tumor samples is usually low, sequencing depth, ensuring sufficient coverage of the variation site, strand bias, whether the variation only occurs in a single strand, known database alignment, excluding common polymorphic sites, and further marking high-confidence variations and calculating the false positive rate of variation detection.

[0030] Preferably, the biological logic layer analysis refers to verifying data from the perspective of biological rationality, including Mendelian inheritance consistency, checking whether the offspring mutation is consistent with the parent genetic law, tumor-normal paired samples, whether somatic mutations do not exist in normal samples, gene expression correlation, checking whether the expression pattern between samples is consistent with the expectation, known biological pathways, such as whether cancer driver gene mutations are within the expected range, and further marking abnormalities that do not conform to biological logic, such as samples with homozygous mutations that neither parent has, which may be contaminated. Finally, based on the results of each layer of analysis, the sample is comprehensively scored and graded, for example, qualified, all levels meet the quality control standards; suspicious, some indicators are abnormal, such as slightly low alignment rate; unqualified, key indicators deviate seriously or there are a large number of mutations that do not conform to biological laws. Through layer-by-layer progression and cross-validation, it is ensured that the sequencing data is reliable from the technical level to the biological level, and finally the accurate quality grade mark is output.

[0031] Further, step S300 further comprises step S310 of calling the original signal processing layer to read the base quality sequence in the feature extraction result and the platform standard offset template; step S320 of performing first-order difference on the base quality sequence to establish a quality change sequence; step S330 of calculating a local offset trend vector of the quality change sequence based on a sliding window; step S340 of comparing the local offset trend vector with the platform standard offset template to establish a read segment granularity offset anomaly score; and step S350 of completing the original signal layer analysis according to the offset anomaly score.

[0032] Preferably, the standardized operation of the sequencer should maintain stable base recognition quality, and if the quality change pattern of a certain sequence deviates from the platform expectation, it may indicate an abnormality, and the raw signal processing layer is called to read the base quality sequence in the feature extraction result and the platform standard offset template, wherein the base quality sequence refers to the Phred quality score of each position of a single read, and the platform standard offset template refers to the pre-defined quality change pattern of the sequencing platform under normal operation, including the expected quality change curve and the allowed fluctuation range of different sequencing cycles. Specifically, the base quality sequence is first-order differentiated to obtain a quality change sequence by calculating the difference between adjacent positions to amplify the mutation points of the quality value and facilitate the detection of abnormal fluctuations; then, a local offset trend vector of the quality change sequence is calculated based on a sliding window, that is, a fixed length (such as 5 bp) window is used to traverse the difference sequence, and statistics in the window are calculated, including the mean of the difference in the window, reflecting the local quality change direction, and the standard deviation of the difference in the window, reflecting the local fluctuation intensity, and then a local offset trend vector is obtained, representing the local pattern of quality change, wherein the local offset trend vector is a fixed length vector. The local offset trend vector is compared with the platform standard offset template, the deviation degree is calculated based on dynamic time warping or Euclidean distance, and then an offset anomaly score of a single read is established, for example, if a read has a sharp quality drop at positions 50-60 bp, the anomaly score of this region will be significantly increased; finally, the raw signal layer analysis is completed according to the offset anomaly score, including outputting the single read offset anomaly score to mark the abnormal read and outputting the global anomaly report to count the abnormal patterns of all reads.

[0033] Further, step S350 further includes step S351, calling the raw signal processing layer to read the sequence end base information at a single read granularity, and loading an adapter template database using the quality control rule template; step S352, constructing a base-Q value two-dimensional pattern graph according to the sequence end base information; step S353, performing similarity evaluation on the base-Q value two-dimensional pattern graph based on the adapter template database to generate an adapter artifact anomaly score; and step S354, completing raw signal layer analysis according to the adapter artifact anomaly score and the offset anomaly score.

[0034] Preferably, the original signal layer analysis process focuses on detecting and identifying data quality abnormalities caused by library adapter contamination or sequencer end artifacts in the sequencing data by combining single read end base pattern analysis and adapter database comparison. Specifically, the original signal processing layer is called to read the sequence end base information of single read granularity, i.e. to extract the base sequence and corresponding Phred quality score (Q value) of a fixed length at the end of each read, and to load the adapter template database from the quality control rule template, which contains known adapter sequences and corresponding expected Q values. Then the frequency of each base at the end position is counted, the mean value curve of the end Q value is drawn, and the base composition of the end sequence is associated with the quality score to form a patternable two-dimensional feature, i.e. a base-Q value two-dimensional pattern graph. For example, for a normal read end, the bases are randomly distributed and the Q value decreases smoothly; for an adapter contaminated end, a specific base appears frequently and the Q value drops sharply.

[0035] Preferably, the base-Q value two-dimensional pattern graph is evaluated based on the adapter template database, i.e. the read end pattern is compared with the templates in the adapter database in multiple dimensions, including calculating the coincidence degree of the end sequence with the known adapter using k-mer matching, and checking whether the Q value suddenly decreases in the adapter matching area to identify the Q value abnormal pattern. Then the sequence matching degree and Q value drop amplitude are weighted to obtain the adapter artifact abnormal score. If the adapter artifact abnormal score is greater than a preset threshold (such as 80 points), the read is marked as adapter contaminated. Finally, the original signal layer analysis is completed by combining the adapter artifact abnormal score and the offset abnormal score. Specifically, if both are normal, the read is marked as high quality; if only the offset is abnormal, it may be a temporary fault of the sequencer; if only the adapter is abnormal, it may indicate a library problem, such as incomplete removal of the adapter; if both are abnormal, it may be a serious technical problem, such as interruption of sequencing causing adapter sequence to mix into the data sequence; and finally the abnormal score is output for sample quality level correction to ensure high quality data quality control management.

[0036] Further, step S300 further includes step S360 of calling the alignment mapping processing layer, reading the alignment configuration parameters in the quality control rule template, selecting a reference genome and an alignment tool, and performing standardized alignment on the sequencing data to generate an alignment result; step S370 of extracting the alignment position, alignment score and alignment state identifier of single read granularity from the alignment result, establishing the association of single read granularity feature extraction results, combining the structural annotation information of the reference genome, and constructing an alignment mapping feature map; and step S380 of performing position fidelity score analysis and mismatch pattern clustering identification analysis using the alignment mapping feature map, establishing an alignment quality index set, and completing alignment mapping layer analysis according to the alignment quality index set.

[0037] Preferably, alignment mapping layer analysis is the core of sequencing data quality control. By aligning the sequencing reads with the reference genome, the accuracy and reliability of read positioning are evaluated to identify potential alignment errors, contamination or technical bias. Specifically, the alignment mapping processing layer is called, the alignment configuration parameters in the quality control rule template are read, including the reference genome version, alignment tools and alignment algorithm parameters such as the number of allowed mismatches and gap penalties, and then the corresponding reference genome is dynamically loaded according to the experimental type, such as human whole genome, microbial sequencing, while the preset alignment tool such as BWA-MEM is called to perform standardized alignment. For long read data, a relaxed gap penalty is enabled, and for short read data, the number of mismatches is strictly limited, and then the alignment results, i.e. BAM / SAM files, are generated, containing the alignment position, CIGAR string, alignment quality score and other information of each read.

[0038] Preferably, according to the alignment results, the alignment position, alignment score and alignment state identifier of single read granularity are extracted, wherein the alignment position is the coordinate of the read on the reference genome; the alignment score refers to a score of 0-60, the higher the score, the more unique and reliable the alignment, and an alignment score of 0 usually indicates multiple alignment; the alignment state identifier includes unique alignment or multiple alignment and soft clipping proportion; and then the association of single read granularity feature extraction results is established, i.e. the alignment position is cross-analyzed with the gene structure coordinates using a genome annotation tool, and combined with the structural annotation information of the reference genome, such as whether the alignment region falls within the exon, intron or non-coding region of a gene, an alignment mapping feature map is constructed, wherein the alignment features of each read, such as alignment quality score, gene region and original signal layer feature Q value are associated, for example, a read with high alignment quality score but high soft clipping proportion may indicate a structural variation.

[0039] Preferably, based on the alignment mapping feature map, position fidelity score analysis and mismatch mode cluster identification analysis are performed. Specifically, the position fidelity score analysis refers to calculating the overall alignment quality score mean of the sample; calculating the unique alignment rate, i.e. the proportion of unique alignment reads to total alignment reads; chain specificity check, whether the alignment direction conforms to the library expectation, such as FR / RF direction of chain-specific library; wherein, low unique alignment rate may indicate reference genome mismatch or sample contamination. The mismatch mode cluster identification analysis refers to mismatch positioning, i.e. statistics of the distribution of non-matching bases in reads, if multiple reads at a certain genomic position all have mismatches, it may be a real variation or a reference genome error; insertion / deletion mode, i.e. detecting whether the length and position of Indel are concentrated in a specific region, such as microsatellite sequence; using machine learning such as DBSCAN to cluster mismatches / Indels, to distinguish random errors from technical bias, and then output an alignment quality index set, including unique alignment rate, average alignment quality score, and mismatch cluster significance. Finally, according to the alignment quality index set, the alignment mapping layer analysis is completed, i.e. high unique alignment rate, high average alignment quality score and no significant mismatch cluster, which indicates that it is qualified; slightly low unique alignment rate but mismatch mode conforms to known technical bias, which indicates that it is suspicious; low unique alignment rate or high frequency mismatch cluster, which indicates that it is unqualified. Further, fine error recognition is realized to distinguish sequencing errors from real biological phenomena, so as to comprehensively evaluate the positioning reliability of sequencing data.

[0040] Further, step S300 further includes step S390 of activating an annotated variant processing layer, calling the alignment result to perform variant detection, and establishing a variant candidate set; step S3100 of calculating the average base quality value based on the variant position range search based on the variant candidate set, to generate a first confidence score; step S3110 of performing annotated variant matching based on the quality control rule template on the variant candidate set, and establishing a second confidence score according to the matching result; and step S3120 of completing annotated variant layer analysis according to the first confidence score and the second confidence score.

[0041] Preferably, the annotated variant layer analysis is a key link in the quality control of sequencing data, through the dual verification of technical indicators and biological annotations, false positive variants are eliminated, the reliability of detected genomic variants is evaluated, and the accuracy of subsequent analysis is ensured. Specifically, the annotated variant processing layer is activated, and the variant detection tool is called to perform variant detection on the alignment result, i.e., the tool detects variants from the alignment result to generate an initial VCF file for storing genetic variant information, including basic information such as variant type, position, genotype, and quality score. Variants in low-complexity regions are removed, and a variant candidate set containing potential variant sites to be evaluated is generated. Then, based on the variant candidate set, the average base quality value is calculated based on the variant position range search to evaluate the reliability of the variant site from the sequencing technology perspective, i.e., for each variant site, the base quality values of all reads supporting the variant at the variant position are extracted and the average value is calculated. Then, the coverage depth and strand balance are checked, and then the first reliability score is output. The higher the score, the stronger the technical reliability.

[0042] Preferably, the annotated variant layer analysis is a key link in the quality control of sequencing data, through the dual verification of technical indicators and biological annotations, false positive variants are eliminated, the reliability of detected genomic variants is evaluated, and the accuracy of subsequent analysis is ensured. Specifically, the annotated variant processing layer is activated, and the variant detection tool is called to perform variant detection on the alignment result, i.e., the tool detects variants from the alignment result to generate an initial VCF file for storing genetic variant information, including basic information such as variant type, position, genotype, and quality score. Variants in low-complexity regions are removed, and a variant candidate set containing potential variant sites to be evaluated is generated. Then, based on the variant candidate set, the average base quality value is calculated based on the variant position range search to evaluate the reliability of the variant site from the sequencing technology perspective, i.e., for each variant site, the base quality values of all reads supporting the variant at the variant position are extracted and the average value is calculated. Then, the coverage depth and strand balance are checked, and then the first reliability score is output. The higher the score, the stronger the technical reliability.

[0043] Further, step S300 further includes step S3130 of activating a biological logic analysis layer, calling a population frequency database according to sample basic information of the sequencing data; step S3140 of performing variant sample distribution proportion analysis in the sequencing data using the population frequency database to establish a proportion deviation value; and step S3150 of completing biological logic layer analysis according to the proportion deviation value.

[0044] Preferably, the biological logic layer analysis is the final key link of sequencing data quality control. By comparing the distribution frequency of variations in the sample with the expected distribution of the known population database, possible technical false positives or biological abnormalities, such as sample contamination, family relationship errors, etc. are identified, and the credibility of the variation data is verified from the perspective of population genetics and biological rationality. Among them, the sample basic information of the sequencing data includes the sequencing type and the sample type, such as tumor / normal pairing, family members, specific ethnic groups, and the population frequency database includes public databases and local databases. The public database contains the thousand genome project, and the local database contains private frequency data of specific populations or diseases, such as the tumor mutation spectrum of the East Asian population. The population frequency database is used for variation sample distribution ratio analysis in the sequencing data. That is, for each variation site, the allele frequency of the corresponding population in the population database is queried, and the deviation of the sample variation frequency from the population frequency is calculated by Z-score standardization, that is, the ratio deviation value is obtained, and the frequency distribution pattern of all variations in the sample is counted. Finally, the biological logic layer analysis is completed according to the ratio deviation value, including identifying technical abnormalities and biological abnormalities through the ratio deviation value. When all variation frequencies meet the population expectation and there is no global distribution anomaly, the analysis passes; if some variations deviate, but can be explained by technical reasons, such as clonal selection of tumor samples, an alert is output; if there is a significant deviation and cannot be reasonably explained, such as the high-frequency occurrence of oncogenic mutations in healthy samples, the analysis fails, thereby ensuring the quality and reliability of sequencing data quality control management.

[0045] Step S400, configure a trigger threshold, use the trigger threshold for trigger analysis of sample quality level identification, perform traceability detection based on a traceability analysis channel according to the trigger analysis result, and establish an abnormal influence path diagram.

[0046] Step S400 further includes step S410 of activating the traceability analysis channel, inputting the trigger analysis result and sample quality level identification as input data into the traceability analysis channel; step S420 of constructing an authentication question chain for each abnormal item, performing evidence query based on the authentication question chain, and establishing an authentication node; and step S430 of completing traceability detection through the authentication node and establishing an abnormal influence path diagram.

[0047] Preferably, according to the different quality control levels, multi-dimensional trigger thresholds are configured for automatic determination of sample quality. The trigger threshold can be automatically adapted according to the experimental type. The raw signal layer threshold includes that the base proportion is greater than or equal to 80% for qualification and the adapter contamination rate is less than or equal to 5% for qualification. The alignment layer threshold includes that the unique alignment rate is greater than or equal to 90% for qualification and the average alignment quality is greater than or equal to 30 for qualification. The annotation variation layer threshold includes that the high-confidence variation proportion is greater than or equal to 95% for qualification and the number of known pathogenic variations is greater than or equal to 1 driver mutation for expectation. The biological logic layer threshold includes that the population frequency deviation value is less than or equal to 3 for qualification and the Mendelian conflict rate of the family is less than or equal to 1% for qualification.

[0048] Preferably, the trigger analysis of sample quality grade identification is performed using trigger thresholds, that is, the results of each layer of quality control analysis are compared with trigger thresholds to determine whether to trigger. Specifically, the output indicators of each layer are compared with the corresponding trigger thresholds to determine whether to trigger a warning or failure. For example, if the Q30 is 75%, the raw signal layer warning is triggered, and if the unique alignment rate is 85%, the alignment layer warning is triggered. The hierarchical weights are set according to experimental requirements, for example, clinical samples pay more attention to variation layer and biological logic layer, and the trigger analysis results are determined comprehensively. If all levels do not trigger the threshold or only the minor level deviates slightly, it is determined to be qualified; if 1-2 major levels trigger a warning but do not reach the failure threshold, it is determined to be suspicious; if any core level triggers a failure or multiple levels deviate seriously, it is determined to be unqualified.

[0049] Preferably, the traceability detection based on the traceability analysis channel is performed according to the trigger analysis results, wherein the traceability analysis channel is an intelligent root cause diagnosis unit in the sequencing data quality control process. Through structured problem chain reasoning and evidence integration, abnormalities found in the quality control process, such as low-quality samples and variation conflicts, are traced back to specific technical or biological roots, and a visual abnormality impact path diagram is generated to guide problem repair or data correction. Specifically, the traceability analysis channel is activated, and the trigger analysis results and sample quality grade identification are input as input data into the traceability analysis channel. An authentication problem chain is automatically generated for each abnormal item, and the root cause is locked through progressive questioning. For example, the example problem chain for low alignment rate may be whether the low alignment rate is concentrated in a specific chromosome region to check the uniformity of alignment distribution; whether the low alignment read has high adapter contamination to associate the raw signal layer result; whether the same batch of samples all have this problem to query the batch metadata.

[0050] Preferably, evidence query is further performed based on the authentication problem chain, that is, the authentication problem is answered through multi-source data retrieval to form an authentication node, that is, an intermediate conclusion supported by evidence. The evidence types may include technical evidence such as abnormal alarms in the sequencer log and low concentration prompts in the library QC report; data evidence such as similar abnormal patterns of other samples in the same batch; and rule evidence such as common end degradation problems of a certain platform marked in the quality control rule template. Finally, the traceability detection is completed through the authentication node, that is, the authentication nodes are integrated to construct a directed acyclic graph, that is, an abnormality impact path diagram, which directly displays the abnormality propagation path. The abnormality impact path diagram includes root causes such as DNA degradation leading to insufficient capture of telomere regions, intermediate impacts such as low alignment rate leading to reduced variation detection sensitivity, and final consequences such as sample degradation to suspicious and reduced tumor mutation detection rate. Thus, discrete quality control abnormalities are converted into actionable attribution conclusions to improve the quality and reliability of sequencing data quality control management.

[0051] Step S500, after the sample quality level identification is corrected according to the abnormal influence path diagram, a quality control analysis result is generated.

[0052] Preferably, in combination with the dynamic feedback of the abnormal influence path diagram, the sample quality level identification is intelligently corrected, that is, according to the explainability and severity of the abnormal root cause, it is determined whether to downgrade, maintain or upgrade the sample quality level, and the analysis results of each level, the abnormal path diagram and the correction basis are integrated to form a traceable quality control analysis result. Specifically, if the abnormality can be clearly attributed to a technical problem, such as adapter contamination or batch effect, and has limited impact on downstream analysis, it is adjusted to suspicious; if the abnormality involves biological contradictions, such as family conflicts or high-frequency driver mutations in normal samples, it is maintained as unqualified; abnormalities that cannot be automatically attributed, such as newly discovered sequencer error patterns, trigger a manual review process, and the report indicates that manual review is required; finally, a quality control analysis result is generated, including the final version of the sample quality identification, a summary of key abnormalities, an abnormal influence path diagram, and recommended measures, such as technical abnormalities suggesting re-sequencing and increasing DNA input. Further reducing data waste, enhancing the credibility of quality control analysis results, and ensuring the quality and reliability of sequencing data quality control management.

[0053] In the foregoing, with reference to Figure 1 A multi-generation sequencing data quality control management method according to an embodiment of the application is described in detail. Next, with reference to Figure 2 A multi-generation sequencing data quality control management system according to an embodiment of the application will be described.

[0054] A multi-generation sequencing data quality control management system according to an embodiment of the application is used to solve the technical problem of single dimension of sequencing data quality control, poor dynamic adaptability, insufficient deep abnormality recognition capability in the prior art, which leads to poor quality and reliability of sequencing data quality control management. It achieves multi-dimensional precision quality control and improves the quality and reliability of sequencing data quality control management. As shown in Figure 2 A multi-generation sequencing data quality control management system includes a file header data reading module 10, a feature extraction result establishing module 20, a quality control analysis module 30, a traceability detection module 40, and a quality control analysis result generating module 50.

[0055] The file header data reading module 10 is configured to read file header data of sequencing data, and load a quality control rule template according to the file header data; the feature extraction result establishing module 20 is configured to perform feature extraction on the sequencing data after import conversion, and establish a feature extraction result at a single read granularity; the quality control analysis module 30 is configured to perform multi-layer nested quality control analysis according to the quality control rule template, the sequencing data and the feature extraction result, establish a sample quality level identifier, and the multi-layer nested quality control analysis includes raw signal layer analysis, alignment mapping layer analysis, annotation variation layer analysis and biological logic layer analysis; the traceability detection module 40 is configured to configure a trigger threshold, perform trigger analysis on the sample quality level identifier by using the trigger threshold, perform traceability detection based on a traceability analysis channel according to a trigger analysis result, and establish an abnormal influence path diagram; and the quality control analysis result generating module 50 is configured to generate a quality control analysis result after correcting the sample quality level identifier according to the abnormal influence path diagram.

[0056] Next, the specific configuration of the quality control analysis module 30 will be described in detail. The quality control analysis module 30 further includes: calling a raw signal processing layer, reading base quality sequences in the feature extraction result and a platform standard offset template; performing first-order difference on the base quality sequences to establish quality change sequences; calculating local offset trend vectors of the quality change sequences based on a sliding window; comparing the local offset trend vectors with the platform standard offset template to establish offset anomaly scores at a single read granularity; and completing the raw signal layer analysis according to the offset anomaly scores.

[0057] Next, the specific configuration of the quality control analysis module 30 will be described in detail. The quality control analysis module 30 further includes: calling a raw signal processing layer, reading base quality sequences in the feature extraction result and a platform standard offset template; performing first-order difference on the base quality sequences to establish quality change sequences; calculating local offset trend vectors of the quality change sequences based on a sliding window; comparing the local offset trend vectors with the platform standard offset template to establish offset anomaly scores at a single read granularity; and completing the raw signal layer analysis according to the offset anomaly scores.

[0058] Next, the specific configuration of the quality control analysis module 30 will be described in detail. The quality control analysis module 30 further comprises: a calling alignment mapping processing layer, reading the alignment configuration parameters in the quality control rule template, selecting a reference genome and an alignment tool, performing standard alignment on the sequencing data, and generating an alignment result; extracting the alignment position, alignment score and alignment state identifier of the single read granularity according to the alignment result, establishing the association of the single read granularity feature extraction result, combining the structure annotation information of the reference genome, and constructing an alignment mapping feature map; using the alignment mapping feature map to perform position fidelity scoring analysis and mismatch pattern clustering identification analysis, establishing an alignment quality index set, and completing alignment mapping layer analysis according to the alignment quality index set.

[0059] Next, the specific configuration of the quality control analysis module 30 will be described in detail. The quality control analysis module 30 further comprises: an activated annotation variation processing layer, calling the alignment result to perform variation detection, and establishing a variation candidate set; performing average base quality value calculation based on variation position range search on the variation candidate set, generating a first confidence score; performing annotation variation matching based on the quality control rule template on the variation candidate set, and establishing a second confidence score according to the matching result; completing annotation variation layer analysis according to the first confidence score and the second confidence score.

[0060] Next, the specific configuration of the quality control analysis module 30 will be described in detail. The quality control analysis module 30 further comprises: activating a biological logic analysis layer, calling a population frequency database according to the sample basic information of the sequencing data; using the population frequency database to perform variation sample distribution proportion analysis in the sequencing data, and establishing a proportion deviation value; completing biological logic layer analysis according to the proportion deviation value.

[0061] Next, the specific configuration of the traceability detection module 40 will be described in detail. The traceability detection module 40 further comprises: activating a traceability analysis channel, inputting the trigger analysis result and sample quality level identifier as input data into the traceability analysis channel; constructing an authentication question chain for each abnormal item, performing evidence query based on the authentication question chain, and establishing an authentication node; completing traceability detection through the authentication node, and establishing an abnormal influence path graph.

[0062] The multi-generation sequencing data quality control management system provided by the embodiment of the application can execute the multi-generation sequencing data quality control management method provided by any embodiment of the application, and has the corresponding function modules and beneficial effects of the execution method.

[0063] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or the server, the various units and modules included are only divided according to the functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific name of each functional unit is only for the convenience of mutual differentiation, and is not used to limit the protection scope of the present application.

[0064] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for quality control management of multi-generation sequencing data, characterized in that, The method includes: Read the header data of the sequencing data file, and load the quality control rule template according to the header data. The header data includes sequencing platform information, sequencing operation parameters, sample information, data generation time, and instrument ID. After importing and converting the sequencing data, feature extraction is performed at the single-read granularity to establish feature extraction results. Each read in the single-read granularity represents a sequenced DNA / RNA fragment. Based on the quality control rule template, the sequencing data, and the feature extraction results, a multi-layer nested quality control analysis is performed to establish a sample quality level identifier. The multi-layer nested quality control analysis includes original signal layer analysis, alignment mapping layer analysis, annotation variation layer analysis, and biological logic layer analysis. Configure a trigger threshold, use the trigger threshold to perform trigger analysis for sample quality level identification, perform source tracing detection based on the source tracing analysis channel according to the trigger analysis results, and establish an anomaly impact path diagram; After correcting the sample quality level label according to the abnormal impact path diagram, the quality control analysis results are generated. Call the original signal processing layer to read the base quality sequence and platform standard offset template from the feature extraction results; The quality sequence of the bases is subjected to first-order difference to establish a quality change sequence; Calculate the local offset trend vector of the quality change sequence based on a sliding window; The local offset trend vector is compared using the platform's standard offset template to establish an offset anomaly score at the single-read segment level. The original signal layer analysis is completed based on the offset anomaly. The alignment mapping processing layer is invoked to read the alignment configuration parameters in the quality control rule template, select the reference genome and alignment tool, perform standardized alignment of the sequencing data, and generate alignment results. Based on the alignment results, the alignment position, alignment score, and alignment status identifier of single read granularity are extracted, the association of single read granularity feature extraction results is established, and the alignment mapping feature map is constructed by combining the structural annotation information of the reference genome. The location fidelity scoring analysis and mismatch pattern clustering identification analysis are performed using the comparison mapping feature map, a comparison quality index set is established, and the comparison mapping layer analysis is completed based on the comparison quality index set. Activate the annotation mutation processing layer, call the alignment results to perform mutation detection, and establish a mutation candidate set; The first confidence score is generated by calculating the average base quality value based on the mutation position range search of the mutation candidate set. The candidate mutation set is matched with annotated mutations based on quality control rule templates, and a second credibility score is established based on the matching results. Complete the annotation variation layer analysis based on the first confidence score and the second confidence score; Activate the biological logic analysis layer and call the population frequency database based on the basic sample information of the sequencing data; The population frequency database was used to analyze the distribution ratio of variant samples within the sequencing data, and a ratio deviation value was established. The biological logic layer analysis is completed based on the stated deviation value.

2. The method for quality control management of multi-generation sequencing data as described in claim 1, characterized in that, The step of performing original signal layer analysis based on the offset anomaly includes: The original signal processing layer is invoked to read the sequence terminal base information at the single-segment granularity, and the adapter template database is loaded using the quality control rule template. A two-dimensional base-Q value pattern diagram is constructed based on the terminal base information of the sequence. Based on the connector template database, a similarity assessment is performed on the base-Q value two-dimensional pattern diagram to generate connector artifact anomaly scores; The original signal layer analysis is completed based on the joint artifact anomaly score and the offset anomaly score.

3. The method for quality control management of multi-generation sequencing data as described in claim 1, characterized in that, The step of performing source tracing detection based on the source tracing analysis channel according to the trigger analysis results and establishing an anomaly impact path graph includes: Activate the source tracing analysis channel and input the trigger analysis results and sample quality level identifier as input data into the source tracing analysis channel; For each anomaly, an authentication question chain is constructed, and evidence is queried based on the authentication question chain to establish an authentication node. The source tracing and detection are completed through the authentication node, and an anomaly impact path diagram is established.

4. A multi-generation sequencing data quality control management system, characterized in that, The system is used to implement the multi-generation sequencing data quality control management method according to any one of claims 1 to 3, the system comprising: The file header data reading module is used to read the file header data of sequencing data and load the quality control rule template according to the file header data. The file header data includes sequencing platform information, sequencing operation parameters, sample information, data generation time, and instrument ID. The feature extraction result establishment module is used to import and convert the sequencing data, perform feature extraction at the single read granularity, and establish feature extraction results. Each read in the single read granularity represents a sequenced DNA / RNA fragment. The quality control analysis module is used to perform multi-layer nested quality control analysis based on the quality control rule template, the sequencing data, and the feature extraction results, and to establish a sample quality level identifier. The multi-layer nested quality control analysis includes original signal layer analysis, alignment mapping layer analysis, annotation variation layer analysis, and biological logic layer analysis. The source tracing and detection module is used to configure trigger thresholds, perform trigger analysis for sample quality level identification using the trigger thresholds, execute source tracing and detection based on the source tracing analysis channel according to the trigger analysis results, and establish an abnormal impact path diagram. The quality control analysis result generation module is used to generate quality control analysis results after correcting the sample quality level identifier according to the abnormal influence path diagram.

Citation Information

Patent Citations

  • Test data management and analysis system

    CN119091967A

  • Accurate comparison and validation of single nucleotide variants

    US20130245958A1