Multi-generation sequencing data quality control management method and system
By employing multi-layered nested quality control analysis and source tracing detection, the problems of single-dimensional quality control of sequencing data and insufficient identification of deep anomalies have been solved, thereby improving the quality and reliability of sequencing data quality control management.
Patent Information
- Application Number
- CN202511332142.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-18
AI Technical Summary
The existing technology has a single dimension for sequencing data quality control, poor dynamic adaptability, and insufficient ability to identify deep abnormalities, resulting in poor quality and reliability of sequencing data quality control management.
This paper provides a method for quality control management of multi-generation sequencing data. By reading the header data of the file, loading the quality control rule template, performing multi-layer nested quality control analysis, including the original signal layer, the alignment mapping layer, the annotation variation layer, and the biological logic layer analysis, configuring the trigger threshold, performing source tracing detection, and generating quality control analysis results.
It achieves multi-dimensional precise quality control, improves the quality and reliability of sequencing data quality control management, and ensures the accuracy and reliability of the data.
Smart Images

Figure CN120833091A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data management, and particularly relates to a multi-generation sequencing data quality control management method and system. BACKGROUND
[0002] Multi-generation sequencing platforms are widely used in the fields of genomics, transcriptomics, epigenetics, etc. However, the quality of sequencing data directly affects the accuracy of bioinformatics analysis, especially in key fields such as precision medicine, cancer genomics and genetic disease diagnosis. Low-quality data may lead to incorrect biological conclusions or clinical misjudgments. Traditional sequencing data quality control is usually based on single-dimensional indicators such as Phred quality score, sequencing depth, GC content, etc., which is difficult to comprehensively evaluate the biological rationality of the data. Moreover, multi-generation sequencing data is highly complex, with multiple noise sources such as base recognition errors, chimeric reads, and amplification bias. Existing quality control strategies cannot accurately and effectively identify deep abnormal data, thereby affecting the reliability of multi-generation sequencing data quality control management.
[0003] Therefore, in the related art, there are technical problems of single dimension of sequencing data quality control, poor dynamic adaptability, insufficient deep abnormality recognition capability, resulting in poor quality and reliability of sequencing data quality control management. SUMMARY
[0004] The present application provides a multi-generation sequencing data quality control management method and system, which solves the technical problems of single dimension of sequencing data quality control, poor dynamic adaptability, insufficient deep abnormality recognition capability, resulting in poor quality and reliability of sequencing data quality control management in the prior art, and achieves the technical effects of multi-dimensional accurate quality control and improved quality and reliability of sequencing data quality control management.
[0005] The present application provides a multi-generation sequencing data quality control management method, which comprises the following steps: reading file header data of sequencing data, loading a quality control rule template according to the file header data; performing feature extraction at a single read granularity after importing and converting the sequencing data, and establishing a feature extraction result; performing multi-layer nested quality control analysis according to the quality control rule template, the sequencing data and the feature extraction result, establishing a sample quality level identifier, wherein the multi-layer nested quality control analysis comprises original signal layer analysis, alignment mapping layer analysis, annotation variation layer analysis and biological logic layer analysis; configuring a trigger threshold, performing trigger analysis of the sample quality level identifier using the trigger threshold, performing traceability detection based on a traceability analysis channel according to the trigger analysis result, and establishing an abnormal influence path diagram; correcting the sample quality level identifier according to the abnormal influence path diagram, and generating a quality control analysis result.
[0006] In a possible implementation, the multi-generation sequencing data quality control management method further performs the following processing: calling a raw signal processing layer, reading base quality sequences and platform standard offset templates in feature extraction results; performing first-order difference on the base quality sequences to establish quality change sequences; calculating local offset trend vectors of the quality change sequences based on a sliding window; comparing the local offset trend vectors with the platform standard offset templates to establish read segment granularity-based offset anomaly scores; and completing raw signal layer analysis according to the offset anomaly scores.
[0007] In a possible implementation, the multi-generation sequencing data quality control management method further performs the following processing: calling a raw signal processing layer, reading sequence end base information at a read segment granularity, and loading an adapter template database by using the quality control rule template; constructing a base-Q value two-dimensional pattern graph according to the sequence end base information; performing similarity evaluation on the base-Q value two-dimensional pattern graph based on the adapter template database to generate adapter artifact anomaly scores; and completing raw signal layer analysis according to the adapter artifact anomaly scores and the offset anomaly scores.
[0008] In a possible implementation, the multi-generation sequencing data quality control management method further performs the following processing: calling an alignment mapping processing layer, reading alignment configuration parameters in the quality control rule template, selecting a reference genome and an alignment tool, performing standardized alignment on the sequencing data to generate an alignment result; extracting alignment positions, alignment scores, and alignment state identifiers at a read segment granularity according to the alignment result, establishing an association of read segment granularity feature extraction results, constructing an alignment mapping feature graph in combination with structural annotation information of the reference genome; performing position fidelity score analysis and mismatch pattern clustering identification analysis by using the alignment mapping feature graph, establishing an alignment quality index set, and completing alignment mapping layer analysis according to the alignment quality index set.
[0009] In a possible implementation, the multi-generation sequencing data quality control management method further performs the following processing: activating an annotated variation processing layer, calling the alignment result to perform variation detection, and establishing a variation candidate set; performing average base quality value calculation based on variation position range search on the variation candidate set to generate a first confidence score; performing annotated variation matching based on the quality control rule template on the variation candidate set, and establishing a second confidence score according to a matching result; and completing annotated variation layer analysis according to the first confidence score and the second confidence score.
[0010] In a possible implementation, the multi-generation sequencing data quality control management method further performs the following processing: activating a biological logic analysis layer, calling a population frequency database according to sample basic information of the sequencing data; performing variation sample distribution proportion analysis in the sequencing data by using the population frequency database, and establishing a proportion deviation value; and completing biological logic layer analysis according to the proportion deviation value.
[0011] In a possible implementation, the multi-generation sequencing data quality control management method further performs the following processing: activating a traceability analysis channel, inputting a trigger analysis result and a sample quality level identifier as input data into the traceability analysis channel; constructing an authentication question chain for each abnormal item, performing evidence query based on the authentication question chain, and establishing an authentication node; and completing traceability detection by using the authentication node, and establishing an abnormal influence path graph.
[0012] The application also provides a multi-generation sequencing data quality control management system, which comprises: a file header data reading module, configured to read file header data of sequencing data, and load a quality control rule template according to the file header data; a feature extraction result establishing module, configured to perform feature extraction on the sequencing data after import conversion, and establish a feature extraction result at a single read segment granularity; a quality control analysis module, configured to perform multi-layer nested quality control analysis according to the quality control rule template, the sequencing data and the feature extraction result, and establish a sample quality level identifier, wherein the multi-layer nested quality control analysis comprises original signal layer analysis, alignment mapping layer analysis, annotation variation layer analysis and biological logic layer analysis; a traceability detection module, configured to configure a trigger threshold, perform trigger analysis on the sample quality level identifier by using the trigger threshold, perform traceability detection based on a traceability analysis channel according to a trigger analysis result, and establish an abnormal influence path graph; and a quality control analysis result generating module, configured to generate a quality control analysis result after correcting the sample quality level identifier according to the abnormal influence path graph.
[0013] The application provides a multi-generation sequencing data quality control management method and system, which read file header data of sequencing data, load a quality control rule template, perform feature extraction on the sequencing data after import conversion at a single read segment granularity, establish a feature extraction result, perform multi-layer nested quality control analysis according to the quality control rule template, the sequencing data and the feature extraction result, establish a sample quality level identifier, configure a trigger threshold, perform trigger analysis on the sample quality level identifier, perform traceability detection, and establish an abnormal influence path graph. The technical problem of poor sequencing data quality control management quality and reliability due to single quality control dimension, poor dynamic adaptability and insufficient deep abnormality recognition capability in the prior art is solved, multi-dimensional accurate quality control is achieved, and the technical effect of improving sequencing data quality control management quality and reliability is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments of the present disclosure will be briefly introduced below. In the present application, flowcharts are used to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the foregoing or the following operations are not necessarily performed in sequence. On the contrary, various steps can be processed in reverse order or simultaneously according to needs. Meanwhile, other operations can be added to these processes, or one or more steps of operations can be removed from these processes.
[0015] Figure 1 A multi-generation sequencing data quality control management method flowchart is provided for the embodiments of the present application.
[0016] Figure 2 A multi-generation sequencing data quality control management system structure diagram is provided for the embodiments of the present application.
[0017] Legend: file header data reading module 10, feature extraction result establishing module 20, quality control analysis module 30, traceability detection module 40, quality control analysis result generating module 50. DETAILED DESCRIPTION
[0018] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application, the content of the specification can be implemented, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described.
[0019] In order to make the purposes, technical solutions and advantages of the present application more clear, the present application will be further described in detail below with reference to the drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without making creative labor are within the scope of protection of the present application.
[0020] In the following description, "some embodiments" are referred to, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict, and the term "first\second" referred to only distinguishes similar objects, and does not represent a specific order for the objects. The terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or modules not clearly listed or inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application.
[0021] The embodiments of the present application provide a multi-generation sequencing data quality control management method, as shown in the method, which comprises: Figure 1 Step S100, reading the file header data of the sequencing data, and loading the quality control rule template according to the file header data.
[0022] Preferably, in the sequencing data analysis process, the sequencing data is usually stored in a standard file format, and the header part of the file contains key metadata information. The file header data of the sequencing data is read and parsed, and a suitable quality control rule template is dynamically selected to ensure that the quality control process can adapt to different sequencing platforms, experimental types or data analysis requirements. The file header data of the sequencing data includes sequencing platform information, sequencing running parameters, sample information, and data generation time and instrument ID. Specifically, different platforms in the sequencing platform information need different quality control strategies; the sequencing running parameters such as read length, sequencing chemistry version and chip type affect the data quality evaluation standard, such as additional detection of chimeric reads for long read data; the sample information such as sample ID and library preparation method needs special quality control rules; the data generation time and instrument ID can be used to identify batch effects or systematic bias. The quality control rule template is loaded and adjusted according to the key fields of the header data, mainly including platform adaptation rules, experimental type rules and dynamic threshold adjustment, for example, Illumina data needs to check the Phred quality value distribution, and Nanopore data needs to pay attention to signal stability; whole genome sequencing pays attention to coverage uniformity, and RNA-seq needs to check strand specificity; tumor low-frequency mutation detection needs a more stringent alignment quality threshold, and population genetics can relax some standards. Further ensure that the multi-layer nested quality control can be based on the correct benchmark, thereby improving the accuracy and reliability of the overall quality control.
[0023] Step S200, after importing and converting the sequencing data, feature extraction is performed at single read granularity to establish a feature extraction result.
[0024] Preferably, importing and converting the sequencing data refers to reading the original sequencing data from a storage format and converting it into a standardized data structure, ensuring data integrity and format uniformity. Specifically, the sequencing data is parsed, the file header and main content are read, metadata such as sequencing platform and sample information are extracted, encoding conversion is performed, and the encoding method of base sequence and quality score data is ensured to meet the analysis standard. At the same time, data verification is performed, including checking file integrity and whether there are damaged reads. Then, the read is objectified, i.e. each read is converted into a data object containing uniform fields, including sequence, quality score, read ID, and alignment information, and then converted into a standardized data structure. Then, feature extraction is performed at single read granularity, wherein the sequencing data usually consists of millions or even billions of reads, each representing a DNA / RNA fragment being sequenced. Single read granularity analysis refers to analyzing each read and performing anomaly detection to find local problems, such as local low quality and alignment anomalies of a read. Then, feature extraction is performed for each read, including but not limited to original signal layer features such as base quality distribution, signal intensity anomaly, and low complexity sequence; alignment mapping layer features such as alignment position information, insertions / deletions, and soft clipping proportion; sequence composition features such as GC content deviation and k-mer frequency anomaly; structural anomaly features such as chimeric reads and read directionality. Finally, the extracted features of each read are structured and stored to form a feature extraction result, which may include a feature matrix, each row representing a read and each column representing a feature; anomaly labels, labeling potential problems for each read; and metadata association with file header information such as sequencing platform and sample ID, thereby significantly improving the accuracy and reliability of sequencing data quality control.
[0025] Step S300, according to the quality control rule template, the sequencing data, and the feature extraction result, multi-layer nested quality control analysis is performed to establish a sample quality level identifier, wherein the multi-layer nested quality control analysis includes original signal layer analysis, alignment mapping layer analysis, annotation variation layer analysis, and biological logic layer analysis.
[0026] Preferably, the multi-layer nested quality control analysis is performed according to the quality control rule template, the sequencing data and the feature extraction result, i.e. the data quality is checked step by step from different levels to finally form a quality grade identification of the sample, such as qualified, suspicious and unqualified, so as to ensure the accuracy and reliability of the data, wherein the multi-layer nested quality control analysis includes raw signal layer analysis, alignment mapping layer analysis, annotation variation layer analysis and biological logic layer analysis. Specifically, the raw signal layer analysis refers to evaluating the quality of the raw signal generated by the sequencer, identifying technical noise, including analyzing base quality distribution, checking Phred quality score and removing low-quality reads, for the FASTQ (base sequence and quality score) and PacBio / Nanopore current signal data.
[0027] Preferably, the alignment mapping layer analysis refers to evaluating the accuracy of the read alignment to the reference genome, identifying abnormal alignment, including analyzing alignment rate, the proportion of reads successfully aligned to the reference genome, low alignment rate may indicate contamination or mismatch of the reference genome, alignment quality, high alignment quality value represents high alignment reliability, low alignment quality value may indicate repeat region or alignment error, insert size distribution, checking whether it meets the library building expectation, soft clipping proportion, too much unaligned part at both ends of the read may indicate structural variation or sequencing error, strand specificity, checking whether it meets the expected direction of strand-specific library, and further marking abnormal alignment reads and calculating effective alignment rate.
[0028] Preferably, the annotation variation layer analysis refers to evaluating whether the detected variation is reliable and excluding technical false positives, including analyzing variation quality, filtering low-confidence variations, allele frequency, checking whether it meets the expectation, such as the allele frequency of somatic mutation of tumor samples is usually low, sequencing depth, ensuring sufficient coverage of the variation site, strand bias, whether the variation only occurs in a single strand, known database alignment, excluding common polymorphic sites, and further marking high-confidence variations and calculating the false positive rate of variation detection.
[0029] Preferably, the biological logic layer analysis refers to verifying data from the perspective of biological rationality, including Mendelian inheritance consistency, checking whether the offspring mutation is consistent with the parent genetic law, tumor-normal paired samples, whether the somatic mutation does not exist in the normal sample, gene expression correlation, checking whether the expression pattern between samples is consistent with the expectation, known biological pathways, such as whether the cancer driver gene mutation is within the expected range, and further marking abnormalities that do not conform to biological logic, such as samples with homozygous mutations that neither parent has, which may be contaminated. Finally, based on the results of each layer of analysis, the sample is comprehensively scored and graded, for example, qualified, all levels meet the quality control standards; suspicious, part of the indicators are abnormal, such as slightly low alignment rate; unqualified, key indicators deviate seriously or there are a large number of mutations that do not conform to biological laws. Through layer-by-layer progression and cross-verification, it is ensured that the sequencing data is reliable from the technical level to the biological level, and finally the accurate quality grade mark is output.
[0030] Further, step S300 further comprises step S310 of calling the original signal processing layer to read the base quality sequence in the feature extraction result and the platform standard offset template; step S320 of performing first-order difference on the base quality sequence to establish a quality change sequence; step S330 of calculating a local offset trend vector of the quality change sequence based on a sliding window; step S340 of comparing the local offset trend vector with the platform standard offset template to establish a read segment granularity offset anomaly score; and step S350 of completing the original signal layer analysis according to the offset anomaly score.
[0031] Preferably, the standardized operation of the sequencer should maintain stable base recognition quality, and if the quality change pattern of a certain sequence deviates from the platform expectation, it may indicate an abnormality, and the raw signal processing layer is called to read the base quality sequence in the feature extraction result and the platform standard offset template, wherein the base quality sequence refers to the Phred quality score of each position of a single read, and the platform standard offset template refers to the pre-defined quality change pattern of the sequencing platform under normal operation, including the expected quality change curve and the allowed fluctuation range of different sequencing cycles. Specifically, the base quality sequence is first-order differentiated to obtain a quality change sequence by calculating the difference between adjacent positions to amplify the mutation points of the quality value and facilitate the detection of abnormal fluctuations; then, a local offset trend vector of the quality change sequence is calculated based on a sliding window, that is, a fixed length (such as 5 bp) window is used to traverse the difference sequence, and statistics in the window are calculated, including the mean of the difference in the window, reflecting the local quality change direction, and the standard deviation of the difference in the window, reflecting the local fluctuation intensity, and then a local offset trend vector is obtained, representing the local pattern of quality change, wherein the local offset trend vector is a fixed length vector. The local offset trend vector is compared with the platform standard offset template, the deviation degree is calculated based on dynamic time warping or Euclidean distance, and then an offset anomaly score of a single read is established, for example, if a read has a sharp quality drop at positions 50-60 bp, the anomaly score of this region will be significantly increased; finally, the raw signal layer analysis is completed according to the offset anomaly score, including outputting the single read offset anomaly score to mark the abnormal read and outputting the global anomaly report to count the abnormal patterns of all reads.
[0032] Further, step S350 further includes step S351, calling the raw signal processing layer to read the sequence end base information of a single read, and loading the adapter template database using the quality control rule template; step S352, constructing a base-Q value two-dimensional pattern graph according to the sequence end base information; step S353, performing similarity evaluation on the base-Q value two-dimensional pattern graph based on the adapter template database to generate an adapter artifact anomaly score; and step S354, completing raw signal layer analysis according to the adapter artifact anomaly score and the offset anomaly score.
[0033] Preferably, the original signal layer analysis process focuses on detecting and identifying data quality abnormalities caused by library adapter contamination or sequencer end artifacts in the sequencing data by combining single read end base pattern analysis and adapter database comparison. Specifically, the original signal processing layer is called to read the sequence end base information of single read granularity, i.e. to extract the base sequence and corresponding Phred quality score (Q value) of a fixed length at the end of each read, and to load the adapter template database from the quality control rule template, which contains known adapter sequences and corresponding expected Q values. Then the frequency of each base at the end position is counted, the mean value curve of the end Q value is drawn, and the base composition of the end sequence is associated with the quality score to form a patternable two-dimensional feature, i.e. a base-Q value two-dimensional pattern graph. For example, for a normal read end, the bases are randomly distributed and the Q value decreases smoothly; for an adapter contaminated end, a specific base appears frequently and the Q value drops sharply.
[0034] Preferably, the base-Q value two-dimensional pattern graph is evaluated based on the adapter template database, i.e. the read end pattern is compared with the templates in the adapter database in multiple dimensions, including calculating the coincidence degree of the end sequence with the known adapter using k-mer matching, and checking whether the Q value suddenly decreases in the adapter matching area to identify the Q value abnormal pattern. Then the sequence matching degree and Q value drop amplitude are weighted to obtain the adapter artifact abnormal score. If the adapter artifact abnormal score is greater than a preset threshold (such as 80 points), the read is marked as adapter contaminated. Finally, the original signal layer analysis is completed by combining the adapter artifact abnormal score and the offset abnormal score. Specifically, if both are normal, the read is marked as high quality; if only the offset is abnormal, it may be a temporary fault of the sequencer; if only the adapter is abnormal, it may indicate a library problem, such as incomplete removal of the adapter; if both are abnormal, it may be a serious technical problem, such as interruption of sequencing causing adapter sequence to mix into the data sequence; and finally the abnormal score is output for sample quality level correction to ensure high quality data quality control management.
[0035] Further, step S300 further includes step S360 of calling the alignment mapping processing layer, reading the alignment configuration parameters in the quality control rule template, selecting a reference genome and an alignment tool, and performing standardized alignment on the sequencing data to generate an alignment result; step S370 of extracting the alignment position, alignment score and alignment state identifier of single read granularity from the alignment result, establishing the association of single read granularity feature extraction results, combining the structural annotation information of the reference genome, and constructing an alignment mapping feature map; and step S380 of performing position fidelity score analysis and mismatch pattern clustering identification analysis using the alignment mapping feature map, establishing an alignment quality index set, and completing alignment mapping layer analysis according to the alignment quality index set.
[0036] Preferably, alignment mapping layer analysis is the core of sequencing data quality control. By aligning the sequencing reads with the reference genome, the accuracy and reliability of read positioning are evaluated to identify potential alignment errors, contamination or technical bias. Specifically, the alignment mapping processing layer is called, the alignment configuration parameters in the quality control rule template are read, including the reference genome version, alignment tools and alignment algorithm parameters such as the number of allowed mismatches and gap penalties, and then the corresponding reference genome is dynamically loaded according to the experimental type, such as human whole genome, microbial sequencing, while the preset alignment tool such as BWA-MEM is called to perform standardized alignment. For long read data, a relaxed gap penalty is enabled, and for short read data, the number of mismatches is strictly limited, and then the alignment results, i.e. BAM / SAM files, are generated, containing the alignment position, CIGAR string, alignment quality score and other information of each read.
[0037] Preferably, according to the alignment results, the alignment position, alignment score and alignment state identifier of single read granularity are extracted, wherein the alignment position is the coordinate of the read on the reference genome; the alignment score refers to a score of 0-60, the higher the score, the more unique and reliable the alignment, and an alignment score of 0 usually indicates multiple alignment; the alignment state identifier includes unique alignment or multiple alignment and soft clipping proportion; and then the association of single read granularity feature extraction results is established, i.e. the alignment position is cross-analyzed with the gene structure coordinates using a genome annotation tool, and combined with the structural annotation information of the reference genome, such as whether the alignment region falls within the exon, intron or non-coding region of a gene, an alignment mapping feature map is constructed, wherein the alignment features of each read, such as alignment quality score, gene region and original signal layer feature Q value are associated, for example, a read with high alignment quality score but high soft clipping proportion may indicate a structural variation.
[0038] Preferably, based on the alignment mapping feature map, position fidelity score analysis and mismatch mode cluster identification analysis are performed. Specifically, the position fidelity score analysis refers to calculating the overall alignment quality score mean of the sample; calculating the unique alignment rate, i.e. the proportion of unique alignment reads to total alignment reads; chain specificity check, whether the alignment direction conforms to the library expectation, such as FR / RF direction of chain-specific library; wherein, low unique alignment rate may indicate reference genome mismatch or sample contamination. The mismatch mode cluster identification analysis refers to mismatch positioning, i.e. statistics of the distribution of non-matching bases in reads, if multiple reads at a certain genomic position all have mismatches, it may be a real variation or a reference genome error; insertion / deletion mode, i.e. detecting whether the length and position of Indel are concentrated in a specific region, such as microsatellite sequence; using machine learning such as DBSCAN to cluster mismatches / Indels, to distinguish random errors from technical bias, and then output an alignment quality index set, including unique alignment rate, average alignment quality score, and mismatch cluster significance. Finally, according to the alignment quality index set, the alignment mapping layer analysis is completed, i.e. high unique alignment rate, high average alignment quality score and no significant mismatch cluster, which indicates that it is qualified; slightly low unique alignment rate but mismatch mode conforms to known technical bias, which indicates that it is suspicious; low unique alignment rate or high frequency mismatch cluster, which indicates that it is unqualified. Further, fine error recognition is realized to distinguish sequencing errors from real biological phenomena, so as to comprehensively evaluate the positioning reliability of sequencing data.
[0039] Further, step S300 further includes step S390 of activating an annotated variant processing layer, calling the alignment result to perform variant detection, and establishing a variant candidate set; step S3100 of calculating the average base quality value based on the variant position range search based on the variant candidate set, to generate a first confidence score; step S3110 of performing annotated variant matching based on the quality control rule template on the variant candidate set, and establishing a second confidence score according to the matching result; and step S3120 of completing annotated variant layer analysis according to the first confidence score and the second confidence score.
[0040] Preferably, the annotated variant layer analysis is a key link in the quality control of sequencing data, through the dual verification of technical indicators and biological annotations, false positive variants are eliminated, the reliability of detected genomic variants is evaluated, and the accuracy of subsequent analysis is ensured. Specifically, the annotated variant processing layer is activated, and the variant detection tool is called to perform variant detection on the alignment result, i.e., the tool detects variants from the alignment result to generate an initial VCF file for storing genetic variant information, including basic information such as variant type, position, genotype, and quality score. Variants in low-complexity regions are removed, and a variant candidate set containing potential variant sites to be evaluated is generated. Then, based on the variant candidate set, the average base quality value is calculated based on the variant position range search to evaluate the reliability of the variant site from the sequencing technology perspective, i.e., for each variant site, the base quality values of all reads supporting the variant at the variant position are extracted and the average value is calculated. Then, the coverage depth and strand balance are checked, and then the first reliability score is output. The higher the score, the stronger the technical reliability.
[0041] Preferably, the annotated variant layer analysis is a key link in the quality control of sequencing data, through the dual verification of technical indicators and biological annotations, false positive variants are eliminated, the reliability of detected genomic variants is evaluated, and the accuracy of subsequent analysis is ensured. Specifically, the annotated variant processing layer is activated, and the variant detection tool is called to perform variant detection on the alignment result, i.e., the tool detects variants from the alignment result to generate an initial VCF file for storing genetic variant information, including basic information such as variant type, position, genotype, and quality score. Variants in low-complexity regions are removed, and a variant candidate set containing potential variant sites to be evaluated is generated. Then, based on the variant candidate set, the average base quality value is calculated based on the variant position range search to evaluate the reliability of the variant site from the sequencing technology perspective, i.e., for each variant site, the base quality values of all reads supporting the variant at the variant position are extracted and the average value is calculated. Then, the coverage depth and strand balance are checked, and then the first reliability score is output. The higher the score, the stronger the technical reliability.
[0042] Further, step S300 further includes step S3130 of activating a biological logic analysis layer, calling a population frequency database according to sample basic information of the sequencing data; step S3140 of performing variant sample distribution proportion analysis in the sequencing data using the population frequency database to establish a proportion deviation value; and step S3150 of completing biological logic layer analysis according to the proportion deviation value.
[0043] Preferably, the biological logic layer analysis is the final key link of sequencing data quality control. By comparing the distribution frequency of variations in the sample with the expected distribution of the known population database, possible technical false positives or biological abnormalities, such as sample contamination, family relationship errors, etc. are identified, and the credibility of the variation data is verified from the perspective of population genetics and biological rationality. Among them, the sample basic information of the sequencing data includes sequencing type and sample type, such as tumor / normal pairing, family members, specific ethnic groups, and the population frequency database includes public databases and local databases. The public database contains the thousand genome project, and the local database contains private frequency data of specific populations or diseases, such as the tumor mutation spectrum of the East Asian population. The population frequency database is used for variation sample distribution ratio analysis in the sequencing data, that is, for each variation site, the allele frequency of the corresponding population in the population database is queried, and the deviation of the sample variation frequency from the population frequency is calculated by Z-score standardization, that is, the ratio deviation value is obtained, and the frequency distribution pattern of all variations in the sample is counted. Finally, the biological logic layer analysis is completed according to the ratio deviation value, including identifying technical abnormalities and biological abnormalities through the ratio deviation value. When all variation frequencies meet the population expectation and there is no global distribution anomaly, the analysis passes; if some variations deviate, but can be explained by technical reasons, such as clonal selection of tumor samples, an alert is output; if there is a significant deviation and cannot be reasonably explained, such as the high-frequency occurrence of oncogenic mutations in healthy samples, the analysis fails, thereby ensuring the quality and reliability of sequencing data quality control management.
[0044] Step S400, configure a trigger threshold, use the trigger threshold for trigger analysis of sample quality level identification, perform provenance detection based on a provenance analysis channel according to the trigger analysis result, and establish an abnormal influence path graph.
[0045] Step S400 further includes step S410 of activating the provenance analysis channel and inputting the trigger analysis result and sample quality level identification as input data to the provenance analysis channel; step S420 of constructing an authentication question chain for each abnormal item and performing evidence query based on the authentication question chain to establish an authentication node; and step S430 of completing provenance detection through the authentication node to establish an abnormal influence path graph.
[0046] Preferably, according to the different quality control levels, multi-dimensional trigger thresholds are configured for automatic determination of sample quality. The trigger threshold can be automatically adapted according to the experimental type. The raw signal layer threshold includes that base percentage ≥80% is qualified and adapter contamination rate ≤5% is qualified. The alignment layer threshold includes that unique alignment rate ≥90% is qualified and average alignment quality ≥30 is qualified. The annotated variation layer threshold includes that high-confidence variation percentage ≥95% is qualified and the number of known pathogenic variations ≥1 driver mutation is expected. The biological logic layer threshold includes that population frequency deviation value ≤3 is qualified and family Mendelian conflict rate ≤1% is qualified.
[0047] Preferably, the trigger analysis of sample quality grade identification is performed using trigger thresholds, that is, the results of each layer of quality control analysis are compared with trigger thresholds to determine whether to trigger. Specifically, the output indicators of each layer are compared with the corresponding trigger thresholds to determine whether to trigger a warning or failure. For example, if the Q30 is 75%, the raw signal layer warning is triggered, and if the unique alignment rate is 85%, the alignment layer warning is triggered. The hierarchical weights are set according to experimental requirements, for example, clinical samples pay more attention to variation layer and biological logic layer, and the trigger analysis results are determined comprehensively. If all levels do not trigger the threshold or only the minor level deviates slightly, it is determined to be qualified; if 1-2 major levels trigger a warning but do not reach the failure threshold, it is determined to be suspicious; if any core level triggers a failure or multiple levels deviate seriously, it is determined to be unqualified.
[0048] Preferably, the traceability detection based on the traceability analysis channel is performed according to the trigger analysis results, wherein the traceability analysis channel is an intelligent root cause diagnosis unit in the sequencing data quality control process. Through structured problem chain reasoning and evidence integration, abnormalities found in the quality control process, such as low-quality samples and variation conflicts, are traced back to specific technical or biological roots, and a visual abnormality impact path diagram is generated to guide problem repair or data correction. Specifically, the traceability analysis channel is activated, and the trigger analysis results and sample quality grade identification are input as input data into the traceability analysis channel. An authentication problem chain is automatically generated for each abnormal item, and the root cause is locked through progressive questioning. For example, the example problem chain for low alignment rate may be whether the low alignment rate is concentrated in a specific chromosome region to check the uniformity of alignment distribution; whether the low alignment read has high adapter contamination to associate the raw signal layer result; whether the same batch of samples all have this problem to query the batch metadata.
[0049] Preferably, evidence query is further performed based on the authentication problem chain, that is, authentication problems are answered through multi-source data retrieval to form authentication nodes, that is, intermediate conclusions supported by evidence. The evidence types may include technical evidence such as abnormal alarms in the sequencer log and low concentration prompts in the library QC report; data evidence such as similar abnormal patterns of other samples in the same batch; and rule evidence such as common end degradation problems of a certain platform marked in the quality control rule template. Finally, the traceability detection is completed through the authentication nodes, that is, the authentication nodes are integrated to construct a directed acyclic graph, that is, an abnormality impact path diagram, which directly displays the abnormality propagation path. The abnormality impact path diagram includes root causes such as DNA degradation leading to insufficient capture of telomere regions, intermediate impacts such as low alignment rate leading to decreased variation detection sensitivity, and final consequences such as sample degradation to suspicious and reduced tumor mutation detection rate. Thus, discrete quality control abnormalities are converted into actionable attribution conclusions to improve the quality and reliability of sequencing data quality control management.
[0050] Step S500, after the sample quality level identification is corrected according to the abnormal influence path diagram, a quality control analysis result is generated.
[0051] Preferably, the sample quality level identification is intelligently corrected in combination with the dynamic feedback of the abnormal influence path diagram, that is, according to the explainability and severity of the abnormal root cause, it is determined whether to downgrade, maintain or upgrade the sample quality level, and the analysis results of each level, the abnormal path diagram and the correction basis are integrated to form a traceable quality control analysis result. Specifically, if the abnormality can be clearly attributed to a technical problem, such as adapter contamination or batch effect, and has limited impact on downstream analysis, it is adjusted to suspicious; if the abnormality involves biological contradictions, such as family conflicts or high-frequency driver mutations in normal samples, it is maintained as unqualified; abnormalities that cannot be automatically attributed, such as newly discovered sequencer error patterns, trigger a manual review process, and the report indicates that manual review is required; finally, a quality control analysis result is generated, including the final version of the sample quality identification, a summary of key abnormalities, an abnormal influence path diagram and recommended measures, such as re-constructing the library and increasing the DNA input for technical abnormalities, and checking the sample label or family relationship for biological abnormalities. Further reducing data waste, enhancing the credibility of the quality control analysis result, and ensuring the quality and reliability of the sequencing data quality control management.
[0052] In the foregoing, with reference to Figure 1 A multi-generation sequencing data quality control management method according to an embodiment of the present application is described in detail. Next, with reference to Figure 2 A multi-generation sequencing data quality control management system according to an embodiment of the present application will be described.
[0053] The multi-generation sequencing data quality control management system according to the embodiment of the present application is used to solve the technical problems of single dimension of sequencing data quality control, poor dynamic adaptability, insufficient deep abnormality recognition capability in the prior art, which leads to poor quality and reliability of sequencing data quality control management. It achieves multi-dimensional precise quality control and improves the quality and reliability of sequencing data quality control management. As shown in Figure 2 A multi-generation sequencing data quality control management system includes a file header data reading module 10, a feature extraction result establishing module 20, a quality control analysis module 30, a traceability detection module 40, and a quality control analysis result generating module 50.
[0054] The file header data reading module 10 is configured to read file header data of sequencing data, and load a quality control rule template according to the file header data; the feature extraction result establishing module 20 is configured to perform feature extraction on the sequencing data after import conversion, and establish a feature extraction result at a single read granularity; the quality control analysis module 30 is configured to perform multi-layer nested quality control analysis according to the quality control rule template, the sequencing data and the feature extraction result, establish a sample quality level identifier, and the multi-layer nested quality control analysis includes raw signal layer analysis, alignment mapping layer analysis, annotation variation layer analysis and biological logic layer analysis; the traceability detection module 40 is configured to configure a trigger threshold, perform trigger analysis on the sample quality level identifier by using the trigger threshold, perform traceability detection based on a traceability analysis channel according to a trigger analysis result, and establish an abnormal influence path diagram; and the quality control analysis result generating module 50 is configured to generate a quality control analysis result after correcting the sample quality level identifier according to the abnormal influence path diagram.
[0055] Next, the specific configuration of the quality control analysis module 30 will be described in detail. The quality control analysis module 30 further includes: calling a raw signal processing layer, reading base quality sequences in the feature extraction result and a platform standard offset template; performing first-order difference on the base quality sequences to establish quality change sequences; calculating local offset trend vectors of the quality change sequences based on a sliding window; comparing the local offset trend vectors with the platform standard offset template to establish offset anomaly scores at a single read granularity; and completing the raw signal layer analysis according to the offset anomaly scores.
[0056] Next, the specific configuration of the quality control analysis module 30 will be described in detail. The quality control analysis module 30 further includes: calling a raw signal processing layer, reading base quality sequences in the feature extraction result and a platform standard offset template; performing first-order difference on the base quality sequences to establish quality change sequences; calculating local offset trend vectors of the quality change sequences based on a sliding window; comparing the local offset trend vectors with the platform standard offset template to establish offset anomaly scores at a single read granularity; and completing the raw signal layer analysis according to the offset anomaly scores.
[0057] Next, the specific configuration of the quality control analysis module 30 will be described in detail. The quality control analysis module 30 further comprises: an alignment mapping processing layer is activated, alignment configuration parameters in the quality control rule template are read, a reference genome and an alignment tool are selected, the sequencing data is standardized aligned to generate an alignment result; alignment positions, alignment scores and alignment state identifiers of a single read granularity are extracted according to the alignment result, a single read granularity feature extraction result is associated, a structure annotation information of the reference genome is combined to construct an alignment mapping feature map; the alignment mapping feature map is used for position fidelity scoring analysis and mismatch mode clustering identification analysis to establish an alignment quality index set, and alignment mapping layer analysis is completed according to the alignment quality index set.
[0058] Next, the specific configuration of the quality control analysis module 30 will be described in detail. The quality control analysis module 30 further comprises: an annotation variation processing layer is activated, the alignment result is called to perform variation detection to establish a variation candidate set; average base quality value calculation based on variation position range search is performed on the variation candidate set to generate a first confidence score; annotation variation matching based on the quality control rule template is performed on the variation candidate set to establish a second confidence score according to the matching result; and annotation variation layer analysis is completed according to the first confidence score and the second confidence score.
[0059] Next, the specific configuration of the quality control analysis module 30 will be described in detail. The quality control analysis module 30 further comprises: a biological logic analysis layer is activated, a population frequency database is called according to sample basic information of the sequencing data; the population frequency database is used to perform variation sample distribution proportion analysis in the sequencing data to establish a proportion deviation value; and biological logic layer analysis is completed according to the proportion deviation value.
[0060] Next, the specific configuration of the traceability detection module 40 will be described in detail. The traceability detection module 40 further comprises: a traceability analysis channel is activated, the trigger analysis result and sample quality level identifier are input as input data to the traceability analysis channel; an authentication question chain is constructed for each abnormal item, evidence query is performed based on the authentication question chain to establish an authentication node; traceability detection is completed through the authentication node to establish an abnormal influence path graph.
[0061] The multi-generation sequencing data quality control management system provided in the embodiment of the application can execute the multi-generation sequencing data quality control management method provided in any embodiment of the application, has the corresponding function modules and beneficial effects of the execution method.
[0062] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or the server, the various units and modules included are only divided according to the functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific name of each functional unit is only for the convenience of mutual differentiation, and is not used to limit the protection scope of the present application.
[0063] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for quality control management of multi-generation sequencing data, characterized in that, The method comprises: reading file header data of the sequencing data, loading quality control rule templates according to the file header data; after importing and converting the sequencing data, performing feature extraction at a single read granularity to establish feature extraction results; performing multi-layer nested quality control analysis according to the quality control rule templates, the sequencing data and the feature extraction results, establishing sample quality grade identification, wherein the multi-layer nested quality control analysis comprises original signal layer analysis, alignment mapping layer analysis, annotation variation layer analysis and biological logic layer analysis; configuring a trigger threshold, performing trigger analysis of the sample quality grade identification by using the trigger threshold, performing traceability detection based on a traceability analysis channel according to the trigger analysis result, and establishing an abnormal influence path diagram; after correcting the sample quality grade identification according to the abnormal influence path diagram, generating a quality control analysis result.
2. The method of claim 1, wherein the method further comprises: The multi-layer nested quality control analysis according to the quality control rule templates, the sequencing data and the feature extraction results comprises: calling an original signal processing layer, reading base quality sequences and platform standard offset templates in the feature extraction results; performing first-order difference on the base quality sequences to establish quality change sequences; calculating local offset trend vectors of the quality change sequences based on a sliding window; aligning the local offset trend vectors with the platform standard offset templates to establish offset anomaly scores at a single read granularity; completing the original signal layer analysis according to the offset anomaly scores.
3. The method of claim 2, wherein the method further comprises: The original signal layer analysis according to the offset anomaly scores comprises: calling an original signal processing layer, reading sequence end base information at a single read granularity, and loading a linker template database by using the quality control rule templates; constructing a base-Q value two-dimensional pattern diagram according to the sequence end base information; performing similarity evaluation on the base-Q value two-dimensional pattern diagram based on the linker template database to generate a linker artifact anomaly score; completing the original signal layer analysis according to the linker artifact anomaly score and the offset anomaly score.
4. The method of claim 1, wherein the method further comprises: The multi-layer nested quality control analysis according to the quality control rule templates, the sequencing data and the feature extraction results further comprises: calling an alignment mapping processing layer, reading alignment configuration parameters in the quality control rule templates, selecting a reference genome and an alignment tool, performing standardized alignment on the sequencing data to generate an alignment result; extracting alignment positions, alignment scores and alignment state identification at a single read granularity according to the alignment result, establishing association of single read granularity feature extraction results, and constructing an alignment mapping feature map in combination with structural annotation information of the reference genome; performing position fidelity score analysis and mismatch pattern clustering identification analysis by using the alignment mapping feature map to establish an alignment quality index set, and completing the alignment mapping layer analysis according to the alignment quality index set.
5. The method of claim 4, wherein the method further comprises: The multi-layer nested quality control analysis according to the quality control rule templates, the sequencing data and the feature extraction results further comprises: activating an annotation variation processing layer, calling the alignment result to perform variation detection, and establishing a variation candidate set; performing average base quality value calculation based on variation position range search according to the variation candidate set to generate a first confidence score; The variant candidate set is subjected to quality control rule template-based annotation variant matching, and a second confidence score is established according to the matching result; According to the first confidence score, the second confidence score, the annotation variant layer analysis is completed.
6. The method of claim 5, wherein the method further comprises: According to the quality control rule template, the sequencing data, and the feature extraction result, the multi-layer nested quality control analysis is further included: Activate the biological logic analysis layer, and call a population frequency database according to the sample basic information of the sequencing data; The proportion deviation value is established by using the population frequency database to analyze the distribution proportion of the variant sample in the sequencing data. According to the proportion deviation value, the biological logic layer analysis is completed.
7. The multi-generation sequencing data quality control management method according to claim 1, characterized in that: According to the trigger analysis result, the traceability detection based on the traceability analysis channel is performed, and an abnormal influence path graph is established, including: Activate the traceability analysis channel, and input the trigger analysis result and sample quality level identifier as input data into the traceability analysis channel; An authentication question chain is constructed for each abnormal item, evidence is queried based on the authentication question chain, and an authentication node is established; The traceability detection is completed through the authentication node, and the abnormal influence path graph is established.
8. A multi-generation sequencing data quality control management system, characterized in that, The system is used to implement the multi-generation sequencing data quality control management method of any one of claims 1 to 7, and the system includes: A file header data reading module is configured to read file header data of sequencing data, and load a quality control rule template according to the file header data; A feature extraction result establishing module is configured to perform feature extraction under a single read segment granularity after importing and converting the sequencing data, and establish a feature extraction result; A quality control analysis module is configured to perform multi-layer nested quality control analysis according to the quality control rule template, the sequencing data, and the feature extraction result, and establish a sample quality level identifier, wherein the multi-layer nested quality control analysis includes original signal layer analysis, alignment mapping layer analysis, annotation variant layer analysis, and biological logic layer analysis; A traceability detection module is configured to configure a trigger threshold, perform trigger analysis of the sample quality level identifier by using the trigger threshold, perform traceability detection based on a traceability analysis channel according to a trigger analysis result, and establish an abnormal influence path graph; A quality control analysis result generating module is configured to generate a quality control analysis result after correcting the sample quality level identifier according to the abnormal influence path graph.
Citation Information
Patent Citations
Test data management and analysis system
CN119091967A
High-efficiency high-throughput gene sequencing data processing system
CN119446266A
Gene sequence variation detection method based on robust statistics and multi-strategy fusion
CN119889432A
Accurate comparison and validation of single nucleotide variants
US20130245958A1
Cited By
Intelligent identification method for transgenic crops based on high-throughput sequencing
CN121544285A
Artificial intelligence assisted trace sample pretreatment data traceability quality control system and artificial intelligence assisted trace sample pretreatment data traceability quality control method
CN122241123A