Computer-implemented method for quality control processing of high-throughput sequenced genomic data

A computer-implemented method for high-throughput genomic data analysis calculates GQ scores and assesses quality compliance using visualizations, addressing precision issues in existing methods to enhance reliability and efficiency.

WO2026074104A1PCT designated stage Publication Date: 2026-04-09INST NAT DE LA SANTE & DE LA RECHERCHE MEDICALE (INSERM) +2
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-02
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Current methods for ensuring quality and managing risks in high-throughput genomic data analysis lack precision, leading to extended analysis times and potential errors due to suboptimal data quality.

Method used

A computer-implemented method that calculates genotype quality (GQ) score values, computes GQ mean-like and variance-like scores, and assesses data quality compliance using two-dimensional visualizations to ensure accurate risk control and quality monitoring.

Benefits of technology

Ensures reliable genomic data quality by minimizing errors and improving analysis efficiency through precise risk control and quality monitoring, particularly in clinical diagnostics and personalized medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000018_0001
    Figure IMGF000018_0001
  • Figure IMGF000025_0001
    Figure IMGF000025_0001
  • Figure IMGF000026_0001
    Figure IMGF000026_0001
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for processing a set of sequenced genomic data, which is based on the use of so-called genotype quality (GQ) score values, which are obtained for each sample of said set of sequenced genomic data, the further computation of a GQ mean-like score and a GQ variance-like score for each sample, and the further determination from these values, whether or not the set of sequenced genomic data is quality compliant either at the level of the set of sequenced genomic data level, or at the level of at least one sample of the set, or both. The invention also relates to a computer-implemented method for determining threshold values for determining whether or not a set of sequenced genomic data is quality compliant, a computer-implemented method for assessing whether a threshold value for GQ mean-like score or a threshold value for GQ variance-like score is suited to evaluate the risk that sequenced genomic data is not quality compliant, a method for monitoring over time the quality of successive sets of sequenced genomic data, a data processing apparatus, computer program product and non-transitory computer-readable medium for carrying out the invention.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] COMPUTER-IMPLEMENTED METHOD FOR QUALITY CONTROL PROCESSING OF HIGH- THROUGHPUT SEQUENCED GENOMIC DATA

[0002] The present invention pertains to the field of bioinformatics and computational genomics in connection with high-throughput genome sequencing data, and more particularly concerns tools for quality control and risk management when high-throughput genome sequencing is implied.

[0003] The invention relates to a computer-implemented method for processing a set of sequenced genomic data, a computer-implemented method for determining threshold values for determining whether or not a set of sequenced genomic data is quality compliant, a computer-implemented method for assessing whether a threshold value is suited to evaluate the risk that sequenced genomic data is not quality compliant, a method for monitoring over time the quality of successive sets of sequenced genomic data, and a data processing apparatus, computer program product and non-transitory computer-readable medium to this end.

[0004] Advancements in genomic analysis have led to an exponential increase in the amount of data generated, particularly from high-throughput techniques. While this comes with clear advantages for the practice of medicine to the benefit of patients, or for exploitation of these results for environmental concerns. The sheer volume of data produced requires effective quality control measures to ensure the accuracy and reliability of the conclusions which can be made from these analyses.

[0005] Current methods for ensuring the quality and managing the risks of high-throughput data analysis may lack the precision necessary to detect suboptimal data quality, especially in due time, i.e., when needed. This deficiency can lead to extended analysis times and the possibility of errors. Consequently, there is a recognized need for more refined tools that can offer accurate assessments of risk control and quality monitoring.

[0006] Present invention arises from experiments devised to address this question. The invention is based on the exploitation of the genome's intrinsic characteristics and genetic variability. It offers an approach to the analysis and visualization of sequenced data, enabling precise risk control assessment and quality monitoring of high-throughput genome sequencing.

[0007] The invention is intended to enable genomics professionals, particularly in clinical diagnostics, can guarantee the reliability of high-throughput genome sequencing data, minimize errors, and improve the quality of results. This technology offers considerable advantages in fields such as personalized medicine, genetic research and other applications requiring risk control and rigorous monitoring of genomic data quality.

[0008] The invention relates to a computer-implemented method for processing a set of sequenced genomic data comprising:

[0009] Obtaining genotype quality (GQ) score values for each sample of said set of sequenced genomic data, and

[0010] Computing a variable indicative of a central tendency of the GQ score values of the variants thereof, GQ mean-like score, and a variable assessing the variability of the said GQ mean-like score values of the variants thereof, GQ variance-like score, for each sample of said set of sequenced genomic data, and Determining, from the values of the GQ mean-like score and the GQ variance-like score, whether or not the set of sequenced genomic data is quality compliant either at the level of the set of sequenced genomic data, or at the level of at least one sample of the set, or both.

[0011] By “sequenced genomic data” it is meant data that can be collected by a sequencing method, especially a so-called high-throughput sequencing (HTS) method, which is a term used to state herein that the sequencing is implemented by a DNA or RNA sequencer. A sequencing method enables the determination of the sequence of a nucleic acid molecule. High throughput sequencing (HTS) is especially suited for genomic data, where the nucleic acid molecule(s) to be sequenced is(are) either a large DNA molecule or is constituted by a large batch of sequences, the latter of which can result from the fragmentation into smaller fragments of a whole genome or large nucleic acid to be sequenced. HTS includes so-called next-generation sequencing (NGS) methods, also termed second and third generation sequencing by contrast, for the purpose of the present description, to “first” generation sequencing. By first generation sequencing it is referred to the initial technologies developed for determining the nucleotide sequence of DNA, the most prominent method of which is Sanger sequencing, also known as chain termination or dideoxy sequencing in the 1970s. Next generation sequencing (NGS) methods, in contrast to first-generation sequencing (like Sanger sequencing), utilize the strategy that is termed “Massive Parallel Sequencing” (MPS), the latter of which allows, after fragmentation of the nucleic acid to be sequenced, for the simultaneous, concurrent, sequencing of millions of nucleic acid fragments. This parallelism significantly increases throughput, reduces the time required for sequencing, and lowers the cost per base. Basically, one main difference between second and third generation sequencing methods is the size of the reads, short reads (50-300 bp) for second generation sequencing methods, and long reads (up to hundreds of kilobases) for third generation sequencing methods. According to another, equivalent, definition of Next generation sequencing (NGS) methods for the purpose of the present description, these are methods in which MPS is implemented, in particular on a nucleic acid sequence which has been “fractioned” in smaller fragments for the purpose of the sequencing. The concept of sequencing a nucleic acid sequence in the context described above is to reconstruct the departure whole sequence through dedicated bioinformatic algorithms, starting from the detected fragments.

[0012] According to a particular embodiment, the sequenced genomic data which is processed in instant invention has been collected by a next-generation sequencing (NGS) method, carried out by a DNA or RNA sequencer, i.e., a method using MPS, notably carried out by a DNA or RNA sequencer.

[0013] Examples of NGS platforms / techniques or type of apparatuses on which MPS can be carried out are, in a non-limitative manner to illustrate the state of the art known to date:

[0014] Illumina (Solexa): sequencing by synthesis, using reversible dye terminators,

[0015] Roche 454 Pyrosequencing: sequencing by synthesis, detecting released pyrophosphate during nucleotide incorporation,

[0016] Ion Torrent Sequencing: sequencing by detecting hydrogen ions released during nucleotide incorporation,

[0017] SOLiD (Sequencing by Oligonucleotide Ligation and Detection): sequencing by ligation. PacBio (Pacific Biosciences): single-molecule real-time (SMRT) sequencing

[0018] Oxford Nanopore: sequencing by detecting changes in electrical conductivity as DNA passes through a nanopore. All these platforms / techniques or type of apparatuses, and others, can readily be used; they implement a next-generation sequencing (NGS) method per the definitions provided herein.

[0019] The applications / types of set of nucleic acids that can be processed by such platforms / techniques or type of apparatuses or in the context of present invention, are also multiple, numerous, and non limitative. They can be for example:

[0020] Whole-Genome Sequencing (WGS), which is the Sequencing the entire genome of an organism, which can assist clinical diagnostics or personalized medicine,

[0021] Whole-Exome Sequencing (WES), which is the sequencing of only the coding regions (exons) of a genome, bearing the focus on regions that are more likely to affect protein function and cause disease,

[0022] Epigenetic Studies carried out on particular nucleic acid molecules, where DNA modifications, such as methylation, which affect gene expression without altering the DNA sequence can be investigated,

[0023] Metagenomics studies, where genetic material recovered directly from environmental samples is sequenced, and

[0024] Generally, clinical diagnostics, where genetic mutations and variations associated with diseases can be identified.

[0025] All applications permitted by NGS can however benefit from present invention.

[0026] Use of Massive Parallel Sequencing, through the generation of a massive volume of fragments, comes with the advantage of providing deep coverage of the nucleic acid sequence(s) to be sequenced. This deep, also termed “high”, coverage of the nucleic acid sequence(s) to be sequenced is important for accurately identifying variants, particularly rare or low-frequency variants, which may be missed with a lower coverage. Of note, the minimal and maximal sizes of the fragments (not whole final, such as genome sequence, which is reconstructed from the fragments) that can be sequenced using High Throughput Sequencing (HTS) technologies varies depending on the specific sequencing platform and method used, but can be as small as 50 bp or 100 bp to 400 bp or 600 bp for second generation sequencers (“short-read” sequencing), or a few kilobases to over 50kb for third generation sequencers (“long-read” sequencing).

[0027] Basically, Massive Parallel Sequencing follows the following scheme:

[0028] 1 . Nucleic acid from a sample previously retrieved or recovered from an individual (an human, an animal, or a microorganism, for example, when the environment is sampled) is fragmented to produce fragments, short or long reads depending upon the technique used, and these fragments are sequenced,

[0029] 2. These fragments (or “reads”) are aligned to a reference sequence, notably a reference genome if relevant, to identify differences between the sequence of a fragment (a “read”) and the corresponding region of the reference sequence, and

[0030] 3. An analysis software is used to identify “variants” over said fragments (“reads”), based on the differences observed with respect to the reference sequence, and optionally classify these variants per categories of variants. This step is also commonly known as “variant call” or “variant calling” in the literature, and

[0031] 4. The software then “annotates” the identified variants according to its findings.

[0032] Therefore, Massive Parallel Sequencing necessarily comes with computer-implemented processing of the recovered data, while highlighting the presence of a biological information in the processed data, which is comprised in the “variants” information thus uncovered. The variants whose presence is thus identified can provide insights into genetic diversity, disease mechanisms, and evolutionary processes, depending on the sought purpose of the sequencing. In the context of NGS, "variants" refer to differences in the sequenced (fragments of) departure nucleic acid sequence sequence when compared to a reference nucleic acid sequence. Variants can be Single Nucleotide Variants (SNVs), Insertion and Deletion Variants (Indels), Copy Number Variants (CNVs), Structural Variants (SVs), Microsatellite Variants, etc... The annotation is proper to the considered sequences and purpose sought.

[0033] It also follows from the above that sequencing, notably NGS, comes with the provision of several parameters associated with the sequenced data, as provided by the instruments / apparatuses operating the sequencing or associated computers processing the data.

[0034] Accordingly, a “set of” sequenced genomic data processed in a computer-implemented method of the invention is the set of data obtained from a next-generation sequencing using the tools known in the art and described herein, used as an entry point for the subsequent analysis to be carried out. It can also be said that the said set of sequenced genomic data results from a “capture” sequencing operation, according to the terminology used in the field. A “capture” corresponds to a peculiar genomic region that is analyzed, or a peculiar set of genomic regions, since for example an “exome capture” is made on several genomic regions, which are non-contiguous within the genome of a patient. A “capture” as used in can comprise sequencing data of several samples as long as these data correspond to a same peculiar genomic region that is analyzed, or a same peculiar set of genomic regions. The data of one sample may correspond to the data of one patient. For instance, as shown in the graphical representation of Figures 2 or 5, one dot corresponds to the data of one sample and to data of one patient, and all the dots of the graphical representation pertain to one “capture”.

[0035] The set of sequenced genomic data encompasses at least the data enabling to characterize all the variants generated in the so-called capture sequencing operation. Notably, this set of data comprises at least data enabling the determination of the peculiar parameters which are the GQ mean-like scores and GQ variance-like scores as discussed hereafter, which, in line with their names encompassing the words “mean” or “variance”, can be determined over a plurality of well-known “GC score values”, also discussed hereafter. The set of sequenced genomic data encompasses data from a plurality of samples, which can pertain to a plurality of patients. In turn, according to the invention, GC score values and parameters which can be determined based on GC score values, e.g., GQ mean-like scores and GQ variance-like scores, are calculated for a particular sample and, when displayed over a graphical visualization, represent a value for a particular sample (unless they correspond to a reference value).

[0036] According to a particular embodiment, a capture sequencing operation is defined by the genomic region it covers (let it be one contiguous genomic region of a genome, or a set of distinct genomic regions of a genome, for example when a set of genes is sequenced) and the sequencing characteristics of the sequencing operation in itself (such as the sequencing depth, also known as read depth or depth of coverage, which refers to the number of times a specific base (nucleotide) in the nucleic acid sequence is read during the sequencing process. In other words, it's the average number of times a given position in the genome is sequenced).

[0037] According to particular embodiments, the sequenced genomic data is NGS generated and is selected among: Whole Genome Sequencing (WGS) data or Whole exome sequencing (WES) data. According to particular embodiments, combinable with other embodiments described herein, the sequenced genomic data is NGS generated and is selected among: clinical exome data and gene panel data.

[0038] According to particular embodiments, combinable with other embodiments described herein, the sequenced genomic data is from samples previously obtained from human, animal or environmental samples.

[0039] The manipulated data being processed by computers, it follows that it has to be handled in a form enabling said computers to obtain information from it. A common, standard format for handling MPS data including variants determination is the so-called Variant Call Format (VCF), a standardized text file format serving as a structured format for storing and sharing the variant information detected during sequencing. Each variant identified in a sequencing process can be recorded in a VCF file with detailed annotations. Typically, each entry in (line of) a VCF file represents a single variant and contains detailed information about its position in the genome, the type of variant, and additional annotations, as allowed by the standard. Of note, there are alternatives to VCF files for storing and exchanging genetic variant information, such as BCF (Binary Call Format) which is a binary version of the VCF format or the GVF (Genome Variation Format), or others, but VCF is a widely known, generally implemented, standard and conversion methods between formats are readily available and known by the skilled person in the art. While researchers and bioinformaticians may choose a format alternative to VCF based on their specific needs, a conversion to VCF is most likely to be available (and known and feasible) given the widespread use of the VCF format.

[0040] The method of the invention therefore requires in a first step that genotype quality (GQ) score values be obtained from the set of sequenced genomic data processed in the method.

[0041] By “obtained” it is meant “retrieved” or “gathered”, “fed to the computer-implemented method”, in particular as an entry (input entry) when the data is provided to a computer or alternatively “determined”, in particular “computed” or “calculated” when the data under its used form has be to calculated in situ. Manipulation or retrieval of the data can be carried out using a VCF file as an intermediate, or any file serving such purpose of handling data. According to a particular embodiment, genotype quality (GQ) score values are determined, i.e., calculated as part of the method steps of the invention, according to the knowledge readily accessible to the skilled person. In particular, genotype quality (GQ) score values are obtained for each sample of the set of sequenced genomic data. Of note, since massive parallel sequencing is at stake, genotype quality (GQ) score values are generated for variants at a peculiar sequenced position.

[0042] According to the invention, one genotype quality (GQ) score value is obtained for one (per) peculiar sample comprised in the set of sequenced genomic data processed in the method of the invention. Indeed, the set of sequenced genomic data is obtained from sequencing data, which derives from a particular instrument which can process samples pertaining to different individuals, yet unique and identifiable. In other words, the method of the invention is sample-centered, when the set of sequenced genomic data in which the considered sample is found, can encompass several (synonym for “a plurality of”) samples. Where a sample pertains to one patient, the method of the invention can be said to be patient-centered. Since the method of the invention enables visualizing values from several samples (and / or patients) as found within a larger batch of sequenced genomic data, the method of the invention enables processing and determining quality compliance either at the level of the whole set of sequenced genomic data, e.g., obtained from a particular device such as a particular sequencing apparatus from a particular sequencing center, or at the level of at least one sample of the set of sequenced genomic data, or both when both can be displayed. Accordingly, a genotype quality (GQ) score value is a value associated to a sample, and is not meant from its definition, to be a value aimed at defining the quality of the sequencing of a variant as made by a sequencer. There are two successive steps when sequencing is carried out on a sample: the first one is the determination of the bases of the nucleic acid sequence, and the second one is the assignation of a genotype corresponding to a particular variant, to the sequenced nucleic acid sequence. The quality of the outcome of these two steps can be appreciated separately. Present invention is concerned with the processing of data at the level of a sample or a whole set of sequenced genomic data, and not with the determination of the bases of a nucleic acid sequence that is being sequenced.

[0043] The Genotype Quality (GQ) score is a metric used across various tools in the field of genomics for assessing the confidence or accuracy of the assigned genotype for a particular variant in a given sample. This step of assignation of a genotype to a variant is commonly known as “genotype call” or “genotype calling” in the literature. The GQ score can thus be used to assess the reliability of the variant calls which have been made. This reliability can be influenced by various factors such as sequencing depth, quality of the reads, and the complexity of the genomic region. However, the GQ score cannot be used to evaluate the quality of the sequencing carried out over the whole set of sequenced genomic data pertaining to a single sample or experiment, as the present invention does.

[0044] The GQ score can be calculated using different tools in the field of genomics according to well documented definitions. Whatever definition is used the GQ score is meant to represent the confidence in the genotype call which has been made for a variant, i.e., the confidence that the assigned phenotype for a given sample at a specific variant position, is correct. For instance, the GQ score is a standard field in VCF files. Since GATK (Genome Analysis Toolkit) is by far the most widely used tools for variant discovery and genotyping, a thorough description is provided herein by reference to this particular toolkit. However, other equivalent tools will also be discussed, for which the invention applies mutatis mutandis, as long as a GQ score is calculated.

[0045] In GATK (Genome Analysis Toolkit), one of the most widely used tools for variant discovery and genotyping, the GQ value represents the Phred-scaled confidence that the genotype assignment (GT) that has been made is correct. The GQ value is derived from the genotype likelihoods, termed PLs in the GATK platform (PL are "Normalized" Phred-scaled likelihoods of the possible genotypes). Specifically, the GQ is the difference between the PL of the second most likely genotype observed at a site in the observed sequence data, and the PL of the most likely genotype observed at the same site in the observed sequence data. PLs values are calculated by the HaplotypeCaller and GenotypeGVCFs tools of the Genome Analysis Toolkit (gatk) software. The values of the PLs are normalized so that the most likely PL is always 0, so the GQ ends up being equal to the second smallest PL, unless that PL is greater than 99. In GATK, the value of GQ is capped at 99 because larger values are not more informative, but they take more space in the file. So if the second most likely PL is greater than 99, the assigned GQ value is still 99. Basically, a GQ value gives the difference between the likelihoods of the two most likely genotypes. If it is low, the signification is that there is not much confidence in the genotype, i.e. there was not enough evidence to confidently choose one genotype over another. Reference is made to GATK documentation regarding how to precisely calculate a PL and a GQ value using GATK, as available through the reference articles cited herein or the technical documentation readily available to the skilled person online https: / / gatk.broadinstitute.org / hc / en- us. In any event, the Genotype Quality (GQ) score associated with each site in the observed sequence data, i.e., each variant as determined a posteriori, is provided in a VCF file, as a metric used across various tools in the field of genomics for assessing the confidence of genotype calls. It is not unique to GATK and included in the calling tools that work with MPS data because the inclusion of GQ scores in VCF files, notably, allows for consistent quality assessment across different datasets and tools, facilitating reliable genetic analysis and interpretation.

[0046] The genotype quality score (GQ) calculated by HaplotypeCaller, a tool in the Genome Analysis Toolkit (GATK), is one common way to assess the confidence in a called genotype. As seen above, it is based on concept of using the difference between the Phred-scaled likelihoods of the most likely and the second most likely genotype. Tools that use the same concept as HaplotypeCaller for GQ score calculation, i.e., basically involve the formula GQ=PL[1]-PL[0]) where PL represents the Phred-scaled likelihoods of the possible genotypes, and PL[0] is the likelihood of the most likely genotype, while PL

[0001] is the likelihood of the second most likely genotype, include, for example, Samtools (mpileup), bcftools (part of Samtools suite), Platypus, VarDict, which also calculates quality scores based on the probability of observing the allele given the sequencing data.

[0047] Other tools use the concept of posterior probabilities of each genotype given the observed data and then convert these probabilities into Phred-scale quality scores; examples of such tools include FreeBayes and DeepVariant. FreeBayes uses a Bayesian model to estimate the posterior probability of each genotype. In FreeBayes, GQ=-10*log10(1 -P(most likely genotype)) where P(most likely genotype)is the posterior probability of the most likely genotype. DeepVariant uses deep learning models to call variants and assign quality scores. The genotype quality scores are derived from the confidence of the neural network in its predictions. The quality scores are based on the probability outputs of the neural network, converted into Phred-scale scores. The exact formula used is proprietary but follows the principle given above for FreeBayes. In any event, the quality scores calculated by FreeBayes and DeepVariant reflect the confidence in the called genotype and therefore the information provided is similar to the information provided by a calculation of GQ scores using GATK.

[0048] According to a particular embodiment, a GQ score value is defined as the difference between a normalized Phred-scaled likelihood (PL) of the second most likely genotype of an allele observed at a site in the observed sequence data and a normalized Phred-scaled likelihood (PL) of the most likely genotype of an allele observed at the same site in the observed sequence data, as calculated by the HaplotypeCaller and GenotypeGVCFs tools of the Genome Analysis Toolkit (gatk) software. A normalized Phred-scaled likelihood (PL) in the Genome Analysis Toolkit (gatk) software is calculated according to the formula: PL = -10 * \log{P(Genotype | Data)} where P(Genotype | Data) is the conditional probability of the Genotype given the sequence Data that we have observed. Reference is made to the technical documentation of gatk for the process according to which P(Genotype | Data) is calculated (it requires the HaplotypeCaller module). Reference is also made to Li H. A statistical framework for SNP calling, mutation discovery, association mapping and population genetical parameter estimation from sequencing data. Bioinformatics. 201 1 Nov 1 ;27(21 ):2987-93. doi: 10.1093 / bioinformatics / btr509. Epub 201 1 Sep 8. PMID: 21903627; PMCID: PMC3198575 for theoretical considerations.

[0049] The method of the invention requires in a second step that two variables be computed for each sample of the set of sequenced genomic data that is analyzed: a variable indicative of a central tendency of the GQ score values of the variants thereof, also referred to as “GQ mean-like score" herein, and a variable assessing the variability of the said GQ mean-like score values of the variants thereof, also referred to as “GQ variance-like score" herein.

[0050] As an alternative wording, which has the same meaning as the wording stated above for the second step, it can be said that the method of the invention requires in a second step that two variables be computed for each sample of the set of sequenced genomic data that is analyzed: a “GQ mean-like score", which is a variable indicative of a central tendency of the GQ score values obtained for a particular site in the observed sequence data to be assigned to a variant, and a “GQ variance-like score", which is a variable assessing the variability of the said GQ mean-like score values obtained for a particular site in the observed sequence data to be assigned to a variant.

[0051] GQ mean-like scores and GQ variance-like scores are parameters determined on the basis of a plurality of genotype quality (GQ) score values, since, as their names indicate, they respectively reflect a “mean” over a set of values for a GQ mean-like score and reflect a “variance” over a set of values for a GQ variance-like score. The said set of values in the expression “over a set of values” is intended to be representative of a sample and / or a patient, in particular is a set of values gathered for a sample pertaining to a (single) patient. According to the invention, the GQ values, the GQ mean-like scores and GQ variancelike scores are parameters determined for each sample of the analyzed set of sequences genomic data, which can comprise data from several samples.

[0052] It is reminded that a “set of sequenced genomic data” processed in the computer-implemented method of the invention encompasses data enabling to characterize at least all the variants generated in a capture sequencing operation, per the definitions provided herein. There is one genotype assignment (GQ value) per assigned variant in the analyzed “set of sequenced genomic data”, and because sequencing is redundant, there are therefore several variants per sample, notably pertaining to one individual (patient).

[0053] According to a particular embodiment, a set of sequenced genomic data” encompasses data of several samples, which may pertain to a single patient if only one patient is used for providing data, but will more commonly pertain to a plurality of patients. According to the invention, the calculations made are sample-centered, in particular patient-centered, even if the sample or patient is found within a larger batch of sequenced genomic data, i.e. , the so-called “set of sequenced genomic data”.

[0054] Therefore, in the context of the present invention, "central tendency" refers to a statistical measure that identifies the center or typical value within a set of numerical data. Specifically, when applied to the genotype quality (GQ) score values of the variants in a sample of sequenced genomic data, the central tendency provides a summary value that represents the overall level of genotype quality for that sample.

[0055] The central tendency can be determined using various statistical approaches, including but not limited to: the arithmetic mean, which is the sum of all GQ score values for the variants in a sample divided by the number of variants, or the median, which is the middle value when all GQ score values are arranged in ascending or descending order, or the trimmed mean, which is calculated by removing a specified proportion of the highest and lowest GQ score values before computing the mean, thereby reducing the influence of outliers.

[0056] In the context of the invention, the variable indicative of central tendency — referred to as the GQ mean-like score — may be any of these measures, with the arithmetic mean being a particular embodiment. This value serves as a representative indicator of the overall confidence in genotype calls for a given sample, facilitating the assessment of sequencing quality and compliance with predetermined quality thresholds.

[0057] Accordingly, it is possible to calculate a measure of central tendency or average value for the GQ values observed for a sample within the set of sequenced genomic data, which can be a value selected among: the arithmetic mean, the median, a trimmed mean of the GQ values observed for all the variants of the observed subset of data, which can also be termed mean-like variable of the GQ score values of the variants thereof (i.e., the variants of the analyzed set of sequenced genomic data), abbreviated under the expression GQ mean-like score herein. According to (nonexclusive) particular embodiment, there is one sample per analyzed patient. A “sample” is a subset of the “set of sequenced genomic data”, which may in particular encompass variants for a sequenced region, corresponding to the sequencing of the genome (or portions thereof) of a same patient.

[0058] According to a particular embodiment, the variable indicative of a central tendency is the arithmetic mean of the GQ values observed for the sample.

[0059] According to a particular embodiment, the variable indicative of a central tendency is the median of the GQ values observed for the sample.

[0060] According to a particular embodiment, the variable indicative of a central tendency is the trimmed mean of the GQ values observed for the sample.

[0061] In the context of the present invention, "the variability of the said GQ mean-like score values" refers to a statistical measure that quantifies the degree of dispersion or spread among the GQ mean-like score values calculated for the variants within a sample of sequenced genomic data. This measure provides insight into how much the individual GQ scores deviate from the central tendency (such as the mean or median) within that sample.

[0062] The variability can be assessed using standard statistical metrics, including but not limited to: the variance, which is the average of the squared differences between each GQ score value and the mean of the GQ scores for the sample. Variance provides a measure of how far the GQ scores are spread out from their average value. the standard deviation, which is the square root of the variance and expresses the dispersion of GQ scores in the same units as the original data.

[0063] In the context of the invention, the variable assessing the variability — referred to as the GQ variance-like score — may be either the variance or the standard deviation of the GQ score values for the variants in a sample. This value serves as an indicator of the consistency or reliability of genotype quality assignment within the sample. A low variability suggests that the GQ scores are consistently high (or low), while a high variability indicates greater fluctuation in genotype quality assignment, which may be a sign of overall suboptimal sequencing quality or any technical issues affecting the data. As a result however, the quality how what is observed is compromised. By “a variable assessing the variability of the said GQ mean-like score values of the variants thereof”, abbreviated under the expression GQ variance-like score herein, it is meant a variable that measures the dispersion of the GQ mean-like score values defined above around their center, which can be a variable selected among: the variance of the GQ mean-like score values defined above, or the standard deviation of the GQ mean-like score values defined above.

[0064] According to a particular embodiment, the variable assessing the variability is the variance of the GQ mean-like score values defined above.

[0065] In a particular embodiment, the variable indicative of a central tendency is the arithmetic mean of the GQ values observed over the analyzed “set of sequenced genomic data”, and the variable assessing the variability discussed above is the variance of the arithmetic mean of the GQ values observed for the sample.

[0066] The method of the invention requires in a third step the determination of whether or not the set of sequenced genomic data is quality compliant either at the level of the set of sequenced genomic data, or at the level of at least one sample of the set, or both, based on the values of the GQ mean-like score and the GQ variance-like score determined in the preceding step. This determination enables to conclude either that the set of sequenced genomic data is quality compliant or that the set of sequenced genomic data is not quality compliant, either at the level of the set of sequenced genomic data, or at the level of at least one sample of the set, or both, for example by observing the values of the GQ mean-like score and the GQ variance-like score determined in the preceding step.

[0067] According to embodiment, the computer-implemented method for processing a set of sequenced genomic data of the invention is alternatively termed method for assessing quality compliance of a set of sequenced genomic data either at the level of the set of sequenced genomic data or at the level of at least one sample comprised in the set of sequenced genomic data, or both.

[0068] For assessing quality compliance of a set of sequenced genomic data, the skilled person will readily apprehend that the latter can be made by comparison of the observed values to so-called predetermined or reference values. When the assessment is made at the level of the set of sequenced genomic data, comparison is made to pre-determined or reference values representative of this level of data. When the assessment is made at the level of one sample comprised in the set of sequenced genomic data, comparison is made to pre-determined or reference values representative of this level of data. The skilled person will readily apprehend that observing values that are not representative of a situation deemed quality-compliant, has the meaning of presence of data of compromised quality. Therefore, by “observing the values of the GQ mean-like score and the GQ variance-like score determined in the preceding step” the skilled person will readily understand that such observation is inherently made within a frame where it can be appreciated the positioning of the observed values by reference to pre-determined or references values.

[0069] Accordingly, according to a particular embodiment, the method of the invention requires in a third step the determination of whether or not the set of sequenced genomic data is quality compliant either at the level of the set of sequenced genomic data, or at the level of at least one sample of the set, or both, based on the values of the GQ mean-like score and the GQ variance-like score determined in the preceding step, by reference to pre-determined or references values. Pre-determined or reference values can conventionally be determined by the skilled person according to his / her needs, given if required the guidance provided herein. It can be observed that pre-determined or reference values can be set differently depending on the purpose sought. Basically, for assessing the quality of a whole set of sequenced genomic data, it can be taken into account the plurality of values found in that set, notably each value being representative of an assayed sample and / or patient, and comparing the overall situation of this plurality of values with respect to pre-determined or reference values, e.g., intrinsically comparing values one by one with respect to pre-determined or reference values and / or comparing the distribution of the plurality of values found in that set, the latter of which is eased by the use of a two-dimensional visualization of the GQ mean-like scores and the GQ variance-like scores of the set of sequenced genomic data.

[0070] According to a particular embodiment, determining whether or not the set of sequenced genomic data is quality compliant, is carried out through a two-dimensional visualization of the GQ mean-like scores and the GQ variance-like scores of the set of sequenced genomic data, in particular by considering the form of the distribution with respect to predetermined threshold GQ mean-like score and GQ variance-like score values. Reference is made to the exemplary Figures 2 and 5, which depict two-dimensional plots of the observed mean and variance of the genotyping quality (GQ) scores. One dot corresponds to one sample (and one patient, in the provided illustration), a sample. One can observe that dots distribution makes the dots present mostly in the lower right quadrant of the diagram, which is the region where, based on appropriate threshold values for defining said lower right quadrant, quality compliant sequencing can be observed (per the rationale introduced in Figure 1 ).

[0071] Accordingly, by “a two-dimensional visualization of the GQ mean-like scores and the GQ variancelike scores of the set of sequenced genomic data” it is in particular meant a two-dimensional visualization of the GQ mean-like scores plotted against the GQ variance-like scores, or a two-dimensional visualization of the GQ variance-like scores plotted against the GQ mean-like scores. The skilled person will readily understand that it is not important which parameter is on the y axis and which parameter is on the x axis, both representations being suited for capturing the situation of the assayed whole set of sequenced genomic data, mutatis mutandis. According to a particular embodiment, the GQ variance-like scores are plotted against the GQ mean-like scores, i.e ., the GQ variance-like scores are on the y axis and the GQ mean-like scores are on the x axis. According to another particular embodiment, the GQ mean-like scores are plotted against the GQ variance-like scores, i.e., the GQ mean-like scores are on the y axis and the GQ variance-like scores are on the x axis.

[0072] According to a particular embodiment, comparing the distribution of the plurality of values found in the set of sequenced genomic data means that the skilled person can observe the dots distribution in the quadrants of a two-dimensional visualization as described above, where the regions of the quadrants of the two-dimensional visualization are defined by pre-determined, reference values, termed threshold values in that context - which can readily be determined by the skilled person as a result of his / her knowledge or purpose and guidance provided herein, or resulting from the intrinsic data of the sequenced set, or set of a visual representation (for example, le middle of the screen) - and the distribution of the plurality of values found in the sequenced set can be observed so that conclusions can be made. As an example, on Figure 1 , the majority of the values are found in the lower right quadrant of the diagram, which is the region where, based on appropriate threshold values for defining said lower right quadrant, quality compliant sequencing can be observed. The skilled person can readily adapt the conclusion on the basis of the plotted or displayed two-dimensional visualization and his / her knowledge, if required using the guidance provided herein. When determination of “threshold” values is concerned, the skilled person can readily apprehend the numerical value of the same depending on its purpose. For example, such a predetermined value can be set to a value generally observed for a batch of sequenced genomic data, for example generally obtained from a particular sequencing center, or for particular patients within such a center, or even generally pertaining to a set of peculiar sequencers, as appropriate. Therefore, the numerical value of such predetermined threshold value can vary without affecting the reproducibility of the method disclosed herein. To the contrary, the possibility to adjust said thresholds values depending on the purpose sought for concluding a quality compliance, is a further advantage of the present invention. In any event, the skilled person can readily determine a threshold value based on the batch of data that are processed and received from laboratories. Such threshold values can be determined depending on the purpose sought. The present disclosure also provides a specific method for determining threshold values, as claimed, which can be used in particular embodiments. It is generally made reference to this method for any threshold value which can be calculated, as an exemplary, non-limitative manner, of determining a threshold value. This remark applies throughout the present description.

[0073] In a specific embodiment, a predetermined threshold value for a GQ mean-like score is defined as the lower limit value of a 90% confidence interval calculated using a method selected among: the Clinical and Laboratory Standards Institute (CLSI) robust method, the Normal distribution method, a non- parametrical percentile method, on the GQ mean-like score values of a reference set of sequenced genomic data.

[0074] In a specific embodiment combinable with any embodiment herein, a predetermined threshold value for the GQ variance-like score is defined as the upper limit value of a 90% confidence interval calculated using a method selected among: the Clinical and Laboratory Standards Institute (CLSI) robust method, the Normal distribution method, a non-parametrical percentile method, on the GQ variance-like score values of a reference set of sequenced genomic data.

[0075] According to these embodiments, a reference set of sequenced genomic data, can be the whole sequenced set of data processed in a method of the invention as disclosed herein, which therefore enables, in that embodiment, the determination of predetermined thresholds by taking into account the whole of the data it contains.

[0076] According to a particular embodiment, it is concluded that the set of sequenced genomic data is quality compliant if the distribution of the GQ mean-like scores and the GQ variance-like scores of the set of sequenced genomic data in the lower right quadrant of a two-dimensional diagram of the GQ mean-like scores in abscissa and the GQ variance-like scores in ordinates, the lower right quadrant being defined by reference predetermined threshold values, is such as at least 51 %, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61 %, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71 %, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or any interval between any of these values, of the GQ mean-like scores and the GQ variance-like scores of the set of sequenced genomic data are in that quadrant.

[0077] It is to be understood, as it can be appreciated by the skilled person in the art, that quality compliance of the set of sequenced genomic data comprising the analyzed samples can visually be appreciated by a general appreciation of the closeness of all the dots defined by a GQ mean-like score and the GQ variance-like score pair from the quadratic line shown in a diagram such as shown in Figures 2 and 5 (exemplary diagrams), as well as the density of the distribution of those dots around that quadratic line, and the peculiar position of a GQ mean-like score and the GQ variance-like score pair with respect to horizontal and vertical lines virtually establishing boundaries (mathematically corresponding to thresholds) as far as that peculiar pair is concerned. For these reasons, the skilled person can readily appreciate that most of the dots of a quality compliant set of sequenced genomic data will be found in the lower right quadrant of a two-dimensional diagram as defined above. The skilled person can also readily appreciate that the threshold to be applied for concluding to a set of sequenced genomic data being quality compliant may depend on the purpose sought or quality level sought.

[0078] According to a particular embodiment, it is concluded that the set of sequenced genomic data is quality compliant if all the GQ mean-like scores and the GQ variance-like scores of the set of sequenced genomic data are in the lower right quadrant of the above-defined two-dimensional diagram of the GQ mean-like scores in abscissa and the GQ variance-like scores in ordinates, defined by predetermined threshold values.

[0079] Apart the fact that the said predetermined threshold values can be predetermined, the predetermined threshold values can be calculated as described in the following, when threshold values calculation is discussed.

[0080] According to a particular embodiment, determining whether or not a sample of the set of sequenced genomic data is quality compliant is carried out according to the following: for a sample of the set of sequenced genomic data, respectively comparing the GQ mean-like score and the GQ variance-like score to predetermined threshold values, and: if the GQ mean-like score of the sample is superior or equal to its predetermined threshold value and the GQ variance-like score is inferior or equal to its predetermined threshold value, classifying the corresponding sample (as sequence data of the set) as quality compliant, if the GC score mean is superior or equal to its predetermined threshold value and the GC score variance is superior to its predetermined threshold value, classifying the corresponding sample (as sequence data of the set) as not quality compliant, if the GQ mean-like score is inferior to its predetermined threshold value and the GQ variance-like score is superior to its predetermined threshold value, classifying the corresponding sample (as sequence data of the set) as not quality compliant, and if the GQ mean-like score is inferior to its predetermined threshold value and the GQ variance-like score is inferior or equal to its predetermined threshold value, classifying the corresponding sample (as sequence data of the set) as not quality compliant.

[0081] As stated above, a sample is, according to a particular embodiment, a subset of the sequence data of the analyzed set of data, in particular a single couple of GQ mean-like score and GQ variance-like score which can be plotted on a two-dimensional visualization as described above, or not, and which generally pertain to one individual (e.g., patient, or individual the genome of which has to be sequenced and analyzed). That situation is therefore different from the situation where it is the quality of the whole of the sequenced set of genomic data, i.e., the quality of a plurality of assayed samples as found in a single sequenced set of genomic data, which has been discussed here above in the context of a two-dimensional visualization for determining the quality of the whole set.

[0082] However, even without the explicit use of any two-dimensional visualization, it is simple for the skilled person to determine, according to the invention, whether or not a sample of the set of sequenced genomic data is quality compliant by comparing the couple of values constituted by the GQ mean-like score and the GQ variance-like score, with respective threshold values per the rules provided here-above and herein, in order to conclude about quality compliance. When determination of “threshold” values is concerned, the same remarks as above in the paragraph beginning by the words “When determination of “threshold” values” apply similarly.

[0083] The set of sequenced genomic data is the set generated in a sequencing operation, especially a single sequencing operation, which corresponds to a capture sequencing operation as defined herein, i.e. , a capture sequencing defined by the genomic region it covers (target genomic region(s)) and the sequencing characteristics of the sequencing operation in itself (parameters used for the specific sequencing carried out such as the sequencing depth etc...). The skilled person can readily apprehend the features which are characteristic of a sequencing operation when the latter is implemented, and is naturally aware that comparison between the results obtained from two sequencing methods, are meaningful when what is compared is similar, i.e., when the compared sequencing operations are at least similar or identical, according to a particular embodiment.

[0084] Therefore, the skilled person will appreciate without being told that for a meaningful conclusion regarding quality compliance of the analyzed set of sequenced genomic data, the predetermined threshold values referred to above should have been previously determined on the basis of at least similar parameters with respect to the sequencing carried out, between the sequencing used for determining threshold values and the sequencing used for analyzing the set of sequenced genomic data fed to a method of the invention, let it be regarding the genomic region the sequencing covers or the sequencing characteristics of the sequencing operation.

[0085] Figure 1 herein and comments associated with it provides theoretical explanations regarding the decision rules defined above. Figures 2 and 5 provide experimental illustrations.

[0086] The skilled person is able to appreciate the predetermined threshold values to be applied, depending upon the results sought, if needed with the assistance of the guidance provided herein.

[0087] As a particular example of predetermined threshold values, which can be applied, a determination that a sample is at risk to not be quality compliant (subject to the appreciation of compliance by the operator) can be made if the GQ mean-like score is inferior to a value selected among: 100, 99, 98, 97, 96, 95, 94, 93, 92, 91 , 90, 89, 88, 87, 86, 85, 84, 83, 82, 81 and 80, or any value in between any range defined using one of these values as upper or lower boundary, in particular is inferior to 89.4 and the GQ variance-like score is superior to a value selected among : 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540 and 550, or any value in between any range defined using one of these values as upper or lower boundary, in particular is superior to 468.

[0088] According to a more particular example, a determination that a sample is at risk to not be quality compliant can be made if the conditions set in the preceding paragraph are met and furthermore there are at least 2 other metrics outside of acceptable ranges, as determined by a method of the invention when the method further comprises a step of calculating for each sample of the set of sequenced genomic data a number of quality metrics, and their confidence intervals, as further described in detail later in the present description.

[0089] Of note, according to a particular embodiment, after it has been determined whether or not one or more samples of the set of sequenced genomic data is(are) quality compliant according to the above, it can also be determined whether or not the set of sequenced genomic data as a whole, or at least the part of the set of sequenced genomic data which comprises the samples, the quality of which has been determined, is quality compliant, based on the overall or sum of results for the individual samples analyzed. This can be done using a two-dimensional visualization where the GQ mean-like scores and the GQ variance-like scores of the analyzed samples are spotted, along with the applied corresponding thresholds or considering the obtained individual results (quality compliant or not).

[0090] According to a particular embodiment, it is concluded that the set of sequenced genomic data comprising the analyzed samples is quality compliant if the distribution of the GQ mean-like scores and the GQ variance-like scores of the analyzed samples of the set of sequenced genomic data in the lower right quadrant of a two-dimensional diagram of the GQ mean-like scores in abscissa and the GQ variance-like scores in ordinates, the lower right quadrant being defined by reference predetermined threshold values, is such as at least 51 %, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61 %, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71 %, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or any interval between any of these values, of the GQ mean-like scores and the GQ variance-like scores of the analyzed samples of the set of sequenced genomic data are in that quadrant.

[0091] According to a particular embodiment, it is concluded that the set of sequenced genomic data comprising the analyzed samples is quality compliant if all the GQ mean-like scores and the GQ variancelike scores of the analyzed samples of the set of sequenced genomic data are in the lower right quadrant of the above-defined two-dimensional diagram of the GQ mean-like scores in abscissa and the GQ variance-like scores in ordinates, defined by predetermined threshold values.

[0092] According to a particular embodiment, a GQ score value as used in any embodiment of a method of the invention as defined herein is the difference between a normalized Phred-scaled likelihood (PL) of the second most likely genotype of an allele observed at a site in the observed sequence data and a normalized Phred-scaled likelihood (PL) of the most likely genotype of an allele observed at the same site in the observed sequence data, as calculated by the HaplotypeCaller and GenotypeGVCFs tools of the Genome Analysis Toolkit (gatk) software. Accordingly, a GQ score value is concerned with assignation of a genotype to a particular variant, as commonly known in the field, the term variant being understood in that sentence as the sequence at a particular site in the observed sequence data, which can be assigned to a variant (among several probabilities).

[0093] According to a particular embodiment, the method of the invention is for determining whether or not a sample of the set of sequenced genomic data is quality compliant, and to this end, the method further comprises a step of calculating for each sample of the set of sequenced genomic data a number of quality metrics, and their confidence intervals, so that where a GQ mean-like score and / or a GQ variance-like score of the sample is equal to the predetermined threshold value to which it is compared, conclusion about quality compliance of the sample is made by taking into consideration the number of quality metrics outside their confidence intervals.

[0094] Exemplary metrics which can be used to this end are shown in Table 1 . They can be, for example, Total Aligned Reads Aligned (NR in millions), Read enrichment (%), Reads enrichment upstream and downstream of the target (%), Base enrichment (%), Database enrichment upstream and downstream of the target (%), Median fragment length, Minimum fragment length, Maximum fragment length, Average depth of genome covered by sequencing, Coverage uniformity (%), Target coverage at 20X (%), Target coverage at 50X (%), This list is not limitative but exemplary. According to a particular embodiment, , the method further comprises a step of calculating for each sample of the set of sequenced genomic data a number of quality metrics, and their confidence intervals, wherein the quality metrics is one or more quality metrics selected amongst: Total Aligned Reads Aligned (NR in millions), Read enrichment (%), Reads enrichment upstream and downstream of the target (%), Base enrichment (%), Database enrichment upstream and downstream of the target (%), Median fragment length, Minimum fragment length, Maximum fragment length, Average depth of genome covered by sequencing, Coverage uniformity (%), Target coverage at 20X (%), Target coverage at 50X (%).

[0095] According to a particular embodiment, 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 1 1 or 12 quality metrics are taken into consideration for assessing whether the sample is quality compliant. According to a particular embodiment, a range of quality metrics having as boundaries any one of the following numbers: 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 1 1 or 12 quality metrics is taken into consideration for assessing whether the sample is quality compliant. According to a specific embodiment 12 quality metrics are taken into consideration for assessing whether the sample is quality compliant, notably the 12 quality metrics as described in the preceding paragraph. Reference is also made to the description of a method for assessing whether a threshold value for a GQ mean-like score or a threshold value for a GQ variance-like score determined according to the above method is suited to evaluate the risk that sequenced genomic data is not quality compliant herein, which elaborates on how quality metrics can be used to this end. The skilled person can profitably use the guidance regarding the relationship between quality metrics and the fitness of the same and their observed values for concluding that a sample is quality compliant;

[0096] Conclusion about quality compliance of the sample can, in a particular embodiment, be made by taking into consideration the number of quality metrics outside their confidence intervals: according to a particular embodiment, the sample is quality compliant if 0, 1 , 2 or 3, notably less than 3 (2 included) quality metrics are within their respective confidence intervals. Conversely, a sample is not quality compliant if 3, 4, 5, 6, 7 or more, notably more than 3 or 4 quality metrics are outside their respective confidence intervals. Data commented in the experimental section in connection with Figures 3 and 4 provide insights regarding these boundaries.

[0097] According to a particular embodiment, use of quality metrics for assessing whether a sample is quality compliant is made where a GQ mean-like score of the sample and / or a GQ variance-like score of the sample is about equal or equal to its predetermined threshold value: in these cases, considering the number of quality metrics of the sample outside their confidence intervals may assist the conclusion. The same considerations as in the above paragraphs apply.

[0098] According to a particular embodiment, the predetermined threshold value for the GQ mean-like score is defined as the lower limit value of a 90% confidence interval calculated using a method selected among: the Clinical and Laboratory Standards Institute (CLSI) robust method, the Normal distribution method, a non-parametrical percentile method, on the GQ mean-like score values of a reference set of sequenced genomic data.

[0099] According to a particular embodiment, combinable with the former embodiment, the predetermined threshold value for the GQ variance-like score is defined as the upper limit value of a 90% confidence interval calculated using a method selected among: the Clinical and Laboratory Standards Institute (CLSI) robust method, the Normal distribution method, a non-parametrical percentile method, on the GQ variancelike score values of a reference set of sequenced genomic data. A reference set of sequenced genomic data is a set of sequenced genomic data, which has at least been sequenced with similar if not identical parameters with respect to the sequencing carried out, as for the set of sequenced genomic data analyzed with the method of the invention described herein. According to a particular embodiment, a reference set of sequenced genomic data is a reference “capture” set of sequenced genomic data according to the definitions provided herein, i.e. a set of data obtained in a capture sequencing operation defined by the genomic region it covers and the sequencing characteristics of the sequencing operation in itself, which can be used as a reference for the purpose of the present invention. A ’’reference” in this context can have the signification that the set of data is a set to which another set of data is to be compared, or is a kind of “ideal” set of data for the considered analysis to be carried out, to which another set of data is to be compared.

[0100] A reference set of sequenced genomic data may also have been sequenced starting from similar biological samples; the skilled person can readily determine that for a pertinent comparison, the compared data should have been obtained in similar conditions. The skilled person has freedom and capability to determine the same, for a meaningful comparison to a “reference”.

[0101] According to a particular embodiment, the reference set of sequenced genomic data is the set of sequenced genomic data analyzed in a method of the invention as described in any embodiment herein. According to a particular embodiment, as further detailed hereafter, the creation of a two-dimensional visualization of the samples analyzed in a series enables to visualize the distribution of means and variances of GQ scores per sample, and, through this visualization, identify samples having poor sequencing quality. In this sense, the "mean vs. variance GQ score diagram" approach enabled by the invention can be used calculate the mean and variance of the GQ score per sample, the reference values of which are not known. Reference is made to the developments herein with respect to the two-dimensional visualization of the GQ mean-like scores and the GQ variance-like scores of the set of sequenced genomic data with the method of the invention, also enabling the determination of appropriate threshold values, although these can also be pre-determined differently.

[0102] A 90% confidence interval is recommended by the CLSI (Clinical and Laboratory Standards Institute) Guidelines C28-A3. It can be calculated using the following 3 methods: (a) using the Normal distribution (b) using a non-parametrical percentile method, and (c) the robust method as described in the CLSI Guidelines C28-A3 available at https: / / www.medcalc.org / manual / referenceinterval.php.

[0103] In the Normal distribution method, the mean, variance, and standard deviation of the sample data are calculated. For a 2-sided reference interval, the limits of normality are :

[0104] Lower Limit = Mean — 1.96 x SD

[0105] Upper Limit = Mean + 1.96 x SD

[0106] The 90% confidence interval for each limit is given by (Bland, 2000):

[0107] Limit + 1.64

[0108] The normal distribution method requires that the data present a normal distribution, possibly after logarithmic or Box-Cox transformation. This method does not require a minimum number of subjects, but a minimum sample size of 40 is recommended (Le Boedic, 2019).

[0109] In the percentile method, the lower and upper limits of normality are given by the 2.5th and 97.5th percentiles for a double-sided reference interval. The 5th percentile for a left sided reference interval, and 95th percentile for a right sided reference interval. Following the CLSI guidelines, 90% confidence intervals are defined using the method of Reed et al. (1971 ). For the calculation of a 90% confidence interval in the percentile method, a minimal sample size of 120 is required (CLSI, 2008).

[0110] The calculation of a reference interval using the robust method (Horn & Pesce, 2005) involves an iterative process, in which the initial central value is estimated by the median and the initial spread by the median absolute deviation about the median (MAD). In the iterative process, actual observations are downweighted according to their distance from the central tendency of the sample. In each iteration, a quantity Tbi, representing the updated estimate of central tendency, is calculated, until the change in consecutive iterative values is negligible. Weighted estimators of variability and spread are calculated to establish the reference limits. See Horn & Pesce (2005) or CLSI (2008) for computational details. 90% confidence intervals for the reference limits are estimated using bootstrapping (percentile interval method, Efron & Tibshirani, 1993); The robust method can be used as an alternative to the percentile method when sample size is less than 120.

[0111] Follow CLSI guidelines for percentiles and their Cis: if the NCCLS and Clinical and Laboratory Standards Institute (CLSI) guidelines C28-A2 and C28-A3 are followed for estimating percentiles and their 90% confidence intervals, percentiles are calculated as the observations corresponding to rank r=p*(n+1 ). Also for the 90% confidence intervals of the reference limits the CLSI guidelines are followed and conservative confidence intervals are calculated using integer ranks (and therefore the confidence intervals are at least 90% wide). Otherwise, MedCalc calculates percentiles as the observations corresponding to rank r=p*n+0.5 (Lentner, 1982; Schoonjans et al., 201 1 ), and calculates a less conservative and more precise confidence interval using an iterative method.

[0112] As an add-on method to the method described above, the invention also relates to a computer- implemented method for determining threshold values for determining whether or not a set of sequenced genomic data is quality compliant, i.e., for determining threshold values appropriately suited for determining whether or not a set of sequenced genomic data is quality compliant, said method comprising steps of: Obtaining GQ score values for the set of sequenced genomic data,

[0113] For each sample of said set of sequenced genomic data, computing (i) a variable indicative of a central tendency of the GQ score values of the variants thereof, also termed GQ mean-like score and (ii) a variable assessing the variability of the said GQ mean-like score values of the variants thereof, also termed GQ variance-like score,

[0114] Determining the lower limit value of a 90% confidence interval calculated using a method selected among: the Clinical and Laboratory Standards Institute (CLSI) robust method, the Normal distribution method, a non-parametrical percentile method, on (all) the GQ mean-like score values of the said set of sequenced genomic data - it is to be understood that the 90% confidence interval is preferably, meaningfully, calculated considering all or most of the GQ mean-like score values of the said set of sequenced genomic data (erratic data may be removed without being detrimental to the end-result sought), and

[0115] Determining the upper limit value of a 90% confidence interval calculated using a method selected among: the Clinical and Laboratory Standards Institute (CLSI) robust method, the Normal distribution method, a non-parametrical percentile method, on (all) the GQ variance-like score values of the said set of sequenced genomic data - it is to be understood that the 90% confidence interval is preferably, meaningfully, calculated considering all or most of the GQ variance-like score values of the said set of sequenced genomic data (erratic data may be removed without being detrimental to the end-result sought),.

[0116] By “appropriately suited” it is meant, which are fitted to this end, in particular which are the best fitted to this end (optimized).

[0117] The definitions provided herein also applies, including in the more precise embodiments which can be also implemented in such a method for determining threshold values.

[0118] The invention also relates to a computer-implemented method for assessing whether a threshold value for a GQ mean-like score or a threshold value for a GQ variance-like score determined according to the above method is suited to evaluate the risk that sequenced genomic data is not quality compliant, said method comprising the steps of: a. For the set of sequenced genomic data to be analyzed, calculating a number of quality metrics, and their confidence intervals, in particular using the CLSI robust method, and b. Determining the number of quality metrics as calculated in step a. that is(are) outside their respective confidence intervals as calculated in step a., and c. Analyzing by logistic regression the association between the number of quality metrics outside their confidence intervals as determined in step b. and the value of the threshold value to be assessed, and d. If a discriminating score can be found on the basis of a sensitivity of at least 80% and a specificity of at least 80%, concluding that the threshold value for GQ mean-like score or the threshold value for GQ variance-like score is suited to evaluate the risk that sequenced genomic data is not quality compliant.

[0119] The metrics or “quality metrics” referred to above, are as described in the present text and Examples section when reference is made to “12 quality metrics”. All features described in connection with these “12 quality metrics” also apply in the context of the above method for assessing whether a threshold value for a GQ mean-like score or a threshold value for a GQ variance-like score is suited to evaluate the risk that sequenced genomic data is not quality compliant.

[0120] According to a particular embodiment, combinable with any embodiment described herein, the method of the invention comprises a step of sequencing the genomic data of the set of sequenced genomic data ahead of (prior to) carrying out any other method step as described herein, such as obtaining GQ score values or other actions or calculations. The sequencing can be carried out according to any method or embodiment described in the present invention or known be the skilled person, which is aware of sequencing methods per his / her knowledge, since the present invention comes in assistance to quality control and risk management when high-throughput genome sequencing is implied.

[0121] The invention also relates to a method for monitoring over time the quality of successive sets of sequenced genomic data, said method comprising the steps of: a. Carrying out at a first point in time on a set of sequenced genomic data a computer- implemented method for determining threshold values as described in any embodiment herein, which returns GQ mean-like score, GQ variance-like score threshold values, and optionally quality metrics as described in any embodiment herein, and b. Carrying out at a second point in time on a set of sequenced genomic data that is different from the set of sequenced genomic data of step a., a computer-implemented method for determining threshold values as described in any embodiment herein, which returns GQ mean-like score, GQ variance-like score threshold values, and optionally quality metrics as described in any embodiment herein, and c. Comparing the evolution, between the first and second points in time, of the GQ mean-like score, GQ variance-like score threshold values, and optionally quality metrics returned in steps a. and b, for example using the Levey Jennings method.

[0122] It will be understood that the sets of sequenced genomic data will be chosen by the skilled person so that they can be compared (in a manner enabling a meaningful conclusion, e.g., same type of data, same type of sequencing conditions, etc...) in the above method.

[0123] A Levey-Jennings (Levey, 1950), also called Shewhart (Shewhart, 1931 ), plot, which is a common presentation for evaluating Quality Control results. This format shows each Quality Control result sequentially over time and allows a quick visual assessment of method performance, including trend detection. Assuming the method is performing in a stable condition consistent with its specifications, ot can depict mean value representing the target (or expected) value for the result, and SD lines representing the expected imprecision for the method. Assuming a Gaussian (normal) distribution of imprecision, the results are distributed uniformly around the mean, with results observed more frequently closer to the mean than near the extremes of the distribution. More details for such a method can be found at https: / / www .westgard.com / lessons / basic-qc-practices-l / 28-lesson12.html.

[0124] According to a particular embodiment, a conclusion is made regarding the quality of successive sets of sequenced genomic data over time, which depends upon the comparison made and the observed evolution of observed values over time.

[0125] The invention also relates to a data processing apparatus configured for carrying out a method as described in any embodiment herein, especially a computer-implemented method for processing a set of sequenced genomic data, in particular when the method is provided with GQ score values for a set of sequenced genomic data. Conversely, the data processing apparatus can comprise a processor that will compute GQ score values for a set of sequenced genomic data, when provided with (fed with) the raw data required to this end.

[0126] The invention also relates to a computer program product comprising software code adapted to cause a computer to perform a method as described in any embodiment herein, especially a computer- implemented method for processing a set of sequenced genomic data, in particular when the method is provided with GQ score values for a set of sequenced genomic data.

[0127] The invention also relates to a non-transitory computer-readable medium having stored thereon the computer program product described above and herein.

[0128] The invention also relates to the use of a data processing apparatus, or a computer program product according, or a non-transitory computer-readable medium as described in any embodiment herein, for carrying out a method as disclosed in any embodiment herein. In particular, the invention also relates to the use of a data processing apparatus, or a computer program product according, or a non-transitory computer-readable medium as described in any embodiment herein, for determining quality compliance of a set of sequenced genomic data either at the level of the set of sequenced genomic data or at the level of at least one sample comprised in the set of sequenced genomic data, or both, or for determining threshold values for determining whether or not a set of sequenced genomic data is quality compliant, or for monitoring over time the quality of successive sets of sequenced genomic data. Other examples and features of the invention will be apparent when reading the examples and the Figures, which illustrate the experiments conducted by the inventors, in complement to the features and definitions given in the present description.

[0129] Legend of the Figures

[0130] Fig. 1 Representation of the mean versus variance of the Genotyping Quality (GQ) score.

[0131] Fig. 2 Two-dimensional plot of the mean and variance of the genotyping quality (GQ) score. The samples analyzed are represented in a two-dimensional diagram reporting the mean and variance of the GQ Score of the 769 samples highlighted by Trusight Flex Expanded capture. Each point on the diagram corresponds to a sample.

[0132] Fig. 3 representation of the dispersion of the average genotyping quality score per sample as a function of the number of quality metrics out of bounds. NOTE. 1 = 1 metric out of bounds; 2 = 2 metrics out of bounds; 3 = 3 metrics out of bounds

[0133] Fig. 4 Dispersion of the variance of genotyping quality scores per sample as a function of the number of out-of-bounds quality metrics. NOTE. 1 = 1 metric out of bounds; 2 = 2 metrics out of bounds; 3 = 3 metrics out of bounds

[0134] Fig. 5 Two-dimensional plot of the mean and variance of the genotyping quality (GQ) score, relative to a sample that is different from the sample analyzed in Figure 2.

[0135] Experimental Section and Results

[0136] When Clinical exome sequencing (CES) was done, it was done according to the guidance provided in Alix T, Chery C, Josse T, Bronowicki JP, Feillet F, Gueant-Rodriguez RM, Namour F, Gueant JL, Oussalah A. Predictors of the utility of clinical exome sequencing as a first-tier genetic test in patients with Mendelian phenotypes: results from a referral center study on 603 consecutive cases. Hum Genomics. 2023 Feb 5;17(1 ):5. CES was performed using TruSight One Panel and the Illumina MiSeq platform (Illumina, Evry, France) or TruSight One expanded panel and the Illumina NextSeq 550 platform. The strategy was based on Illumina TruSight captures that cover a large number of disease-associated genes instead of a full- exome-based approach. TruSight One Sequencing Panel provides comprehensive coverage of > 4800 disease-associated genes; T ruSight One Expanded Sequencing Panel targets ~ 1900 additional genes with recent disease associations in the scientific literature. We used Nextera and Nextera Flex enrichment solutions with TruSight One and TruSight One Expanded captures, respectively. TruSight One Expanded and TruSight One cover 100% and 82% (145 / 176) of the 176 mitochondrial nuclear genes reported (see additional files of the cited reference). Bioinformatics analyses are detailed in the Supplemental Methods of the cited reference, which is incorporated herein in its entirety.

[0137] 1. Risk Management and Quality Control in High-Throughput Genome Sequencing: From ISO 15189 to the GQ Score Mathematical Approach

[0138] High-throughput genome sequencing is a molecular biology technique enabling the rapid sequencing of thousands or even millions of DNA molecules simultaneously, enabling the detection of genetic variants, the number of which varies according to the type of capture and genome coverage. For example, clinical exome sequencing generates on average between 13,000 and 15,000 genetic variants. The high-throughput sequencing process generates a large amount of data, such as quality metrics on the level of enrichment of reads, bases, fragment size and sequencing coverage. Among the generated quality metrics , each generated variant is associated with a genotyping quality score called the "GQ Score", whose value ranges from 0 to 99 (see section "Description of the concept of mean versus variance of the genotyping quality score (GQ score)"). Bioinformatics analysis of the data generated by high-throughput sequencing is therefore essential.

[0139] Given the criticality of the decisions to be taken on the basis of data of sequencing, it is essential to ensure the quality and risk control of the technique as implemented. This is recommended by the COFRAC, the French Accreditation Committee with regards to the NF EN ISO 15189 norm. In this context, laboratories must provide documentation and raw data recording to enable potential exploitation of results. In addition, analyses must be carried out to identify risks at the various stages of NGS pre-, per- and postanalysis, in order to define the most appropriate strategy.

[0140] The high-throughput genome sequencing process includes rigorous quality control steps to ensure that data analysis leads to a reliable diagnosis. High-throughput genome sequencing is a widely used analysis method in medical biology, and as such must comply with the ISO 15189 norm.

[0141] However, to date there is a lack of tools for addressing the problem of how to evaluate the quality of sequencing, especially over time, and the quality of sequenced batches in a given context.

[0142] The invention here comes with an approach based on risk control using quality metrics, combined or not with complex scores, as described in the present application. The invention constitutes a new risk management tool based on a mathematical hypothesis linked to the genotyping quality of identified variants. Taking into consideration all the variants generated during a capture for a given sample, the invention is formulated on the following hypothesis: when all the called-up variants of good quality, the average of the genotyping quality (GQ) scores for all the variants will be 99, and the calculated variance will be zero. On the other hand, if the DNA sequencing quality of a sample is poor, the average GQ score of the sample will be low and the variance will be high (Figure 1).

[0143] The concept of calculating the mean and variance of "GQ scores" per sample enables the creation of a two-dimensional visualization of the samples analyzed in a series. This visualization identifies the distribution of means and variances of GQ scores per sample, enabling the identification of samples likely to have poor sequencing quality. According to a particular embodiment, the "mean vs. variance GQ score diagram" approach is also used to calculate the mean and variance of the GQ score per sample, the reference values of which are not known.

[0144] 2. Formalizing and using GQ Score in High-Throughput Sequencing: the concept of GQ Score mean versus variance

[0145] On the GATK website (the Broad Institute's reference website, which develops international reference pipelines for high-throughput sequencing, see https: / / gatk.broadinstitute.org / hc / en-us), the GQ score is based on the probability of detecting the AA, AB, BB genotypes (reference allele: A; alternative allele: B). The GQ corresponds to the difference in probabilities expressed in the phred score (PL) for detection of the AB and BB genotypes (GQ score = PL [AB] - PL [BB]). The PL value is a sample-level annotation calculated by HaplotypeCaller and GenotypeGVCFs software (integrated into the NGS bioinformatics pipeline), and is recorded in the sample-level columns of the variant records in VCF files. The value of the GQ score varies from 0 to infinity, but is set to 99 for ease of use. In the situation where all highlighted variants would have a maximum genotyping quality score (i.e. GQ = 99), the mean GQ score of all samples would be 99 and the variance 0. Samples with a low mean GQ score and a high GQ score variance may indicate poor sequencing quality linked to several causes (e.g. poor sample quality, poor preparation, alteration of the sequencing process).

[0146] 3. Development of the working database

[0147] In this database, we used anonymized data from 32 analysis runs, covering 759 samples analyzed by high-throughput clinical exome sequencing. The working database is composed of 3 sub-datasets. Dataset #1 corresponds to the raw quality metrics data generated by the NGS sequencing process. Subdataset #2 was created by performing a principal component analysis, taking into account all the quality metrics data in order to calculate the Eigenvalues for all 759 samples for the 3 top vectors (PC1 , PC2 and PC3). The aim of the analysis is to define the vectors that can explain the variance of the matrix, called Eigen vectors. The 12 quality metrics have been summarized into 3 main vectors: PC1 , PC2 and PC3. Subdataset #3 corresponds to the data associated with the GQ score. To produce this sub-dataset, the 759 VCFs were imported and a GQ score matrix was constructed, taking into account the 759 samples and all variants identified by NGS sequencing. On average, NGS exome sequencing generates between 13,000 and 15,000 genetic variants. For each sample analyzed, we calculated the mean and variance of the GQ score, taking into account all the genetic variants identified. We performed a "merge" procedure to merge the three sub-datasets to obtain a database that includes for each of the 759 samples the raw quality metrics data from the NGS sequencing (12 columns), the principal components from the quality metrics data (3 columns) and the mean and variance of the GQ score (2 columns).

[0148] 4. Statistical analysis

[0149] Continuous data were presented as median with interquartile range (p25-p75). To identify the three main vectors, we performed a principal component analysis on the 12 quality metrics derived from NGS sequencing. To establish reference values for the NGS quality metrics, we applied the "Robust" method in line with CLSI C28-A3 (Clinical and Laboratory Standards Institute) recommendations.

[0150] The "Robust" method is a statistical approach for defining upper and lower limits for the values of a variable, together with their 90% confidence intervals. It is used to detect outliers or extreme values that could influence the accuracy of the results of a statistical analysis. The "Robust" method is validated by CLSI, a non-profit organization that develops standards and guidelines for clinical laboratories and healthcare facilities. CLSI works closely with healthcare professionals, scientists and laboratory experts to develop standards and guidelines aimed at improving the quality of healthcare. These standards are widely adopted around the world to guide laboratory practices and improve the quality of laboratory test results.

[0151] We also used multivariate regression analysis to assess which quality metrics were independently correlated with the principal components, the mean of the GQ score and the variance of the GQ score. This method enabled us to examine the relationship between a dependent variable (PC1 , PC2, PC3, the mean and variance of the genotyping quality score) and several independent variables (the 12 NGS-derived quality metrics). For each model, we reported correlation coefficients, associated P values, and intercorrelation coefficients to assess collinearity. In addition, we provided the coefficient of determination (R2) for each model, enabling us to assess overall model performance. We used logistic regression analysis to estimate the association between the presence of an anomaly on quality metrics and the risk of poor sequencing quality as identified by the mean-versus-variance method. All statistical analyses were performed using MedCalc, and SVS software (v8.8.1 ; Golden Helix, Inc., Bozeman, Montana, USA). 5. Results

[0152] 5.1. Setting Reference Bounds for Quality Metrics, Principal Component Analyses, and Evaluation of Genotyping Quality Means and Variances We applied the "Robust" method as recommended by CLSI to establish a robust confidence interval for the quality metrics and the mean and variance of the genotyping quality score. These reference limits are reported in Table 1. It is interesting to note that the lower limit of the genotyping quality score per sample is set at 88.6 and the upper limit of the variance of the genotyping quality score per sample is set at 496 (Table 1).

[0153] Table 1. Reference bounds for quality metrics, mean and variance of genotyping quality

[0154] NOTE: NR = Number of reads in millions

[0155] * : The confidence interval was established using the CLSI (Clinical and Laboratory Standards Institute) method (reference). t: the limits shown were used to establish the lower and upper bounds for the reference values of the quality metrics.

[0156] 5.2. Comparison of quality approaches based on principal components and mean-variance analysis of the GQ score

[0157] 5.2.1. Multivariate modeling of principal component analysis (PC1)

[0158] In multivariate analysis, 6 quality metrics were independently associated with the first principal component. The multivariate model had a coefficient of determination R2of 0.50. Metrics 2, 4 and 5 are strongly collinear.

[0159] 5.2.2. Multivariate modeling of principal component analysis 2 (PC2)

[0160] In multivariate analysis, 4 quality metrics were independently associated with the second principal component. The multivariate model had a coefficient of determination R2of 0.05. The intercorrelation coefficient for each of the variables is not high, so we can deduce that the variables are not correlated with each other.

[0161] 5.2.3. Multivariate modeling of principal component analysis (PC3)

[0162] Using multivariate analysis, we were able to identify 6 quality metrics that were independently associated with the third principal component. The multivariate model has a coefficient of determination R2of 0.14. We also found that metrics 2 and 3 were highly collinear.

[0163] 5.2.4. Multivariate modelling of average genotyping quality

[0164] By performing a multivariate analysis, we identified 5 quality metrics that were associated with the mean genotyping quality score. The multivariate model has a coefficient of determination R2of 0.96. The intercorrelation coefficient for each of the variables is low, so we can deduce that the variables are not correlated with each other.

[0165] 5.2.5. Multivariate modelling of genotyping quality variance

[0166] We used multivariate analysis to identify 9 quality metrics that were independently related to the variance of the genotyping quality score. The multivariate model had a coefficient of determination R2of 0.93. We noted that metrics 2 and 4 were strongly collinear.

[0167] Figure 2 shows the samples analyzed in a two-dimensional diagram highlighting the mean and variance of the genotyping quality scores for the 769 samples defined in section 3. This figure shows the samples in relation to the lower limit of the mean of the genotyping quality score per sample, set at 88.6, and the upper limit of the variance of the genotyping quality score per sample, set at 496.

[0168] 5.3. Analysis of the mean and variance of genotyping quality as a function of quality metrics outside the set limits

[0169] Reference is made to Figure 3, which represents the dispersion of the average genotyping quality score per sample as a function of the number of quality metrics out of bounds, and to Figure 4, which represents the dispersion of the variance of genotyping quality scores per sample as a function of the number of out-of-bounds quality metrics.

[0170] 5.4. Association between the number of out-of-bounds quality metrics and the risk of poor sequencing quality according to the mean versus variance approach.

[0171] We assessed the association between the number of out-of-bounds points and the presence of a genotyping quality score below the established threshold of 88.6. To do this, we performed an ROC analysis, which revealed a discriminant threshold greater than 2, enabling us to detect an average GQ score with a sensitivity of 92%, a specificity of 86% and an area under the ROC curve of 0.913. We then performed a logistic regression analysis in order to estimate the association between a number of out-of- bounds quality metrics greater than 2 and an average GQ score below 88.6. Logistic regression analysis showed that a sample with at least two out-of-bounds quality metrics had a high risk of obtaining a mean GQ score below 88.6, with an odds ratio of 73.8 (95% Cl: 15.6 - 350.4; P < 0.0001 ).

[0172] We assessed the association between the number of out-of-bounds points and the presence of a variance in the genotyping quality score above the threshold set at 496. To do this, we performed an ROC analysis, which revealed a discriminant threshold greater than 2, enabling us to detect an average GQ score with a sensitivity of 90%, a specificity of 82% and an area under the ROC curve of 0.904. We then performed a logistic regression analysis to estimate the association between a number of out-of-bounds metrics greater than 2 and a mean GQ score greater than 496. The logistic regression analysis showed that a sample with a number of out-of-bounds quality metrics greater than 2 had a high risk of obtaining a mean GQ score greater than 496, with an odds ratio of 36.40 (95% Cl: 11 .56 - 114.59; P < 0.0001 ).

[0173] 6. Example based on a different dataset on a small panel of 8 genes

[0174] These results have been obtained on a set of sequenced genomic data that is different from the set of sequenced genomic data analyzed in Table 1 , i.e., the sequencing capture is different. Table 2 Reference bounds for quality metrics, mean and variance of genotyping quality

[0175] Results of the global analysis

[0176] 95% Reference interval, Double-sided a Bootstrap confidence interval (10000 iterations; random number seed: 978).

[0177] 95% Reference interval, Double-sided a Bootstrap confidence interval (10000 iterations; random number seed: 978).

[0178] See Figure 5 for an illustration of a two-dimensional vizualisation of this data.

[0179] Integrating 12 Quality Metrics with the 'Mean vs. Variance' Approach for Enhanced Sequencing Quality Assessment

[0180] The "mean vs. variance" approach of the GQ score allows for the simple identification of samples with good or poor sequencing quality. Some samples exhibit borderline values in their mean and / or variance of the GQ score, making the decision less clear. In such cases, using the 12 proposed quality metrics helps in making a decision, as the presence of 3 out of 12 metrics outside the acceptable range indicates suboptimal sequencing quality.

[0181] References

[0182] Van der Auwera GA & O'Connor BD. (2020). Genomics in the Cloud: Using Docker, GATK, and WDL in Terra (1st Edition). O'Reilly Media.

[0183] Poplin R, Ruano-Rubio V, DePristo MA, Fennell TJ, Carneiro MO, Van der Auwera GA, Kling DE, Gauthier LD, Levy-Moonshine A, Roazen D, Shakir K, Thibault J, Chandran S, Whelan C, Lek M, Gabriel S, Daly MJ, Neale B, MacArthur DG, Banks E. (2017). Scaling accurate genetic variant discovery to tens of thousands of samples bioRxiv, 201 178. DOI: 10.1101 / 201 178

[0184] Van der Auwera GA, Carneiro M, Hartl C, Poplin R, del Angel G, Levy-Moonshine A, Jordan T, Shakir K, Roazen D, Thibault J, Banks E, Garimella K, Altshuler D, Gabriel S, DePristo M. (2013). From FastQ Data to High- Confidence Variant Calls: The Genome Analysis Toolkit Best Practices Pipeline. Curr Protoc Bioinformatics, 43:1 1.10.1 -1 1 .10.33. DOI: 10.1002 / 047125O953.bi11 10s43.

[0185] DePristo M, Banks E, Poplin R, Garimella K, Maguire J, Hartl C, Philippakis A, del Angel G, Rivas MA, Hanna M, McKenna A, Fennell T, Kernytsky A, Sivachenko A, Cibulskis K, Gabriel S, Altshuler D, Daly M. (201 1 ). A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nat Genet, 43:491 -498. DOI: 10.1038 / ng.806.

[0186] McKenna A, Hanna M, Banks E, Sivachenko A, Cibulskis K, Kernytsky A, Garimella K, Altshuler D, Gabriel S, Daly M, DePristo MA. (2010). The Genome Analysis Toolkit: a MapReduce framework for analyzing nextgeneration DNA sequencing data. Genome Res, 20:1297-303. DOI: 10.1 101 / gr.107524.1 10.

[0187] Bland M (2000) An introduction to medical statistics, 3rd ed. Oxford: Oxford University Press.

[0188] CLSI (2008) Defining, establishing, and verifying reference intervals in the clinical laboratory: approved guideline - third edition. CLSI Document C28-A3. Wayne, PA: Clinical and Laboratory Standards Institute.

[0189] Efron B, Tibshirani RJ (1993) An introduction to the Bootstrap. Chapman & Hall / CRC.

[0190] Le Boedec K (2019) Reference interval estimation of small sample sizes: A methodologic comparison using a computer-simulation study. Veterinary Clinical Pathology 48:335-346.

[0191] Horn PS, Pesce AJ (2005). Reference Intervals. A User's Guide. Washington, DC: AACC Press.

[0192] Lentner C (Ed) (1982) Geigy Scientific Tables, 8th edition, Volume 2. Basle: Ciba-Geigy Limited.

[0193] NCCLS (2000) How to define and determine reference intervals in the clinical laboratory: approved guideline. 2nd edition. NCCLS document C28-A2. Wayne, PA: NCCLS.

[0194] Reed AH, Henry RJ, Mason WB (1971 ) Influence of statistical method used on the resulting estimate of normal range. Clinical Chemistry, 17:275-284.

[0195] Schoonjans F, De Bacquer D, Schmid P (201 1 ) Estimation of population percentiles. Epidemiology 22: 750- 751.

[0196] Sheskin DJ (201 1 ) Handbook of parametric and nonparametric statistical procedures. 5th ed. Boca Raton: Chapman & Hall / CRC.

[0197] Tukey JW (1977) Exploratory data analysis. Reading, Mass: Addison-Wesley Publishing Company.

[0198] Westfall PH (2014) Kurtosis as Peakedness, 1905 - 2014. R.I.P. The American Statistician 68:191 -195.

[0199] Alix T, Chery C, Josse T, Bronowicki JP, Feillet F, Gueant-Rodriguez RM, Namour F, Gueant JL, Oussalah A. Predictors of the utility of clinical exome sequencing as a first-tier genetic test in patients with Mendelian phenotypes: results from a referral center study on 603 consecutive cases. Hum Genomics. 2023 Feb 5;17(1 ):5. doi: 10.1 186 / S40246-023-00455-X. PMID: 36740706; PMCID: PMC9899384.

[0200] Mergnac JP, Wiedemann A, Chery C, Ravel JM, Namour F, Gueant JL, Feillet F, Oussalah A. Diagnostic yield of clinical exome sequencing as a first-tier genetic test for the diagnosis of genetic disorders in pediatric patients: results from a referral center study. Hum Genet. 2022 Jul;141 (7):1269-1278. doi: 10.1007 / s00439-021 -02358-0. Epub 2021 Sep 8. PMID: 34495415.

[0201] Oussalah A, Jeannesson-Thivisol E, Chery C, Perrin P, Rouyer P, Josse T, Cano A, Barth M, Fouilhoux A, Mention K, Labarthe F, Arnoux JB, Maillot F, Lenaerts C, Dumesnil C, Wagner K, Terral D, Broue P, De Parscau L, Gay C, Kuster A, Bedu A, Besson G, Lamireau D, Odent S, Masurel A, Rodriguez-Gueant RM, Feillet F, Gueant JL, Namour F. Population and evolutionary genetics of the PAH locus to uncover overdominance and adaptive mechanisms in phenylketonuria: Results from a multiethnic study. EBioMedicine. 2020 Jan;51 :102623. doi: 10.1016 / j.ebiom.2019.102623. Epub 2020 Jan 7. PMID: 31923802; PMCID: PMC7000351.

[0202] Wiedemann A, Chery C, Coelho D, Flayac J, Gueguen N, Desquiret-Dumas V, Feillet F, Lavigne C, Neau JP, Fowler B, Baumgartner MR, Reynier P, Gueant JL, Oussalah A. Mutations in MTHFR and POLG impaired activity of the mitochondrial respiratory chain in 46-year-old twins with spastic paraparesis. J Hum Genet. 2020 Jan;65(2):91 -98. doi: 10.1038 / sl 0038-019-0689-y. Epub 2019 Oct 23. PMID: 31645654.

Claims

CLAIMS1 . A computer-implemented method for processing a set of sequenced genomic data, in particular for determining quality compliance of a set of sequenced genomic data either at the level of the set of sequenced genomic data or at the level of at least one sample comprised in the set of sequenced genomic data, or both, the method comprising:Obtaining genotype quality (GQ) score values for each sample of said set of sequenced genomic data, andComputing a variable indicative of a central tendency of the GQ score values of the variants thereof, GQ mean-like score, and a variable assessing the variability of the said GQ meanlike score values of the variants thereof, GQ variance-like score, for each sample of said set of sequenced genomic data, andDetermining, from the values of the GQ mean-like score and the GQ variance-like score, whether or not the set of sequenced genomic data is quality compliant either at the level of the set of sequenced genomic data level, or at the level of at least one sample of the set, or both.

2. The method according to claim 1 , wherein determining whether or not the set of sequenced genomic data is quality compliant, is carried out through a two-dimensional visualization of the GQ mean-like scores and the GQ variance-like scores of the set of sequenced genomic data, in particular by considering the form of the distribution with respect to predetermined threshold GQ mean-like score and GQ variance-like score values.

3. The method according to claim 2, wherein the two-dimensional visualization is a two-dimensional visualization of the GQ mean-like scores plotted against the GQ variance-like scores, or is a two- dimensional visualization of the GQ variance-like scores plotted against the GQ mean-like scores.

4. The method according to any one of claims 1 to 3, wherein determining whether or not a sample of the set of sequenced genomic data is quality compliant, is carried out according to the following: for a sample of the set of sequenced genomic data, respectively comparing the GQ mean-like score and the GQ variance-like score io predetermined threshold values, and: if the GQ mean-like score of the sample is superior or equal to its predetermined threshold value and the GQ variance-like score is inferior or equal to its predetermined threshold value, classifying the corresponding sample as quality compliant, if the GC score mean of the sample is superior or equal to its predetermined threshold value and the GC score variance is superior to its predetermined threshold value, classifying the corresponding sample as not quality compliant, if the GQ mean-like score of the sample is inferior to its predetermined threshold value and the GQ variance-like score is superior to its predetermined threshold value, classifying the corresponding sample as not quality compliant, and32if the GQ mean-like score of the sample is inferior to its predetermined threshold value and the GQ variance-like score is inferior or equal to its predetermined threshold value, classifying the corresponding sample as not quality compliant.

5. The method according to any one of claims 1 to 4, wherein a GQ score value is the difference between a normalized Phred-scaled likelihood (PL) of the second most likely genotype of an allele observed at a site in the observed sequence data and a normalized Phred-scaled likelihood (PL) of the most likely genotype of an allele observed at the same site in the observed sequence data, as calculated by the HaplotypeCaller and GenotypeGVCFs tools of the Genome Analysis Toolkit (gatk) software.

6. The method according to any one of claims 4 to 5, which is for determining whether or not a sample of the set of sequenced genomic data is quality compliant, further comprising a step of calculating for each sample of the set of sequenced genomic data a number of quality metrics, and their confidence intervals, and wherein when a GQ mean-like score and / or a GQ variance-like score of the sample is equal to the predetermined threshold value to which it is compared, concluding about quality compliance of the sample by taking into consideration the number of quality metrics outside their confidence intervals.

7. The method according to any one of claims 4 to 6, wherein: the predetermined threshold value for the GQ mean-like score is defined as the lower limit value of a 90% confidence interval calculated using a method selected among: the Clinical and Laboratory Standards Institute (CLSI) robust method, the Normal distribution method, a non-parametrical percentile method, on the GQ mean-like score values of a reference set of sequenced genomic data, and / or the predetermined threshold value for the GQ variance-like score is defined as the upper limit value of a 90% confidence interval calculated using a method selected among: the Clinical and Laboratory Standards Institute (CLSI) robust method, the Normal distribution method, a non-parametrical percentile method, on the GQ variance-like score values of a reference set of sequenced genomic data.

8. The method of claim 7, wherein the reference set of sequenced genomic data is the set of sequenced genomic data analyzed in the method of claims 1 to 6.

9. A computer-implemented method for determining threshold values for determining whether or not a set of sequenced genomic data is quality compliant, said method comprising steps of:Obtaining GQ score values for the set of sequenced genomic data,For each sample of said set of sequenced genomic data, computing (i) a variable indicative of a central tendency of the GQ score values of the variants thereof, GQ mean-like score and (ii) a variable assessing the variability of the said GQ mean-like score values of the variants thereof, GQ variance-like score,Determining the lower limit value of a 90% confidence interval calculated using a method selected among: the Clinical and Laboratory Standards Institute (CLSI) robust method, the Normal distribution method, a non-parametrical percentile method, on the GQ mean-like score values of the said set of sequenced genomic data, andDetermining the upper limit value of a 90% confidence interval calculated using a method selected among: the Clinical and Laboratory Standards Institute (CLSI) robust method, the Normal distribution method, a non-parametrical percentile method, on the GQ variance-like score values of the said set of sequenced genomic data.

10. A computer-implemented method for assessing whether a threshold value for GQ mean-like score or a threshold value for GQ variance-like score determined according to claim 9 is suited to evaluate the risk that sequenced genomic data is not quality compliant, said method comprising the steps of: a. For the set of sequenced genomic data analyzed with the method according to any one of claims 1 to 8, calculating a number of quality metrics, and their confidence intervals, in particular using the CLSI robust method, and b. Determining the number of quality metrics as calculated in step a. that is(are) outside their respective confidence intervals as calculated in step a., and c. Analyzing by logistic regression the association between the number of quality metrics outside their confidence intervals as determined in step b. and the value of the threshold value to be assessed, and d. If a discriminating score can be found on the basis of a sensitivity of at least 80% and a specificity of at least 80%, concluding that the threshold value for GQ mean-like score or the threshold value for GQ variance-like score is suited to evaluate the risk that sequenced genomic data is not quality compliant.1 1 . The method according to any one of claims 1 to 10 further comprises sequencing the genomic data ahead of obtaining GQ score values.

12. A method for monitoring over time the quality of successive sets of sequenced genomic data, said method comprising the steps of: a. Carrying out at a first point in time on a set of sequenced genomic data a computer- implemented method for determining threshold values according to claim 9, which returns GQ mean-like score, GQ variance-like score threshold values, and optionally quality metrics according to claim 10, and b. Carrying out at a second point in time on a set of sequenced genomic data that is different from the set of sequenced genomic data of step a., a computer-implemented method for determining threshold values according to claim 8, which returns GQ mean-like score, GQ variance-like score threshold values, and optionally quality metrics according to claim 9, andc. Comparing the evolution, between the first and second points in time, of the GQ mean-like score, GQ variance-like score threshold values, and optionally quality metrics returned in steps a. and b.

13. The method according to any one of claims 1 to 12, wherein the sequenced genomic data is Next- Generation Sequencing (NGS) generated and: a. is selected among: Whole Genome Sequencing (WGS) data or Whole exome sequencing (WES) data, and / or b. is selected among: clinical exome data and gene panel data, and / or c. is from samples previously obtained from human, animal or environmental samples.

14. A data processing apparatus configured for carrying out the method of any one of claims 1 to 13, in particular when provided with GQ score values or equivalent metrics of a set of sequenced genomic data.

15. A computer program product comprising software code adapted to cause a computer to perform a method according to any one of claims 1 to 13, in particular when provided with GQ score values of a set of sequenced genomic data.

16. A non-transitory computer-readable medium having stored thereon the computer program product of claim 15.

17. Use of a data processing apparatus according to claim 14, or a computer program product according to claim 15, or a non-transitory computer-readable medium according to claim 16, for carrying out a method according to any one of claims 1 to 13.

18. Use of a data processing apparatus according to claim 14, or a computer program product according to claim 15, or a non-transitory computer-readable medium according to claim 16, for determining quality compliance of a set of sequenced genomic data either at the level of the set of sequenced genomic data or at the level of at least one sample comprised in the set of sequenced genomic data, or both, or for determining threshold values for determining whether or not a set of sequenced genomic data is quality compliant, or for monitoring over time the quality of successive sets of sequenced genomic data.

Citation Information

Patent Citations

  • System and method for process control of gene sequencing

    US20170228496A1