Genomic sequencing selection system

BR122026018155A2Pending Publication Date: 2026-08-25
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
BR122026018155
Authority / Receiving Office
BR · BR
Patent Type
Applications
Publication Date
2026-08-25

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

1 / 27 GENOMIC SEQUENCING SELECTION SYSTEM DIVIDED FROM ORDER BR 112021007293-4, FILED ON 10 / 16 / 2019 CROSS-REFERENCE TO RELATED REQUESTS

[001] This application claims priority to the benefit of Provisional Patent Application No. US 62 / 766,432, entitled “GENOMIC SEQUENCING SELECTION SYSTEM” and filed on October 17, 2018, the content of which is incorporated in its entirety by reference by the present invention for all purposes. BACKGROUND OF THE INVENTION

[002] Genomic sequencing systems, including next-generation sequencing (NGS) systems (sometimes referred to as massively parallel sequencing systems or by similar terms), can produce large amounts of sequencing data of varying quality. Specifically, in many implementations, an NGS system can fragment a genome into a plurality of small segments. These small segments can be sequenced in parallel, reducing processing requirements compared to sequencing the entire genome as a whole, and then recombined to generate a complete sequence. Sequence metrics can be calculated on the sequencing data.

[003] NGS systems provide faster and less expensive sequencing compared to first-generation sequencing techniques such as Sanger sequencing. However, NGS systems suffer from inaccuracy or noise due to errors in identifying base sequences or base calling, or errors introduced during sample preparation. Error rates in base readings can be 10% or more, sometimes as high as 25% or more. Given the immense amount of data that can be obtained in a short time by an NGS system, even moderate error rates Petition 870260072528, dated 07 / 21 / 2026, page 15 / 56 2 / 27 can result in data with hundreds of thousands or even millions of incorrect base pairs. SUMMARY OF THE INVENTION

[004] The systems and methods disclosed in the present invention provide measurement of error rates and read quality on a read-by-read basis and, in some implementations, can filter out or exclude low-quality reads or extract high-quality reads and provide detailed metrics. This can reduce processing requirements compared to analyzing all datasets including low-quality or erroneous data and can increase computational speeds of sequence metric determination by reducing the amount of computational time spent on data that may yield inaccurate results. In many implementations, these systems and methods can also reduce memory and bandwidth consumption compared to processing or transferring datasets with high error rates.

[005] In some implementations, the present solution can calculate sequencing statistics, such as depth of coverage. The present solution can determine read statistics, such as variant frequencies, and identify clinically relevant variants. The present solution can read BAM and VCF input files and Phred scale quality scores. The present solution can select high-quality reads based on quality scores and can calculate reference and alternative allele counts for single nucleotide polymorphisms (SNPs), insertions and deletions (INDELs), and structural variants. The present solution can calculate sequencing metrics for different strands to measure read strands. The present solution can also determine minimum, maximum, and average depths for each region of the sequence data. Petition 870260072528, dated 07 / 21 / 2026, page 16 / 56 3 / 27

[006] According to at least one aspect of the disclosure, a method for filtering sequencing data may involve receiving, by a data processing system, data that may include a plurality of genetic sequences. Each of the plurality of genetic sequences may include a chromosome indication, a position indication, a base value, and a quality score. The method may involve selecting, by the data processing system, a subset of the plurality of genetic sequences. Each of the subset of the plurality of genetic sequences may have the same chromosome indication. The method may involve filtering, by the data processing system, from the subset of the plurality of genetic sequences, genetic sequences comprising base values ​​that have a quality score above a predetermined threshold.The method may include determining, by the data processing system, an aggregate count for each position of the filtered genetic sequences. The method may include determining, by the data processing system, an alternative base count for each position of the filtered genetic sequences. The method may include generating, by the data processing system, an identification of a genetic sequence variant based on a ratio of the alternative base count for each position to the aggregate count for each position that exceeds a threshold.

[007] In some implementations, the method may include determining an alternative count for a deletion sequence in the filtered subset of the plurality of genetic sequences where the base values ​​have a quality score above the predetermined threshold. The deletion sequence may start at an index neighboring the position.

[008] The method may include determining an alternative count for an insertion sequence in the filtered subset of the plurality of sequences. Petition 870260072528, dated 07 / 21 / 2026, page 17 / 56 4 / 27 genetic sequences where the base values ​​have a quality score above the predetermined threshold. The method may additionally include determining the alternative count for the insertion sequence by identifying an alternative sequence match. The method may also include identifying a structural variant in the filtered plurality of genetic sequences.

[009] In some implementations, the alternative base count can be determined based on the structural variant identified in the plurality of genetic sequences. The determination of the aggregate count may include counting a match in each of the filtered subset of the plurality of genetic sequences with a CIGAR string.

[010] In some implementations, the determination of the aggregate count may include counting a deletion, insertion, reference jump, soft clip, or hard clip in each of the filtered subset of the plurality of genetic sequences. The method may include calculating at least one of an average read coverage, a minimum read coverage, or a maximum read coverage for the filtered plurality of genetic sequences based on the aggregate count and the alternate base count.

[011] In some implementations, the method may include calculating a reading strip for the plurality of genetic sequences based on the aggregate count and the alternate base count.

[012] According to at least one aspect of the disclosure, a system for filtering sequencing data may include a data processing system. The system may receive data that may include a plurality of genetic sequences. Each of the plurality of genetic sequences may include a chromosome indication, a position indication, a base value, and a quality score. The system may select a subset of the plurality of genetic sequences. Each of the subset Petition 870260072528, dated 07 / 21 / 2026, page 18 / 56 5 / 27 of the plurality of genetic sequences may have the same chromosome indication. The system can filter, from the subset of the plurality of genetic sequences, genetic sequences in which the base values ​​have a quality score above a predetermined threshold. The system can determine an aggregate count for each position in the filtered subset of the plurality of genetic sequences where the base values ​​have a quality score above the predetermined threshold. The system can determine an alternative base count for each position in the filtered plurality of genetic sequences where the base values ​​have a quality score above the predetermined threshold. The system can identify genetic sequence variants based on a ratio of the alternative base count for each position to the aggregate count for each position, and can generate an identifier for the genetic sequence variants.

[013] In some implementations, the system may determine an alternative count for a deletion sequence in the subset of the plurality of genetic sequences where the base values ​​have a quality score above the predetermined threshold. The system may determine an alternative count for an insertion sequence in the filtered subset of the plurality of genetic sequences where the base values ​​have a quality score above the predetermined threshold.

[014] In some implementations, the system can determine the alternative count for the insertion sequence by identifying an alternative sequence match. The system can identify a structural variant in the plurality of genetic sequences.

[015] The system can determine the aggregate count by counting a match in each of the filtered subset of the plurality of genetic sequences with a CIGAR chain. The system can determine the Petition 870260072528, dated 07 / 21 / 2026, page 19 / 56 6 / 27 aggregate count when counting one deletion, insertion, reference jump, soft clip, or hard clip in each of the subset of the plurality of genetic sequences.

[016] The system can calculate at least one of an average read coverage, a minimum read coverage, or a maximum read coverage for the filtered plurality of genetic sequences based on the aggregate count and the alternate base count. The system can calculate a read tape for the plurality of genetic sequences based on the aggregate count and the alternate base count.

[017] The above general description and the following description of the drawings are illustrative and explanatory and are intended to provide further explanation of the invention as claimed. Other innovative objects, advantages and features will be readily apparent to those skilled in the art from the brief description of the drawings and the detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[018] The attached drawings are not intended to be drawn to scale. Similar reference numbers and designations in the various drawings indicate similar elements. For clarity, not every component can be identified in every drawing. In the drawings:

[019] Figure 1 illustrates a block diagram of an exemplary system for computing NGS read depth statistics.

[020] Figure 2 illustrates a block diagram of an exemplary method for determining sequencing data coverage metrics using the system illustrated in Figure 1.

[021] Figure 3 illustrates example sequence listings for a given chromosome.

[022] Figure 4 illustrates a block diagram of a system Petition 870260072528, dated 07 / 21 / 2026, page 20 / 56 7 / 27 computational example. DETAILED DESCRIPTION

[023] The various concepts introduced and discussed in more detail below can be implemented in any of numerous ways, as the concepts described are not limited to any particular method of implementation. Examples of specific implementations and applications are provided primarily for illustrative purposes.

[024] The present solution can calculate sequencing statistics, such as depth of coverage. The present solution can determine variant frequencies and identify clinically relevant variants based on variant frequencies. The present solution can read BAM and VCF input files and Phred scale quality scores. The present solution can select relatively high-quality reads from input files based on quality scores and can calculate reference and alternative allele counts for SNPs, insertions and deletions (INDELs), and structural variants. The present solution can calculate sequencing metrics for different strands to measure read strands. The present solution can also determine minimum, maximum, and average depths for each region of the sequence data.The present solution can use quality scores to select and analyze only readings of relatively high quality, which can increase computational speeds for determining sequence metrics by reducing the amount of computational time spent on data that may provide inaccurate results.

[025] Figure 1 illustrates a block diagram of an exemplary system 100 for computing NSG read depth statistics. The system 100 may include a sequencing system 102. The sequencing system 102 may include a data analyzer 110 that reads files. Petition 870260072528, dated 07 / 21 / 2026, page 21 / 56 8 / 27 of data 114 from a data repository 116. The data analyzer 110 can load the data into a buffer 106. The sequencing system 102 may include a reporting mechanism 104, a filtering mechanism 108, and an analytical mechanism 112. The system 100 may include an NGS sequencer 118 that can provide the data files 114 to the sequencing system 102.

[026] System 100 may include a sequencing system 102. The sequencing system 102 may include at least one server or computer that has at least one processor. For example, the sequencing system 102 may include a plurality of servers located in at least one data center or server tower, or the sequencing system 102 may be a desktop computer. The processor may include a microprocessor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), other special-purpose logic circuits, or combinations thereof. The sequencing system 102 may be a data processing system as described in relation to Figure 4. For example, the sequencing system 102 may include one or more processors and memory.The 102 sequencing system may include a user interface (e.g., a graphical user interface) that is rendered and displayed to the user via a display coupled to the 102 sequencing system. One or more input / output (I / O) devices may be coupled to the 102 sequencing system.

[027] The sequencing system 102 may include the data repository 116. The data repository 116 may include one or more local or distributed databases. The data repository 116 may include computer or memory data storage and may store one or more files of Petition 870260072528, dated 07 / 21 / 2026, page 22 / 56 9 / 27 data 114. The data repository 116 may include non-volatile memory, such as one or more hard disk drives (HDDs) or other magnetic or optical storage media, one or more solid-state drives (SSDs), such as a flash drive or other solid-state storage media, one or more hybrid magnetic or solid-state drives, one or more virtual storage volumes, such as cloud storage, or a combination thereof.

[028] The sequencing system 102 can store one or more data files 114 in the data repository 116. Each of the data files 114 can include a plurality of genetic sequence data. The genetic sequence data can include a chromosome indication, a position indication, a base value, and a quality score.

[029] Data files 114 can be data files that are in variant call format (VCF), sequence alignment mapping format (SAM), binary sequence alignment mapping (BAM), or other archive data file formats used in bioinformatics. For example, data files 114 may include text data or binary data. In some implementations, data files 114 may include sequencing data strings. In some implementations, data files 114 may include sequencing data that identifies the differences between a reference sequence and a sample sequence.

[030] For example, the VCF file format can be used to store sequence variations. The VCF file format can be used to store single nucleotide polymorphisms (SNPs), short insertions or deletions (e.g., less than 10 base pairs), and large structural variants. The VCF file format (and other file formats) may include a header section and a body section. The header section may Petition 870260072528, dated 07 / 21 / 2026, page 23 / 56 10 / 27 Include metadata that further describes the data contained in the body of the VCF file format. The body of the VCF file format may include a plurality of columns. Each row may indicate a variation. The columns may identify the chromosome on which the variation is named; a position of the variation in the sequence; an identifier of the variation; a reference base value for the position; an alternative base value for the position (e.g., which base other than the reference base was read at the position); a score; and a flag indicating which of a given set of filters the variation passed.

[031] The sequencing system 102 may include an NGS sequencer 118. The NGS sequencer 118 may generate the data files 114. The system 100 may include a plurality of NGS sequencers 118. The NGS sequencer 118 may provide samples from which the NGS sequencer 118 generates sequencing data. The NGS sequencer 118 may save the data in one of the file formats described above. In some implementations, the NGS sequencer 118 may transmit the data files 114 to the sequencing system 102 via a network. In some implementations, the NGS sequencer 118 may transmit the data files 114 to an intermediate device, such as cloud-based storage or a removable hard drive. The data files 114 may be transferred from the intermediate device to the sequencing system 102.

[032] The sequencing system 102 may include a data analyzer 110. The data analyzer 110 may be any script, file, program, application, set of instructions, or executable computer code that is configured to enable a computing device on which the data analyzer 110 runs to read and extract data from the data repository 116. The data analyzer 110 may read the files of Petition 870260072528, dated 07 / 21 / 2026, page 24 / 56 11 / 27 data 114 from data repository 116. In some implementations, data files 114 may be stored in data repository 116 in a compressed format. Data analyzer 110 may decompress data files 114 before extracting sequencing data from data files 114. Data analyzer 110 may read data files 114 from data repository 116, which may be stored on the sequencing system's hard drive 102. Data analyzer 110 may load data files 114 and store the data from data files 114 in buffer 106.

[033] In some implementations, data parser 110 may load one or more data files 114 into buffer 106. Data parser 110 may parse or process the data before data parser 110 loads the data into buffer 106. For example, data parser 110 may parse the body of the VCF file format into one or more dictionaries or other structural file formats.

[034] The sequencing system 102 may include a buffer 106. The buffer may be stored in random access memory (RAM) or other cached memory. The buffer may be stored in volatile memory. In some implementations, reading from and writing to buffer 106 may be faster than reading from and writing to the data store 116. The data analyzer 110 may load data files 114 into buffer 106 to reduce the number of reads and writes performed on the data store 116 to improve the overall computation speeds of the sequencing system 102.

[035] The 102 sequencing system may include a 108 filtering mechanism. The 108 filtering mechanism may be any script, file, program, application, set of instructions, or executable code. Petition 870260072528, dated 07 / 21 / 2026, page 25 / 56 12 / 27 computer that is configured to enable a computing device on which the filtering mechanism 108 runs to select variants from the sequencing data loaded into buffer 106. As described above, each variant may include a score. The score may be a quality score. The quality score may be a Phred quality score. The quality score may be an indication of the quality of the base identified during the sequencing process. For example, the quality score may be an indication of the probability that the base at the given position was not correctly identified and was not a sequencing error.

[036] The 108 filtering mechanism can select only variations that have a quality score above a predetermined threshold. For example, the 108 filtering mechanism can discard variations with a quality score below the predetermined threshold from buffer 106 or further analysis. In some implementations, the 108 filtering mechanism does not use any variations with a Phred quality score less than 60, less than 50, less than 40, less than 30, or less than 20. In some implementations, the quality score may be based on the average reads per base in the sequencing data. For example, the quality score threshold may initially be set to 30 and then lowered if the average reads per base are above 100.

[037] The sequencing system 102 may include an analytical engine 112. The analytical engine 112 may be any script, file, program, application, set of instructions, or executable computer code that is configured to enable a computing device on which the analytical engine 112 runs to calculate sequencing statistics. Petition 870260072528, dated 07 / 21 / 2026, page 26 / 56 13 / 27

[038] Analytical engine 112 can calculate alternative base frequencies at each of the positions (P) indicated in the data files 114. The alternative base frequencies can be based on a count of all reads at a given position. For example, analytical engine 112 can determine the number of times each base occurs at each position in the genetic sequence (or portion thereof), which can be referred to as an ALT base count for the given base. Analytical engine 112 can determine an aggregate count for each position in the genetic sequence (or portion thereof). In some implementations, analytical engine 112, when determining the ALT base count and the aggregate base count, can only include or count bases with a quality score above a predetermined threshold.

[039] Analytic engine 112 can calculate alternative base frequencies for insertions and deletions. In some implementations, insertions or deletions are less than 10 base pairs in length. For deletions, analytic engine 112 can determine the ALT count by identifying each of the deletions of a given length K that starts at position P+1. For insertions, analytic engine 112 can determine the ALT count by counting the number of occurrences of an insertion of a given length that corresponds to a CIGAR string. For large structural variants, analytic engine 112 can determine a reference count (REF), an ALT count, and an aggregate or total count. Analytic engine 112 can determine the REF count as the number of occurrences that analytic engine 112 identifies that corresponds to a CIGAR string across an event boundary.Analytical mechanism 112 can determine ALT counts as the number of deletions, insertions, reference jumps, discrete cleavages, or coarse cleavages in the ALT. Petition 870260072528, dated 07 / 21 / 2026, page 27 / 56 14 / 27 CIGAR chain across the event threshold. The total count may be the sum of the REF count and the ALT count. Based on statistics and other data determined by analytical mechanism 112, analytical mechanism 112 can identify clinically relevant variants of several common variants.

[040] The sequencing system 102 may include a reporting mechanism 104. The reporting mechanism 104 may be any script, file, program, application, instruction set, or executable computer code that is configured to enable a computing device on which the reporting mechanism 104 runs to generate reports based on the data generated by the analytical mechanism 112. The reporting mechanism 104 may receive the data generated by the analytical mechanism 112, such as ALT count, REF count, and ALT frequencies. The reporting mechanism 104 may generate reports based on the data. The reporting mechanism 104 may determine and include in the report coverage frequencies; tape reading; and intermediate, maximum, and average coverage.

[041] Figure 2 illustrates a block diagram of an exemplary method 200 for determining sequencing data coverage metrics. Method 200 may include receiving data (BLOCK 202). Also with reference to Figure 1, sequencing system 102 may receive the data. Sequencing system 102 may receive the data from the NGS sequencer 118 or sequencing system 102 may retrieve the data from the data repository 116. Sequencing system 102 may receive the data as BAM, VCF, text, or other file format that may contain sequencing data. Sequencing system 102 may also receive Phred scale quality scores for the received data. The data may include a plurality of genetic sequences. The data may indicate a chromosome to genetic sequence, position data, Petition 870260072528, dated 07 / 21 / 2026, page 28 / 56 15 / 27 base values ​​in each of the positions and quality scores for the base values. In some implementations, the 102 sequencing system can receive and open the data files. The 102 sequencing system can read the data files into buffer 106. Reading the data files into buffer 106 can reduce the number of reads that are made from data repository 116.

[042] Method 200 may include selecting a genetic sequence (BLOCK 204). The 102 sequencing system may select one or more genetic sequences that belong to the same chromosome. In some implementations, the 102 sequencing system may select one or more genetic sequences that also belong to the same general location on the chromosome or to the same specific location. For example, genetic sequences may be received in data files that include a plurality of columns. One of the plurality of columns may indicate a chromosome for the sequence data contained in another column of the data file. The 102 sequencing system may filter through the data to select genetic sequences that are below a predetermined chromosome.

[043] Method 200 may include determining whether each base value has a threshold above a threshold (BLOCK 206). The 102 sequencing system can identify base values ​​in the sequence data that include base values ​​at a given position that are below the quality threshold. The 102 sequencing system can discard loaded data for the given position where the base value has a quality score below the predetermined threshold. The 102 sequencing system can save the base values ​​for a given position that have a quality score above the predetermined threshold in a data structure, such as a dictionary, that is Petition 870260072528, dated 07 / 21 / 2026, page 29 / 56 16 / 27 saved in buffer 106.

[044] Method 200 may include identifying a variant type in the sequence data (BLOCK 208). Sequencing system 102 may determine whether the variant is a single nucleotide polymorphism (SNP) and continues in BLOCK 210, an insertion or deletion and continues in BLOCK 212, or a large structural variant and continues in BLOCK 226. In some implementations, insertions or deletions are smaller than 10 base pairs (bp), and large structural variants are larger than 10 base pairs.

[045] If the 102 sequencing system determines that the variant is an SNP, the 200 method may include determining an aggregate count for the position (BLOCK 216). Also with reference to Figure 3, among others, Figure 3 illustrates four sequence listings 300(1) to 300(4) (which are generally referred to as sequence listings 300) for a given chromosome. Each of the sequence listings 300 may include a plurality of base pairs 302. Each of the selected sequence listings 300 may overlap a given base pair position 304. In general, the location of a base pair 302 can be described with the variable P where the next base pair 302 has location P+1 and the previous base pair 302 has location P-1. In this example, the data files may indicate that the SNP occurs at base pair position 304, which can be referred to as P.For example, sequence listing 300(1) and sequence listing 300(2) indicate that the base pair at base pair position 304 should be G, and sequence listing 300(3) and sequence listing 300(4) indicate that the base pair at base pair position 304 should be C. Each of the base pairs 302 at base pair position 304 may have an associated quality score.

[046] The aggregate count for position P may be the number of sequence listings 300 that include position P with a quality score above Petition 870260072528, dated 07 / 21 / 2026, page 30 / 56 17 / 27 of the predetermined limit. For example, and continuing with the example illustrated above in Figure 3, if base pair 302 in sequence listing 300(4) at base pair position 304 has a quality score below the predetermined limit, the aggregate count for base pair position 304 may be 3.

[047] Method 200 may include determining the alternative count (ALT) for the position (BLOCK 218). The 102 sequencing system may determine an ALT count for each base pair (e.g., C, G, G, and T). The ALT count for each base pair location 304 may be the aggregate count or number of occurrences of the base pair at the base pair location 304. The 102 sequencing system may only include base pairs 302 in the ALT count that have a quality score above the predetermined threshold. For example, and with reference to the example illustrated in Figure 3, the 102 sequencing system may determine that the ALT count for G at the base pair location 304 is 2 and the ALT count for C at the base pair location 304 is 1.The ALT count for C at base pair location 304 is not 2 because, as discussed above, in this example, base pair 302 at base pair location 304 in sequence listing 300(4) has a quality score below the predetermined quality score threshold and is not considered in the calculations made by sequencing system 102.

[048] If, in BLOCK 208, sequencing system 102 determines that the variant type is an insertion or deletion, method 200 may continue in BLOCK 212. Method 200 may include determining an aggregate count for each position (BLOCK 220). As described in relation to BLOCK 216 and BLOCK 218, sequencing system 102 may only count base pairs with a quality score above the predetermined threshold when determining the aggregate count for each position. Petition 870260072528, dated 07 / 21 / 2026, page 31 / 56 18 / 27

[049] Method 200 may include determining the ALT count (BLOCK 222). For a deletion, the ALT count may be determined for the location of P+1. For example, the ALT count may be the number of deletions with a deletion length of K at position P+1 of the CIGAR string. For an insertion, the ALT count may be the count of the number of reads with length L at the initial position P+1 of the CIGAR string and an alternate sequence match that corresponds to the base pair read at P+1.

[050] If, in BLOCK 208, sequencing system 102 determines that the variant type is a structural variant, method 200 can continue in BLOCK 226. Method 200 can then include determining a reference count (REF) (BLOCK 228). When determining the REF count, sequencing system 102 can only count reads with a quality score above the predetermined threshold. The structural variant can span an event boundary that starts at an event start in the genetic sequence and ends at an event end in the genetic sequence. Sequencing system 102 can determine the REF count as the number of reads that match in CIGAR along the event boundary.

[051] Method 200 may include determining an ALT count (BLOCK 230). When the variant type is a structural variant, sequencing system 102 may determine the ALT count as the occurrences of deletions, insertions, reference jumps, discrete cleavages, or coarse cleavages in CIGAR across the event boundary.

[052] Method 200 may include determining the aggregate count (BLOCK 232). Sequencing system 102 may sum the REF count and the ALT count to determine the aggregate count when the variant type is a structural variant.

[053] Method 200 may include determining sequence metrics Petition 870260072528, dated 07 / 21 / 2026, page 32 / 56 19 / 27 Genetics (BLOCK 234). Genetic sequence metrics may include determining an ALT frequency. The 102 sequencing system may determine the ALT frequency as the ALT count divided by the aggregate count for the position. In some implementations, the genetic sequence metric may include determining an intermediate, maximum, minimum, or average coverage depth for the sequence. The sequence metric may include determining a count of each nucleotide count and insertion and deletion counts for each base. Also with reference to Figure 3, the 102 sequencing system may determine the intermediate, maximum, or average coverage or read depth for each base pair 302 relative to each of the sequence listings 300. The 102 sequencing system may only count base pairs 302 that score above the predetermined threshold.In some implementations, the 102 sequencing system can identify counts per tape to identify the read tape. The 102 sequencing system can also identify clinically relevant variants by identifying alternative calls at the base pair location that occur with ALT frequency.

[054] In some implementations, method 200 may include sequencing system 102 that transmits genetic sequence metrics to a client device. For example, sequencing system 102 may transmit genetic sequencing metrics to a laptop computer or other user computing device. In some implementations, sequencing system 102 may run as a component of a user computing device (e.g., a laptop computer), and sequencing system 102 may render or display the genetic sequence metrics to the user.

[055] Figure 4 illustrates a block diagram of a computer system. Petition 870260072528, dated 07 / 21 / 2026, page 33 / 56 Example 400. The computing system or computing device 400 may include or be used to implement the system 100 or its components, such as the sequencing system 102. For example, the data analyzer 110, the analytical engine 112, the reporting engine 104, the filtering engine 108 may be components stored in main memory 415. The computing system 400 includes a bus 405 or other communication component for communicating information and a processor 410 or processing circuit coupled to the bus 405 for processing information. The computing system 400 may also include one or more processors 410 or processing circuits coupled to the bus for processing information.The computing system 400 also includes main memory 415, such as random access memory (RAM) or other dynamic storage device coupled to bus 405 to store information and instructions to be executed by processor 410. Main memory 415 may be or include data store 116. Main memory 415 may also be used to store position information, temporary variables, or other intermediate information during instruction execution by processor 410. The computing system 400 may additionally include read-only memory (ROM) 420 or other static storage device coupled to bus 405 to store static information and instructions for processor 410. A storage device 425, such as a solid-state device, magnetic disk, or optical disk, may be coupled to bus 405 to persistently store information and instructions.Storage device 425 may include or be part of data repository 116.

[056] The 400 computing system can be coupled via bus 405 to a 435 display, such as a liquid crystal display or a matrix display. Petition 870260072528, dated 07 / 21 / 2026, page 34 / 56 21 / 27 activates to display information to a user. An input device 430, such as a keyboard including alphanumeric keys and other keys, can be coupled to the bus 405 to communicate information and command selections to the processor 410. The input device 430 may include a touch-screen display 435. The input device 430 may also include a cursor control, such as a mouse, a command ball, or cursor direction keys to communicate direction information and command selections to the processor 410 and to control cursor movement on the display 435. The display 435 may be part of the sequencing system 102 or another component of Figure 1, for example.

[057] The processes, systems, and methods described in the present invention can be implemented by the computing system 400 in response to the processor 410 executing an array of instructions contained in the main memory 415. Such instructions can be read from the main memory 415 or from another computer-readable medium, such as the storage device 425. The execution of the array of instructions contained in the main memory 415 causes the computing system 400 to perform the illustrative processes described in the present invention. One or more processors in a multiprocessing array can also be employed to execute the instructions contained in the main memory 415. Wired circuit assemblies can be used in place of or in combination with software instructions along with the systems and methods described in the present invention.The systems and methods described in the present invention are not limited to any specific combination of hardware and software circuitry.

[058] Although an exemplary computing system has been described in Figure 4, the subject matter including the operations described in this descriptive report can be implemented in other types of circuit assemblies. Petition 870260072528, dated 07 / 21 / 2026, page 35 / 56 22 / 27 digital electronics, or in computer software, firmware or hardware, including the structures disclosed in this descriptive report and their structural equivalents or combinations thereof.

[059] The subject matter and operations described in this descriptive report may be implemented in a set of digital electronic circuits or in computer software, firmware, or hardware, including the structures disclosed in this descriptive report and their structural equivalents, or in combinations of one or more thereof. The subject matter described in this descriptive report may be implemented as one or more computer programs, for example, one or more computer program instruction circuits, encoded in one or more computer storage media for execution by or to control the operation of data processing devices.Alternatively or additionally, program instructions may be encoded in an artificially generated propagated signal, for example, a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to a suitable receiving apparatus for execution by a data processing apparatus. A computer storage medium may be or be included in a computer-readable storage device, a computer-readable storage substrate, a random-access or serial-access memory array or device, or a combination of one or more of the same. Although a computer storage medium is not a propagated signal, a computer storage medium may be a source or destination of computer program instructions encoded in an artificially generated propagated signal.Computer storage media can also be contained within one or more separate components or media (e.g., multiple CDs, disks, or other storage devices). Petition 870260072528, dated 07 / 21 / 2026, page 36 / 56 23 / 27 operations described in this descriptive report can be implemented as operations performed by a data processing device on data stored in one or more computer-readable storage devices or received from other sources.

[060] The terms “data processing system”, “computing device”, “component” or “data processing apparatus” encompass various apparatuses, devices and machines for processing data, including, by way of example, a programmable processor, a computer, a system on a chip, or multiples or combinations thereof. The apparatus may include a set of special-purpose logic circuits, for example, an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus may also include, in addition to hardware, code that creates an execution environment for the computer program in question, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of these.The appliance or execution environment can implement several different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures. The components of the system 100 may include or share one or more data processing appliances, systems, computing devices, or processors.

[061] A computer program (also known as a program, software, software application, application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and Petition 870260072528, dated 07 / 21 / 2026, page 37 / 56 24 / 27 can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program can correspond to a file in a file system. A computer program can be stored in a portion of a file that contains other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subprograms, or code portions). A computer program can be deployed to run on one computer or on multiple computers that are located in one location or distributed across multiple locations and interconnected by a communication network.

[062] The processes and logical flows described in this descriptive report can be performed by one or more programmable processors that execute one or more computer programs (e.g., components of the 102 sequencing system) to perform actions when operating on input data and generating output. The processes and logical flows can also be performed by, and devices can also be implemented as, a set of special-purpose logic circuits, for example, an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit).Suitable devices for storing computer program instructions and data include all forms of non-volatile memory, memory media and devices, including, by way of example, semiconductor memory devices, for example, EPROM, EEPROM and flash memory devices; magnetic disks, for example, internal hard disks or removable disks; magneto-optical disks; and CD-ROM type disks. Petition 870260072528, dated 07 / 21 / 2026, page 38 / 56 25 / 27 and DVD-ROM. The processor and memory may be supplemented by, or incorporated into, a set of special-purpose logic circuits.

[063] Although the operations are depicted in the drawings in a particular order, it is not required that such operations be performed in the particular order shown or in sequential order, and it is not required that all illustrated operations be performed. The actions described in the present invention may be performed in a different order.

[064] Separating various system components does not require separation in all implementations, and the program components described can be included in a single hardware or software product.

[065] Having now described some illustrative implementations, it is evident that the foregoing is illustrative and not limiting, having been presented by way of example. In particular, although many of the examples presented in the present invention involve specific combinations of method acts or system elements, those acts and those elements can be combined in other ways to achieve the same objective. It is not intended that acts, elements and features discussed together with one implementation are excluded from a similar role in other implementations or implementations.

[066] The phraseology and terminology used in the present invention are for descriptive purposes only and should not be considered limiting. The use of including, comprising, having, containing, involving, “characterized by, characterized by the fact that” and variations thereof in the present invention shall encompass the items listed below, equivalents thereof and additional terms, as well as alternative implementations consisting exclusively of the items listed below. In one implementation, the systems and methods described in the present invention Petition 870260072528, dated 07 / 21 / 2026, page 39 / 56 26 / 27 consist of one, each combination of more than one, or all of the elements, acts, or components described.

[067] As used in the present invention, the terms about and substantially will be understood by those skilled in the art and will vary to some extent depending on the context in which they are used. If there are uses of the term that are not clear to those of common skill in the art given the context in which it is used, about will mean more or less 10% of the particular term.

[068] Any references to implementations or elements or acts of the systems and methods referred to in this invention in the singular may encompass implementations including a plurality of such elements, and any references in the plural to any implementation or element or act in the present invention may also encompass implementations including only a single element. It is not intended that references in the singular or plural form limit currently disclosed systems or methods, their components, acts or elements to configurations in the singular or plural. References to any act or element based on any information, the act or element may include implementations in which the act or element is based at least in part on any information, act or element.

[069] Any implementation disclosed in the present invention may be combined with any other implementation or embodiment, and references to an implementation, some implementations, an implementation, or the like are not necessarily mutually exclusive and are intended to indicate that a particular feature, structure, or characteristic described together with the implementation may be included in at least one implementation or embodiment. Such terms as used in the present invention do not necessarily all refer to the same Petition 870260072528, dated 07 / 21 / 2026, pp. 40 / 56 27 / 27 implementation. Any implementation may be combined without any other implementation, inclusively or exclusively, in any manner consistent with the aspects and implementations disclosed in the present invention.

[070] The indefinite articles “a” and “an” as used in the present invention, in this descriptive report and in the claims, unless clearly indicated otherwise, should be understood as meaning “at least one”.

[071] References to “or” can be interpreted as inclusive so that any terms using “or” can indicate any one of, more than one of, or all of the terms described. For example, a reference to “at least one of A and B” can include only A, only B, as well as both A and B. Such references used in conjunction with “including” or other open terminology can include additional terms.

[072] Where technical features in the drawings, detailed description or any claim are followed by reference signs, the reference signs have been included to enhance the intelligibility of the drawings, detailed description and claims. Consequently, neither the reference signs nor their absence have any limiting effect on the scope of any claim element.

[073] The systems and methods described in the present invention can be incorporated in other specific ways without departing from their characteristics. The aforementioned implementations are illustrative rather than limiting to the systems and methods described. The scope of the systems and methods described in the present invention is thus indicated by the appended claims, rather than the aforementioned description, and changes that are within the meaning and range of equivalence of the claims are encompassed in the present invention. Petition 870260072528, dated 07 / 21 / 2026, pp. 41 / 56

Claims

1 / 5 CLAIMS 1. A method for filtering sequencing data characterized in that it comprises: receiving (202), by a data processing system, data comprising a plurality of genetic sequences, wherein each of the plurality of genetic sequences comprises a chromosome indication, a position indication, a base value and a quality score; selecting (204), by the data processing system, a subset of the plurality of genetic sequences, wherein each subset of the plurality of genetic sequences has the same chromosome indication; filtering (206), by the data processing system, from the subset of the plurality of genetic sequences, the genetic sequences comprising base values ​​that have an associated quality score above a predetermined threshold;To determine, using the data processing system, an aggregate count for each position in the filtered genetic sequences; to determine, using the data processing system, an alternative base count for each position in the filtered genetic sequences; and to generate, using the data processing system, an identifier for a variant of the genetic sequence, responsive to a relationship between the alternative base count for each position and the aggregate count for each position that exceeds a limit.

2. A method according to claim 1, characterized in that it further comprises determining an alternative count for a deletion sequence in the filtered genetic sequences.

3. Method, according to claim 2, characterized by the fact that the deletion sequence begins at an index adjacent to the position.

4. Method according to claim 1, characterized in that it further comprises determining an alternative count for an insertion sequence in the filtered genetic sequences.

5. Method according to claim 4, characterized in that the determination of the alternative count for the insertion sequence further comprises identifying an alternative sequence match.

6. Method according to claim 1, characterized in that it further comprises identifying a structural variant in a plurality of genetic sequences.

7. Method according to claim 6, characterized in that it further comprises determining the alternative base count based on the structural variant identified in the plurality of genetic sequences.

8. Method according to claim 6, characterized in that the determination of the aggregate count further comprises counting a match in each of the filtered genetic sequences with the CIGAR chain.

9. Method according to claim 6, characterized in that the determination of the aggregate count further comprises counting a deletion, insertion, reference jump, soft clip or hard clip in each of the subsets of the plurality of genetic sequences.

10. Method, according to claim 1, characterized in that it further comprises calculating at least one of an average read coverage, a high read coverage or a maximum read coverage for the plurality of genetic sequences based on the aggregate count and the alternate base count.

11. Method, according to claim 1, characterized in that it further comprises calculating a reading strip for the plurality of genetic sequences based on aggregate counting and alternative base counting.

12. System for filtering sequencing data characterized in that it comprises: a processor (410) communicating with a memory device (116, 415, 425), the processor (410) executing a data analyzer (110) and a filtering mechanism (108); wherein the data analyzer (110) is configured to: receive, from the memory device (116, 415, 425), data comprising a plurality of genetic sequences, wherein each of the plurality of genetic sequences comprises a chromosome indication, a position indication, a base value and a quality score, and select a subset of the plurality of genetic sequences, wherein each subset of the plurality of genetic sequences has the same chromosome indication;and wherein the filtering mechanism (108) is configured to: filter, from the subset of the plurality of genetic sequences, the genetic sequences comprising base values ​​that have an associated quality score above a predetermined threshold, determine an aggregate count for each position of the filtered genetic sequences, determine an alternative base count for each position of the filtered genetic sequences, and generate a variant identifier for the genetic sequence, responsive to a ratio between the alternative base count for each position to the aggregate count for each position that exceeds a threshold.

13. System according to claim 12, characterized in that the filtering mechanism (108) is further configured to determine an alternative count for a deletion sequence in the filtered genetic sequences.

14. System according to claim 12, characterized in that the filtering mechanism (108) is further configured to determine an alternative count for an insertion sequence in the filtered genetic sequences.

15. System according to claim 14, characterized in that the filtering mechanism (108) is additionally configured to determine the alternative count for the insertion sequence when identifying an alternative sequence match.

16. System according to claim 12, characterized in that the filtering mechanism (108) is further configured to identify a structural variant in the plurality of genetic sequences.

17. System according to claim 16, characterized in that the filtering mechanism (108) is further configured to determine the aggregate by counting a match in each of the filtered genetic sequences with a CIGAR chain.

18. System according to claim 16, characterized in that the filtering mechanism (108) is further configured to determine the aggregate count by counting a deletion, insertion, reference jump, soft clip or hard clip in each of the subsets of the plurality of genetic sequences.

19. System according to claim 12, characterized in that Petition 870260072528, dated 07 / 21 / 2026, page 45 / 56 5 / 5 that the filtering mechanism (108) is further configured to calculate at least one of an average read coverage, a high read coverage or a maximum read coverage for the plurality of genetic sequences based on aggregate counting and alternative base counting.

20. System according to claim 12, characterized in that the filtering mechanism (108) is additionally configured to calculate a reading strip for the plurality of genetic sequences based on aggregate counting and alternative base counting.

21. Invention of a product, process, system, kit, means or use, characterized by the fact that it comprises one or more elements described in this patent application. Petition 870260072528, dated 07 / 21 / 2026, pp. 46 / 56