Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

72 results about "Base calling" patented technology

Base calling is the process of assigning nucleobases to chromatogram peaks. One computer program for accomplishing this job is Phred base-calling, which is a widely used basecalling software program by both academic and commercial DNA sequencing laboratories because of its high base calling accuracy.

Method for improving nucleic acid sequencing quality by eliminating nucleic acids with deaminated bases from library and method for sequencing in which complexes of primers, polymerases and labelled probes are bound to concatemers

The present disclosure provides methods for reducing sequencing errors comprising one or any combination of: (i) removing deaminated bases in any nucleic acid molecule throughout a library preparation workflow which includes immobilised splints which bind to the library, the use of a compaction oligonucleotide, optionally with an intervening sequence, formation of closed circular nucleic acids, creating gaps using glycosylase and lyase activities at positions with deaminated bases. The library may be sequenced using pairwise sequencing, e.g. with dark sequencing and / or sequencing using a multivalent labelled probe for the formation of an avidity molecule and soluble primer and polymerase. Method for sequencing concatemers in which the concatermers are contacted with polymerases, soluble primers and a multivalent labelled molecule which forms a complex with the polymerase. Detecting polymerase position and nucleobase bound to the polymerase in the complex. These methods generate higher quality base calls during downstream sequencing workflows.
Owner:ELEMENT BIOSCIENCES INC

High throughput inramolecular consensus reads

Described herein are methods, systems, and apparatuses for determining a partial consensus sequence of a double-stranded nucleic acid molecule. Both strands of the nucleic acid molecule may be sequenced to generate a sequence of base calls. The base calls may be used to identify sets of concordant and discordant positions. A partial consensus sequence may be generated by using concordant values derived from the concordant positions and discordant values derived from the discordant positions. Also described herein are methods, systems, and apparatuses for determining a consensus sequence of a double-stranded nucleic acid molecule. Both strands of the nucleic acid molecule may be sequenced to generate base calls and quality scores corresponding to each strand. Concordant and discordant positions may be identified using the sequences of base calls. Discordant positions may also use the quality scores and weights. A consensus sequence may be determined using the concordant and discordant positions.
Owner:ROCHE SEQUENCING SOLUTIONS INC

Primary analysis in next generation sequencing

ActiveUS12505571B2Image analysisRecognition of DNA microarray patternFlow cellBase calling
Image data analysis, and particularly identifying cluster or polony locations for performing base-calling in a digital image of a flow cell during DNA sequencing is described. A method may include generating a first plurality of flow cell images of a cellular sample immobilized on a support by conducting one or more cycles of sequencing reactions. The cellular sample may include a plurality of concatemer molecules therewithin. For the first plurality of flow cell image, pixel intensities, and a respective color purity of each of the pixel intensities may be determined. A base calling template may include base calling locations based on the pixel intensities and the respective color purity of the pixel intensities. The base calling template may be for registering a second plurality of flow cell images of the support in one or more subsequent cycles of the one or more cycles.
Owner:ELEMENT BIOSCIENCES INC

Nanopore sequencing base calling

Disclosed herein are systems and methods for nanopore sequencing basecalling. In one embodiment, the method can include: receiving raw nanopore sequencing data comprising a plurality of continuous data acquisition (DAC) values corresponding to a biomolecule; normalizing the raw nanopore sequencing data to generate normalized nanopore sequencing data comprising a plurality of normalized DAC values; generating, using a first neural network (NN) and a normalized DAC value of the plurality of normalized DAC values, a vector of transformed probability values, segmenting the plurality of normalized DAC values into a plurality of discrete events; generating, using a second neural network and the event vector, an element determination of the biomolecule.
Owner:RGT UNIV OF CALIFORNIA

Linked ligation

The invention generally relates to capturing, amplifying, and sequencing nucleic acids. In certain embodiments, copies of the sense and antisense strands of a duplex template nucleic acid are captured using linked capture probes and multiple binding and extension steps to improve specificity over traditional single binding target capture techniques. Methods of seeding sequencing clusters with sense and antisense strands of a target nucleic acid are also disclosed including identifying the strands using sense-specific barcodes and confirming base calls using two sense-specific sequencing reads. Linked adapters may be used to increase adapter ligation selectively or efficiency and yield.
Owner:NCAN GENOMICS INC

Methods for detecting and suppressing alignment errors caused by fusion events

Methods and systems for producing a filtered read sequence information data set by identifying one or more split sequence reads in a set of test sequence reads obtained from cell-free nucleic acid (cfNA) in a biological sample obtained from a subject, wherein each split sequence read comprises at least one breakpoint; and, suppressing, in the set of test sequence reads, (i) at least a portion of one or more of the split sequence reads and / or at least a portion of one or more of the test sequence reads that comprise at least one sequence variant within a selected number of nucleotides from a given breakpoint, thereby producing the filtered sequence information data set, or, (ii) one or more base calls of the split sequence reads and / or one or more base calls of the test sequence reads that comprise at least one sequence variant within a selected number of nucleotides from a given breakpoint, thereby producing the filtered sequence information data set.
Owner:GUARDANT HEALTH INC

Systems and methods for sequencing image analysis

The technology disclosed relates to equalizer-based intensity correction for base calling. In particular, the technology disclosed relates to accessing an image whose pixels depict intensity emissions from a target cluster and intensity emissions from additional adjacent clusters, selecting a lookup table that contains pixel coefficients that are configured to increase a signal-to-noise ratio, applying the pixel coefficients to intensity values of the pixels in the image to produce an output, and base calling the target cluster based on the output.
Owner:ILLUMINA INC

Phasing correction

Methods determine corrected image data acquired by a nucleic acid sequencer during a cycle. Such methods may: (a) obtain an image of a substrate including a plurality of sites where nucleic acid bases are read; (b) measure color values of the plurality of sites from the image of the substrate; (c) store the color values in a processor buffer; (d) retrieve partially phase-corrected color values of the plurality of sites, where the partially phase-corrected color values were stored in the sequencer's memory during an immediately preceding base calling cycle; (e) determine a prephasing correction; and (f) determine the corrected color values. In implementations, these operations are all performed during a single base calling cycle. In embodiments, the methods additionally include using the corrected color values to make base calls for the plurality of sites.
Owner:ILLUMINA INC

Three-dimensional base calling in next generation sequencing analysis

Disclosed herein are system, apparatus, method, and / or computer program product embodiments, and / or combinations and sub-combinations thereof which enables 3D base calling using flow cell images of samples such as in situ cells or tissue to ensure accurate base calling and sequencing analysis of 3D samples. Embodiments of the methods, systems, and media for 3D base calling of flow cell images includes image intensity, location, size, and / or of clusters or polonies to be relied on for accurate base calling.
Owner:ELEMENT BIOSCIENCES INC

Linked ligation

The invention generally relates to capturing, amplifying, and sequencing nucleic acids. In certain embodiments, copies of the sense and antisense strands of a duplex template nucleic acid are captured using linked capture probes and multiple binding and extension steps to improve specificity over traditional single binding target capture techniques. Methods of seeding sequencing clusters with sense and antisense strands of a target nucleic acid are also disclosed including identifying the strands using sense-specific barcodes and confirming base calls using two sense-specific sequencing reads. Linked adapters may be used to increase adapter ligation selectively or efficiency and yield.
Owner:NCAN GENOMICS INC

Efficient base matching method based on software and hardware collaboration

The invention relates to quantum key distribution, in particular to an efficient base matching method based on software and hardware collaboration, which comprises the following steps: temporarily storing massive base matching information in DDR (Double Data Rate) through collaboration among modules such as software at a transmitting end and a receiving end, an FPGA (Field Programmable Gate Array) and the DDR, firstly carrying out hardware base matching on the FPGA according to a detection result, and obviously reducing the data volume preprocessed by the FPGA; the method comprises the following steps of: preprocessing data of a hardware module, transmitting the preprocessed data to software, and finishing base matching by the software on the basis, thereby effectively reducing data interaction pressure between the hardware module and the software, performing data compression on high-redundancy information at a receiving end and a transmitting end, reducing data communication pressure, and providing technical support for high-speed quantum key distribution and long-distance transmission. The working frequency and the maximum transmission distance of the quantum key distribution system are possible to be improved; according to the technical scheme provided by the invention, the processing pressure of massive base data of the quantum key distribution system under the conditions of high-speed working frequency and long-distance transmission can be effectively overcome.
Owner:ANHUI QASKY QUANTUM SCI & TECH CO LTD

Systems and methods for nucleic acid mismatch error detection

A double-stranded template nucleic acid molecule may include an adaptor that includes a mismatched moiety. The mismatch moiety may include a first mismatch sequence in a first chain and a second mismatch sequence in a second chain that is not complementary to the first mismatch sequence. When the double-stranded template nucleic acid molecule is amplified to produce a cluster of amplified strands and the cluster is sequenced to produce a sequenced read, the double-stranded template nucleic acid molecule is amplified to produce the sequenced read. A portion of the sequencing read corresponding to the mismatch portion can be analyzed to determine whether the cluster of amplified strands originates only from one strand or both strands of the double-stranded template nucleic acid molecule. A sequencing signal between two chain derivatives in the cluster, a base recognition, or an inconsistency at one or more loci in a sequencing read can be used to correct for sequencing errors and improve sequencing accuracy.
Owner:ULTIMA GENOMICS INC

Systems and methods for sequencing image analysis

A system, a method and a non-transitory computer readable storage medium for base calling are described. The base calling method includes processing through a neural network first image data comprising images of clusters and their surrounding background captured by a sequencing system for one or more sequencing cycles of a sequencing run. The base calling method further includes producing a base call for one or more of the clusters of the one or more sequencing cycles of the sequencing run.
Owner:ILLUMINA INC

Base recognition method and device for nanopore sequencing and storage medium

The invention relates to the technical field of bioinformatics, and discloses a base recognition method and device for nanopore sequencing and a storage medium. The method comprises the following steps: inputting current data obtained by nanopore sequencing into a pre-trained convolutional neural network model to obtain current feature data; inputting the current feature data into a pre-trained encoder model to determine an attention score of the current feature data, and determining a semantic vector based on the attention score; the encoder model has a multi-head self-attention structure to calculate an attention score, and the multi-head self-attention structure fuses the rotation position codes to fuse the position dependency relationship between different positions in the attention score; inputting a semantic vector output by the encoder model into a CTC decoder model to obtain a base recognition result; the method not only can be applied to identification of DNA molecules, but also can be applied to identification of RNA molecules, and the universality of the base identification method can be improved while the accuracy of base identification is improved.
Owner:BEIJING POLYSEQ BIOTECH CO LTD +2

Machine learning model for detecting air bubbles within a nucleotide sample slide for sequencing

Methods, systems, and non-transitory computer-readable media are disclosed for accurately and efficiently detecting when a bubble is affecting a nucleic acid sequencing run based on data captured (or derived from) during base callings during the sequencing run. Specifically, in one or more embodiments, the disclosed systems receive data identifying nucleobase callings and data identifying quality indicators for the nucleobase callings during a sequencing cycle. Based on a particular nucleobase calling and a threshold marking for the quality indicator, the disclosed systems utilize a machine learning model to detect the presence of a bubble in a nucleotide sample slide. In addition to simply detecting the presence of a bubble, the disclosed systems can also classify different detected bubbles, such as air bubbles, oil bubbles, or ghost bubbles, or other outputs during sequencing. By utilizing the call data and quality indicators, the disclosed systems can use readily available sequencing data in a platform-agnostic approach to detect bubbles using a uniquely trained machine learning model.
Owner:ILLUMINA INC +1

Machine-learning model for recalibrating nucleotide-base calls

This disclosure describes methods, non-transitory computer readable media, and systems that can utilize a machine learning model to recalibrate nucleotide-base calls (e.g., variant calls) of a call-generation model. For instance, the disclosed systems can train and utilize a call-recalibration-machine-learning model to generate a set of predicted variant-call classifications based on sequencing metrics associated with a sample nucleotide sequence. Leveraging the set of variant-call classifications, the disclosed systems can further update or modify nucleotide-base calls (e.g., variant calls) corresponding to genomic coordinates. Indeed, the disclosed systems can generate an initial nucleotide-base call based on sequencing metrics for nucleotide reads of a sample sequence utilizing a call-generation model and further utilize a call-recalibration-machine-learning model to generate classification predictions for updating or recalibrating the initial nucleotide-base call from a subset of the same sequencing metrics or other sequencing metrics.
Owner:ILLUMINA INC

Neural network parameter quantization for base calling

A method of quantizing parameters of a neural network includes grouping a plurality of parameters of a neural network in a plurality of groups. Each group of the plurality of groups includes corresponding two or more parameters of the plurality of parameters. In an example, for each group, a corresponding quantization format is selected from a plurality of available quantization formats, such that a first quantization format selected for at least a first group is different from a second quantization format selected for at least a second group. For each group, individual parameters within the corresponding group are quantized using the quantization format selected for the corresponding group. The quantized parameters of the plurality of groups are stored in a memory.
Owner:ILLUMINA INC

Blind equalization systems for base calling applications

PCT designated stageWO2025240924A1BiostatisticsSequence analysisAlgorithmBase calling
This disclosure describes embodiments of methods, systems, and non-transitory computer readable media that can quickly and accurately determine equalizer coefficients for an equalizer based on estimated point-spread-function values and estimated noise values derived from expected response signals of oligonucleotide clusters. For example, the disclosed systems can receive signal values for expected response signals from one or more clusters of oligonucleotides incorporating labeled nucleobases. Based on the signal values, the disclosed systems can determine estimated point-spread-function values for one or more such clusters of oligonucleotides and estimated noise values within a channel. From the estimated point-spread-function values and the estimated noise values, the disclosed systems can determine equalizer coefficients that compensate for the estimated point-spread-function values and the estimated noise values. The disclosed systems can further determine a base call for one or more such clusters of oligonucleotides by utilizing the equalizer coefficients.
Owner:ILLUMINA INC

Methods and apparatus for determining sequencing context parameters for nanopore sequencing data

PendingCN122637890ABase callingData mining
Embodiments of the present specification provide methods and apparatuses for determining sequencing environment parameters of nanopore sequencing data. In determining the sequencing environment parameters, base quality score features are extracted from the nanopore sequencing data, the extracted base quality score features including global features of base quality scores. Subsequently, sequencing environment parameters of the nanopore sequencing data are determined according to the base quality score features, the determined sequencing environment parameters including a sequencing chip type and a base calling software configuration, the base calling software configuration including a base calling software type and a base calling software version. Since different sequencing environment parameters produce base quality score patterns with different base quality score features, the sequencing environment parameters of the nanopore sequencing data can be deduced by analyzing these patterns, thereby providing effective guidance for downstream analysis of the nanopore sequencing data.
Owner:XIN HUA HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE +1

Linked ligation

The invention generally relates to capturing, amplifying, and sequencing nucleic acids. In certain embodiments, copies of the sense and antisense strands of a duplex template nucleic acid are captured using linked capture probes and multiple binding and extension steps to improve specificity over traditional single binding target capture techniques. Methods of seeding sequencing clusters with sense and antisense strands of a target nucleic acid are also disclosed including identifying the strands using sense-specific barcodes and confirming base calls using two sense-specific sequencing reads. Linked adapters may be used to increase adapter ligation selectively or efficiency and yield.
Owner:NCAN GENOMICS INC

Base calling method and apparatus, electronic device, and storage medium

The present application discloses a base calling method and apparatus, an electronic device, and a storage medium. The method for base calling includes: determining a first sequencing information based on intensity features of a first spot in images corresponding to a consecutive cycles of base extension reactions including a designated cycle of base extension reaction, where a is a natural number greater than or equal to 1; and determining a base type of the designated cycle of base extension reaction based on the first sequencing information and a basecall model, where the basecall model is determined based on a second sequencing information corresponding to the a consecutive cycles of base extension reactions in a training sample and base type of at least one cycle of base extension reaction in the a consecutive cycles of base extension reactions, and the second sequencing information includes the first sequencing information. The technical schemes of the examples of the present application improve the accuracy and efficiency of base calling.
Owner:GENEMIND BIOSCIENCES CO LTD

Primary analysis in next generation sequencing

ActiveUS12469162B2Image analysisRecognition of DNA microarray patternBase callingData profiling
Image data analysis, particularly identifying cluster locations for performing base-calling in a digital flow cell image during DNA sequencing, is described. Each nucleic acid template molecule immobilized on a support may include an insert sequence and a sample index sequence. The sample index sequence may include a k-mer sequence. A sequencing system may conduct k cycles of sequencing reactions of the k-mer sequence before conducting one or more cycles of the insert sequence sequencing reactions and generate a first plurality of flow cell images. Pixel intensities may be determined for pixels of the first plurality of flow cell images. A base calling template may be determined and include base calling locations based on the pixel intensities and respective color purities of the pixel intensities. The base calling template may register a second plurality of flow cell images of the support in one or more cycles subsequent to the k cycles.
Owner:ELEMENT BIOSCIENCES INC

Nanopore base calling method based on libtorch and c++

The application discloses a nanopore base recognition method based on Libtorch and C++, which comprises the following steps: obtaining nanopore sequencing data of a target sample and preprocessing the data, wherein the preprocessing comprises normalization and overlapping segmentation; delivering the preprocessed data to a model inference module based on a convolutional neural network and a bidirectional long short-term memory network in an asynchronous mode, performing GPU inference on each time step, and outputting base sequence weights of each time step; decoding the base sequence weights of each time step output by the model inference module through a CTC decoding module based on a greedy search and a prefix beam search; reassembling the decoded result; and outputting the result. The model inference module is optimized, efficient data loading and multi-GPU inference are realized, the inference time is significantly shortened, the memory occupation is effectively reduced, the memory consumption is reduced while the high accuracy is maintained, and the running speed and throughput of base recognition are accelerated.
Owner:ZHONGSHAN OPHTHALMIC CENT SUN YAT SEN UNIV

Multi-pair multi-base interpretation based on artificial intelligence

The technology disclosed herein relates to base interpretation based on artificial intelligence. The disclosed technology relates to accessing the progress of a set of per-cycle analyte channels generated for a sequencing cycle of a sequencing run; a window of a set of analyte channels per cycle in progress is processed by a neural network based base interpreter (NNBC) for a window of a sequencing cycle of a sequencing run such that the NNBC processes a subject window of the set of analyte channels per cycle in progress for a subject window of a sequencing cycle of the sequencing run, and using the NNBC to generate a temporary base interpretation prediction for three or more sequencing cycles in a target window of a sequencing cycle from a plurality of windows at which a particular sequencing cycle appears at different positions, to generate a temporary base interpretation prediction for the particular sequencing cycle; and determining a base interpretation for the particular sequencing cycle based on the plurality of base interpretation predictions.
Owner:ILLUMINA INC

Cluster intensity variation correction and base calling

The technology disclosed corrects inter-cluster intensity profile variation for improved base calling on a cluster-by-cluster basis. The technology disclosed accesses current intensity data and historic intensity data of a target cluster, where the current intensity data is for a current sequencing cycle and the historic intensity data is for one or more preceding sequencing cycles. A first accumulated intensity correction parameter is determined by accumulating distribution intensities measured for the target cluster at the current and preceding sequencing cycles. A second accumulated intensity correction parameter is determined by accumulating intensity errors measured for the target cluster at the current and preceding sequencing cycles. Based on the first and second accumulated intensity correction parameters, next intensity data for a next sequencing cycle is corrected to generate corrected next intensity data, which is used to base call the target cluster at the next sequencing cycle.
Owner:ILLUMINA INC

Hardware execution and acceleration of artificial intelligence-based base caller

A system for analysis of base call sensor output has memory accessible by the runtime program storing tile data including sensor data for a tile from sensing cycles of a base calling operation. A neural network processor having access to the memory is configured to execute runs of a neural network using trained parameters to produce classification data for sensing cycles. A run of the neural network operates on a sequence of N arrays of tile data from respective sensing cycles of N sensing cycles, including a subject cycle, to produce the classification data for the subject cycle. Data flow logic moves tile data and the trained parameters from the memory to the neural network processor for runs of the neural network using input units including data for spatially aligned patches of the N arrays from respective sensing cycles of N sensing cycles.
Owner:ILLUMINA INC