Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

21 results about "Short read" patented technology

A method for detecting microbial structural variations, mutations and indels based on next generation sequencing

ActiveCN120954507BBiostatisticsSequence analysisInsertion deletionMicroorganism
The present application relates to the technical field of bioinformatics, in particular to a method for detecting microbial structural variation, mutation and insertion deletion based on second-generation sequencing. The present application is based on the short read Illumina second-generation sequencing strategy, through testing, using the current mainstream biological tools and combining the self-developed program, the automatic analysis of microbial metagenome structural variation (SV), single nucleotide variation (SNV / mutation) and small fragment insertion deletion (Indel) is realized, and the problems of low sensitivity, high false positive rate and low efficiency in the prior art for detecting variation under the metagenome background are significantly solved.
Owner:ACADEMY OF MILITARY MEDICAL SCIENCES

Method for identifying large insertions in target genomic regions and uses thereof

ActiveCN121260248BProteomicsGenomicsGenomic sequencingSequence Insertions
The application discloses a method for identifying large fragment sequence insertion of a target genomic region and application thereof, and belongs to the technical field of bioinformatics. In view of the problem that large fragment insertion variation is difficult to be accurately recognized in clinical metagenomic sequencing due to short sequencing read length, insufficient coverage and other factors, the application proposes to construct a reference sequence which can represent the insertion variation by means of manual construction, and to realize efficient identification of the insertion event of the target genomic region by combining a short read-based fast alignment process. The method overcomes the dependence of existing structural variation detection tools on high sequencing depth and long read length, has the advantages of fast identification speed, high sensitivity and high accuracy, and is suitable for rapid screening of large fragment insertion related to drug resistance mechanism in clinical samples. Meanwhile, the method can be popularized for insertion variation analysis of other pathogen drug resistance related genes or genomic regions, and has a good clinical application prospect.
Owner:BEIJING GOLDEN KEY MEDICAL LAB CO LTD +2

Preparation of long read nucleic acid libraries

Some embodiments of the methods and compositions provided herein relate to obtaining long read information from short reads of a target nucleic acid. Some embodiments include steps to selectively generate, mark, and amplify long nucleic acid fragments. Some embodiments include enriching for certain sequences in the long fragments with selection probes directed to certain challenging medically relevant genes (CMRG). Some embodiments also include fragmenting the long nucleic acid fragments into shorter fragments for sequencing, and informatically reconstructing a sequence of the target nucleic acid.
Owner:ILLUMINA INC

A method and system for coronavirus transcriptome identification analysis

The application discloses a coronavirus transcriptome identification and analysis method, which comprises the following steps: using bwa software to perform sequence alignment on coronavirus second-generation sequencing original fastq files obtained from a short read archive SRA and NCBI reference genomes, to generate BAM files; calling and filtering single nucleotide polymorphisms SNPs through bcftools, and annotating the SNPs by using a vcf-annotator; performing CIGAR string analysis and breakpoint identification operations on the BAM files, to perform sgRNA identification; generating an expression matrix from the sgRNA identification result and performing transcriptome analysis. The application further discloses a coronavirus transcriptome identification and analysis system, which comprises a storage and a processor; the storage is provided with a computer program, and when the computer program is executed by the processor, the method described above is realized. The coronavirus transcriptome identification and analysis method and system provide a new cognitive means for coronavirus biology, and provide valuable resources for the development of future treatment methods.
Owner:SHANGHAI INST OF IMMUNOLOGY

A method, primer set and kit thereof for adenovirus whole genome sequencing

The application discloses a method for adenovirus whole genome sequencing, a primer group and a kit thereof, and belongs to the technical field of DNA detection.The method for adenovirus whole genome sequencing comprises the following steps: S1, first round amplification is carried out by using first round amplification primers with sequences shown in SEQ ID NO.1 to SEQ ID NO.93 in Table 1; and S2, second round amplification is carried out by using second round amplification primers with sequences shown in SEQ ID NO.94 to SEQ ID NO.285 in Table 2.The average length of data sequenced by the method is 0.8 kb, compared with short read length of a second-generation platform, the uniformity of adenovirus coverage can be greatly improved, meanwhile, the adenovirus types are comprehensively covered, and the method can be widely promoted.
Owner:PEKING UNION MEDICAL COLLEGE HOSPITAL +1

Preparation of long read nucleic acid libraries

PendingUS20250327064A1Microbiological testing/measurementDNA preparationMajor histocompatibilityShort read
Some embodiments of the methods and compositions provided herein relate to obtaining long read information from short reads of a target nucleic acid. Some embodiments include steps to selectively generate, mark, and amplify long nucleic acid fragments. Some embodiments include enriching for certain sequences in the long fragments with selection probes directed to major histocompatibility complex (MHC) genes. Some embodiments also include fragmenting the long nucleic acid fragments into shorter fragments for sequencing, and informatically reconstructing a sequence of the target nucleic acid.
Owner:ILLUMINA INC

Method for identifying large-fragment sequence insertion in target genome region and application of method

ActiveCN121260248AProteomicsGenomicsGenomic sequencingSequence Insertions
The invention discloses a method for identifying large-fragment sequence insertion in a target genome region and application of the method, and belongs to the technical field of bioinformatics. In order to solve the problem that large-fragment insertion variation is difficult to accurately identify due to factors such as short sequencing reading length and insufficient coverage in clinical metagenome sequencing, the invention proposes that a reference sequence capable of representing the insertion variation is artificially constructed and is combined with a rapid comparison process based on short reads to realize efficient identification of an insertion event in a target genome region. The method overcomes the dependence of an existing structure variation detection tool on high sequencing depth and long reading length, has the advantages of high identification speed, high sensitivity, high accuracy and the like, and is suitable for rapid screening of large fragment insertion related to a drug resistance mechanism in a clinical sample; meanwhile, the method can be popularized and applied to insertion variation analysis of other pathogen drug resistance related genes or genome areas, and has a good clinical application prospect.
Owner:BEIJING GOLDEN KEY MEDICAL LAB CO LTD +2

Integrating variant calls from multiple sequencing pipelines using a machine learning architecture

The present disclosure describes methods, non-transitory computer-readable media, and systems that can generate genotype calls from a combined pipeline for processing nucleotide reads from multiple read types / sources for robust and accurate genotype calling. For example, the disclosed systems can train and / or utilize a genotype call integration machine learning model to generate predictions of genotype calls based on data associated with a first type of nucleotide read (e.g., short reads) and a second type of nucleotide read (e.g., long reads). As disclosed, the disclosed systems can determine sequencing metrics and utilize a genotype call integration machine learning model to generate predictions (e.g., genotype probabilities, variant call classifications) for generating output genotype calls based on the sequencing metrics. The disclosed systems can utilize multiple such genotype call integration machine learning models to generate genotype calls for different variants, such as SNPs (Single Nucleotide Polymorphisms) and indels, with the genotype call integration machine learning model generating different predictions for each variant.
Owner:ILLUMINA INC

Method and system for haplotype level transposon insertion identification based on physical typing

ActiveCN120526846BData visualisationProteomicsGeneticsShort read
The application discloses a haplotype level transposon insertion identification method and system based on physical typing, and the method comprises the following steps: respectively aligning sequencing data to two sets of haplotype genomes, obtaining alignment scores and mismatch numbers after the sequencing data are aligned to the two sets of haplotype genomes; for single-end long read sequencing data, distributing the sequencing sequence to a haplotype with a larger alignment score; for double-end short read sequencing data, adding the alignment scores and the mismatch numbers of the double-end short reads as the alignment score and the mismatch number of each pair of sequences, and distributing the sequencing sequence to a haplotype with a larger alignment score; based on the physical typing result, separating the alignment files of the two sets of haplotype genomes, and identifying the transposon insertion at the haplotype level based on the alignment files. The application identifies the transposon insertion at the haplotype level, can more clearly distinguish the insertion differences of the transposon in the two sets of haplotypes, and improves the resolution of identification.
Owner:HUAZHONG AGRI UNIV

Aspergillus telomere to telomere genome assembly methods, apparatuses, devices, and storage media

PendingCN122392622AGenomic sequencingContig
The application discloses an aspergillus telomere-to-telomere genome assembly method, device, equipment and storage medium. The method comprises the following steps: using a plurality of sequencing sequence assembly tools to assemble target aspergillus long read genome sequencing data from scratch to obtain a first assembled genome; selecting a first assembled genome meeting a preset condition as an initial assembled genome; integrating other first assembled genomes to fill gaps between repeat regions of the initial assembled genome to obtain a second assembled genome; aligning the long read genome sequencing data to the second assembled genome, identifying abnormal coverage regions and correcting sequences to obtain a third assembled genome; aligning a reference genome to the third assembled genome, connecting and orienting different contigs, and mounting the contigs to chromosomes to obtain a fourth assembled genome; and aligning the long read genome sequencing data and short read sequencing data to the fourth assembled genome for correction to obtain an aspergillus telomere-to-telomere genome assembly result.
Owner:PEKING UNION MEDICAL COLLEGE HOSPITAL +2

A method for quantifying srbdv based on sequencing of the transcriptome of infected tissues of rice

The application discloses a kind of based on rice infected tissue transcriptome sequencing of SRBSDV Quantitative method of virus, comprising: 1) obtain the genome assembly and gene annotation file of SRBSDV virus, and the high-throughput transcriptome sequencing original data of rice virus infected tissue;2) construct SRBSDV virus genome index module;3) obtain the non-redundant length of each gene in SRBSDV virus genome annotation file;4) obtain high-quality transcriptome sequencing data;5) obtain the total number of short reads of each sample in step 4);6) obtain individual level BAM file;7) obtain the original expression of different genes in each sample Virus;8) obtain gene expression standardization function FPKB.The application can more accurately evaluate the infection level of virus by means of standardized bioinformatics analysis framework;And more virology information can be mined, the maximization of data utilization is realized, support is provided for virus research and prevention and control, and the accuracy of virus quantification is improved.
Owner:RICE RES INST GUANGDONG ACADEMY OF AGRI SCI

A method for strain-level classification of metagenomic data based on pan-genome graph

ActiveCN118866126BBiostatisticsProteomicsGenomic sequencingShort read
The application discloses a method for strain-level classification of metagenomic data based on a pan-genome graph, which first captures the genomes of multiple strains of each species as a reference through the technology of the pan-genome graph, instead of using only the representative genome of each species as a reference, expands the strain diversity, and enhances the alignment accuracy of metagenomic sequencing reads and the estimation of composition at the strain level; then, by using the corresponding short read alignment tool Giraffe and the long read alignment tool GraphAligner in the alignment of NGS and TGS metagenomic data with the pan-genome graph, two types of data can be efficiently processed at the same time; finally, by further optimizing the path abundance of the species-level classification results based on the pan-genome graph, multi-species strain-level classification can be realized. The application can solve the technical problem that the strain-level classification capability is limited due to the strategy of a single reference genome in the prior art.
Owner:HUNAN UNIV

Ciliate-specific primers, ciliate-specific long primers suitable for amplicon sequencing technology, kits and applications thereof

The application provides ciliate-specific primers, ciliate-specific long primers and kits suitable for amplicon sequencing technology and applications, and specifically belongs to the field of molecular biology detection and high-throughput gene sequencing technology.The ciliate-specific primers include a forward primer CS322F and a reverse primer CS819R;the ciliate-specific long primers suitable for amplicon sequencing include a forward long primer LCS322F and a reverse long primer LCS819R.The ciliate-specific primers have broad-spectrum generality for ciliophora, are suitable for a short read length high-throughput sequencing method, and can be used for comprehensively analyzing the diversity and ecological functions of a ciliate community.The application can realize accurate detection and efficient analysis of the ciliate community, and provides technical support for environmental monitoring, germplasm resource development and ecological assessment.
Owner:OCEAN UNIV OF CHINA

A method for automatic typing of HLA based on third generation sequencing

The application relates to the field of information technology and discloses a method for automatically analyzing and typing HLA based on third-generation sequencing; first, the HLA types of first sequencing read data of different data types of a to-be-tested sample are preliminarily judged, then second sequencing read data is selected, then the identified second sequencing read data of different HLA types is compared with corresponding HLA reference sequences for verification and screening, so that the HLA types existing in the to-be-tested sample are determined; the method can effectively process long read length data generated based on third-generation sequencing technology, reduces the calculation amount of accurate comparison, can realize accurate analysis and report generation of typical third-generation HLA data within 15 minutes, and meets the demand of rapid HLA detection and analysis; meanwhile, the method can also analyze short read length data obtained by second-generation sequencing, and has good data compatibility.
Owner:JIANGSU COWIN BIOTECH CO LTD +1

Composition and method for detection of multiple drug resistance in tuberculosis by multiplex short-amplicon based targeted NGS (TNGS) using ribonucleotide bases in the primers to increase target specificity

PCT designated stageWO2026074580A1Microbiological testing/measurementMultiplexTargeted ngs
This invention relates to a composition and method for the detection of Tuberculosis drug resistance by multiplex small amplicon based, targeted Next-Gen sequencing (tNGS) using ribonucleotide bases in the primers. It also provides a bioinformatic data analysis tool, for the prediction of drug resistance / susceptibility towards anti- tuberculosis therapy drugs. The invention provides an efficient, cost-effective solution for identifying drug resistance / susceptibility right after tuberculosis diagnosis. In addition, a short read NGS platform with massive parallel sequencing, along with data analysis of this invention, offers high sensitivity and specificity for detection of drug resistance.
Owner:MIBIOME THERAPEUTICS LLP

An AI-based virus-host RNA sequence classification method and device

ActiveCN120977392BImprove analytical accuracyeasy to identifyBiostatisticsBiological modelsHost genomeRNA Sequence
The application discloses a virus-host RNA sequence classification method and device based on AI, and relates to the field of biological detection.The method comprises the following steps: mapping preprocessed short read RNA sequences to a host genome twice, assembling short read RNA sequences which are not mapped to the host genome into continuous RNA sequences, and screening RNA sequences with a length greater than 1000bp from the continuous RNA sequences; and performing AI classification on the RNA sequences with a length greater than 1000bp to obtain virus RNA sequences.The application combines host filtering, rapid assembly and AI classification, can significantly reduce the calculation amount and hardware pressure, realizes efficient and accurate virus sequence classification, and can quickly distinguish unknown viruses.
Owner:BEIJING LINGWEI TECHNOLOGY DEVELOPMENT CO LTD

Merging duplicate marking to optimize computer operations for gene sequencing pipeline

ActiveUS12620455B2Other databases indexingSequence analysisReference genome sequenceEngineering
In accordance with embodiments, a processing unit performs alignment of a short read (SR) against a reference genome sequence. The processing unit determines whether the SR is aligned. If the SR is not aligned, the processing unit receives the next SR and processes the next SR by repeating. If the SR is aligned, in response to the determination that the SR is aligned with the reference genome sequence at a first position in the reference genome sequence, the processing unit generates a new SR metadata entry corresponding to the SR. The processing unit finds a linked list in a SR metadata collection. The first position of the linked list in the SR metadata collection corresponds to the first position of the reference genome sequence where the SR is aligned. The processing unit performs duplicate marking based on the SR and the linked list.
Owner:HUAWEI TECH CO LTD

Methods of in-solution positional co-barcoding for sequencing long DNA molecules

PendingUS20260009071A1Microbiological testing/measurementGenomic SegmentSingle strand
The methods and compositions disclosed herein relate to preparing libraries to sequence long molecules in their entirety using massively parallel short read sequencing. The methods disclosed herein generate a nested set of nucleic acid constructs for each genomic fragment and generate a plurality of nested sets for a plurality of genomic fragments. The nucleic acid constructs may be single-stranded or double-stranded. Each nucleic acid construct in each nested set comprises a barcode and target sequence portion, and nucleic acid constructs within each nested set have different lengths. The nucleic acid constructs in each nested set share a unique barcode sequence.
Owner:MGI TECH CO LTD

Preparation of long read nucleic acid libraries

PendingUS20260022371A1Microbiological testing/measurementDNA preparationShort readMedical genetics
Some embodiments of the methods and compositions provided herein relate to obtaining long read information from short reads of a target nucleic acid. Some embodiments include steps to selectively generate, mark, and amplify long nucleic acid fragments. Some embodiments include enriching for certain sequences in the long fragments with selection probes directed to an American College of Medical Genetics (ACMG) panel of genes. Some embodiments also include fragmenting the long nucleic acid fragments into shorter fragments for sequencing, and informatically reconstructing a sequence of the target nucleic acid.
Owner:ILLUMINA INC

Reiterative short read sequencing inside a cellular sample

The present disclosure provides methods for conducting in situ multiplex and multi-omics detection and identification using coded padlocks probes. The methods comprise simultaneous use of RNA-specific padlock probes and polypeptide-specific padlock probes to detect both RNA and polypeptides in a cellular sample. Both types of probes include a barcode that unique identifies the RNA or polypeptide that that padlock probe detects. Both types of probes also include a batch-specific sequencing primer binding site to enable sequencing a desired subset of concatemer template molecules. Use of the batch-specific sequencing primers reduces overcrowding signals and images, to produces optical images that are intense and resolvable. By conducting multiple rounds of sequencing on the same cellular sample using different batch-specific sequencing primers enables multiplex and multi-omics sequencing to reveal numerous target RNAs and their encoded polypeptides.
Owner:ELEMENT BIOSCIENCES INC