Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

16 results about "Sequence assembly" patented technology

In bioinformatics, sequence assembly refers to aligning and merging fragments from a longer DNA sequence in order to reconstruct the original sequence. This is needed as DNA sequencing technology cannot read whole genomes in one go, but rather reads small pieces of between 20 and 30000 bases, depending on the technology used. Typically the short fragments, called reads, result from shotgun sequencing genomic DNA, or gene transcript (ESTs).

DNA sequence assembly method and system based on dynamic variable-order unitg-level graph

The invention discloses a DNA sequence assembly method and system based on a dynamic variable order unitg-level graph, and relates to the technical field of bioinformatics and DNA storage. According to the method, a pseudo genome and a source perception k-mer index are constructed, a Mid-Max and Min-Mid two-stage variable order expansion strategy is adopted, a k value is dynamically adjusted to enhance the graph structure connectivity, and the problems that under the condition of low coverage rate or high error rate, an existing de Bruijn graph method is prone to breakage and path fuzziness is prone to being generated are effectively solved. According to the method, a hidden path is accurately repaired through node connection and splitting operation, and redundancy k-mer is filtered in combination with index continuity and prefix similarity, so that the assembly integrity and accuracy are improved, and meanwhile, the robustness of DNA data reconstruction is remarkably enhanced; the method is suitable for various high-noise and low-coverage-rate scenes such as genome assembly and DNA storage and reconstruction, and particularly shows excellent performance in practical application with high requirements on data integrity and reliability.
Owner:DALIAN UNIV

Pit mud metagenome data automatic analysis method and system

The invention relates to the technical field of metagenomics, discloses an automatic analysis method and system for pit mud metagenomic data, and aims at solving the problem that an existing method is poor in efficiency and accuracy, and the scheme mainly comprises the steps that a sequencing data type, a file path and analysis parameters are received; performing quality control on the original offline data; sequence assembly is carried out, and a contigs file is generated; carrying out assembly quality evaluation on the contigs file; carrying out genome binning by using at least two binning tools; integrating output results of the binning tool, and performing optimization based on a preset integrity threshold value and a preset pollution degree threshold value to obtain an optimized binning genome data set; evaluating and optimizing the integrity, the pollution degree and the strain heterogeneity of the binning genome; calculating coverage and relative abundance; performing species classification annotation and function annotation; and integrating the result data of the previous steps to generate an analysis report. According to the method, automatic analysis of metagenome data is realized, and the analysis efficiency and accuracy are improved.
Owner:WULIANGYE +1

A method for analyzing macroviral group data

ActiveCN116682492BSequence analysisHybridisationMedicineSequence clustering
The application discloses a macrovirus group data analysis method, and belongs to the technical field of macrovirology. The method comprises the following steps: sequence quality control, sequence assembly, sequence clustering, virus sequence identification, virus sequence checking, virus abundance calculation, species annotation, virus lifestyle judgment, virus host prediction and virus auxiliary metabolism gene analysis. The application uses Trimmomatic software, BWA-MEN, Megahit software, CD-HIT and other tools to execute the analysis process of macrovirus group data. Practice proves that the application can accurately identify and annotate virus species, comprehensively and systematically deeply analyze and mine macrovirus group data, the steps are simple and clear, the analysis time is short, and the effect of macrovirus identification research is greatly improved.
Owner:JIANGNAN UNIV

Populus hybrid parent source identification method, system and equipment based on single-copy gene sequence and medium

The invention belongs to the technical field of gene sequencing, and particularly relates to a populus hybrid parent source identification method, system, equipment and medium based on a single-copy gene sequence, the method comprises the following steps: firstly, constructing a single-copy gene sequence reference library containing five populus reference species, and each sequence comprises a gene coding region and upstream and downstream extension sequences thereof; secondly, acquiring second-generation sequencing data of a to-be-detected sample, comparing and assembling by taking the representative species sequence as reference, and reconstructing a single-copy gene sequence of the to-be-detected sample; and finally, determining a parent source through sequence alignment and statistical analysis. According to the method disclosed by the invention, a single-copy gene sequence can be accurately reconstructed from a hybrid sample through an optimally designed sequence assembly and comparison process, so that the coverage degree and the resolution ratio of a genetic marker are remarkably improved; in addition, a comparison screening rule and a layering judgment standard based on a complete matching length are designed, and precise identification of hybrid parent sources is realized by quantitatively analyzing the matching proportion of each reference species.
Owner:NANJING FORESTRY UNIV

System and method for storing and sharing genomic data using blockchain

A method of compressing genomic data. The method has the steps of: aligning the genomic data with reference data, obtaining difference between the genomic data and the reference data, and compressing the difference using a statistical compression method to obtain compressed genomic data. In some embodiments, the statistical compression method may be an arithmetic coding method. In some embodiments, the method may further has a step of processing the difference using one or more statistical modeling methods, and compressing the processed difference using the statistical compression method. In some embodiments, the method further has a step of assembling a plurality of reads to form the reference data. In some embodiments, the method further has a step of storing compressed genomic data in a blockchain.
Owner:CARDIAI TECH LTD

Methods and systems for proximity enhanced sequence assembly

Described herein are methods and systems for assembling sequence reads from a genomic nucleic acid sample. In some embodiments, the sequence reads are read from clusters of nucleic acids on a flow cell. In some embodiments, the methods and systems use information relating to the proximity of clusters determined from their flow cell locations to assemble the sequence reads. For example, flow cell proximity information may be used during steps of generating a first assembly, recruiting sequence reads, generating a second assembly, and / or scaffolding an assembly.
Owner:ILLUMINA INC

A method for rapid identification of exogenous insertion sequences based on whole genome sequencing data

This invention provides a method for rapidly identifying exogenous inserted sequences based on whole-genome sequencing data, comprising the following steps: sequencing data quality control, preliminary data alignment, screening of preliminary alignment results, evaluation of secondary alignment results, homology detection, preliminary insertion region detection, precise insertion site identification, assembly of inserted and flanking sequences, acquisition of genes affected by the insertion site, and primer design for the insertion site. Compared with traditional experimental detection techniques, this invention is less time-consuming, reproducible, and presents more comprehensive transgenic insertion characteristics; compared with other existing whole-genome resequencing methods, it is more accurate and intuitive, the results are easier to understand, and it is more conducive to the safety risk assessment of transgenic plants and animals, thus contributing to the development and application of transgenic or gene editing technologies.
Owner:WUHAN WANMO TECH CO LTD

Peptide sequence assembly method and apparatus based on de bruijn graph

The application provides a de Bruijn graph-based peptide sequence assembly method and device, the method comprises the following steps: obtaining light and heavy chain sequence data, and creating a sequence alignment database using the light and heavy chain sequence data; performing k-mer processing on peptide sequence test data to obtain a k-mer sequence set; aligning the k-mer sequence set with the sequence alignment database to obtain an alignment score; if the k-mer sequence set is mixed light and heavy chain data, the alignment score is divided into two categories of light chain and heavy chain, and a de Bruijn graph is constructed respectively, if the k-mer sequence set is light chain or heavy chain data, a de Bruijn graph is directly constructed; performing sequence assembly in the de Bruijn graph to obtain assembled peptide sequences. The application can perform peptide sequence assembly under the condition of distinguishing light and heavy chains, the assembly efficiency is greatly improved, and the assembly result is effective and reliable.
Owner:JIANGSU UNIV OF TECH

Aspergillus telomere to telomere genome assembly methods, apparatuses, devices, and storage media

PendingCN122392622AGenomic sequencingContig
The application discloses an aspergillus telomere-to-telomere genome assembly method, device, equipment and storage medium. The method comprises the following steps: using a plurality of sequencing sequence assembly tools to assemble target aspergillus long read genome sequencing data from scratch to obtain a first assembled genome; selecting a first assembled genome meeting a preset condition as an initial assembled genome; integrating other first assembled genomes to fill gaps between repeat regions of the initial assembled genome to obtain a second assembled genome; aligning the long read genome sequencing data to the second assembled genome, identifying abnormal coverage regions and correcting sequences to obtain a third assembled genome; aligning a reference genome to the third assembled genome, connecting and orienting different contigs, and mounting the contigs to chromosomes to obtain a fourth assembled genome; and aligning the long read genome sequencing data and short read sequencing data to the fourth assembled genome for correction to obtain an aspergillus telomere-to-telomere genome assembly result.
Owner:PEKING UNION MEDICAL COLLEGE HOSPITAL +2

Antibody sequencing method and device

The invention belongs to the technical field of biological analysis, and relates to an antibody sequencing method and device, in particular to an antibody peptide sequence assembling method and device based on beam search. Specifically, the invention relates to an antibody sequencing method which comprises the following steps: S1, obtaining a homologous template of an antibody to be detected; s2, cleaning the de novo sequencing data of the peptide fragment to obtain a cleaned peptide fragment sequence; s3, segmenting the cleaned peptide fragment sequence into a short peptide sequence with a fixed length; and S4, based on the amino acid information on the homologous template and the signal intensity of the oligopeptide sequence, constructing a Debrueine map through beam search, and carrying out sequence assembly to obtain an antibody sequence. The flux of the monoclonal antibody from the beginning sequencing can be improved, so that the mass spectrometric detection time and the cost of experimental consumables are reduced, and the method has a good application prospect.
Owner:XIANG AN BIOMEDICINE LABORATORY +1

Tagging nucleic acids for sequence assembly

PendingUS20250305028A1Microbiological testing/measurementContigMoiety
Various approaches for generating long-distance contiguity information to facilitate contig assembly and phase determination are disclosed. Nucleic acids are assembled into complexes using binding moieties such that, when the nucleic acid backbones are cleaved, the ensuing fragments remain bound. Exposed ends are tagged and ligated either to one another or to tagging moieties such as oligo labels. Ligated junctions are sequenced, and the sequence information is used to assemble contigs into common scaffolds or to assign phase information. Various approaches to tagging the exposed ends are presented.
Owner:DOVETAIL GENOMICS LLC

Method and apparatus for pooled sequencing, electronic device and storage medium

The present disclosure provides a method and device for mixed sample sequencing, an electronic device and a storage medium, wherein after a first to-be-sequenced library is sequenced to obtain a first test sequence set, a second to-be-sequenced library is sequenced to obtain a second test sequence set; the test sequences in the test sequence set are assembled; each assembled sequence set is evaluated to obtain an assembled sequence corresponding to each to-be-sequenced sample; and the assembled sequences corresponding to each to-be-sequenced sample are determined based on the similarity between the assembled sequences. According to the present disclosure, after the first to-be-sequenced library and the second to-be-sequenced library are continuously sequenced in the same sequencing chip, the sequences obtained by sequencing are assembled and evaluated to obtain the assembled sequence corresponding to the first to-be-sequenced library and the assembled sequence corresponding to the second to-be-sequenced library; and the present disclosure can reduce the time spent on cleaning the sequencing chip when sequencing two adjacent batches, thereby improving the sequencing efficiency.
Owner:BGI TECH SOLUTIONS CO LTD

Nucleic acid sequence assembly

Disclosed herein are compositions, systems and methods related to sequence assembly, such as nucleic acid sequence assembly of single reads and contigs into larger contigs and scaffolds through the use of read pair sequence information, such as read pair information indicative of nucleic acid sequence phase or physical linkage.
Owner:DOVETAIL GENOMICS LLC

CRISPR-Cas12a-based plasmid pool DNA database management method and system

The invention discloses a CRISPR-Cas12a-based plasmid pool DNA database management method and a CRISPR-Cas12a-based plasmid pool DNA database management system. The method comprises the following steps: S1, data writing: encoding data required to be stored by a user, assembling an encoded DNA sequence and a designed multi-layer address sequence into a plasmid map, and synthesizing and cloning to form a plasmid pool database; s2, data retrieval: S21, analyzing a target address; s22, assembling the one-way guide RNA and a Cas12a protein into a ribonucleoprotein complex; s23, carrying out specific cleavage on the plasmid pool database by utilizing a ribonucleoprotein complex; s24, enriching a cutting product; s25, performing amplification and sequencing on the enriched product, and restoring target data after decoding; s3, data deletion: S31, designing a one-way guide RNA (Ribonucleic Acid) aiming at a to-be-deleted data address sequence; s32, cutting a target plasmid through a Cas12a protein; and S33, separating and removing the cut plasmids through a gel electrophoresis method so as to realize data deletion, and promoting DNA storage from a static archiving mode to a dynamic database management mode.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Analysis method of norovirus high-throughput sequencing sequence

The invention provides an analysis method of a norovirus high-throughput sequencing sequence, which comprises the following steps of: 1) performing quality control on sequencing data to obtain a filtering data set; 2) comparing with a self-built virus read comparison database to respectively obtain the number of target pathogen reads, the proportion of the target pathogen reads and a target pathogen comparison matching data bam file in the sequencing sample, and converting the bam file into a fastq file; 3) screening and assembling the read segments from the beginning to obtain an assembled sequence file; and 4) comparing with a self-built virus identification database, carrying out manual correction, and outputting a final assembly sequence result. The invention belongs to the technical field of pathogenic microorganism high-throughput sequencing, establishes a complete norovirus sequence assembly and identification sequencing offline processing method, increases the utilization rate of virus effective reads in metagenome data, and improves the sensitivity and specificity of virus genome sequence identification and typing.
Owner:SHANGHAI CUSTOMS COLLEGE

An AI-based virus-host RNA sequence classification method and device

ActiveCN120977392BImprove analytical accuracyeasy to identifyBiostatisticsBiological modelsHost genomeRNA Sequence
The application discloses a virus-host RNA sequence classification method and device based on AI, and relates to the field of biological detection.The method comprises the following steps: mapping preprocessed short read RNA sequences to a host genome twice, assembling short read RNA sequences which are not mapped to the host genome into continuous RNA sequences, and screening RNA sequences with a length greater than 1000bp from the continuous RNA sequences; and performing AI classification on the RNA sequences with a length greater than 1000bp to obtain virus RNA sequences.The application combines host filtering, rapid assembly and AI classification, can significantly reduce the calculation amount and hardware pressure, realizes efficient and accurate virus sequence classification, and can quickly distinguish unknown viruses.
Owner:BEIJING LINGWEI TECHNOLOGY DEVELOPMENT CO LTD