Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

23 results about "Contig" patented technology

A contig (from contiguous) is a set of overlapping DNA segments that together represent a consensus region of DNA. In bottom-up sequencing projects, a contig refers to overlapping sequence data (reads); in top-down sequencing projects, contig refers to the overlapping clones that form a physical map of the genome that is used to guide sequencing and assembly. Contigs can thus refer both to overlapping DNA sequence and to overlapping physical segments (fragments) contained in clones depending on the context.

Automatic analysis method and device for phytophagous insect food web DNA molecular data based on high-pass sequencing and storage medium

PendingCN120998298ABiostatisticsProteomicsDNA databaseA-DNA
The invention provides a phytophagous insect food web DNA molecular data automatic analysis method and device based on high-pass sequencing and a storage medium, and relates to the field of molecular biological information detection.The method comprises the steps that sequence splicing, screening and species identification are carried out on obtained double-end sequencing data and local and downloaded DNA databases through an automatic system, and a DNA molecular database is obtained; generating an Excel table containing species names and a DNA bar code sequence file; performing comparative analysis on the double-end sequencing data by adopting matching splicing, and generating a contiguous group sequence based on a local DNA database; if the matching splicing cannot generate the effective sequence, generating a new gene file by adopting non-parameter splicing, and performing gene annotation in combination with the downloaded DNA database; all analysis steps are connected in series through standardized parameter input, including gene screening through threshold values and generation of insect recipe identification results. According to the method, the sequencing data can be subjected to full-process automatic analysis through a one-key command, and the efficiency of food web authentication high-throughput sequencing data processing is greatly improved.
Owner:HEBEI NORMAL UNIV

Rongchang pig T2T genome assembly method

PendingCN121227691ADNA preparationContigGenomic annotation
The invention discloses a Rongchang pig T2T genome assembly method. The method comprises the following steps: 1) collecting and sequencing a sample; 2) genome investigation and assembly; 3) genome annotation; wherein in the sequencing step, three sequencing technical means, namely, a three-generation gene sequencing technology PacBio, Nanopore PromethION 48 short reading and chromatin conception capture (HiC), are adopted, and the Rongchang pig genome is subjected to sequencing and sequence splicing together. The Contig N50 value of the genome is nearly three times that of Sscrofa11.1, and the improvement is mainly embodied in a complex genome region (centromere and telomere regions), so that the genome becomes the most complete genome available at present.
Owner:CHONGQING ACAD OF ANIMAL SCI

A second-generation de novo assembly method and system based on gene numerical expression

The present invention discloses a second-generation de novo assembly method and system based on digital gene expression. The method comprises the following steps: S1: implementing base sequencing data management through a flexible disk array (RAID); S2: gene base sequence analysis and custom numbering; S3: calculating the front and back centroid values ​​of two groups of specific lengths for each sequencing read; S4: constructing a B+ tree index based on the base sequence and read ID of each sequencing read; S5: generating an ID docking table using an artificial intelligence algorithm or a numerical fast sort algorithm; S6: multi-threaded assembly of small contigs based on the ID docking table, outputting a thread identification table, small contig fragments and their numbers, and a single nucleotide variant (SNP) information table; S7: constructing a De Brujin Graph based on the thread identification and pathfinding to assemble the small contigs into contigs; and S8: assembling a scaffold sequence based on the contigs. The present invention completely solves the problem of excessive memory usage in bioinformatics assembly and reduces the risk of assembly errors.
Owner:SICHUAN INNOVATION RES INST OF TIANJIN UNIV +1

A gap-filling method based on third-generation sequencing data

An embodiment of the present invention discloses a gap-filling method based on third-generation sequencing data, comprising the following steps: obtaining an assembled sequence, dividing the assembled sequence into multiple overlapping groups according to gaps; aligning the multiple overlapping groups with each other, and removing excessively long overlapping sequences at the ends of the multiple overlapping groups; aligning reads with the overlapping group sequences, retaining reads that successfully align with the ends of the overlapping group sequences and reads that fail to align with the overlapping group sequences, with the retained reads being referred to as UTreads; aligning the UTreads with themselves, finding dovetail overlap relationships between reads, and deleting dovetail overlap relationships and reads with high depth; constructing a directed graph with the ends of the overlapping groups and both ends of the UTreads as nodes based on the alignments between the ends of the multiple overlapping groups, the alignments between the reads and the ends of the overlapping group sequences, and the dovetail overlap relationships obtained by aligning the UTreads with themselves; finding an optimal path between the ends of the overlapping groups on both sides of the gap based on the directed graph with the ends of the overlapping groups and both ends of the reads as nodes; and generating a sequence to fill the gap based on the optimal path.
Owner:WUHAN GRANDOMICS BIOSCIENCES CO LTD

A method and system for virus detection based on tumor RNA sequencing data

ActiveCN117746985BContigMedicine
A method and system for virus detection in tumor RNA sequencing data are disclosed. The method includes preprocessing the raw tumor RNA sequencing data; inputting the processed tumor RNA sequencing data into sequence-information-based channels and codon-based channels for feature extraction to generate a feature matrix; constructing a sequence information prediction model and a codon information prediction model, and inputting the feature matrices generated from the sequence information-based channels and codon-based channels into the sequence information prediction model and codon information prediction model, respectively, for training and optimization; predicting the virus probability of each sequencing read to obtain a model score; and selecting viral sequencing reads based on the model scores to assemble viral contigs. This invention improves the accuracy and robustness of virus monitoring by introducing a multimodal deep learning method, and can adaptively process sequencing data from different sources and of different lengths, thereby better meeting the needs of medical and research fields for virus identification in tumor sequencing data.
Owner:XIAMEN UNIV

Protein function prediction method and system based on Contig perception

PendingCN121075418ABiostatisticsProteomicsProtein targetProtein function prediction
The invention relates to a Contig perception-based protein function prediction method and system. The method comprises the following steps: obtaining a protein amino acid sequence, a nucleotide sequence corresponding to Contig and arrangement information of CDS on the Contig; splicing the protein-level features corresponding to the CDS sequence on the same Contig with the k-mer frequency vector of the Contig nucleotide sequence to generate enhanced protein-level features; setting the length of a sliding window, acquiring a plurality of fragments with fixed window lengths, which have protein function tags at the central positions of enhanced protein-level features according to a CDS sequence on the same Contig, taking one fragment as a sample, and training a function prediction network consisting of a bidirectional long-short-term memory network BiLSTM and a multilayer perceptron, and obtaining a target protein prediction probability value of each fragment. And the accuracy of protein function annotation is obviously improved.
Owner:HEBEI UNIV OF TECH

Method and device for recognizing a new-born chromatin loop of HPV integration

PendingCN122290694Arecognition stabilityaccurate identificationLocalization systemContig
This application provides a method and apparatus for identifying HPV integration into newly formed chromatin loops, relating to the field of bioinformatics. The method includes: acquiring interaction sequencing read data and determining HPV reference sequences and host reference sequences; identifying HPV integration breakpoints based on chimeric read characteristics, split read characteristics, and abnormal pairing end characteristics; constructing fusion contigs based on breakpoint directions to obtain a breakpoint-aware extended reference set; comparing the interaction sequencing read data with the breakpoint-aware extended reference set, and obtaining high-confidence cross-genomic anchor pairs based on breakpoint-aware alignment constraints; clustering the anchor pairs and screening target HPV-loops; calculating the newborn score of the target HPV-loop and outputting the chromatin loop identification result. This application solves the problem of low accuracy and identification bias that often occurs in traditional interaction localization systems during the identification of newly formed chromatin loops.
Owner:HUAZHONG AGRI UNIV

Pathogenic microorganism genome database, construction method, computer system and application

PendingCN120600129ABiostatisticsInstrumentsPathogenic microorganismStrain specificity
The invention relates to the technical field of bioinformatics, in particular to a pathogenic microorganism genome, a construction method, a computer system and application. According to the method, a genome data quality control strategy in the pathogenic microorganism genome database creation process is innovated, ANI and AF double indexes are creatively adopted for classification error recognition, contig pollution sequence recognition and k-mer short sequence pollution recognition are combined for sequence pollution recognition and processing on genome data, the number of genomes is reserved to the maximum, and the number of the genomes is reduced to the maximum. According to the method, contig level sequence errors are accurately recognized, strain specific sequences are prevented from being mistakenly deleted, the diversity of species genomes is guaranteed, and the constructed database can improve the identification accuracy / sensitivity of pathogen infected species.
Owner:AUTOBIO DIAGNOSTICS CO LTD

Metagenome function annotation method and system based on heterogeneous graph network

The invention relates to a metagenome function annotation method and system based on a heterogeneous graph network. Human metagenome sequencing information is obtained and filtered; the method comprises the following steps: forming a metagenome heterogeneous graph according to metagene internal information: respectively modeling contig nodes and gene nodes, wherein the contig nodes and the gene nodes are connected through an affiliation relationship; determining whether there is an edge relationship among the gene nodes through sequence similarity; if an edge exists between the contig node Ci and the gene node Pi, splicing a contig node feature to a gene node feature; the gene nodes after feature enhancement use a multi-head attention mechanism to carry out message transmission between the gene nodes; and performing gene function classification in a classifier. And high-precision prediction of metagenome functions is realized.
Owner:HEBEI UNIV OF TECH

A method for constructing and analyzing cfDNA-TCR library based on liquid phase probe and application thereof

ActiveCN120708729BBiostatisticsProteomicsMolecular ImmunologyImmune repertoire
The application provides a liquid-phase probe-based cfDNA-TCR library construction analysis method and application thereof, and belongs to the field of molecular immunology, and comprises the following steps: plasma cfDNA extraction, cfDNA UMI-UDI library construction, liquid-phase probe hybridization capture, enrichment of cfDNA library containing TCR genes, second-generation sequencing, off-machine data filtering and cleaning, merging of double-end sequencing data, alignment of sequencing data to a reference genome, removal of non-specific capture sequences, PCR deduplication according to UMI molecules, input of data into TRUST4 for contig assembly and annotation, and further immune repertoire analysis according to annotated CDR3 results. According to the method, TCR molecules in cfDNA can be more accurately detected by distinguishing PCR repeats and cloning repeats.
Owner:HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES

Methods and systems for long-range methylation profiling

Methods for determining an epigenetic profile of a cell in a cell population are described herein. The method can include sequencing DNA molecules obtained from cells in a. cell population to provide a plurality of sequence reads comprising a methylation status for a plurality of bases in each sequence read; and assembling a plurality of contigs based on the plurality of sequence reads. Sequence reads having the same nucleobase sequence and methylation statuses within overlapping portions are joined together to form the same contig. Contigs having substantially the same nucleobase sequence and different methylation profiles are identified as being associated with different cells in the cell population.
Owner:MOONWALK BIOSCIENCES INC

Aspergillus telomere to telomere genome assembly methods, apparatuses, devices, and storage media

PendingCN122392622AGenomic sequencingContig
The application discloses an aspergillus telomere-to-telomere genome assembly method, device, equipment and storage medium. The method comprises the following steps: using a plurality of sequencing sequence assembly tools to assemble target aspergillus long read genome sequencing data from scratch to obtain a first assembled genome; selecting a first assembled genome meeting a preset condition as an initial assembled genome; integrating other first assembled genomes to fill gaps between repeat regions of the initial assembled genome to obtain a second assembled genome; aligning the long read genome sequencing data to the second assembled genome, identifying abnormal coverage regions and correcting sequences to obtain a third assembled genome; aligning a reference genome to the third assembled genome, connecting and orienting different contigs, and mounting the contigs to chromosomes to obtain a fourth assembled genome; and aligning the long read genome sequencing data and short read sequencing data to the fourth assembled genome for correction to obtain an aspergillus telomere-to-telomere genome assembly result.
Owner:PEKING UNION MEDICAL COLLEGE HOSPITAL +2

A chromosome-level gapless genome assembly system and method based on a multi-layer computation graph

This invention discloses a chromosome-level gapless genome assembly system and method based on a multi-layer computational graph, belonging to the field of bioinformatics. The system includes a data preprocessing module, a multi-layer computational graph construction module, an inter-layer communication module, a pathfinding module, a sequence generation module, and a quality assessment module. The multi-layer computational graph structure comprises four layers: a first-layer sequence overlap graph handles read-level overlap relationships; a second-layer fragment connection graph handles contig-level connection relationships; a third-layer scaffold construction graph utilizes Hi-C data for chromosome-level assembly; and a fourth-layer gap-filling graph employs differentiated filling strategies for different types of gaps. The inter-layer communication module enables bidirectional information transfer and conflict resolution. This invention achieves true chromosome-level gapless genome assembly, improving assembly continuity by 3-5 times and reducing the number of gaps by more than 90%.
Owner:CHINA AGRI UNIV

Molecular marker for identifying sex of jellyfish and application of molecular marker

The invention relates to a molecular marker for identifying the sex of jellyfish and application thereof. According to the invention, through whole genome association analysis and research, it is found that the 219111th basic group on the jellyfish Contig REG5524 has an SNP locus, and Ggt occurs; a mutation; the 384007th basic group on the Contig REGS678 has an SNP (Single Nucleotide Polymorphism) site, and Ggt is generated; a mutation; through analysis and verification, the nucleotide sequences of the molecular markers are shown as SEQ ID No.1 and SEQ ID No.2, the molecular markers are SNP markers and are GG genotypes or GA genotypes at the 301st basic group positions of the sequences shown as SEQ ID No.1 and SEQ ID No.2, the GG genotypes are female jellyfish genotypes, and the GA genotypes are male jellyfish genotypes.
Owner:LIAONING ACAD OF MARINE FISHERIES SCI (DALIAN INST OF BIOTECHNOLOGY LIAONING ACAD OF AGRI SCI LIAONING MARINE ENVIRONMENT MONITORING STATION)

Construction and analysis method of cfDNA-TCR library based on liquid phase probe and application of construction and analysis method

ActiveCN120708729ABiostatisticsProteomicsMolecular ImmunologyImmune repertoire
The invention provides a construction and analysis method of a cfDNA-TCR library based on a liquid-phase probe and application thereof, and belongs to the field of molecular immunology, and the construction and analysis method comprises the following steps: plasma cfDNA extraction, cfDNA UMI-UDI library construction, liquid-phase probe hybridization capture, enrichment of a cfDNA library containing a TCR gene, next-generation sequencing, offline data filtration and cleaning, double-end sequencing data merging, and analysis of the cfDNA-TCR library based on the liquid-phase probe. Sequencing data is compared to a reference genome, a non-specific capture sequence is removed, PCR deduplication is carried out according to UMI molecules, the data is input into TRUST4 for contig assembly and annotation, and further immune repertoire analysis is carried out according to an annotated CDR3 result. According to the method disclosed by the invention, the TCR molecules in the cfDNA can be more accurately detected by distinguishing PCR repetition from cloning repetition.
Owner:HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES

Analysis method and analysis system for automatically processing metagenome high-throughput sequencing data binning to obtain viral genome and application

PendingCN121096424AEnsemble learningBiostatisticsOriginal dataSequence clustering
The invention discloses an analysis method for automatically processing metagenome high-throughput sequencing data binning to obtain a viral genome, which comprises the following steps: acquiring second-generation sequencing original data, and performing data quality control to obtain Reads to be analyzed; a plurality of Reads are assembled into an overlapping group through overlapping of fragments, statistics is conducted on the overlapping group, and a length distribution diagram is drawn; reconstructing a viral genome based on contig binning, and generating a box after sequence clustering; constructing a three-dimensional matrix of all boxes according to the attributes of the sequence, and predicting and identifying the species category of each box; carrying out quality identification on the virus box bin, and detecting and rejecting virus host genes; counting the abundance of the virus box bin in each sample to generate a virus box bin abundance table; constructing a virus annotation database, completing species annotation, constructing a virus evolutionary tree, and finally generating a report. The invention further discloses an analysis system and application for implementing the analysis method.
Owner:SHANGHAI OE BIOTECH CO LTD

Optimization method for analyzing environmental microbial community by metagenome binning

The invention discloses an optimization method for analyzing environmental microbial communities by metagenome binning. The optimization method comprises the following steps: collecting environmental microbial community samples, extracting DNA (Deoxyribonucleic Acid), sequencing to generate reads, and carrying out k-mer feature tag calculation and clustering analysis; performing quality control on the reads of each sample, and assembling the clean reads to generate long contigs; comparing the sequencing reads to the contigs to generate a BAM (Business Administration and Maintenance) file containing the coverage information of each contig read; and binning the sample sets in each group in combination with the coverage information and the sequence composition features. According to the method, on one hand, the accuracy of box separation is improved by comparing the coverage modes of different samples, pollutants Contigs and chimeric bins are effectively identified, and hidden pollution is reduced; on the other hand, by integrating various bioinformatics tools and algorithms, efficient and accurate analysis of metagenome data of environmental microbial communities is realized.
Owner:SHANGHAI JIAOTONG UNIV +1

Method for identifying species and functional information of marine eukaryotic microflora at high throughput

The invention discloses a method for identifying species and functional information of marine eukaryotic microbial communities at high throughput, belongs to the technical field of biological information, and aims to solve the problems of inaccurate annotation and low efficiency. The method comprises the following steps of: performing quality control on Mlaspina2010 original high-throughput data by using BBDuk (Broadband Duk) to obtain clean reads; the MEGAHIT is assembled to be contigs; filtering the prokaryotic sequence by combining EukRep with Kaiju, so as to obtain eukaryotic contigs; metaEuk is used for predicting a protein sequence, and DIAMOND is combined with KEGG to carry out function annotation; and calculating contigs and gene abundance by using CoverM and Salmon, and integrating species and function information. Through multi-software collaboration and false positive filtering, full-process analysis from data to ecological analysis is accurately achieved, key technical support is provided for diversity research and resource development of marine eukaryotic microorganisms, and the method has the advantages of being complete in process, high in precision and the like.
Owner:SUN YAT SEN UNIV

Tagging nucleic acids for sequence assembly

PendingUS20250305028A1Microbiological testing/measurementContigMoiety
Various approaches for generating long-distance contiguity information to facilitate contig assembly and phase determination are disclosed. Nucleic acids are assembled into complexes using binding moieties such that, when the nucleic acid backbones are cleaved, the ensuing fragments remain bound. Exposed ends are tagged and ligated either to one another or to tagging moieties such as oligo labels. Ligated junctions are sequenced, and the sequence information is used to assemble contigs into common scaffolds or to assign phase information. Various approaches to tagging the exposed ends are presented.
Owner:DOVETAIL GENOMICS LLC

Nucleic acid sequence assembly

Disclosed herein are compositions, systems and methods related to sequence assembly, such as nucleic acid sequence assembly of single reads and contigs into larger contigs and scaffolds through the use of read pair sequence information, such as read pair information indicative of nucleic acid sequence phase or physical linkage.
Owner:DOVETAIL GENOMICS LLC

Pathogenic fungus generic genome analysis method, device and equipment and readable storage medium

PendingCN121601026ABiostatisticsProteomicsContigFungal gene
The invention relates to a pathogenic fungus generic genome analysis method, device and equipment and a readable storage medium. The method comprises the following steps: obtaining to-be-analyzed sequencing genome data containing a plurality of contig sequences, removing the contig sequences of human and bacteria in the sequencing genome data to obtain cleaned sequencing genome data, performing gene prediction on the cleaned sequencing genome data to obtain a gff protein sequence file corresponding to the cleaned sequencing genome data, and analyzing the gff protein sequence file according to the gff protein sequence file. And finally, performing generic genome clustering analysis on the gff protein sequence file to obtain a clustering analysis result of sequencing genome data. According to the method, the to-be-analyzed sequencing genome is subjected to cleaning of human and bacterial sequences, only fungal gene sequences are reserved, the cleaned data are further predicted, so that the corresponding gff protein sequence file is obtained, clustering analysis is executed on the basis of the gff protein sequence file, and the accuracy and integrity of analysis are improved.
Owner:CHINA TOBACCO SICHUAN IND CO LTD

Method, device and equipment for calculating gene abundance of metagenome

PendingCN121565241AProteomicsGenomicsContigData mining
The invention provides a gene abundance calculation method, device and equipment for metagenomes, and the method comprises the following steps: processing a sample to obtain clean reads and contigs; traversing the clean reads, performing mapping processing on related contigs in the contigs, storing all the clean reads of which the scores meet a first preset threshold value and comparison information of the corresponding contigs into a first SAM file, and splitting all the clean reads of which the scores do not meet the first preset threshold value into a plurality of subsets according to a preset numerical value; traversing each subset, carrying out mapping processing on the subset and each contig in the contigs, and storing all the clean read of which the comparison quality score meets a second preset threshold in each subset and the comparison information of the corresponding contig in a second SAM file; combining the SAM files, sorting the SAM files according to the comparison quality scores, and determining contig corresponding to the highest score of each clear read and corresponding comparison information to obtain a coverage result; and determining the gene abundance according to the obtained gene prediction result and coverage result of the metagenome. And the gene abundance calculation accuracy can be improved.
Owner:SHANGHAI PASSION BIOTECHNOLOGY CO LTD

A multi-threaded method and system for gene assembly

The present invention discloses a multithreaded method and system for gene assembly, which includes the following steps: S1: generating an ID docking table; S2: using a B+ tree index to extract base sequences corresponding to numbered read IDs from a hard disk or virtual memory; S3: matching multiple sequences after the docking relationship with a reference sequence in sequence; S4: reading base sequences of the next batch of numbered read IDs that can match the reference sequence; S5: outputting thread identifiers and matching small contigs after merging the output; S6: after the ID docking table is read, traversing the ID docking table, replacing the base sequences of scattered numbered read IDs in the thread identifier table with corresponding base sequence IDs, and outputting a single nucleotide variant (SNP) information table. The present invention performs calculations by forming a server cluster, using one of the servers as shared storage, and allowing the other servers to individually perform tasks, thereby achieving multi-machine parallel computing on the server, increasing the linear performance of the assembly algorithm, and infinitely reducing the calculation time.
Owner:SICHUAN INNOVATION RES INST OF TIANJIN UNIV +1