Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

14 results about "Contig" patented technology

A contig (from contiguous) is a set of overlapping DNA segments that together represent a consensus region of DNA. In bottom-up sequencing projects, a contig refers to overlapping sequence data (reads); in top-down sequencing projects, contig refers to the overlapping clones that form a physical map of the genome that is used to guide sequencing and assembly. Contigs can thus refer both to overlapping DNA sequence and to overlapping physical segments (fragments) contained in clones depending on the context.

Automatic analysis method and device for phytophagous insect food web DNA molecular data based on high-pass sequencing and storage medium

PendingCN120998298ABiostatisticsProteomicsDNA databaseA-DNA
The invention provides a phytophagous insect food web DNA molecular data automatic analysis method and device based on high-pass sequencing and a storage medium, and relates to the field of molecular biological information detection.The method comprises the steps that sequence splicing, screening and species identification are carried out on obtained double-end sequencing data and local and downloaded DNA databases through an automatic system, and a DNA molecular database is obtained; generating an Excel table containing species names and a DNA bar code sequence file; performing comparative analysis on the double-end sequencing data by adopting matching splicing, and generating a contiguous group sequence based on a local DNA database; if the matching splicing cannot generate the effective sequence, generating a new gene file by adopting non-parameter splicing, and performing gene annotation in combination with the downloaded DNA database; all analysis steps are connected in series through standardized parameter input, including gene screening through threshold values and generation of insect recipe identification results. According to the method, the sequencing data can be subjected to full-process automatic analysis through a one-key command, and the efficiency of food web authentication high-throughput sequencing data processing is greatly improved.
Owner:HEBEI NORMAL UNIV

Rongchang pig T2T genome assembly method

PendingCN121227691ADNA preparationContigGenomic annotation
The invention discloses a Rongchang pig T2T genome assembly method. The method comprises the following steps: 1) collecting and sequencing a sample; 2) genome investigation and assembly; 3) genome annotation; wherein in the sequencing step, three sequencing technical means, namely, a three-generation gene sequencing technology PacBio, Nanopore PromethION 48 short reading and chromatin conception capture (HiC), are adopted, and the Rongchang pig genome is subjected to sequencing and sequence splicing together. The Contig N50 value of the genome is nearly three times that of Sscrofa11.1, and the improvement is mainly embodied in a complex genome region (centromere and telomere regions), so that the genome becomes the most complete genome available at present.
Owner:CHONGQING ACAD OF ANIMAL SCI

A method and system for virus detection based on tumor RNA sequencing data

ActiveCN117746985BContigMedicine
A method and system for virus detection in tumor RNA sequencing data are disclosed. The method includes preprocessing the raw tumor RNA sequencing data; inputting the processed tumor RNA sequencing data into sequence-information-based channels and codon-based channels for feature extraction to generate a feature matrix; constructing a sequence information prediction model and a codon information prediction model, and inputting the feature matrices generated from the sequence information-based channels and codon-based channels into the sequence information prediction model and codon information prediction model, respectively, for training and optimization; predicting the virus probability of each sequencing read to obtain a model score; and selecting viral sequencing reads based on the model scores to assemble viral contigs. This invention improves the accuracy and robustness of virus monitoring by introducing a multimodal deep learning method, and can adaptively process sequencing data from different sources and of different lengths, thereby better meeting the needs of medical and research fields for virus identification in tumor sequencing data.
Owner:XIAMEN UNIV

Protein function prediction method and system based on Contig perception

PendingCN121075418ABiostatisticsProteomicsProtein targetProtein function prediction
The invention relates to a Contig perception-based protein function prediction method and system. The method comprises the following steps: obtaining a protein amino acid sequence, a nucleotide sequence corresponding to Contig and arrangement information of CDS on the Contig; splicing the protein-level features corresponding to the CDS sequence on the same Contig with the k-mer frequency vector of the Contig nucleotide sequence to generate enhanced protein-level features; setting the length of a sliding window, acquiring a plurality of fragments with fixed window lengths, which have protein function tags at the central positions of enhanced protein-level features according to a CDS sequence on the same Contig, taking one fragment as a sample, and training a function prediction network consisting of a bidirectional long-short-term memory network BiLSTM and a multilayer perceptron, and obtaining a target protein prediction probability value of each fragment. And the accuracy of protein function annotation is obviously improved.
Owner:HEBEI UNIV OF TECH

Method and device for recognizing a new-born chromatin loop of HPV integration

PendingCN122290694Arecognition stabilityaccurate identificationLocalization systemContig
This application provides a method and apparatus for identifying HPV integration into newly formed chromatin loops, relating to the field of bioinformatics. The method includes: acquiring interaction sequencing read data and determining HPV reference sequences and host reference sequences; identifying HPV integration breakpoints based on chimeric read characteristics, split read characteristics, and abnormal pairing end characteristics; constructing fusion contigs based on breakpoint directions to obtain a breakpoint-aware extended reference set; comparing the interaction sequencing read data with the breakpoint-aware extended reference set, and obtaining high-confidence cross-genomic anchor pairs based on breakpoint-aware alignment constraints; clustering the anchor pairs and screening target HPV-loops; calculating the newborn score of the target HPV-loop and outputting the chromatin loop identification result. This application solves the problem of low accuracy and identification bias that often occurs in traditional interaction localization systems during the identification of newly formed chromatin loops.
Owner:HUAZHONG AGRI UNIV

Metagenome function annotation method and system based on heterogeneous graph network

The invention relates to a metagenome function annotation method and system based on a heterogeneous graph network. Human metagenome sequencing information is obtained and filtered; the method comprises the following steps: forming a metagenome heterogeneous graph according to metagene internal information: respectively modeling contig nodes and gene nodes, wherein the contig nodes and the gene nodes are connected through an affiliation relationship; determining whether there is an edge relationship among the gene nodes through sequence similarity; if an edge exists between the contig node Ci and the gene node Pi, splicing a contig node feature to a gene node feature; the gene nodes after feature enhancement use a multi-head attention mechanism to carry out message transmission between the gene nodes; and performing gene function classification in a classifier. And high-precision prediction of metagenome functions is realized.
Owner:HEBEI UNIV OF TECH

A method for constructing and analyzing cfDNA-TCR library based on liquid phase probe and application thereof

ActiveCN120708729BBiostatisticsProteomicsMolecular ImmunologyImmune repertoire
The application provides a liquid-phase probe-based cfDNA-TCR library construction analysis method and application thereof, and belongs to the field of molecular immunology, and comprises the following steps: plasma cfDNA extraction, cfDNA UMI-UDI library construction, liquid-phase probe hybridization capture, enrichment of cfDNA library containing TCR genes, second-generation sequencing, off-machine data filtering and cleaning, merging of double-end sequencing data, alignment of sequencing data to a reference genome, removal of non-specific capture sequences, PCR deduplication according to UMI molecules, input of data into TRUST4 for contig assembly and annotation, and further immune repertoire analysis according to annotated CDR3 results. According to the method, TCR molecules in cfDNA can be more accurately detected by distinguishing PCR repeats and cloning repeats.
Owner:HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES

Methods and systems for long-range methylation profiling

Methods for determining an epigenetic profile of a cell in a cell population are described herein. The method can include sequencing DNA molecules obtained from cells in a. cell population to provide a plurality of sequence reads comprising a methylation status for a plurality of bases in each sequence read; and assembling a plurality of contigs based on the plurality of sequence reads. Sequence reads having the same nucleobase sequence and methylation statuses within overlapping portions are joined together to form the same contig. Contigs having substantially the same nucleobase sequence and different methylation profiles are identified as being associated with different cells in the cell population.
Owner:MOONWALK BIOSCIENCES INC

Aspergillus telomere to telomere genome assembly methods, apparatuses, devices, and storage media

PendingCN122392622AGenomic sequencingContig
The application discloses an aspergillus telomere-to-telomere genome assembly method, device, equipment and storage medium. The method comprises the following steps: using a plurality of sequencing sequence assembly tools to assemble target aspergillus long read genome sequencing data from scratch to obtain a first assembled genome; selecting a first assembled genome meeting a preset condition as an initial assembled genome; integrating other first assembled genomes to fill gaps between repeat regions of the initial assembled genome to obtain a second assembled genome; aligning the long read genome sequencing data to the second assembled genome, identifying abnormal coverage regions and correcting sequences to obtain a third assembled genome; aligning a reference genome to the third assembled genome, connecting and orienting different contigs, and mounting the contigs to chromosomes to obtain a fourth assembled genome; and aligning the long read genome sequencing data and short read sequencing data to the fourth assembled genome for correction to obtain an aspergillus telomere-to-telomere genome assembly result.
Owner:PEKING UNION MEDICAL COLLEGE HOSPITAL +2

A chromosome-level gapless genome assembly system and method based on a multi-layer computation graph

This invention discloses a chromosome-level gapless genome assembly system and method based on a multi-layer computational graph, belonging to the field of bioinformatics. The system includes a data preprocessing module, a multi-layer computational graph construction module, an inter-layer communication module, a pathfinding module, a sequence generation module, and a quality assessment module. The multi-layer computational graph structure comprises four layers: a first-layer sequence overlap graph handles read-level overlap relationships; a second-layer fragment connection graph handles contig-level connection relationships; a third-layer scaffold construction graph utilizes Hi-C data for chromosome-level assembly; and a fourth-layer gap-filling graph employs differentiated filling strategies for different types of gaps. The inter-layer communication module enables bidirectional information transfer and conflict resolution. This invention achieves true chromosome-level gapless genome assembly, improving assembly continuity by 3-5 times and reducing the number of gaps by more than 90%.
Owner:CHINA AGRI UNIV

Analysis method and analysis system for automatically processing metagenome high-throughput sequencing data binning to obtain viral genome and application

PendingCN121096424AEnsemble learningBiostatisticsOriginal dataSequence clustering
The invention discloses an analysis method for automatically processing metagenome high-throughput sequencing data binning to obtain a viral genome, which comprises the following steps: acquiring second-generation sequencing original data, and performing data quality control to obtain Reads to be analyzed; a plurality of Reads are assembled into an overlapping group through overlapping of fragments, statistics is conducted on the overlapping group, and a length distribution diagram is drawn; reconstructing a viral genome based on contig binning, and generating a box after sequence clustering; constructing a three-dimensional matrix of all boxes according to the attributes of the sequence, and predicting and identifying the species category of each box; carrying out quality identification on the virus box bin, and detecting and rejecting virus host genes; counting the abundance of the virus box bin in each sample to generate a virus box bin abundance table; constructing a virus annotation database, completing species annotation, constructing a virus evolutionary tree, and finally generating a report. The invention further discloses an analysis system and application for implementing the analysis method.
Owner:SHANGHAI OE BIOTECH CO LTD

Nucleic acid sequence assembly

Disclosed herein are compositions, systems and methods related to sequence assembly, such as nucleic acid sequence assembly of single reads and contigs into larger contigs and scaffolds through the use of read pair sequence information, such as read pair information indicative of nucleic acid sequence phase or physical linkage.
Owner:DOVETAIL GENOMICS LLC

Pathogenic fungus generic genome analysis method, device and equipment and readable storage medium

PendingCN121601026ABiostatisticsProteomicsContigFungal gene
The invention relates to a pathogenic fungus generic genome analysis method, device and equipment and a readable storage medium. The method comprises the following steps: obtaining to-be-analyzed sequencing genome data containing a plurality of contig sequences, removing the contig sequences of human and bacteria in the sequencing genome data to obtain cleaned sequencing genome data, performing gene prediction on the cleaned sequencing genome data to obtain a gff protein sequence file corresponding to the cleaned sequencing genome data, and analyzing the gff protein sequence file according to the gff protein sequence file. And finally, performing generic genome clustering analysis on the gff protein sequence file to obtain a clustering analysis result of sequencing genome data. According to the method, the to-be-analyzed sequencing genome is subjected to cleaning of human and bacterial sequences, only fungal gene sequences are reserved, the cleaned data are further predicted, so that the corresponding gff protein sequence file is obtained, clustering analysis is executed on the basis of the gff protein sequence file, and the accuracy and integrity of analysis are improved.
Owner:CHINA TOBACCO SICHUAN IND CO LTD

Method, device and equipment for calculating gene abundance of metagenome

PendingCN121565241AProteomicsGenomicsContigData mining
The invention provides a gene abundance calculation method, device and equipment for metagenomes, and the method comprises the following steps: processing a sample to obtain clean reads and contigs; traversing the clean reads, performing mapping processing on related contigs in the contigs, storing all the clean reads of which the scores meet a first preset threshold value and comparison information of the corresponding contigs into a first SAM file, and splitting all the clean reads of which the scores do not meet the first preset threshold value into a plurality of subsets according to a preset numerical value; traversing each subset, carrying out mapping processing on the subset and each contig in the contigs, and storing all the clean read of which the comparison quality score meets a second preset threshold in each subset and the comparison information of the corresponding contig in a second SAM file; combining the SAM files, sorting the SAM files according to the comparison quality scores, and determining contig corresponding to the highest score of each clear read and corresponding comparison information to obtain a coverage result; and determining the gene abundance according to the obtained gene prediction result and coverage result of the metagenome. And the gene abundance calculation accuracy can be improved.
Owner:SHANGHAI PASSION BIOTECHNOLOGY CO LTD