Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

15 results about "Genomic annotation" patented technology

DNA annotation or genome annotation is the process of identifying the locations of genes and all of the coding regions in a genome and determining what those genes do. An annotation (irrespective of the context) is a note added by way of explanation or commentary.

Rongchang pig T2T genome assembly method

PendingCN121227691ADNA preparationContigGenomic annotation
The invention discloses a Rongchang pig T2T genome assembly method. The method comprises the following steps: 1) collecting and sequencing a sample; 2) genome investigation and assembly; 3) genome annotation; wherein in the sequencing step, three sequencing technical means, namely, a three-generation gene sequencing technology PacBio, Nanopore PromethION 48 short reading and chromatin conception capture (HiC), are adopted, and the Rongchang pig genome is subjected to sequencing and sequence splicing together. The Contig N50 value of the genome is nearly three times that of Sscrofa11.1, and the improvement is mainly embodied in a complex genome region (centromere and telomere regions), so that the genome becomes the most complete genome available at present.
Owner:CHONGQING ACAD OF ANIMAL SCI

A method for serialization extraction of highly variable exons

PendingCN122290698AInformation densityExon
This invention discloses an efficient RNA data preprocessing method to address the problems of low processing efficiency and low information density in high-throughput sequencing data. Its core steps include: (1) introducing a parallel processing scheme for high-throughput sequence data, rapidly mapping RNA-seq data to a reference genome to generate a BAM file; (2) extracting base sequences and expression levels and storing them as compact PKL format files; (3) extracting all exon position information by parsing the genome annotation file; (4) combining multi-sample expression level data to screen for highly variable exons and constructing a high-information-density feature list based on the sample set; and (5) accurately extracting target sequences from the preprocessed file based on this list. Compared to traditional methods, this innovative approach achieves triple optimization: full-process parallel processing for accelerated computation, high-compression data storage, and adaptive feature selection. Processing speed is increased by 3-5 times, and data volume is reduced by more than 90%, making it suitable for high-throughput RNA-seq data analysis with large sample sizes.
Owner:TIANJIN UNIV

Strain specific culture medium prediction method, medium and computer equipment

The invention discloses a strain specific culture medium prediction method, a medium and computer equipment. The method comprises the following steps: acquiring strain gene function annotation data and encoding the data into a high-dimensional feature vector; performing principal component analysis dimension reduction on the high-dimensional feature vector; obtaining candidate culture medium component information and encoding the candidate culture medium component information into a binary vector; splicing the low-dimensional strain feature representation and the medium component vector to form a joint input feature vector; inputting the joint input feature vector into a pre-trained deep residual network model; and outputting a growth compatibility probability score and sorting the candidate culture media. By integrating high-throughput genome annotation data and a known culture medium formula, a data-driven intelligent prediction model is constructed and is used for efficiently recommending an optimal culture medium composition suitable for a specific microbial strain, and the prediction precision and the screening speed of microbial isolated culture are remarkably improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Genome annotation method and electronic device

PendingCN121306282ASequence analysisInstrumentsGene AnnotationGenomic annotation
The invention provides a genome annotation method and an electronic device. The genome annotation method comprises the following steps: S1) performing gene structure prediction on a genome by adopting multiple modes to obtain multiple prediction gene sets; s2) performing gene integration on the multiple predictive gene sets by using an EVM tool to obtain an integrated gene set; s3) performing BUSCO evaluation on the integrated gene set to obtain an integrated gene set evaluation file; s4) performing BUSCO evaluation on the genome to obtain a genome evaluation file; and S5) correcting the integrated gene set evaluation file by using the genome evaluation file to obtain a corrected gene set, wherein the gene structure prediction comprises transcriptome prediction, de novo prediction and homologous prediction.The genome annotation method can significantly improve the accuracy and integrity of gene annotation.
Owner:YUN SI TUO (TIAN JIN) SHENG WU KE JI YOU XIAN GONG SI

Construction and Analysis Methods of Genome-Scale Metabolic Network Model of Paranitrogenous Denitrifying Cocci

ActiveCN119626320Befficient designEfficient transformationBiostatisticsProteomicsMetabolic network modelGenomic information
This invention discloses a method for constructing and analyzing a genome-scale metabolic network model of *Paragonimus denitrifyingus*, belonging to the field of systems biology. The method includes: whole-genome annotation; obtaining global metabolic response data of *Paragonimus denitrifyingus*; automatically retrieving genomic information and constructing Model 1 based on the species code and genome annotation results of *Paragonimus denitrifyingus* in the KEGG database; constructing Model 2 by identifying homologous proteins in the *Paragonimus denitrifyingus* genome through homology searching of proteins in the target organism based on a pre-trained Hidden Markov Model; and integrating Model 1 and Model 2. This invention allows for the efficient design and modification of denitrification engineering, achieving precise control of nitrogen degradation processes. Compared to existing metabolic engineering methods, this invention effectively reduces the workload of exploratory experiments and greatly advances a deeper understanding of the nitrogen degradation characteristics of *Paragonimus denitrifyingus*.
Owner:JIANGNAN UNIV

Cytochrome P450 reductase coding gene CPR2 from aspergillus nidulans and application thereof

PendingCN121737167AFungiMicroorganism based processesCytochrome P450 reductaseNucleotide
The invention belongs to the technical field of gene engineering and microbial fermentation, and particularly discloses a cytochrome P450 reductase coding gene CPR2 from aspergillus nidulans and application thereof. The nucleotide sequence of the gene is shown as SEQ ID NO: 1, and the amino acid sequence of the protein coded by the gene is shown as SEQ ID NO: 2. According to the invention, the CPR2 gene (OECPR2) is identified as an optimal electron transfer regulatory element by performing function screening on eight CPR genes annotated in a genome, and the CPR2 gene is over-expressed in an echinocandin B production strain aspergillus nidulans, so that the intracellular electron transfer efficiency is remarkably enhanced, the enzyme catalytic activity of cytochrome P450 is driven, and the cytochrome P450 can be used for preparing an echinocandin B cell. And finally, the yield of echinocandin B is efficiently increased. Experiments show that the engineering strain for over-expression of CPR2 can enable the yield of echinocandin B to reach 1869.97 + / -96.98 mg / L in shake flask fermentation, which is 24.37% higher than that of the original strain. The invention provides a key gene element and an engineering strain for industrial efficient production of echinocandin B, and has important economic value.
Owner:ZHEJIANG UNIV OF TECH +1

An improved method for whole genome selection of corn hybrids

PendingCN122290693AGenotypingHaplotype
This application discloses an improved method for whole-genome selection of maize hybrids, belonging to the field of genetic breeding technology. To address the shortcomings of existing technologies, this application provides a computer device to implement the following steps: (A1) Construction of a maize pan-genome: screening core planting and breeding backbone parents, constructing libraries and sequencing, loading reference genomes and genome annotations, and iteratively assembling to construct a pan-genome to obtain a sequence / gene-based pan-genome; (A2) Construction of a maize haplotype library: setting chromosomal reference segments / genes and statistically analyzing the haplotypes of each reference segment to obtain a whole-genome haplotype library; (A3) Haplotype genotyping of the training population: obtaining haplotype data of the training population; (A4) Whole-genome prediction: training the training model data through evaluation dimensions such as haplotype effect value, parental combining ability, and / or prediction accuracy.
Owner:CHINA AGRI UNIV

A method for constructing a gene regulatory network and related devices

The application discloses a gene regulation network construction method and related device, and relates to the technical field of gene regulation network construction. The method comprises the following steps: downloading position data of candidate regulation elements from a SCREEN database, combining a genome annotation file and scATAC-seq data to accurately position promoters and enhancers, introducing a variational autoencoder to accurately determine enhancer-promoter pairs that exist potential interaction, associating enhancers with target genes, processing a co-expression gene network based on potential binding sites of each transcription factor of a to-be-detected object on an enhancer, enhancer-promoter pairs that exist potential interaction, and target genes corresponding to each promoter, and obtaining a gene regulation network. The application can solve the problems of failing to accurately position promoters and enhancers, being difficult to associate enhancers with target genes, being difficult to distinguish direct and indirect relationships, and having too many false positive regulation relationships.
Owner:INNER MONGOLIA UNIVERSITY

A Transposon Hierarchical Classification Method Based on Dynamically Gated Multi-Scale Features

This invention belongs to the field of bioinformatics, specifically relating to a hierarchical transposon classification method based on dynamically gated multi-scale features. First, K-mer frequencies and one-hot encoded bimodal features are extracted from the DNA sequence. Second, these features are encoded separately using a gated multi-scale convolutional network and a DNA-specific convolutional pyramid. Then, adaptive feature fusion is performed using an attention-gating mechanism conditioned on sequence features. Finally, a top-down hierarchical decision is made in the classification tree based on probability thresholds. This invention overcomes the shortcomings of existing methods in feature capture, fusion, and hierarchical decision-making, achieving high-precision and high-efficiency transposon classification, providing a more accurate and reliable analytical tool for genome annotation, evolutionary research, and functional genomics.
Owner:LUDONG UNIVERSITY

Method for improving genome selection accuracy by integrating genome annotation information

PendingCN121938453AMathematical modelsBiostatisticsNucleotide diversityGenome resequencing
The invention discloses a method for improving genome selection accuracy by integrating genome annotation information, and belongs to the technical field of molecular breeding. The method comprises the following steps: constructing a reference group and accurately determining target traits; acquiring high-density genotype data by using a whole genome re-sequencing technology; constructing an SNP function annotation matrix containing multi-dimensional information such as a gene structure, codon degeneracy in a coding region, nucleotide diversity and the like; estimating genetic variance weights of different functional regions by using a random Heismann-Elstarton regression (RHE) algorithm and a Monte Carlo REML algorithm; and constructing a weighted genome prediction model based on mixed prior distribution to estimate a breeding value. According to the method, differential modeling is carried out on functional sites and background noise markers through biological prior information, the genome selection accuracy of complex characters and the robustness of cross-population prediction are remarkably improved, and the method has a good application prospect.
Owner:OCEAN UNIV OF CHINA

Methods and systems for creating methylation-based scores for cancer prediction and tumor tissue of origin determination

PendingUS20260179721A1BiostatisticsMedical automated diagnosisGenomic intervalDisease
Methods for generating and evaluation methylation meta scores and their use for the detection of disease and prediction of tumor tissue of origin are described. The methods may comprise, e.g., generating annotated genomic interval data comprising at least one of: (i) genomic coordinates, and (ii) associated genomic annotation data; providing annotated patient sample data comprising at least one of: (i) methylation sequencing data, and (ii) associated clinical annotation data; generating an aggregated methylation meta score matrix comprising candidate methylation meta scores; inputting (i) the aggregated methylation meta score matrix, and (ii) the associated clinical annotation data for at least a subset of patient samples as training data for training a machine learning model; and training the model using the training data to identify an optimal methylation meta score for disease detection and / or prediction of disease tissue of origin (TOO) for at least one disease type based on methylation sequencing data.
Owner:FOUNDATION MEDICINE INC

Systems and methods for machine learning-based genome annotation

PCT designated stageWO2025191449A8BiostatisticsProteomicsNucleotidePromoter
The present disclosure, among other things, provides machine-learning technologies for identifying and localizing particular genomic elements (e.g., gene elements and / or regulatory elements) within nucleotide sequences, such as DNA and / or RNA sequences. In certain embodiments, similar to the manner in which image processing methods can be used to localize particular objects in images at pixel level resolution, referred to as "segmentation," systems and methods of the present disclosure predict presence and locations of certain genomic elements within nucleotide sequences, thereby "segmenting" nucleotide sequences. Accordingly, genomic element segmentation technologies described herein may be used to generate annotations that identify and label portions of nucleotide sequences according to their predicted (e.g., via machine learning models described herein) function – e.g., as protein-coding genes, untranslated regions, splice sites, promotors, enhancers, etc. Among other things, these genomic annotations may be used to inform underlying biological processes driving diseases and facilitate development of new therapies.
Owner:INSTADEEP LTD +1

Genomic language model training method and device, electronic equipment and storage medium

The invention provides a genome language model training method and device, electronic equipment and a storage medium, and relates to the technical field of computers. The method comprises the following steps: acquiring original genome data; constructing a natural language expression sequence of a gene product for the original genome data based on genome annotation; pre-training a basic language model by using the natural language expression sequence of the gene product to obtain a pre-trained model; and performing supervision and fine tuning on the pre-trained model for different gene function tasks to obtain a genome language model subjected to supervision and fine tuning. According to the method, the problems that a current genome language model takes nucleotide or k-mer as a token and has no explicit mapping with biological functions, a prediction result cannot be explained and a nucleotide sequence modeling is too long, so that computing resources required by training are very huge are solved.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Dynamic gating multi-scale feature-based transposon hierarchical classification method

The invention belongs to the field of bioinformatics, and particularly relates to a transposon hierarchical classification method based on dynamic gating multi-scale features. The method comprises the following steps: firstly, extracting K-mer frequency and One-hot coding bimodal features from a DNA sequence; secondly, coding is carried out through a gated multi-scale convolutional network and a DNA specific convolutional pyramid; then, self-adaptive feature fusion is carried out by utilizing an attention gating mechanism taking sequence features as conditions; and finally, performing top-down hierarchical decision-making in the classification tree based on a probability threshold. According to the method, the defects of an existing method in feature capture, fusion and hierarchical decision making are overcome, high-precision and high-efficiency transposon classification is achieved, and a more accurate and reliable analysis tool is provided for genome annotation, evolutionary research and functional genomics.
Owner:LUDONG UNIVERSITY

Compost N2O emission path identification and environmental risk quantification method based on multiple omics

The invention belongs to the technical field of biological information, and particularly relates to a compost N2O emission path identification and environmental risk quantification method based on multiple omics. The method comprises the following steps: extracting and sequencing total DNA (Deoxyribonucleic Acid) and total RNA (Ribonucleic Acid) of a collected compost sample; by analyzing metagenome data, reconstructing an N2O-generated microbial genome and performing function annotation on the N2O-generated microbial genome; according to a functional gene set annotated by a genome, dividing four generation ways of nitrification, nitrifying bacteria denitrification, heterotrophic denitrification and reduction of dissimilatory nitrate into ammonium; and analyzing gene expression dynamics in the genome by combining with a metatranscriptome and carrying out risk quantification. Compared with targeted gene detection technologies such as PCR (polymerase chain reaction) and the like, the method has the advantages that the generation way and the emission potential of N2O in a composting system are accurately recognized by fusing the metagenome and the metatranscriptome, the limitation of a traditional method on the gene coverage degree is broken through, and the method has important application value on accurate management and control of organic solid waste engineering greenhouse gases.
Owner:SUN YAT SEN UNIV