Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

58 results about "Functional annotation" patented technology

Enzyme activity site prediction method based on graph neural network

The invention belongs to the technical field of active site prediction, and particularly relates to an enzyme active site prediction method based on a graph neural network. In order to realize high-precision prediction of active sites, the enzyme active site prediction model HFGN adopts a double-branch architecture: enzyme branches integrate a three-dimensional structure, PLM embedding, homology score and EC function annotation, and realize multi-source feature fusion through an attention mechanism; the reaction branch is based on molecular maps of a substrate and a product, chemical reaction specificity is modeled through a map neural network, and finally enzyme-reaction context sensing embedding is realized through a cross-modal attention mechanism, so that the prediction precision and generalization ability of residue-level active sites are improved.
Owner:SHANXI UNIV

Multi-omics data spatial integration and analysis method and system of Wuzhishan pig organs

InactiveCN121117660ABiostatisticsBiological modelsData spacePathway enrichment
The invention provides a Wuzhishan pig organ multi-omics data space integration and analysis method, and belongs to the technical field of data statistics and analysis, the method comprises the following steps: S1, experiment design and Wuzhishan pig organ sample standardization preparation; s2, independently collecting and preprocessing multi-omics data of Wuzhishan pig organs; s3, multi-omics data standardization and batch effect correction; s4, labeling and calibrating spatial dimension information; s5, spatial specificity correlation modeling of the multi-omics data is carried out; s6, carrying out integrated clustering analysis on the spatial multi-omics data; s7, performing function annotation and path enrichment analysis on the clustering feature clusters; and S8, analyzing the spatial specific molecular mechanism and constructing a multi-omics data spatial integration database. According to the method, spatialization and systematization analysis of the Wuzhishan pig organ multi-omics data is achieved through collaborative optimization of the HHO-IK-means-BI three algorithms, and technical support is provided for experimental zoology research and conversion of medical application.
Owner:SANYA RESEARCH INSTITUTE OF HAINAN ACADEMY OF AGRICULTURAL SCIENCES (HAINAN EXPERIMENTAL ANIMAL RESEARCH CENTER)

Porcine SNP chip construction method based on gene regulatory network characteristics and application

The invention discloses a pig SNP chip construction method based on gene regulatory network characteristics, and relates to the technical field of animal genetic breeding, and the pig SNP chip construction method comprises the following steps: constructing a pig tissue specificity multi-level gene regulatory network related to target traits and tissue types; performing function annotation and network comprehensive feature extraction on the whole genome SNP based on the gene regulation network; and carrying out score sorting and screening on the extracted network comprehensive characteristics, screening SNP variation sites with regulation function potential, constructing a pig SNP site set, and preparing the SNP chip. By introducing tissue / character related gene regulatory network information, functional priority ranking and screening are performed on SNP loci in a whole genome range, so that the character interpretation ability and breeding value estimation accuracy of SNP in a chip are remarkably improved.
Owner:AGRI GENOMICS INST CHINESE ACADEMY OF AGRI SCI

Factor analysis system based on long-chain non-coding RNA multi-omics integration analysis

The invention relates to the technical field of bioinformatics and molecular biology, and discloses a factor analysis system based on long-chain non-coding RNA multi-omics integration analysis. The system comprises: a multi-omics data acquisition module configured to acquire multi-modal omics data related to long-chain non-coding RNA; the data pre-processing module is configured to pre-process the multi-modal omics data to generate a data set in a unified format; the unsupervised factor analysis module is configured to integrate and analyze the data set in the unified format and identify potential factors; the heterogeneity analysis module is configured to construct a mapping relation between the potential factors and multiple omics data features; and the factor annotation and expansion analysis module is configured to perform function annotation on the analyzed potential factors based on biological function enrichment analysis, regulation and control network inference and cross-modal data association, perform missing value estimation on the multi-omics data in combination with the mapping relationship, and output an analysis result containing factor annotation information and complete data.
Owner:BOCE BIOMEDICAL (TIANJIN) CO LTD

A macrovirus group analysis method based on second-generation and third-generation sequencing technology

The application discloses a macrovirus group analysis method based on second-generation and third-generation sequencing technologies, which comprises the following steps: performing quality control and filtering on second-generation sequencing raw data to obtain clean reads, performing single-sample and multi-sample assembly on the quality-controlled data to obtain contigs sequences; performing correction and quality control on third-generation sequencing raw data to obtain clean long sequences, performing assembly on the quality-controlled third-generation data to obtain contigs sequences; performing mixed assembly on the quality-controlled data of the second-generation and third-generation sequencing to obtain contigs sequences; merging all contigs to construct a non-redundant contigs set; and finally performing virus identification and determination, virus species annotation and functional annotation. The application provides a reliable macrovirus group analysis method based on second-generation and third-generation sequencing technologies, and the implementation method is simple and the application range is wide.
Owner:HUAZHONG UNIV OF SCI & TECH

Cell infiltration inference method and system fusing go function annotation and ppi network information

ActiveCN121075448BBiostatisticsInference methodsCellCell function
The application relates to a cell infiltration inference method and system fusing GO function annotation and PPI network information, and the method comprises the following steps: collecting gene expression data, GO function annotation data and PPI network data; constructing a cell-cell function correlation network and a cell-cell physical interaction network respectively; performing weighted fusion processing on the two networks to obtain a comprehensive cell relationship network; calculating a final cell infiltration score through a restart walk algorithm, and inferring the infiltration degree in a tumor microenvironment according to the final cell infiltration score. The application innovatively fuses GO function annotation information and PPI network data, comprehensively considers the functional similarity and physical or signal interaction between cells, enables the model to understand cell synergy from the biological pathway level and analyze cell direct interaction from the protein interaction level, avoids one-sidedness of a single perspective, and provides a more stereoscopic cognitive framework for tumor microenvironment analysis.
Owner:GUANGZHOU UNIVERSITY

Cell infiltration inference method and system fusing GO function annotation and PPI network information

ActiveCN121075448ABiostatisticsInference methodsCellCell function
The invention relates to a cell infiltration inference method and system fusing GO function annotation and PPI network information. The method comprises the following steps: collecting gene expression data, GO function annotation data and PPI network data; respectively constructing a cell * cell function association network and a cell * cell physical interaction network; carrying out weighted fusion processing on the two to obtain a comprehensive cell relation network; and calculating a final cell infiltration fraction through a restart migration algorithm, and deducing the infiltration degree in the tumor microenvironment according to the final cell infiltration fraction. According to the method, GO function annotation information and PPI network data are creatively fused, functional similarity and physical or signal interaction between cells are comprehensively considered, the model can understand cell synergy from the biological pathway level and can analyze cell direct interaction from the protein interaction level, one-sidedness of a single view angle is avoided, and the method has the advantages of being simple in structure and convenient to operate. And a more three-dimensional cognitive framework is provided for tumor microenvironment analysis.
Owner:GUANGZHOU UNIVERSITY

Method for researching influence on hemolytic activity of marine microalgae based on transcriptome technology

The invention discloses a method for researching influence on hemolytic activity of marine microalgae based on a transcriptome technology, and relates to the field of hemolytic activity analys.The method comprises the steps that pretreatment is conducted on the marine microalgae based on culture requirements, and cystic morphological cell density and hemolytic activity of the marine microalgae are measured according to experimental design rules after pretreatment is completed; carrying out transcriptome sample collection operation by utilizing the marine microalgae subjected to pretreatment, constructing a library according to a transcriptome sample collection result, and obtaining a gene function annotation and a differential expression gene result; the growth conditions of the marine microalgae under different temperature conditions are analyzed according to cystic morphological cell density and hemolytic activity, and the regulatory gene influencing the hemolytic activity of the marine microalgae is obtained by combining gene function annotation and differential expression gene results. According to the method, the hemolytic activity of the marine microalgae cultured under different temperature conditions is extracted, so that the aim of researching related metabolic pathways possibly participating in toxin synthesis and hemolytic activity regulation is fulfilled.
Owner:GUANGXI ACAD OF SCI

A method for constructing a genetic risk prediction model integrating functional annotation information and its application

The present invention belongs to the field of genetic risk prediction technology and discloses a method for constructing a genetic risk prediction model integrating functional annotation information and its application, comprising: calculating the heritability of the specific annotations of each cell type corresponding to each tissue of the sample to be tested based on the marginal chi-square statistic of each SNP, and then obtaining the heritability of the j-th SNP in the current tissue f and estimating the joint effect size b of M SNPs in tissue f. f , and based on b f The covariance between the phenotype vector y of the n samples to be tested is b f The optimal linear unbiased estimate of the genetic risk prediction model is then obtained. Furthermore, the genetic risk scores corresponding to each tissue are integrated, and the functional annotation information of each tissue is incorporated into the genetic risk prediction. A genetic risk prediction method integrating functional annotation information is also provided. The present invention can improve the accuracy of genetic risk prediction.
Owner:HUAZHONG UNIV OF SCI & TECH

Protein function prediction model generation method and system

PendingCN121662146ABiostatisticsBiological modelsProtein targetProtein function prediction
The invention discloses a protein function prediction model generation method and system, and belongs to the technical field of bioinformatics, and the method specifically comprises the steps: firstly, extracting a natural disordered region of a target protein and experimental condition parameters thereof; retrieving a conformation database based on the regional features to obtain a reference protein set with multi-condition conformations and function annotations; thirdly, a condition dependent network fusing the target and the reference conformation is constructed, and the network edge weight is jointly determined by the conformation conversion difficulty and the experiment condition matching degree; if the target conformation cannot be effectively connected with the reference conformation, virtual folding simulation is started to generate supplementary conformation nodes; integrating all nodes to generate a conformation state relation graph, and training a prediction model according to the conformation state relation graph; and finally outputting a model which can receive a new protein sequence and condition parameters thereof and predict possible conformation states and corresponding functions thereof. According to the method, condition-dependent conformation and function association prediction can be realized aiming at natural disordered region protein.
Owner:LANJIATANG BIOLOGICAL MEDICINE FUJIAN CO LTD

A High-Throughput Intelligent Screening Method for Multi-Track Ginseng Breeding Materials Based on Deep Learning

ActiveCN122090927ABiostatisticsBiological modelsMultiple traitsScreening method
This invention relates to the field of intelligent breeding technology, specifically disclosing a high-throughput intelligent screening method for multi-trait ginseng breeding materials based on deep learning. The method includes breeding data acquisition, high-throughput acquisition of hyperspectral phenotypes, construction of a multi-trait prediction model, model interpretability analysis, comprehensive screening of multiple traits, and verification of screening results. This scheme utilizes UAV hyperspectral remote sensing technology to achieve rapid and non-destructive determination of saponin content, significantly improving the throughput of phenotypic data acquisition. A multi-task learning architecture and deep neural networks are used to construct a multi-trait prediction model. By sharing representation layers, genetic associations between traits are captured, while retaining the specific information of each trait, achieving collaborative and accurate prediction of multiple traits. Through integrated gradient and attention weight analysis, key marker sites are accurately located, and their biological significance is revealed by functional annotation, enhancing the model's transparency and credibility.
Owner:CHANGCHUN UNIV OF CHINESE MEDICINE

Biological function feature fusion method based on layered weight perception attention mechanism

PendingCN121811977ABiostatisticsBiological modelsFunctional semanticsFeature fusion
The invention relates to the technical field of artificial intelligence and bioinformatics, and discloses a biological function feature fusion method based on a layered weight perception attention mechanism. The method comprises the following steps: acquiring gene sequence data and multi-source function annotations, and converting into vector representation by using a pre-training model; fusing the functional semantics in the library into a library-level summary vector by adopting a hierarchical fusion mechanism of weight perception; constructing a hierarchical fusion architecture, and realizing deep adaptive fusion of a gene sequence and multi-source knowledge through a shared projection layer, a cross attention mechanism and a gating network; and the model performance is improved through end-to-end multi-objective optimization training. According to the method, the problems of semantic gap, insufficient weight utilization and biological logic deficiency in multi-source functional data fusion are solved, the unified gene function feature vector with high biological characterization capability is generated, and the accuracy and interpretability of gene function annotation are improved.
Owner:CHONGQING JIAOTONG UNIV

Pit mud metagenome data automatic analysis method and system

The invention relates to the technical field of metagenomics, discloses an automatic analysis method and system for pit mud metagenomic data, and aims at solving the problem that an existing method is poor in efficiency and accuracy, and the scheme mainly comprises the steps that a sequencing data type, a file path and analysis parameters are received; performing quality control on the original offline data; sequence assembly is carried out, and a contigs file is generated; carrying out assembly quality evaluation on the contigs file; carrying out genome binning by using at least two binning tools; integrating output results of the binning tool, and performing optimization based on a preset integrity threshold value and a preset pollution degree threshold value to obtain an optimized binning genome data set; evaluating and optimizing the integrity, the pollution degree and the strain heterogeneity of the binning genome; calculating coverage and relative abundance; performing species classification annotation and function annotation; and integrating the result data of the previous steps to generate an analysis report. According to the method, automatic analysis of metagenome data is realized, and the analysis efficiency and accuracy are improved.
Owner:WULIANGYE +1

Gene data processing method and device, computer device and storage medium

The application discloses a gene data processing method and device, computer equipment and a storage medium, and belongs to the technical field of computers. Through a gene function query request of a to-be-tested gene, the application can call a gene association model corresponding to a cell type to which the to-be-tested gene belongs, mine a nonlinear association degree between the to-be-tested gene and a known candidate gene, and use the candidate gene with a higher nonlinear association degree to label function annotation information of the to-be-tested gene. The way of calling the gene association model to extract the nonlinear relationship is completely different from the way of extracting the linear relationship in traditional statistics, can deeply mine the candidate gene with a higher similarity to the to-be-tested gene, and the similarity is not a linear similarity but an implicit nonlinear similarity, so that the accuracy of the gene data processing process is greatly improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Research method and system for immunomodulatory effect of TPD52

The embodiment of the invention provides a TPD52 immunomodulatory effect research method and system, and the method comprises the steps: obtaining single cell transcriptome data of a breast cancer tumor microenvironment; on the basis of the data, identifying the TPD52 as the immune regulation key dangerous gene by utilizing a machine learning algorithm; verifying the association of TPD52 expression and patient prognosis in a multi-center queue; and based on the function annotation and the body appearance type experiment of the TPD52, confirming the immunomodulatory effect of the TPD52, and outputting a comprehensive evaluation report of the immunomodulatory effect of the TPD52. According to the invention, a machine learning method and scRNA-seq are utilized to explore the effect of TPD52 as a key immunomodulatory factor in BRCA, and the application has important significance on tumor behavior and patient prognosis.
Owner:TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH

Gene function prediction method based on semantic correspondence of regulatory region

The invention discloses a gene function prediction method based on semantic correspondence of a regulatory region. The method comprises the following steps: firstly, constructing an inter-species regulation semantic correspondence relationship data set, constructing an artificial intelligence model structure, then, constructing a cross-species semantic correspondence network, and finally, carrying out function annotation on a target gene or identifying a candidate gene with a specific function in a target species. The accuracy of the PhytoBabel model constructed by the method is obviously higher than that of other model structures. By utilizing the method disclosed by the invention, the genes ZmERF104 and ZmGRF16 for promoting the regeneration of the corn somatic embryos and the gene ZmNAC17 for inhibiting the regeneration of the somatic embryos, which cannot be found by the traditional method, are successfully identified.
Owner:CHINA AGRI UNIV

A method of predicting pathogenic germline mutations

PendingCN122157767AMathematical modelsBiostatisticsGermline mutationRisk classification
The application provides a prediction method of pathogenic embryonic mutation, comprising the following steps: obtaining a variation site set of an embryonic mutation to be predicted, and performing population frequency screening on the variation site; filtering variation sites which are not annotated or are only annotated as clinically uncertain by using a clinical variation database; calculating a pathogenic posterior probability score of a candidate pathogenic variation by using a variation Bayesian inference (VBI) model with multiple types of function annotations; and cross- verifying the posterior probability score, external pathogenic prediction tool scores and variation types to determine a final pathogenicity discrimination result. Compared with an existing method based on single site frequency or single function score, the present application provides a posterior probability evaluation model with more biological interpretation in the field of rare variation judgment, effectively improves the prediction accuracy in a complex scene, significantly improves the risk classification accuracy based on multi-level evidence fusion discrimination, and effectively reduces the misdiagnosis and missed diagnosis rate in clinical stratification.
Owner:NANJING MEDICAL UNIV

A protein characterization learning method and device, computer equipment and storage medium

PendingCN122455104AData setSmall sample
The application relates to a protein characterization learning method and device, computer equipment and a storage medium. The method comprises the following steps: collecting protein sequence, structure and function annotation information, and constructing a multi-modal feature data set; performing multi-modal feature extraction on the multi-modal feature data set by using a MASSA framework to obtain protein multi-modal representation; constructing a graph network by using the protein multi-modal representation, and storing the representation and a historical model of an existing protein task in the graph network; in the training process of a new protein task, a memory enhancement mechanism is used to retrieve a similar node of the new protein task from the graph network, a historical model of an existing protein task corresponding to the similar node is obtained, model parameters of the historical model are transferred by using a transfer learning framework, and model training is performed on the new protein task. The application effectively alleviates the data scarcity problem and improves the accuracy and generalization ability of the model in a small sample learning task.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Genome sequence generation and model training methods and related products

This disclosure provides a method and related products for genome sequence generation and model training. During the training phase, pre-trained multimodal data, including sequence data, functional annotation data, omics data, and species evolution information, is acquired. A fused representation is obtained through heterogeneous coding and multimodal fusion, and input into the sequence generation model for feature calculation, outputting the predicted base probability distribution for each base position to be generated. The model is optimized by combining a multi-task training mode of autoregressive generation and intermediate padding, and a weighted loss function based on biological annotation. During the inference phase, sequences are generated progressively based on multimodal generation context information. Candidate bases are obtained through sampling and evaluated using biological constraints, omics validation, structural validation, and functional validation. Resampling is performed when conditions are not met until preset conditions are satisfied. This method can improve the biological rationality and application value of the generated sequences.
Owner:BIOMAP (BEIJING) INTELLIGENCE TECH LTD

Enzyme catalytic conversion number prediction method and system based on function annotation and hierarchical structure

The invention belongs to the technical field of enzyme catalytic conversion number prediction, and discloses an enzyme catalytic conversion number prediction method and system based on function annotation and a hierarchical structure, and the prediction method comprises the steps: extracting protein unique identifier information based on a protein resource database, analyzing protein dynamic information and gene ontology functions according to the unique identifier information, and obtaining a prediction result; obtaining a relation table of the gene ontology and the enzymatic conversion coefficient; combining the hierarchical structure of the gene ontology with the relation table, capturing the mutual relation between the gene and the gene product, and constructing a hierarchical total tree between the gene ontology and the enzymatic conversion coefficient according to a relation capturing result; and extracting target gene ontology information according to the gene code of the target object, and matching an enzymatic conversion coefficient corresponding to the target gene ontology information from the hierarchical total tree as an enzymatic conversion number prediction result. According to the invention, gene and protein function annotations are provided by using the gene ontology, so that the function similarity of enzymes can be measured based on the gene ontology, and the enzymatic conversion value of unknown enzymes is speculated.
Owner:TIANJIN INST OF IND BIOTECH CHINESE ACADEMY OF SCI

Single-cell rna sequencing annotation method and device based on dynamic hypergraph

The application provides a single-cell RNA sequencing annotation method and device based on a dynamic hypergraph, relates to the technical field of bioinformatics, and comprises the following steps: obtaining a data set, extracting a low-dimensional embedding vector of each cell from a gene expression vector, and constructing a dynamic hypergraph with cells as nodes and gene pathways as hyperedges; extracting a pathway feature of each hyperedge from the dynamic hypergraph; calculating the importance weight of each cell in each hyperedge based on the low-dimensional embedding vector and the pathway feature of the hyperedge; inputting the dynamic hypergraph, the pathway feature and the importance weight into a preset hypergraph neural network for message aggregation and feature learning to generate a prediction label of a cell type; and training the hypergraph neural network through an optimization algorithm to obtain a trained cell annotation model. The application realizes more accurate and more biologically interpretable cell type and function annotation by introducing a hypergraph structure and a metabolic pathway activity dynamic modeling mechanism.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Integrated method of combining target gene set and kegg weight for downstream focusing analysis of differential proteins

The application provides a proteomics data mining-function annotation analysis method combining a target gene set and KEGG weight, the method focuses on the differential proteins related to the target gene set, and most of the interference information is excluded through the target gene set. The method can reduce the noise of irrelevant proteins, more efficiently screen key target proteins and perform preliminary function annotation.
Owner:LOTUSLAKE BIOMEDICAL TECH CO LTD

Pseudoxanthomonas strain JC1303 and application thereof

The invention relates to the technical field of new strains for degrading cellulose, in particular to a pseudoxanthomonas strain JC1303 and application thereof. Through whole genome sequencing and function annotation, it is systematically revealed for the first time that the strain carries a complete cellulose degrading enzyme system including incision beta-1, 4-glucanase, cellulase and beta-glucosidase, has a plurality of central metabolic pathways supporting efficient degradation and utilization of cellulose, and can be used for degrading cellulose. And specific gene resources are determined through generic genome analysis. The strain has the remarkable technical effects that the strain shows higher cellulase activity after being cultured for 5 days, and cellulose materials such as agricultural wastes and the like can be stably and efficiently degraded in a salt-containing environment by virtue of the unique genetic background and marine source characteristics of the strain.
Owner:ZHEJIANG OCEAN UNIV

Virus and disease association prediction model, construction method of prediction model and prediction method

The invention relates to the technical field of interactive prediction, in particular to a virus and disease associated prediction model, a construction method of the prediction model and a prediction method. Comprising the following steps: S1, collecting and processing virus and disease data; s2, performing feature extraction of the independent prediction model A; s3, performing feature extraction of the independent prediction model B; s4, based on five-fold cross validation and grid search, determining an optimal machine learning algorithm, and constructing an independent prediction model A and an independent prediction model B; and S5, based on a bagging integration strategy, integrating the independent prediction model A and the independent prediction model B, and generating a virus and disease associated prediction model. According to the invention, the VDP provides a new thought of multi-modal biological information integration; meanwhile, molecular sequence information of viruses, semantic structure information of diseases and GO function annotation information of the two parties are considered, more comprehensive feature representation is achieved, and the problems that an existing method is single in information source and insufficient in prediction accuracy are solved.
Owner:HUNAN UNIV

Method, device, equipment and medium for screening and detecting key genes of beef production of Xianan cattle

The invention discloses a method, a device, equipment and a medium for screening and detecting key genes of beef produced by Xianan cattle. The method comprises the following steps: preprocessing genome data, apparent group data and beef phenotype data; screening specific mononucleotide polymorphic sites differentiated from other beef cattle varieties through a dynamic population differentiation index threshold value, performing function annotation, performing correlation analysis on the sites subjected to the function annotation and meat production phenotype data through a three-order analysis model combined with captured data verification to obtain candidate key genes, and then identifying the candidate key genes according to the candidate key genes. Detecting the relative expression quantity of the candidate key gene in the muscular tissue of the Xianan cattle through quantitative polymerase chain reaction, and analyzing and verifying the relevance between the candidate key gene and the meat production phenotype in combination with Pearson correlation; specific polymerase chain reaction primers are designed for the key genes passing verification, and detection of the key genes is achieved through fluorescent quantitative polymerase chain reaction.
Owner:河南省种业发展中心

Function annotation abundance sequence-based base model training method and device

The invention relates to a base model training method and device based on a functional annotation abundance sequence. The method comprises the following steps: S1, carrying out function annotation on an open reading frame of a genome or metagenome sample; s2, counting the occurrence frequency of each annotation and constructing a sequence according to an abundance descending order; s3, after the sequence is subjected to token processing, inputting the sequence into a model based on Transform, and adopting joint training of language modeling, comparative learning and classification loss to obtain species-level and token-level fixed dimension embedding; s4, on the basis of the embedding, completing downstream tasks such as phylogenetic tree construction, species identification and phenotype prediction, BGC / MGC recognition and key gene positioning in three levels, namely a genome level, a gene cluster level and a gene / protein level. In the embodiment of the invention, good uniformity, expandability and interpretability are displayed, and the dependence on a reference database is reduced. The corresponding device comprises a data processing module, a model training module and an application module, and can be realized by program instructions in a computer readable storage medium.
Owner:ZHEJIANG LAB

Integrated genome analysis method based on low-depth sequencing

PendingCN121999865AProteomicsGenomicsGeneticsSequence variation
The invention relates to the field of biological medicine, and discloses an integrated genome analysis method based on low-depth sequencing, and the method comprises the following steps: obtaining whole genome low-depth sequencing data; after quality control and comparison, SNV, Indel and CNV are calculated; performing multi-dimensional function annotation by combining genome position, coding influence, splicing disturbance, conservative property, regulatory element and three-dimensional chromatin interaction; database information such as ClinVar and HGMD is integrated, and according to a phenotype-driven rule engine, a clinical interpretable report is generated according to the ACMG / AMP standard. According to the method, a complete analysis chain covering sequence variation detection, multi-dimensional function annotation, three-dimensional genome association, public database integration and phenotype driven interpretation is constructed, so that the fundamental defect that only an original variation list is output and a clinical action basis cannot be provided in traditional low-depth sequencing is overcome.
Owner:PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY)

Method for analyzing functional information of marine microorganisms in marine samples

The present application relates to a kind of functional information analysis method of marine microorganism in marine sample, based on liquid mass spectrometry obtains mass spectrum RAW file, utilize the peptide sequence obtained by macrogenomics sequencing, the database obtained by self-defined configuration or the database published in public is combined as macro proteome data search database.Combined with multiple sources of databases, such as databases published in public, databases obtained by self-defined configuration, and peptide sequences obtained by macrogenomics sequencing, as macro proteome data search database.The database is split according to the classification level of species, and the sub-database is reduced by iterative search method, and then the sub-database is combined, and the qualitative and quantitative analysis of macro proteome is carried out by using proteome search analysis software, and the function annotation software is used for function annotation based on public database.The proteome quantitative analysis software screens different environmental difference proteins, and deeply excavates the community composition and metabolic activity difference of different environmental microorganisms.
Owner:DALIAN INSTITUTE OF CHEMICAL PHYSICS CHINESE ACADEMY OF SCIENCES

A method for screening cross-species HGT

The present invention belongs to the field of bioinformatics technology and specifically relates to a method for screening cross-species HGT. The method includes the steps of constructing a background database, inferring homology groups, constructing a gene phylogenetic tree, inferring HGT based on the gene tree, verifying candidate HGT inferred based on the gene tree, and assigning HGT to a timeline. By integrating a multi-step process such as phylogenetic analysis, sequence alignment, gene function annotation, and statistical verification, the method achieves high-throughput, automated screening of HGT events in large genetic datasets. The method of the present invention is particularly suitable for studying HGT across extremely long time periods and extremely long evolutionary distances across the biological kingdom, and can provide an important scientific tool for understanding gene flow and functional enhancement in the course of biological evolution.
Owner:YELLOW SEA FISHERIES RES INST CHINESE ACAD OF FISHERIES SCI

Graph attention-based lncrna function prediction method, storage medium and device

The lncRNA function prediction method based on graph attention, storage medium and equipment belong to the field of biological information technology. In order to solve the problem that the prior lncRNA function annotation data is little in the existing lncRNA function prediction method, the protein function annotation data needs to be borrowed to indirectly predict the function of lncRNA, resulting in low accuracy. The application first extracts various lncRNA similarity information by using a graph contrast learning method based on a cross attention mechanism, generates a comprehensive similarity matrix; then extracts the features of GO terms in the GO graph by using a knowledge graph embedding model, generates a GO semantic similarity matrix; finally, a graph representation learning model combining GCN and GAT is constructed to learn the features of lncRNA and GO, and a KAN classifier is used to predict the GO function annotation information of lncRNA. The method can effectively improve the accuracy of lncRNA function prediction.
Owner:NORTHEAST FORESTRY UNIV