Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

82 results about "Functional annotation" patented technology

Low-altitude airspace three-dimensional traffic light management method, system and device and computer readable storage medium

The invention provides a low-altitude airspace three-dimensional traffic light management method, system and device and a computer readable storage medium. The method comprises the steps of obtaining genome sequence data to be annotated; setting annotation parameters, and carrying out structure annotation on the genome sequence data; and performing function annotation on the genome subjected to structure annotation. In the scheme provided by the embodiment of the invention, by dividing the airspace into the grids, calculating the airspace complexity entropy of each grid in real time, mapping the complexity entropy value into the traffic light state, and controlling the flight state of the unmanned aerial vehicle in the grids, the traffic flow of the three-dimensional airspace is effectively managed and controlled, the required parameters are simple, and the implementation is easy. The method can adapt to complex scenes, has the advantages of real-time performance, three-dimensional performance, wide application range and the like, and can remarkably improve the efficiency and safety of low-altitude traffic organization.
Owner:BEI DOU FU XI XIN XI JI SHU YOU XIAN GONG SI

Method and system for clustering transcriptome sequencing data

The invention provides a clustering method and system for transcriptome sequencing data. The method is applied to the technical field of data processing, and comprises the following steps: collecting cervical adenocarcinoma and para-carcinoma tissue specimens of a plurality of patients, and carrying out data preprocessing to obtain a standardized gene expression matrix; performing clustering operation on the standardized gene expression matrix based on a clustering analysis method combining a potential category model and sub-alliance division to obtain a clustered gene set; performing gene function annotation on the clustered gene set, and performing correlation analysis on a clustering result after function annotation and clinical characteristics of cervical adenocarcinoma to obtain correlation between gene expression and the clinical characteristics; and on the basis of clustering analysis and correlation analysis results, constructing a deep clustering prediction model based on a variational auto-encoder and a Gamma hybrid model so as to predict the prognosis risk or chemotherapy sensitivity of the patient. The problems that high-dimensional transcriptome data is poor in clustering stability and low in prognosis prediction precision are solved.
Owner:THE AFFILIATED HOSPITAL OF SOUTHWEST MEDICAL UNIV +1

Big model technology-based biological information analysis system

The invention relates to the technical field of bioinformatics, in particular to a biological information analysis system based on a large model technology, which comprises a data analysis calibration module, a problem disassembly module, an analysis task arrangement module, a result mapping module and a feedback iteration module. According to the method, genome comparison and clinical phenotype timestamps are dynamically calibrated, time sequence dislocation deviation is eliminated, base complementary pairing is combined with protein network anomaly screening, low-abundance collaborative variation capture is enhanced, genotype-phenotype discrete distribution quantifies and unifies multi-modal data benchmark, and the problem of multi-source heterogeneous standardization deficiency is solved; the method comprises the following steps: classifying and integrating pathogenic gene semantic weights by structural variation, balancing a statistical threshold and a biological function, dynamically optimizing an analysis sequence, synchronously covering a key mutation region, improving function annotation of a non-coding region, integrating gene expression clustering and protein network topology in a three-dimensional distribution manner, breaking through two-dimensional space limitation, performing closed-loop feedback to correct a threshold iteration elimination rule, and finally obtaining a high-quality gene expression cluster. And the genetic heterogeneity false positive rate is reduced.
Owner:GUANXUN (HANGZHOU) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Method for automated topology recognition and functional annotation of analog / mixed-signal circuits

A system and method for automated topology recognition and functional annotation of a mixed-signal circuit is disclosed. The method includes extracting structural information available from the mixed-signal circuit. The method includes identifying functionality of a sub-circuit of the mixed-signal circuit based on the extracted structural information by comparing the sub-circuit with a library cell. The method includes annotating the extracted structural information with the identified functionality corresponding to the sub-circuit. The method includes automatically generating a plurality of configurations corresponding to the annotated structural information.
Owner:SYNOPSYS INC

Enzyme activity site prediction method based on graph neural network

The invention belongs to the technical field of active site prediction, and particularly relates to an enzyme active site prediction method based on a graph neural network. In order to realize high-precision prediction of active sites, the enzyme active site prediction model HFGN adopts a double-branch architecture: enzyme branches integrate a three-dimensional structure, PLM embedding, homology score and EC function annotation, and realize multi-source feature fusion through an attention mechanism; the reaction branch is based on molecular maps of a substrate and a product, chemical reaction specificity is modeled through a map neural network, and finally enzyme-reaction context sensing embedding is realized through a cross-modal attention mechanism, so that the prediction precision and generalization ability of residue-level active sites are improved.
Owner:SHANXI UNIV

Multi-omics data spatial integration and analysis method and system of Wuzhishan pig organs

InactiveCN121117660ABiostatisticsBiological modelsData spacePathway enrichment
The invention provides a Wuzhishan pig organ multi-omics data space integration and analysis method, and belongs to the technical field of data statistics and analysis, the method comprises the following steps: S1, experiment design and Wuzhishan pig organ sample standardization preparation; s2, independently collecting and preprocessing multi-omics data of Wuzhishan pig organs; s3, multi-omics data standardization and batch effect correction; s4, labeling and calibrating spatial dimension information; s5, spatial specificity correlation modeling of the multi-omics data is carried out; s6, carrying out integrated clustering analysis on the spatial multi-omics data; s7, performing function annotation and path enrichment analysis on the clustering feature clusters; and S8, analyzing the spatial specific molecular mechanism and constructing a multi-omics data spatial integration database. According to the method, spatialization and systematization analysis of the Wuzhishan pig organ multi-omics data is achieved through collaborative optimization of the HHO-IK-means-BI three algorithms, and technical support is provided for experimental zoology research and conversion of medical application.
Owner:SANYA RESEARCH INSTITUTE OF HAINAN ACADEMY OF AGRICULTURAL SCIENCES (HAINAN EXPERIMENTAL ANIMAL RESEARCH CENTER)

Porcine SNP chip construction method based on gene regulatory network characteristics and application

The invention discloses a pig SNP chip construction method based on gene regulatory network characteristics, and relates to the technical field of animal genetic breeding, and the pig SNP chip construction method comprises the following steps: constructing a pig tissue specificity multi-level gene regulatory network related to target traits and tissue types; performing function annotation and network comprehensive feature extraction on the whole genome SNP based on the gene regulation network; and carrying out score sorting and screening on the extracted network comprehensive characteristics, screening SNP variation sites with regulation function potential, constructing a pig SNP site set, and preparing the SNP chip. By introducing tissue / character related gene regulatory network information, functional priority ranking and screening are performed on SNP loci in a whole genome range, so that the character interpretation ability and breeding value estimation accuracy of SNP in a chip are remarkably improved.
Owner:AGRI GENOMICS INST CHINESE ACADEMY OF AGRI SCI

Factor analysis system based on long-chain non-coding RNA multi-omics integration analysis

The invention relates to the technical field of bioinformatics and molecular biology, and discloses a factor analysis system based on long-chain non-coding RNA multi-omics integration analysis. The system comprises: a multi-omics data acquisition module configured to acquire multi-modal omics data related to long-chain non-coding RNA; the data pre-processing module is configured to pre-process the multi-modal omics data to generate a data set in a unified format; the unsupervised factor analysis module is configured to integrate and analyze the data set in the unified format and identify potential factors; the heterogeneity analysis module is configured to construct a mapping relation between the potential factors and multiple omics data features; and the factor annotation and expansion analysis module is configured to perform function annotation on the analyzed potential factors based on biological function enrichment analysis, regulation and control network inference and cross-modal data association, perform missing value estimation on the multi-omics data in combination with the mapping relationship, and output an analysis result containing factor annotation information and complete data.
Owner:BOCE BIOMEDICAL (TIANJIN) CO LTD

A macrovirus group analysis method based on second-generation and third-generation sequencing technology

The application discloses a macrovirus group analysis method based on second-generation and third-generation sequencing technologies, which comprises the following steps: performing quality control and filtering on second-generation sequencing raw data to obtain clean reads, performing single-sample and multi-sample assembly on the quality-controlled data to obtain contigs sequences; performing correction and quality control on third-generation sequencing raw data to obtain clean long sequences, performing assembly on the quality-controlled third-generation data to obtain contigs sequences; performing mixed assembly on the quality-controlled data of the second-generation and third-generation sequencing to obtain contigs sequences; merging all contigs to construct a non-redundant contigs set; and finally performing virus identification and determination, virus species annotation and functional annotation. The application provides a reliable macrovirus group analysis method based on second-generation and third-generation sequencing technologies, and the implementation method is simple and the application range is wide.
Owner:HUAZHONG UNIV OF SCI & TECH

Cell infiltration inference method and system fusing go function annotation and ppi network information

ActiveCN121075448BBiostatisticsInference methodsCellCell function
The application relates to a cell infiltration inference method and system fusing GO function annotation and PPI network information, and the method comprises the following steps: collecting gene expression data, GO function annotation data and PPI network data; constructing a cell-cell function correlation network and a cell-cell physical interaction network respectively; performing weighted fusion processing on the two networks to obtain a comprehensive cell relationship network; calculating a final cell infiltration score through a restart walk algorithm, and inferring the infiltration degree in a tumor microenvironment according to the final cell infiltration score. The application innovatively fuses GO function annotation information and PPI network data, comprehensively considers the functional similarity and physical or signal interaction between cells, enables the model to understand cell synergy from the biological pathway level and analyze cell direct interaction from the protein interaction level, avoids one-sidedness of a single perspective, and provides a more stereoscopic cognitive framework for tumor microenvironment analysis.
Owner:GUANGZHOU UNIVERSITY

Cell infiltration inference method and system fusing GO function annotation and PPI network information

ActiveCN121075448ABiostatisticsInference methodsCellCell function
The invention relates to a cell infiltration inference method and system fusing GO function annotation and PPI network information. The method comprises the following steps: collecting gene expression data, GO function annotation data and PPI network data; respectively constructing a cell * cell function association network and a cell * cell physical interaction network; carrying out weighted fusion processing on the two to obtain a comprehensive cell relation network; and calculating a final cell infiltration fraction through a restart migration algorithm, and deducing the infiltration degree in the tumor microenvironment according to the final cell infiltration fraction. According to the method, GO function annotation information and PPI network data are creatively fused, functional similarity and physical or signal interaction between cells are comprehensively considered, the model can understand cell synergy from the biological pathway level and can analyze cell direct interaction from the protein interaction level, one-sidedness of a single view angle is avoided, and the method has the advantages of being simple in structure and convenient to operate. And a more three-dimensional cognitive framework is provided for tumor microenvironment analysis.
Owner:GUANGZHOU UNIVERSITY

Method for researching influence on hemolytic activity of marine microalgae based on transcriptome technology

The invention discloses a method for researching influence on hemolytic activity of marine microalgae based on a transcriptome technology, and relates to the field of hemolytic activity analys.The method comprises the steps that pretreatment is conducted on the marine microalgae based on culture requirements, and cystic morphological cell density and hemolytic activity of the marine microalgae are measured according to experimental design rules after pretreatment is completed; carrying out transcriptome sample collection operation by utilizing the marine microalgae subjected to pretreatment, constructing a library according to a transcriptome sample collection result, and obtaining a gene function annotation and a differential expression gene result; the growth conditions of the marine microalgae under different temperature conditions are analyzed according to cystic morphological cell density and hemolytic activity, and the regulatory gene influencing the hemolytic activity of the marine microalgae is obtained by combining gene function annotation and differential expression gene results. According to the method, the hemolytic activity of the marine microalgae cultured under different temperature conditions is extracted, so that the aim of researching related metabolic pathways possibly participating in toxin synthesis and hemolytic activity regulation is fulfilled.
Owner:GUANGXI ACAD OF SCI

A method for constructing a genetic risk prediction model integrating functional annotation information and its application

The present invention belongs to the field of genetic risk prediction technology and discloses a method for constructing a genetic risk prediction model integrating functional annotation information and its application, comprising: calculating the heritability of the specific annotations of each cell type corresponding to each tissue of the sample to be tested based on the marginal chi-square statistic of each SNP, and then obtaining the heritability of the j-th SNP in the current tissue f and estimating the joint effect size b of M SNPs in tissue f. f , and based on b f The covariance between the phenotype vector y of the n samples to be tested is b f The optimal linear unbiased estimate of the genetic risk prediction model is then obtained. Furthermore, the genetic risk scores corresponding to each tissue are integrated, and the functional annotation information of each tissue is incorporated into the genetic risk prediction. A genetic risk prediction method integrating functional annotation information is also provided. The present invention can improve the accuracy of genetic risk prediction.
Owner:HUAZHONG UNIV OF SCI & TECH

Protein function prediction model generation method and system

PendingCN121662146ABiostatisticsBiological modelsProtein targetProtein function prediction
The invention discloses a protein function prediction model generation method and system, and belongs to the technical field of bioinformatics, and the method specifically comprises the steps: firstly, extracting a natural disordered region of a target protein and experimental condition parameters thereof; retrieving a conformation database based on the regional features to obtain a reference protein set with multi-condition conformations and function annotations; thirdly, a condition dependent network fusing the target and the reference conformation is constructed, and the network edge weight is jointly determined by the conformation conversion difficulty and the experiment condition matching degree; if the target conformation cannot be effectively connected with the reference conformation, virtual folding simulation is started to generate supplementary conformation nodes; integrating all nodes to generate a conformation state relation graph, and training a prediction model according to the conformation state relation graph; and finally outputting a model which can receive a new protein sequence and condition parameters thereof and predict possible conformation states and corresponding functions thereof. According to the method, condition-dependent conformation and function association prediction can be realized aiming at natural disordered region protein.
Owner:LANJIATANG BIOLOGICAL MEDICINE FUJIAN CO LTD

A High-Throughput Intelligent Screening Method for Multi-Track Ginseng Breeding Materials Based on Deep Learning

ActiveCN122090927ABiostatisticsBiological modelsMultiple traitsScreening method
This invention relates to the field of intelligent breeding technology, specifically disclosing a high-throughput intelligent screening method for multi-trait ginseng breeding materials based on deep learning. The method includes breeding data acquisition, high-throughput acquisition of hyperspectral phenotypes, construction of a multi-trait prediction model, model interpretability analysis, comprehensive screening of multiple traits, and verification of screening results. This scheme utilizes UAV hyperspectral remote sensing technology to achieve rapid and non-destructive determination of saponin content, significantly improving the throughput of phenotypic data acquisition. A multi-task learning architecture and deep neural networks are used to construct a multi-trait prediction model. By sharing representation layers, genetic associations between traits are captured, while retaining the specific information of each trait, achieving collaborative and accurate prediction of multiple traits. Through integrated gradient and attention weight analysis, key marker sites are accurately located, and their biological significance is revealed by functional annotation, enhancing the model's transparency and credibility.
Owner:CHANGCHUN UNIV OF CHINESE MEDICINE

Biological function feature fusion method based on layered weight perception attention mechanism

PendingCN121811977ABiostatisticsBiological modelsFunctional semanticsFeature fusion
The invention relates to the technical field of artificial intelligence and bioinformatics, and discloses a biological function feature fusion method based on a layered weight perception attention mechanism. The method comprises the following steps: acquiring gene sequence data and multi-source function annotations, and converting into vector representation by using a pre-training model; fusing the functional semantics in the library into a library-level summary vector by adopting a hierarchical fusion mechanism of weight perception; constructing a hierarchical fusion architecture, and realizing deep adaptive fusion of a gene sequence and multi-source knowledge through a shared projection layer, a cross attention mechanism and a gating network; and the model performance is improved through end-to-end multi-objective optimization training. According to the method, the problems of semantic gap, insufficient weight utilization and biological logic deficiency in multi-source functional data fusion are solved, the unified gene function feature vector with high biological characterization capability is generated, and the accuracy and interpretability of gene function annotation are improved.
Owner:CHONGQING JIAOTONG UNIV

A genomic data analysis method

The present invention relates to the field of data processing technology, and in particular to a genomic data analysis method. The method comprises the following steps: preprocessing genomic data, performing intelligent compression processing on the preprocessed genomic data using an adaptive multidimensional space compression algorithm to obtain compressed genomic data; performing gene variation detection and filtering on the compressed genomic data to obtain variant genomic data, and performing functional annotation to obtain genomic data with annotation results; performing gene expression analysis on the genomic data with annotation results to obtain gene expression data; and performing association analysis based on the genomic data with annotation results and gene expression data using a gene variation-phenotype association analysis method to screen out potential disease markers. The method solves the technical problems of inaccurate genomic data processing, low computational efficiency, and low analysis accuracy in the genomic data analysis process of traditional genomic data storage and processing technologies.
Owner:XIDIAN GRP HOSPITAL

Pit mud metagenome data automatic analysis method and system

The invention relates to the technical field of metagenomics, discloses an automatic analysis method and system for pit mud metagenomic data, and aims at solving the problem that an existing method is poor in efficiency and accuracy, and the scheme mainly comprises the steps that a sequencing data type, a file path and analysis parameters are received; performing quality control on the original offline data; sequence assembly is carried out, and a contigs file is generated; carrying out assembly quality evaluation on the contigs file; carrying out genome binning by using at least two binning tools; integrating output results of the binning tool, and performing optimization based on a preset integrity threshold value and a preset pollution degree threshold value to obtain an optimized binning genome data set; evaluating and optimizing the integrity, the pollution degree and the strain heterogeneity of the binning genome; calculating coverage and relative abundance; performing species classification annotation and function annotation; and integrating the result data of the previous steps to generate an analysis report. According to the method, automatic analysis of metagenome data is realized, and the analysis efficiency and accuracy are improved.
Owner:WULIANGYE +1

Gene data processing method and device, computer device and storage medium

The application discloses a gene data processing method and device, computer equipment and a storage medium, and belongs to the technical field of computers. Through a gene function query request of a to-be-tested gene, the application can call a gene association model corresponding to a cell type to which the to-be-tested gene belongs, mine a nonlinear association degree between the to-be-tested gene and a known candidate gene, and use the candidate gene with a higher nonlinear association degree to label function annotation information of the to-be-tested gene. The way of calling the gene association model to extract the nonlinear relationship is completely different from the way of extracting the linear relationship in traditional statistics, can deeply mine the candidate gene with a higher similarity to the to-be-tested gene, and the similarity is not a linear similarity but an implicit nonlinear similarity, so that the accuracy of the gene data processing process is greatly improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Research method and system for immunomodulatory effect of TPD52

The embodiment of the invention provides a TPD52 immunomodulatory effect research method and system, and the method comprises the steps: obtaining single cell transcriptome data of a breast cancer tumor microenvironment; on the basis of the data, identifying the TPD52 as the immune regulation key dangerous gene by utilizing a machine learning algorithm; verifying the association of TPD52 expression and patient prognosis in a multi-center queue; and based on the function annotation and the body appearance type experiment of the TPD52, confirming the immunomodulatory effect of the TPD52, and outputting a comprehensive evaluation report of the immunomodulatory effect of the TPD52. According to the invention, a machine learning method and scRNA-seq are utilized to explore the effect of TPD52 as a key immunomodulatory factor in BRCA, and the application has important significance on tumor behavior and patient prognosis.
Owner:TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH

Gene function prediction method based on semantic correspondence of regulatory region

The invention discloses a gene function prediction method based on semantic correspondence of a regulatory region. The method comprises the following steps: firstly, constructing an inter-species regulation semantic correspondence relationship data set, constructing an artificial intelligence model structure, then, constructing a cross-species semantic correspondence network, and finally, carrying out function annotation on a target gene or identifying a candidate gene with a specific function in a target species. The accuracy of the PhytoBabel model constructed by the method is obviously higher than that of other model structures. By utilizing the method disclosed by the invention, the genes ZmERF104 and ZmGRF16 for promoting the regeneration of the corn somatic embryos and the gene ZmNAC17 for inhibiting the regeneration of the somatic embryos, which cannot be found by the traditional method, are successfully identified.
Owner:CHINA AGRI UNIV

A method of predicting pathogenic germline mutations

PendingCN122157767AMathematical modelsBiostatisticsGermline mutationRisk classification
The application provides a prediction method of pathogenic embryonic mutation, comprising the following steps: obtaining a variation site set of an embryonic mutation to be predicted, and performing population frequency screening on the variation site; filtering variation sites which are not annotated or are only annotated as clinically uncertain by using a clinical variation database; calculating a pathogenic posterior probability score of a candidate pathogenic variation by using a variation Bayesian inference (VBI) model with multiple types of function annotations; and cross- verifying the posterior probability score, external pathogenic prediction tool scores and variation types to determine a final pathogenicity discrimination result. Compared with an existing method based on single site frequency or single function score, the present application provides a posterior probability evaluation model with more biological interpretation in the field of rare variation judgment, effectively improves the prediction accuracy in a complex scene, significantly improves the risk classification accuracy based on multi-level evidence fusion discrimination, and effectively reduces the misdiagnosis and missed diagnosis rate in clinical stratification.
Owner:NANJING MEDICAL UNIV

A protein characterization learning method and device, computer equipment and storage medium

PendingCN122455104AData setSmall sample
The application relates to a protein characterization learning method and device, computer equipment and a storage medium. The method comprises the following steps: collecting protein sequence, structure and function annotation information, and constructing a multi-modal feature data set; performing multi-modal feature extraction on the multi-modal feature data set by using a MASSA framework to obtain protein multi-modal representation; constructing a graph network by using the protein multi-modal representation, and storing the representation and a historical model of an existing protein task in the graph network; in the training process of a new protein task, a memory enhancement mechanism is used to retrieve a similar node of the new protein task from the graph network, a historical model of an existing protein task corresponding to the similar node is obtained, model parameters of the historical model are transferred by using a transfer learning framework, and model training is performed on the new protein task. The application effectively alleviates the data scarcity problem and improves the accuracy and generalization ability of the model in a small sample learning task.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Genome sequence generation and model training methods and related products

This disclosure provides a method and related products for genome sequence generation and model training. During the training phase, pre-trained multimodal data, including sequence data, functional annotation data, omics data, and species evolution information, is acquired. A fused representation is obtained through heterogeneous coding and multimodal fusion, and input into the sequence generation model for feature calculation, outputting the predicted base probability distribution for each base position to be generated. The model is optimized by combining a multi-task training mode of autoregressive generation and intermediate padding, and a weighted loss function based on biological annotation. During the inference phase, sequences are generated progressively based on multimodal generation context information. Candidate bases are obtained through sampling and evaluated using biological constraints, omics validation, structural validation, and functional validation. Resampling is performed when conditions are not met until preset conditions are satisfied. This method can improve the biological rationality and application value of the generated sequences.
Owner:BIOMAP (BEIJING) INTELLIGENCE TECH LTD

A Clustering Method and System for Transcriptome Sequencing Data

The present invention provides a clustering method and system for transcriptome sequencing data. Applied to the technical field of data processing, the method includes: collecting cervical adenocarcinoma and adjacent tissue specimens of a number of patients, and performing data preprocessing to obtain a standardized gene expression matrix; performing clustering operations on the standardized gene expression matrix based on a clustering analysis method combining a latent class model and sub-coalition partitioning to obtain a clustered gene set; performing gene function annotation on the clustered gene set, and performing correlation analysis between the functionally annotated clustering results and the clinical characteristics of cervical adenocarcinoma to obtain the correlation between gene expression and clinical characteristics; constructing a deep clustering prediction model based on a variational autoencoder and a Gamma mixture model based on the clustering analysis and correlation analysis results to predict the prognosis risk or chemotherapy sensitivity of patients. The present invention solves the problems of poor clustering stability of high-dimensional transcriptome data and low prognostic prediction accuracy.
Owner:THE AFFILIATED HOSPITAL OF SOUTHWEST MEDICAL UNIV +1

Rice heat-resistant candidate gene screening method based on population inheritance and transcriptome integration

The invention provides a rice heat-resistant candidate gene screening method based on population inheritance and transcriptome integration, which comprises the following steps: acquiring a rice population, performing population differentiation analysis on the rice population, and screening out highly differentiated sites. And carrying out heterozygosity screening on the rice population to obtain a target SNP site. And taking a target SNP site which is obviously differentiated among the subgroups and is homozygous in the subgroups as a candidate SNP site. False positive signals caused by genetic background differences can be effectively eliminated through combined screening of genetic differentiation between subgroups and homozygosity in groups. And performing function annotation association on the candidate SNP sites, and performing transcriptome differential expression analysis on the rice population. And determining a final candidate gene according to the potential regulatory gene and the differential expression gene. Through collaborative analysis of population genetics and transcriptomics, genetic differentiation and functional expression are verified at the same time, the accuracy and reliability of candidate genes can be remarkably improved, and limitation of single-dimension screening is avoided.
Owner:SOUTH CHINA AGRICULTURAL UNIVERSITY

Enzyme catalytic conversion number prediction method and system based on function annotation and hierarchical structure

The invention belongs to the technical field of enzyme catalytic conversion number prediction, and discloses an enzyme catalytic conversion number prediction method and system based on function annotation and a hierarchical structure, and the prediction method comprises the steps: extracting protein unique identifier information based on a protein resource database, analyzing protein dynamic information and gene ontology functions according to the unique identifier information, and obtaining a prediction result; obtaining a relation table of the gene ontology and the enzymatic conversion coefficient; combining the hierarchical structure of the gene ontology with the relation table, capturing the mutual relation between the gene and the gene product, and constructing a hierarchical total tree between the gene ontology and the enzymatic conversion coefficient according to a relation capturing result; and extracting target gene ontology information according to the gene code of the target object, and matching an enzymatic conversion coefficient corresponding to the target gene ontology information from the hierarchical total tree as an enzymatic conversion number prediction result. According to the invention, gene and protein function annotations are provided by using the gene ontology, so that the function similarity of enzymes can be measured based on the gene ontology, and the enzymatic conversion value of unknown enzymes is speculated.
Owner:TIANJIN INST OF IND BIOTECH CHINESE ACADEMY OF SCI

A method for constructing a metagenomic functional annotation correction model

ActiveCN117253551BBiostatisticsInstrumentsSequence DeletionsSequence annotation
The invention discloses a method for constructing a metagenomic functional annotation correction model, relates to the technical field of metagenomic functional annotation, and discloses the following steps: step 1: dividing a reference amino acid sequence of known function according to the types of front-end deletion, back-end deletion, and double-end deletion; step 2: using an HMMsearch tool to import an incomplete amino acid sequence set into an HMM functional annotation model to be corrected for annotation, and obtaining scores, domain coverage, and sequence coverage results corresponding to the sequences; step 3: inputting four factors, namely, scores, domain coverage, sequence coverage results, and amino acid sequence deletion types, in the HMMsearch annotation results into a machine learning model as features, and taking whether the function of the reference amino acid is the same as that of the model as a category, performing model training, and finally obtaining a correction model of the HMM model for the incomplete amino acid sequence annotation results.
Owner:SHANGHAI PASSION BIOTECHNOLOGY CO LTD

Single-cell rna sequencing annotation method and device based on dynamic hypergraph

The application provides a single-cell RNA sequencing annotation method and device based on a dynamic hypergraph, relates to the technical field of bioinformatics, and comprises the following steps: obtaining a data set, extracting a low-dimensional embedding vector of each cell from a gene expression vector, and constructing a dynamic hypergraph with cells as nodes and gene pathways as hyperedges; extracting a pathway feature of each hyperedge from the dynamic hypergraph; calculating the importance weight of each cell in each hyperedge based on the low-dimensional embedding vector and the pathway feature of the hyperedge; inputting the dynamic hypergraph, the pathway feature and the importance weight into a preset hypergraph neural network for message aggregation and feature learning to generate a prediction label of a cell type; and training the hypergraph neural network through an optimization algorithm to obtain a trained cell annotation model. The application realizes more accurate and more biologically interpretable cell type and function annotation by introducing a hypergraph structure and a metabolic pathway activity dynamic modeling mechanism.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Database construction method, device, equipment and medium based on Nephila clavata spiders

The present application discloses a method, apparatus, device and medium for constructing a database based on Nephila clavata spiders, relating to the field of databases, including: assembling the chromosomal genome of Nephila clavata spiders that meets the first preset high-quality condition; annotating the chromosomal genome to obtain an initial annotation file, and deleting the annotations of gene sequences that do not meet the screening criteria in the initial annotation file to obtain a target annotation file; determining the expression value and co-expression relationship of each gene sequence based on the ordinary transcriptome; performing functional annotation on the classified cell clusters obtained by classifying the single-cell transcriptome after sequencing to obtain an annotated transcriptome; determining the cell distribution within the spatial structure based on the spatial transcriptome after sequencing; constructing a database based on the target annotation file, expression value and co-expression relationship, annotated transcriptome and cell distribution within the spatial structure. A comprehensive spider database is constructed to facilitate the rapid acquisition of comprehensive information on the genome and other omics of spiders.
Owner:SOUTHWEST UNIV