Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

22 results about "Genetic Sequence Databases" patented technology

Records in sequence databases are deposited from a wide range of sources, from individual researchers to large genome sequencing centers. As a result, the sequences themselves, and especially the biological annotations attached to these sequences, may vary in quality.

A method and system for mining and analyzing the correlation of clinical comorbidities of discharged patients

The application relates to a method and system for mining and analyzing the correlation of clinical comorbidity of discharged patients. The method comprises the following steps: constructing an individual diagnosis and treatment narrative graph for each patient according to the discharge medical record text, generating a time sequence transaction sequence corresponding to each patient, and constructing a sequence database; based on the database, the original time sequence frequent pattern set is constructed by analyzing through an improved generalized sequence pattern algorithm; each pattern in the set is classified to obtain multiple classification clusters, and the original time sequence frequent pattern in each classification cluster is processed through multi-sequence alignment to construct a generalized clinical path graph, and the information in the generalized clinical path graph is extracted to generate a natural language abstract. The method improves the time sequence logic, knowledge abstraction degree and clinical interpretability of the clinical comorbidity correlation mining by constructing a diagnosis and treatment narrative graph, mining a time sequence frequent pattern, and constructing a generalized clinical path graph through clustering and multi-sequence alignment.
Owner:WEST CHINA HOSPITAL SICHUAN UNIV

Method for collecting active proteins of water leeches, and detection method, analysis method and application thereof

The application discloses a method for collecting active proteins of water leeches, a detection method, an analysis method and application thereof, and relates to the technical field of biological medicines. The application provides the water leech seedlings treated by hunger and the living host treated by emptying, so that the water leech seedlings adsorb and suck the host, and then the active proteins are injected into the body of the host; the sample of the host after being sucked is collected, and then is made into a freeze-dried powder after being frozen and dried, so that the active proteins of the water leeches are obtained; the proteins are extracted from the freeze-dried powder, and then are subjected to enzymolysis treatment to generate peptide segments; the liquid chromatography-mass spectrometry technology is used to detect the peptide segments; and the mass spectrometry data are subjected to database searching analysis based on the protein sequence database of the water leeches and the host, so that the active proteins from the water leeches are identified and the relative content is determined. The application utilizes the characteristic that the active proteins are injected into the body of the host in the sucking process of the water leeches, and realizes efficient enrichment of the target proteins. In combination with the proteomics technology, high-throughput and high-sensitivity protein identification is realized.
Owner:JINGGANGSHAN UNIVERSITY +1

Method, system and device for identifying age of dark duck based on multi-feature fusion

PendingCN121214494AMachine learningMultiple biometrics useJuvenileAythya baeri
The invention discloses a method, a system and a device for identifying the age of a green-head diving duck based on multi-feature fusion. The method comprises the following steps: shooting a high-definition image of a target individual; obtaining one or more of feather state information of an ear feather region, a wing region, a hypochondrium region, a shoulder and back region, a tail region, a head and neck region, a chest and abdomen region, eye iris and beak color change information and chest and tail distance and beak length data information of the target individual in the high-definition image, and performing quantitative processing to obtain feature parameters corresponding to each feature; performing multi-dimensional fusion analysis on the characteristic parameters through an integrated machine learning model to realize accurate matching with corresponding week age / month age data in a feather moulting standard time sequence database of the green-head diving ducklings / young birds; and outputting a week age / month age identification result of the target individual. According to the method, a moulting standard time sequence database covering a young bird stage of the green-head diving ducks for the first time is constructed, high-precision and high-efficiency non-injury age identification is realized through combination of the database and a machine learning model, and the dependence on professionals is reduced.
Owner:BEIJING ZOO

Bird species identification method based on multi-gene joint amplification and nanopore sequencing

This invention relates to the fields of molecular biology and forensic identification, specifically to a method for bird species identification based on multi-gene co-amplification and nanopore sequencing. The method involves extracting genomic DNA from avian biological samples; using the extracted DNA as a template, amplification is performed using four independent PCR primer pools that specifically target the avian mitochondrial cytochrome C oxidase subunit I gene, cytochrome B gene, 12S rRNA gene, and 16S rRNA gene; after purification of the amplification products, a nanopore sequencing library is constructed and sequenced; finally, bioinformatics analysis, including read length clustering, draft consensus sequence generation, and polishing, is performed on the raw sequencing data to obtain consensus sequences for each target gene. Species identification is then completed by comparing the sequence with a avian sequence database. This method effectively solves the problem that existing technologies cannot simultaneously meet the practical needs of large-scale, rapid, and high-success-rate identification of avian biological samples.
Owner:BEIJING JIANWEI MEDICAL LAB CO LTD +1

A lung cancer early screening detection device based on TCR sequencing and a construction method and application thereof

ActiveCN119517157BHealth-index calculationBiostatisticsSequence databaseOncology
The application discloses a kind of based on TCR sequencing lung cancer early screening model construction method, first based on the TCR sequencing data of lung cancer patient and healthy person sample is constructed TCR sequence database enrichment, then the TCR sequencing of each sample is carried out and obtains the TCR characteristics of each sample, including V gene and J gene The proportion characteristics of gene, the statistical characteristics of TCR immune group, convergence sequence characteristics, TCR sequence amino acid proportion characteristics, different length TCR sequence proportion characteristics, different frequency TCR sequence proportion characteristics and TCR enrichment sequence characteristics;The TCR characteristics of each sample are screened using feature selection method, and a machine learning model is constructed by combining other information of the sample to obtain a lung cancer early screening model.Based on the lung cancer early screening model, a lung cancer early screening detection device is constructed, which can realize lung cancer early screening based on TCR sequencing information of the sample.
Owner:SHENZHEN HAPLOX BIOTECH

Method for early diagnosis of cancer based on artificial intelligence using cell-free DNA distribution of tissue-specific regulatory region

Provided are an artificial intelligence-based early cancer diagnosis method and device using a method of inputting information on a cell-free DNA distribution of a tissue-specific regulatory region to an artificial intelligence model learned to early diagnose cancer and analyzing the information, and an information providing device and a storage medium.SOLUTION: A method for providing information for early cancer diagnosis based on artificial intelligence includes extracting a nucleic acid from a biological sample to obtain sequence information, arranging the obtained sequence information in a reference chromosome sequence database, selecting a nucleic acid fragment of a regulatory region based on the arranged sequence information, generating the selected nucleic acid fragment as image data, and inputting the generated image data to an artificial intelligence model learned to distinguish a normal image and a cancer image, analyzing the image data, and comparing the image data with a reference value to determine the presence or absence of cancer.SELECTED DRAWING: Figure 1
Owner:GREEN CROSS GENOME CORP

Methods of polypeptide design using combined masked language modeling and yeast surface display and sequences

The polypeptide design method and sequence of the present application combine mask language modeling and yeast surface display, comprising the following steps: cleaning a protein sequence database, selecting protein sequences meeting the requirements as a training set of a language model, and performing mask language modeling on the protein sequences contained in the training set; designing a set of downstream tasks on the basis of a pre-trained model, and fine-tuning the downstream tasks; randomly masking residues of a selected reference sequence, and predicting the masked residues; performing virtual screening on polypeptide candidates generated by the model; and determining the protein expression level and affinity of the screened polypeptides through yeast display technology. The present application can generate a large number of polypeptide candidates that may have specific properties or functions by transferring natural language processing technology to the polypeptide generation field, and combines artificial intelligence generation, virtual screening and wet experimental characterization to build a relatively complete "dry-wet combination" design process.
Owner:GUANGXI ZHONGMA PENCHENG PHARMACEUTICAL IND GROUP CO LTD

Pattern mining method, device and equipment for critical path data and storage medium

PendingCN121996713Aadd depthstatistically significantDigital data information retrievalSpecial data processing applicationsAlgorithmSequence database
The invention provides a mode mining method, device and equipment for critical path data and a storage medium, and relates to a data mining technology, the method comprises the following steps: obtaining an analysis processing request; and reading a sequence database corresponding to the sequence database identifier according to the analysis processing request. Generating a candidate mode set; determining a candidate mode greater than or equal to a preset minimum support degree threshold in a plurality of candidate modes in the candidate mode set based on the plurality of sequences in the sequence database; taking the candidate modes greater than or equal to a preset minimum support degree threshold as high-support-degree modes, and generating a high-support-degree mode set; and if it is determined that the high-support-degree mode set is null or the second length is greater than the preset maximum mode length, stopping high-support-degree mode mining. According to the method, the technical problem that the accuracy of mode mining of the key path of the processor microstructure correlation diagram is low is solved.
Owner:LOONGSON TECH CORP

IGH gene fusion identification method and device, storage medium and program product

PendingCN121237218AMicrobiological testing/measurementBiostatisticsSequence databaseA-DNA
The invention discloses an IGH gene fusion identification method and device, a storage medium and a program product, and the method comprises the following steps: obtaining sequencing data, and preprocessing the obtained sequencing data; the preprocessed data is compared to an IGH library, one or more candidate sequences are obtained, the candidate sequences are sequences containing IGH sections and unknown sections, and the unknown sections are sequences with the length within a preset first length range and are not compared to IGH fragments; acquiring the unknown sections, and comparing the unknown sections to a DNA sequence database to obtain a plurality of comparison results; and analyzing the plurality of comparison results, and identifying the gene corresponding to the optimal comparison result meeting a preset similarity condition as the gene fused with the IGH segment.
Owner:BOE TECHNOLOGY GROUP CO LTD +1

An unsupervised KPI anomaly detection method based on double cross-coupling correlation

The application discloses a kind of based on double cross coupling correlation's unsupervised KPI anomaly detection method, first construct by historical KPI sequence database, preprocessing module, input module, based on double cross coupling correlation's KPI fragment feature extraction model, feature conversion module, feature space distance calculation module, training optimization module and online anomaly detection module constitute based on double cross coupling correlation's unsupervised KPI anomaly detection system.First to the unsupervised KPI anomaly detection system is trained, KPI fragment feature extraction model simultaneously models the coupling correlation of KPI data internal time sequence dimension and the coupling correlation between KPI sequence, extracts the feature vector of KPI data to be detected, feature space distance calculation module calculates the distance of feature vector and feature space center point, and online anomaly detection module determines whether KPI data to be detected is abnormal according to distance.The application can effectively improve the accuracy of unsupervised KPI anomaly detection.
Owner:NAT UNIV OF DEFENSE TECH

Data quality evaluation method and system based on real-sequence database

The invention discloses a data quality evaluation method and system based on a real-sequence database, and the method comprises the steps: collecting factory bit number time sequence data from at least one data source, and storing the time sequence data in a time sequence data storage layer; based on preset multi-dimensional data quality evaluation indexes, the time series data are evaluated, a corresponding quality evaluation result is generated, and the multi-dimensional data quality evaluation indexes at least comprise data integrity, data consistency, data accuracy, data credibility and data stability; according to the quality evaluation result, a factory data quality space is constructed and maintained, and the data quality space is used for centralized management and visual display of the data quality state of the factory position number; and providing an access interface for the data quality space for a user through the client. A statistical analysis technology is utilized to research factory data quality, and a user is helped to better understand and solve a data quality problem.
Owner:SUPCON TECH CO LTD +1

Proteome mass spectrum data quality evaluation method based on self-attention model encoder

PendingCN121393573ABiostatisticsBiological modelsData setSequence database
The invention provides a proteome tandem mass spectrum quality evaluation method based on a machine learning self-attention model encoder, which is used for evaluating the quality of proteome mass spectrum experimental data. The method can be used for identifying high-quality mass spectra (such as search parameters which are not optimized, unspecified post-translational modifications and incomplete sequence databases) which are not identified due to data analysis reasons, can also be used for evaluating the overall data quality of one experiment, provides an overall statistical overview of the proteome mass spectrum experiment data quality, and can be used for analyzing the mass spectrum of the proteome. The method is used for constructing a proteome standard data set with consistent data quality and the like, and the reliability and reproducibility of proteomics analysis conclusions are promoted.
Owner:CHINA JILIANG UNIV

Enzyme with function of catalyzing formaldehyde to synthesize acetyl phosphate and application of enzyme in CO2 biosynthesis of ethanol

PendingCN121450626ABacteriaBiofuelsSequence databasePhosphoric acid
The invention relates to phosphoketolase MaPKT and SzPKT with a function of catalyzing formaldehyde to synthesize acetyl phosphate and application of phosphoketolase MaPKT and SzPKT in CO2 biosynthesis of ethanol, and belongs to the technical field of enzyme mining and biological catalysis application. A phosphoketolase specific protein sequence database and an oligopeptide module sequence library are established through an oligopeptide module recognition (PPR) technology, two phosphoketolase PKTs capable of catalyzing formaldehyde to synthesize acetyl phosphate are efficiently and accurately excavated, further research finds that the formaldehyde tolerance of MaPKT is improved, and the catalytic efficiency of MaPKT on formaldehyde is 8 times that of glycolaldehyde. Based on this, the invention also constructs an efficient condensation approach for synthesizing ethanol by artificially biologically converting CO2, and has potential application prospects.
Owner:INSTITUTE OF PROCESS ENGINEERING CHINESE ACADEMY OF SCIENCES

Multiplex PCR amplicon sequencing data detection method based on deep learning

The invention relates to the technical field of bioinformatics and computational biology, in particular to a multiple PCR amplicon sequencing data detection method based on deep learning. The technical problem that the accuracy of overall detection of complex sequencing data is low in the prior art is solved. The method comprises the following steps: acquiring a to-be-detected sequencing data set; obtaining a reference sequence database; constructing a feature extraction model for extracting amplicon features based on the historical amplification condition information and the multiple groups of historical sequencing data of the target tag; and processing a to-be-detected sequencing data set based on the feature extraction model, and determining a target category corresponding to each sequencing sequence in the sequencing data set. The multi-PCR amplicon sequencing method is used for a multi-PCR amplicon sequencing scene.
Owner:THE FIRST HOSPITAL OF HEBEI MEDICAL UNIV +1

Multi-dimensional efficient DNA sequence pattern mining method

PendingCN121506258ABiostatisticsProteomicsSequence analysisSequence database
The invention discloses a multi-dimensional efficient DNA sequence pattern mining method, and aims to identify a pattern with biological significance in a DNA sequence. According to the method, firstly, an original database containing DNA sequences and dimension information is divided into a DNA sequence database and a dimension database on the basis of a database analysis technology, and secondly, base combinations with remarkable effectiveness under different conditions are extracted from DNA sequence data through an efficient sequence pattern mining algorithm; and finally, generating a multi-dimensional efficient DNA sequence mode in combination with a multi-dimensional data mining technology. According to the method, experimental measurement values or importance of all basic groups are quantified into utility values, deep mining is carried out on multi-dimensional DNA data, and high-utility gene segments under specific species, tissue types and experimental conditions are identified. According to the method disclosed by the invention, the efficiency and depth of DNA sequence analysis are remarkably improved, and a functional sequence mode which is difficult to find by a traditional method and has condition specificity can be systematically revealed, so that the scientific preciseness and the practical application value of research results are improved in multiple aspects of gene function analysis, disease marker discovery, evolution research and the like.
Owner:QINGDAO UNIV OF TECH

Waste mine geothermal monitoring data management method and system based on time sequence database

The invention provides a waste mine geothermal monitoring data management method and system based on a time sequence database, and belongs to the technical field of data management, and the method comprises the steps: obtaining waste mine geothermal monitoring data, and constructing a time sequence database architecture of a multi-stage association table structure; starting time sequence database extension and performing optimal configuration, converting a time sequence data table in the time sequence database into a supertable, and setting a partitioning strategy, index optimization and a data compression strategy; importing the monitoring data into a time sequence database and controlling the quality in real time; intelligent analysis is carried out based on a built-in function of a time sequence database, multi-time granularity aggregation analysis, trend analysis and anomaly detection are realized, a multi-dimensional intelligent early warning mechanism is constructed, and real-time early warning message pushing is realized through a trigger; and storing an analysis result to realize data life cycle management. According to the invention, efficient storage, rapid query and time series data management with intelligent analysis capability are realized.
Owner:JINING MINING GRP CO LTD +1

Rotavirus detection primer pair and typing analysis method and system thereof

The invention discloses a rotavirus detection primer pair and a typing analysis method and system thereof. The invention relates to a primer pair (containing a forward primer of SEQ ID NO: 1 and a reverse primer of SEQ ID NO: 2) for genotyping of rotavirus VP7, a primer pair (containing a forward primer of SEQ ID NO: 3 and a reverse primer of SEQ ID NO: 4) for genotyping of VP4 and a composition containing the primer pairs. The invention also provides a rotavirus genotyping method which comprises the following steps: extracting rotavirus RNA in a sample, performing reverse transcription to obtain cDNA, performing PCR amplification by using the primer pair, sequencing a product, and comparing a sequencing sequence with a rotavirus reference sequence database to obtain G-type and P-type genotyping results; the invention also provides an automatic typing system for typing analysis of the sequencing sequence of the PCR product. The primer can amplify main epidemic genotypes only through one round of PCR, is convenient to operate and excellent in specificity and detection capacity, is accurate in typing method result, effectively solves the problems of insufficient primer sensitivity, tedious typing and the like in the prior art, and is suitable for rotavirus genotyping and epidemiological monitoring.
Owner:SHENZHEN CHILDRENS HOSPITAL

Cell type annotation method based on multi-feature contrast learning

PendingCN121999887AExcellent accuracyExcellent F1 scoreBiostatisticsProteomicsData setSequence database
The invention relates to the technical field of cell type annotation, in particular to a cell type annotation method based on multi-feature contrast learning, which comprises the following steps: acquiring a gene expression matrix to extract a gene name, extracting a gene base sequence from a transcriptome sequence database based on the gene name, and matching the gene base sequence with a motif sequence, obtaining a gene motif matching matrix; multiplying the gene expression matrix by the gene motif matching matrix to obtain a cell motif matrix; preprocessing the gene expression matrix and the cell motif matrix, and constructing a data set based on the preprocessed gene expression matrix and cell motif matrix; training a double-feature coding model by using the training set and the verification set, wherein the double-feature coding model comprises a gene expression encoder, a motif feature encoder and a classifier; and inputting the test set into the trained double-feature coding model to obtain the annotated cell type. According to the method, the annotation result has biological significance, and downstream function analysis is facilitated.
Owner:DALIAN MARITIME UNIVERSITY

Amino acid sequence generation and screening method and system used for amino acid sequence generation and screening method

A virus amino acid sequence generation and screening method comprises the following steps: S1, a pre-training step: constructing a data set for a pre-training model based on a label-free amino acid sequence database, and carrying out preliminary training on the pre-training model to learn grammar and semantic information of an amino acid sequence; s2, conditional generation: based on a model obtained by pre-training, preliminarily generating an amino acid sequence by using a virus amino acid sequence data set; s3, a rational design step: based on the'conditionally generated 'amino acid sequence, performing optimization design on the amino acid sequence through a rational design module, and based on existing positive sample data with required functions, performing further optimization and transformation on the amino acid sequence output by the generation model by utilizing a heuristic algorithm, so as to obtain an optimized amino acid sequence; the sequence space is evolved towards the positive sample direction, and the sequence is optimized; and S4, a filtering step: performing function prediction on the virus amino acid sequences in the target virus amino acid sequence library, performing function prediction according to the functions of the required amino acid sequences, and filtering out the amino acid sequences with poor functions. The invention also provides a system for implementing the method, and a key amino acid sequence screened out based on the system and used as an AAV capsid sequence.
Owner:ZHEJIANG UNIV

Intelligent typing method, system, and media for chikungunya virus

PendingCN122455122ASemantic alignmentSequence database
The present disclosure provides a method, system and medium for intelligent typing of chikungunya virus, obtains CHIKV sequence data from a public database, performs quality control, uses an artificial intelligence model to extract typing reference sequences of CHIKV, and then compares the sequence data with the typing reference sequences one by one. If the comparison result of the sequence data and a certain typing reference sequence is similarity>99% and sequence coverage>99%, the sequence information of the typing reference sequence is supplemented to the sequence data, and a typing sequence database is generated. The present disclosure uses an artificial intelligence model to realize automatic extraction, semantic alignment and standardization of multi-source data, solves the bottleneck of data dispersion and metadata loss, realizes intelligent typing of chikungunya virus, encapsulates complex bioinformatics processes into a "one-key" online tool, significantly reduces the technical threshold, enables grassroots personnel to independently complete high-level analysis, and provides real-time, visual molecular evidence and risk prompts for epidemic prevention and control.
Owner:INST OF MICROBIOLOGY CHINESE ACAD OF SCI

Diversity detection method for dinoflagellate and cysts thereof based on autonomous LSU rDNA and ITS double databases and double molecular markers

The invention discloses a dinoflagellate and its cyst diversity detection method based on autonomous LSU rDNA and ITS double databases and double molecular markers, relates to the technical field of molecular biology and marine ecological monitoring, and solves the problems of insufficient coverage of existing public databases and limited resolution of single molecular markers. According to the invention, an LSU rDNA sequence database and an ITS sequence database are constructed, and molecular markers of an LSU D1-D2 region and an ITS1 region are correspondingly amplified respectively, so that combined application of double databases and double markers is realized, and complementarity, comprehensiveness and high-resolution molecular identification is carried out on dinoflagellate and cyst class groups thereof. According to the method, the defects that an existing public database is insufficient in coverage and the resolution ratio of a single molecular marker is limited can be effectively overcome, the detection rate and identification accuracy of the dinoflagellate community are remarkably improved, and scientific support and technical guarantee are provided for dinoflagellate red tide monitoring, marine ecosystem evaluation and environmental risk early warning.
Owner:NANJING UNIV OF INFORMATION SCI & TECH +2

A negative sequence pattern analysis system and method based on influence degree

The application discloses a negative sequence pattern analysis system based on influence degree, which comprises a data preprocessing module, a frequent pattern mining module, an influence degree analysis module and a graphical interface representation module. The data preprocessing module is used for obtaining data, preprocessing and storing into a database. The frequent pattern mining module is used for mining frequent patterns in the database according to a minimum support threshold set by a user. The influence degree analysis module is used for calculating the influence degree of positive and negative frequent patterns obtained by the frequent pattern mining module respectively, and selecting frequent patterns meeting the influence degree constraint. The graphical interface representation module is used for representing the mining result of the influence degree analysis module on the graphical interface of the system. The application also discloses a corresponding analysis method. According to the live shopping transaction sequence database, the application uses the negative sequence pattern mining algorithm based on the influence degree to provide the live shopping information which is really interesting to the user in the form of the graphical interface, and the user can select the interesting sequence pattern in the mining result according to the will of the user to show.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)