Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

35 results about "Biosequence" patented technology

A BioSequence is an object representation of a DNA, RNA, or protein sequence. It can be represented by a Clone, Gene, or the sequence. (caMAGE)

Systems and methods for multimodal conversational agents for biological sequence analysis

Provided herein are technologies for framing and evaluating biological sequence-based analysis tasks in a unified, natural-language-based, text in and text out format. Among other things, methods and systems of the present disclosure provide machine-learning technologies for combining biological sequence data, representing, for example, DNA, RNA, and protein sequences, with natural language, conversational style prompts that set out particular analysis tasks to be performed on the biological sequence data. This approach, for example, allows complex analysis tasks, including, but not limited to, identification of various sequence modifications, genes, and regulatory elements in DNA sequences, and quantification of properties such as degradation propensity of RNA and protein stability, to be input to a machine learning model in a uniform text-based format and for output to be generated in a same, unified, text-based format.
Owner:INSTADEEP LTD +1

Group health evaluation method and system based on intestinal flora

The invention relates to the technical field of health evaluation, in particular to a group image health evaluation method and system based on intestinal flora, and the method comprises the steps: collecting intestinal flora samples of a target group through a standardized process, and obtaining microbial sequence information through a high-throughput sequencing technology; microflora diversity characteristics and core flora characteristics are extracted through strain identification and relative abundance analysis, individual flora characteristics are compared with a healthy population reference database, health state scores are calculated, health grades are divided, population health grade distribution is counted, and the population health grade distribution is calculated. And generating a group health evaluation report containing a visual chart and text analysis. According to the method, a complete technical system from sample collection to health strategy making is established, systematic evaluation of the group health state from the perspective of intestinal flora is achieved, and the limitation of a traditional method in group health early evaluation is overcome. The method is suitable for health monitoring of different scales of groups, and provides effective technical support for public health management and health service.
Owner:SECOND MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL

Dynamic multi-modal biological sequence staged pre-training system and method for precise breeding

The invention provides a dynamic multi-modal biological sequence staged pre-training system for precise breeding, which comprises a multi-modal data organization unit used for collecting a single-modal sequence, generating a two-modal pairing sequence, constructing a three-modal interpenetrating sequence and simulating biological information transmission; the unified sequence representation unit is used for performing unified word segmentation on each modal sequence and adding modal marks; the progressive pre-training strategy unit is used for training in three stages and dynamically adjusting the modal mixing proportion in combination with a simulated annealing strategy; the cross-modal autoregression unit realizes prediction conversion between modals; the sequence feature extraction and prediction unit is used for obtaining a quantitative prediction value. The invention further provides a dynamic multi-mode biological sequence staged pre-training method for precise breeding. Therefore, the biological sequence feature prediction precision can be remarkably improved, the training convergence speed is increased, the performance equivalent to that of a large model is achieved, the cost is low, flexible deployment in various environments can be achieved, and efficient, accurate and low-cost technical support is provided for biological sequence analysis.
Owner:AGRI INFORMATION INST OF CHINESE ACAD OF AGRI SCI

Biological sequence compression using sequence alignment

Compressing files is disclosed. An DNA sequence to be compressed is first aligned. Aligning the DNA sequence includes splitting the DNA sequences into smaller sequences or portions that can be aligned. After the DNA sequence is spilt one or more time and aligned, a compression matrix is generated. Each row of the compression matrix corresponds to part of the DNA sequence. A consensus sequence is determined from the compression matrix. Using the consensus sequence, pointer pairs are generated. Each pointer pair identifies a subsequence of the consensus matrix. The compressed file includes the pointer pairs and the consensus sequence.
Owner:DELL PROD LP

Method for constructing an online prediction website based on biological sequence cis-acting regulatory elements

This invention discloses a method for constructing an online prediction website based on cis-regulatory elements of biological sequences. This method is applicable to building online prediction websites from models trained using machine learning on biological sequences. With the continuous development of artificial intelligence technology, machine learning algorithms have flourished and been successfully applied in multiple fields. In particular, in the fields of biological and medical research, researchers have constructed numerous prediction models to serve various research tasks, such as the prediction of cis-regulatory elements, biomedical pathogenic elements, and proteins. Although many prediction models have been successfully constructed, professionals who frequently use computers are more concerned with how to apply these models to practical work. Therefore, constructing the prediction models in this paper into user-friendly and easy-to-use online websites can greatly enhance the practical application value of the models.
Owner:GUILIN UNIV OF ELECTRONIC TECH

IncRNA and disease association prediction method based on deep learning

The invention belongs to the technical field of bioinformatics, and particularly relates to an lncRNA and disease association prediction method based on deep learning, which comprises the following specific steps: firstly, according to known lncRNA-disease association information, disease-miRNA association information and lncRNA-miRNA association information, constructing an lncRNA-disease association matrix, a disease-miRNA association matrix and an lncRNA-miRNA association matrix; and constructing an lncRNA comprehensive similarity matrix LS, a disease comprehensive similarity matrix DS and a miRNA comprehensive similarity matrix MS. The biological sequence selective compression network adopted by the invention can effectively model a long-distance dependency relationship in an ultra-long sequence and dynamically distinguish key functional fragments from redundant fragments according to context.
Owner:GUANGDONG UNIV OF TECH

Synthetic biological data intelligent extraction system based on multi-modal large language model

The invention discloses a synthetic biological data intelligent extraction system based on a multi-modal large language model, and belongs to the technical field of biological data intelligent extraction, and the system comprises a multi-modal data preprocessing and standardization module which is used for carrying out the analysis, classification and structured output of input heterogeneous synthetic biological data; and the multi-modal large language model processing module is connected with the preprocessing module and is used for receiving the processed structured data and carrying out joint understanding and information extraction through knowledge enhancement and deep fusion technologies. The heterogeneous synthetic biological data can be efficiently processed, different types of data such as scientific literatures, chart images and biological sequences are accurately analyzed and classified through the multi-modal data preprocessing and standardization module, the data format barrier is broken, scattered biological information is systematically integrated, and the data processing efficiency is improved. A standard and unified data foundation is laid for subsequent information extraction, and the comprehensiveness and effectiveness of data processing are greatly improved.
Owner:AIXBIO (HANGZHOU) BIOTECHNOLOGY CO LTD

Method for constructing database, method for retrieving document and computer device

Disclosed are a method for constructing a database, a method for labeling an association degree of biological sequences, a method for retrieving a document, and a computer device. In the solution of this application, a biological sequence and attribute information are extracted from a target document, and an entry in a database is constructed based on the extracted biological sequence and the attribute information. When a user conducts retrieval based on the database, a server can match an entry for the user by means of the biological sequence and the attribute information in the entry or a combination of the two. Therefore, when applied to a retrieval platform, the database of this application can provide the user with various types of retrieval support, such as biological sequence retrieval, biological sequence attribute retrieval, and comprehensive biological sequence and biological sequence attribute retrieval, and the like.
Owner:PATSNAP LIMITED

Fitness space roughness evaluation method and device, electronic equipment and storage medium

The application provides a roughness evaluation method and device of fitness space, electronic equipment and storage medium, and relates to the technical field of biology. The method comprises the following steps: acquiring a data set; the data set comprises a plurality of biological sequences; based on the mutation sites between each biological sequence in the data set, a plurality of adjacent sequence pairs in the data set are determined; the mutation sites between the two biological sequences in each adjacent sequence pair meet the set number of mutation sites; and the roughness of the fitness space is estimated according to each adjacent sequence pair, to obtain a roughness evaluation result. The application solves the problem of inaccurate roughness evaluation of the fitness space in the related art.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

A machine learning-guided biological sequence engineering modification method and device

ActiveCN115249514BBiostatisticsProteomicsScalable computingEngineering
The present invention provides a method and apparatus for biosequence engineering guided by machine learning. Specifically, a Bayesian optimization-guided evolutionary algorithm (BO-EVO) is provided, which combines Bayesian optimization (BO) and evolutionary algorithm (EVO). EVO solves the problem of excessive computational complexity caused by BO locating the global optimum by brute force search of the entire design space. At the same time, the exploratory nature of BO is used to neutralize the greed and under-exploration shortcomings of EVO to achieve efficient iteration between machine learning models and robotic experiments, so as to economically obtain new protein variants of high value. The method of the present invention is expected to achieve efficient and scalable computing and exploration, and provide an efficient biosequence engineering solution.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Phylogenetic tree construction method and system based on deep learning and beam search

This invention discloses a phylogenetic tree construction method and system based on deep learning and beam search. It predefines evolutionary scenarios, setting evolutionary parameters for each scenario with reference to real-world biological sequence attributes. Training, validation, and test sets are created based on simulated phylogenetic trees and corresponding multiple sequence alignment data according to the predefined parameters. A deep learning classifier with convolutional neural networks and long short-term memory neural networks as its core is constructed. The deep learning classifier is trained and validated using the training and validation sets, and its accuracy is tested using the test set. Based on the trained deep learning classifier and the sliding window method, classification predictions are performed on all sub-quad-sequence trees of the four-sequence data. A phylogenetic tree reconstruction is performed on the multiple sequence data using an improved stepwise addition method and the quad-sequence tree classification prediction results, resulting in a complete reconstruction. This enables phylogenetic tree construction under conditions of different species numbers and sequence lengths.
Owner:CHINESE INST FOR BRAIN RES BEIJING +1

Compression Method and Device for Biological Sequence Identifiers, Decompression Method and Device

The present invention discloses a compression method and device for biological sequence identifiers, as well as a decompression method and device. For each identifier in a gene sequencing file, the identifier is split into several sub-identifiers; encoding rules for several windows are defined, and the encoding rules match the text format of the sub-identifiers; the sub-identifiers with the same referential meaning are divided into the same window; for each window, all the sub-identifiers in the window are encoded according to the corresponding encoding rule, and the encoding results of each window are aggregated into the compression result of the identifier. These methods maximize the compression ratio of all identifier data while ensuring compatibility with special data as much as possible, and at the same time ensure the encoding and decoding performance.
Owner:MGI HLDG CO LTD

A method and system for RNA secondary structure prediction and folding path deduction

The application belongs to the technical field of biological sequence data processing, and discloses a method and system for predicting and deducing folding paths of RNA secondary structure, which comprises: inputting a to-be-predicted RNA sequence into a deep learning pre-training model to perform feature extraction and pattern recognition, and obtaining a plurality of local structure modules each composed of at least three nucleotide pairs; regarding all the local structure modules as an action space of a reinforcement learning agent, and enabling the reinforcement learning agent to perform filtering and pairing actions on the local structure modules, gradually output a prediction result of the RNA secondary structure, and record the local structure modules and pairing sequences selected at each step to obtain a folding path, wherein the reinforcement learning agent adopts a composite reward function comprising a reward structure module formation, a punishment topological conflict, and an evaluation of local stability, and adaptively adjusts the weights of the composite reward function according to a folding stage and a current state. The application can improve the calculation efficiency and generalization ability, and reproduce the RNA folding path.
Owner:BEIJING NEOCURNA BIOTECHNOLOGY CORP

Systems and methods for multimodal conversational agents for biological sequence analysis

Provided herein are technologies for framing and evaluating biological sequence-based analysis tasks in a unified, natural-language-based, text in and text out format. Among other things, methods and systems of the present disclosure provide machine-learning technologies for combining biological sequence data, representing, for example, DNA, RNA, and protein sequences, with natural language, conversational style prompts that set out particular analysis tasks to be performed on the biological sequence data. This approach, for example, allows complex analysis tasks, including, but not limited to, identification of various sequence modifications, genes, and regulatory elements in DNA sequences, and quantification of properties such as degradation propensity of RNA and protein stability, to be input to a machine learning model in a uniform text-based format and for output to be generated in a same, unified, text-based format.
Owner:INSTADEEP LTD +1

Method, device, electronic equipment and program product for predicting siRNA silencing efficiency

PendingCN122658423ALinguistic modelEngineering
The present disclosure provides a siRNA silencing efficiency prediction method, device, electronic equipment and program product, comprising the following steps: obtaining a pseudo-label feature value of a to-be-predicted siRNA sequence based on a language model; searching for an off-target effect score of the to-be-predicted siRNA sequence based on a search model; obtaining a rule score of the to-be-predicted siRNA sequence based on a preset rule; performing feature extraction based on the to-be-predicted siRNA sequence to obtain sequence features; and obtaining a silencing efficiency of the to-be-predicted siRNA sequence based on a machine learning model according to the pseudo-label feature value, the off-target effect score, the rule score and the sequence features. The present disclosure comprehensively utilizes the biological sequence features of traditional rules and the strong learning ability of deep learning models, improves the accuracy and robustness of siRNA silencing efficiency prediction, reduces the dependence on a large amount of labeled data, and enhances the generalization ability among different biological samples, thereby providing a more reliable and practical tool for siRNA design.
Owner:SHANGHAI PHARMACEUTICALS HOLDING CO LTD

Cancer resistance peptide generation and screening method and system based on diffusion model and seed synchronous autoencoder

PendingCN122117054AImprove biological effectivenessSolve mapping collapse problemEnsemble learningBiostatisticsRandom seedCancer resistance
The application discloses an anticancer peptide generation and screening method and system based on a diffusion model and a seed synchronous autoencoder, and comprises the following steps: preprocessing original anticancer peptide and non-anticancer peptide sequences and extracting multi-dimensional physicochemical features; training a diffusion model in a continuous numerical space based on sequence mapping based on a global random seed, adding noise through forward diffusion and learning the potential distribution of anticancer peptide data features through reverse denoising; based on a reverse seed synchronization strategy, training an autoencoder using the same random seed and environment state, and establishing a mapping from continuous random noise to discrete biological sequences; using the diffusion model to sample and generate potential vectors in the latent space and mapping them into candidate anticancer peptide sequences through the decoded synchronously trained decoder; inputting an integrated learning classifier based on grouping features, selecting the optimal feature group and predicting the anticancer activity probability, and outputting the anticancer characteristic peptide chain; the application solves the mapping contradiction between continuous space and discrete sequences and can efficiently screen anticancer peptides with high activity sequences.
Owner:CHONGQING UNIV

Architectures for training neural networks using biological sequences, conservation, and molecular phenotypes

ActiveUS12626782B2BiostatisticsProteomicsMolecular phenotypeData set
The present disclosure provides methods and systems that can ascertain how genetic variants impact molecular phenotypes. Such methods and systems may use additional conservation information. In an aspect, the present disclosure provides a method for training a molecular phenotype neural network (MPNN), comprising: (a) providing a molecular phenotype neural network (MPNN) comprising one or more parameters; (b) providing a training data set comprising (i) a set of one or more inputs comprising biological sequences and (ii) for each input in the set of one or more inputs, a set of one or more molecular phenotypes corresponding to the input; (c) configuring the one or more parameters of the MPNN based on the training data set to minimize a total loss of the training data set, thereby training the MPNN; and (d) outputting the one or more parameters of the MPNN.
Owner:DEEP GENOMICS INC

Parallel-processing systems and methods for highly scalable analysis of biological sequence data

An apparatus includes a memory configured to store a sequence. The sequence includes an estimation of a biological sequence. The sequence includes a set of elements. The apparatus also includes a set of hardware processors. Each hardware processor is configured to implement a segment processing module. The apparatus also includes an assignment module implemented in a hardware processor. The assignment module is configured to receive the sequence from the memory, and assign each element to at least one segment from a set of segments, including, when an element maps to at least a first segment and a second segment, assigning the element to both the first segment and the second segment. The segment processing module is configured to, for each segment from a set of segments specific to that hardware processor, and substantially simultaneous with the remaining hardware processors, remove at least a portion of duplicate elements in that segment to generate a deduplicated segment. The segment processing module is further configured to reorder the elements in the deduplicated segment to generate a realigned segment that has a reduced likelihood for alignment errors.
Owner:RES INST AT NATIONWIDE CHILDRENS HOSPITAL

Method for analysis and query-based interrogation of biological sequence data

PCT designated stageWO2026132228A1ProteomicsGenomicsData decompositionProtein
The invention relates to a computer-implemented method for creating a searchable indexed data structure for biological sequence data. The invention further relates to searching in said data structure for any given biological sequence data or associated information. The invention relates to a computer-implemented method comprising: a. providing biological sequence input data, b. decomposing the biological input data into non-overlapping portions, c. indexing each portion of the biological input data with at least one associated information of said portion comprised within or provided with the biological input data, thereby creating a two-tiered index for each portion. The method may further comprise d. receiving at least one query from a user, wherein the query comprises biological sequence information, language-based prompts and / or queries regarding the metadata comprised within or provided with the biological input data. The method may further comprise e. providing at least one output comprising at least one biological sequence information, biological sequence identifier and / or metadata of one or more portions of the biological sequence input data corresponding to the query. The invention further relates to the use of the method of the invention for identifying tissue and / or cell localized expression of at least one nucleic acid, such as a gene, and / or protein of interest. The method further relates to a data structure and a computer-readable medium comprising a data structure, such as a database, constructed according to the invention.
Owner:MAX DELBRUECK CENT FUER MOLEKULARE MEDIZIN

Methods for classifying biological sequences, computer programs, and computer implementations (identification of unknown genomes and recently known genomes).

To provide a method, a computer program and a computer-implemented method for classifying biological sequences (identification of unknown genomes and closest known genomes).SOLUTION: A method for classifying biological sequences comprises: providing a training set related to a biological sequence; dividing the training set and the biological sequence into sequence fragments; extracting feature vectors from the sequence fragments of the training set and from the sequence fragments of the biological sequence; identifying feature vectors extracted from the sequence fragments of the biological sequence that correspond to feature vectors extracted from the sequence fragments of the training set; establishing a threshold value comprising an average degree of divergence of the feature vectors from the sequence fragments of the biological sequence, from the feature vectors from the sequence fragments of the training set; and detecting an anomaly among the extracted features of the sequence fragments of the biological sequence.SELECTED DRAWING: None
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Bioinformatics processing

In a first aspect, the present invention relates to a computer-implemented method for obtaining information about a biological entity based on at least one biological sequence, comprising: (a) providing a repository of fingerprint data strings for a biological sequence database, each fingerprint data string representing a characteristic biological subsequence composed of sequence units, each characteristic biological subsequence having a combination number less than the total number of different sequence units available in the biological sequence database, the combination number of the biological subsequence being defined as the number of different sequence units that appear as consecutive sequence units of the biological subsequence in the biological sequence database; (b) determining one or more fingerprint data strings representative of the biological entity; (c) searching a repository including information associated with the fingerprint data strings for information associated with the one or more representative fingerprint data strings; and (d) processing the information.
Owner:BIO BEACH INC

Digital to biological converter

The present invention provides a system for receiving biological sequence information and activating the synthesis of a biological entity. The system has a receiving unit for receiving a signal encoding biological sequence information transmitted from a transmitting unit. The transmitting unit can be present at a remote location from the receiving unit. The system also has an assembly unit connected to the receiving unit, and the assembly unit assembles the biological entity according to the biological sequence information. Thus, according to the present invention biological sequence information can be digitally transmitted to a remote location and the information converted into a biological entity, for example a protein useful as a vaccine, immediately upon being received by the receiving unit and without further human intervention after preparing the system for receipt of the information. The invention is useful, for example, for rapidly responding to viral and other biological threats that are specific to a particular locale.
Owner:TELESIS BIO INC

Method for obtaining and correcting biological sequence information

This disclosure provides methods for sequencing biomolecules such as nucleic acid molecules, and methods for detecting and / or correcting sequencing errors in sequencing results. Reagent kits and systems based on the methods disclosed herein are also provided.
Owner:CYGNUS BIOSCI BEIJING CO LTD

Data processing method and device, electronic equipment, computer readable storage medium and computer program product

The invention provides a data processing method and device, electronic equipment, a computer readable storage medium and a computer program product. The method is applied to artificial intelligence. The method comprises the following steps: determining a monomics sequence of a biomolecule based on a molecular sequence of the biomolecule; performing sequence conversion on the molecular sequence to obtain a multi-omics sequence of the biomolecule; fusing the single omics sequence and the multi-omics sequence to obtain a biological sequence; coding the biological sequence to obtain biological sequence coding characteristics of the biological molecules; and performing prediction processing on the biological sequence coding features to obtain molecular prediction information of the biological molecules. According to the invention, the accuracy of molecular prediction information can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Systems and methods for multimodal conversational agents for biological sequence analysis

Provided herein are technologies for framing and evaluating biological sequence-based analysis tasks in a unified, natural-language-based, text in and text out format. Among other things, methods and systems of the present disclosure provide machine-learning technologies for combining biological sequence data, representing, for example, DNA, RNA, and protein sequences, with natural language, conversational style prompts that set out particular analysis tasks to be performed on the biological sequence data. This approach, for example, allows complex analysis tasks, including, but not limited to, identification of various sequence modifications, genes, and regulatory elements in DNA sequences, and quantification of properties such as degradation propensity of RNA and protein stability, to be input to a machine learning model in a uniform text-based format and for output to be generated in a same, unified, text-based format.
Owner:INSTADEEP LTD +1

Sequence recognition method, sequence recognition model training method, sequence recognition model training system, sequence recognition equipment and storage medium

PendingCN121884946ABiostatisticsBiological modelsRepresentation ComponentGene
The invention relates to the technical field of biological sequence information identification, in particular to a sequence identification method, a sequence identification model training method, a sequence identification system, equipment and a storage medium. In order to solve the problem that efficiency and precision are difficult to consider when high-throughput sequencing data is processed in the prior art, the invention provides a model comprising a global representation component and a local representation component. During training, a pre-training strategy based on dynamic destruction and reconstruction is adopted for the global representation component to learn global features of the sequence, and a contrast learning strategy is adopted for the local representation component to distinguish local differences of the sequence. During identification, the input sequence is rapidly pre-screened through the global characterization component, and then the candidate sequence is accurately re-screened through the local characterization component, so that the sequence category is determined. According to the method, highly polymorphic biological sequences such as HLA, TCR, KIR and gene fusion can be efficiently and accurately recognized, and the recognition sensitivity and accuracy are improved.
Owner:YICON (BEIJING) BIOMEDICAL TECH INC