Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

25 results about "Biosequence" patented technology

A BioSequence is an object representation of a DNA, RNA, or protein sequence. It can be represented by a Clone, Gene, or the sequence. (caMAGE)

Group health evaluation method and system based on intestinal flora

The invention relates to the technical field of health evaluation, in particular to a group image health evaluation method and system based on intestinal flora, and the method comprises the steps: collecting intestinal flora samples of a target group through a standardized process, and obtaining microbial sequence information through a high-throughput sequencing technology; microflora diversity characteristics and core flora characteristics are extracted through strain identification and relative abundance analysis, individual flora characteristics are compared with a healthy population reference database, health state scores are calculated, health grades are divided, population health grade distribution is counted, and the population health grade distribution is calculated. And generating a group health evaluation report containing a visual chart and text analysis. According to the method, a complete technical system from sample collection to health strategy making is established, systematic evaluation of the group health state from the perspective of intestinal flora is achieved, and the limitation of a traditional method in group health early evaluation is overcome. The method is suitable for health monitoring of different scales of groups, and provides effective technical support for public health management and health service.
Owner:SECOND MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL

Biological sequence compression using sequence alignment

Compressing files is disclosed. An DNA sequence to be compressed is first aligned. Aligning the DNA sequence includes splitting the DNA sequences into smaller sequences or portions that can be aligned. After the DNA sequence is spilt one or more time and aligned, a compression matrix is generated. Each row of the compression matrix corresponds to part of the DNA sequence. A consensus sequence is determined from the compression matrix. Using the consensus sequence, pointer pairs are generated. Each pointer pair identifies a subsequence of the consensus matrix. The compressed file includes the pointer pairs and the consensus sequence.
Owner:DELL PROD LP

Method for constructing an online prediction website based on biological sequence cis-acting regulatory elements

This invention discloses a method for constructing an online prediction website based on cis-regulatory elements of biological sequences. This method is applicable to building online prediction websites from models trained using machine learning on biological sequences. With the continuous development of artificial intelligence technology, machine learning algorithms have flourished and been successfully applied in multiple fields. In particular, in the fields of biological and medical research, researchers have constructed numerous prediction models to serve various research tasks, such as the prediction of cis-regulatory elements, biomedical pathogenic elements, and proteins. Although many prediction models have been successfully constructed, professionals who frequently use computers are more concerned with how to apply these models to practical work. Therefore, constructing the prediction models in this paper into user-friendly and easy-to-use online websites can greatly enhance the practical application value of the models.
Owner:GUILIN UNIV OF ELECTRONIC TECH

IncRNA and disease association prediction method based on deep learning

The invention belongs to the technical field of bioinformatics, and particularly relates to an lncRNA and disease association prediction method based on deep learning, which comprises the following specific steps: firstly, according to known lncRNA-disease association information, disease-miRNA association information and lncRNA-miRNA association information, constructing an lncRNA-disease association matrix, a disease-miRNA association matrix and an lncRNA-miRNA association matrix; and constructing an lncRNA comprehensive similarity matrix LS, a disease comprehensive similarity matrix DS and a miRNA comprehensive similarity matrix MS. The biological sequence selective compression network adopted by the invention can effectively model a long-distance dependency relationship in an ultra-long sequence and dynamically distinguish key functional fragments from redundant fragments according to context.
Owner:GUANGDONG UNIV OF TECH

Synthetic biological data intelligent extraction system based on multi-modal large language model

The invention discloses a synthetic biological data intelligent extraction system based on a multi-modal large language model, and belongs to the technical field of biological data intelligent extraction, and the system comprises a multi-modal data preprocessing and standardization module which is used for carrying out the analysis, classification and structured output of input heterogeneous synthetic biological data; and the multi-modal large language model processing module is connected with the preprocessing module and is used for receiving the processed structured data and carrying out joint understanding and information extraction through knowledge enhancement and deep fusion technologies. The heterogeneous synthetic biological data can be efficiently processed, different types of data such as scientific literatures, chart images and biological sequences are accurately analyzed and classified through the multi-modal data preprocessing and standardization module, the data format barrier is broken, scattered biological information is systematically integrated, and the data processing efficiency is improved. A standard and unified data foundation is laid for subsequent information extraction, and the comprehensiveness and effectiveness of data processing are greatly improved.
Owner:AIXBIO (HANGZHOU) BIOTECHNOLOGY CO LTD

Method for constructing database, method for retrieving document and computer device

Disclosed are a method for constructing a database, a method for labeling an association degree of biological sequences, a method for retrieving a document, and a computer device. In the solution of this application, a biological sequence and attribute information are extracted from a target document, and an entry in a database is constructed based on the extracted biological sequence and the attribute information. When a user conducts retrieval based on the database, a server can match an entry for the user by means of the biological sequence and the attribute information in the entry or a combination of the two. Therefore, when applied to a retrieval platform, the database of this application can provide the user with various types of retrieval support, such as biological sequence retrieval, biological sequence attribute retrieval, and comprehensive biological sequence and biological sequence attribute retrieval, and the like.
Owner:PATSNAP LIMITED

Fitness space roughness evaluation method and device, electronic equipment and storage medium

The application provides a roughness evaluation method and device of fitness space, electronic equipment and storage medium, and relates to the technical field of biology. The method comprises the following steps: acquiring a data set; the data set comprises a plurality of biological sequences; based on the mutation sites between each biological sequence in the data set, a plurality of adjacent sequence pairs in the data set are determined; the mutation sites between the two biological sequences in each adjacent sequence pair meet the set number of mutation sites; and the roughness of the fitness space is estimated according to each adjacent sequence pair, to obtain a roughness evaluation result. The application solves the problem of inaccurate roughness evaluation of the fitness space in the related art.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

A method and system for RNA secondary structure prediction and folding path deduction

The application belongs to the technical field of biological sequence data processing, and discloses a method and system for predicting and deducing folding paths of RNA secondary structure, which comprises: inputting a to-be-predicted RNA sequence into a deep learning pre-training model to perform feature extraction and pattern recognition, and obtaining a plurality of local structure modules each composed of at least three nucleotide pairs; regarding all the local structure modules as an action space of a reinforcement learning agent, and enabling the reinforcement learning agent to perform filtering and pairing actions on the local structure modules, gradually output a prediction result of the RNA secondary structure, and record the local structure modules and pairing sequences selected at each step to obtain a folding path, wherein the reinforcement learning agent adopts a composite reward function comprising a reward structure module formation, a punishment topological conflict, and an evaluation of local stability, and adaptively adjusts the weights of the composite reward function according to a folding stage and a current state. The application can improve the calculation efficiency and generalization ability, and reproduce the RNA folding path.
Owner:BEIJING NEOCURNA BIOTECHNOLOGY CORP

Systems and methods for multimodal conversational agents for biological sequence analysis

Provided herein are technologies for framing and evaluating biological sequence-based analysis tasks in a unified, natural-language-based, text in and text out format. Among other things, methods and systems of the present disclosure provide machine-learning technologies for combining biological sequence data, representing, for example, DNA, RNA, and protein sequences, with natural language, conversational style prompts that set out particular analysis tasks to be performed on the biological sequence data. This approach, for example, allows complex analysis tasks, including, but not limited to, identification of various sequence modifications, genes, and regulatory elements in DNA sequences, and quantification of properties such as degradation propensity of RNA and protein stability, to be input to a machine learning model in a uniform text-based format and for output to be generated in a same, unified, text-based format.
Owner:INSTADEEP LTD +1

Method, device, electronic equipment and program product for predicting siRNA silencing efficiency

PendingCN122658423ALinguistic modelEngineering
The present disclosure provides a siRNA silencing efficiency prediction method, device, electronic equipment and program product, comprising the following steps: obtaining a pseudo-label feature value of a to-be-predicted siRNA sequence based on a language model; searching for an off-target effect score of the to-be-predicted siRNA sequence based on a search model; obtaining a rule score of the to-be-predicted siRNA sequence based on a preset rule; performing feature extraction based on the to-be-predicted siRNA sequence to obtain sequence features; and obtaining a silencing efficiency of the to-be-predicted siRNA sequence based on a machine learning model according to the pseudo-label feature value, the off-target effect score, the rule score and the sequence features. The present disclosure comprehensively utilizes the biological sequence features of traditional rules and the strong learning ability of deep learning models, improves the accuracy and robustness of siRNA silencing efficiency prediction, reduces the dependence on a large amount of labeled data, and enhances the generalization ability among different biological samples, thereby providing a more reliable and practical tool for siRNA design.
Owner:SHANGHAI PHARMACEUTICALS HOLDING CO LTD

Cancer resistance peptide generation and screening method and system based on diffusion model and seed synchronous autoencoder

PendingCN122117054AImprove biological effectivenessSolve mapping collapse problemEnsemble learningBiostatisticsRandom seedCancer resistance
The application discloses an anticancer peptide generation and screening method and system based on a diffusion model and a seed synchronous autoencoder, and comprises the following steps: preprocessing original anticancer peptide and non-anticancer peptide sequences and extracting multi-dimensional physicochemical features; training a diffusion model in a continuous numerical space based on sequence mapping based on a global random seed, adding noise through forward diffusion and learning the potential distribution of anticancer peptide data features through reverse denoising; based on a reverse seed synchronization strategy, training an autoencoder using the same random seed and environment state, and establishing a mapping from continuous random noise to discrete biological sequences; using the diffusion model to sample and generate potential vectors in the latent space and mapping them into candidate anticancer peptide sequences through the decoded synchronously trained decoder; inputting an integrated learning classifier based on grouping features, selecting the optimal feature group and predicting the anticancer activity probability, and outputting the anticancer characteristic peptide chain; the application solves the mapping contradiction between continuous space and discrete sequences and can efficiently screen anticancer peptides with high activity sequences.
Owner:CHONGQING UNIV

Architectures for training neural networks using biological sequences, conservation, and molecular phenotypes

ActiveUS12626782B2BiostatisticsProteomicsMolecular phenotypeData set
The present disclosure provides methods and systems that can ascertain how genetic variants impact molecular phenotypes. Such methods and systems may use additional conservation information. In an aspect, the present disclosure provides a method for training a molecular phenotype neural network (MPNN), comprising: (a) providing a molecular phenotype neural network (MPNN) comprising one or more parameters; (b) providing a training data set comprising (i) a set of one or more inputs comprising biological sequences and (ii) for each input in the set of one or more inputs, a set of one or more molecular phenotypes corresponding to the input; (c) configuring the one or more parameters of the MPNN based on the training data set to minimize a total loss of the training data set, thereby training the MPNN; and (d) outputting the one or more parameters of the MPNN.
Owner:DEEP GENOMICS INC

Parallel-processing systems and methods for highly scalable analysis of biological sequence data

An apparatus includes a memory configured to store a sequence. The sequence includes an estimation of a biological sequence. The sequence includes a set of elements. The apparatus also includes a set of hardware processors. Each hardware processor is configured to implement a segment processing module. The apparatus also includes an assignment module implemented in a hardware processor. The assignment module is configured to receive the sequence from the memory, and assign each element to at least one segment from a set of segments, including, when an element maps to at least a first segment and a second segment, assigning the element to both the first segment and the second segment. The segment processing module is configured to, for each segment from a set of segments specific to that hardware processor, and substantially simultaneous with the remaining hardware processors, remove at least a portion of duplicate elements in that segment to generate a deduplicated segment. The segment processing module is further configured to reorder the elements in the deduplicated segment to generate a realigned segment that has a reduced likelihood for alignment errors.
Owner:RES INST AT NATIONWIDE CHILDRENS HOSPITAL

Method for analysis and query-based interrogation of biological sequence data

PCT designated stageWO2026132228A1ProteomicsGenomicsData decompositionProtein
The invention relates to a computer-implemented method for creating a searchable indexed data structure for biological sequence data. The invention further relates to searching in said data structure for any given biological sequence data or associated information. The invention relates to a computer-implemented method comprising: a. providing biological sequence input data, b. decomposing the biological input data into non-overlapping portions, c. indexing each portion of the biological input data with at least one associated information of said portion comprised within or provided with the biological input data, thereby creating a two-tiered index for each portion. The method may further comprise d. receiving at least one query from a user, wherein the query comprises biological sequence information, language-based prompts and / or queries regarding the metadata comprised within or provided with the biological input data. The method may further comprise e. providing at least one output comprising at least one biological sequence information, biological sequence identifier and / or metadata of one or more portions of the biological sequence input data corresponding to the query. The invention further relates to the use of the method of the invention for identifying tissue and / or cell localized expression of at least one nucleic acid, such as a gene, and / or protein of interest. The method further relates to a data structure and a computer-readable medium comprising a data structure, such as a database, constructed according to the invention.
Owner:MAX DELBRUECK CENT FUER MOLEKULARE MEDIZIN

Methods for classifying biological sequences, computer programs, and computer implementations (identification of unknown genomes and recently known genomes).

To provide a method, a computer program and a computer-implemented method for classifying biological sequences (identification of unknown genomes and closest known genomes).SOLUTION: A method for classifying biological sequences comprises: providing a training set related to a biological sequence; dividing the training set and the biological sequence into sequence fragments; extracting feature vectors from the sequence fragments of the training set and from the sequence fragments of the biological sequence; identifying feature vectors extracted from the sequence fragments of the biological sequence that correspond to feature vectors extracted from the sequence fragments of the training set; establishing a threshold value comprising an average degree of divergence of the feature vectors from the sequence fragments of the biological sequence, from the feature vectors from the sequence fragments of the training set; and detecting an anomaly among the extracted features of the sequence fragments of the biological sequence.SELECTED DRAWING: None
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Digital to biological converter

The present invention provides a system for receiving biological sequence information and activating the synthesis of a biological entity. The system has a receiving unit for receiving a signal encoding biological sequence information transmitted from a transmitting unit. The transmitting unit can be present at a remote location from the receiving unit. The system also has an assembly unit connected to the receiving unit, and the assembly unit assembles the biological entity according to the biological sequence information. Thus, according to the present invention biological sequence information can be digitally transmitted to a remote location and the information converted into a biological entity, for example a protein useful as a vaccine, immediately upon being received by the receiving unit and without further human intervention after preparing the system for receipt of the information. The invention is useful, for example, for rapidly responding to viral and other biological threats that are specific to a particular locale.
Owner:TELESIS BIO INC

Method for obtaining and correcting biological sequence information

This disclosure provides methods for sequencing biomolecules such as nucleic acid molecules, and methods for detecting and / or correcting sequencing errors in sequencing results. Reagent kits and systems based on the methods disclosed herein are also provided.
Owner:CYGNUS BIOSCI BEIJING CO LTD

Data processing method and device, electronic equipment, computer readable storage medium and computer program product

The invention provides a data processing method and device, electronic equipment, a computer readable storage medium and a computer program product. The method is applied to artificial intelligence. The method comprises the following steps: determining a monomics sequence of a biomolecule based on a molecular sequence of the biomolecule; performing sequence conversion on the molecular sequence to obtain a multi-omics sequence of the biomolecule; fusing the single omics sequence and the multi-omics sequence to obtain a biological sequence; coding the biological sequence to obtain biological sequence coding characteristics of the biological molecules; and performing prediction processing on the biological sequence coding features to obtain molecular prediction information of the biological molecules. According to the invention, the accuracy of molecular prediction information can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Sequence recognition method, sequence recognition model training method, sequence recognition model training system, sequence recognition equipment and storage medium

PendingCN121884946ABiostatisticsBiological modelsRepresentation ComponentGene
The invention relates to the technical field of biological sequence information identification, in particular to a sequence identification method, a sequence identification model training method, a sequence identification system, equipment and a storage medium. In order to solve the problem that efficiency and precision are difficult to consider when high-throughput sequencing data is processed in the prior art, the invention provides a model comprising a global representation component and a local representation component. During training, a pre-training strategy based on dynamic destruction and reconstruction is adopted for the global representation component to learn global features of the sequence, and a contrast learning strategy is adopted for the local representation component to distinguish local differences of the sequence. During identification, the input sequence is rapidly pre-screened through the global characterization component, and then the candidate sequence is accurately re-screened through the local characterization component, so that the sequence category is determined. According to the method, highly polymorphic biological sequences such as HLA, TCR, KIR and gene fusion can be efficiently and accurately recognized, and the recognition sensitivity and accuracy are improved.
Owner:YICON (BEIJING) BIOMEDICAL TECH INC

Multi-biological sequence model unified management method and system, device, medium and product

PendingCN121687193ASequence analysisInference methodsMemory orderingData mining
The invention discloses a multi-biological sequence model unified management method, system and device, a medium and a product. The method comprises the steps that in response to a biological sequence reasoning requirement, biological sequence task information is obtained, and the biological sequence task information comprises multiple pieces of sequence information; analyzing each piece of sequence information to obtain a corresponding biological sequence structure feature and a sequence type thereof; obtaining a biological sequence target model corresponding to each sequence type; grouping the multiple pieces of sequence information to obtain a plurality of subgroup sequence sets; memory requirements corresponding to each subgroup sequence set are obtained and sorted, a required memory sorting result is obtained, and batch strategies corresponding to the multiple pieces of sequence information are determined according to the current actual memory; and dynamically determining a model loading strategy according to the batch strategy, and executing a reasoning task according to the model loading strategy. According to the method, when a large amount of diversified biological sequence data is processed, appropriate models can be efficiently and intelligently selected and managed, and effective utilization of resources and efficient execution of tasks are achieved.
Owner:BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD

Predicting labels for biological sequences using neural networks conditioned on positive and negative examples

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for predicting labels for biological sequences. One of the methods includes, in response to receiving a request to identify labels associated with an input biological sequence: determining, for each of a plurality of candidate labels, a score characterizing a likelihood that the input biological sequence is associated with the candidate label. Each score is determined by identifying a plurality of positive biological sequences that are each associated with the candidate label; and processing a network input including the input biological sequence and the plurality of positive biological sequences using a neural network to generate the score characterizing the likelihood that the input biological sequence is associated with the candidate label. The method includes selecting one or more of the candidate labels as labels for the input biological sequence based on the scores.
Owner:DEEPMIND TECH LTD

Parallel-processing systems and methods for highly scalable analysis of biological sequence data

An apparatus includes a set of hardware processors and a memory configured to store a sequence. The sequence includes a set of elements. Hardware processors are configured to implement a segment processing module, and an assignment module. The assignment module receives the sequence and assigns each element to at least one segment from a set of segments, including, when an element maps to at least a first segment and a second segment, assigning the element to both the first segment and the second segment. A segment processing module is configured to substantially simultaneously, for each segment from a set of segments specific to that one or more hardware processors, remove at least a portion of duplicate elements to generate a deduplicated segment. The segment processing module reorders the elements in the deduplicated segment to generate a realigned segment that has a reduced alignment errors.
Owner:RES INST AT NATIONWIDE CHILDRENS HOSPITAL

Systems and methods for modulating a target gene

Provided herein are compositions, methods, and systems for modulating expression of a target gene (e.g., a target endogenous gene). In some embodiments, an engineered genetic effector is provided, the engineered genetic effector comprising a first peptide having a length of 75 to 95 amino acids and a second peptide having a length of 75 to 95 amino acids. The engineered gene effector can facilitate modulation of the expression level or activity level of a target gene when approaching the target gene or target gene regulatory sequence in a complex with a targeting moiety, such as a heterologous endonuclease. Also provided are computer-implemented methods for producing functional biological sequences, as well as functional biological sequences, such as engineered genetic effectors made by the methods.
Owner:EPICRISPR BIOTECHNOLOGIES INC