Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

38 results about "Molecular sequence" patented technology

Molecular Sequence Data. Descriptions of specific amino acid, carbohydrate, or nucleotide sequences which have appeared in the published literature and/or are deposited in and maintained by databanks such as GENBANK, European Molecular Biology Laboratory (EMBL), National Biomedical Research Foundation (NBRF), or other sequence repositories.

Multi-modal characterization molecular property prediction method based on layered bidirectional cross attention

The invention provides a multi-modal characterization molecular property prediction method based on hierarchical bidirectional cross attention, and relates to the technical field of machine learning assisted organic chemistry, and the method comprises the following steps: S10, generating same-molecule multiple sequences for data enhancement; s20, coding the sequence features through a pre-trained molecular language model MolBERT; s30, performing multi-modal feature fusion through a layered bidirectional cross attention mechanism; s40, establishing a prediction head; s50, in the reasoning stage, only the feature extraction and fusion steps are executed, and a molecular property prediction result is output through the trained prediction head. According to the method, the molecular sequence, the topological graph structure and the fingerprint features are effectively integrated, so that the prediction precision of the model on a plurality of MoleculeNet (molecular network benchmark) public data sets is superior to that of an existing method.
Owner:NANTONG UNIV

Molecular multi-modal characterization method based on attention fusion

The invention discloses a molecular multi-mode characterization method based on attention fusion, and the method comprises the following steps: firstly, carrying out the independent feature extraction of the one-dimensional, two-dimensional and three-dimensional modes of a molecule through a molecular sequence encoder, a molecular map encoder and a molecular conformation encoder, and carrying out the feature extraction of the one-dimensional, two-dimensional and three-dimensional features of the molecule; based on natural language text description, multi-level contrast learning is adopted, and gradual alignment of each mode and the text is achieved; secondly, multi-modal consistency constraint is introduced, the characterization distance of the same molecule in different modals is shortened, and semantic consistency is ensured; thirdly, adaptively integrating multi-modal features through a modal attention fusion module to obtain uniform molecular semantic representation, and further performing final alignment with a text embedding space; and finally, completing tasks such as molecule-text bidirectional retrieval and molecule attribute prediction by utilizing fusion representation. According to the method, under the condition that the model does not need to be retrained for different tasks, natural language-driven molecular retrieval, editing and property prediction can be realized, and the method has relatively high robustness, expansibility and universality.
Owner:SHENYANG INST OF COMPUTING TECH CO LTD THE CHINESE ACAD OF SCI

System and method for detecting real-time anomalies in transmission path of resources within an entity network via DNA computing

PendingUS20260133864A1Non-redundant fault processingDNA computingEngineering
Embodiments of the present invention provide a system for detecting real-time anomalies in transmission path of resources within an entity network via DNA computing. The system is configured for determining that an application received a resource from a source system to be transmitted to an end application via an entity network, extracting metadata from the resource, wherein the metadata is in a binary format, converting the metadata in the binary format to a DNA format, generating predicted molecular sequence associated with predicted transmission path of the resource within the entity network based on the DNA format of the metadata, monitoring real-time transmission path of the resource, generating real-time molecular sequence associated with the real-time transmission path of the resource, comparing the predicted molecular sequence with the real-time molecular sequence to determine an anomaly, and transmitting alerts associated with the anomaly.
Owner:BANK OF AMERICA CORP

Primer and method for identifying sex of grassland caterpillar

The invention discloses a molecular sequence and a primer for sex identification of grassland caterpillars and application of the molecular sequence and the primer. The nucleotide sequences of the molecular marker are as shown in SEQ ID NO: 1 and SEQ ID NO: 2; wherein a nucleotide sequence as shown in SEQ ID NO: 1 is female, and a nucleotide sequence as shown in SEQ ID NO: 2 is male. The invention also provides a primer group for amplifying the molecular fragment, and the nucleotide sequences of the primer group are shown as SEQ ID NO: 3 and SEQ ID NO: 4. Wherein the size of a strip capable of being amplified by a female individual is 484 bp, and the size of a strip capable of being amplified by a male individual is 199 bp. According to the molecular marker and the primer group provided by the invention, sex identification can be carried out on grassland caterpillars, larvae, incomplete pupae and cocoons with unknown sex, the accuracy is high, the speed is high, and a reliable method is provided for monitoring the field sex ratio of the grassland caterpillars.
Owner:LANZHOU UNIV

Application of rice zos2-02 gene in regulating salt tolerance

The purpose of this invention is to disclose a rice salt tolerance-related gene ZOS2-02, its encoded protein, and its applications. The gene ZOS2-02 is a DNA molecule as described in either 1) or 2) below: 1) a DNA molecule with a genomic sequence as shown in SEQ ID NO.1; 2) a DNA molecule with a CDS sequence as shown in SEQ ID NO.2; 3) a DNA molecule that hybridizes to the DNA sequence defined in 1) or 2) under stringent conditions and encodes the protein. The genetic engineering application of the gene ZOS2-02 provided by this invention in regulating rice salt tolerance specifically involves knocking out the aforementioned gene ZOS2-02 to improve rice salt tolerance.
Owner:NANJING AGRICULTURAL UNIVERSITY

Hybrid variant calling

PCT designated stageWO2026043987A1BiostatisticsProteomicsGeneticsVariome
A computer-implemented method for identifying a genomic variant is provided. The method includes obtaining one or more reference molecular sequences and sequencing data pertaining to a biological sample. One or more candidate variant positions are determined from the sequence reads for the biological sample. A classifier model is applied to each candidate variant position for selecting between a haplotype-aware or haplotype-agnostic variant analysis respectively. The classifier model is trained from haplotype structure and / or sequence reads identified with germline variants in a plurality of regions. Based on the application of the classifier model, respectively applying a haplotype-aware or haplotype-agnostic variant analysis to generate a variant identification for each candidate variant position.
Owner:ROCHE SEQUENCING SOLUTIONS INC

A drug interaction prediction method and system based on graph attention sampling

The application discloses a drug interaction prediction method based on graph attention sampling, which firstly acquires drug pair data to be predicted, pre-processes the drug pair data to be predicted, obtains pre-processed drug pair data and drug molecular sequences, converts the pre-processed drug molecular sequences into molecular structure graphs, atomic vector matrices and molecular-atomic relationship lists through a molecular sequence conversion model, constructs an atomic adjacency matrix, a drug interaction network and a drug adjacency matrix, inputs the atomic vector matrix, the atomic adjacency matrix, the molecular-atomic relationship list, the drug interaction network and the drug adjacency matrix into a pre-trained drug interaction prediction model, and obtains a final drug interaction prediction result. The application can solve the technical problem that the existing drug interaction prediction method based on graph representation learning cannot consider both the drug molecular structure characteristics and the relationship information contained in the interaction network.
Owner:HUNAN UNIV

Primer group for detecting PARMS molecular marker of rice plant height gene OsPH4 and application of primer group

The invention belongs to the field of rice breeding, and particularly discloses a primer group for detecting a PARMS molecular marker of a rice plant height gene OsPH4 and application of the primer group, dwarf type OsPH4 gene detection is carried out on a rice material through a molecular sequence of the rice plant height gene OsPH4 by applying a PARMS technology, and the rice plant height gene OsPH4 is obtained. The rice dwarf type OsPH4 gene can be rapidly and accurately detected in different germplasm resources such as indica rice and japonica rice. According to the method, tedious procedures such as enzyme digestion, electrophoresis and sequencing are not needed in the detection process, pollution of PCR product aerosol and use of toxic substances such as EB are reduced, meanwhile, foreground selection and background selection can be conducted in the early stage of molecular marker assisted breeding, the background recovery rate is increased, the scale of breeding populations is reduced, the breeding process is accelerated, and the breeding efficiency is improved. The efficient and environment-friendly application of the OsPH4 gene in commercial molecular breeding of rice is facilitated.
Owner:LIAONING RICE RES INST

Protein drug response prediction method based on em fusion learning

PendingCN122455077AData setProtein
The present application relates to a protein drug reaction prediction method based on EM fusion learning, comprising: 1) constructing a data set containing protein nodes, drug nodes and intermediate nodes into a text attribute graph; 2) embedding the protein amino acid sequence and the drug molecule SMILES sequence to obtain corresponding sequence representations; 3) constructing a transformer-based protein-drug reaction prediction model; 4) constructing a GNN-based protein-drug reaction prediction model; 5) using a variational EM framework to alternately update the above two models to obtain an EM fusion protein-drug reaction prediction model; 6) inputting the to-be-tested protein amino acid sequence and the to-be-tested drug SMILES sequence into the above prediction model to output the reaction prediction result of the to-be-tested protein-drug pair. The present application learns the topological information in the protein-drug network to realize protein-drug reaction prediction.
Owner:LIAONING UNIVERSITY

Virus and disease association prediction model, construction method of prediction model and prediction method

The invention relates to the technical field of interactive prediction, in particular to a virus and disease associated prediction model, a construction method of the prediction model and a prediction method. Comprising the following steps: S1, collecting and processing virus and disease data; s2, performing feature extraction of the independent prediction model A; s3, performing feature extraction of the independent prediction model B; s4, based on five-fold cross validation and grid search, determining an optimal machine learning algorithm, and constructing an independent prediction model A and an independent prediction model B; and S5, based on a bagging integration strategy, integrating the independent prediction model A and the independent prediction model B, and generating a virus and disease associated prediction model. According to the invention, the VDP provides a new thought of multi-modal biological information integration; meanwhile, molecular sequence information of viruses, semantic structure information of diseases and GO function annotation information of the two parties are considered, more comprehensive feature representation is achieved, and the problems that an existing method is single in information source and insufficient in prediction accuracy are solved.
Owner:HUNAN UNIV

Biomolecular sequence searching method and apparatus, device, and storage medium

PCT designated stageWO2026114095A1BiostatisticsSequence analysisAlgorithmSequence search
The present disclosure provides a biomolecular sequence searching method and apparatus, a device, and a storage medium. The method comprises: acquiring an unknown biomolecular sequence; inputting the unknown biomolecular sequence into a sequence encoding model for sequence encoding, to obtain an unknown biomolecular sequence representation; performing similarity searching on different candidate vector sets in a biomolecular sequence vector database by means of the unknown biomolecular sequence representation, wherein candidate biomolecular sequences of candidate vectors in the candidate vector sets have a same sequence length; on the basis of the similarity searching result, searching each candidate vector set for a preset number of known biomolecular sequence representations; and screening known biomolecular sequences corresponding to the plurality of known biomolecular sequence representations for homologous biomolecular sequences corresponding to the unknown biomolecular sequence.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Consensus calling with a neural network

A method and systems for determining a consensus molecular sequence from sequence data of a sample, the sequence data including base calls and quality scores (q-scores). One or more feature vectors are generated and input into a convolutional neural network (CNN) model, which outputs a consensus call sequence. In order to generate the feature vectors, sequence reads are aligned and clustered. For each cluster, one or more feature vectors is then generated by either concatenating base calls and quality scores from the plurality of sequence reads in the cluster, or calculating aggregation statistics for the plurality of sequence reads and concatenating the aggregation statistics and corresponding quality metrics for the plurality of sequence reads in the cluster.
Owner:ROCHE SEQUENCING SOLUTIONS INC

MHC-II type molecular antigen presentation prediction method, electronic equipment and program product

The invention discloses an MHC-II type molecular antigen presentation prediction method. The method comprises the following steps: inputting an MHC-II molecular sequence, an oligopeptide sequence and a context sequence thereof; the MHC-II molecule sequence and the oligopeptide sequence are jointly input into a forward BICL encoder containing a forward BICL convolution kernel and a reverse BICL encoder containing a reverse BICL convolution kernel so as to capture the interaction characteristics of the MHC-II molecule and the oligopeptide under forward binding and reverse binding, and the interaction characteristics under the actual optimal binding direction are determined through score comparison. The method is used for predicting the presentation of the MHC-II molecular antigen. Each of the forward BICL encoder and the reverse BICL encoder comprises a plurality of layers of sensors and a maximum pooling layer, and corresponding interaction representation vectors are obtained after interaction matrixes output by the forward BICL encoder and the reverse BICL encoder are processed.
Owner:FUDAN UNIVERSITY

SgRNA molecule, bovine DDX58 gene single base mutation system, construction method and application

The invention relates to the technical field of animal genetic breeding and genetic engineering, in particular to an sgRNA molecule, a cattle DDX58 gene single base mutation system, a construction method and application. And the molecular sequence of the sgRNA is as shown in SEQ ID NO. 3. According to the method, sgRNA shown in SEQ ID NO.3 is subjected to specific chemical modification, CBE mRNA obtained through in-vitro transcription is jointly delivered to a cattle fertilized egg, the cattle congenital immune key gene DDX58 is precisely edited, finally, the gene edited disease-resistant cattle is obtained, the editing efficiency is high, and the off-target risk is reduced. According to the method, the toxicity problem related to DNA delivery does not exist, a simple, convenient and effective way is provided for simplifying the genome editing development process, the application value of chemical synthesis of sgRNA in genome editing is proved, and wide application of the CRISPR-Cas technology in the fields of biotechnology and treatment is expected to be accelerated.
Owner:NORTHWEST A & F UNIV

Soybean salt-tolerant related gene GmABCG14, and coding protein and application thereof

The application discloses a soybean salt-tolerant related gene GmABCG14, a coding protein and application thereof. The gene GmABCG14 is an NDA molecule as described in 1), 2) or 3) below: 1) a DNA molecule with a genomic sequence as shown in SEQ ID NO. 2; 2) a DNA molecule with a CDS sequence as shown in SEQ ID NO. 2; 3) a DNA molecule hybridized with the DNA sequence defined in 1) or 2) under stringent conditions and encoding the DNA molecule. The application provides a genetic engineering application of the gene GmABCG14 in regulating soybean salt tolerance, specifically overexpressing the aforementioned gene GmABCG14 to improve the salt tolerance of soybean.
Owner:NANJING AGRICULTURAL UNIVERSITY

A processing method and device of a molecular generation model

This invention relates to a method and apparatus for processing molecular generation models. The method includes: configuring a molecular generation instruction template; selecting a large language model and configuring a corresponding post-processing model for it; constructing first, second, and third datasets; training the post-processing model based on the first dataset; performing inference training on the first large language model based on the molecular generation instruction template and the second dataset; after inference training, further strengthening training on the first large language model based on the molecular generation instruction template and the third dataset using a population relative strategy optimization mechanism; after both models are trained, constructing a molecular generation model with the first large language model and the post-processing model as the core; substituting the molecular description input by the user into the molecular generation instruction template to obtain molecular generation instructions; inputting the molecular generation instructions into the molecular generation model for processing to obtain the corresponding molecular sequence and feeding it back to the current user. This invention can improve the chemical accuracy of generated molecules.
Owner:BEIJING DP TECH CO LTD

A new drug molecule design method and device fusing convex optimization and evolutionary learning

ActiveCN119943205Bhigh similarityStrong effectivenessMolecular designBiological modelsEvolutionary learningAlgorithm
The application provides a new drug molecule design method and device fusing convex optimization and evolutionary learning, and belongs to the field of biological information. The method comprises the following steps: in a generative adversarial network, a generator is used to generate a small molecule sequence, the generator is established by adding an attention mechanism model in a long short-term memory network; a discriminator is used to evaluate the authenticity of the small molecule sequence, the discriminator is established by using a convolutional neural network based on convex optimization improvement; a strategy gradient method is used to update the parameters of the generator, and a gradient descent method is used to update the parameters of the discriminator; and a final generator corresponding to the final generator parameters is used to generate a final small molecule sequence. The application can generate small molecules with high similarity to original samples, high effectiveness and strong innovation, improves the diversity and quality of generated molecules, and especially performs well in generating molecules with specific functions.
Owner:NORTHEASTERN UNIV CHINA

Method for predicting structure of compound model, method for training model, and related apparatuses

A method for predicting a structure of a compound includes: obtaining a combination of biomolecular sequences by combining specified biomolecular sequences; predicting a first probability distribution for the combination of biomolecular sequences, in which the first probability distribution is used for indicating first probabilities of candidate structural unit groups in the combination of biomolecular sequences, a candidate structural unit group includes structural units of at least two biomolecular sequences, and a first probability is used for indicating a possibility of interaction between structural units in a respective candidate structural unit group; determining, from the plurality of candidate structural unit groups, at least one first structural unit group based on the first probability distribution; and predicting a target structure of a biomolecular compound based on structural units interacted with each other in the at least one first structural unit group.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Systems and methods for predicting biological responses

Systems and methods can apply machine learning techniques to predict biological responses. One of the methods is performed by at least one processor executing executable logic including at least one machine learning model trained to predict biological responses. The method includes receiving first sequence data of a first molecular sequence, receiving second sequence data of a second molecular sequence, and predicting a biological response for the second molecular sequence based at least partly on the received first and second sequence data.
Owner:SANOFI PASTEUR INC

Preparation method of scar-free annular nucleic acid molecule and application of scar-free annular nucleic acid molecule in targeted delivery

The invention provides a preparation method of a scar-free annular nucleic acid molecule and application of the scar-free annular nucleic acid molecule in targeted delivery. The precursor nucleic acid molecule for preparing the circular nucleic acid molecule comprises, along the direction from 5'to 3 ', a. 3' I group intron or a mutant fragment thereof, b. A first unit fragment II of which the 5'end comprises a ribozyme recognition fragment II, c. A functional unit (comprising IRES, an RNA aptamer, a protein binding sequence, a protein coding region, a non-coding region and the like or a combination thereof), e. A first unit fragment I of which the 3 'end comprises a ribozyme recognition fragment I, and e. A second unit fragment II of which the 5' end comprises a ribozyme recognition fragment II. And f.5 'I group introns or mutant fragments thereof. According to the invention, scar-free cyclization of RNA is realized by using the nucleic acid aptamer sequence, and the targeting property of the circular nucleic acid molecule sequence is improved.
Owner:HANGZHOU INSTITUTE OF MEDICAL SCIENCES CHINESE ACADEMY OF SCIENCES

Systems and methods for predicting biological responses

PendingUS20260187450A1Data miningBioinformatics
Systems and methods can apply machine learning techniques to predict biological responses. One of the methods is performed by at least one processor executing executable logic including at least one machine learning model trained to predict biological responses. The method includes receiving first sequence data of a first molecular sequence, receiving second sequence data of a second molecular sequence, and predicting a biological response for the second molecular sequence based at least partly on the received first and second sequence data.
Owner:SANOFI PASTEUR INC

A molecular structure semantic search method, device, equipment, medium and product

This application discloses a molecular structure semantic search method, apparatus, device, medium, and product, relating to the field of data processing technology. The method includes obtaining the molecular sequence representation of a target molecule to be queried; extracting bimodal features from the molecular sequence representation to obtain an implicit deep semantic vector and an explicit topological fingerprint vector; fusing the implicit deep semantic vector and the explicit topological fingerprint vector to form a multimodal representation vector of the target molecule; using the multimodal representation vector as the query vector; performing a similarity search in a vectorized molecular knowledge base to obtain related factual information of one or more reference molecules; enhancing the retrieval of the related factual information of the reference molecules and the molecular sequence representation of the target molecule through a large language model; and generating a consistency analysis report of the target molecule's structure and spectrum. This application improves the retrieval accuracy and semantic matching capability for complex molecules, ensures the scientific validity and reliability of the output results, and significantly improves the efficiency of chemical data analysis.
Owner:SHANGHAI DEV CENT OF COMP SOFTWARE TECH

Soybean salt-tolerant related gene GmNAT12, and coding protein and application thereof

The application discloses a soybean salt-tolerant related gene GmNAT12, a coding protein thereof and application thereof. The gene GmNAT12 is an NDA molecule as described in 1) or 2) or 3) below: 1) a DNA molecule with a genomic sequence as shown in SEQ ID NO. 1; 2) a DNA molecule with a CDS sequence as shown in SEQ ID NO. 2; 3) a DNA molecule hybridized with the DNA sequence defined in 1) or 2) under stringent conditions and encoding the DNA molecule. The application provides a genetic engineering application of the gene GmNAT12 in regulating soybean salt tolerance, specifically overexpressing the aforementioned gene GmNAT12 to improve the salt tolerance of soybean.
Owner:NANJING AGRICULTURAL UNIVERSITY

Sequence generation method and device, electronic equipment, storage medium and program product

The invention provides a sequence generation method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: coding a molecular sequence to obtain a first molecular coding feature, the molecular sequence comprising a heavy chain sequence and a light chain sequence; determining a heavy chain coding feature and a light chain coding feature based on the heavy chain sequence, the light chain sequence and the first molecular coding feature; fusing the heavy chain coding feature and the light chain coding feature to obtain a target coding feature; and performing sequence generation on the target coding feature to obtain a target molecular sequence. According to the method and the device, the complex sequence relationship between the paired light chain sequence and heavy chain sequence can be captured, and the accuracy of sequence generation is improved.
Owner:SHENZHEN TENCENT COMP SYST CO LTD +1

Genome scale metabolic network model deletion reaction prediction and filling method and system

PendingCN122090962AFully capture dynamic charactersFully capture interactionsBiostatisticsBiological modelsMetaboliteData set
The invention discloses a genome scale metabolic network model deletion reaction prediction and filling method and system, and the method comprises the steps: obtaining a plurality of genome scale metabolic network models from a metabolic network database, and constructing a data set with a supervision label; extracting and fusing the molecular sequence and molecular map characteristics of the metabolite to obtain an initial characteristic vector; constructing a metabolic directed graph and a metabolic reaction hypergraph based on a positive reaction sample, performing feature directivity enhancement by using the directed graph, and extracting high-order topological information through a hypergraph convolutional neural network to obtain final feature representation of metabolites; on the basis of the feature representation, adopting an attention mechanism to predict candidate reaction confidence and training a model; and finally, screening a high-confidence reaction from the candidate reaction pool and filling the target model with the high-confidence reaction. According to the method, the metabolite multi-dimensional molecular characteristics and the network high-order topological information are deeply fused, the prediction accuracy is remarkably improved, the method does not depend on experimental data, and the method is suitable for efficient metabolic network model correction and optimization.
Owner:JIANGNAN UNIV

Preparation method for scarless circular nucleic acid molecule and application thereof in targeted delivery

Provided are a preparation method for a scarless circular nucleic acid molecule and an application thereof in targeted delivery. The precursor nucleic acid molecule used for preparing the circular nucleic acid molecule comprises, in the direction from 5' to 3': a. a 3' group I intron or a mutant fragment thereof, b. a first unit fragment II, comprising a recognition fragment of ribozyme II at the 5' end thereof, c. a functional unit (comprising IRES, an aptamer, a protein binding sequence, a protein coding region, a non-coding region, etc. or a combination thereof), d. a first unit fragment I, comprising a recognition fragment of ribozyme I at the 3' end thereof, and e. 5' group I intron or a mutant fragment thereof. The use of the aptamer sequence achieves scarless cyclization of RNA, and improves the targeting property of the circular nucleic acid molecule sequence.
Owner:HANGZHOU INSTITUTE OF MEDICAL SCIENCES CHINESE ACADEMY OF SCIENCES

LAMP primer set, kit and application for detecting peony yellow spot pathogen

This invention relates to LAMP primer sets, kits, and applications for detecting *Pseudomonas aeruginosa*, belonging to the field of molecular biology. Based on the conserved internal transcribed spacer (ITS) region of the ribosomal DNA of *Pseudomonas aeruginosa*, this invention designs in vitro DNA amplification primers. Through color reactions and electrophoretic banding after in vitro amplification, a simple and rapid method for detecting *Pseudomonas aeruginosa* is established. This method solves some technical problems existing in the early diagnosis of peony yellow spot disease based on morphological characteristics combined with molecular sequence identification. Compared with existing technologies, the LAMP detection technology for *Pseudomonas aeruginosa* established in this invention is rapid, efficient, low-cost, highly sensitive, highly specific, and convenient for product detection, providing a theoretical basis for the early diagnosis and timely control of peony yellow spot disease in production.
Owner:HENAN UNIV OF SCI & TECH

Generalizable transformer for disease state classification from liquid biopsies

A system and method for training a generalizable transformer model for disease state classification. The system obtains general training data from subjects non-specific to a target disease. The system trains an attention subspace as a latent representation of sequencing data in the general training data. The system obtains targeted training data from subjects associated with the target disease. For the targeted training data, the system generates a molecule embedding for each molecular sequence read in each sample based on the molecular information, generates a genome-wide positional encoding for each molecular embedding based on positional information of the molecular sequence read, applies an inverse attention head to determine a weight vector for each molecule embedding, generates a sample matrix as a product of the molecule embeddings, the positional encodings, and the weight vectors, and applies a feed-forward network to classify disease state of the target disease with the sample matrix.
Owner:HEPTA BIO INC

A processing method and device of a molecular description generation model

This invention relates to a method and apparatus for processing molecular description generation models. The method includes: constructing a 3D structure and physicochemical feature recognition model, selecting a large language model, configuring a molecular description instruction template, and constructing a molecular description generation model with the 3D structure and physicochemical feature recognition model and the large language model as the core; first training the 3D structure and physicochemical feature recognition model, then performing a first-stage targeted fine-tuning of the molecular description generation model's reasoning ability, and finally performing a second-stage enhancement of the molecular description generation model based on a population relative strategy optimization algorithm; after training, substituting the user-input molecular sequence S into the molecular description instruction template to generate an initial instruction X0, inputting X0 into the molecular description generation model for processing to obtain the corresponding molecular description D, and feeding it back to the current user. This invention can improve the model's understanding of the essence of molecules, improve the model's chemical reasoning ability, and improve the chemical accuracy of molecular descriptions.
Owner:BEIJING DP TECH CO LTD

Zero-sample universal affinity prediction model and construction method thereof

The invention relates to a zero-sample general affinity prediction model and a construction method thereof. The construction method of the zero-sample general affinity prediction model comprises the following steps: constructing a cross-target molecular dynamics simulation dynamic mode library; constructing a molecular sequence feature extraction module, wherein the molecular sequence feature extraction module is used for performing molecular sequence feature extraction based on the large language model; and constructing a fusion module, wherein the fusion module fuses the cross-target molecular dynamics simulation dynamic mode library with the molecular sequence features. According to the method, high-precision prediction of the protein-ligand affinity can be realized under a new target spot (zero sample scene) which is not seen at all, the dependence of a traditional model on specific data is broken through, the starting cost of new target spot research is remarkably reduced, and the method is particularly suitable for prediction requirements of rare disease target spots and emerging target spots, so that the whole research and development process is accelerated.
Owner:DIVAMICS INC