Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

72 results about "Molecular sequence" patented technology

Molecular Sequence Data. Descriptions of specific amino acid, carbohydrate, or nucleotide sequences which have appeared in the published literature and/or are deposited in and maintained by databanks such as GENBANK, European Molecular Biology Laboratory (EMBL), National Biomedical Research Foundation (NBRF), or other sequence repositories.

Multi-modal characterization molecular property prediction method based on layered bidirectional cross attention

The invention provides a multi-modal characterization molecular property prediction method based on hierarchical bidirectional cross attention, and relates to the technical field of machine learning assisted organic chemistry, and the method comprises the following steps: S10, generating same-molecule multiple sequences for data enhancement; s20, coding the sequence features through a pre-trained molecular language model MolBERT; s30, performing multi-modal feature fusion through a layered bidirectional cross attention mechanism; s40, establishing a prediction head; s50, in the reasoning stage, only the feature extraction and fusion steps are executed, and a molecular property prediction result is output through the trained prediction head. According to the method, the molecular sequence, the topological graph structure and the fingerprint features are effectively integrated, so that the prediction precision of the model on a plurality of MoleculeNet (molecular network benchmark) public data sets is superior to that of an existing method.
Owner:NANTONG UNIV

Molecular property prediction method and system based on multi-task pre-training and multi-modal fusion

The invention relates to a molecular property prediction method and system based on multi-task pre-training and multi-modal fusion. The method comprises the following steps: collecting a data set and performing molecular conversion; performing multi-task pre-training, which comprises the following steps: generating a heterogeneous enhanced view; constructing a pseudo label; performing comparative learning according to the structure enhanced view and the heterogeneous enhanced view, and constructing a maximum similar task; capturing semantic differences among molecules according to the pseudo labels; carrying out multi-modal fusion, namely introducing functional group structure information, and extracting molecular sequence characteristics based on Transform and Mamba2 to obtain fused molecular multi-modal representation; and analyzing and predicting the classification or regression task. A heterogeneous enhanced view is established in a multi-task pre-training stage, multi-task self-supervision is performed in combination with multi-granularity features of molecular fingerprints, and functional group structure information is introduced in a multi-modal fusion stage, so that deep cross-modal interaction is realized, downstream prediction performance is improved, and candidate drug screening and molecular property evaluation processes are accelerated.
Owner:HAINAN UNIV

Intermolecular consensus of partially read sequences

PCT designated stageWO2025235888A1Sequence analysisInstrumentsBinding siteData mining
A computer-implemented method for identifying a consensus molecular sequence from full length and partial length read sequencing data of a sample. The method includes sorting the sequencing data into full length and partial length reads with 5' and / or 3' binding sites, clustering the full length reads and the partial length reads, the clustering based on a position of binding site(s), and comparing a partial length read cluster to a full length read cluster to determine a positional distance between the compared clusters. In response to determining that the positional distance between the compared clusters falls within a redetermined threshold, merging the compared clusters, and identifying a consensus molecular sequence of the sample based on the merged clusters.
Owner:ROCHE SEQUENCING SOLUTIONS INC

High-precision molecular property prediction method and device based on multi-modal information

The invention discloses a high-precision molecular property prediction method and device based on multi-modal information. The method comprises the following steps: acquiring a molecular map, a molecular sequence and a molecular three-dimensional geometric configuration of a molecule; decomposing each molecular graph into a group of motifs, constructing a multi-stage pre-training framework to capture multi-scale information in the molecule, and performing molecular graph training on the molecule to obtain a molecular graph feature vector; designing an information transmission mechanism based on distance and angle, and carrying out learning training on the molecular three-dimensional geometric configuration to obtain a molecular geometric feature vector; preprocessing the molecular sequence feature vector of the molecule, and learning the molecular sequence information of the molecule to obtain a molecular sequence feature vector; and according to the molecular map feature vector, the molecular geometric feature vector and the molecular sequence feature vector, designing a molecular property prediction model based on a cross-modal cross attention mechanism to perform molecular property prediction. The method can enhance the robustness and accuracy of molecular property prediction.
Owner:ZHEJIANG UNIV

Molecular multi-modal characterization method based on attention fusion

The invention discloses a molecular multi-mode characterization method based on attention fusion, and the method comprises the following steps: firstly, carrying out the independent feature extraction of the one-dimensional, two-dimensional and three-dimensional modes of a molecule through a molecular sequence encoder, a molecular map encoder and a molecular conformation encoder, and carrying out the feature extraction of the one-dimensional, two-dimensional and three-dimensional features of the molecule; based on natural language text description, multi-level contrast learning is adopted, and gradual alignment of each mode and the text is achieved; secondly, multi-modal consistency constraint is introduced, the characterization distance of the same molecule in different modals is shortened, and semantic consistency is ensured; thirdly, adaptively integrating multi-modal features through a modal attention fusion module to obtain uniform molecular semantic representation, and further performing final alignment with a text embedding space; and finally, completing tasks such as molecule-text bidirectional retrieval and molecule attribute prediction by utilizing fusion representation. According to the method, under the condition that the model does not need to be retrained for different tasks, natural language-driven molecular retrieval, editing and property prediction can be realized, and the method has relatively high robustness, expansibility and universality.
Owner:SHENYANG INST OF COMPUTING TECH CO LTD THE CHINESE ACAD OF SCI

Soybean salt-tolerant related gene GmRLP15, and coding protein and application thereof

The application discloses a soybean salt-tolerant related gene GmRLP15, a coding protein thereof and application thereof. The gene GmRLP15 is an NDA molecule as described in 1), 2) or 3) below: 1) a DNA molecule with a genomic sequence as shown in SEQ ID NO. 1; 2) a DNA molecule with a CDS sequence as shown in SEQ ID NO. 2; 3) a DNA molecule hybridized with the DNA sequence defined in 1) or 2) under stringent conditions and encoding the DNA molecule. The application provides a genetic engineering application of the gene GmRLP15 in regulating soybean salt tolerance, specifically overexpressing the aforementioned gene GmRLP15 to improve the salt tolerance of soybean.
Owner:NANJING AGRICULTURAL UNIVERSITY

A method for predicting sequence structure of biological macromolecule

The application discloses a biological macromolecule sequence structure prediction method, and relates to the technical field of computers. In the diffusion process, a geometric position coding gating mechanism is introduced, and a multi-scale attention broadcast mechanism is used to broadcast the attention weights corresponding to different attention types of the target information processed in different broadcast range modes, and the attention weights in different time steps in the broadcast range are reused, so that the same attention weight is used in the steps of the broadcast range, the calculation amount is reduced, the model efficiency performance is improved, the perception ability of the model to the amino acid position information in the protein sequence can be effectively improved, the model can fully consider the position relationship between amino acids when calculating the attention weight, the model can better identify the long-chain amino acid interaction, the structure prediction error caused by the missing position information is reduced, and the prediction precision is improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

System and method for detecting real-time anomalies in transmission path of resources within an entity network via DNA computing

PendingUS20260133864A1Non-redundant fault processingDNA computingEngineering
Embodiments of the present invention provide a system for detecting real-time anomalies in transmission path of resources within an entity network via DNA computing. The system is configured for determining that an application received a resource from a source system to be transmitted to an end application via an entity network, extracting metadata from the resource, wherein the metadata is in a binary format, converting the metadata in the binary format to a DNA format, generating predicted molecular sequence associated with predicted transmission path of the resource within the entity network based on the DNA format of the metadata, monitoring real-time transmission path of the resource, generating real-time molecular sequence associated with the real-time transmission path of the resource, comparing the predicted molecular sequence with the real-time molecular sequence to determine an anomaly, and transmitting alerts associated with the anomaly.
Owner:BANK OF AMERICA CORP

Primer and method for identifying sex of grassland caterpillar

The invention discloses a molecular sequence and a primer for sex identification of grassland caterpillars and application of the molecular sequence and the primer. The nucleotide sequences of the molecular marker are as shown in SEQ ID NO: 1 and SEQ ID NO: 2; wherein a nucleotide sequence as shown in SEQ ID NO: 1 is female, and a nucleotide sequence as shown in SEQ ID NO: 2 is male. The invention also provides a primer group for amplifying the molecular fragment, and the nucleotide sequences of the primer group are shown as SEQ ID NO: 3 and SEQ ID NO: 4. Wherein the size of a strip capable of being amplified by a female individual is 484 bp, and the size of a strip capable of being amplified by a male individual is 199 bp. According to the molecular marker and the primer group provided by the invention, sex identification can be carried out on grassland caterpillars, larvae, incomplete pupae and cocoons with unknown sex, the accuracy is high, the speed is high, and a reliable method is provided for monitoring the field sex ratio of the grassland caterpillars.
Owner:LANZHOU UNIV

Application of rice zos2-02 gene in regulating salt tolerance

The purpose of this invention is to disclose a rice salt tolerance-related gene ZOS2-02, its encoded protein, and its applications. The gene ZOS2-02 is a DNA molecule as described in either 1) or 2) below: 1) a DNA molecule with a genomic sequence as shown in SEQ ID NO.1; 2) a DNA molecule with a CDS sequence as shown in SEQ ID NO.2; 3) a DNA molecule that hybridizes to the DNA sequence defined in 1) or 2) under stringent conditions and encodes the protein. The genetic engineering application of the gene ZOS2-02 provided by this invention in regulating rice salt tolerance specifically involves knocking out the aforementioned gene ZOS2-02 to improve rice salt tolerance.
Owner:NANJING AGRICULTURAL UNIVERSITY

Hybrid variant calling

PCT designated stageWO2026043987A1BiostatisticsProteomicsGeneticsVariome
A computer-implemented method for identifying a genomic variant is provided. The method includes obtaining one or more reference molecular sequences and sequencing data pertaining to a biological sample. One or more candidate variant positions are determined from the sequence reads for the biological sample. A classifier model is applied to each candidate variant position for selecting between a haplotype-aware or haplotype-agnostic variant analysis respectively. The classifier model is trained from haplotype structure and / or sequence reads identified with germline variants in a plurality of regions. Based on the application of the classifier model, respectively applying a haplotype-aware or haplotype-agnostic variant analysis to generate a variant identification for each candidate variant position.
Owner:ROCHE SEQUENCING SOLUTIONS INC

Nucleic acid detection compositions, kits and methods based on template and primer role switching

The present application relates to a kind of for nucleic acid amplification, composition or kit and method of capture oligonucleotide.It specifically, the capture oligonucleotide includes successively from 5' end to 3' end: first universal sequence (U2a), folding sequence (1s), second universal sequence (U1a) and binding capture sequence (2a);Wherein, (1) the folding sequence is at least partially identical with the 5' end sequence of target molecule;(2) the binding capture sequence is complementary with the 3' end sequence of target molecule;(3) the capture oligonucleotide further includes nucleic acid extension blocking modification located in the 3' end of binding capture sequence;And (4) the capture oligonucleotide further includes nucleic acid extension blocking modification between folding sequence and second universal sequence, the universal sequence is irrelevant with the sequence of target molecule.The product and method of the present application can realize specific, multiplex nucleic acid detection with high sensitivity.
Owner:SHANGHAI SCI-TECH INNO CENTER FOR INFECTION & IMMUNITY

A drug interaction prediction method and system based on graph attention sampling

The application discloses a drug interaction prediction method based on graph attention sampling, which firstly acquires drug pair data to be predicted, pre-processes the drug pair data to be predicted, obtains pre-processed drug pair data and drug molecular sequences, converts the pre-processed drug molecular sequences into molecular structure graphs, atomic vector matrices and molecular-atomic relationship lists through a molecular sequence conversion model, constructs an atomic adjacency matrix, a drug interaction network and a drug adjacency matrix, inputs the atomic vector matrix, the atomic adjacency matrix, the molecular-atomic relationship list, the drug interaction network and the drug adjacency matrix into a pre-trained drug interaction prediction model, and obtains a final drug interaction prediction result. The application can solve the technical problem that the existing drug interaction prediction method based on graph representation learning cannot consider both the drug molecular structure characteristics and the relationship information contained in the interaction network.
Owner:HUNAN UNIV

A soybean salt tolerance-related gene GmCIPK12, its encoded protein, and its applications.

The purpose of this invention is to disclose a soybean salt tolerance-related gene GmCIPK12, its encoded protein, and its applications. The gene GmCIPK12 is an DNA molecule as described in 1), 2), or 3) below: 1) a DNA molecule with a genomic sequence as shown in SEQ ID NO.1; 2) a DNA molecule with a CDS sequence as shown in SEQ ID NO.2; 3) a DNA molecule that hybridizes to the DNA sequence defined in 1) or 2) under stringent conditions and encodes the aforementioned DNA molecule. The genetic engineering application of the gene GmCIPK12 provided by this invention in regulating soybean salt tolerance specifically involves overexpressing the aforementioned gene GmCIPK12 to improve soybean salt tolerance.
Owner:NANJING AGRICULTURAL UNIVERSITY

Preparation method, device, equipment and product of multi-target targeted capture probe

The embodiment of the present application provides a method, device, equipment and product for preparing a multi-target targeted capture probe, which relates to the field of medical biotechnology. On the basis of obtaining the full-length molecular sequence and molecular identifier of the target molecule, M original probe sequences are generated; N connecting arm sequences are obtained; an initial probe sequence is generated based on the M original probe sequences and the N connecting arm sequences; a feature analysis is performed on the initial probe sequence to generate a probe sequence feature; based on the probe sequence feature, at least one base in the initial probe sequence is replaced to generate a target probe sequence; based on the probe application information and the target probe sequence, a multi-target targeted capture probe is prepared; the target molecule is captured by the multi-target targeted capture probe, which solves the problem of spatial arrangement conflict and kinetic competition between the mixed single-target targeted capture probes, that is, solves the problem of low capture efficiency of target molecules caused by the solution of the prior art.
Owner:SHANGHAI JINFUKANG PHARMACEUTICAL ENGINEERING TECHNOLOGY CO LTD

Method for analyzing molecules of interest, method for automatically selecting electrical signals related to a set of known barcode molecules, signal processing unit and computer program for executing the method

PCT designated stageWO2025224313A1BiostatisticsSequence analysisTarget signalBarcode
A method for analyzing molecules of interest by use of an adaptor attached to the molecule of interest is described, wherein said adaptor comprises a barcode molecule, wherein the barcode molecule is one instance of a set of defined sequences of molecules. The method comprises the steps of: - partitioning the electrical signal derived by electrically measuring at least of the part of a molecule into segments, - temporarily aligning segments of the partitioned electrical signal related to an unknown barcode molecule derived by electrically measuring at least of the part of the molecule of interest comprising the barcode molecule and a set of respective signals for known barcode molecules stored in a data memory, and - predicting the instance of barcode molecule from the similarity between the temporarily aligned segments of the electrical signal and the respective signals for known barcode molecules, wherein the electrical signal is temporarily aligned by compressing or stretching the temporal axis and the instance of barcode molecule is predicted automatically in a data processing unit. The method is also directed to a process of barcode design based on the similarity of temporarily aligned segments of electrical signals related to the candidate barcode molecules and a set of designed target signals.
Owner:HELMHOLTZ ZENTRUM FUER INFEKTIONSFORSCHUNG GMBH +2

Biomacromolecule sequence structure prediction method

The invention discloses a biomacromolecule sequence structure prediction method, and relates to the technical field of computers. The attention weights corresponding to different attention types of the processed target information are broadcasted in different broadcast range modes by using a multi-scale attention broadcast mechanism, and the attention weights in different time steps in the broadcast range are repeatedly used, so that the same attention weight is adopted in the steps of the broadcast range, the calculation amount is reduced, and the efficiency is improved. The efficiency performance of the model is improved, and the perception ability of the model to amino acid position information in a protein sequence can be effectively improved, so that the model can fully consider the position relationship between amino acids when calculating the attention weight, and the model can be helped to better identify the interaction of long-chain amino acids; the structure prediction error caused by the missing of the position information is reduced, and the prediction precision is improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Nucleic acid detection composition, kit and method based on template and primer role interchange

The present invention relates to a capture oligonucleotide, composition or kit and method for nucleic acid amplification. Specifically, the capture oligonucleotide comprises, from the 5' end to the 3' end, a first universal sequence (U2a), a folding sequence (1s), a second universal sequence (U1a) and a binding capture sequence (2a); wherein, (1) the folding sequence is at least partially identical to the 5' end sequence of the target molecule; (2) the binding capture sequence is complementary to the 3' end sequence of the target molecule; (3) the capture oligonucleotide further comprises a nucleic acid extension blocking modification located at the 3' end of the binding capture sequence; and (4) the capture oligonucleotide further comprises a nucleic acid extension blocking modification located between the folding sequence and the second universal sequence, and the universal sequence is unrelated to the target molecule sequence. The product and method of the present invention can achieve multiple nucleic acid detection with good specificity and high sensitivity.
Owner:SHANGHAI SCI-TECH INNO CENTER FOR INFECTION & IMMUNITY

A soybean salt-tolerance-related gene GmCNGC29, its encoded protein, and application

The present invention aims to disclose a soybean salt-tolerance-related gene GmCNGC29, its encoded protein, and its application. The gene GmCNGC29 is a DNA molecule as described in 1) or 2) or 3) below: 1) a DNA molecule whose genomic sequence is shown in SEQ ID NO.1; 2) a DNA molecule whose CDS sequence is shown in SEQ ID NO.2; 3) a DNA molecule that hybridizes with the DNA sequence defined in 1) or 2) under stringent conditions and encodes the DNA molecule. The present invention provides a genetic engineering application of the gene GmCNGC29 in regulating soybean salt tolerance, specifically overexpressing the aforementioned gene GmCNGC29 to improve soybean salt tolerance.
Owner:NANJING AGRICULTURAL UNIVERSITY

Primer group for detecting PARMS molecular marker of rice plant height gene OsPH4 and application of primer group

The invention belongs to the field of rice breeding, and particularly discloses a primer group for detecting a PARMS molecular marker of a rice plant height gene OsPH4 and application of the primer group, dwarf type OsPH4 gene detection is carried out on a rice material through a molecular sequence of the rice plant height gene OsPH4 by applying a PARMS technology, and the rice plant height gene OsPH4 is obtained. The rice dwarf type OsPH4 gene can be rapidly and accurately detected in different germplasm resources such as indica rice and japonica rice. According to the method, tedious procedures such as enzyme digestion, electrophoresis and sequencing are not needed in the detection process, pollution of PCR product aerosol and use of toxic substances such as EB are reduced, meanwhile, foreground selection and background selection can be conducted in the early stage of molecular marker assisted breeding, the background recovery rate is increased, the scale of breeding populations is reduced, the breeding process is accelerated, and the breeding efficiency is improved. The efficient and environment-friendly application of the OsPH4 gene in commercial molecular breeding of rice is facilitated.
Owner:LIAONING RICE RES INST

Scientific and technological achievement dynamic valuation method based on multi-source heterogeneous data fusion

The invention discloses a multi-source heterogeneous data fusion scientific and technological achievement dynamic valuation method, and relates to the technical field of scientific research evaluation, and the method comprises the following steps: S1, data collection: obtaining subjective cognition data of experts on scientific and technological achievements through brain-computer interface equipment, technical application scene data and market terminal feedback data generated by Internet of Things equipment are collected in real time in combination with an edge computing node, and distributed preprocessing and encrypted transmission of the data are achieved through federal edge learning; and S2, data cleaning and feature extraction: encoding data into a molecular sequence, and performing data de-duplication and abnormal value detection by using DNA calculation parallel processing characteristics. Through the combination of the brain-computer interface equipment, the edge computing node and federal edge learning, a real-time acquisition and preprocessing system of multi-modal heterogeneous data is constructed, and expert subjective cognitive data, technical application scene data of Internet of Things equipment and market terminal feedback data can be effectively integrated.
Owner:SICHUAN YEXIN TECH SERVICE GRP CO LTD

Protein drug response prediction method based on em fusion learning

PendingCN122455077AData setProtein
The present application relates to a protein drug reaction prediction method based on EM fusion learning, comprising: 1) constructing a data set containing protein nodes, drug nodes and intermediate nodes into a text attribute graph; 2) embedding the protein amino acid sequence and the drug molecule SMILES sequence to obtain corresponding sequence representations; 3) constructing a transformer-based protein-drug reaction prediction model; 4) constructing a GNN-based protein-drug reaction prediction model; 5) using a variational EM framework to alternately update the above two models to obtain an EM fusion protein-drug reaction prediction model; 6) inputting the to-be-tested protein amino acid sequence and the to-be-tested drug SMILES sequence into the above prediction model to output the reaction prediction result of the to-be-tested protein-drug pair. The present application learns the topological information in the protein-drug network to realize protein-drug reaction prediction.
Owner:LIAONING UNIVERSITY

Virus and disease association prediction model, construction method of prediction model and prediction method

The invention relates to the technical field of interactive prediction, in particular to a virus and disease associated prediction model, a construction method of the prediction model and a prediction method. Comprising the following steps: S1, collecting and processing virus and disease data; s2, performing feature extraction of the independent prediction model A; s3, performing feature extraction of the independent prediction model B; s4, based on five-fold cross validation and grid search, determining an optimal machine learning algorithm, and constructing an independent prediction model A and an independent prediction model B; and S5, based on a bagging integration strategy, integrating the independent prediction model A and the independent prediction model B, and generating a virus and disease associated prediction model. According to the invention, the VDP provides a new thought of multi-modal biological information integration; meanwhile, molecular sequence information of viruses, semantic structure information of diseases and GO function annotation information of the two parties are considered, more comprehensive feature representation is achieved, and the problems that an existing method is single in information source and insufficient in prediction accuracy are solved.
Owner:HUNAN UNIV

Magnetic bead enrichment SERS (Surface Enhanced Raman Scattering) probe sensing platform based on lung cancer marker protein-DNA (Deoxyribonucleic Acid) signal conversion as well as preparation method and application thereof

The invention discloses a magnetic bead enrichment SERS (Surface Enhanced Raman Scattering) probe sensing platform based on lung cancer marker protein-DNA (Deoxyribose Nucleic Acid) signal conversion as well as a construction method and application of the magnetic bead enrichment SERS probe sensing platform. The target protein aptamer is an aptamer specifically bound with a carcino-embryonic antigen (CEA) protein; the SERS probe comprises gold nanoparticles, and the surfaces of the gold nanoparticles are covalently connected with a Raman reporter molecule sequence and capture DNA (Deoxyribose Nucleic Acid); the SERS probe is hybridized with the target protein aptamer on the first magnetic carrier compound by capturing DNA to form a DNA double-chain structure; the second magnetic carrier compound comprises a second magnetic carrier fixed with enriched DNA, and the enriched DNA sequence and the captured DNA are completely complementary. According to the invention, ultrasensitive and high-specificity quantitative detection of lung cancer marker protein (such as CEA) is realized.
Owner:NANJING UNIV

Application of rice oszos1-18 gene in regulating salt tolerance

ActiveCN119876264BPlant peptidesFermentationSalt resistanceA-DNA
The application discloses a salt-tolerant related gene ZOS1-18 of rice, and a coding protein and application thereof. The gene ZOS1-18 is a DNA molecule as described in 1) or 2) or 3) below: 1) a DNA molecule with a genomic sequence as shown in SEQ ID NO. 1; 2) a DNA molecule with a CDS sequence as shown in SEQ ID NO. 2; and 3) a DNA molecule hybridized with the DNA sequence defined in 1) or 2) under stringent conditions and encoding the protein. The application provides a genetic engineering application of the gene ZOS1-18 in regulating the salt tolerance of rice, specifically, knocking out the aforementioned gene ZOS1-18 to improve the salt sensitivity of rice, and overexpressing the aforementioned gene ZOS1-18 to improve the salt tolerance of rice.
Owner:NANJING AGRICULTURAL UNIVERSITY

Biomolecular sequence searching method and apparatus, device, and storage medium

PCT designated stageWO2026114095A1BiostatisticsSequence analysisAlgorithmSequence search
The present disclosure provides a biomolecular sequence searching method and apparatus, a device, and a storage medium. The method comprises: acquiring an unknown biomolecular sequence; inputting the unknown biomolecular sequence into a sequence encoding model for sequence encoding, to obtain an unknown biomolecular sequence representation; performing similarity searching on different candidate vector sets in a biomolecular sequence vector database by means of the unknown biomolecular sequence representation, wherein candidate biomolecular sequences of candidate vectors in the candidate vector sets have a same sequence length; on the basis of the similarity searching result, searching each candidate vector set for a preset number of known biomolecular sequence representations; and screening known biomolecular sequences corresponding to the plurality of known biomolecular sequence representations for homologous biomolecular sequences corresponding to the unknown biomolecular sequence.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Consensus calling with a neural network

A method and systems for determining a consensus molecular sequence from sequence data of a sample, the sequence data including base calls and quality scores (q-scores). One or more feature vectors are generated and input into a convolutional neural network (CNN) model, which outputs a consensus call sequence. In order to generate the feature vectors, sequence reads are aligned and clustered. For each cluster, one or more feature vectors is then generated by either concatenating base calls and quality scores from the plurality of sequence reads in the cluster, or calculating aggregation statistics for the plurality of sequence reads and concatenating the aggregation statistics and corresponding quality metrics for the plurality of sequence reads in the cluster.
Owner:ROCHE SEQUENCING SOLUTIONS INC

MHC-II type molecular antigen presentation prediction method, electronic equipment and program product

The invention discloses an MHC-II type molecular antigen presentation prediction method. The method comprises the following steps: inputting an MHC-II molecular sequence, an oligopeptide sequence and a context sequence thereof; the MHC-II molecule sequence and the oligopeptide sequence are jointly input into a forward BICL encoder containing a forward BICL convolution kernel and a reverse BICL encoder containing a reverse BICL convolution kernel so as to capture the interaction characteristics of the MHC-II molecule and the oligopeptide under forward binding and reverse binding, and the interaction characteristics under the actual optimal binding direction are determined through score comparison. The method is used for predicting the presentation of the MHC-II molecular antigen. Each of the forward BICL encoder and the reverse BICL encoder comprises a plurality of layers of sensors and a maximum pooling layer, and corresponding interaction representation vectors are obtained after interaction matrixes output by the forward BICL encoder and the reverse BICL encoder are processed.
Owner:FUDAN UNIVERSITY

SgRNA molecule, bovine DDX58 gene single base mutation system, construction method and application

The invention relates to the technical field of animal genetic breeding and genetic engineering, in particular to an sgRNA molecule, a cattle DDX58 gene single base mutation system, a construction method and application. And the molecular sequence of the sgRNA is as shown in SEQ ID NO. 3. According to the method, sgRNA shown in SEQ ID NO.3 is subjected to specific chemical modification, CBE mRNA obtained through in-vitro transcription is jointly delivered to a cattle fertilized egg, the cattle congenital immune key gene DDX58 is precisely edited, finally, the gene edited disease-resistant cattle is obtained, the editing efficiency is high, and the off-target risk is reduced. According to the method, the toxicity problem related to DNA delivery does not exist, a simple, convenient and effective way is provided for simplifying the genome editing development process, the application value of chemical synthesis of sgRNA in genome editing is proved, and wide application of the CRISPR-Cas technology in the fields of biotechnology and treatment is expected to be accelerated.
Owner:NORTHWEST A & F UNIV