Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

19 results about "Multiple sequence alignment" patented technology

A multiple sequence alignment (MSA) is a sequence alignment of three or more biological sequences, generally protein, DNA, or RNA. In many cases, the input set of query sequences are assumed to have an evolutionary relationship by which they share a linkage and are descended from a common ancestor. From the resulting MSA, sequence homology can be inferred and phylogenetic analysis can be conducted to assess the sequences' shared evolutionary origins. Visual depictions of the alignment as in the image at right illustrate mutation events such as point mutations (single amino acid or nucleotide changes) that appear as differing characters in a single alignment column, and insertion or deletion mutations (indels or gaps) that appear as hyphens in one or more of the sequences in the alignment. Multiple sequence alignment is often used to assess sequence conservation of protein domains, tertiary and secondary structures, and even individual amino acids or nucleotides.

A method for constructing a multiplex PCR reaction system for detecting and identifying pycnospora and its application

ActiveCN116179735BOptimizing Multiplex PCR Reaction ConditionsMultiplexGene cluster
The application discloses a construction method of a multiplex PCR reaction system for detecting and identifying Podosphaera species and application thereof. The method comprises the following steps: obtaining whole genome sequences of multiple Podosphaera species, and performing multiple sequence alignment to obtain specific genes and specific gene clusters of each species; taking the genes in the specific gene clusters of each species as templates, designing multiple upstream primers and downstream primers according to a conventional primer design method to generate primer pairs, and excluding unreasonable primer pairs and primer pairs with low sensitivity and specificity; according to the size of the amplification products, the primers are freely combined and matched to generate four pairs of mixed multiplex primers, and the multiplex PCR reaction conditions are optimized to obtain a multiplex PCR system capable of specifically detecting Podosphaera and simultaneously identifying single and multiple species. The multiplex PCR method for detecting and identifying Podosphaera provided by the application can quickly and accurately complete the detection of Podosphaera and identify the species identity of Podosphaera.
Owner:NANJING AGRICULTURAL UNIVERSITY

An indel molecular marker for identifying old crow petal of wanyu and application thereof

The present application relates to the technical field of molecular marker, in particular to Indel molecular marker of Wan-Yu old crow petal and application thereof. The Indel molecular marker of the present application is shown as SEQ ID NO. 1-3. The Indel molecular marker can be used to identify Wan-Yu old crow petal, and the identification can be carried out by electrophoresis, multiple sequence alignment or reads alignment of high-throughput sequencing.
Owner:YIHU BIOTECHNOLOGY (ANHUI) CO LTD

Targeting superantigen fusion protein based on improved se(3)-transformer and implementation method

PendingCN122266440AAchieve collaborative structure optimizationImprove targetingMicroorganism based processesBiostatisticsPattern recognitionAntigen epitope
The application discloses a targeting superantigen fusion protein based on an improved SE(3)-Transformer and an implementation method. In an offline stage, amino acid sequences are first converted into one-hot encoding or language model embedding (such as ESM-2), and are spliced with multiple sequence alignment (MSA) features for geometric initialization. A neural network (improved SE(3)-Transformer) combined with multiple sequence alignment (MSA) and an attention mechanism is constructed to predict the coordinates of C alpha, C, N and O atoms for main chain prediction, and the neural network is trained through a gradient descent method based on a physical heuristic potential item. In a verification stage, the improved SE(3)-Transformer after training is used to generate a predicted structure, conformational stability is verified through a simplified force field, and fine tuning is performed based on a confidence score. The application can accurately predict the structure of a target antigen epitope and an antibody variable region, optimize a superantigen functional domain in combination with a graph neural network, and dynamically design a flexible connecting peptide to realize modular fusion.
Owner:SHANGHAI JIAOTONG UNIV

Protein fitness prediction method, device, terminal and storage medium

ActiveCN121393550BSolve Application Bottleneckseffective nonadditive effectBiostatisticsBiological modelsWild typeProtein engineering
The application relates to the technical field of protein engineering and bioinformatics, and specifically provides a protein fitness prediction method and device, a terminal and a storage medium. The method comprises the following steps: extracting wild-type sequence-level representation and mutant sequence-level representation by using a pre-trained protein language model, and calculating a representation difference vector between the two; deducing coevolution coupling information from multiple sequence alignment information, and constructing pairwise evolutionary constraint information reflecting spatial proximity relationship; then modeling the interaction between the representation difference vector and the pairwise evolutionary constraint information through an attention mechanism, so that the pairwise coevolution information guides the propagation and weighting of the representation difference vector between residues to generate enhanced representation; and inputting the enhanced representation into a downstream prediction head to regress a scalar value as a fitness prediction result. The application models the interaction between the sequence-level difference vector and the pairwise evolutionary constraint, so that the non-additive effect between mutations can be accurately captured without relying on any experimental structure.
Owner:XIDIAN UNIV

Primers, kits, methods, and applications for detecting pathogenic *Pseudomonas proteus* strains based on the *oprL* gene.

This invention discloses primers, kits, methods, and applications for detecting pathogenic *Pseudomonas proteus* strains based on the *oprL* gene, belonging to the field of microbial detection technology. This invention provides specific primers for detecting pathogenic *P. proteus* strains based on the *oprL* gene. The primers are designed based on the differential patterns identified after multiple sequence alignment, specifically by analyzing the results of multiple sequence alignments of all *P. proteus* *oprL* gene sequences in the NCBI database. The nucleotide sequences of the primers are shown in SEQ ID No. 1, SEQ ID No. 2, and SEQ ID No. 3. The *oprL* gene of *P. proteus* is amplified by PCR, and the PCR amplification products are then subjected to electrophoresis. The pathogenicity of *P. proteus* can be effectively distinguished based on the electrophoresis image, enabling rapid detection of *P. proteus* infecting large yellow croaker. The primers have good specificity, and the detection method is simple and intuitive.
Owner:FUJIAN AGRI & FORESTRY UNIV

Protein sequence-structure co-generation method, system, device, and storage medium

This application provides a method, system, device, and storage medium for co-generating protein sequence-structure, relating to the field of bioinformatics. The generation method includes: acquiring multiple sequence alignment data and encoding it to obtain an explicit evolutionary prior representation; initializing the current sequence state and current structural state; performing at least one round of co-iterative iterative generation based on the explicit evolutionary prior representation; generating the final three-dimensional structure and outputting the protein sequence and the final three-dimensional structure. The method provided in this application directly guides the generation trajectory by extracting multiple sequence alignment data as an explicit evolutionary prior, ensuring the natural feasibility of the molecule. Simultaneously, it performs alternating and interwoven co-iterative updates of the sequence and structural states, breaking the limitations of sequential fragmentation and achieving bidirectional communication between spatial conformation and amino acid prediction. This results in a high degree of consistency between the final sequence and structure, significantly improving the functional hit rate and experimental success rate of the protein.
Owner:ACADEMY OF MILITARY MEDICAL SCIENCES

Screening method for rational mutation sites of heat-sensitive udg based on strain culture temperature

PendingCN122337320ACold adaptedPrimary screening
This invention discloses a method for screening thermosensitive UDG rational mutation sites based on strain culture temperature. The method first divides strains into heat-tolerant and cold-adapted groups according to temperature thresholds based on the culture temperature of the strain preservation database, constructing corresponding UDG sequence sets and three-dimensional structure sets. Then, a first candidate mutation site set is generated through multiple sequence alignment and amino acid distribution difference analysis. A second candidate mutation site set is generated through rigid body alignment and functional protection rules under a unified coordinate system. Finally, multiple-source indicators such as structural dispersion, contact and salt bridge networks, and kinetic coupling are calculated for the second candidate mutation site set. Combined with effect-safety dual-dimensional scoring, spatial clustering, and cross-validation of the first candidate mutation site set, a preferred mutation site set is output. This invention achieves a systematic transformation from temperature ecotags to engineered sites, improving the reproducibility and interpretability of site screening without requiring a large-scale mutation library.
Owner:PULUOMAIGE BIOLOGICAL PRODS SHANGHAI

Primer group, kit and detection method for detecting xanthomonas campestris pv. Campestris

The invention discloses a primer group for detecting xanthomonas campestris pv.campestris. The primer group consists of an upstream primer and a downstream primer, wherein the nucleotide sequence of the upstream primer is as shown in SEQ ID NO.1, and the nucleotide sequence of the downstream primer is as shown in SEQ ID NO.2. The invention further discloses a kit for detecting the xanthomonas campestris pv.campestris. According to the present invention, multiple sequence comparison analysis is performed on the whole genome sequences of the Xanthomonas campestris pv. Campestris pathogen and other related bacteria, such that a pair of primers is designed according to the genome specificity region of the Xanthomonas campestris pv. Campestris pv. Campestris pv. Campestris pv. Campestris pv. Campestris pv. Campestris pv. Campestris pv. Campestris; the method is used for high-sensitivity rapid molecular detection of plants with the brassica oleracea, and a rapid, simple, high-specificity and high-sensitivity monitoring technology system for the brassica oleracea is established.
Owner:JIANGSU ACAD OF AGRI SCI

Construction and application of polyphosphate kinase bsppk mutant and its producing strain

ActiveCN122012453BArginineNucleotide
The application discloses a polyphosphate kinase BsPPK mutant and construction and application of a producing strain thereof, and belongs to the technical field of genetic engineering. The application carries out site-directed mutation on polyphosphate kinase BsPPK from the genus Bredia through homologous modeling, molecular docking and multiple sequence alignment, mutates lysine at the 92th position of the wild type BsPPK into alanine, threonine at the 95th position into alanine, and serine at the 205th position into arginine, and obtains a combined mutant BsPPK-K92A / T95A / S205R. Enzyme activity determination results show that the specific enzyme activity of the mutant is increased by 324.4% compared with the wild type, and the catalytic efficiency on AMP is significantly improved. When the mutant is applied to synthesis of UDP-Gal and derivative products thereof, only 15 mM AMP is needed to achieve a similar yield obtained by using 30 mM AMP for the wild type, and the nucleotide consumption is reduced by 50%.
Owner:OCEAN UNIV OF CHINA

Mutant of protein glutaminase with improved heat resistance as well as construction method and application of mutant

The invention discloses a heat-resistant mutant of protein glutaminase (PG enzyme) as well as a construction method and application of the heat-resistant mutant, and belongs to the technical field of protein engineering. The amino acid sequence of the PG enzyme is obtained by performing multi-sequence comparison on the PG enzyme from C. proteolyticum YF810 and other PG enzymes which are from different strains and have improved heat resistance, screening heat-resistant single-point mutations, and performing combined mutation on two or more of all beneficial single-point mutations to obtain a combined mutant of the PG enzyme. According to the five-combination mutant of the PG enzyme with the heat resistance improved to the maximum degree, the T5010min value is improved by 12.80 DEG C, the t1 / 260 DEG C value is improved by 55.13 times, and compared with wild type PG enzyme, the enzyme activity is not reduced. Therefore, compared with the wild type PG enzyme, the mutant of the PG enzyme provided by the invention has better heat resistance and better industrial application prospect.
Owner:EAST CHINA NORMAL UNIV

Protein design methods, apparatuses, devices, and media

The present disclosure provides a protein design method, device, equipment and medium, relates to the field of artificial intelligence, in particular to the technical field of deep learning, biological computing and large language model. The generation method comprises the following steps: constructing a plurality of candidate proteins, each of which comprises a first chain of an original protein and a non-natural sequence constructed based on a second chain of the original protein; retrieving a first multiple sequence alignment of the first chain and a second multiple sequence alignment of the non-natural sequence; matching the first multiple sequence alignment and the second multiple sequence alignment to obtain a cross-chain homologous sequence by using a pre-trained initial protein language model; predicting the structure and a first score of the candidate protein by using a protein structure prediction model; determining a reward value based on the first score and performing reinforcement learning training on the initial protein language model; determining a second score of each of the plurality of candidate proteins by using the trained target protein language model and the protein structure prediction model, so as to obtain a protein design result.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

A method of co-engineering a multimeric fluorinase

The application relates to the technical field of bioengineering, and particularly relates to a synergistic engineering modification method of a multimeric fluorinase, which comprises step-by-step synergistic modification of an N-terminal sequence, a distal region and a core region of a target enzyme. The N-terminal sequence is optimized through ancestral sequence reconstruction or hydrophobicity analysis, the distal mutation site is screened by using multiple sequence alignment and free energy calculation, and the catalytic activity of the core region is improved by combining conservativeness analysis and saturation mutation, so as to finally construct a multi-site combined mutant. The application can significantly improve the overall structural stability and local catalytic efficiency of the multimeric fluorinase, and meanwhile, the thermal stability is enhanced, a universal methodology is provided for modification of other highly conservative multimeric enzymes, and has a wide application prospect.
Owner:BEIJING UNIV OF CHEM TECH

Polyphosphate kinase BsPPK mutant as well as construction and application of producing strain of polyphosphate kinase BsPPK mutant

The invention discloses a polyphosphate kinase BsPPK mutant and construction and application of a producing strain thereof, and belongs to the technical field of genetic engineering. According to the invention, site-directed mutagenesis is carried out on polyphosphate kinase BsPPK from Breyd bacteria through methods of homologous modeling, molecular docking, multiple sequence alignment and the like, lysine at the 92th site of wild BsPPK is mutated into alanine, threonine at the 95th site of wild BsPPK is mutated into alanine, serine at the 205th site of wild BsPPK is mutated into arginine, and the combined mutant BsPPK-K92A / T95A / S205R is obtained. An enzyme activity determination result shows that the specific enzyme activity of the mutant is improved by 324.4% compared with that of a wild type, and the catalytic efficiency on AMP is remarkably improved. When the mutant is applied to synthesis of UDP-Gal and derivative products thereof, only 15 mM of AMP is needed, the similar yield which can be obtained by 30 mM of AMP of a wild type can be achieved, and the use amount of nucleotide is reduced by 50%.
Owner:OCEAN UNIV OF CHINA

Guide editing efficiency prediction method and system based on deep learning

The invention relates to the technical field of nucleic acid sequence analysis, in particular to a guide editing efficiency prediction method and system based on deep learning. The method comprises the following steps: acquiring a plurality of key nucleic acid sequences in a guide editing process, and constructing a multi-sequence alignment matrix as input; generating a latent matrix fusing the sequence features and the position information by using a word embedding and rotating position coding technology; respectively capturing inter-sequence association and intra-sequence context dependence through an axial attention encoder; a deformable convolution and pooling alternating structure is adopted to carry out spatial feature refining; and finally, outputting an editing efficiency prediction value through attention weighting and multi-layer perceptron regression. According to the method, the defects that a traditional method depends on artificial features and neglects interaction among sequences are overcome, end-to-end high-precision efficiency prediction is achieved, and a reliable calculation tool is provided for optimization design of a guiding editing system.
Owner:CHONGQING MEDICAL UNIVERSITY

Soft decision information decoding method and encoding method for DNA storage

The embodiment of the present disclosure discloses a soft decision information decoding method and encoding method for DNA storage. The soft decision information decoding method for DNA storage comprises: clustering sequencing sequences in obtained sequencing data; obtaining a consensus sequence for each cluster to obtain a plurality of consensus sequences, taking the support degree of multiple sequence alignment recorded by each base in the consensus sequence as the quality value of each base on the consensus sequence; arranging the plurality of consensus sequences according to sequence indexes to obtain a decoding matrix block; and decoding the decoding matrix block. By obtaining a consensus sequence for the sequencing sequence, most random errors are removed, and then the support degree of multiple sequence alignment is used to predict and sort the remaining errors, so that the sequencing information is fully mined to improve the prediction accuracy and reduce the calculation complexity of prediction. The soft decision decoding can greatly improve the error correction capability, so that the DNA storage can be applied to a larger data scale and has stronger fidelity.
Owner:AGRICULTURAL GENOMICS INSTITUTE AT SHENZHEN CHINESE ACADEMY OF AGRICULTURAL SCIENCES (SHENZHEN BRANCH GUANGDONG LABORATORY FOR LINGNAN MODERN AGRICULTURE)

Protein structure prediction

PendingUS20260100244A1BiostatisticsSequence analysisAlgorithmResidue coding
Disclosed herein are methods, systems, and apparatus, including computer programs encoded on computer storage media, for antibody structure prediction. In an example method, a target antibody sequence of a target antibody that includes a sequence of amino acids is received. The target antibody sequence is processed by an antibody language model (ALM) to obtain a residue encoding and an attention weight encoding without performing multiple sequence alignment (MSA), wherein the ALM is a protein language model trained from antibody sequences, and the ALM comprises a plurality of self-attention layers. The residue encoding and the attention weight encoding are transformed into a single representation and a pair representation that are input into a structure prediction model. A predicted structure of the target antibody is determined using the structure prediction model.
Owner:BIOMAP (BEIJING) INTELLIGENCE TECH LTD

A method and apparatus for predicting and analyzing RNA-binding residues

ActiveCN118866080BBiostatisticsProteomicsProtein DatabasesAlgorithm
A method and apparatus for predicting and analyzing RNA-binding residues are disclosed. The method constructs a GDRBind model using PyTorch and the DGL framework; inputs a protein sequence of a predetermined length into the GDRBind model and obtains the protein structure data of the protein sequence; generates a multiple sequence alignment (MSA) file using the protein sequence search software HHblits; collects several RNA-binding protein sequences from the UniProtKB protein database, performs clustering processing to obtain a pre-training dataset, and uses it to train a general protein language model to obtain an ESM-RBP representation model; obtains a first embedding matrix and a second embedding matrix through processing, and concatenates them to obtain a residue node feature representation matrix; calculates the edge features of all residue pairs in the protein sequence to obtain an edge set; and predicts the RNA-binding residues of the protein sequence using an equivariant graph neural network (EGNN) prediction model, outputting the predicted results. This invention can solve problems such as insufficient domain feature mining, low segmentation accuracy, and lack of interpretability.
Owner:HUNAN UNIV

Optical quantum computer-based protein structure prediction method, system and apparatus

The present invention relates to the technical field of protein structure prediction, and relates to an optical quantum computer-based protein structure prediction method, system and apparatus. The method comprises: 1) obtaining a multiple sequence alignment (MSA) matrix of a target sequence on the basis of target sequence alignment of a protein needing to be predicted; 2) encoding the obtained MSA matrix into a (0,1) matrix; 3) converting interaction of amino acids of the protein into an undirected graph model of (0,1) state nodes, and using a Boltzmann machine training mechanism and an optical quantum computer to perform training to obtain weight coefficients of edges connecting nodes in the undirected graph model; and 4) obtaining coefficients of interaction between different amino acids of the protein on the basis of the weight coefficients. The present invention improves the computation efficiency and solution result of protein structure prediction, and solves the problems in the prior art that it is difficult to perform computation for complex scheduling problems and accurate solutions cannot be obtained.
Owner:BEIJING QBOSON QUANTUM TECH CO LTD

A method for multi-domain protein assembly guided by flexible residues

A flexible residue-guided multi-domain protein assembly method, belonging to the fields of bioinformatics and computational intelligence, firstly utilizes a flexible residue prediction network to directionally decouple co-evolutionary information in multiple sequence alignment (MSA), thereby generating multiple distance maps that reflect conformational heterogeneity. Subsequently, these distance maps are introduced into a multi-objective optimization model as energy functions to constrain the conformational sampling process, guiding it to search for optimal solutions in a broader conformational space, ultimately obtaining a high-precision multi-domain protein structure model. This invention achieves accurate modeling of protein structural diversity by deeply mining the dynamic information implicit in the protein static database (PDB) and combining it with a large language model to analyze the multiple conformational features contained in the protein amino acid sequence. This method can significantly improve the prediction accuracy of multi-domain proteins, providing strong support for a deeper understanding of protein functional mechanisms.
Owner:ZHEJIANG UNIV OF TECH