Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

67 results about "Multiple sequence alignment" patented technology

A multiple sequence alignment (MSA) is a sequence alignment of three or more biological sequences, generally protein, DNA, or RNA. In many cases, the input set of query sequences are assumed to have an evolutionary relationship by which they share a linkage and are descended from a common ancestor. From the resulting MSA, sequence homology can be inferred and phylogenetic analysis can be conducted to assess the sequences' shared evolutionary origins. Visual depictions of the alignment as in the image at right illustrate mutation events such as point mutations (single amino acid or nucleotide changes) that appear as differing characters in a single alignment column, and insertion or deletion mutations (indels or gaps) that appear as hyphens in one or more of the sequences in the alignment. Multiple sequence alignment is often used to assess sequence conservation of protein domains, tertiary and secondary structures, and even individual amino acids or nucleotides.

Protein structure prediction

PendingCN120092293ABiostatisticsSequence analysisAlgorithmResidue coding
Disclosed herein are methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for antibody structure prediction. In one example method, a target antibody sequence for a target antibody comprising an amino acid sequence is received. The target antibody sequence is processed by an antibody language model (ALM) to obtain residue coding and attention weight coding without multiple sequence alignment (MSA), where the ALM is a protein language model trained according to the antibody sequence, and the ALM comprises a plurality of self-attention layers. The residue coding and attention weight coding are converted into a single representation and paired representations and input into a structure prediction model. And determining a prediction structure of the target antibody by using the structure prediction model.
Owner:BIOMAP (BEIJING) INTELLIGENCE TECH LTD

Method, device and equipment for identifying binding sites of protein and metal ions

The invention provides a protein and metal ion binding site identification method, device and equipment, and belongs to the field of protein detection.The method comprises the steps that feature extraction is conducted on known metal ion binding protein, and multiple sample evolution information features are obtained; the method comprises the following steps: for an unknown protein sequence, determining candidate distant homologous metal ion binding proteins through multi-sequence comparison and cosine similarity screening, and further constructing a training set; a composite framework in which a bidirectional long-short-term memory network and a full-connection neural network are connected in series is adopted, input features comprise evolutionary information features and physicochemical attribute features, and probability values of binding sites of different metal ions are output; training the composite framework through the training set to obtain a site prediction model; and predicting a binding site of an unknown protein sequence by using the prediction model. The stable and efficient prediction on the metal ion binding site is realized.
Owner:YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA

Ancestor nitrilase design based on multi-sequence alignment and machine learning and application of mutant

ActiveCN120636526AChemical property predictionBiostatisticsMutantEvolutionary landscape
The invention discloses an ancestor nitrilase design based on multi-sequence alignment and machine learning and application of a mutant, and belongs to the technical field of enzyme engineering. The invention designs an ancestor enzyme sequence-structure-molecular dynamics strategy so as to obtain information for eradicating ancestor enzyme in an evolutionary landscape, specifically, a nitrilase sequence is collected for evolutionary tree analysis to obtain the evolutionary landscape, an ancestor enzyme reconstruction algorithm is combined to obtain an amino acid sequence of the ancestor enzyme, and an ancestor enzyme primary structure sequence library is constructed; predicting the tertiary structures of all ancestor enzymes to obtain a structural library; and finally, carrying out molecular dynamics simulation on all ancestor enzyme tertiary structures to obtain a kinetic parameter library, and screening through specific kinetic parameters. Based on the strategy, ancestor nitrilase capable of tolerating 90 DEG C is obtained, the thermal stability of the ancestor nitrilase is evolved, and a series of dominant mutants are obtained.
Owner:JIANGNAN UNIV

Method and system for protein three-dimensional structure prediction

The application provides a protein three-dimensional structure prediction method and system, comprising the following steps: S1: obtaining a multiple sequence alignment matrix containing protein coevolution information and a protein template, and performing conversion of MSA sequence and residue pair coding to obtain MSA sequence coding and residue pair coding; S2: the MSA sequence coding and the residue pair coding are updated by a recurrent attention neural network to generate the latest MSA sequence coding and the residue pair coding; S3: based on the current residue pair coding, target sequence coding and starting main chain framework T i The current target sequence coding is updated by the invariant point attention neural network; and S4: the main chain framework is updated based on the current target sequence coding, and the torsion angle between amino acids is calculated, so that the protein main chain structure is obtained, and the side chain atom angle and the final three-dimensional structure are calculated through the residue network.
Owner:SHANGHAI TIANRANG NETWORK TECH CO LTD

Sequence diffusion method based on multi-sequence alignment and electronic equipment thereof

The invention discloses a sequence diffusion method based on multiple sequence alignment and electronic equipment thereof. According to the method, two-stage training is adopted to split different-stage training, so that the interpretability of the model is improved, and the training of the field sub-model is flexibly carried out. The method focuses on feature extraction of the MSA, uses a deeper attention layer to extract paired information hidden in the MSA, and can effectively capture biological correlation between sequences. According to the method, the sequence diffusion mode is utilized, sequence generation guided by evolutionary information is achieved, and high-accuracy and diversified sequences can be generated.
Owner:HANGZHOU LEVINTHAL BIOTECHNOLOGY CO LTD

A method for constructing a multiplex PCR reaction system for detecting and identifying pycnospora and its application

ActiveCN116179735BOptimizing Multiplex PCR Reaction ConditionsMultiplexGene cluster
The application discloses a construction method of a multiplex PCR reaction system for detecting and identifying Podosphaera species and application thereof. The method comprises the following steps: obtaining whole genome sequences of multiple Podosphaera species, and performing multiple sequence alignment to obtain specific genes and specific gene clusters of each species; taking the genes in the specific gene clusters of each species as templates, designing multiple upstream primers and downstream primers according to a conventional primer design method to generate primer pairs, and excluding unreasonable primer pairs and primer pairs with low sensitivity and specificity; according to the size of the amplification products, the primers are freely combined and matched to generate four pairs of mixed multiplex primers, and the multiplex PCR reaction conditions are optimized to obtain a multiplex PCR system capable of specifically detecting Podosphaera and simultaneously identifying single and multiple species. The multiplex PCR method for detecting and identifying Podosphaera provided by the application can quickly and accurately complete the detection of Podosphaera and identify the species identity of Podosphaera.
Owner:NANJING AGRICULTURAL UNIVERSITY

An indel molecular marker for identifying old crow petal of wanyu and application thereof

The present application relates to the technical field of molecular marker, in particular to Indel molecular marker of Wan-Yu old crow petal and application thereof. The Indel molecular marker of the present application is shown as SEQ ID NO. 1-3. The Indel molecular marker can be used to identify Wan-Yu old crow petal, and the identification can be carried out by electrophoresis, multiple sequence alignment or reads alignment of high-throughput sequencing.
Owner:YIHU BIOTECHNOLOGY (ANHUI) CO LTD

Blood protein characteristic polypeptide, detection kit and application of blood protein characteristic polypeptide in identifying adulteration and weight gain of hirudo nipponia

The invention belongs to the field of biological detection, and particularly relates to a blood protein characteristic polypeptide, a detection kit and application of the blood protein characteristic polypeptide to identification of adulteration and weight gain of hirudo nipponia. And the sequence of the blood protein marker polypeptide is LLGNVIVVVLAR. On the basis of constructing a species marker polypeptide identification method, a heterologous blood protein traceability analysis strategy is introduced, beta globin in blood of common livestock such as pigs, cattle, sheep, horses, donkeys and the like is subjected to multiple sequence alignment, and highly conservative species-independent marker polypeptide is screened out as a chemical marker; and a multi-reaction monitoring (MRM) non-standard quantitative peak area ratio judgment and detection method based on a liquid chromatography-triple quadrupole mass spectrometry technology is established, so that a scientific judgment basis is provided for the problem of blood doping and weight increment.
Owner:SHANDONG INST FOR FOOD & DRUG CONTROL +1

A HBV S protein-specific monoclonal antibody HBV-S-4H7 and its use in preparing a detection kit

The present invention relates to a HBV S protein-specific monoclonal antibody, HBV-S-4H7, and its use in the preparation of a detection kit. Based on the amino acid sequences of different HBV S proteins, multiple sequence alignment, and online epitope screening software, the present invention ultimately selects a preferred epitope peptide and immunizes mice to prepare a broad-spectrum monoclonal antibody. The antibody has good binding properties. Based on the role of the S protein in the virus, neutralization experiments confirm that the monoclonal antibody prepared by the present invention also has good virus neutralization effect, and has good application prospects.
Owner:SHAANXI ZHUOJIEMU BIOTECHNOLOGY CO LTD

Targeting superantigen fusion protein based on improved se(3)-transformer and implementation method

PendingCN122266440AAchieve collaborative structure optimizationImprove targetingMicroorganism based processesBiostatisticsPattern recognitionAntigen epitope
The application discloses a targeting superantigen fusion protein based on an improved SE(3)-Transformer and an implementation method. In an offline stage, amino acid sequences are first converted into one-hot encoding or language model embedding (such as ESM-2), and are spliced with multiple sequence alignment (MSA) features for geometric initialization. A neural network (improved SE(3)-Transformer) combined with multiple sequence alignment (MSA) and an attention mechanism is constructed to predict the coordinates of C alpha, C, N and O atoms for main chain prediction, and the neural network is trained through a gradient descent method based on a physical heuristic potential item. In a verification stage, the improved SE(3)-Transformer after training is used to generate a predicted structure, conformational stability is verified through a simplified force field, and fine tuning is performed based on a confidence score. The application can accurately predict the structure of a target antigen epitope and an antibody variable region, optimize a superantigen functional domain in combination with a graph neural network, and dynamically design a flexible connecting peptide to realize modular fusion.
Owner:SHANGHAI JIAOTONG UNIV

Primer group, kit and detection method for mink coronavirus detection

The invention discloses a primer group, a kit and a detection method for detecting mink coronavirus, and belongs to the technical field of virus detection. According to the invention, 1b fragment and N fragment genome sequences of mink coronavirus are selected as target regions, and genome sequence conserved regions are selected for designing two groups of detection primers through multiple sequence alignment. Specific reaction conditions of the two groups of detection primers are respectively optimized, and the sensitivity of the two groups of detection primers is verified through experiments. 135 high-throughput sequencing positive samples are selected, the two sets of detection primers in the method are used for sample detection, and the result is consistent with the high-throughput sequencing result. The RT-qPCR detection method based on SYBR-Green disclosed by the invention can be used for detecting and quantifying the mink coronavirus.
Owner:SHANDONG FIRST MEDICAL UNIV & SHANDONG ACADEMY OF MEDICAL SCI

Protein fitness prediction method, device, terminal and storage medium

ActiveCN121393550BSolve Application Bottleneckseffective nonadditive effectBiostatisticsBiological modelsWild typeProtein engineering
The application relates to the technical field of protein engineering and bioinformatics, and specifically provides a protein fitness prediction method and device, a terminal and a storage medium. The method comprises the following steps: extracting wild-type sequence-level representation and mutant sequence-level representation by using a pre-trained protein language model, and calculating a representation difference vector between the two; deducing coevolution coupling information from multiple sequence alignment information, and constructing pairwise evolutionary constraint information reflecting spatial proximity relationship; then modeling the interaction between the representation difference vector and the pairwise evolutionary constraint information through an attention mechanism, so that the pairwise coevolution information guides the propagation and weighting of the representation difference vector between residues to generate enhanced representation; and inputting the enhanced representation into a downstream prediction head to regress a scalar value as a fitness prediction result. The application models the interaction between the sequence-level difference vector and the pairwise evolutionary constraint, so that the non-additive effect between mutations can be accurately captured without relying on any experimental structure.
Owner:XIDIAN UNIV

Prediction of pathogenicity of protein mutations using amino acid fraction profiles

Methods, systems, and apparatus are disclosed, including computer programs encoded on a computer storage medium, for generating a pathogenicity score characterizing a likelihood that a protein mutation is a pathogenicity mutation, wherein the mutation modifies the amino acid sequence of the protein by replacing the original amino acid with a substitute amino acid at a mutation position in the amino acid sequence of the protein. In one aspect, a method includes generating a network input for a pathogenicity prediction neural network, where the network input includes a multiple sequence alignment (MSA) representation representing a protein; processing the network input using a pathogenicity prediction neural network to generate a fraction distribution over the set of amino acids; and generating a pathogenic fraction using the fraction distribution over the set of amino acids.
Owner:DEEPMIND TECH LTD

Protein structure prediction method, system and device based on optical quantum computer

The present invention belongs to the field of protein structure prediction technology and relates to a protein structure prediction method, system, and device based on an optical quantum computer. The method comprises: 1) obtaining a multiple sequence alignment (MSA) matrix of the target sequence based on the target sequence alignment of the protein to be predicted; 2) encoding the obtained MSA matrix into a {0,1} matrix; 3) converting the protein's amino acid interactions into an undirected graph model with {0,1} state nodes and training the model using an optical quantum computer using a Boltzmann machine training mechanism to obtain weight coefficients for the edges connecting the nodes in the undirected graph model; and 4) obtaining interaction coefficients between different amino acids in the protein based on the weight coefficients. This method improves the computational efficiency and solution results of protein structure prediction, and solves the problem that the existing technology is difficult to calculate and cannot obtain accurate solutions for complex scheduling problems.
Owner:BEIJING QBOSON QUANTUM TECH CO LTD

A privacy protection method for genome multiple sequence alignment based on secret sharing

The application provides a privacy protection method for genome multi-sequence alignment based on secret sharing, comprising the following steps: step one, a query party splits a genome sequence in a genome dataset; step two, the query party constructs a local public sub-sequence set; step three, a computing node A acquires a seed sequence set by alignment; step four, the query party locally reconstructs a seed sequence order; step five, the query party splits the genome sequence to obtain a sequence to be aligned; step six, the query party locally constructs secret sharing data slices; step seven, the computing nodes A and B perform multi-sequence alignment on the secret sharing data slices; and step eight, the query party restores and obtains a multi-sequence alignment calculation result. The application protects the privacy information of the genome sequence in the multi-sequence alignment process when realizing high-precision approximate multi-sequence alignment, and realizes the privacy protection of the genome multi-sequence alignment.
Owner:BEIHANG UNIV +1

Method for identifying DNA fragments of male parent nibea albiflora in gynogenesis pseudosciaena crocea genome

The invention discloses a method for identifying DNA (deoxyribonucleic acid) fragments of male parent nibea albiflora in a gynogenesis pseudosciaena crocea genome. According to the scheme, the pseudosciaena crocea genome is used as a reference coordinate system; male parent positive sites in a gynogenesis pseudosciaena crocea genome are screened by identifying the specific variation sites of the nibea albiflora and analyzing the genotypes of the gynogenesis pseudosciaena crocea at the specific variation sites of the nibea albiflora; performing multi-sequence comparison, calculating the number of positive sites in comparison blocks and the proportion, filtering candidate segments and the like to establish the method for identifying the genetic material from the male parent nibea albiflora in the gynogenetic large yellow croaker genome. At present, related identification is mainly carried out through comparative analysis and microsatellite marking; the methods have the problems of high cost, low positive rate and the like; according to the scheme, genetic materials derived from male parent nibea albiflora in a gynogenetic pseudosciaena crocea genome are comprehensively and accurately screened on the whole genome level, the economic character source of gynogenetic progeny can be defined, and application such as variety identity identification and variety right protection is facilitated by developing specific molecular markers.
Owner:FUJIAN AGRI & FORESTRY UNIV

Primers, kits, methods, and applications for detecting pathogenic *Pseudomonas proteus* strains based on the *oprL* gene.

This invention discloses primers, kits, methods, and applications for detecting pathogenic *Pseudomonas proteus* strains based on the *oprL* gene, belonging to the field of microbial detection technology. This invention provides specific primers for detecting pathogenic *P. proteus* strains based on the *oprL* gene. The primers are designed based on the differential patterns identified after multiple sequence alignment, specifically by analyzing the results of multiple sequence alignments of all *P. proteus* *oprL* gene sequences in the NCBI database. The nucleotide sequences of the primers are shown in SEQ ID No. 1, SEQ ID No. 2, and SEQ ID No. 3. The *oprL* gene of *P. proteus* is amplified by PCR, and the PCR amplification products are then subjected to electrophoresis. The pathogenicity of *P. proteus* can be effectively distinguished based on the electrophoresis image, enabling rapid detection of *P. proteus* infecting large yellow croaker. The primers have good specificity, and the detection method is simple and intuitive.
Owner:FUJIAN AGRI & FORESTRY UNIV

Protein sequence-structure co-generation method, system, device, and storage medium

This application provides a method, system, device, and storage medium for co-generating protein sequence-structure, relating to the field of bioinformatics. The generation method includes: acquiring multiple sequence alignment data and encoding it to obtain an explicit evolutionary prior representation; initializing the current sequence state and current structural state; performing at least one round of co-iterative iterative generation based on the explicit evolutionary prior representation; generating the final three-dimensional structure and outputting the protein sequence and the final three-dimensional structure. The method provided in this application directly guides the generation trajectory by extracting multiple sequence alignment data as an explicit evolutionary prior, ensuring the natural feasibility of the molecule. Simultaneously, it performs alternating and interwoven co-iterative updates of the sequence and structural states, breaking the limitations of sequential fragmentation and achieving bidirectional communication between spatial conformation and amino acid prediction. This results in a high degree of consistency between the final sequence and structure, significantly improving the functional hit rate and experimental success rate of the protein.
Owner:ACADEMY OF MILITARY MEDICAL SCIENCES

Phylogenetic tree construction method and system based on deep learning and beam search

This invention discloses a phylogenetic tree construction method and system based on deep learning and beam search. It predefines evolutionary scenarios, setting evolutionary parameters for each scenario with reference to real-world biological sequence attributes. Training, validation, and test sets are created based on simulated phylogenetic trees and corresponding multiple sequence alignment data according to the predefined parameters. A deep learning classifier with convolutional neural networks and long short-term memory neural networks as its core is constructed. The deep learning classifier is trained and validated using the training and validation sets, and its accuracy is tested using the test set. Based on the trained deep learning classifier and the sliding window method, classification predictions are performed on all sub-quad-sequence trees of the four-sequence data. A phylogenetic tree reconstruction is performed on the multiple sequence data using an improved stepwise addition method and the quad-sequence tree classification prediction results, resulting in a complete reconstruction. This enables phylogenetic tree construction under conditions of different species numbers and sequence lengths.
Owner:CHINESE INST FOR BRAIN RES BEIJING +1

Cooperative engineering transformation method of polymer fluorinase

The invention relates to the technical field of bioengineering, in particular to a collaborative engineering transformation method of polymer fluorinase, which comprises the step of performing step-by-step collaborative modification on an N-terminal sequence, a far-end region and a core region of a target enzyme. An N-terminal sequence is optimized through ancestor sequence reconstruction or hydrophobicity analysis, far-end mutation sites are screened through multi-sequence alignment and free energy calculation, the catalytic activity of a core region is improved in combination with conservative analysis and saturated mutation, and finally the multi-site combined mutant is constructed. According to the method, the overall structure stability and local catalytic efficiency of the polymer fluorinase can be remarkably improved, meanwhile, the thermal stability is enhanced, a universal methodology is provided for modification of other highly conservative polymer enzymes, and the method has a wide application prospect.
Owner:BEIJING UNIV OF CHEM TECH

A urease catalytic performance control method, device, equipment and storage medium

PendingCN122637885AUrocaninaseWild type
The application relates to the technical field of bioinformatics, and discloses a urease catalytic performance control method, device, equipment and storage medium. The method comprises the following steps: obtaining a wild-type amino acid sequence of a target urease and a three-dimensional structure corresponding to the wild-type amino acid sequence; mapping the amino acid sequence into a sequence feature unit, and mapping a local structure feature extracted from the three-dimensional structure into a structure feature unit; jointly encoding the sequence feature unit and the structure feature unit by using a decoupling multi-head cross attention module to obtain original logic values; constructing a multiple sequence alignment file based on homologous sequences, counting the amino acid frequency distribution of each residue position, and generating evolution logic values; determining an initial candidate mutant set; constructing an initial prediction model and a target prediction model; and predicting the initial candidate mutant set based on the target prediction model to obtain a target mutant with target catalytic performance. The scheme improves the control stability of urease catalytic performance.
Owner:LONGYAN UNIV

Predicting protein structures by sharing information between multiple sequence alignments and pair embeddings

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for predicting a structure of a protein comprising one or more chains. In one aspect, a method comprises: obtaining an initial multiple sequence alignment (MSA) representation; obtaining a respective initial pair embedding for each pair of amino acids in the protein; processing an input comprising the initial MSA representation and the initial pair embeddings using an embedding neural network to generate an output that comprises a final MSA representation and a respective final pair embedding for each pair of amino acids in the protein; and determining a predicted structure of the protein using the final MSA representation, the final pair embeddings, or both.
Owner:GDM HOLDING LLC

Tumor multiple detection kit and application thereof

The invention relates to the field of medicine, in particular to a tumor multiple detection kit and application thereof, and the kit comprises reagents for detecting HOTTIP, CCAT1, PVT1, MEG3, LINC01419, LINC01123, LINC01559, PCAT1, LINC02381, BCAR4, SNHG16 and LINC00665. The reagent comprises a primer group, and the sequences of the primer group are shown as SEQ ID NO: 1 to SEQ ID NO: 24. Compared with the prior art, the invention at least has the following beneficial effects: (1) target screening: based on literatures and databases, screening 12 kinds of trans-cancer high-expression lncRNAs with clinical value; (2) primer design: carrying out multi-sequence alignment by adopting Clustal X, and designing a specific primer in combination with a Gexp eXpress Profilter tool, so as to avoid a cross reaction; (3) technical optimization: verifying the sensitivity (LOD reaches 0.01%) through a gradient dilution experiment, and verifying the result accuracy by using fluorescent quantitative PCR or sequencing; and (4) clinical verification: cooperating with a hospital to obtain multiple cancer samples, comparing pathological diagnosis results, and ensuring that the detection consistency is greater than or equal to 95%.
Owner:SHANGHAI LINGEN BIOTECHNOLOGY CO LTD

A method, system and storage medium for designing and evaluating whole-virus primers

The present invention discloses a method, system and storage medium for designing and evaluating whole-virus primers. The method includes: obtaining a viral genome sequence library of a virus; designing usable primers for species identification using scheme one or two, performing inter-species verification and host verification on the usable primers for species identification, and evaluating each pair of usable primers for species identification; designing usable primers for subtype identification using scheme one or two, performing non-target virus subtype verification, inter-species verification and host verification on the usable primers for subtype identification, and evaluating each pair of usable primers for subtype identification. Scheme two is: performing multiple sequence alignment on the genomic sequences in the viral genome sequence library to identify the conserved regions of the virus, designing primers for the conserved regions, performing specific verification on the primers for the conserved regions, and using primers for the conserved regions with a passing rate that meets the verification standards as usable primers. The present invention improves the accuracy of primers, but is not universally applicable to the design of primers for viral subtype identification.
Owner:INST OF MICROBIOLOGY CHINESE ACAD OF SCI

Screening method for rational mutation sites of heat-sensitive udg based on strain culture temperature

PendingCN122337320ACold adaptedPrimary screening
This invention discloses a method for screening thermosensitive UDG rational mutation sites based on strain culture temperature. The method first divides strains into heat-tolerant and cold-adapted groups according to temperature thresholds based on the culture temperature of the strain preservation database, constructing corresponding UDG sequence sets and three-dimensional structure sets. Then, a first candidate mutation site set is generated through multiple sequence alignment and amino acid distribution difference analysis. A second candidate mutation site set is generated through rigid body alignment and functional protection rules under a unified coordinate system. Finally, multiple-source indicators such as structural dispersion, contact and salt bridge networks, and kinetic coupling are calculated for the second candidate mutation site set. Combined with effect-safety dual-dimensional scoring, spatial clustering, and cross-validation of the first candidate mutation site set, a preferred mutation site set is output. This invention achieves a systematic transformation from temperature ecotags to engineered sites, improving the reproducibility and interpretability of site screening without requiring a large-scale mutation library.
Owner:PULUOMAIGE BIOLOGICAL PRODS SHANGHAI

A method for protein generation based on position-specific weight matrix

The application discloses a protein generation method based on a position-specific weight matrix, comprising the following steps: obtaining a training sample set; training a transformer model based on the training sample set through a cross-entropy loss function, simultaneously outputting an amino acid sequence probability distribution, obtaining a predicted amino acid sequence through top-k sampling, performing multiple sequence alignment on the predicted amino acid sequence by adopting a psi-blast method to obtain a position-specific weight matrix, performing information entropy calculation on the position-specific weight matrix to obtain a weight-specific matrix information value; constructing a total loss function through the weight-specific matrix information value and a cross loss function, updating model parameters based on the training sample set through the total loss function to obtain an amino acid sequence generation model; taking a leading sequence as input and sequentially passing through the amino acid sequence generation model to generate an amino acid sequence, and folding the amino acid sequence through a trRosetta model to obtain protein secondary and tertiary structures.
Owner:ZJU HANGZHOU GLOBAL SCI & TECH INNOVATION CENT

A degenerate primer pair for detecting the tet(X) gene and its homologous genes and a method thereof

The present invention belongs to the field of molecular biology, and specifically discloses a degenerate primer pair for detecting the tet(X) gene and its homologous genes and a method therefor. The degenerate primers are designed based on the multiple sequence alignment of the tet(X) gene and its homologous genes, and have good specificity and sensitivity. By optimizing the PCR parameters, a PCR method for rapidly detecting the tet(X) gene and its homologous genes is established. Repeated experiments have proven that this method is rapid, sensitive, and can detect multiple tet(X) homologous genes with a pair of primers, with low detection cost, short time consumption, and good universality. It provides an efficient and reliable method for the epidemiological investigation of drug-resistant genes.
Owner:FOSHAN UNIVERSITY

MHC-presented peptide prediction method and system based on multimodal deep learning

The present invention provides a method and system for predicting MHC-presented peptides based on multimodal deep learning. The method comprises: constructing a multi-species MHC-peptide binding dataset; performing single genotype analysis on the multi-allele MHC-peptide data in the multi-species MHC-peptide binding dataset to determine unique MHC-peptide correspondences; performing multiple sequence alignment and feature alignment on cross-species MHC sequences, and using a frequency-weighted amino acid feature distance algorithm to extract core site information characterizing MHC molecule polymorphisms; and constructing a multimodal deep learning model that integrates MHC sequence features and peptide sequence features to predict the binding ability of MHC molecules and peptides. This method can more comprehensively and effectively capture the high-order interaction features between MHC and antigen peptides.
Owner:南昌大学第一附属医院 +1

Protein structure prediction from amino acid sequences using self-attention neural networks

Methods, systems and apparatus, including computer programs encoded on computer storage media, for determining a predicted structure of a protein specified by an amino acid sequence. In one aspect, the method comprises: obtaining a multiple sequence alignment for the protein; determining, from the multiple sequence alignment and for each amino acid pair in the amino acid sequence of the protein, a corresponding initial embedding of the amino acid pair; processing the initial embeddings of the amino acid pairs using a pairwise embedding neural network including a plurality of self-attention neural network layers to generate a final embedding for each amino acid pair; and determining a predicted structure of the protein based on the final embeddings for each amino acid pair.
Owner:GDM HOLDINGS LTD

A primer set, a kit and a detection method for mink coronavirus detection

The application discloses a primer group, a kit and a detection method for mink coronavirus detection, and belongs to the technical field of virus detection. The 1b fragment and N fragment genome sequences of the mink coronavirus are selected as target regions, multiple sequence alignment is carried out, and a conservative region of the genome sequence is selected to design two groups of detection primers. The specific reaction conditions of the two groups of detection primers are optimized respectively, and the sensitivity of the two groups of detection primers is verified through experiments. 135 high-throughput sequencing positive samples are selected, and the two groups of detection primers in the method are used for sample detection respectively, and the results are consistent with the high-throughput sequencing results. The RT-qPCR detection method based on SYBR-Green can be used for mink coronavirus detection and quantification.
Owner:SHANDONG FIRST MEDICAL UNIV & SHANDONG ACADEMY OF MEDICAL SCI