Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

44 results about "Multiple sequence alignment" patented technology

A multiple sequence alignment (MSA) is a sequence alignment of three or more biological sequences, generally protein, DNA, or RNA. In many cases, the input set of query sequences are assumed to have an evolutionary relationship by which they share a linkage and are descended from a common ancestor. From the resulting MSA, sequence homology can be inferred and phylogenetic analysis can be conducted to assess the sequences' shared evolutionary origins. Visual depictions of the alignment as in the image at right illustrate mutation events such as point mutations (single amino acid or nucleotide changes) that appear as differing characters in a single alignment column, and insertion or deletion mutations (indels or gaps) that appear as hyphens in one or more of the sequences in the alignment. Multiple sequence alignment is often used to assess sequence conservation of protein domains, tertiary and secondary structures, and even individual amino acids or nucleotides.

Method and system for protein three-dimensional structure prediction

The application provides a protein three-dimensional structure prediction method and system, comprising the following steps: S1: obtaining a multiple sequence alignment matrix containing protein coevolution information and a protein template, and performing conversion of MSA sequence and residue pair coding to obtain MSA sequence coding and residue pair coding; S2: the MSA sequence coding and the residue pair coding are updated by a recurrent attention neural network to generate the latest MSA sequence coding and the residue pair coding; S3: based on the current residue pair coding, target sequence coding and starting main chain framework T i The current target sequence coding is updated by the invariant point attention neural network; and S4: the main chain framework is updated based on the current target sequence coding, and the torsion angle between amino acids is calculated, so that the protein main chain structure is obtained, and the side chain atom angle and the final three-dimensional structure are calculated through the residue network.
Owner:SHANGHAI TIANRANG NETWORK TECH CO LTD

A method for constructing a multiplex PCR reaction system for detecting and identifying pycnospora and its application

ActiveCN116179735BOptimizing Multiplex PCR Reaction ConditionsMultiplexGene cluster
The application discloses a construction method of a multiplex PCR reaction system for detecting and identifying Podosphaera species and application thereof. The method comprises the following steps: obtaining whole genome sequences of multiple Podosphaera species, and performing multiple sequence alignment to obtain specific genes and specific gene clusters of each species; taking the genes in the specific gene clusters of each species as templates, designing multiple upstream primers and downstream primers according to a conventional primer design method to generate primer pairs, and excluding unreasonable primer pairs and primer pairs with low sensitivity and specificity; according to the size of the amplification products, the primers are freely combined and matched to generate four pairs of mixed multiplex primers, and the multiplex PCR reaction conditions are optimized to obtain a multiplex PCR system capable of specifically detecting Podosphaera and simultaneously identifying single and multiple species. The multiplex PCR method for detecting and identifying Podosphaera provided by the application can quickly and accurately complete the detection of Podosphaera and identify the species identity of Podosphaera.
Owner:NANJING AGRICULTURAL UNIVERSITY

An indel molecular marker for identifying old crow petal of wanyu and application thereof

The present application relates to the technical field of molecular marker, in particular to Indel molecular marker of Wan-Yu old crow petal and application thereof. The Indel molecular marker of the present application is shown as SEQ ID NO. 1-3. The Indel molecular marker can be used to identify Wan-Yu old crow petal, and the identification can be carried out by electrophoresis, multiple sequence alignment or reads alignment of high-throughput sequencing.
Owner:YIHU BIOTECHNOLOGY (ANHUI) CO LTD

Blood protein characteristic polypeptide, detection kit and application of blood protein characteristic polypeptide in identifying adulteration and weight gain of hirudo nipponia

The invention belongs to the field of biological detection, and particularly relates to a blood protein characteristic polypeptide, a detection kit and application of the blood protein characteristic polypeptide to identification of adulteration and weight gain of hirudo nipponia. And the sequence of the blood protein marker polypeptide is LLGNVIVVVLAR. On the basis of constructing a species marker polypeptide identification method, a heterologous blood protein traceability analysis strategy is introduced, beta globin in blood of common livestock such as pigs, cattle, sheep, horses, donkeys and the like is subjected to multiple sequence alignment, and highly conservative species-independent marker polypeptide is screened out as a chemical marker; and a multi-reaction monitoring (MRM) non-standard quantitative peak area ratio judgment and detection method based on a liquid chromatography-triple quadrupole mass spectrometry technology is established, so that a scientific judgment basis is provided for the problem of blood doping and weight increment.
Owner:SHANDONG INST FOR FOOD & DRUG CONTROL +1

A HBV S protein-specific monoclonal antibody HBV-S-4H7 and its use in preparing a detection kit

The present invention relates to a HBV S protein-specific monoclonal antibody, HBV-S-4H7, and its use in the preparation of a detection kit. Based on the amino acid sequences of different HBV S proteins, multiple sequence alignment, and online epitope screening software, the present invention ultimately selects a preferred epitope peptide and immunizes mice to prepare a broad-spectrum monoclonal antibody. The antibody has good binding properties. Based on the role of the S protein in the virus, neutralization experiments confirm that the monoclonal antibody prepared by the present invention also has good virus neutralization effect, and has good application prospects.
Owner:SHAANXI ZHUOJIEMU BIOTECHNOLOGY CO LTD

Targeting superantigen fusion protein based on improved se(3)-transformer and implementation method

PendingCN122266440AAchieve collaborative structure optimizationImprove targetingMicroorganism based processesBiostatisticsPattern recognitionAntigen epitope
The application discloses a targeting superantigen fusion protein based on an improved SE(3)-Transformer and an implementation method. In an offline stage, amino acid sequences are first converted into one-hot encoding or language model embedding (such as ESM-2), and are spliced with multiple sequence alignment (MSA) features for geometric initialization. A neural network (improved SE(3)-Transformer) combined with multiple sequence alignment (MSA) and an attention mechanism is constructed to predict the coordinates of C alpha, C, N and O atoms for main chain prediction, and the neural network is trained through a gradient descent method based on a physical heuristic potential item. In a verification stage, the improved SE(3)-Transformer after training is used to generate a predicted structure, conformational stability is verified through a simplified force field, and fine tuning is performed based on a confidence score. The application can accurately predict the structure of a target antigen epitope and an antibody variable region, optimize a superantigen functional domain in combination with a graph neural network, and dynamically design a flexible connecting peptide to realize modular fusion.
Owner:SHANGHAI JIAOTONG UNIV

Protein fitness prediction method, device, terminal and storage medium

ActiveCN121393550BSolve Application Bottleneckseffective nonadditive effectBiostatisticsBiological modelsWild typeProtein engineering
The application relates to the technical field of protein engineering and bioinformatics, and specifically provides a protein fitness prediction method and device, a terminal and a storage medium. The method comprises the following steps: extracting wild-type sequence-level representation and mutant sequence-level representation by using a pre-trained protein language model, and calculating a representation difference vector between the two; deducing coevolution coupling information from multiple sequence alignment information, and constructing pairwise evolutionary constraint information reflecting spatial proximity relationship; then modeling the interaction between the representation difference vector and the pairwise evolutionary constraint information through an attention mechanism, so that the pairwise coevolution information guides the propagation and weighting of the representation difference vector between residues to generate enhanced representation; and inputting the enhanced representation into a downstream prediction head to regress a scalar value as a fitness prediction result. The application models the interaction between the sequence-level difference vector and the pairwise evolutionary constraint, so that the non-additive effect between mutations can be accurately captured without relying on any experimental structure.
Owner:XIDIAN UNIV

A privacy protection method for genome multiple sequence alignment based on secret sharing

The application provides a privacy protection method for genome multi-sequence alignment based on secret sharing, comprising the following steps: step one, a query party splits a genome sequence in a genome dataset; step two, the query party constructs a local public sub-sequence set; step three, a computing node A acquires a seed sequence set by alignment; step four, the query party locally reconstructs a seed sequence order; step five, the query party splits the genome sequence to obtain a sequence to be aligned; step six, the query party locally constructs secret sharing data slices; step seven, the computing nodes A and B perform multi-sequence alignment on the secret sharing data slices; and step eight, the query party restores and obtains a multi-sequence alignment calculation result. The application protects the privacy information of the genome sequence in the multi-sequence alignment process when realizing high-precision approximate multi-sequence alignment, and realizes the privacy protection of the genome multi-sequence alignment.
Owner:BEIHANG UNIV +1

Method for identifying DNA fragments of male parent nibea albiflora in gynogenesis pseudosciaena crocea genome

The invention discloses a method for identifying DNA (deoxyribonucleic acid) fragments of male parent nibea albiflora in a gynogenesis pseudosciaena crocea genome. According to the scheme, the pseudosciaena crocea genome is used as a reference coordinate system; male parent positive sites in a gynogenesis pseudosciaena crocea genome are screened by identifying the specific variation sites of the nibea albiflora and analyzing the genotypes of the gynogenesis pseudosciaena crocea at the specific variation sites of the nibea albiflora; performing multi-sequence comparison, calculating the number of positive sites in comparison blocks and the proportion, filtering candidate segments and the like to establish the method for identifying the genetic material from the male parent nibea albiflora in the gynogenetic large yellow croaker genome. At present, related identification is mainly carried out through comparative analysis and microsatellite marking; the methods have the problems of high cost, low positive rate and the like; according to the scheme, genetic materials derived from male parent nibea albiflora in a gynogenetic pseudosciaena crocea genome are comprehensively and accurately screened on the whole genome level, the economic character source of gynogenetic progeny can be defined, and application such as variety identity identification and variety right protection is facilitated by developing specific molecular markers.
Owner:FUJIAN AGRI & FORESTRY UNIV

Primers, kits, methods, and applications for detecting pathogenic *Pseudomonas proteus* strains based on the *oprL* gene.

This invention discloses primers, kits, methods, and applications for detecting pathogenic *Pseudomonas proteus* strains based on the *oprL* gene, belonging to the field of microbial detection technology. This invention provides specific primers for detecting pathogenic *P. proteus* strains based on the *oprL* gene. The primers are designed based on the differential patterns identified after multiple sequence alignment, specifically by analyzing the results of multiple sequence alignments of all *P. proteus* *oprL* gene sequences in the NCBI database. The nucleotide sequences of the primers are shown in SEQ ID No. 1, SEQ ID No. 2, and SEQ ID No. 3. The *oprL* gene of *P. proteus* is amplified by PCR, and the PCR amplification products are then subjected to electrophoresis. The pathogenicity of *P. proteus* can be effectively distinguished based on the electrophoresis image, enabling rapid detection of *P. proteus* infecting large yellow croaker. The primers have good specificity, and the detection method is simple and intuitive.
Owner:FUJIAN AGRI & FORESTRY UNIV

Protein sequence-structure co-generation method, system, device, and storage medium

This application provides a method, system, device, and storage medium for co-generating protein sequence-structure, relating to the field of bioinformatics. The generation method includes: acquiring multiple sequence alignment data and encoding it to obtain an explicit evolutionary prior representation; initializing the current sequence state and current structural state; performing at least one round of co-iterative iterative generation based on the explicit evolutionary prior representation; generating the final three-dimensional structure and outputting the protein sequence and the final three-dimensional structure. The method provided in this application directly guides the generation trajectory by extracting multiple sequence alignment data as an explicit evolutionary prior, ensuring the natural feasibility of the molecule. Simultaneously, it performs alternating and interwoven co-iterative updates of the sequence and structural states, breaking the limitations of sequential fragmentation and achieving bidirectional communication between spatial conformation and amino acid prediction. This results in a high degree of consistency between the final sequence and structure, significantly improving the functional hit rate and experimental success rate of the protein.
Owner:ACADEMY OF MILITARY MEDICAL SCIENCES

Phylogenetic tree construction method and system based on deep learning and beam search

This invention discloses a phylogenetic tree construction method and system based on deep learning and beam search. It predefines evolutionary scenarios, setting evolutionary parameters for each scenario with reference to real-world biological sequence attributes. Training, validation, and test sets are created based on simulated phylogenetic trees and corresponding multiple sequence alignment data according to the predefined parameters. A deep learning classifier with convolutional neural networks and long short-term memory neural networks as its core is constructed. The deep learning classifier is trained and validated using the training and validation sets, and its accuracy is tested using the test set. Based on the trained deep learning classifier and the sliding window method, classification predictions are performed on all sub-quad-sequence trees of the four-sequence data. A phylogenetic tree reconstruction is performed on the multiple sequence data using an improved stepwise addition method and the quad-sequence tree classification prediction results, resulting in a complete reconstruction. This enables phylogenetic tree construction under conditions of different species numbers and sequence lengths.
Owner:CHINESE INST FOR BRAIN RES BEIJING +1

Cooperative engineering transformation method of polymer fluorinase

The invention relates to the technical field of bioengineering, in particular to a collaborative engineering transformation method of polymer fluorinase, which comprises the step of performing step-by-step collaborative modification on an N-terminal sequence, a far-end region and a core region of a target enzyme. An N-terminal sequence is optimized through ancestor sequence reconstruction or hydrophobicity analysis, far-end mutation sites are screened through multi-sequence alignment and free energy calculation, the catalytic activity of a core region is improved in combination with conservative analysis and saturated mutation, and finally the multi-site combined mutant is constructed. According to the method, the overall structure stability and local catalytic efficiency of the polymer fluorinase can be remarkably improved, meanwhile, the thermal stability is enhanced, a universal methodology is provided for modification of other highly conservative polymer enzymes, and the method has a wide application prospect.
Owner:BEIJING UNIV OF CHEM TECH

A urease catalytic performance control method, device, equipment and storage medium

PendingCN122637885AUrocaninaseWild type
The application relates to the technical field of bioinformatics, and discloses a urease catalytic performance control method, device, equipment and storage medium. The method comprises the following steps: obtaining a wild-type amino acid sequence of a target urease and a three-dimensional structure corresponding to the wild-type amino acid sequence; mapping the amino acid sequence into a sequence feature unit, and mapping a local structure feature extracted from the three-dimensional structure into a structure feature unit; jointly encoding the sequence feature unit and the structure feature unit by using a decoupling multi-head cross attention module to obtain original logic values; constructing a multiple sequence alignment file based on homologous sequences, counting the amino acid frequency distribution of each residue position, and generating evolution logic values; determining an initial candidate mutant set; constructing an initial prediction model and a target prediction model; and predicting the initial candidate mutant set based on the target prediction model to obtain a target mutant with target catalytic performance. The scheme improves the control stability of urease catalytic performance.
Owner:LONGYAN UNIV

Predicting protein structures by sharing information between multiple sequence alignments and pair embeddings

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for predicting a structure of a protein comprising one or more chains. In one aspect, a method comprises: obtaining an initial multiple sequence alignment (MSA) representation; obtaining a respective initial pair embedding for each pair of amino acids in the protein; processing an input comprising the initial MSA representation and the initial pair embeddings using an embedding neural network to generate an output that comprises a final MSA representation and a respective final pair embedding for each pair of amino acids in the protein; and determining a predicted structure of the protein using the final MSA representation, the final pair embeddings, or both.
Owner:GDM HOLDING LLC

A method, system and storage medium for designing and evaluating whole-virus primers

The present invention discloses a method, system and storage medium for designing and evaluating whole-virus primers. The method includes: obtaining a viral genome sequence library of a virus; designing usable primers for species identification using scheme one or two, performing inter-species verification and host verification on the usable primers for species identification, and evaluating each pair of usable primers for species identification; designing usable primers for subtype identification using scheme one or two, performing non-target virus subtype verification, inter-species verification and host verification on the usable primers for subtype identification, and evaluating each pair of usable primers for subtype identification. Scheme two is: performing multiple sequence alignment on the genomic sequences in the viral genome sequence library to identify the conserved regions of the virus, designing primers for the conserved regions, performing specific verification on the primers for the conserved regions, and using primers for the conserved regions with a passing rate that meets the verification standards as usable primers. The present invention improves the accuracy of primers, but is not universally applicable to the design of primers for viral subtype identification.
Owner:INST OF MICROBIOLOGY CHINESE ACAD OF SCI

Screening method for rational mutation sites of heat-sensitive udg based on strain culture temperature

PendingCN122337320ACold adaptedPrimary screening
This invention discloses a method for screening thermosensitive UDG rational mutation sites based on strain culture temperature. The method first divides strains into heat-tolerant and cold-adapted groups according to temperature thresholds based on the culture temperature of the strain preservation database, constructing corresponding UDG sequence sets and three-dimensional structure sets. Then, a first candidate mutation site set is generated through multiple sequence alignment and amino acid distribution difference analysis. A second candidate mutation site set is generated through rigid body alignment and functional protection rules under a unified coordinate system. Finally, multiple-source indicators such as structural dispersion, contact and salt bridge networks, and kinetic coupling are calculated for the second candidate mutation site set. Combined with effect-safety dual-dimensional scoring, spatial clustering, and cross-validation of the first candidate mutation site set, a preferred mutation site set is output. This invention achieves a systematic transformation from temperature ecotags to engineered sites, improving the reproducibility and interpretability of site screening without requiring a large-scale mutation library.
Owner:PULUOMAIGE BIOLOGICAL PRODS SHANGHAI

A method for protein generation based on position-specific weight matrix

The application discloses a protein generation method based on a position-specific weight matrix, comprising the following steps: obtaining a training sample set; training a transformer model based on the training sample set through a cross-entropy loss function, simultaneously outputting an amino acid sequence probability distribution, obtaining a predicted amino acid sequence through top-k sampling, performing multiple sequence alignment on the predicted amino acid sequence by adopting a psi-blast method to obtain a position-specific weight matrix, performing information entropy calculation on the position-specific weight matrix to obtain a weight-specific matrix information value; constructing a total loss function through the weight-specific matrix information value and a cross loss function, updating model parameters based on the training sample set through the total loss function to obtain an amino acid sequence generation model; taking a leading sequence as input and sequentially passing through the amino acid sequence generation model to generate an amino acid sequence, and folding the amino acid sequence through a trRosetta model to obtain protein secondary and tertiary structures.
Owner:ZJU HANGZHOU GLOBAL SCI & TECH INNOVATION CENT

A primer set, a kit and a detection method for mink coronavirus detection

The application discloses a primer group, a kit and a detection method for mink coronavirus detection, and belongs to the technical field of virus detection. The 1b fragment and N fragment genome sequences of the mink coronavirus are selected as target regions, multiple sequence alignment is carried out, and a conservative region of the genome sequence is selected to design two groups of detection primers. The specific reaction conditions of the two groups of detection primers are optimized respectively, and the sensitivity of the two groups of detection primers is verified through experiments. 135 high-throughput sequencing positive samples are selected, and the two groups of detection primers in the method are used for sample detection respectively, and the results are consistent with the high-throughput sequencing results. The RT-qPCR detection method based on SYBR-Green can be used for mink coronavirus detection and quantification.
Owner:SHANDONG FIRST MEDICAL UNIV & SHANDONG ACADEMY OF MEDICAL SCI

Synthetic Augmentation of Multiple Sequence Alignment of Protein-Protein Interactions

The present disclosure provides a method of predicting a structure of an interface between a target peptide and a targeting peptide. The method leverages test pairs of variants of a target peptide and variants of a targeting peptide and their binding affinities measured by a high-throughput analysis. Synergistic pairs among the test pairs are selected and multiple sequence alignment (MSA) of the selected pairs is performed to predict a structure of the protein complex formed with the target peptide and the targeting peptide. Structure prediction using MSA of the synergistic pairs provides for improved results, thereby paving the path for downstream analyses, e.g., small molecule design for molecular glues or antibody design.
Owner:A ALPHA BIO INC

Bio-sequence alignment method and system based on parameter adaptive growing optimizer

The present disclosure provides a biological multiple sequence alignment method and system based on a parameter adaptive growth optimizer, relating to the technical field of biological multiple sequence alignment, comprising initializing a hidden Markov model, obtaining a gene sequence file to be aligned, and determining the length of the gene sequence; setting the parameters of the hidden Markov model according to the length of the gene sequence, and then obtaining an alignment result based on the hidden Markov model; wherein in the hidden Markov model, a four-parameter adaptive growth optimizer algorithm is used to adaptively update individuals, a Jensen-Shannon divergence balancing factor is introduced to balance the adaptive optimization process of the mutually antagonistic parameters in the antagonistic characteristics, so that the population highly adapts to evolution, and then the individuals are subjected to boundary constraints, and the out-of-bound components in a certain dimension are reinitialized within the effective range. The present disclosure can fully utilize the current known information to adaptively adjust the settings of its parameters.
Owner:SHANDONG NORMAL UNIV

Primer group, kit and detection method for detecting xanthomonas campestris pv. Campestris

The invention discloses a primer group for detecting xanthomonas campestris pv.campestris. The primer group consists of an upstream primer and a downstream primer, wherein the nucleotide sequence of the upstream primer is as shown in SEQ ID NO.1, and the nucleotide sequence of the downstream primer is as shown in SEQ ID NO.2. The invention further discloses a kit for detecting the xanthomonas campestris pv.campestris. According to the present invention, multiple sequence comparison analysis is performed on the whole genome sequences of the Xanthomonas campestris pv. Campestris pathogen and other related bacteria, such that a pair of primers is designed according to the genome specificity region of the Xanthomonas campestris pv. Campestris pv. Campestris pv. Campestris pv. Campestris pv. Campestris pv. Campestris pv. Campestris pv. Campestris; the method is used for high-sensitivity rapid molecular detection of plants with the brassica oleracea, and a rapid, simple, high-specificity and high-sensitivity monitoring technology system for the brassica oleracea is established.
Owner:JIANGSU ACAD OF AGRI SCI

Accurate sequence analysis method and system based on nanopore length and read length sequencing data

The invention discloses an accurate sequence analysis method and system based on nanopore length and read length sequencing data. The method comprises the following steps: comparing nanopore sequencing data with a reference genome through a minimap2 algorithm, and screening out high-quality long-read-length sequences in a target area and specific upstream and downstream ranges; extracting a target area local sequence of each read length, calculating an editing distance matrix, performing Box-Cox transformation, performing clustering analysis by adopting a Gaussian mixture model, and determining an optimal clustering number through a BIC criterion; performing multi-sequence alignment and polishing on each type of sequences to generate a high-precision consensus sequence; and finally, recognizing homologous conserved sequence fragments through re-comparison and global comparison, and outputting candidate sequences for targeting PCR (Polymerase Chain Reaction) primer or CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) sgRNA design. According to the method, the problem of high error rate of nanopore length and read length data can be effectively solved, the accuracy and reliability of sequence analysis are improved, and reliable technical support is provided for precise gene editing and molecular diagnosis.
Owner:CHENGDU SEVENTH PEOPLES HOSPITAL

Construction and application of polyphosphate kinase bsppk mutant and its producing strain

ActiveCN122012453BArginineNucleotide
The application discloses a polyphosphate kinase BsPPK mutant and construction and application of a producing strain thereof, and belongs to the technical field of genetic engineering. The application carries out site-directed mutation on polyphosphate kinase BsPPK from the genus Bredia through homologous modeling, molecular docking and multiple sequence alignment, mutates lysine at the 92th position of the wild type BsPPK into alanine, threonine at the 95th position into alanine, and serine at the 205th position into arginine, and obtains a combined mutant BsPPK-K92A / T95A / S205R. Enzyme activity determination results show that the specific enzyme activity of the mutant is increased by 324.4% compared with the wild type, and the catalytic efficiency on AMP is significantly improved. When the mutant is applied to synthesis of UDP-Gal and derivative products thereof, only 15 mM AMP is needed to achieve a similar yield obtained by using 30 mM AMP for the wild type, and the nucleotide consumption is reduced by 50%.
Owner:OCEAN UNIV OF CHINA

Mutant of protein glutaminase with improved heat resistance as well as construction method and application of mutant

The invention discloses a heat-resistant mutant of protein glutaminase (PG enzyme) as well as a construction method and application of the heat-resistant mutant, and belongs to the technical field of protein engineering. The amino acid sequence of the PG enzyme is obtained by performing multi-sequence comparison on the PG enzyme from C. proteolyticum YF810 and other PG enzymes which are from different strains and have improved heat resistance, screening heat-resistant single-point mutations, and performing combined mutation on two or more of all beneficial single-point mutations to obtain a combined mutant of the PG enzyme. According to the five-combination mutant of the PG enzyme with the heat resistance improved to the maximum degree, the T5010min value is improved by 12.80 DEG C, the t1 / 260 DEG C value is improved by 55.13 times, and compared with wild type PG enzyme, the enzyme activity is not reduced. Therefore, compared with the wild type PG enzyme, the mutant of the PG enzyme provided by the invention has better heat resistance and better industrial application prospect.
Owner:EAST CHINA NORMAL UNIV

Protein design methods, apparatuses, devices, and media

The present disclosure provides a protein design method, device, equipment and medium, relates to the field of artificial intelligence, in particular to the technical field of deep learning, biological computing and large language model. The generation method comprises the following steps: constructing a plurality of candidate proteins, each of which comprises a first chain of an original protein and a non-natural sequence constructed based on a second chain of the original protein; retrieving a first multiple sequence alignment of the first chain and a second multiple sequence alignment of the non-natural sequence; matching the first multiple sequence alignment and the second multiple sequence alignment to obtain a cross-chain homologous sequence by using a pre-trained initial protein language model; predicting the structure and a first score of the candidate protein by using a protein structure prediction model; determining a reward value based on the first score and performing reinforcement learning training on the initial protein language model; determining a second score of each of the plurality of candidate proteins by using the trained target protein language model and the protein structure prediction model, so as to obtain a protein design result.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

T-Cell Receptor Repertoire Selection Prediction with Physical Model Augmented Pseudo-Labeling for Personalized Medicine Decision Making

Systems and methods for predicting T-Cell receptor (TCR)-peptide interaction, including training a deep learning model for the prediction of TCR-peptide interaction by determining a multiple sequence alignment (MSA) for TCR-peptide pair sequences from a dataset of TCR-peptide pair sequences using a sequence analyzer, building TCR structures and peptide structures using the MSA and corresponding structures from a Protein Data Bank (PDB) using a MODELLER, and generating an extended TCR-peptide training dataset based on docking energy scores determined by docking peptides to TCRs using physical modeling based on the TCR structures and peptide structures built using the MODELLER. TCR-peptide pairs are classified and labeled as positive or negative pairs using pseudo-labels based on the docking energy scores, and the deep learning model is iteratively retrained based on the extended TCR-peptide training dataset and the pseudo-labels until convergence.
Owner:NEC LABORATORIES AMERICA INC

A method of co-engineering a multimeric fluorinase

The application relates to the technical field of bioengineering, and particularly relates to a synergistic engineering modification method of a multimeric fluorinase, which comprises step-by-step synergistic modification of an N-terminal sequence, a distal region and a core region of a target enzyme. The N-terminal sequence is optimized through ancestral sequence reconstruction or hydrophobicity analysis, the distal mutation site is screened by using multiple sequence alignment and free energy calculation, and the catalytic activity of the core region is improved by combining conservativeness analysis and saturation mutation, so as to finally construct a multi-site combined mutant. The application can significantly improve the overall structural stability and local catalytic efficiency of the multimeric fluorinase, and meanwhile, the thermal stability is enhanced, a universal methodology is provided for modification of other highly conservative multimeric enzymes, and has a wide application prospect.
Owner:BEIJING UNIV OF CHEM TECH

A sequence diffusion method based on multiple sequence alignment and electronic device thereof

The application discloses a sequence diffusion method based on multiple sequence alignment and an electronic device thereof. The application adopts two-stage training to split the purposes of training in different stages, which is more convenient for the interpretability of the model and the flexible training of the field sub-model. The application focuses on the feature extraction of MSA, extracts the hidden pair information in MSA by using a deeper attention layer, and can effectively capture the biological correlation between sequences. The application uses the sequence diffusion mode to realize the sequence generation guided by the evolutionary information, and can generate sequences with high correctness and diversity.
Owner:HANGZHOU LEVINTHAL BIOTECHNOLOGY CO LTD

T-cell receptor repertoire selection prediction with physical model augmented pseudo-labeling for personalized medicine decision making

Systems and methods for predicting T-Cell receptor (TCR)-peptide interaction, including training a deep learning model for the prediction of TCR-peptide interaction by determining a multiple sequence alignment (MSA) for TCR-peptide pair sequences from a dataset of TCR-peptide pair sequences using a sequence analyzer, building TCR structures and peptide structures using the MSA and corresponding structures from a Protein Data Bank (PDB) using a MODELLER, and generating an extended TCR-peptide training dataset based on docking energy scores determined by docking peptides to TCRs using physical modeling based on the TCR structures and peptide structures built using the MODELLER. TCR-peptide pairs are classified and labeled as positive or negative pairs using pseudo-labels based on the docking energy scores, and the deep learning model is iteratively retrained based on the extended TCR-peptide training dataset and the pseudo-labels until convergence.
Owner:NEC CORP