Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

11 results about "Representative sequences" patented technology

Representative sequences are short regions within protein sequences that can be used to approximate the evolutionary relationships of those proteins, or the organisms from which they come. Representative sequences are contiguous subsequences (typically 300 residues) from ubiquitous, conserved proteins, such that each orthologous family of representative sequences taken alone gives a distance matrix in close agreement with the consensus matrix.

Heuristic biological sequence clustering method based on semi-global comparison

PendingCN120299521ABiostatisticsSequence analysisSequence alignment algorithmSequence clustering
The invention provides a heuristic biological sequence clustering method based on semi-global comparison. The method comprises the following steps: firstly, reading all biological sequences, removing repetitive sequences, and carrying out descending sorting on the biological sequences according to sequence lengths; a first sequence is used as a representative sequence of a first category, then a next sequence is read, the similarity between the current sequence and the representative sequence is calculated by adopting a semi-global sequence comparison algorithm, each sequence can be compared to the most similar area in the representative sequence through semi-global sequence comparison, the similarity between the sequences is found to the maximum extent, and the similarity between the sequences and the representative sequence is calculated. Obtaining a relatively high comparison similarity value between the sequences; if the similarity meets a clustering threshold value, adding the similarity into a category with the same representative sequence, otherwise, taking the similarity as a new representative sequence, and generating a new category; repeating the steps until all the sequences are processed; the number of the final representative sequences is the number of clustering categories, and the categories which are the same as the representative sequences are member sequences of each category.
Owner:BAOJI UNIV OF ARTS & SCI

Edge computing-based methods for the detection and drug resistance analysis of Mycobacterium tuberculosis.

PendingCN122314109AFast and effective detectionRapid and effective drug resistance analysisData setRepresentative sequences
This application relates to the field of microbial detection technology, specifically to a method for detecting and analyzing the drug resistance of Mycobacterium tuberculosis based on edge computing. The method includes: classifying sequencing data of biological samples based on a metagenomic reference database to obtain a matching dataset aligned to Mycobacterium tuberculosis; extracting representative sequences of each amplified fragment from multiple amplified fragments in the first region based on the matching dataset; and performing a first alignment of the representative sequences of each amplified fragment with a mycobacterium reference database to identify Mycobacterium tuberculosis complexes based on sequence similarity. This method is suitable for operation on edge devices. The method proposed in this application, through process optimization, enables accurate, efficient, and low-cost detection of mycobacteria even with limited computing resources, thereby helping to prevent and control the spread of tuberculosis.
Owner:MGI TECH CO LTD

Primer design methods, systems, and computer storage media based on nucleic acid identity

This application discloses a primer design method, system, and computer storage medium based on nucleic acid consistency. The method includes: obtaining a target nucleic acid sequence; classifying the target nucleic acid sequence to obtain multiple taxonomic units; using at least one target nucleic acid sequence from each taxonomic unit as a representative sequence; designing candidate primers for the representative sequence; sequentially selecting at most one pair of candidate primers from each taxonomic unit to establish a candidate primer pool; and excluding candidate primers in the candidate primer pool that exhibit non-specific amplification to obtain the final primer pool. This application is simple to operate and can handle large nucleic acid libraries, making it particularly suitable for universal primer design for pathogen libraries. Compared to similar software, the method in this application is more systematic and comprehensive, resulting in a primer pool with high coverage and low host contamination rate.
Owner:HUGOBIOTECH BEIJING CO LTD +1

Primer design method, system and computer storage medium based on minimum degeneracy

The application discloses a primer design method, system and computer storage medium based on minimum degeneracy, the method comprising: obtaining nucleotide sequences; dividing the nucleotide sequences into multiple classification units; selecting a representative sequence from each classification unit; designing a corresponding primer pair set for the representative sequence based on the base frequency in the representative sequence; determining the coverage of the primer pair set relative to all nucleotide sequences, and selecting multiple primer pairs in each primer pair set as a candidate primer pair set according to the coverage of each primer pair set; and selecting multiple candidate primers in the candidate primer pair set to establish a primer pool. The application considers the mismatch between primer pairs and the coverage in the primer design process, thereby allowing the primer pairs to have extremely high coverage and lower degeneracy.
Owner:HUGOBIOTECH BEIJING CO LTD +1

A high-throughput sequencing data sequence clustering method and system based on cyclic self-blast

ActiveCN122067616BBarcodeSequence clustering
The application discloses a high-throughput sequencing data sequence clustering method and system based on cyclic self-blast: the high-throughput sequencing data is de-duplicated; unique sequences with a number of repetitions lower than a threshold value a are filtered, and the filtered unique sequences are sorted according to the number of repetitions; the sorted unique sequences are divided into subsets and a parent set; the subsets are subjected to blast multiple alignment to obtain a de-redundant subset; the parent set and the de-redundant subset are subjected to cyclic blast alignment to obtain a de-redundant parent set; all representative sequences in a preliminary clustering set are combined, sorted according to the number of repetitions and distributed with identification tags to generate a final OTU clustering result set; in the prior art, when OTU clustering and ASV methods are used to process environmental DNA macro-barcode sequencing data, different OTU or ASV sequences are annotated to the same species, so that a large amount of redundancy still exists in the classified units after clustering, and the reliability of species analysis is improved.
Owner:NANJING NORMAL UNIVERSITY +1

High-throughput sequencing data sequence clustering method and system based on circulating self-blast

ActiveCN122067616ABiostatisticsSequence analysisBarcodeSequence clustering
The invention discloses a high-throughput sequencing data sequence clustering method and system based on circulating self-blast. The method comprises the following steps: carrying out duplicate removal processing on high-throughput sequencing data; filtering the unique sequences of which the repetition number is lower than a threshold value a, and sorting the filtered unique sequences according to the repetition number; dividing the sorted unique sequence into a subset and a mother set; performing blast multiple comparison on the subsets to obtain redundancy-removed subsets; performing cyclic blast comparison on the mother set and the redundancy-removed subset to obtain a redundancy-removed mother set; combining all representative sequences in the preliminary clustering set, sorting according to the number of repetitions and distributing identification labels, and generating a final OTU clustering result set; in the prior art, when an OTU clustering method and an ASV method are used for processing environment DNA macro bar code sequencing data, different OTU or ASV sequences are annotated to the same species, so that a large amount of redundancy still exists in a classified unit after clustering, and the reliability of species analysis is improved.
Owner:NANJING NORMAL UNIVERSITY +1

Messenger Ribonucleic Acid Sequence Design Method, Device, Computing Device and Storage Medium

The present invention relates to the field of biopharmaceutical technology, and discloses a method, device, computing device and storage medium for messenger ribonucleic acid sequence design. Based on a given 5' UTR sequence as the starting input, the method adds new codons one by one using the growth method, and calculates and scores the full-length minimum free energy and codon usage efficiency of the overall sequence for sorting. The elimination method is used to exclude sequences containing restriction enzyme sites / undesirable properties / or poor scores, and then the clustering search method is used to select representative sequences to increase diversity, and finally the sequences are obtained for recommended synthesis. The present invention proposes an mRNA sequence optimization design method in the environment of 5' UTR and 3' UTR, taking into account the full-length minimum free energy of the mRNA sequence and the codon usage efficiency of the CDS region, which can optimize these two objectives simultaneously, can optimize its stability from the perspective of the overall molecule, and at the same time optimize the expressibility of the CDS region, ensuring the stability and expressibility of the mRNA sequence.
Owner:ZHIYAO TECH

Classification method, classification device, classification system, classification program, and recording medium

PendingCN121368801ABiostatisticsSequence analysisData miningRepresentative sequences
This classification method comprises: a measurement step for measuring the nucleotide sequence of a nucleic acid molecule in a measurement sample as a measurement sequence; and a classification step for classifying the measurement sequence into each of a plurality of groups on the basis of the degree of similarity between a representative sequence set as a base sequence representing each of the groups and the measurement sequence measured in the measurement step. The plurality of groups are obtained by grouping the base sequences of a plurality of nucleic acid molecules according to a specific rule.
Owner:ARKRAY INC

Immunological entity sequence data processing

Immunological entity sequences are clustered by a two-step process based on their similarity, i.e., their ability to react to the same antigen. In the first step, the sequences are converted into numerical vectors, and these numerical vectors are clustered into initial clusters. In the second step, the initial clusters are refined into final clusters, which can be done using the sequence data itself rather than numerical vectors. Each final cluster may be assigned a representative sequence to reduce the dimensionality of the cluster.
Owner:OMNISCOPE LTD

Genomic genome database and construction method and system thereof

The invention belongs to the technical field of bioinformatics, and discloses a generic genome database and a construction method and system thereof. According to the method, microorganisms or parasites provided by a plurality of public databases and / or own databases are classified, genome sequences of the microorganisms or parasites are screened out by using different screening methods, the genome sequences are further cleaned to remove polluted sequences, and finally representative sequences are screened out from the cleaned genome sequences, so that the microorganisms or parasites are screened out. And summarizing all representative sequences to complete the construction of the generic genome database. According to the generic genome database constructed by the method, the accuracy and efficiency of metagenome sequencing are remarkably improved, and the method is suitable for rapid identification of clinical infection pathogens.
Owner:HUGOBIOTECH BEIJING CO LTD

Protein sequence design method based on biomolecular interaction domain enhancement

ActiveCN119601074BBiostatisticsNeural learning methodsSequence designSolvent accessibility
The application discloses a protein sequence design method based on biomolecular interaction domain enhancement, which comprises the following steps: inputting a protein main chain skeleton three-dimensional coordinate information of a size of L*N*3 to be subjected to sequence design; obtaining a protein sequence in contact with a biomolecule and an interaction domain interval; clustering the obtained sequence and taking a representative sequence of each cluster as a training set; extracting three-dimensional structure, secondary structure, solvent accessibility and functional annotation feature representation of each training sample; using a LoRA algorithm to fine-tune the last ten transformer modules of a general multi-modal protein language model ESM3, and giving greater weight to the loss of a mask residue located in the interaction domain interval; and inputting the atomic coordinates of the protein main chain skeleton to be subjected to sequence design into the trained model to obtain a target sequence. On the one hand, the application utilizes multi-modal information of a large amount of proteins; and on the other hand, the application can generate a more robust and reasonable functional protein sequence.
Owner:HUNAN UNIV