Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.
43 results about "Reference genome sequence" patented technology
Filter
Efficacy Topic
Property
Owner
Technical Advancement
Application Domain
Technology Topic
Technology Field Word
Patent Country/Region
Patent Type
Patent Status
Application Year
Inventor
A reference genome is the initial sequence to which all subsequent sequences are ultimately compared, and therefore must be as complete as possible given fiscal and technical constraints. Reference genome sequences result from the de novo sequencing and assembly of a haploid complement of an organism’s genome.
The invention discloses a conservative non-coding element identification method and system based on multi-species genome alignment, and the method comprises the following steps: establishing an index database based on reference genome sequences of multiple species, completing whole genome alignment, and further processing to obtain a high-credibility chain alignment result; the multi-species chain type comparison results are integrated into multi-sequence comparison data in a unified format; on this basis, a neutral evolution model is constructed based on quadruple degenerate sites, and candidate conservative regions are predicted through conservative scoring; in combination with genomeannotation information, a length threshold is set, and a coding region and a UTR region are rejected, so that a high-confidence non-coding conservative element is obtained; and finally, displaying a cross-species conservative distribution diagram of the CNE by utilizing a visual tool. According to the method, the CNE with a potential regulation function can be accurately, efficiently and automatically identified, and technical support is provided for regulation system analysis, functional gene mining and molecular breeding of various organisms.
The invention discloses a compression and decompression method based on generic genome expression, and relates to the technical field of compression and decompression of DNA next-generation sequencing data, in particular to the compression and decompression method based on generic genome expression. The method aims at solving the problems that in the prior art, the capacity of processingpopulationgenetic diversity is insufficient, original sequencing quality information cannot be effectively restored during decompression, and memory occupation is too high during large-scale data processing. Obtaining a to-be-compressed sequencing sequence data file, a reference genome sequence and a thousand-person genome variation sample; obtaining a haplotypelist, a variation list and a haplotype offset list corresponding to each window block; storing the window number, the haplotype number, the haplotype offset, the head and tail unmatched sequences, the current sequence name and the quality score character string into a single compression block; carrying out binding storage; completing the compression processing of the mass fraction; and obtaining each to-be-compressed sequencing sequence based on the result of the compressed part.
The invention provides a probioticscreening method based on a fine tuning DNABERT model and a convolutional neural network. The method comprises the following steps: acquiring a reference genome sequence and a target sequence sample; preprocessing the reference genome sequence and the target sequence sample; segmenting the preprocessed reference gene sequence, and inputting the segmented reference gene sequence into a DNABERT model for fine tuning training; expanding the number of samples of the preprocessed target sequence, segmenting the sequence by using a sliding window, and inputting the segmented sequence into the fine-tuned DNABERT model for decoding to obtain a representation vector; inputting the representation vector into a CNN model for training, sequencing a target sequence to obtain a to-be-detected sequence sample, encoding the to-be-detected sequence sample, inputting the encoded to-be-detected sequence sample into the CNN model, judging whether the strain is the probiotics or not, and completing screening of the probiotics; according to the method, the data enhancement technology of introducing the reverse complementary sequence and the randomly intercepted sub-sequence is adopted, so that the diversity and the quantity of training data are increased, and the overfitting problem is effectively relieved.
The invention relates to the technical field of molecular biology, in particular to a betel nut InDel marker and application thereof in provenance detection. According to the invention, on the basis of the results of re-sequencing of one betel nut germplasm resource and comparative analysis of a reference genome sequence, InDel sites are excavated to develop molecular markers, 46 pairs of InDelmolecular marker primer pairs with high polymorphism are obtained, the variation range of Shannon's diversity index is 0.361-1.713, and the variation range of polymorphism information content is 0.201-0.752. The clustering analysis result shows that the genetic similarity coefficient is 0.36, 211 parts of areca-nut materials are divided into four groups, the InDel marker pair provided by the invention can be effectively used for detecting and analyzing the genetic background of areca-nut germplasm resources, the genetic relationship between the tested areca-nut materials is accurately identified, and the method has the advantages of high specificity, high accuracy and high accuracy. And a foundation is laid for genetic diversity analysis of betel nut germplasm resources and detection of seed fruit and seedling sources.
Systems and methods for determining one or more sequences corresponding to one or more nucleic acid molecules from a plurality of sequence reads are provided herein. In some cases, sequence reads may be obtained and mapped to a reference genomic sequence. Sequence reads may be grouped by one or more features of the sequence reads. The groups of sequence reads may be further grouped to generate one or more subgroups. One or more consensus sequences may be determined corresponding to one or more nucleic acid molecules of a biological sample.
The invention provides a poplar whole genome SNP molecular marker combination, a liquid phasechip prepared from the poplar whole genome SNP molecular marker combination and application of the poplar whole genome SNP molecular marker combination, the whole genome liquid phasechip comprises a probe used for detecting the poplar SNP molecular marker combination, and the SNP molecular marker combination comprises 6549 SNP molecular markers. The physical positions of the 6549 SNP loci are determined based on comparison of a populus simonii haplotype A reference genome sequence, and the specific SNP molecular marker condition is shown in the specification table 1. According to the invention, the probe formed by using a few SNP marker combinations can efficiently and accurately distinguish five branch poplar varieties, the detection cost is reduced, the detection time is shortened, and the SNP marker combinations can provide theoretical and data support for the identification of new poplar varieties, group division and other work, and have important theoretical value and application significance.
The application provides a genetic marker associated with chicken intestinal length in a KCNIP4 gene and an application thereof, and belongs to the fields of animal genetics and breeding and biotechnology.The genetic marker comprises IL_tag1 or IL_tag2; the Ensembl number of the IL_tag1 is rs316532738, corresponds to the sequence of a positive strand of a chromosome No.4 of a chicken reference genome bGalGal1.mat.broiler.GRCg7b published by NCBI, is located in the 1st intron of a gene KCNIP4, and the base at the position is T or C; the Ensembl number of the IL_tag2 is rs316953671.The genetic marker is helpful to genetically improve the intestinal length of a laying hen, is applied to the genetic breeding of a chicken, and is favorable to improving intestinal traits and obtaining a laying hen variety with better nutrient absorption.
The application relates to the technical field of gene variation detection, in particular to a copy number variation detection method based on semi-supervised learning. The method comprises the following steps: obtaining read depth signals and mapping quality signals of each normal window of a reference genome sequence from alignment information of sequencing reads, correcting the read depth signals of all normal windows in terms of GC content bias, adopting a cyclic binary segmentationalgorithm to divide all normal windows into segmented regions with uniform read depth signals, identifying copy number variationbreakpoint positions in combination with a split read strategy, performing normalization processing on the mapping quality signals, performing smoothing and noise reduction processing on the read depth signals, labeling pseudo labels for corresponding segmented regions, performing clustering analysis on all segmented regions through an improved density clustering algorithm, integrating and determining the variation types of abnormal segmented regions, and outputting copy number variation detection results, so that efficient detection of copy number variation is realized, and the accuracy and reliability of the detection results are significantly improved.
The invention discloses a sequence alignment method and a seed search method and device thereof. The seed searching method comprises the steps that an index structure is constructed based on a reference genome sequence S, and the constructed index structure comprises an FM-Index index, a TBWT array, a TOCC array and a CNT array. And taking out the last base from the base fragments to be compared as an initial matching target, and locking the current matching region according to the CNT array and the initial matching target. And starting from the last but one basic group in the basic group fragment, matching two basic groups with S according to the current matching area and the constructed index structure in a reverse order every time until the number of the matched basic groups in the basic group fragment exceeds a preset threshold value, and taking the matched base sequence in the base fragment as a seed sequence. According to the method, the number of times of irregular memory access is reduced, so that the efficiency of a seed search task is improved, and a more efficient genesequence comparison scheme is provided.
The invention discloses a centromere sequence extraction and chromosome classification method and a related device, and belongs to the technical field of bioinformatics. The method comprises the following steps: acquiring a second-generation sequencing sequence data file and a corresponding reference genome sequence file, and segmenting the second-generation sequencing sequence data file into N sub-files according to a set parallel thread count; creating N parallel thread units, extracting feature vectors of the sequencing sequences in the N sub-files by adopting a DNAsequence feature extraction model, inputting the feature vectors into a pre-constructed centromere sequence recognition model, and screening out candidate centromere sequences; and converting each candidate centromere sequence into a feature vector by using the DNAsequence feature extraction model, and inputting the feature vector into a pre-constructed chromosome classification model to obtain a chromosome attribution result. According to the method, by combining data parallel preprocessing, machine learning and a deep learning model, the centromere region sequence can be identified in large-scale massive next-generation sequencing data, and the chromosome to which the centromere region sequence belongs can be further predicted.
In one aspect, the disclosed technology relates to systems and methods for determining a length of a tandem repeat region in each haplotype of a diploid genome. In some embodiments, the method may include obtaining paired-end sequencing reads of the diploid genome; aligning the paired-end sequencing reads to a tandem repeat region in a reference genome sequence; classifying each of the paired-end sequencing reads that overlaps the tandem repeat region based on the alignments into a plurality of classes and counting the number of paired-end sequencing reads in each class; providing a set of hypotheses of a length of the tandem repeat region in each haplotype of the diploid genome; and evaluating which hypothesis has the highest likelihood of generating the counted or observed number of paired-end sequencing reads in the plurality of classes to determine the length of the tandem repeat region in each haplotype of the diploid genome.
Provided is a computer implemented method for mapping transgene integration into n organism by providing a custom genome comprising a reference genome sequence and a transgene sequence, aligning sequence reads of the transgenic organism to the custom genome, identifying candidate junction reads, wherein the candidate junction reads correspond to sequence reads comprising sequence regions that align to the transgene sequence and to the host sequence, categorizing the candidate junction reads into junction clusters, wherein each of the junction clusters comprises a plurality of the candidate junction reads each with an alignment segment terminus within a cluster proximity region of the custom genome, determining a transgene junction based on the junction clusters that possesses at least three of the candidate junction reads within the cluster proximity region of the custom genome and mapping the integration of the transgene into the transgenic organism based on the transgene junctions.
Systems and methods for determining one or more sequences corresponding to one or more nucleic acid molecules from a plurality of sequence reads are provided herein. In some cases, sequence reads may be obtained and mapped to a reference genomic sequence. Sequence reads may be grouped by one or more features of the sequence reads. The groups of sequence reads may be further grouped to generate one or more subgroups. One or more consensus sequences may be determined corresponding to one or more nucleic acid molecules of a biological sample.
The present application relates to the technical field of genesequencing dataprocessing, and particularly relates to a method for calculating chromosome copy number based on software and hardware cooperation, comprising the following steps: S1, separately forming an odd item or an even item of a reference genome sequence R into a subset Rc; S2, indexing the subset Rc through a hash function; S3, mapping a sequencing sequence obtained by gene sequencing back to a position of the reference genome; S4, extracting a seed set So of the subset Qo, finding all candidate positions Lp in an index table SI, and determining an exact position of Q mapping to R; and S5, after completing all sequencing dataprocessing, counting the number of sequencing sequences in each block, and performing normalization. The present application innovatively reduces a standard three-step calculation process completed by three independent programs into a one-step method, that is, a program directly outputs a result without additional hard disk IO reading and writing, thereby shortening the calculation steps.
The application discloses a quality control method and device for clinical detection samples, electronic equipment and a storage medium, and belongs to the technical field of high-throughput sequencing technology. The quality control method for the clinical detection samples comprises the following steps: performing quality control filtering on an original FASTQ data file to obtain a target FASTQ data file; performing alignment on the target FASTQ data file and a reference genome sequence file to obtain a BAM file; determining an alignment rate, rRNA content and globinRNA content according to the BAM file; performing sequencing depth detection by using the BAM file to obtain gene 3' end sequencing depth and gene 5' end sequencing depth; performing consistency checking on multiple sequencing data obtained by using multiple sequencing strategies on the clinical detection samples to obtain a consistency checking result; and comprehensively judging whether the clinical detection samples are qualified. The application can accurately detect the quality of the clinical detection samples and improve the effectiveness of the data.
This invention discloses a gene controlling spike length and grain number per spike in wheat. This gene is located on chromosome 2B using a recombinant inbred line population and high-density SNP markers, within the 657.82–664.40 Mb region of chromosome 2B in the wheat Chinese spring reference genome. A candidate gene, named TaeEF1A, was identified based on the wheat reference genome sequence, RNA-seq, and gene function. Gene editing of TaeEF1A further validated its function in regulating spike length, spikelet number, and grain number per spike. Through gene sequencing and analysis of the gene sequence in a large number of germplasm resources, key variations affecting spike length and grain number per spike were identified in the 3'-UTR of TaeEF1A. Based on this, a detection marker was designed for this gene, which can be used to detect superior allelic variants of the TaeEF1A gene. This marker can be easily, rapidly, and with high throughput applied to marker-assisted breeding to increase spike length, spikelet number, and grain number per spike, thereby increasing yield.
This invention provides a three-generation sequence alignment method based on longest path search. First, a hash index of the reference genome sequence is constructed. Then, each k-mer of the sequence to be aligned is extracted, and its position in the genome is found using the hash index. Each matching k-mer is treated as a node, and a directed acyclic graph of k-mer l-neighborhoods is constructed. Based on the position information of the matching k-mer in the sequence to be aligned, it can be determined whether nodes are connected by edges and their directions. Isolated nodes and small isolated networks are then filtered out. A dynamic scoring strategy is designed to determine the predecessor node of each node, select the highest score from the predecessor nodes, and record the score path. The longest path is selected. The sequence to be aligned and the reference genome can be divided into seed regions and non-seed regions. For non-seed regions, a traditional double sequence alignment method is used to obtain detailed base alignment results, which are finally merged with the seed regions to obtain the alignment results for the entire sequence.
The application discloses a method for identifying duck palm weight and a related molecular marker application, relates to the technical field of molecular marker assisted selection, and a primer pair is used for amplifying a DNA fragment containing a polymorphic site of a duck reference genome IASCAAS_PekinDuck_T2T chromosome 4 from the 5' end 63823029 base position; the duck reference genome IASCAAS_PekinDuck_T2T is a duck reference genome sequence in a GenBankdatabase; the duck is a Zhongxinbaihu meat duck, and the primer pair consists of a DNA molecule shown in SEQ ID No. 2 and a DNA molecule shown in SEQ ID No. 3. The method for identifying the duck palm weight and the related molecular marker application can determine the genotype of the duck by detecting the genomic DNA of the duck, realize early in-vivo screening of the duck palm weight trait, do not need to wait for post-slaughter determination of a phenotype, significantly shorten a breeding cycle, and reduce breeding cost.
The invention relates to a high-speed classification system and a high-speed classification method for special variant DNA fragments of pine trees, in particular to a high-speed classification system and a high-speed classification method for special variant DNA fragments of pine trees based on an FPGA (Field Programmable Gate Array). The invention aims to solve the problems of low calculation speed, high energy consumption and poor real-time performance when a CPU (Central Processing Unit) or a GPU (GraphicsProcessing Unit) is used for analyzing the pine tree variation DNA fragment in the prior art. A data input interface module of the classification system is used for receiving to-be-analyzed DNA sequence data and reference genome sequence data from an external memory or a host and loading the to-be-analyzed DNA sequence data and reference genome sequence data into an on-chip memory; the FPGA processing unit is connected with the data input interface module and the on-chip memory, and is used for carrying out high-speed parallel processing on the DNA sequence data and generating a classification result of variant DNA fragments; and the result output interface module is used for transmitting the classification result to an external host or storage equipment from the FPGA processing unit. The invention belongs to the technical field of bioinformatics and computer hardware acceleration.
The invention discloses a data comparison method and device based on a space transcriptome and a program product. The method comprises the steps of obtaining a reference position bar code; acquiring spatial transcription sequencing data, wherein the spatial transcription sequencing data comprises a sequencing position bar code; based on each reference position bar code and each sequencing position bar code, dividing the spatial transcription sequencing data into a plurality of groups of sequencing data; comparing each group of sequencing data with a reference genome sequence to obtain a comparison result corresponding to each group of sequencing data; and combining the comparison results corresponding to each group of sequencing data to obtain a combined result.
The invention discloses a metagenome classification method and system. The method comprises the following steps: generating l-mer sequences and establishing a mapping table of the l-mer sequences and biological semantic scores; acquiring a set of reference genome sequences; sliding the windows on the reference genome sequence along a preset step length, and intercepting a sequence fragment in each window; performing feature extraction operation on each window; taking 64-bit integer hash values of the features as keys, taking species classification IDs of the reference genome sequences to which the keys belong as values, and storing the keys and the values into a hash table; receiving original sequencing read length data; sliding the windows on the read length along a preset step length, intercepting a sequence fragment in each window, and performing feature extraction operation on each window to obtain a feature query set; taking 64-bit integer hash values of the features as keys, and retrieving in a hash value table to obtain species classification IDs of the features; and based on the species classification ID of each feature, obtaining species classification of the read length. According to the invention, rapid and accurate species classification can be realized.