Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

6 results about "Repetitive Regions" patented technology

Genome assembling method

PendingCN120727104ASequence analysisInstrumentsGeneticsRepetitive Regions
The invention relates to a genome assembling method, in particular to a genome assembling method. The genome assembly method disclosed by the invention comprises the following steps: 1, independently assembling by utilizing a PacBio HiFi technology based on an Oxford Nanopore platform; 2, taking an assembly result of the HiFi data as a reference; 3, carrying out fusion by using QuickMerge v0.3, and carrying out fusion by using QuickMerge v0.3; 4, implementing multiple rounds of high-precision correction; 5, generating a paf file by using minimap2v2.17, and removing a redundant sequence by using Purge-Dups v1.2. 6 according to the file; 6, three-wheel structure optimization and manual correction are carried out; and 7, completing gap filling of the remaining area. According to the assembling method, the analysis capability of the ultra-long repeated region is remarkably improved, and the accuracy of a base layer surface and the overall quality of a genome are remarkably improved.
Owner:HEILONGJIANG RIVER FISHERY RES INST CHINESE ACADEMY OF FISHERIES SCI

Sequencing data error correction method and device, electronic equipment and storage medium

PendingCN121983117AEfficient error correctionUniversal error correctionBiostatisticsProteomicsEngineeringData error
The invention provides a sequencing data error correction method and device, electronic equipment and a storage medium. The method comprises the steps that long-read-long sequencing data are obtained, and the long-read-long sequencing data comprise tandem repeat areas; coding the tandem repeat region and the upstream and downstream sequencing data of the tandem repeat region by adopting a first coding mode to obtain a base continuity characteristic; encoding the upstream and downstream sequencing data of the tandem repeat region by adopting a second encoding mode to obtain upstream and downstream features; and inputting the base continuity feature and the upstream and downstream features into a pre-trained correction model to perform error correction on a tandem repeat region in the long read length sequencing data, thereby realizing combination of the tandem repeat region in the long read length sequencing data and upstream and downstream sequencing data of the tandem repeat region in a depth model. Error correction is carried out on a long-read-length sequencing sequence, so that efficient and universal sequencing data error correction is achieved, the accuracy of sequencing data is improved, and subsequent assembly performance is improved.
Owner:HANGZHOU HUADA XUFENG TECHNOLOGY CO LTD

Systems and methods for tandem repeat mapping

Systems and methods for mapping a plurality of sequence reads to a genomic region are provided. A plurality of sequence reads mappable to the genomic region are obtained. An initial Markov model for the genomic region is obtained. The initial Markov model comprises at least (i) a first repeat for a first repeat region, (ii) a second repeat for a second repeat region, and (iii) an intermediate region linking the first repeat to the second repeat. The initial Markov model is refined using the plurality of sequence reads, thereby obtaining a refined Markov model. For each respective sequence read in the plurality of sequences, the respective sequence read is used to find a highest probability path through the Markov model. This highest probability path is then used to map the respective sequence read to the genomic region.
Owner:PACIFIC BIOSCIENCES OF CALIFORNIA INC

Aspergillus telomere to telomere genome assembly methods, apparatuses, devices, and storage media

PendingCN122392622AGenomic sequencingContig
The application discloses an aspergillus telomere-to-telomere genome assembly method, device, equipment and storage medium. The method comprises the following steps: using a plurality of sequencing sequence assembly tools to assemble target aspergillus long read genome sequencing data from scratch to obtain a first assembled genome; selecting a first assembled genome meeting a preset condition as an initial assembled genome; integrating other first assembled genomes to fill gaps between repeat regions of the initial assembled genome to obtain a second assembled genome; aligning the long read genome sequencing data to the second assembled genome, identifying abnormal coverage regions and correcting sequences to obtain a third assembled genome; aligning a reference genome to the third assembled genome, connecting and orienting different contigs, and mounting the contigs to chromosomes to obtain a fourth assembled genome; and aligning the long read genome sequencing data and short read sequencing data to the fourth assembled genome for correction to obtain an aspergillus telomere-to-telomere genome assembly result.
Owner:PEKING UNION MEDICAL COLLEGE HOSPITAL +2

Preparation method of chromatin conformation capture SMRT sequencing library

The invention provides a chromatin conformation capture SMRT sequencing library preparation method, which comprises: S100, treating a cell to be detected to obtain a DNA connection product; s200, repairing the DNA connection product and adding an A tail at the 3'tail end to obtain DNA with a cohesive tail end; s300, adding amplification linkers to two ends of the DNA with the cohesive ends, and performing amplification enrichment purification; s400, carrying out PacBio library construction on the DNA subjected to amplification, enrichment and purification; and S500, carrying out Qubit quantification and Femto Pulse fragment analysis on the constructed library, and carrying out sequencing on a PacBio Revio system. According to the method, the HiFi reads with extremely high accuracy can be obtained, compared with a traditional method, a high-repetition region or an allelic region can be analyzed more accurately, and by means of the advantage of long read length, information of simultaneous interaction among multiple chromatin sites can be obtained.
Owner:WUHAN FRASERGEN CO LTD

Identification, construction and application of tandem repeat sequences of plant centromere

This invention discloses the identification, construction, and application of plant centromere tandem repeat sequences. The method includes: obtaining the target plant genome sequence; detecting tandem repeats through prefix sum vectorization to obtain candidate repeat regions and periods; extracting repeat units as candidate monomers and eliminating low-complexity noise; calculating a multi-feature fusion score for candidate monomers, the score including period consistency, genome enrichment, sequence complexity, and GC content shift, and screening based on the score; clustering the screened candidate monomers to obtain monomer subtypes; performing phase correction on monomers within subtypes; and constructing a consensus sequence based on the corrected sequence. This application achieves quantitative evaluation of centromere attribution probability through multi-feature fusion scoring, combined with exhaustive cyclic shift phase correction, to obtain a consensus sequence without relying on external tools or known motifs. Taking the Arabidopsis Col-CEN genome as an example, the optimal period is 178 bp, the column consistency rate is 0.9272, and the alignment consistency with the reference sequence exceeds 98%.
Owner:NANTONG UNIV