Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

57 results about "Reference genome sequence" patented technology

A reference genome is the initial sequence to which all subsequent sequences are ultimately compared, and therefore must be as complete as possible given fiscal and technical constraints. Reference genome sequences result from the de novo sequencing and assembly of a haploid complement of an organism’s genome.

Use of SNP site combination of litopenaeus vannamei in trait evaluation of mixed-family culture, and probe and kit

The use of an SNP site combination of Litopenaeus vannamei in trait evaluation of mixed-family culture, a probe and a kit. The SNP site combination of Litopenaeus vannamei comprises 1125 SNP sites, wherein the physical positions are determined on the basis of alignment with a reference genome sequence of Litopenaeus vannamei. Also disclosed are a molecular probe for capturing the SNP site combination and a kit comprising the molecular probe. By means of using the SNP site combination for paternity identification, the present invention can achieve the mixed culture of individuals of different families in the early stage, so that the common environmental effect caused by separate family culture in the early stage is effectively reduced, and the accuracy of a trait evaluation result of mixed culture is improved; moreover, the genetic relationship between individuals can be accurately calculated. Compared with physical marks, the present invention avoids the phenomena of falling and dislocation of marks on individuals and human identification errors, thus obtaining accurate pedigrees.
Owner:YELLOW SEA FISHERIES RES INST CHINESE ACAD OF FISHERIES SCI

Pathogen interpretation method, equipment and medium

The invention provides a pathogen interpretation method and device and a medium, and the method comprises the steps: carrying out the data preprocessing of received initial sequencing data, and obtaining first sequence data; comparing the first sequence data with a reference genome sequence in a preset pathogen database to obtain a comparison result; according to a comparison result, obtaining a read parameter set of the initial sequencing data; performing pathogen interpretation processing on the read parameter set to obtain a target pathogen type corresponding to the initial sequencing data; performing drug-resistant gene detection on the target pathogen type, and judging whether a target drug-resistant gene corresponding to the target pathogen exists or not; if the target drug-resistant gene exists, analyzing the target drug-resistant gene to obtain drug-resistant data; and generating a target report according to the target pathogen type, the target drug-resistant gene and the drug-resistant data. According to the invention, the type of the infected pathogen can be rapidly and accurately interpreted, and the drug resistance condition of the infected pathogen can be accurately analyzed, so as to assist clinicians in formulating a reasonable target report.
Owner:JILIN JINYU MEDICAL SCI INSPECTION CO LTD

Method for identifying weight of duck webs and related molecular marker application thereof

The invention discloses a method for identifying duck palm weight and related molecular marker application thereof, and relates to the technical field of molecular marker-assisted selection, a primer pair is used for amplifying a DNA fragment containing a 63823029th base polymorphic site from the 5'terminal on a fourth chromosome of a duck reference genome IASAASPekinDuckT2T, and the primer pair is used for amplifying a DNA fragment containing the 63823029th base polymorphic site from the 5 '-terminal on the fourth chromosome of the duck reference genome IASAASPekinDuckT2T. The duck reference genome IAASAASPekinDuckT2T is a duck reference genome sequence in a GenBank database, and the duck reference genome sequence is a duck reference genome sequence in the GenBank database; the duck is a Chinese and new white feather meat duck, and the primer pair consists of a DNA (deoxyribonucleic acid) molecule as shown in SEQ ID No.2 and a DNA molecule as shown in SEQ ID No.3. According to the method for identifying the duck foot weight and the application of the related molecular marker, the genotype of the duck can be judged by detecting the genome DNA of the duck, early living screening of the duck foot weight character is achieved, phenotype determination after slaughtering is not needed, the breeding period is remarkably shortened, and the breeding cost is reduced.
Owner:INSTITUTE OF ANIMAL SCIENCES OF CHINESE ACADEMY OF AGRICULTURAL SCIENCES +1

Conservative non-coding element identification method and system based on multi-species genome comparison

The invention discloses a conservative non-coding element identification method and system based on multi-species genome alignment, and the method comprises the following steps: establishing an index database based on reference genome sequences of multiple species, completing whole genome alignment, and further processing to obtain a high-credibility chain alignment result; the multi-species chain type comparison results are integrated into multi-sequence comparison data in a unified format; on this basis, a neutral evolution model is constructed based on quadruple degenerate sites, and candidate conservative regions are predicted through conservative scoring; in combination with genome annotation information, a length threshold is set, and a coding region and a UTR region are rejected, so that a high-confidence non-coding conservative element is obtained; and finally, displaying a cross-species conservative distribution diagram of the CNE by utilizing a visual tool. According to the method, the CNE with a potential regulation function can be accurately, efficiently and automatically identified, and technical support is provided for regulation system analysis, functional gene mining and molecular breeding of various organisms.
Owner:WUHAN FRASERGEN CO LTD

Compression and decompression method based on generic genome representation

The invention discloses a compression and decompression method based on generic genome expression, and relates to the technical field of compression and decompression of DNA next-generation sequencing data, in particular to the compression and decompression method based on generic genome expression. The method aims at solving the problems that in the prior art, the capacity of processing population genetic diversity is insufficient, original sequencing quality information cannot be effectively restored during decompression, and memory occupation is too high during large-scale data processing. Obtaining a to-be-compressed sequencing sequence data file, a reference genome sequence and a thousand-person genome variation sample; obtaining a haplotype list, a variation list and a haplotype offset list corresponding to each window block; storing the window number, the haplotype number, the haplotype offset, the head and tail unmatched sequences, the current sequence name and the quality score character string into a single compression block; carrying out binding storage; completing the compression processing of the mass fraction; and obtaining each to-be-compressed sequencing sequence based on the result of the compressed part.
Owner:HARBIN INST OF TECH

Probiotic screening method based on fine tuning DNABERT model and convolutional neural network

The invention provides a probiotic screening method based on a fine tuning DNABERT model and a convolutional neural network. The method comprises the following steps: acquiring a reference genome sequence and a target sequence sample; preprocessing the reference genome sequence and the target sequence sample; segmenting the preprocessed reference gene sequence, and inputting the segmented reference gene sequence into a DNABERT model for fine tuning training; expanding the number of samples of the preprocessed target sequence, segmenting the sequence by using a sliding window, and inputting the segmented sequence into the fine-tuned DNABERT model for decoding to obtain a representation vector; inputting the representation vector into a CNN model for training, sequencing a target sequence to obtain a to-be-detected sequence sample, encoding the to-be-detected sequence sample, inputting the encoded to-be-detected sequence sample into the CNN model, judging whether the strain is the probiotics or not, and completing screening of the probiotics; according to the method, the data enhancement technology of introducing the reverse complementary sequence and the randomly intercepted sub-sequence is adopted, so that the diversity and the quantity of training data are increased, and the overfitting problem is effectively relieved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

SNP (Single Nucleotide Polymorphism) molecular site, detection primer thereof and application of SNP molecular site in predicting egg laying character of chicken

The invention discloses an SNP (Single Nucleotide Polymorphism) molecular site, a detection primer thereof and application of the SNP molecular site to prediction of egg laying traits of chickens. SNP (Single Nucleotide Polymorphism) typing data comparative analysis of chicken varieties with different egg laying performances is utilized, correlation analysis is performed by combining phenotypic data related to egg laying traits, finally, SNP marker loci remarkably related to the egg laying number of the chicken at the age of 300 days are identified and obtained, the SNP marker loci correspond to chromosomes 61, 577 and 523bp of chicken reference genome GRC7b version sequence information 1 published in an Enmbl website, a basic group is A or G, and an SNP number is rs740376220; the invention further provides a detection primer for detecting the genotype of the SNP molecular site. By adopting the SNP marker site and the detection primer thereof provided by the invention, the egg laying traits of the chickens can be accurately predicted, a technical support can be provided for molecular marker-assisted selective breeding of the egg laying traits of the chickens, the breeding period is shortened, and the breeding efficiency is improved.
Owner:INSTITUTE OF ANIMAL SCIENCES OF CHINESE ACADEMY OF AGRICULTURAL SCIENCES

Areca InDel marker and application thereof in provenance detection

The invention relates to the technical field of molecular biology, in particular to a betel nut InDel marker and application thereof in provenance detection. According to the invention, on the basis of the results of re-sequencing of one betel nut germplasm resource and comparative analysis of a reference genome sequence, InDel sites are excavated to develop molecular markers, 46 pairs of InDel molecular marker primer pairs with high polymorphism are obtained, the variation range of Shannon's diversity index is 0.361-1.713, and the variation range of polymorphism information content is 0.201-0.752. The clustering analysis result shows that the genetic similarity coefficient is 0.36, 211 parts of areca-nut materials are divided into four groups, the InDel marker pair provided by the invention can be effectively used for detecting and analyzing the genetic background of areca-nut germplasm resources, the genetic relationship between the tested areca-nut materials is accurately identified, and the method has the advantages of high specificity, high accuracy and high accuracy. And a foundation is laid for genetic diversity analysis of betel nut germplasm resources and detection of seed fruit and seedling sources.
Owner:COCONUT RES INST OF CHINESE ACAD OF TROPICAL AGRI SCI

Methods and systems for analyzing sequence reads

Systems and methods for determining one or more sequences corresponding to one or more nucleic acid molecules from a plurality of sequence reads are provided herein. In some cases, sequence reads may be obtained and mapped to a reference genomic sequence. Sequence reads may be grouped by one or more features of the sequence reads. The groups of sequence reads may be further grouped to generate one or more subgroups. One or more consensus sequences may be determined corresponding to one or more nucleic acid molecules of a biological sample.
Owner:FORESIGHT DIAGNOSTICS INC

Application of SNP genetic marker affecting chicken body weight at first egg in genetic breeding of laying hens

The application provides application of a SNP genetic marker affecting chicken body weight at the onset of lay in genetic breeding of laying hens, and belongs to the field of animal genetic breeding and biotechnology.The SNP genetic marker affecting chicken body weight at the onset of lay comprises bw1egg_1 and / or bw1egg_2; the Ensembl number of the bw1egg_1 is rs318020581, corresponding to the 76311052th position of the positive strand of chromosome 4 in the chicken reference genome bGalGal1.mat.broiler.GRCg7b sequence published in NCBI, belonging to the 3rd intron of the gene C1QTNF7, and the base at the position is T or G; the Ensembl number of the bw1egg_2 is rs313708699, corresponding to the 74731889th position of the positive strand of chromosome 4 in the chicken reference genome bGalGal1.mat.broiler.GRCg7b sequence published in NCBI, belonging to the 8th intron of the gene SLIT2, and the base at the position is T or C.Both the bw1egg_1 and the bw1egg_2 are helpful to genetically improve the body weight at the onset of lay, and when applied to genetic breeding of chickens, are favorable to improving the body weight at the onset of lay of the laying hens and obtaining a laying hen breed with excellent performance in uniformity of the body weight at the onset of lay.
Owner:JIANGSU INST OF POULTRY SCI

Poplar whole genome SNP (Single Nucleotide Polymorphism) molecular marker combination, liquid phase chip prepared from same and application of liquid phase chip

The invention provides a poplar whole genome SNP molecular marker combination, a liquid phase chip prepared from the poplar whole genome SNP molecular marker combination and application of the poplar whole genome SNP molecular marker combination, the whole genome liquid phase chip comprises a probe used for detecting the poplar SNP molecular marker combination, and the SNP molecular marker combination comprises 6549 SNP molecular markers. The physical positions of the 6549 SNP loci are determined based on comparison of a populus simonii haplotype A reference genome sequence, and the specific SNP molecular marker condition is shown in the specification table 1. According to the invention, the probe formed by using a few SNP marker combinations can efficiently and accurately distinguish five branch poplar varieties, the detection cost is reduced, the detection time is shortened, and the SNP marker combinations can provide theoretical and data support for the identification of new poplar varieties, group division and other work, and have important theoretical value and application significance.
Owner:INST OF FORESTRY CHINESE ACAD OF FORESTRY

Primer combination for amplifying whole genome DNA of sheep embryo cells and application thereof

The present application relates to a primer combination for amplifying whole genome DNA of sheep embryo cells and application thereof, and the primer combination is composed of 20 or 30 primers. The present application designs random primers based on sheep reference genome sequence, optimizes the reaction system based on MDA whole genome amplification technology, so that the whole genome DNA amplification product with high coverage and low mismatch rate can be generated by using 10 sheep embryo cells or as low as pg level genomic DNA, which can be directly adapted to downstream applications such as second-generation sequencing and targeted capture sequencing, and the detection cost is reduced.
Owner:CHINA AGRI UNIV

A method for SNP molecular marker for detecting the pH value of goose meat

The present invention discloses a method for SNP molecular markers for detecting the pH value of goose meat, which relates to the field of SNP molecular marker breeding. The technical solution of the present invention includes the following steps: S1, aligning the whole genome resequencing data of the goose to be detected to the reference genome sequence, combining the pH value data of the goose meat to be detected, and screening candidate SNP sites for the pH value of the goose meat to be detected through genome-wide association analysis; S2, using the technology of flight mass spectrometry to test the distribution frequency of SNP sites in goose individuals and performing association analysis with the pH value trait of goose meat. The identified SNPs can provide theoretical support for the screening of goose meat quality, provide important references for the improvement of goose meat quality and breeding, and provide a new direction for the development of breeding markers.
Owner:CHONGQING ACAD OF ANIMAL SCI

Sample cross contamination evaluation method, apparatus, device, and medium

The present application belongs to the technical field of sample pollution evaluation, and discloses a sample cross-contamination evaluation method, device, equipment and medium, comprising: performing quality control on batch sequencing data, and performing mutation site detection and genetic marker typing according to a human reference genome sequence and a site interval file; according to the genetic marker typing result and the site interval file, the number of heterozygous genotype sites is counted and the sample heterozygosity is calculated, according to the sample heterozygosity and a preset threshold, the sample contamination state is determined; according to the copy number estimation of the sample, the allele frequency is diploid standardized correction; based on the allele frequency in the genetic marker typing result and the mutation site detection result, the cross-contamination relationship between each two samples is analyzed by using identification rules to determine the pollution source; the site for calculating the pollution proportion is selected, and the weighted average algorithm is used to calculate the pollution proportion. The sample pollution degree can be accurately estimated and the pollution source sample can be identified.
Owner:JINAN JINYU MEDICINE JIANYAN CENT CO LTD

Genetic marker associated with chicken intestinal length in kcnip4 gene and application thereof

The application provides a genetic marker associated with chicken intestinal length in a KCNIP4 gene and an application thereof, and belongs to the fields of animal genetics and breeding and biotechnology.The genetic marker comprises IL_tag1 or IL_tag2; the Ensembl number of the IL_tag1 is rs316532738, corresponds to the sequence of a positive strand of a chromosome No.4 of a chicken reference genome bGalGal1.mat.broiler.GRCg7b published by NCBI, is located in the 1st intron of a gene KCNIP4, and the base at the position is T or C; the Ensembl number of the IL_tag2 is rs316953671.The genetic marker is helpful to genetically improve the intestinal length of a laying hen, is applied to the genetic breeding of a chicken, and is favorable to improving intestinal traits and obtaining a laying hen variety with better nutrient absorption.
Owner:JIANGSU INST OF POULTRY SCI

A copy number variation detection method based on semi-supervised learning

The application relates to the technical field of gene variation detection, in particular to a copy number variation detection method based on semi-supervised learning. The method comprises the following steps: obtaining read depth signals and mapping quality signals of each normal window of a reference genome sequence from alignment information of sequencing reads, correcting the read depth signals of all normal windows in terms of GC content bias, adopting a cyclic binary segmentation algorithm to divide all normal windows into segmented regions with uniform read depth signals, identifying copy number variation breakpoint positions in combination with a split read strategy, performing normalization processing on the mapping quality signals, performing smoothing and noise reduction processing on the read depth signals, labeling pseudo labels for corresponding segmented regions, performing clustering analysis on all segmented regions through an improved density clustering algorithm, integrating and determining the variation types of abnormal segmented regions, and outputting copy number variation detection results, so that efficient detection of copy number variation is realized, and the accuracy and reliability of the detection results are significantly improved.
Owner:深圳立专志华科技有限公司

Haplotype identification marker of fertility restorer gene rf4 in rice

The application belongs to the field of plant molecular breeding, and specifically discloses haplotype identification markers of a rice fertility restoration gene Rf4 Based on the high-quality rice reference genome sequences of 33 different varieties published by Professor Qin Peng's team of Sichuan Agricultural University, the published gene sequences or functional gene sites are compared to distinguish the varieties containing target genes, and then the bioinformatics method is used to quickly develop a haplotype identification marker combination of a rice fertility restoration gene Rf4 .
Owner:WUHAN GREENFAFA INST OF NOVEL GENECHIP R&D CO LTD +2

Sequence alignment method and seed search method and device thereof

PendingCN121963880AImprove efficiencyThe number of irregular memory accesses is reducedSequence analysisInstrumentsReference genome sequenceAlgorithm
The invention discloses a sequence alignment method and a seed search method and device thereof. The seed searching method comprises the steps that an index structure is constructed based on a reference genome sequence S, and the constructed index structure comprises an FM-Index index, a TBWT array, a TOCC array and a CNT array. And taking out the last base from the base fragments to be compared as an initial matching target, and locking the current matching region according to the CNT array and the initial matching target. And starting from the last but one basic group in the basic group fragment, matching two basic groups with S according to the current matching area and the constructed index structure in a reverse order every time until the number of the matched basic groups in the basic group fragment exceeds a preset threshold value, and taking the matched base sequence in the base fragment as a seed sequence. According to the method, the number of times of irregular memory access is reduced, so that the efficiency of a seed search task is improved, and a more efficient gene sequence comparison scheme is provided.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Method for centromere sequence extraction and chromosome classification and related device

The invention discloses a centromere sequence extraction and chromosome classification method and a related device, and belongs to the technical field of bioinformatics. The method comprises the following steps: acquiring a second-generation sequencing sequence data file and a corresponding reference genome sequence file, and segmenting the second-generation sequencing sequence data file into N sub-files according to a set parallel thread count; creating N parallel thread units, extracting feature vectors of the sequencing sequences in the N sub-files by adopting a DNA sequence feature extraction model, inputting the feature vectors into a pre-constructed centromere sequence recognition model, and screening out candidate centromere sequences; and converting each candidate centromere sequence into a feature vector by using the DNA sequence feature extraction model, and inputting the feature vector into a pre-constructed chromosome classification model to obtain a chromosome attribution result. According to the method, by combining data parallel preprocessing, machine learning and a deep learning model, the centromere region sequence can be identified in large-scale massive next-generation sequencing data, and the chromosome to which the centromere region sequence belongs can be further predicted.
Owner:XI AN JIAOTONG UNIV

Genotyping for tandem repeats

PCT designated stageWO2025250322A1ProteomicsGenomicsReference genome sequenceDiploid genome
In one aspect, the disclosed technology relates to systems and methods for determining a length of a tandem repeat region in each haplotype of a diploid genome. In some embodiments, the method may include obtaining paired-end sequencing reads of the diploid genome; aligning the paired-end sequencing reads to a tandem repeat region in a reference genome sequence; classifying each of the paired-end sequencing reads that overlaps the tandem repeat region based on the alignments into a plurality of classes and counting the number of paired-end sequencing reads in each class; providing a set of hypotheses of a length of the tandem repeat region in each haplotype of the diploid genome; and evaluating which hypothesis has the highest likelihood of generating the counted or observed number of paired-end sequencing reads in the plurality of classes to determine the length of the tandem repeat region in each haplotype of the diploid genome.
Owner:ILLUMINA INC

Methods and systems for analyzing nucleic acid molecules

To provide processes and materials that can demonstrate improved sensitivity, specificity, and / or reliability in the detection of cancer-derived nucleic acids. [Solution] A method for analyzing cell-free nucleic acids from a subject includes: (a) obtaining sequencing data from a plurality of cell-free nucleic acid molecules obtained from or derived from the subject using a computer system; (b) processing the sequencing data using the system to identify one or more cell-free nucleic acid molecules from the plurality of cell-free nucleic acid molecules, each of which comprises a plurality of stepwise variants with respect to a reference genome sequence, and at least about 10% of the cell-free nucleic acid molecules comprising a first stepwise variant of the plurality of stepwise variants and a second stepwise variant of the plurality of stepwise variants separated by at least one nucleotide; and (c) analyzing the cell-free nucleic acid molecules using the system to determine the state of the subject.
Owner:THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV

Copy number variation detection method based on semi-supervised learning

The invention relates to the technical field of gene variation detection, in particular to a copy number variation detection method based on semi-supervised learning. Comprising the following steps: acquiring a read depth signal and a mapping quality signal of each normal window of a reference genome sequence from comparison information of sequencing reads, and performing GC content deviation correction on the read depth signals of all the normal windows; all normal windows are segmented into segmented areas with uniform read depth signals by adopting a cyclic binary segmentation algorithm, and copy number variation breakpoint positions are identified in combination with a read splitting strategy; performing normalization processing on the mapping quality signal; performing smooth noise reduction processing on the read depth signal; labeling pseudo labels for the corresponding segmented areas; according to the method, clustering analysis is carried out on all segmented areas through an improved density clustering algorithm, integration and variation type judgment are carried out on abnormal segmented areas, and a copy number variation detection result is output, so that efficient detection of copy number variation is realized, and the accuracy and reliability of the detection result are remarkably improved.
Owner:深圳立专志华科技有限公司

Gene variation screening method and device and electronic equipment

The invention discloses a gene variation screening method and device and electronic equipment. The method comprises the following steps: mapping a to-be-identified sequencing sequence to a reference genome sequence, and collecting a mapping quality index of the to-be-identified sequencing sequence; identifying a plurality of gene variations in the sequencing sequence to be identified according to the mapping quality index; filtering the plurality of genetic variations according to a preset filtering rule to obtain a to-be-evaluated variation set, the to-be-evaluated variation set comprising remaining genetic variations after filtering the plurality of genetic variations; determining the score of each gene variation in the to-be-evaluated variation set according to a preset scoring rule; and determining the target gene variation in the variation set to be evaluated according to the score. The technical problem that the required target gene variation cannot be determined from a large amount of gene variation of the target object due to the fact that quantitative analysis cannot be performed on the gene variation in related technologies is solved.
Owner:SHENZHEN HUADA YONGSHENG INTELLIGENT TECHNOLOGY CO LTD

Method of mapping transgene integration

Provided is a computer implemented method for mapping transgene integration into n organism by providing a custom genome comprising a reference genome sequence and a transgene sequence, aligning sequence reads of the transgenic organism to the custom genome, identifying candidate junction reads, wherein the candidate junction reads correspond to sequence reads comprising sequence regions that align to the transgene sequence and to the host sequence, categorizing the candidate junction reads into junction clusters, wherein each of the junction clusters comprises a plurality of the candidate junction reads each with an alignment segment terminus within a cluster proximity region of the custom genome, determining a transgene junction based on the junction clusters that possesses at least three of the candidate junction reads within the cluster proximity region of the custom genome and mapping the integration of the transgene into the transgenic organism based on the transgene junctions.
Owner:TACONIC BIOSCIENCES INC

A method for establishing a cervical disease progression prediction model based on low-depth WGS

The present invention discloses a method for establishing a cervical disease progression prediction model based on low-depth whole-genome sequencing technology, which uses cervical scraping or vaginal swabs to collect exfoliated cells, extract DNA, and perform whole-genome sequencing to obtain the original offline data of low-depth whole-genome sequencing of the DNA sample; the above-mentioned original data is subjected to standard quality control, and then compared with the human reference genome sequence and the repetitive sequences are marked; the genome instability index is calculated by the following formula:; In addition, the present invention also relates to a method for constructing a cervical disease prediction model; the present invention is conducive to improving the accuracy and sensitivity of detection, reducing the cost of detection, and improving the accessibility of cervical disease screening and prediction, while contributing to the early diagnosis of cervical disease and the formulation of treatment decisions, and improving the management and treatment effects of patients with cervical disease.
Owner:SHENYOU GENOME RES INST (NANJING) CO LTD

Methods and systems for analyzing sequence reads

Systems and methods for determining one or more sequences corresponding to one or more nucleic acid molecules from a plurality of sequence reads are provided herein. In some cases, sequence reads may be obtained and mapped to a reference genomic sequence. Sequence reads may be grouped by one or more features of the sequence reads. The groups of sequence reads may be further grouped to generate one or more subgroups. One or more consensus sequences may be determined corresponding to one or more nucleic acid molecules of a biological sample.
Owner:FORESIGHT DIAGNOSTICS INC

Systems and methods for secondary analysis of nucleotide sequencing data

Disclosed herein are systems and methods for performing secondary analysis of nucleotide sequencing data in a time-efficient manner. Some embodiments include iteratively performing secondary analysis as sequence reads are generated by a sequencing system. Secondary analysis can include alignment of sequence reads to a reference sequence (e.g., a human reference genome sequence) and use of the alignment to detect differences between a sample and the reference. Secondary analysis can be capable of detecting genetic differences, variant calling and genotyping, identifying single nucleotide polymorphisms (SNPs), small insertions and deletions (indels), and structural changes in DNA, such as copy number variations (CNVs) and chromosomal rearrangements.
Owner:ILLUMINA INC

Method for computing chromosome copy number based on hardware-software collaboration

The present application relates to the technical field of gene sequencing data processing, and particularly relates to a method for calculating chromosome copy number based on software and hardware cooperation, comprising the following steps: S1, separately forming an odd item or an even item of a reference genome sequence R into a subset Rc; S2, indexing the subset Rc through a hash function; S3, mapping a sequencing sequence obtained by gene sequencing back to a position of the reference genome; S4, extracting a seed set So of the subset Qo, finding all candidate positions Lp in an index table SI, and determining an exact position of Q mapping to R; and S5, after completing all sequencing data processing, counting the number of sequencing sequences in each block, and performing normalization. The present application innovatively reduces a standard three-step calculation process completed by three independent programs into a one-step method, that is, a program directly outputs a result without additional hard disk IO reading and writing, thereby shortening the calculation steps.
Owner:NANJING GEZHI GEMONICS CO LTD