Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

151 results about "Reference genome" patented technology

A reference genome (also known as a reference assembly) is a digital nucleic acid sequence database, assembled by scientists as a representative example of a species' set of genes. As they are often assembled from the sequencing of DNA from a number of donors, reference genomes do not accurately represent the set of genes of any single person. Instead a reference provides a haploid mosaic of different DNA sequences from each donor. For example, GRCh37, the Genome Reference Consortium human genome (build 37) is derived from thirteen anonymous volunteers from Buffalo, New York. The ABO blood group system differs among humans, but the human reference genome contains only an O allele (although the other alleles are annotated).

A whole genome 20k liquid breeding chip for apostichopus japonicus and application thereof

PendingCN122279058ABiotechnologyGenomics
This invention relates to the fields of genomics, molecular biology, bioinformatics, and genome-wide selection breeding, specifically a 20k liquid-phase breeding chip for the whole genome of *S. esculenta* and its applications. The liquid-phase chip contains background SNPs and functional SNPs located on the *S. esculenta* reference genome; wherein the background SNPs are uniformly distributed within the genome; and the functional SNPs are associated with important economic traits of *S. esculenta*; these important economic traits include one or more of the following: saponin content, polysaccharide content, and heat tolerance. The chip can be applied to the assessment of genetic diversity in *S. esculenta*, identification of germplasm resources and phylogenetic relationships, genome-wide association analysis of important economic traits, and genome-wide selection breeding. This chip has advantages such as high throughput, high region coverage, high locus detection rate, and high flexibility, providing powerful tool support for molecular breeding of *S. esculenta*.
Owner:INST OF OCEANOLOGY - CHINESE ACAD OF SCI

A method for serialization extraction of highly variable exons

PendingCN122290698AInformation densityExon
This invention discloses an efficient RNA data preprocessing method to address the problems of low processing efficiency and low information density in high-throughput sequencing data. Its core steps include: (1) introducing a parallel processing scheme for high-throughput sequence data, rapidly mapping RNA-seq data to a reference genome to generate a BAM file; (2) extracting base sequences and expression levels and storing them as compact PKL format files; (3) extracting all exon position information by parsing the genome annotation file; (4) combining multi-sample expression level data to screen for highly variable exons and constructing a high-information-density feature list based on the sample set; and (5) accurately extracting target sequences from the preprocessed file based on this list. Compared to traditional methods, this innovative approach achieves triple optimization: full-process parallel processing for accelerated computation, high-compression data storage, and adaptive feature selection. Processing speed is increased by 3-5 times, and data volume is reduced by more than 90%, making it suitable for high-throughput RNA-seq data analysis with large sample sizes.
Owner:TIANJIN UNIV

A SNP molecular marker located on chromosome 8 of pigs and associated with intramuscular fat traits in pigs and its application

ActiveCN120330348BMedicineGenetics
This invention belongs to the fields of molecular biology and molecular marker technology, specifically relating to a SNP molecular marker located on chromosome 8 of pigs and associated with the intramuscular fat trait, and its application. The SNP site of this molecular marker on chromosome 8 corresponds to the T>C mutation at positions 43,738,941 on chromosome 8 in the International Swine Reference Genome Version 11.1. By selecting the dominant allele of this SNP, this invention can increase the frequency of the dominant allele generation by generation, thereby increasing intramuscular fat content, improving meat quality, and breeding superior pigs with the aforementioned traits. This contributes to accelerating the progress of pig genetic improvement and effectively improving the economic benefits of pig breeding.
Owner:NORTHWEST A & F UNIV

Systems and methods for detecting fusion genes from sequencing data

In some embodiments, a computer-implemented method of detecting a presence of a predetermined fusion gene in a biological sample is provided. A computing system generates an alignment of a read sequence to a reference genome. The alignment includes a first alignment result and a second alignment result. The computing system determines a breakpoint location indicated by the first alignment result and the second alignment result, distances between coordinates of the breakpoint location and coordinates of one or more expected breakpoint locations associated with the predetermined fusion gene, a gap size value and an overlap size value. In response to determining that the gap size value is less than a gap size value threshold, the overlap size value is less than an overlap size value threshold, and the distances are less than a breakpoint distance threshold, the computing system generates an indication of the presence of the predetermined fusion gene.
Owner:UNIV OF WASHINGTON +1

A molecular marker of wheat stripe rust snp homozygous site based on whole genome sequence and application thereof

ActiveCN119799953BGenetic linkage disequilibriumAllele frequency
The application discloses a kind of molecular markers of wheat stripe rust SNP homozygous site based on whole genome sequence and application thereof.The application is analyzed to 28 strains of global wheat stripe rust genome sequence, compared with reference genome, and more than a million SNP sites are mined out.Sequence depth screening, linkage disequilibrium analysis and heterozygosity detection, identify 1076 SNP homozygous sites.Further based on the frequency of secondary allele, screening of deletion rate and adjacent simple repeat sequence, finally determine 37 core SNP markers.Through primer design to these SNP sites, and add fluorescent linker, develop KASP-SNP molecular marker, which can be used for accurate genotyping identification of wheat stripe rust population.The molecular marker developed based on global wheat stripe rust whole genome is suitable for wheat stripe rust population research in all regions of the world.Based on the development of molecular marker of homozygous site, polymorphism difference caused by heterozygous site can be effectively excluded, to ensure the accuracy of detection site.The set of KASP-SNP molecular marker has high polymorphism, good repeatability and high detection efficiency, and can be widely applied to genetic research of wheat stripe rust population.
Owner:NORTHWEST A & F UNIV

Splitting method and splitting device for single-cell pooled sample sequencing data

ActiveCN117079714BRealize traceabilityRealize processProteomicsGenomicsCell trappingSingle cell suspension
This invention provides a method and apparatus for splitting single-cell mixed sample sequencing data, relating to the field of biotechnology. The splitting method includes: capturing and sequencing single-cell suspensions using a single-cell platform; then performing reference genome alignment, cell identification, and gene expression level quantification on the sequencing data using Cellranger; splitting the cell data identified in step a into two groups of cell data for different sexes based on SNP locus information from the 1000 Genomes Project; and distinguishing the two groups of cell data from male or female samples based on the proportion of sex-specific genes expressed in the two groups of cell data. This splitting method eliminates the need for additional experimental operations such as protein labeling and genome sequencing, and can provide accurate and reliable data splitting even when individual SNP information is unavailable.
Owner:TIANJIN NUOHEZHIYUAN BIO-INFORMATION TECH CO LTD

A molecular marker closely linked to soybean seed protein content QTLqPro14 and its application

ActiveCN120905431BBiotechnologyWhole Genome Association Analysis
The application belongs to the technical field of molecular biology and genetic breeding, and discloses a molecular marker closely linked to a soybean seed protein content QTL qPro14 and application. A seed protein content QTL qPro14 is identified on chromosome 14 of soybean by using whole genome association analysis. A significant SNP associated with the QTL is located at the 1,360,542th base of chromosome 14 of a reference genome Glycine_max_v2.1, and can explain 0.32% of phenotypic variation. A PARMS marker developed by using the SNP is clear in genotyping and simple in operation, and is suitable for molecular breeding of soybean seed protein content.
Owner:INST OF FOOD CROPS HUBEI ACAD OF AGRI SCI +1

Data compression and decompression methods, apparatus, and storage system

PCT designated stageWO2026138024A1Data compressionRadiology
The present application provides data compression and decompression methods, an apparatus, and a storage system. The data compression method comprises: acquiring FastQ data and BAM data, the FastQ data comprising a plurality of sequences, the BAM data comprising a plurality of alignment items corresponding to the plurality of sequences, and the plurality of alignment items being obtained by aligning the plurality of sequences separately with a reference genome; determining a first alignment item corresponding to a first sequence among the plurality of sequences, the first alignment item being one of the plurality of alignment items, the first alignment item comprising an alignment result of the first sequence, and the alignment result of the first sequence comprising position information of the first sequence in the reference genome; and on the basis of the alignment result of the first sequence comprised in the first alignment item, performing reference compression on the first sequence, so as to obtain compressed data of the FastQ data. The method can improve the compression rate of data and also improve the compression performance of the data.
Owner:HUAWEI TECH CO LTD

A method and system for determining the chromosome base number of macrobrachium rosenbergii based on multi-omics joint analysis

ActiveCN121811964BClimate change adaptationBiostatisticsGenomicsMulti omics
The present application belongs to the field of biotechnology and genomics, and particularly relates to a method and system for determining the chromosome base number of Macrobrachium rosenbergii based on multi-omics joint analysis. The method obtains de novo assembly sequencing data and Hi-C sequencing data, generates a chromosome-level candidate assembly without presetting the number of chromosomes by using Hi-C interaction signals after primary assembly, and performs whole-genome collinearity alignment with no less than two published reference genomes; in combination with quality constraints such as collinearity continuity, Hi-C boundary characteristics and BUSCO / LAI, the candidate chromosome boundary is comprehensively judged and iteratively converged, and finally the chromosome base number and reviewable evidence chain are output. The embodiments show that the present application can identify and correct the number redundancy caused by over-splitting of the reference genome, determine the base number of Macrobrachium rosenbergii as n=57 (2n=114), and improve the objectivity and reliability of base number determination.
Owner:ZHEJIANG DANSHUI FISHERY RESEARCH INSTITUTE (ZHEJIANG DANSHUI FISHERY ENVIRONMENTAL MONITORING STATION)

Web-based visualization analysis method and system for tumor gene mutation detection by whole exome sequencing

The application relates to the technical field of gene sequencing data processing and bioinformation analysis, in particular to a Web-based whole-exome sequencing tumor gene mutation detection visual analysis method and system. The system collects user sequencing data and a reference genome version through a Web interactive interface; a program is called to perform quality control cleaning and evaluation on the data, and a visual report is generated; sequence alignment is completed based on the reference genome, and a variation site is identified through algorithm iteration; biological annotation of the variation is combined with a database, a candidate pathogenic mutation set is screened out in multiple levels according to a strategy, the candidate set is projected to a visual interface, a site state is confirmed or removed in response to a manual checking instruction, and a final gene mutation detection report is generated. The application greatly simplifies the whole-exome sequencing data processing procedure, makes it easy for clinical doctors or researchers without bioinformation background to start, and improves the popularization rate and work efficiency of tumor gene detection work.
Owner:DELIFU (XIAMEN) BIOTECHNOLOGY CO LTD

A SSR molecular marker primer related to leek dormancy and a detection method and application thereof

ActiveCN120624700Brapid identificationaccurate identificationBiotechnologyGermplasm
The application discloses a SSR molecular marker primer related to Chinese chives dormancy and a detection method and application thereof. In view of the problems that traditional Chinese chives dormancy trait identification relies on field morphological observation for 4 months, is inefficient and lacks special molecular markers, 7 specific SSR primers (SSR005 / 170 / 249 / 325 / 423 / 478 / 610) are developed based on cross-species genome resequencing (taking Chinese onion as a reference genome, and the alignment rate is 18.76%-18.89%) for the first time. Through the establishment of a matching detection system: CTAB method for rapid DNA extraction, optimization of the PCR system (20 muL contains 30 ng of template DNA, 0.4 U of Taq enzyme, 0.2 muM of primer, 0.2 mM of Mg 2+ 2.0 mM of dNTPs, 57 DEG C annealing), 8% polyacrylamide gel electrophoresis combined with silver staining development, the dormancy trait can be accurately identified within 7 days. For example, the SSR423 primer has no band at 177 bp, and the non-dormant / dormant germplasm cannot be distinguished. The detection results of 52 germplasms (such as dormancy germplasm 21-1 and non-dormancy germplasm 22-7) are consistent with the field phenotypes at a rate of 100%.
Owner:河北省农林科学院经济作物研究所

Determining microsatellite instability (MSI) status based on fragmentomic features

PCT designated stageWO2026111789A1Microbiological testing/measurementProteomicsMicrosatellite StableDNA
Techniques for identifying a microsatellite instability (MSI) status of a sample are described. An example method includes identifying sequence read data indicating sequences of DNA fragments of a sample obtained from a subject with cancer and determining, based on the sequence read data, endpoint positions of the DNA fragments with respect to a reference genome. Input features are determined based on the endpoint positions of the DNA fragments with respect to the reference genome. Using a classifier, and based on the input features, an MSI status (e.g., MSI-High, MSI-Low, MSI-Stable, etc.) of a sample is determined.
Owner:FOUNDATION MEDICINE INC

A SNP marker related to chicken age at first egg and its application

The application discloses a SNP marker related to a chicken age of first egg laying character and application thereof, and belongs to the technical field of molecular genetics and gene breeding. The SNP marker is located at the position of 3298230 bp of chicken chromosome 27, the reference genome version is GRCg7b, has G / A polymorphism, and the individual with the AA genotype at the position has the early age of first egg laying character compared with other genotypes. The SNP marker can be applied to the breeding of excellent breeding hens, can significantly shorten the age of first egg laying of a population, and has great economic application value.
Owner:FUJIAN SHENGZE BIOLOGICAL TECH DEV CO LTD +1

Specific snp site primer combination for identifying taihe and wuji breeds and application thereof

The application discloses a specific SNP site primer combination for identifying Taihe black-bone chicken varieties and application thereof. A chicken pan-genome containing a chicken reference genome GRCg7b and 2.49 Gb pan-sequence is taken as reference, WGS data of Taihe black-bone chicken and related chicken varieties is analyzed and screened to obtain two specific SNP sites, namely novel_th1 (G / T, dominant allele G) and novel_th2 (C / T, dominant allele C), wherein the T allele is medium to high frequency in most non-Taihe black-bone chicken varieties / strains. A FAM / VIC double-labeled TaqMan probe and a matching primer are designed based on the two sites to construct a real-time fluorescent quantitative PCR detection kit for joint typing of the two sites, and an amplification system and a Ct value judgment standard are established, so that 96 varieties / strains such as Taihe black-bone chicken, Zhushi chicken, other black-bone chicken, local chicken, white-feathered broiler and egg chicken can be quickly and accurately identified. The method has the advantages of high specificity, high sensitivity, simple operation and low cost, and is suitable for identification of authenticity of Taihe black-bone chicken and products thereof, purity monitoring and molecular-assisted breeding.
Owner:ZHEJIANG UNIV

Sweet potato snp molecular marker combination, snp chip and application thereof

This invention discloses a combination of SNP molecular markers for sweet potato, an SNP chip, and their applications, relating to the fields of plant biotechnology and plant molecular breeding. The chip contains 16,730 SNP loci located on chromosome 15 of the sweet potato reference genome "Y22". This sweet potato 16K liquid-phase SNP chip, SweetpotatoGBTS16K, is applied to genotyping, variety identification, gene mapping, and genome-wide association analysis of sweet potato varieties. The SNP loci on the liquid-phase chip of this invention were screened from large-scale sweet potato genome resequencing data, totaling 16,730 SNP loci, encompassing associated loci for major agronomic traits such as yield, quality, and resistance in sweet potato, and has broad application prospects in multiple fields of sweet potato breeding.
Owner:CROP RES INST GUANGDONG ACAD OF AGRI SCI

A SNP molecular marker combination for predicting the tryptophan content of fresh corn kernels and application thereof

ActiveCN121249959BTryptophanGenotype
The application provides a SNP molecular marker combination for predicting the content of tryptophan in fresh corn kernels and application thereof, and belongs to the technical field of selection and breeding of fresh corn. The SNP molecular marker combination provided by the application is composed of 20000 SNP loci located on the B73 reference genome version v5. The SNP molecular marker combination provided by the application can be used for molecular marker assisted breeding of high-quality protein (high tryptophan content) fresh corn, and can also be used for multi-traits breeding of fresh sweet corn. The method for predicting the content of tryptophan in fresh corn kernels provided by the application only needs to detect the genotype of breeding materials, and can predict the content of high-quality protein (tryptophan) in fresh corn kernels, and has the advantages of clear selection target and being not affected by the environment. Meanwhile, the prediction of the content of high-quality protein (tryptophan) in kernels is mainly carried out at the harvesting period of fresh corn, and early detection can be realized.
Owner:SHANGHAI ACAD OF AGRI SCI

Chicken weight-related structural variation molecular marker and application thereof

PendingCN122357743AChromosome localisationChromosome 12
This application belongs to the field of molecular biological breeding and provides molecular markers for chicken weight-related structural variations and their applications. The molecular markers for chicken weight-related structural variations are: SV1 located at position 29007968 on chromosome 3, based on the chicken reference genome GRCg7b, with a reference allele of SEQ ID NO.1 and a substitute allele of A; or SV2 located at position 65895344 on chromosome 1, with a reference allele of G and a substitute allele of SEQ ID NO.2; SV3 located at position 1199450 on chromosome 12, with a reference allele of C and a substitute allele of SEQ ID NO.3; or SV4 located at position 37121335 on chromosome Z, with a reference allele of C and a substitute allele of SEQ ID NO.4. Experiments and verification have demonstrated that the above markers are significantly correlated with chicken weight, providing new molecular marker resources for the genetic improvement of broiler weight traits.
Owner:CHINA AGRI UNIV

A method for rapid detection of conventional and hybrid varieties of industrial chili peppers and its application

PendingCN122081467AAccurately distinguish common speciesAccurately distinguish hybridsMicrobiological testing/measurementProteomicsBiotechnologyCapsicum chinense
This invention discloses a method for rapidly detecting conventional and hybrid varieties of industrial chili peppers and its application, relating to the field of molecular breeding technology. The method includes: extracting DNA from fresh leaves of the industrial chili pepper to be tested and performing whole-genome resequencing at a sequencing depth of 8×–12×; aligning the sequencing data to the chili pepper reference genome (Capsicum chinense cv. 'PBC932') to screen for high-quality heterozygous SNP loci; using a 1 Mb sliding window to statistically analyze the distribution density of heterozygous SNPs across the entire genome; and determining the variety type based on density thresholds: an average density <300 SNPs / Mb indicates a conventional variety, and >2000 SNPs / Mb indicates a hybrid variety. This invention eliminates the need for phenotypic observation in subsequent generations, requiring only leaves from current-generation plants to rapidly and accurately distinguish conventional and hybrid varieties at the molecular level. It features a short detection cycle, objective and reliable results, and is not limited by the number of molecular markers, providing strong technical support for efficient identification and early screening of industrial chili pepper germplasm resources.
Owner:INST OF HORTICULTURAL CROPS YUNNAN ACAD OF AGRI SCI

Relevant pdgf gene snp molecular marker of lambing number of experienced tianhua mutton sheep, screening method and application thereof

PendingCN122326765APhysiologyGenome resequencing
The application discloses a screening method and application of a SNP molecular marker of a PDGF gene related to the number of lambing of a Tianhua mutton sheep flock. F ST The signal is detected, and a molecular marker affecting the number of lambing of the Tianhua mutton sheep flock is screened in combination with gene annotation and site annotation. The molecular marker is located at 3850004 bp of a PDGFD gene of a chromosome 15 of a sheep reference genome (ARS-UI_Ramb_v2.0), and the base mutation is T or C. The molecular marker has a significant influence on the number of lambing of the Tianhua mutton sheep flock, can be applied to the molecular marker for detecting the number of lambing of the Tianhua mutton sheep flock, accelerates the breeding process of the Tianhua mutton sheep, and has the advantages of simple operation, high speed, high sensitivity and the like.
Owner:LANZHOU UNIV

Methods and systems for phased genome assembly

PCT designated stageWO2026136378A1ProteomicsGenomicsHaplotypeBioinformatics
Provided herein are methods of assembling phased genomes. The methods may comprise using a reference genome panel. The methods may comprises generating a set of k-mers corresponding to nucleic acid variants in a reference genome. Sequencing may be performed on a subject to identify long range linkage information. Sequence reads may be analyzed for the presence of k-mers, and genome may be assembled via determination of an order of k-mers.
Owner:DOVETAIL GENOMICS LLC

Methods for identifying buffalo breeds and methods and apparatus for constructing predictive models.

PendingCN122135775ABiostatisticsProteomicsWater buffaloGenetic stock
This invention relates to the field of buffalo breed identification, specifically to a method for buffalo breed identification and a method and apparatus for constructing a prediction model. The prediction model construction method includes: acquiring second-generation NGS sequencing data of a buffalo population; comparing the sequencing data with a reference genome and performing variant detection to obtain variant site data for each buffalo; merging and filtering the variant site data of all buffaloes to obtain a variant site set; and training the buffalo breed prediction model using the variant site set data to obtain the prediction model for the buffalo breed. The prediction model of this invention achieves high accuracy in buffalo breed identification.
Owner:GUANGXI ZHUANG AUTONOMOUS REGION BUFFALO INST

System and Method for BAM-Assisted Compression of FASTQ Files

This invention presents a system and method for compressing unaligned genomic sequences in a FASTQ file using alignment information from corresponding in a BAM file and a reference genome. This is useful in cases where both FASTQ and BAM data are available, but the user wishes to store the FASTQ file for the long term while discarding the BAM. Using this invention the BAM file can be used to improve the speed and efficiency of compressing the FASTQ ahead of the BAM file's removal. Key modules include a “populator” for indexing alignments from BAM files, a “consumer” for identifying and compressing matching FASTQ reads, and a “reconstructor” for uncompressing reads. The invention addresses challenges such as differences in read sequences between FASTQ and BAM due to trimming. The invention significantly enhances compression efficiency, reducing storage and transmission requirements while preserving the ability to reconstruct original FASTQ files.
Owner:LAN DIVON MORDECHAI

A SNP molecular marker related to immune characteristics of yaks and application thereof

ActiveCN119876421BMarker-assisted selectionGenetics
This invention relates to the field of biological detection technology, and more particularly to a SNP molecular marker related to the immune characteristics of yaks and its application. The invention provides an SNP molecular marker for detecting the immunoglobulin content in yaks. The SNP molecular marker is located at position 107100737 on chromosome 2 of the yak reference genome LU_Bosgru_v3.0, with a mutated base of A or G. Based on the molecular marker alleles, yaks exhibit three genotypes: AA, AG, or GG. This invention provides a new SNP molecular marker resource for marker-assisted selection of yak immune traits for non-diagnostic purposes, and provides a basis for yak assisted breeding.
Owner:GANSU AGRI UNIV +1

A method and system for rapid screening analysis of homologous genes

The application discloses a method and system for rapid screening and analysis of homologous genes, and relates to the technical field of biological information and molecular breeding. The method comprises the following steps: obtaining multi-source input files; the multi-source input files comprise a reference genome file, a sample genome file, a reference annotation file and a target gene list file; pre-processing the multi-source input files to obtain pre-processed multi-source input files; the pre-processing comprises standardization processing, integrity detection, format consistency detection, repeated record detection and sequence file information improvement; performing sequence alignment, variation site extraction and variation marking on the pre-processed multi-source input files to obtain gene screening analysis results; and performing visual processing and data management on the gene screening analysis results. The application can reduce manual intervention, improve the efficiency and consistency of data analysis, and realize rapid standardized processing and intuitive result display of genomic variation information.
Owner:HUAZHONG AGRI UNIV +2

Methods for classifying, detecting and treating biological diseases

PendingAU2024399715A1DiseaseData set
The current disclosure provides for methods and compositions for classifying subjects having different biological states. The disclosure describes a method comprising: filtering sequence data obtained from a sample from a subject based on long non-coding RNA (lncRNA) and / or pseudogene RNA (pgRNA), and / or the reference genome; determining a biological state classification of the subject by providing the filtered sequence data to one or more machine learning classifiers as input, wherein the one or more machine learning classifiers is trained to output biological state classifications based on filtered sequence data of a training data set.
Owner:FBB BIOMED INC

Molecular markers for improving feed utilization efficiency in chickens and their applications

PendingCN122357736ABiotechnologyFeed conversion ratio
This invention discloses a SNP marker associated with chicken feed conversion ratio and its application, belonging to the fields of molecular genetics and gene breeding technology. The SNP marker is located at position 35867073 bp on chicken chromosome 7, with a reference genome version of GRCg7b, and exhibits T / A polymorphism. Individuals with the AA genotype at this position have lower feed conversion ratios than those with other genotypes. The SNP molecular marker of this invention can be applied to the breeding of superior chicken breeds, significantly improving feed utilization efficiency and possessing high economic application value.
Owner:CHINA AGRI UNIV

Information processing device

PendingJP2026100866AProteomicsGenomicsInformation processingLower limit
Numerous molecular data points, consisting of numerical values ​​representing label locations, are aligned to a reference genome to detect structural variations. [Solution] When the size of the label interval of a reference nucleic acid sequence or the label interval of a target nucleic acid sequence is less than or equal to a lower limit, the information processing device according to the present invention uses a correction function composed of a polynomial that takes the interval as an argument to correct the interval so that it becomes a value greater than the lower limit.
Owner:HITACHI HIGH TECH CORP

A cnv detection method based on an isolation forest algorithm, medium and equipment

ActiveCN118230825BGenomic sequencingVariable window
The application relates to the technical field of gene sequencing, in particular to a CNV detection method based on an isolated forest algorithm, a medium and equipment. The method comprises the following steps: obtaining genomic sequencing data of a CNV event introduced in a reference genome from a Bam file of single-cell DNA sequencing; adopting a variable window strategy to divide the reference genome into windows, and adjusting the windows through a consistent read quantity to obtain a segmentation site information file; calculating RD signal values in each window according to the segmentation site information file and the genomic sequencing data, and extracting PEM signals to obtain PEM signal values; performing multi-feature calculation and analysis on the RD signal values and the PEM signal values based on the isolated forest algorithm; and identifying CNV events in each window according to an analysis result. The application can effectively identify CNV events in single-cell DNA sequencing data, not only overcomes the limitations of traditional methods in single-cell DNA sequencing data, but also improves the accuracy and reliability of CNV event detection.
Owner:XI AN JIAOTONG UNIV

Determining an immune signature based on fragmentomic features

Techniques for identifying an immune signature of a subject are described. Embodiments can include identifying sequence read data indicating sequences of DNA fragments of a sample obtained from a subject; determining, based on the sequence read data, endpoint positions of the DNA fragments with respect to a reference genome; determining input features based on the endpoint positions of the DNA fragments with respect to the reference genome; and determining, using a classifier and based on the input features, an immune signature of the subject, wherein, in certain instances, the immune signature provides a prognostic, diagnostic, or therapeutic indicator for the subject.
Owner:FOUNDATION MEDICINE INC

A reference-based gene compression method

ActiveCN114520025BBioinformaticsInstrumentsAlgorithmCompression method
The application discloses a reference-based gene compression method, and steps of the method comprise the following steps: S1, obtaining a direct compression cost cost0 of base row data to be compressed currently; S2, comparing the direct compression cost cost0 with a reference genome-based compression cost cost2, if the cost0 is less than or equal to the cost2, the current base row is compressed by a direct compression method; otherwise, the current base row is compressed by a reference genome-based compression method. The application has the advantages of simple principle, simple operation, small overall cost, high efficiency and the like.
Owner:GENETALKS BIO TECH (CHANGSHA) CO LTD