A chromosome karyotyping method, device, equipment and medium

CN120998296BActive Publication Date: 2026-08-21SUZHOU SAIFU MEDICAL LAB CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511149390.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2026-08-21
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

[0002]核型分析主要用于检测染色体数目变异,例如整倍体变异,非整倍体变异,还有染色体结构变异,例如染色体片段缺失与重复,染色体结构的到位和易位等;常规的核型分析需要先进行样本采集,细胞培养,制片,染色体染色,微观分析及评估等较为复杂且繁琐的生物学实验,并需要使用专业的相关实验设备,因此,常规的核型检测分析不仅周期长,而且检测成本较高,且常规核型分析的最低分辨率为5~10Mb

Benefits of technology

[0037]As can be seen, this application discloses a method for chromosome karyotype analysis, comprising: performing genome alignment on whole-exome sequencing data of a biological sample to be tested based on a preset mapping relationship file to obtain a target counting matrix of sequencing reads of the biological sample to be tested; and performing variant detection on the whole-exome sequencing data to obtain a variant detection result file of the biological sample to be tested; wherein, the preset mapping relationship file is a file containing mapping relationships between different probe capture intervals and cytogenetic bands; determining the sequencing read difference information and statistical significance information of each cytogenetic band interval according to the target counting matrix and the standard counting matrix to obtain the chromosome karyotype analysis results of each cytogenetic band interval. The first karyotype result; wherein, the standard counting matrix is ​​the counting matrix of sequencing reads of the biological sample with standard chromosome karyotype; based on the variant detection result file, the number of variant sites on each chromosome of the biological sample to be tested is counted, and the allele frequency density distribution of the target chromosome whose number of variant sites reaches a preset variant site number threshold is calculated, so as to analyze whether the target chromosome has chromosomal euploidy variation according to the allele peak characteristics corresponding to the allele frequency density distribution, so as to obtain the second karyotype result at the chromosome level; the target chromosome karyotype analysis result of the biological sample to be tested is determined according to the first karyotype result and the second karyotype result. Therefore, direct analysis of whole-exome sequencing data eliminates the need for complex sample processing such as cell culture/slide preparation. Bioinformatics algorithms, including genome alignment, counting matrix generation, variant detection, and dual-dimensional analysis, automate the process, compressing the analysis cycle. Specifically, by dividing regions based on cytogenetic bands, the first karyotype result is obtained, refining the analysis to the cytogenetic band level, resulting in higher karyotype resolution. Finally, dual-dimensional parallel analysis and cross-validation ensure that a single analysis covers two types of variants, avoiding redundant experiments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998296B_ABST
    Figure CN120998296B_ABST
Patent Text Reader

Abstract

The application discloses a chromosome karyotype analysis method, device, equipment and medium, and relates to the technical field of chromosome analysis, and comprises the following steps: performing genome alignment on whole exome sequencing data of a biological sample to be measured based on a preset mapping relationship file, obtaining a target count matrix of sequencing reads, performing variation detection on the whole exome sequencing data, and obtaining a variation detection result file; determining sequencing read difference information and statistical significance information according to the target count matrix and a standard count matrix, and obtaining a first karyotype result; based on the variation detection result file, counting the number of variation sites of each chromosome of the biological sample to be measured, calculating the allele frequency value density distribution of a target chromosome whose variation site number reaches a preset variation site number threshold, and analyzing the target chromosome to obtain a second karyotype result; and determining a target chromosome karyotype analysis result according to the first karyotype result and the second karyotype result. The whole exome sequencing data is directly and automatically processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of chromosome analysis technology, and in particular to a method, apparatus, equipment and medium for chromosome karyotype analysis. Background Technology

[0002] Karyotype analysis is mainly used to detect chromosome number variations, such as euploidy and aneuploidy, as well as chromosome structural variations, such as chromosome segment deletions and duplications, chromosome placement and translocation. Conventional karyotype analysis requires complex and tedious biological experiments, including sample collection, cell culture, slide preparation, chromosome staining, microscopic analysis and evaluation, and requires the use of specialized experimental equipment. Therefore, conventional karyotype analysis is not only time-consuming but also costly, and the minimum resolution of conventional karyotype analysis is 5-10 Mb.

[0003] In summary, how to achieve accurate chromosome karyotype analysis at low cost and in a short time is a technical problem that needs to be solved in this field. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a method, apparatus, device, and medium for chromosome karyotype analysis, enabling accurate chromosome karyotype analysis at low cost and in a short time. The specific solution is as follows:

[0005] In a first aspect, this application discloses a method for chromosome karyotype analysis, including:

[0006] Genome alignment is performed on the whole exome sequencing data of the biological sample to be tested based on a preset mapping relationship file to obtain the target counting matrix of the sequencing reads of the biological sample to be tested, and variant detection is performed on the whole exome sequencing data to obtain the variant detection result file of the biological sample to be tested; wherein, the preset mapping relationship file is a file containing the mapping relationship between different probe capture intervals and cytogenetic bands.

[0007] Based on the target counting matrix and the standard counting matrix, the sequencing read difference information and statistical significance information of each cytogenetic band interval are determined to obtain the first karyotype result of each cytogenetic band interval; wherein, the standard counting matrix is ​​the counting matrix of sequencing reads of biological samples with standard chromosome karyotype;

[0008] Based on the mutation detection result file, the number of mutation sites on each chromosome of the biological sample to be tested is counted, and the allele frequency density distribution of the target chromosome whose number of mutation sites reaches the preset mutation site number threshold is calculated. Based on the allele peak characteristics corresponding to the allele frequency density distribution, the chromosomal euploidy variation of the target chromosome is analyzed to obtain the second karyotype result at the chromosome level.

[0009] The target chromosome karyotype analysis results of the biological sample to be tested are determined based on the first karyotype result and the second karyotype result.

[0010] Optionally, before performing genome alignment on the whole exome sequencing data of the biological sample to be tested based on the preset mapping relationship file, the method further includes:

[0011] The whole exome sequencing probe capture regions are mapped to cytogenetic band regions, generating a preset mapping relationship file containing the mapping relationship between probe capture regions and corresponding cytogenetic band regions.

[0012] Optionally, the step of mapping whole-exome sequencing probe capture regions to cytogenetic band regions and generating a preset mapping relationship file containing the mapping relationship between probe capture regions and corresponding cytogenetic band regions includes:

[0013] Obtain a cytogenetic band information file; the cytogenetic band information file includes chromosomes, start positions, end positions, cytogenetic band names, and relevant intervals for each chromosome;

[0014] Obtain the whole exome sequencing probe file, wherein the whole exome sequencing probe file is the probe interval file corresponding to the capture probe kit used for sequencing, and the whole exome sequencing probe file includes chromosome, capture interval start position and capture interval end position;

[0015] Using preset interval mapping software, the overlap between the whole exome sequencing probe capture interval in the whole exome sequencing probe file and the genomic coordinate interval of the cytogenetic band interval in the cytogenetic band information file is mapped to generate a preset mapping relationship file;

[0016] The preset mapping relationship file contains information such as chromosome, start position of capture interval, end position of capture interval, corresponding cytogenetic band interval, and name of cytogenetic band interval.

[0017] Optionally, the step of performing genome alignment on the whole exome sequencing data of the biological sample to be tested based on a preset mapping relationship file to obtain a target counting matrix of sequencing reads of the biological sample to be tested includes:

[0018] Based on the preset mapping relationship file, the counting information of sequencing reads in the cytogenetic band interval is statistically analyzed, and a corresponding target counting matrix is ​​generated based on the counting information.

[0019] Optionally, the step of performing variant detection on the whole exome sequencing data to obtain a variant detection result file for the biological sample to be tested includes:

[0020] The whole exome sequencing data and the human reference genome are aligned to obtain a first genome alignment result file containing the position, orientation, and alignment quality information of the reads on the genome.

[0021] The variant sites in the first genome comparison result file are scanned, and the allele frequency of each site is calculated based on the number of reads at the variant sites and the total number of reads covered by the variant sites, so as to obtain a variant detection result file containing genotypic information of all sites of the biological sample to be tested.

[0022] Optionally, before determining the sequencing read difference information and statistical significance information of each cytogenetic band interval based on the target counting matrix and the standard counting matrix, the method further includes:

[0023] Standard biological samples characterized by normal chromosome karyotypes were selected, and reads were compared between the whole exome sequencing data of the standard biological samples and the human reference genome to obtain a second genome alignment result file containing the position information, orientation information and alignment quality information of the reads on the genome.

[0024] Based on the mapping relationship file, the sequencing read counts of each cytogenetic band interval in the second genome alignment result file are statistically analyzed to generate a standard counting matrix.

[0025] Optionally, determining the sequencing read difference information and statistical significance information of each cytogenetic band interval based on the target counting matrix and the standard counting matrix includes:

[0026] Based on the target counting matrix and the standard counting matrix, the sequencing read difference value and significance p-value of each cytogenetic band interval are calculated.

[0027] Secondly, this application discloses a chromosome karyotype analysis device, comprising:

[0028] The genome alignment module is used to perform genome alignment on the whole exome sequencing data of the biological sample to be tested based on a preset mapping relationship file, so as to obtain the target counting matrix of the sequencing reads of the biological sample to be tested; wherein, the preset mapping relationship file is a file containing the mapping relationship between different probe capture intervals and cytogenetic bands;

[0029] The variant detection module is used to perform variant detection on the whole exome sequencing data to obtain a variant detection result file of the biological sample to be tested.

[0030] The first analysis results module is used to determine the sequencing read difference information and statistical significance information of each cytogenetic band interval based on the target counting matrix and the standard counting matrix, so as to obtain the first karyotype result of each cytogenetic band interval; wherein, the standard counting matrix is ​​the counting matrix of sequencing reads of biological samples with standard chromosome karyotype;

[0031] The second analysis result module is used to count the number of variant sites on each chromosome of the biological sample to be tested based on the variant detection result file, and to calculate the allele frequency density distribution of the target chromosome whose number of variant sites reaches a preset variant site number threshold, so as to analyze whether the target chromosome has chromosomal euploidy variation according to the allele peak characteristics corresponding to the allele frequency density distribution, so as to obtain the second karyotype result at the chromosome level.

[0032] The comprehensive analysis results module is used to determine the target chromosome karyotype analysis results of the biological sample to be tested based on the first karyotype result and the second karyotype result.

[0033] Thirdly, this application discloses an electronic device, including:

[0034] Memory, used to store computer programs;

[0035] A processor is used to execute the computer program to implement the steps of the aforementioned disclosed chromosome karyotype analysis method.

[0036] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the aforementioned disclosed chromosome karyotype analysis method.

[0037] As can be seen, this application discloses a method for chromosome karyotype analysis, comprising: performing genome alignment on whole-exome sequencing data of a biological sample to be tested based on a preset mapping relationship file to obtain a target counting matrix of sequencing reads of the biological sample to be tested; and performing variant detection on the whole-exome sequencing data to obtain a variant detection result file of the biological sample to be tested; wherein, the preset mapping relationship file is a file containing mapping relationships between different probe capture intervals and cytogenetic bands; determining the sequencing read difference information and statistical significance information of each cytogenetic band interval according to the target counting matrix and the standard counting matrix to obtain the chromosome karyotype analysis results of each cytogenetic band interval. The first karyotype result; wherein, the standard counting matrix is ​​the counting matrix of sequencing reads of the biological sample with standard chromosome karyotype; based on the variant detection result file, the number of variant sites on each chromosome of the biological sample to be tested is counted, and the allele frequency density distribution of the target chromosome whose number of variant sites reaches a preset variant site number threshold is calculated, so as to analyze whether the target chromosome has chromosomal euploidy variation according to the allele peak characteristics corresponding to the allele frequency density distribution, so as to obtain the second karyotype result at the chromosome level; the target chromosome karyotype analysis result of the biological sample to be tested is determined according to the first karyotype result and the second karyotype result. Therefore, direct analysis of whole-exome sequencing data eliminates the need for complex sample processing such as cell culture / slide preparation. Bioinformatics algorithms, including genome alignment, counting matrix generation, variant detection, and dual-dimensional analysis, automate the process, compressing the analysis cycle. Specifically, by dividing regions based on cytogenetic bands, the first karyotype result is obtained, refining the analysis to the cytogenetic band level, resulting in higher karyotype resolution. Finally, dual-dimensional parallel analysis and cross-validation ensure that a single analysis covers two types of variants, avoiding redundant experiments. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0039] Figure 1 This is a flowchart of a chromosome karyotype analysis method disclosed in this application;

[0040] Figure 2 This is a flowchart of a method for generating a pre-defined relationship file disclosed in this application;

[0041] Figure 3 This is a flowchart of a method for generating a standard counting matrix for a comparison set, as disclosed in this application.

[0042] Figure 4 This is a flowchart of a method for generating a target counting matrix and a variation detection result file for a biological sample to be tested, as disclosed in this application.

[0043] Figure 5 This is a flowchart of a method for generating karyotype analysis results of a small fragment interval of a biological sample disclosed in this application;

[0044] Figure 6 This is a flowchart of a method for generating chromosome-level karyotype analysis results of a biological sample to be tested, as disclosed in this application.

[0045] Figure 7(a) is a ch21 chromosome AF density distribution map of a negative sample (female) disclosed in this application; Figure 7(b) is a ch21 chromosome AF density distribution map of internal number MS23080616P1.

[0046] Figure 8(a) is a distribution map of AF density of chr19 chromosome variation in a negative sample (male) disclosed in this application, and Figure 8(b) is a distribution map of AF density of chr19 chromosome variation in internal number MS24031736P1-TT.

[0047] Figure 9(a) is a distribution map of the AF density of the chr18 chromosome variation in a negative sample (male) disclosed in this application, and Figure 9(b) is a distribution map of the AF density of the chr18 chromosome variation in the internal number MS24050045P1-TT.

[0048] Figure 10(a) is a chrX chromosome variation AF density distribution map of a negative sample (female) disclosed in this application, and Figure 10(b) is a chrX chromosome variation AF density distribution map of internal number MS23100189P1.

[0049] Figure 11(a) is a distribution map of AF density of chrX chromosome variation in a negative sample (female) disclosed in this application, and Figure 11(b) is a distribution map of AF density of chrX chromosome variation in internal number MS24080479P1.

[0050] Figure 12(a) is a distribution map of AF density of chrX chromosome variation in a negative sample (male) disclosed in this application, and Figure 12(b) is a distribution map of AF density of chrX chromosome variation in internal number MS22083806P1.

[0051] Figure 13 This is a schematic diagram of the first karyotype analysis results for an internal number MS22110048P1 disclosed in this application;

[0052] Figure 14This is a schematic diagram of the first karyotype analysis results for an internal number MS23040546P1 disclosed in this application;

[0053] Figure 15 This is a schematic diagram of the first karyotype analysis results for an internal number MS23120853P1 disclosed in this application;

[0054] Figure 16 This is a schematic diagram of the first karyotype analysis results for an internal number MS24051012P1 disclosed in this application;

[0055] Figure 17 This is a schematic diagram of the first karyotype analysis results for an internal number MS24070268P1 disclosed in this application;

[0056] Figure 18 This is a schematic diagram of the first karyotype analysis results for an internal number MS22100557P1 disclosed in this application;

[0057] Figure 19 This application discloses a flowchart of a method for generating target chromosome karyotype analysis results for a biological sample to be tested.

[0058] Figure 20 This is a flowchart of a specific chromosome karyotype analysis method disclosed in this application;

[0059] Figure 21 This is a schematic diagram of the structure of a chromosome karyotype analysis device disclosed in this application;

[0060] Figure 22 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0061] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0062] Karyotype analysis is mainly used to detect chromosome number variations, such as euploidy and aneuploidy, as well as chromosome structural variations, such as chromosome segment deletions and duplications, chromosome placement and translocation. Conventional karyotype analysis requires complex and tedious biological experiments, including sample collection, cell culture, slide preparation, chromosome staining, microscopic analysis and evaluation, and requires the use of specialized experimental equipment. Therefore, conventional karyotype analysis is not only time-consuming but also costly, and the minimum resolution of conventional karyotype analysis is 5-10 Mb.

[0063] Therefore, this invention provides a chromosome karyotype analysis scheme that enables accurate chromosome karyotype analysis at low cost and in a short time.

[0064] Reference Figure 1 As shown, this embodiment of the invention discloses a method for chromosome karyotype analysis, comprising:

[0065] Step S11: Perform genome alignment on the whole exome sequencing data of the biological sample to be tested based on the preset mapping relationship file to obtain the target counting matrix of the sequencing reads of the biological sample to be tested, and perform variant detection on the whole exome sequencing data to obtain the variant detection result file of the biological sample to be tested; wherein, the preset mapping relationship file is a file containing the mapping relationship between different probe capture intervals and cytogenetic bands.

[0066] In this embodiment, before performing genome alignment on the whole exome sequencing data of the biological sample to be tested based on the preset mapping relationship file, the method further includes: mapping the whole exome sequencing probe capture interval to the cytogenetic band interval, and generating a preset mapping relationship file containing the mapping relationship between the probe capture interval and the corresponding cytogenetic band interval. Specifically, the process involves: obtaining a cytogenetic band information file, which includes chromosomes, start positions, end positions, cytogenetic band names, and relevant intervals for each chromosome; obtaining a whole-exome sequencing probe file, which is the probe interval file corresponding to the capture probe kit used for sequencing, and includes chromosomes, capture interval start positions, and capture interval end positions; using preset interval mapping software, mapping the overlap between the capture intervals of the whole-exome sequencing probes in the whole-exome sequencing probe file and the genomic coordinate intervals of the cytogenetic bands in the cytogenetic band information file to generate a preset mapping relationship file; the preset mapping relationship file includes chromosomes, capture interval start positions, capture interval end positions, corresponding cytogenetic band intervals, and cytogenetic band interval names. Understandably, the process involves obtaining cytogenetic band (CytoBand) information files. These files contain chromosomes, start and end positions, CytoBand names, and Giemsa staining results, retaining only the relevant regions of chromosomes chr1-chr22, chrX, and chrY, and sorting them according to chromosome and start coordinates. Next, whole exome sequencing (WES) probe BED (Browser Extensible Data) files are obtained. These files contain chromosomes, capture interval start and end positions, and are the probe interval files corresponding to the capture probe kits used for WES sequencing. Using interval mapping software (bedtools), based on the overlap of genomic coordinate intervals, the probe capture intervals in the probe BED files are mapped to the cytogenetic band intervals in the CytoBand information files, generating a pre-defined mapping relationship file. This pre-defined mapping relationship file contains chromosomes, capture interval start and end positions, the corresponding CytoBand intervals, and CytoBand names.

[0067] like Figure 2As shown, based on the coordinate intervals in the CytoBand information file and the coordinate intervals in the WES probe file, the WES probe intervals are mapped to specific CytoBand intervals, forming a mapping Map file. This mapping file will be used to classify and count CytoBand intervals in the Reads of the BAM (Binary Alignment Map, the genome alignment result file of the sequencing data obtained after WES / WGS high-throughput sequencing of the sample) file by aligning the sequencing data with the human reference genome. Specifically, the CytoBand information file is downloaded, the chromosome-related intervals of chr1~chr22, chrX, and chrY are retained, and sorted according to chromosome and starting coordinates to generate the CytoBand information file. The file format of some data in the CytoBand information file is shown in Table 1 below:

[0068] Table 1. Different CytoBand information and staining results for chromosome chr1 in CytoBand information files.

[0069]

[0070] Next, prepare the WES probe BED file, which is the basic attribute file for WES sequencing. Different capture kits have their own corresponding probe interval files. The file format is shown in Table 2 below. Table 2 is a partial data screenshot.

[0071] Table 2. Probe region files for different capture kits for chromosome chr1.

[0072]

[0073] Based on the overlap of genomic coordinate intervals, the probe BED files are mapped to CytoBand intervals, and a Map relation file is constructed. Information from some of the Map relation files is shown in Table 3 below.

[0074] Table 3 Information data of the Map relational file

[0075]

[0076] Therefore, the input files are the CytoBand information file and the captured probe BED file, which are sent to the interval mapping software (bedtools), and the output file is the probe interval and CytoBand mapping Map file.

[0077] This approach, through the construction of a pre-defined mapping file, provides a precise basis for subsequent karyotype analysis based on whole-exome sequencing data, ensuring that read counts are accurately correlated with cytogenetic bands and establishing a reliable data mapping framework for subsequent analysis. It achieves precise correlation between whole-exome sequencing probe capture regions and cytogenetic band regions, enabling karyotype analysis of small fragment regions to be conducted based on clearly defined region divisions. It also standardizes the region criteria for data processing, allowing for comparative analysis of read count matrices between the control set and the test samples within the same region dimension, ensuring the accuracy and consistency of subsequent karyotype analysis results.

[0078] In this embodiment, standard biological samples characterized by normal chromosome karyotypes are selected. The whole-exome sequencing data of these standard biological samples are aligned with the human reference genome to obtain a second genome alignment result file containing the position, orientation, and alignment quality information of the reads on the genome. Based on the mapping relationship file, the sequencing read counts for each cytogenetic band interval in the second genome alignment result file are statistically analyzed to generate a standard count matrix. It can be understood that biological samples with standard chromosome karyotypes of 46,XY without genomic variation (control set samples) from approximately 20 or more male samples and 20 or more female samples are selected. After genome alignment, a Reads count matrix file is generated based on the Map mapping file to create a control set for the samples to be tested. Specifically, as shown... Figure 3 As shown, male and female samples were divided into two groups, and at least 20 negative samples without genomic variations were selected from each group. WES sequencing raw data (Fastq) was prepared for the samples. The Fastq data was aligned with the human reference genome version hg38 to obtain a genome alignment result BAM file. Alignment software included, but not limited to, the following: bwa, bowtie, SOAP2, etc. Based on the intervals in the Map mapping file, the reads in the BAM file were counted to generate a control set read count matrix file. Input files: Fastq file of control set samples, Map mapping file, hg38 reference genome file; Related software: Genome alignment software (bwa, bowtie, SOPA, etc.), read count statistics (bedtools); Output file: Control set read count matrix file, i.e., the standard count matrix.

[0079] In this embodiment, the counting information of sequencing reads within the cytogenetic band interval is statistically analyzed based on the preset mapping relationship file, and a corresponding target counting matrix is ​​generated based on the counting information. It is understood that, following the method described above for obtaining the standard counting matrix of the control set samples, the target counting matrix is ​​generated by statistically analyzing the counting information of test reads located within the cytogenetic band interval based on the preset mapping relationship file; this will not be elaborated further.

[0080] In this embodiment, read alignment is performed on the whole-exome sequencing data and the human reference genome to obtain a first genome alignment result file containing the position, orientation, and alignment quality information of the reads on the genome. Variant sites in the first genome alignment result file are scanned, and the allele frequency of each site is calculated based on the number of reads at the variant sites and the total number of reads covered by the variant sites, to obtain a variant detection result file containing genotypic information for all sites of the biological sample to be tested. It can be understood that the process of generating a VCF (Variant Call Format) file for variant detection is as follows: The input data is the raw WES sequencing data (Fastq file) of the sample to be tested, which contains short read sequences and quality scores from the sequencer. The input data also includes the human reference genome (hg38 version) to provide a genomic coordinate baseline. Then, a data quality control and preprocessing workflow is performed on the input data. First, the quality of the Fastq file is checked (Q20 / Q30 scores, GC content, adapter contamination), and low-quality reads are filtered out (bases with a quality value <20 are pruned or removed). The preprocessed data is then subjected to genome alignment, which involves aligning the quality-controlled Fastq reads to the reference genome and outputting a BAM file. This file contains information such as the location, orientation, and alignment quality of the stored reads on the genome. Post-processing of the BAM file involves: first, sorting the BAM files by genomic coordinates (using samtools sort); removing repetitive sequences introduced by PCR amplification; and correcting for sequencing system errors (base quality). Then, a genome-wide scan for variant sites (single-base variants, SNVs, and insertions / deletions, InDels) is performed on the corrected file, and the allele frequency (AF) for each site is calculated. Multiple samples are then merged to obtain a VCF file (variation detection result file). It is important to note that AF represents the specific variant site described in the VCF file, and is the ratio of the sequencing depth X of the mutation to the total sequencing depth of the mutated base, expressed by the formula:

[0081] AF = Sequencing depth of site mutation / Total sequencing depth of site; VCF file specifically refers to the file format in which the BAM file is analyzed for each base of the genome using variant detection software, and the specific location and type of variants that occur in the sample throughout the whole genome are obtained. There are two types of variants: single nucleotide variants (SNV) and multiple nucleotide insertion or deletion variants (InDel).

[0082] Reference Figure 4As shown, a variant VCF file is generated based on the sample to be tested, and a Reads count matrix is ​​generated based on the Map mapping file; after aligning the WES sequencing data of the sample to be tested with the reference genome, variant detection is performed to generate a VCF file, and a Reads count matrix of the sample to be tested is generated based on the Map information file.

[0083] Construction process: Prepare WES sequencing Fastq data of the sample to be tested; align the Fastq data with the human reference genome version hg38 to obtain the genome alignment result BAM file. Alignment software includes, but is not limited to, the following types: bwa, bowtie, SOAP2, etc.; based on the intervals of the map mapping file, count the reads in the BAM file to generate a control set read count matrix file; perform variant detection on the BAM file to generate a VCF variant file.

[0084] Input files: WES sequencing Fastq file of the sample to be tested, hg38 reference genome sequence file; related software: genome alignment software (bwa, bowtie, SOPA, etc.), reads counting and statistics (bedtools); variant detection software: GATK, sentieon, etc.

[0085] Output files: Reads count matrix file of the test sample, VCF file of the test sample variation.

[0086] Step S12: Determine the sequencing read difference information and statistical significance information of each cytogenetic band interval based on the target counting matrix and the standard counting matrix to obtain the first karyotype result of each cytogenetic band interval; wherein, the standard counting matrix is ​​the counting matrix of sequencing reads of biological samples with standard chromosome karyotype.

[0087] In this embodiment, based on the target counting matrix and the standard counting matrix, the sequencing read difference value and significance p-value of each cytogenetic band interval are calculated to obtain the first karyotype result for each cytogenetic band interval. Figure 5As shown, karyotype analysis of small fragment variations is performed based on the Reads count matrix of the control set and the test samples. Specifically, based on the control set data, the karyotype results of small chromosomal fragment intervals are finally determined by statistical analysis of the log2 Ratio and p-value for each CytoBand interval in the clinical test samples. Construction process: Prepare test samples: Reads count matrix and Reads count matrix file of the sex-matched control set; Count the log2 Ratio value and significance p-value for each CytoBand interval; Perform copy number analysis on each CytoBand interval based on the p-value and log2 Ratio, and perform interval karyotype analysis according to the "CytoBand Interval Copy Number Genotyping Table"; Relevant software: Statistical analysis tools; Input files: Reads count matrix of the test samples, Reads count matrix of the control set; Output file: Karyotype analysis results of small fragment intervals. It is important to note that the above-mentioned log2 ratio is specifically calculated by dividing the number of reads in each interval of the test sample by the number of reads in the corresponding interval of the control set to obtain the ratio value, and then performing a log2 transformation on this ratio value, denoted as log2 ratio; the formula is described as follows:

[0088] log2 Ratio = log2(number of Reads in the test sample interval / number of Reads in the control set interval);

[0089] Taking the male / female chromosome analysis scenario as an example, the CytoBand region copy number genotyping results obtained by performing the above steps are shown in Table 4:

[0090] Table 4 CytoBand Interval Copy Number Genotyping Table (First Karyotype Results)

[0091]

[0092] Step S13: Based on the mutation detection result file, count the number of mutation sites on each chromosome of the biological sample to be tested, and calculate the allele frequency density distribution of the target chromosome whose number of mutation sites reaches the preset mutation site number threshold. Analyze whether the target chromosome has chromosomal euploidy variation based on the allele peak characteristics corresponding to the allele frequency density distribution, so as to obtain the second karyotype result at the chromosome level.

[0093] like Figure 6As shown, prepare the VCF variant file of the sample to be tested, and count the number of variant sites and the AF file of the variant proportion; perform density statistics based on the AF file of the variant proportion: plot the AF values ​​based on the density distribution (Kernel Density Estimation, KDE), count the number of peaks and determine the value range of the AF value corresponding to the peak; perform chromosome karyotype analysis based on the "Variation and AF Peak Plot Typing Table" according to the number of variant sites, the number of peaks in the variant proportion distribution and the AF value corresponding to the peak; related software: statistical analysis tools; input file: VCF file of the sample variants; output file: chromosome hierarchical karyotype analysis results (variation count, AF distribution plot, karyotype determination results, etc.).

[0094] Taking the male / female chromosome analysis scenario as an example (including autosomes and sex chromosomes), the chromosome hierarchical karyotype analysis results and AF peak diagram results obtained by performing the above steps are shown in Table 5:

[0095] Table 5. Chromosomal hierarchical variation and AF peak plot typing table

[0096]

[0097] The following are examples of second karyotype results: Case 1: Internal number MS23080616P1, gender "female". The clinical karyotype diagnosis of this sample is trisomy 21 syndrome. Figure 7(a) is the AF density distribution map of the ch21 chromosome in the negative sample (female). Figure 7(b) is the AF density distribution map of the ch21 chromosome in the internal number MS23080616P1. According to the statistical analysis of the AF distribution of the chr21 chromosome variation in the WES data of this sample, the main peak of heterozygosity at AF~0.5 disappeared and two new heterozygosity peaks were added. The AF values ​​corresponding to the peak values ​​are distributed in the intervals of [0.20, 0.45] and [0.55, 0.80]. According to the "Variation and AF Peak Map Typing Table", the final karyotype result is chr21 chromosome triploid duplication, which is consistent with the clinical diagnosis karyotype. Case 2: Internal number MS24031736P1-TT, gender "unknown". This sample is a clinical miscarriage sample with no karyotype result. Figure 8(a) is the AF density distribution map of the chr19 chromosome variation in the negative sample (male). Figure 8(b) is the AF density distribution map of the chr19 chromosome variation in the internal number MS24031736P1-TT. According to the statistical analysis of the AF distribution of the chr19 chromosome variation in the WES data of this sample, the heterozygous peak of AF ~ 0.5 disappeared and two new heterozygous peaks were added. The AF values ​​corresponding to the peak values ​​are distributed in the intervals of [0.20, 0.45] and [0.55, 0.80]. According to the "Variation and AF Peak Map Typing Table", the final karyotype result is determined to be chr19 chromosome triploid duplication. Because the abortion sample was not suitable for routine karyotyping in a clinical setting, no karyotyping results could be provided. However, this method can analyze the karyotype of abortion samples. Case 3: Internal number MS24050045P1-TT, gender "unknown". This sample was a clinical abortion sample with no karyotype results. Figure 9(a) shows the AF density distribution of the chr18 chromosome variation in the negative sample (male). Figure 9(b) shows the AF density distribution of the chr18 chromosome variation in the internal number MS24050045P1-TT. This method statistically analyzed the AF distribution of the chr18 chromosome variation in the WES data of this sample. The heterozygous peak of AF~0.5 disappeared, and two new heterozygous peaks were added. The AF values ​​corresponding to the peak values ​​were distributed in the intervals of [0.20, 0.45] and [0.55, 0.80]. According to the "Variation and AF Peak Map Typing Table", the final karyotype result was determined to be a chr18 chromosome triploid duplication.Because the abortion sample cannot be routinely karyotyped in the clinical setting, the karyotype results cannot be provided. However, this method can analyze the karyotype of abortion samples. Case 4: Internal number MS23100189P1, gender "female", clinical Turner syndrome, karyotype 46,X,deI(X)(q24-28)

[28] / 45,X[2]. Figure 10(a) is the AF density distribution map of the chrX chromosome variation in the negative sample (female), and Figure 10(b) is the AF density distribution map of the chrX chromosome variation in the internal number MS23100189P1. This method statistically analyzed the AF distribution of the chrX chromosome variation in the WES data of this sample and found that the heterozygous peak of AF~0.5 disappeared and the homozygous peak of AF~[0.85,1] became the main peak. According to the "Variation and AF Peak Map Typing Table", the final karyotype result was determined to be chrX chromosome 1-ploid deletion, which is consistent with the clinical diagnosis karyotype. Case 5: Internal ID MS24080479P1, gender "female", karyotype result of this sample is 47,XXX. Figure 11(a) is the AF density distribution map of chrX chromosome variation in the negative sample (female), and Figure 11(b) is the AF density distribution map of chrX chromosome variation in the internal ID MS24080479P1. According to the statistical analysis of AF distribution of chrX chromosome variation in the WES data of this sample, the heterozygous main peak of AF~0.5 disappeared and two new heterozygous peaks were added. The peak values ​​corresponding to the AF values ​​are distributed in the interval of [0.20, 0.45] and [0.55, 0.80]. According to the "Variation and AF Peak Map Typing Table", the final karyotype result is determined to be chrX chromosome triploid duplication, which is consistent with the clinical diagnosis karyotype. Case 6: Internal ID MS22083806P1, male. The karyotype of this sample is 47,XXY. Figure 12(a) shows the AF density distribution of the chrX chromosome variation in the negative sample (male), and Figure 12(b) shows the AF density distribution of the chrX chromosome variation in the internal ID MS22083806P1. Based on the statistical analysis of the AF distribution of the chrX chromosome variation in the WES data of this sample, one new heterozygous peak with AF values ​​around 0.5 was observed. According to the "Variation and AF Peak Map Typing Table," the final karyotype result was determined to be a chrX chromosome diploid duplication, consistent with the clinical diagnosis. All case information and results are summarized in Table 6.

[0098] Table 6 Summary Results of Multiple Cases

[0099]

[0100] Furthermore, the following are examples of small-fragment hierarchical variations in the first karyotype results: Case 7: Internal number MS22110048P1, the karyotype result of this sample is "47,XX,+mar", +mar: there is an additional small marker chromosome, the source of which is uncertain. This method analyzes the interval log2 ratio of the WES data for this sample and determines it to be a small-fragment copy number variation, such as... Figure 13 As shown, the karyotype analysis result is 15q11.2q13.1(23,326,915-28,553,786)×3, with a variation interval size of 5.22 Mb. Case 8: Internal number MS23040546P1, the karyotype result of this sample is "47,XX,+mar", where +mar indicates the presence of an additional small marker chromosome of uncertain origin. This method analyzes the interval log2 ratio of the WES data for this sample and identifies it as a small fragment copy number variation, such as... Figure 14 As shown, the karyotype analysis results are 22q11.1q11.21(16,397,011-18,172,802)×3, with a variation interval of 1.77Mb.

[0101] Case 9: Internal ID MS23120853P1, the karyotype result of this sample is "46, XX, 15ps+", 15ps+: growth of the short arm of chromosome 15, but the specific interval is not confirmed. This method analyzes the interval log2 ratio of the WES data of this sample, such as... Figure 15 As shown, it was determined to be a small fragment copy number variation. The karyotype analysis result was 22q11.21(20,363,588-21,060,548)×1, with a variation interval size of 0.69Mb. Case 10: Internal number MS24051012P1, the karyotype result of this sample was "46,XY", and no chromosomal abnormalities were detected in the karyotype result. This method analyzes the interval log2 Ratio of the WES data of this sample, as shown... Figure 16 As shown, it was determined to be a small fragment copy number variation. The karyotype analysis result was 11q12.3q13.2(62,705,281-66,527,041)×3, with a variation interval size of 3.82Mb. Case 11: Internal number MS24070268P1, the karyotype result of this sample was "46,XY", and no chromosomal abnormalities were detected in the karyotype result. This method analyzes the interval log2 Ratio of the WES data of this sample, as shown... Figure 17 As shown, it was determined to be a small fragment copy number variation. The karyotype analysis result was 7q11.23(73,303,404-74,734,011)×1, with a variation interval size of 1.43 Mb. Case 12: Internal number MS22100557P1, the karyotype result of this sample was "46,XY", and no chromosomal abnormalities were detected in the karyotype result. This method analyzed the interval log2 Ratio of the WES data of this sample, as shown... Figure 18 As shown, it was determined to be a small fragment copy number variation. The karyotype analysis result was 16p13.3(3,536,102-4,272,761)×3, with a variation interval size of 0.74 Mb. The comprehensive summary results of the above cases are shown in Table 7 below:

[0102] Table 7 Summary of Cases of First Karyotype Results

[0103]

[0104] Step S14: Determine the target chromosome karyotype analysis result of the biological sample to be tested based on the first karyotype result and the second karyotype result.

[0105] like Figure 19 As shown, the final karyotype of the sample is comprehensively evaluated based on the karyotype analysis results of small fragment variations and chromosome-level karyotype analysis. Specifically, the final karyotype analysis result of the sample is obtained by comprehensively judging the results of chromosome-level karyotype analysis and small fragment karyotype analysis. Construction process: Input small fragment karyotype analysis results and chromosome-level karyotype analysis result files; comprehensively evaluate the overall karyotype result of the sample; Input files: small fragment karyotype analysis result file, chromosome-level karyotype analysis; Output file: final karyotype analysis result of the sample.

[0106] like Figure 20 As shown, this invention provides a specific overall method for chromosome karyotype analysis, specifically:

[0107] A Map mapping relationship file is constructed based on CytoBand interval files and WES probe BED files;

[0108] Construct a Reads counting matrix based on the control set samples and the Map mapping file;

[0109] A variant VCF file is generated based on the sample to be tested, and a Reads count matrix is ​​generated based on the Map mapping file;

[0110] Based on the Reads count matrix of the control set and the test sample, karyotype analysis of small fragment variations was performed;

[0111] Based on the VCF file of the sample to be tested, perform karyotype analysis at the chromosome level;

[0112] The final karyotype of the sample under test is comprehensively evaluated based on the results of karyotype analysis of small fragment variations and karyotype analysis at the chromosome level.

[0113] Therefore, this invention analyzes the karyotype of samples based on high-throughput whole-exome sequencing data, eliminating the need for additional karyotype analysis experiments and reducing testing costs; it performs chromosome-level karyotype analysis based on high-throughput data, including deletions / duplications of euploidy, and deletions / duplications of single or multiple chromosomes; and it performs small-fragment karyotype analysis based on high-throughput data, improving the resolution of karyotype analysis results.

[0114] As can be seen, this application discloses a method for chromosome karyotype analysis, comprising: performing genome alignment on whole-exome sequencing data of a biological sample to be tested based on a preset mapping relationship file to obtain a target counting matrix of sequencing reads of the biological sample to be tested; and performing variant detection on the whole-exome sequencing data to obtain a variant detection result file of the biological sample to be tested; wherein, the preset mapping relationship file is a file containing mapping relationships between different probe capture intervals and cytogenetic bands; determining the sequencing read difference information and statistical significance information of each cytogenetic band interval according to the target counting matrix and the standard counting matrix to obtain the chromosome karyotype analysis results of each cytogenetic band interval. The first karyotype result; wherein, the standard counting matrix is ​​the counting matrix of sequencing reads of the biological sample with standard chromosome karyotype; based on the variant detection result file, the number of variant sites on each chromosome of the biological sample to be tested is counted, and the allele frequency density distribution of the target chromosome whose number of variant sites reaches a preset variant site number threshold is calculated, so as to analyze whether the target chromosome has chromosomal euploidy variation according to the allele peak characteristics corresponding to the allele frequency density distribution, so as to obtain the second karyotype result at the chromosome level; the target chromosome karyotype analysis result of the biological sample to be tested is determined according to the first karyotype result and the second karyotype result. Therefore, direct analysis of whole-exome sequencing data eliminates the need for complex sample processing such as cell culture / slide preparation. Bioinformatics algorithms, including genome alignment, counting matrix generation, variant detection, and dual-dimensional analysis, automate the process, compressing the analysis cycle. Specifically, by dividing regions based on cytogenetic bands, the first karyotype result is obtained, refining the analysis to the cytogenetic band level, resulting in higher karyotype resolution. Finally, dual-dimensional parallel analysis and cross-validation ensure that a single analysis covers two types of variants, avoiding redundant experiments.

[0115] Reference Figure 21 As shown, the present invention also discloses a chromosome karyotype analysis device, comprising:

[0116] The genome alignment module 11 is used to perform genome alignment on the whole exome sequencing data of the biological sample to be tested based on a preset mapping relationship file, so as to obtain the target counting matrix of the sequencing reads of the biological sample to be tested; wherein, the preset mapping relationship file is a file containing the mapping relationship between different probe capture intervals and cytogenetic bands;

[0117] The variant detection module 12 is used to perform variant detection on the whole exome sequencing data to obtain a variant detection result file of the biological sample to be tested.

[0118] The first analysis result module 13 is used to determine the sequencing read difference information and statistical significance information of each cytogenetic band interval based on the target counting matrix and the standard counting matrix, so as to obtain the first karyotype result of each cytogenetic band interval; wherein, the standard counting matrix is ​​the counting matrix of sequencing reads of biological samples with standard chromosome karyotype;

[0119] The second analysis result module 14 is used to count the number of variant sites on each chromosome of the biological sample to be tested based on the variant detection result file, and calculate the allele frequency density distribution of the target chromosome whose number of variant sites reaches a preset variant site number threshold, so as to analyze whether the target chromosome has chromosomal euploidy variation according to the allele peak characteristics corresponding to the allele frequency density distribution, so as to obtain the second karyotype result at the chromosome level.

[0120] The comprehensive analysis results module 15 is used to determine the target chromosome karyotype analysis results of the biological sample to be tested based on the first karyotype result and the second karyotype result.

[0121] As can be seen, this application discloses a method for performing genome alignment on whole-exome sequencing data of a biological sample to be tested based on a preset mapping relationship file, to obtain a target counting matrix of sequencing reads of the biological sample to be tested, and to perform variant detection on the whole-exome sequencing data to obtain a variant detection result file of the biological sample to be tested; wherein, the preset mapping relationship file is a file containing the mapping relationship between different probe capture intervals and cytogenetic bands; based on the target counting matrix and the standard counting matrix, the sequencing read difference information and statistical significance information of each cytogenetic band interval are determined to obtain the first karyotype result of each cytogenetic band interval; its In this method, the standard counting matrix is ​​a counting matrix of sequencing reads of biological samples with standard chromosome karyotypes. Based on the variant detection result file, the number of variant sites on each chromosome of the biological sample to be tested is counted, and the allele frequency density distribution of the target chromosome with a number of variant sites reaching a preset threshold is calculated. The presence of chromosomal euploidy variants in the target chromosome is analyzed based on the allele peak characteristics corresponding to the allele frequency density distribution, resulting in a second karyotype result at the chromosome level. The target chromosome karyotype analysis result of the biological sample to be tested is determined based on the first and second karyotype results. Therefore, direct analysis of whole-exome sequencing data eliminates the need for complex sample processing such as cell culture / slide preparation. Then, bioinformatics algorithms for genome alignment, counting matrix generation, variant detection, and two-dimensional analysis are used for automated processing, compressing the detection and analysis cycle. Specifically, the first karyotype result is obtained by dividing the analysis into cytogenetic bands, refining the analysis level to the cytogenetic band level, resulting in higher karyotype resolution. Finally, through two-dimensional parallel analysis and cross-validation, a single analysis covers two types of variants, avoiding repeated experiments.

[0122] Furthermore, embodiments of this application also disclose an electronic device, Figure 22 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0123] Figure 22 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the chromosome karyotype analysis method disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0124] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0125] The processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 21 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0126] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0127] The operating system 221 manages and controls the various hardware devices and computer programs 222 on the electronic device 20 to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. The operating system 221 can be Windows Server, Netware, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of performing the chromosome karyotype analysis method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the electronic device from external devices, as well as data collected by its own input / output interface 25.

[0128] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed chromosome karyotype analysis method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0129] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0130] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software module may be located in random access memory (RAM), memory, read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs (Compact Disc-Read Only Memory), or any other form of storage medium known in the art.

[0131] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0132] The solution provided by the present invention has been described in detail above. Specific examples have been used to illustrate the principle and implementation of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation and application scope based on the idea of ​​the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for chromosome karyotype analysis, characterized in that, include: Genome alignment is performed on the whole exome sequencing data of the biological sample to be tested based on a preset mapping relationship file to obtain the target counting matrix of the sequencing reads of the biological sample to be tested, and variant detection is performed on the whole exome sequencing data to obtain the variant detection result file of the biological sample to be tested; wherein, the preset mapping relationship file is a file containing the mapping relationship between different probe capture intervals and cytogenetic bands. Based on the target counting matrix and the standard counting matrix, the sequencing read difference information and statistical significance information of each cytogenetic band interval are determined to obtain the first karyotype result of each cytogenetic band interval; wherein, the standard counting matrix is ​​the counting matrix of sequencing reads of biological samples with standard chromosome karyotype; Based on the mutation detection result file, the number of mutation sites on each chromosome of the biological sample to be tested is counted, and the allele frequency density distribution of the target chromosome whose number of mutation sites reaches the preset mutation site number threshold is calculated. Based on the allele peak characteristics corresponding to the allele frequency density distribution, the chromosomal euploidy variation of the target chromosome is analyzed to obtain the second karyotype result at the chromosome level. The target chromosome karyotype analysis results of the biological sample to be tested are determined based on the first karyotype results and the second karyotype results. Before performing genome alignment on the whole exome sequencing data of the biological sample to be tested based on the preset mapping relationship file, the process also includes: Map the whole exome sequencing probe capture regions to cytogenetic band regions to generate a preset mapping relationship file containing the mapping relationship between probe capture regions and corresponding cytogenetic band regions; The step of mapping whole-exome sequencing probe capture regions to cytogenetic band regions, generating a preset mapping relationship file containing the mapping relationship between probe capture regions and corresponding cytogenetic band regions, includes: Obtain a cytogenetic band information file; the cytogenetic band information file includes chromosomes, start positions, end positions, cytogenetic band names, and relevant intervals for each chromosome; Obtain the whole exome sequencing probe file, wherein the whole exome sequencing probe file is the probe interval file corresponding to the capture probe kit used for sequencing, and the whole exome sequencing probe file includes chromosome, capture interval start position and capture interval end position; Using preset interval mapping software, the overlap between the whole exome sequencing probe capture interval in the whole exome sequencing probe file and the genomic coordinate interval of the cytogenetic band interval in the cytogenetic band information file is mapped to generate a preset mapping relationship file; The preset mapping relationship file includes information such as chromosome, start position of capture interval, end position of capture interval, corresponding cytogenetic band interval, and name of cytogenetic band interval; The process of performing variant detection on the whole exome sequencing data to obtain a variant detection result file for the biological sample to be tested includes: The whole exome sequencing data and the human reference genome are aligned to obtain a first genome alignment result file containing the position, orientation, and alignment quality information of the reads on the genome. The variant sites in the first genome comparison result file are scanned, and the allele frequency of each site is calculated based on the number of reads at the variant sites and the total number of reads covered by the variant sites, so as to obtain a variant detection result file containing genotypic information of all sites of the biological sample to be tested.

2. The chromosome karyotype analysis method according to claim 1, characterized in that, The process of performing genome alignment on the whole exome sequencing data of the biological sample to be tested based on a preset mapping relationship file to obtain a target counting matrix of sequencing reads of the biological sample to be tested includes: Based on the preset mapping relationship file, the counting information of sequencing reads in the cytogenetic band interval is statistically analyzed, and a corresponding target counting matrix is ​​generated based on the counting information.

3. The chromosome karyotype analysis method according to claim 1, characterized in that, Before determining the sequencing read difference information and statistical significance information of each cytogenetic band interval based on the target counting matrix and the standard counting matrix, the method further includes: Standard biological samples characterized by normal chromosome karyotypes were selected, and reads were compared between the whole exome sequencing data of the standard biological samples and the human reference genome to obtain a second genome alignment result file containing the position information, orientation information and alignment quality information of the reads on the genome. Based on the mapping relationship file, the sequencing read counts of each cytogenetic band interval in the second genome alignment result file are statistically analyzed to generate a standard counting matrix.

4. The chromosome karyotype analysis method according to claim 3, characterized in that, The determination of sequencing read difference information and statistical significance information for each cytogenetic band interval based on the target counting matrix and the standard counting matrix includes: Based on the target counting matrix and the standard counting matrix, the sequencing read difference value and significance p-value of each cytogenetic band interval are calculated.

5. A chromosome karyotype analysis device, characterized in that, include: The genome alignment module is used to perform genome alignment on the whole exome sequencing data of the biological sample to be tested based on a preset mapping relationship file, so as to obtain the target counting matrix of the sequencing reads of the biological sample to be tested; wherein, the preset mapping relationship file is a file containing the mapping relationship between different probe capture intervals and cytogenetic bands; The variant detection module is used to perform variant detection on the whole exome sequencing data to obtain a variant detection result file of the biological sample to be tested. The first analysis results module is used to determine the sequencing read difference information and statistical significance information of each cytogenetic band interval based on the target counting matrix and the standard counting matrix, so as to obtain the first karyotype result of each cytogenetic band interval; wherein, the standard counting matrix is ​​the counting matrix of sequencing reads of biological samples with standard chromosome karyotype; The second analysis result module is used to count the number of variant sites on each chromosome of the biological sample to be tested based on the variant detection result file, and to calculate the allele frequency density distribution of the target chromosome whose number of variant sites reaches a preset variant site number threshold, so as to analyze whether the target chromosome has chromosomal euploidy variation according to the allele peak characteristics corresponding to the allele frequency density distribution, so as to obtain the second karyotype result at the chromosome level. The comprehensive analysis results module is used to determine the target chromosome karyotype analysis results of the biological sample to be tested based on the first karyotype result and the second karyotype result. The chromosome karyotype analysis device is also used to map the whole exome sequencing probe capture interval to the cytogenetic band interval, and generate a preset mapping relationship file containing the mapping relationship between the probe capture interval and the corresponding cytogenetic band interval; The chromosome karyotype analysis device is also used to acquire cytogenetic band information files; the cytogenetic band information files include chromosomes, start positions, end positions, cytogenetic band names, and related intervals for each chromosome; acquire whole-exome sequencing probe files, the whole-exome sequencing probe files being probe interval files corresponding to the capture probe kit used for sequencing, and the whole-exome sequencing probe files including chromosomes, capture interval start positions, and capture interval end positions; using preset interval mapping software, the overlap between the whole-exome sequencing probe capture intervals in the whole-exome sequencing probe files and the genomic coordinate intervals of the cytogenetic band intervals in the cytogenetic band information files is mapped to generate a preset mapping relationship file; the preset mapping relationship file includes chromosomes, capture interval start positions, capture interval end positions, corresponding cytogenetic band intervals, and cytogenetic band interval names; The variant detection module is specifically used to perform read alignment on the whole exome sequencing data and the human reference genome to obtain a first genome alignment result file containing the position information, orientation information, and alignment quality information of the reads on the genome; scan the variant sites in the first genome alignment result file, and calculate the allele frequency of each site based on the number of reads at the variant sites and the total number of reads covered by the variant sites, to obtain a variant detection result file containing genotype information of all sites of the biological sample to be tested.

6. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the chromosome karyotype analysis method as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the chromosome karyotype analysis method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Methods and systems for detection of abnormal karyotypes

    CA3014292A1

  • Methods and systems for detection of abnormal karyotypes

    CN109074426A