Method, device and application for detecting tumor molecular typing based on low-depth whole genome sequencing

Through low-deep whole genome sequencing method, the data were divided into different bin groups, and the copy number and mutation region ratio were calculated using CNVkit software, and a scoring model was established, which solved the accuracy and cost of brain tumor molecular typing detection in the existing technology, and achieved accurate molecular typing of multiple cancer species.

CN119132388BActive Publication Date: 2025-08-15HENAN CANCER HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411153968.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2025-08-15
Estimated Expiration
2044-08-21

AI Technical Summary

Technical Problem

In the molecular typing detection of brain tumors, it is difficult to accurately detect copy number variations in multiple cancer types and chromosomes throughout the region, and it is costly and difficult to extract samples. Machine learning modeling has subjectivity and data accumulation problems.

Method used

The low-deep whole genome sequencing method was used to divide the data into unfixed bin groups and fixed bin groups. The CNVkit software was used to calculate the copy number and mutation region ratio, establish a scoring model, and judge the copy number variation types at each structural level of the chromosome to achieve accurate tumor molecular typing.

Benefits of technology

The accuracy of chromosome multi-level copy number variation detection of different tumor subtypes is achieved, the detection cost is reduced, and it is suitable for molecular typing detection of various cancer species.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119132388B_ABST
    Figure CN119132388B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, and application for detecting tumor molecular typing based on low-depth whole-genome sequencing. The method comprises: independently dividing low-depth whole-genome sequencing data into non-fixed bin groups and fixed bin groups; using the copy numbers of the non-fixed bin groups and the fixed bin groups to calculate the average copy number and mutation region ratio of each group; establishing a scoring model based on the average copy number and mutation region ratio to determine the copy number variation type at each structural level of the chromosome, and then deriving the corresponding molecular typing; wherein the non-fixed bin group includes a first target region belonging to the exon and an anti-target region outside the exon; the fixed bin group includes a second target region of 1M in length and a third target region of 10K in length obtained by continuous division. In this way, the copy number variation type of the chromosome at different hierarchical structural positions is obtained, thereby achieving more accurate molecular typing detection of the size of tumor chromosomes at each level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of gene detection technology, and specifically to a method, device and application for detecting tumor molecular typing based on low-depth whole genome sequencing. Background Art

[0002] Brain tumors include gliomas, meningiomas, pituitary tumors, ependymomas, medulloblastomas, central nervous system lymphomas, hereditary brain tumors, etc. With the release of the latest version of the WHO CNS5 guidelines in 2021, molecular diagnosis has become increasingly important in the field of brain tumors. The test can obtain the subtype of the tumor to determine the patient's prognosis and guide the patient's clinical treatment and medication plan. The complex molecular typing requirements of various types of cancer include methylation, gene mutations, and mutations in different regions of chromosomes. Common detection schemes for second-generation sequencing (targeted sequencing or whole exome sequencing) have the defect of insufficient detection range.

[0003] Currently available detection methods are limited to copy number detection of a certain level of size in a chromosome region, or are limited to the molecular typing of a specific cancer type. To complete copy number detection of multiple cancer types and entire chromosome regions, additional high time and financial costs are required. In addition, since most clinical samples are FFPE samples, it is difficult to extract nucleic acids from samples, and some experimental-based detection schemes may be difficult to implement; and various combination schemes require too large a sample size, which objectively limits the feasibility of detection. Using a large number of samples for machine learning modeling to identify molecular typing, through the process of bioinformatics and statistical modeling, will be affected by the sample data situation and parameter adjustment, has a certain degree of subjectivity, and requires sufficient prior data accumulation, and it is difficult to expand to brain tumors.

[0004] Therefore, how to provide a method for detecting multi-level chromosome copy number variations in different tumor subtypes is very important for realizing tumor molecular typing detection. Summary of the Invention

[0005] The main purpose of the present invention is to provide a method, device and application for detecting tumor molecular typing based on low-depth whole genome sequencing, so as to solve the problem of poor accuracy in detecting copy number variations of multi-level chromosome structures in the existing technology.

[0006] In order to achieve the above-mentioned purpose, according to the first aspect of the present invention, a method for detecting tumor molecular typing based on low-depth whole-genome sequencing is provided, the method comprising: S1, independently dividing the low-depth whole-genome sequencing data into a non-fixed bin group and a fixed bin group; S2, using the copy numbers of the non-fixed bin group and the fixed bin group to calculate the average copy number and mutation region ratio of each group; S3, establishing a scoring model based on the average copy number and mutation region ratio to judge the copy number variation type of each structural level of the chromosome, and then deriving the corresponding molecular typing; wherein, the non-fixed bin group includes a first target region belonging to the exon and an anti-target region outside the exon; the fixed bin group includes a second target region with a length of 1M and a third target region with a length of 10K obtained by continuous division; each structural level of the chromosome includes: the entire chromosome, chromosome arm, chromosome region and chromosome band.

[0007] Furthermore, before partitioning, the inaccessible portion of the second-generation sequencing in the low-depth whole-genome sequencing data is removed; preferably, the low-depth whole-genome sequencing data is a BAM file obtained by comparing with the human reference genome after quality control.

[0008] Furthermore, the copy numbers of the non-fixed bin group and the fixed bin group are calculated respectively using CNVkit software; preferably, before calculating the average copy number and mutation region ratio of each group, the method further includes: annotating the position of the region in the chromosome whose copy number is calculated using CNVkit software in the non-fixed bin group and the fixed bin group; judging the copy number variation type of each region in the non-fixed bin group and the fixed bin group based on the copy number variation threshold; the copy number variation types include: loss, loss risk Loss_risk, amplification Gain and amplification risk Gain_risk.

[0009] Furthermore, calculating the average copy number and mutation region ratio of each group includes: calculating the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio of each region whose copy number is calculated by CNVkit software in the non-fixed bin group at the chromosome whole and chromosome arm levels; calculating the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio of the region whose copy number is calculated by CNVkit software in the second target region and the third target region in the fixed bin group respectively; average copy number = (the length of the region whose copy number is calculated by CNVkit software in the object * the corresponding region copy number) / the sum of the lengths of all regions in the object for which copy numbers are calculated; the first mutation region ratio = the sum of the lengths of the Loss regions in the object / the sum of the lengths of all regions in the object for which copy numbers are calculated; the second mutation region ratio = the sum of the lengths of the Loss_risk regions in the object / the sum of the lengths of all regions in the object for which copy numbers are calculated; the third mutation region ratio = the sum of the lengths of the Gain regions in the object / the sum of the lengths of all regions in the object for which copy numbers are calculated; the fourth mutation region ratio = the sum of the lengths of the Gain_risk regions in the object / the sum of the lengths of all regions in the object for which copy numbers are calculated; wherein the object is: the entire chromosome, the chromosome arm, the second target region or the third target region.

[0010] Furthermore, S3 includes: establishing a first scoring model using the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio at the chromosome overall level to determine the copy number variation type of the chromosome as a whole; establishing a second scoring model using the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio at the chromosome arm level to determine the copy number variation type of the chromosome arm; establishing a third scoring model based on the copy number variation type of the chromosome as a whole and the copy number variation type of the chromosome arm, using the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio of the second target region, and then determining the copy number variation type of the chromosome region; establishing a fourth scoring model based on the copy number variation type of the chromosome as a whole and the copy number variation type of the chromosome arm, using the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio of the third target region, and then determining the copy number variation type of the chromosome band.

[0011] Furthermore, the copy number variation type of the entire chromosome and chromosome arm is determined according to the following judgment criteria: when the average copy number of the judgment object is ≤ the baseline average copy number * 0.8 and the proportion of the first mutation region of the judgment object is ≥ 85%, the copy number variation type of the judgment object is Loss; when the average copy number of the judgment object is ≤ the baseline average copy number * 0.9 and the sum of the proportion of the first mutation region of the judgment object and the proportion of the second mutation region of the judgment object is ≥ 85%, the copy number variation type of the judgment object is Loss_risk; when the average copy number of the judgment object is ≥ the baseline average copy number * 1.2 and the proportion of the third mutation region of the judgment object is ≥ 85%, the copy number variation type of the judgment object is Gain; when the average copy number of the judgment object is ≥ the baseline average copy number * 1.1 and the proportion of the third mutation region of the judgment object and the proportion of the fourth mutation region of the judgment object are ≥ 85%, the copy number variation type of the judgment object is Gain_risk; the judgment object is the entire chromosome or chromosome arm.

[0012] Furthermore, the copy number variation type of the chromosome region is determined according to the following judgment criteria: when the copy number variation type of the chromosome arm in the chromosome region is not amplification Gain, when the average copy number of the second target region is ≤ the baseline average copy number * 0.8 and the ratio of the first mutation region in the second target region is ≥ 85%, then the copy number variation type of the chromosome region is deletion Loss; when the copy number variation type of the chromosome arm in the chromosome region is not amplification Gain, when the average copy number of the second target region is ≤ the baseline average copy number * 0.9 and the sum of the ratio of the first mutation region in the second target region and the ratio of the second mutation region in the second target region is ≥ 85%, then the copy number variation type of the chromosome region is The value of the copy number variation type in the chromosome arm of the chromosome region is Loss_risk; when the copy number variation type in the chromosome arm of the chromosome region is not Loss, when the average copy number of the second target region is ≥ the baseline average copy number * 1.2 and the ratio of the third mutation region in the second target region is ≥ 85%, then the copy number variation type in the chromosome region is Gain; when the copy number variation type in the chromosome arm of the chromosome region is not Loss, when the average copy number of the second target region is ≥ the baseline average copy number * 1.1 and the ratio of the third mutation region in the second target region to the fourth mutation region in the second target region is ≥ 85%, then the copy number variation type in the chromosome region is Gain_risk.

[0013] Furthermore, the copy number variation type of the chromosome band is determined according to the following judgment criteria: when the copy number variation type of the chromosome arm of the chromosome band is not amplification Gain, when the average copy number of the third target region is ≤ the baseline average copy number * 0.8 and the proportion of the first mutation region of the third target region is ≥ 85%, then the copy number variation type of the chromosome band is deletion Loss; when the copy number variation type of the chromosome arm of the chromosome band is not amplification Gain, when the average copy number of the third target region is ≤ the baseline average copy number * 0.9 and the sum of the proportion of the first mutation region of the third target region and the proportion of the second mutation region of the third target region is ≥ 85%, then the copy number variation type of the chromosome band is It is deletion risk Loss_risk; when the copy number variation type of the chromosome arm of the chromosome band is not deletion Loss, when the average copy number of the third target region is ≥ the baseline average copy number * 1.2 and the proportion of the third mutation region of the third target region is ≥ 85%, then the copy number variation type of the chromosome band is amplification Gain; when the copy number variation type of the chromosome arm of the chromosome band is not deletion Loss, when the average copy number of the third target region is ≥ the baseline average copy number * 1.1 and the proportion of the third mutation region of the third target region and the proportion of the fourth mutation region of the third target region are ≥ 85%, then the copy number variation type of the chromosome band is amplification risk Gain_risk.

[0014] In order to achieve the above-mentioned purpose, according to the second aspect of the present invention, a device for detecting tumor molecular typing based on low-depth whole-genome sequencing is provided, which includes: a bin division module, which is configured to independently divide low-depth whole-genome sequencing data into non-fixed bin groups and fixed bin groups; a calculation module, which is configured to use the copy numbers of the non-fixed bin group and the fixed bin group to calculate the average copy number and mutation region ratio of each group; a judgment module, which is configured to establish a scoring model based on the average copy number and mutation region ratio, judge the copy number variation type of each structural level of the chromosome, and then derive the corresponding molecular typing; wherein, the non-fixed bin group includes a first target region belonging to the exon and an anti-target region outside the exon; the fixed bin group includes a second target region with a length of 1M and a third target region with a length of 10K obtained by continuous division; each structural level of the chromosome includes: the entire chromosome, chromosome arm, chromosome region and chromosome band.

[0015] Furthermore, in the bin partitioning module, before partitioning, the inaccessible part of the second-generation sequencing in the low-depth whole-genome sequencing data is removed; preferably, the low-depth whole-genome sequencing data is a BAM file obtained by comparing with the human reference genome after quality control.

[0016] Furthermore, the calculation module includes: a CNVkit software calculation unit, which is configured to use the CNVkit software to calculate the copy numbers of the non-fixed bin group and the fixed bin group respectively; preferably, before calculating the average copy number and mutation region ratio of each group, the calculation module also includes: an annotation unit, which is configured to annotate the position of the region in the chromosome in the non-fixed bin group and the fixed bin group whose copy number is calculated using the CNVkit software; a threshold judgment unit, which is configured to judge the copy number variation type of each region in the non-fixed bin group and the fixed bin group based on the threshold of copy number variation; the copy number variation types include: loss, loss risk Loss_risk, amplification Gain and amplification risk Gain_risk.

[0017] Furthermore, the calculation module also includes: a first calculation unit, which is configured to calculate the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio of each region whose copy number is calculated by CNVkit software in the non-fixed bin group at the chromosome whole level and the chromosome arm level respectively; a second calculation unit, which is configured to calculate the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio of the region whose copy number is calculated by CNVkit software in the second target region and the third target region in the fixed bin group respectively; average copy number = (length of the region whose copy number is calculated by CNVkit software in the object * corresponding region copy number) The first mutation region ratio = the sum of the lengths of the Loss regions in the object / the sum of the lengths of all regions in the object for which copy numbers are calculated; the second mutation region ratio = the sum of the lengths of the Loss_risk regions in the object / the sum of the lengths of all regions in the object for which copy numbers are calculated; the third mutation region ratio = the sum of the lengths of the Gain regions in the object / the sum of the lengths of all regions in the object for which copy numbers are calculated; the fourth mutation region ratio = the sum of the lengths of the Gain_risk regions in the object / the sum of the lengths of all regions in the object for which copy numbers are calculated; wherein the object is: the entire chromosome, the chromosome arm, the second target region or the third target region.

[0018] Furthermore, the judgment module includes: a chromosome overall copy number variation type judgment unit, which is configured to use the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio at the chromosome overall level to establish a first scoring model to judge the copy number variation type of the chromosome as a whole; a chromosome arm copy number variation type judgment unit, which is configured to use the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio at the chromosome arm level to establish a second scoring model to judge the copy number variation type of the chromosome arm; a chromosome region copy number variation type judgment unit, which is configured to use the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio at the chromosome arm level to establish a second scoring model to judge the copy number variation type of the chromosome arm; The copy number variation type of the entire chromosome and the copy number variation type of the chromosome arm are determined by using the average copy number of the second target region, the proportion of the first mutation region, the proportion of the second mutation region, the proportion of the third mutation region and the proportion of the fourth mutation region to establish a third scoring model, thereby determining the copy number variation type of the chromosome region; the chromosome band copy number variation type judgment unit is configured to be based on the copy number variation type of the entire chromosome and the copy number variation type of the chromosome arm, and by using the average copy number of the third target region, the proportion of the first mutation region, the proportion of the second mutation region, the proportion of the third mutation region and the proportion of the fourth mutation region to establish a fourth scoring model, thereby determining the copy number variation type of the chromosome band.

[0019] Furthermore, the copy number variation type judgment criteria in the chromosome whole copy number variation type judgment unit and the chromosome arm copy number variation type judgment unit include: when the average copy number of the judgment object is ≤ the baseline average copy number * 0.8 and the first mutation region ratio of the judgment object is ≥ 85%, the copy number variation type of the judgment object is Loss; when the average copy number of the judgment object is ≤ the baseline average copy number * 0.9 and the sum of the first mutation region ratio of the judgment object and the second mutation region ratio of the judgment object is ≥ 85%, the copy number variation type of the judgment object is Loss_risk; when the average copy number of the judgment object is ≥ the baseline average copy number * 1.2 and the third mutation region ratio of the judgment object is ≥ 85%, the copy number variation type of the judgment object is Gain; when the average copy number of the judgment object is ≥ the baseline average copy number * 1.1 and the third mutation region ratio of the judgment object and the fourth mutation region ratio of the judgment object are ≥ 85%, the copy number variation type of the judgment object is Gain_risk; the judgment object is the entire chromosome or the chromosome arm.

[0020] Furthermore, the judgment criteria for the copy number variation type of the chromosome region in the chromosome region copy number variation type judgment unit include: when the copy number variation type of the chromosome arm in the chromosome region is not amplification Gain, when the average copy number of the second target region is ≤ the baseline average copy number * 0.8 and the first mutation region ratio of the second target region is ≥85%, then the copy number variation type of the chromosome region is deletion Loss; when the copy number variation type of the chromosome arm in the chromosome region is not amplification Gain, when the average copy number of the second target region is ≤ the baseline average copy number * 0.9 and the sum of the first mutation region ratio of the second target region and the second mutation region ratio of the second target region is ≥85%, then the copy number variation type of the chromosome region is deletion Loss. The copy number variation type is deletion risk Loss_risk; when the copy number variation type of the chromosome arm in the chromosome region is not deletion Loss, when the average copy number of the second target region is ≥ the baseline average copy number * 1.2 and the proportion of the third mutation region in the second target region is ≥ 85%, then the copy number variation type of the chromosome region is amplification Gain; when the copy number variation type of the chromosome arm in the chromosome region is not deletion Loss, when the average copy number of the second target region is ≥ the baseline average copy number * 1.1 and the proportion of the third mutation region in the second target region and the proportion of the fourth mutation region in the second target region are ≥ 85%, then the copy number variation type of the chromosome region is amplification risk Gain_risk.

[0021] Furthermore, the judgment criteria for the copy number variation type of the chromosome band in the chromosome band copy number variation type judgment unit include: when the copy number variation type of the chromosome arm of the chromosome band is not amplification Gain, when the average copy number of the third target region is ≤ the baseline average copy number * 0.8 and the first mutation region ratio of the third target region is ≥ 85%, then the copy number variation type of the chromosome band is deletion Loss; when the copy number variation type of the chromosome arm of the chromosome band is not amplification Gain, when the average copy number of the third target region is ≤ the baseline average copy number * 0.9 and the sum of the first mutation region ratio of the third target region and the second mutation region ratio of the third target region is ≥ 85%, then the copy number variation type of the chromosome band is deletion Loss. The copy number variation type is deletion risk Loss_risk; when the copy number variation type of the chromosome arm of the chromosome band is not deletion Loss, when the average copy number of the third target region is ≥ baseline average copy number * 1.2 and the proportion of the third mutation region of the third target region is ≥ 85%, then the copy number variation type of the chromosome band is amplification Gain; when the copy number variation type of the chromosome arm of the chromosome band is not deletion Loss, when the average copy number of the third target region is ≥ baseline average copy number * 1.1 and the proportion of the third mutation region of the third target region and the proportion of the fourth mutation region of the third target region are ≥ 85%, then the copy number variation type of the chromosome band is amplification risk Gain_risk.

[0022] In order to achieve the above-mentioned purpose, according to the third aspect of the present invention, a computer-readable storage medium is provided, which includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the above-mentioned method for detecting tumor molecular typing based on low-depth whole genome sequencing.

[0023] In order to achieve the above-mentioned object, according to the fourth aspect of the present invention, a processor is provided, which is used to run a program, wherein the program runs the above-mentioned method for detecting tumor molecular typing based on low-depth whole genome sequencing.

[0024] By applying the technical solution of the present invention, sWGS sequencing data are divided into different bin combinations according to different standards for subsequent copy number analysis, including being divided into non-fixed bin groups based on whether they are exons and being divided into fixed bin groups based on 1M and 10K lengths. The copy number data of different bin groups are judged to determine the copy number variation types of chromosomes at different hierarchical structural positions, and then the corresponding molecular typing is obtained, thereby achieving more accurate molecular typing detection of the sizes of tumor chromosomes at each level, and further being suitable for the detection of molecular typing of different cancer subspecies. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0026] Figure 1 A flow chart of the detection method of the present application is shown;

[0027] Figure 2 The figure shows the detection results of the whole chromosomes of 12 samples using sWGS in Example 1;

[0028] Figure 3 The figure shows the result of sWGS detection of sample 1 in Example 1;

[0029] Figure 4 The results of Panel+Marker SNP detection for sample 1 in Example 1 are shown, wherein Chromosome: chromosome; Allele Frequency: genotype frequency; Chromosomal Location: chromosome location; type: type; abnormal: abnormal; normal: normal;

[0030] Figure 5 The results of Panel+Marker SNP detection for sample 1 in Example 1 are shown, wherein CopyNumber (log2 ratio): copy number (log2 ratio); gain: amplification; deletion: deletion; cnvType: cnv type;

[0031] Figure 6 The figure shows the results of the gene chip test on Sample 1 in Example 1, where 1-22 and X correspond to the chromosome names. The copy number of the large fragment of the entire chromosome is displayed on the left side. Red represents amplification, blue represents deletion, and white represents neutrality. The darker the color, the more severe the copy number variation. The gene chip test will set a threshold to determine whether amplification has occurred. In the figure, A represents chr7 amplification (dark red), B represents chr12 amplification (light red), and C represents chr18 amplification (light red);

[0032] Figure 7 The result of sWGS detection of sample 11 in Example 1 is shown;

[0033] Figure 8The results of the gene chip test on sample 11 in Example 1 are shown; among them, the labels (1-22, X and Y) in the white box in the lower left corner of the nuclear view correspond to the chromosome name, and the copy number variation occurring in the chromosome segment is displayed to the right of each chromosome. Red represents deletion, blue represents amplification, and no annotation represents neutrality. In the figure, A represents 1p neutrality, B represents chr7 neutrality, C represents chr10 neutrality, D represents a small fragment deletion on chr19, E represents a small fragment amplification on chr19, and F represents 19q neutrality. DETAILED DESCRIPTION

[0034] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0035] As mentioned in the background technology, the existing method for detecting molecular typing of tumors is limited to a certain specific cancer or part of the hierarchical position of the chromosome, and it is impossible to detect more accurate molecular typing results of the whole region chromosome, and the cost is relatively high. The existing method exists by traditional fluorescence in situ hybridization (FISH) to verify the copy number variation of the whole chromosome and the arm region; by combining a gene chip with a certain bioinformatics method, the copy number variation of the chromosome region and the band region is detected; by professionally customized targeted capture gene package (Panel) combined with marker single nucleotide polymorphism site (SNP) to detect the molecular typing of specific cancers. Therefore, in order to obtain accurate molecular typing of the whole region chromosome and expand the scope of application of cancer, the present application is intended to provide a detection method for molecular typing based on low-depth whole genome sequencing results.

[0036] In a first typical embodiment of the present application, a method for detecting tumor molecular typing based on low-depth whole-genome sequencing is provided, the method comprising: S1, independently dividing the low-depth whole-genome sequencing data into a non-fixed bin group and a fixed bin group; S2, using the copy numbers of the non-fixed bin group and the fixed bin group to calculate the average copy number and mutation region ratio of each group; S3, establishing a scoring model based on the average copy number and mutation region ratio to determine the copy number variation type at each structural level of the chromosome, and then deriving the corresponding molecular typing; wherein the non-fixed bin group includes a first target region belonging to the exon and an anti-target region outside the exon; the fixed bin group includes a second target region of 1M in length and a third target region of 10K in length obtained by continuous division; each structural level of the chromosome includes: the entire chromosome, chromosome arm, chromosome region, and chromosome band. wherein the sequencing depth of the tumor sample is 1-5×.

[0037] This application divides the low-depth whole-genome sequencing results into bins with different standards, obtaining a non-fixed bin group divided by whether it is an exon and a fixed bin group divided continuously by length 1M and 10K. The two different bin groupings are used for copy number calculation, and the average copy number and the proportion of mutation areas are obtained based on the copy number calculation. These serve as the basis for judging the copy number variation type of each chromosome level (whole, arm, region, band), establishing a scoring model, determining the copy number variation type of the entire chromosome, and further obtaining accurate tumor molecular typing results.

[0038] Among them, copy number variation at different chromosome levels is one of the criteria for molecular typing of brain tumors. Based on the WHO guidelines, 353 brain tumor-related variations were screened (covering molecular markers related to diagnosis, grading, and prognosis of different cancer types), including chromosome level (such as chr7_AMP, chr10_DEL), broken arm long arm level (such as 1p_DEL, 19q_DEL), gene or above level (such as MET amplification, CDK4 amplification, MYC amplification, etc.). The copy number variation at different chromosome levels can be combined with key gene mutations (Snp+Indel), methylation, etc. as diagnostic criteria for molecular typing of brain tumors, and each standard is independent and parallel to each other.

[0039] In order to further improve the efficiency of bin division and obtain more valid data for data analysis, it is necessary to clean and filter the low-depth whole-genome sequencing data. In a preferred embodiment, before binning, the inaccessible part of the second-generation sequencing in the low-depth whole-genome sequencing data is removed; preferably, the low-depth whole-genome sequencing data is a BAM file obtained by comparison with the human reference genome after quality control.

[0040] Due to the short read lengths of next-generation sequencing (NGS), many non-uniquely aligned regions lack specificity. Variants within these regions, whether point mutations or CNVs / SVs, are not highly reliable. Even within uniquely aligned regions, library construction can favor regions with GC aberrations. Github summarizes the following characteristics of these NGS-inaccessible regions: 1. Repetitive sequence regions; 2. GC aberrations (high and low GC content); 3. Telomeres and centrosomes; 4. Pseudogenes / highly homologous genes; and 5. Regions with low alignment scores.

[0041] The Encode data website has a section that evaluates the uniqueness of alignments of sequences of varying lengths across the genome. There's also an open-source project on GitHub for this purpose (https: / / github.com / Boyle-Lab / Blacklist). The CNVkit manual for the open-source software provides access files for high-quality accessible regions and recommends using the software in preliminary analysis to remove regions with lower-quality second-generation sequencing results to obtain more accurate analysis results (https: / / cnvkit.readthedocs.io / en / stable / pipeline.html). This database can be used to remove inaccessible regions from second-generation sequencing.

[0042] Any method capable of calculating copy number data is applicable to the present application. In a preferred embodiment, the copy numbers of the non-fixed bin group and the fixed bin group are calculated using CNVkit software.

[0043] Among them, this application completed the open source software testing (CNVkit, Control-Free, Sequeza) during the testing phase, among which CNVkit: 17 direct citations in Pubmed within 5 years, and 700 total citations; Control-Free: 63 direct citations in Pubmed within 5 years, and 471 total citations; Sequeza: 12 direct citations in Pubmed within 5 years, and 320 total citations. CNVkit can support low-depth WGS data and is more accurate in detecting large fragment copy numbers at the chromosome region level; Control-Free has poor accuracy in detecting results in special regions when calculating allele frequencies due to its low depth and single sample; Sequeza has difficulty obtaining valid data when calculating allele frequencies due to its low depth and single sample, and the deviation is large. In the end, CNVkit was selected as the solution that meets our detection range expectations.

[0044] Utilizing the existing threshold standards for determining copy number variation, the copy number variation type of the regions in each bin group where the copy number has been detected is determined, facilitating the subsequent determination of the copy number variation type at all levels of the entire chromosome. In a preferred embodiment, before calculating the average copy number and mutation region ratio of each group, the method further comprises: annotating the positions of the regions in the chromosomes where the copy numbers are calculated using the CNVkit software in the non-fixed bin group and the fixed bin group; determining the copy number variation type of each region in the non-fixed bin group and the fixed bin group based on the copy number variation threshold; the copy number variation types include: loss, loss risk Loss_risk, amplification Gain, and amplification risk Gain_risk.

[0045] For chromosomes at different levels, the copy numbers of different bin groups are calculated to obtain relevant data that can determine the types of chromosome copy number variation at different levels. In a preferred embodiment, calculating the average copy number and mutation region ratio of each group includes: calculating the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio of each region whose copy number is calculated by CNVkit software in the non-fixed bin group at the chromosome whole and chromosome arm levels respectively; calculating the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio of the region whose copy number is calculated by CNVkit software in the second target region and the third target region in the fixed bin group respectively;

[0046] Average copy number = (the sum of the length of the region in the subject where the copy number was calculated by CNVkit software * the copy number of the corresponding region) / the sum of the lengths of all regions in the subject where the copy number was calculated;

[0047] The first mutation region ratio = the sum of the lengths of the Loss regions in the subject / the sum of the lengths of all regions with calculated copy numbers in the subject;

[0048] Second mutation region ratio = sum of the lengths of the Loss_risk regions in the object / sum of the lengths of all regions with calculated copy numbers in the object;

[0049] The third mutation region ratio = the sum of the Gain region lengths in the subject / the sum of the lengths of all regions with calculated copy numbers in the subject;

[0050] Fourth, the ratio of mutation regions = the sum of the lengths of the Gain_risk regions in the subject / the sum of the lengths of all regions with calculated copy numbers in the subject;

[0051] The target is: the entire chromosome, a chromosome arm, a second target region or a third target region.

[0052] The copy number-related data of different chromosome levels (whole, arm) or different target regions (second and third target regions) are used to judge the copy number variation type, and a related scoring model is established to obtain the copy number variation type of the chromosome in the entire region, and further obtain the tumor molecular typing. In a preferred embodiment, S3 includes: using the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio at the chromosome overall level to establish a first scoring model to judge the copy number variation type of the chromosome as a whole; using the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio at the chromosome arm level to establish a second scoring model to judge the copy number variation type of the chromosome arm.

[0053] Based on the copy number variation type of the entire chromosome and the copy number variation type of the chromosome arm, a third scoring model is established using the average copy number of the second target region, the proportion of the first mutation region, the proportion of the second mutation region, the proportion of the third mutation region, and the proportion of the fourth mutation region, to further determine the copy number variation type of the chromosome region; based on the copy number variation type of the entire chromosome and the copy number variation type of the chromosome arm, a fourth scoring model is established using the average copy number of the third target region, the proportion of the first mutation region, the proportion of the second mutation region, the proportion of the third mutation region, and the proportion of the fourth mutation region, to further determine the copy number variation type of the chromosome band.

[0054] In a preferred embodiment, the copy number variation type of the entire chromosome and chromosome arm is determined according to the following judgment criteria: when the average copy number of the determined object is ≤ the baseline average copy number * 0.8 and the ratio of the first mutation region of the determined object is ≥ 85%, the copy number variation type of the determined object is Loss; when the average copy number of the determined object is ≤ the baseline average copy number * 0.9 and the sum of the ratio of the first mutation region of the determined object and the ratio of the second mutation region of the determined object is ≥ 85%, the copy number variation type of the determined object is Loss_risk; when the average copy number of the determined object is ≥ the baseline average copy number * 1.2 and the ratio of the third mutation region of the determined object is ≥ 85%, the copy number variation type of the determined object is Gain; when the average copy number of the determined object is ≥ the baseline average copy number * 1.1 and the ratio of the third mutation region of the determined object and the ratio of the fourth mutation region of the determined object are ≥ 85%, the copy number variation type of the determined object is Gain_risk; the determination object is the entire chromosome or a chromosome arm.

[0055] The baseline average copy number is calculated based on a baseline established using WGS data from healthy individuals according to the official CNVkit software manual. The 0.8x ratio is derived from the set of positive and negative results in the reference test sample. The 85% ratio is determined because the largest long arm of all chromosomes in the human reference genome does not exceed 85% of the total chromosome count. This ratio is further applied to chromosome determination at all levels.

[0056] In order to further accurately judge the copy number variation type of the chromosome region, the strong positive results of the chromosome arms obtained above are used for filtering. That is, if the copy number variation type of the chromosome arm is Gain, the copy number variation type of the chromosome region containing the chromosome arm is unlikely to be a Loss.

[0057] In a preferred embodiment, the copy number variation type of the chromosome region is determined according to the following judgment criteria: when the copy number variation type of the chromosome arm in the chromosome region is not amplification Gain, when the average copy number of the second target region is ≤ the baseline average copy number * 0.8 and the first mutation region ratio of the second target region is ≥85%, then the copy number variation type of the chromosome region is deletion Loss; when the copy number variation type of the chromosome arm in the chromosome region is not amplification Gain, when the average copy number of the second target region is ≤ the baseline average copy number * 0.9 and the sum of the first mutation region ratio of the second target region and the second mutation region ratio of the second target region is ≥85%, then the copy number variation type of the chromosome region is deletion Loss. The aberrant type is deletion risk Loss_risk; when the copy number variation type of the chromosome arm in the chromosome region is not deletion Loss, when the average copy number of the second target region is ≥ the baseline average copy number * 1.2 and the ratio of the third mutation region of the second target region is ≥ 85%, then the copy number variation type of the chromosome region is amplification Gain; when the copy number variation type of the chromosome arm in the chromosome region is not deletion Loss, when the average copy number of the second target region is ≥ the baseline average copy number * 1.1 and the ratio of the third mutation region of the second target region to the fourth mutation region of the second target region is ≥ 85%, then the copy number variation type of the chromosome region is amplification risk Gain_risk.

[0058] In order to further accurately judge the type of chromosome band copy number variation, the strong positive results of the chromosome arms obtained above are used for filtering, that is, if the copy number variation type of the chromosome arm is Gain, then the copy number variation type of the chromosome band containing the chromosome arm is unlikely to be a Loss.

[0059] In a preferred embodiment, the copy number variation type of the chromosome band is determined according to the following judgment criteria: when the copy number variation type of the chromosome arm of the chromosome band is not amplification Gain, when the average copy number of the third target region is ≤ the baseline average copy number * 0.8 and the first mutation region ratio of the third target region is ≥ 85%, then the copy number variation type of the chromosome band is deletion Loss; when the copy number variation type of the chromosome arm of the chromosome band is not amplification Gain, when the average copy number of the third target region is ≤ the baseline average copy number * 0.9 and the sum of the first mutation region ratio of the third target region and the second mutation region ratio of the third target region is ≥ 85%, then the copy number variation type of the chromosome band is deletion Loss. The aberrant type is deletion risk Loss_risk; when the copy number variation type of the chromosome arm of the chromosome band is not deletion Loss, when the average copy number of the third target region is ≥ the baseline average copy number * 1.2 and the proportion of the third mutation region of the third target region is ≥ 85%, then the copy number variation type of the chromosome band is amplification Gain; when the copy number variation type of the chromosome arm of the chromosome band is not deletion Loss, when the average copy number of the third target region is ≥ the baseline average copy number * 1.1 and the proportion of the third mutation region of the third target region and the proportion of the fourth mutation region of the third target region are ≥ 85%, then the copy number variation type of the chromosome band is amplification risk Gain_risk.

[0060] In a second typical embodiment of the present application, a device for detecting tumor molecular typing based on low-depth whole-genome sequencing is provided, which includes: a bin division module, which is configured to independently divide low-depth whole-genome sequencing data into a non-fixed bin group and a fixed bin group; a calculation module, which is configured to use the copy numbers of the non-fixed bin group and the fixed bin group to calculate the average copy number and mutation region ratio of each group; a judgment module, which is configured to establish a scoring model based on the average copy number and mutation region ratio, judge the copy number variation type of each structural level of the chromosome, and then derive the corresponding molecular typing; wherein, the non-fixed bin group includes a first target region belonging to the exon and an anti-target region outside the exon; the fixed bin group includes a second target region with a length of 1M and a third target region with a length of 10K obtained by continuous division; each structural level of the chromosome includes: the entire chromosome, chromosome arm, chromosome region and chromosome band.

[0061] The device divides the low-depth whole-genome sequencing results into bins with different standards through the bin division module, obtaining a non-fixed bin group divided by whether it is an exon and a fixed bin group divided continuously by lengths of 1M and 10K. The calculation module uses the two different bin groupings to calculate the copy number, and obtains the average copy number and the proportion of mutation areas based on the copy number calculation. The judgment module uses this as the basis for judging the copy number variation type of each chromosome level (whole, arm, region, band), establishes a scoring model, determines the copy number variation type of the entire chromosome, and further obtains accurate tumor molecular typing results.

[0062] In order to further improve the efficiency of bin division and obtain more valid data for data analysis, it is necessary to clean and filter the low-depth whole-genome sequencing data. In a preferred embodiment, in the bin division module, before division, the inaccessible part of the second-generation sequencing in the low-depth whole-genome sequencing data is removed; preferably, the low-depth whole-genome sequencing data is a BAM file obtained by comparison with the human reference genome after quality control.

[0063] Any method capable of calculating copy number data is applicable to the present application. In a preferred embodiment, the calculation module includes: a CNVkit software calculation unit, which is configured to use the CNVkit software to calculate the copy numbers of the non-fixed bin group and the fixed bin group respectively.

[0064] By using the existing threshold standard for judging copy number variation, the copy number variation type of the region in each bin group where the copy number is detected is judged, which facilitates the subsequent judgment of the copy number variation type at all levels of the whole chromosome. In a preferred embodiment, before calculating the average copy number and mutation region ratio of each group, the calculation module also includes: an annotation unit, which is configured to annotate the position of the region in the chromosome whose copy number is calculated using the CNVkit software in the non-fixed bin group and the fixed bin group; a threshold judgment unit, which is configured to judge the copy number variation type of each region in the non-fixed bin group and the fixed bin group based on the copy number variation threshold; the copy number variation types include: loss, loss risk Loss_risk, amplification Gain and amplification risk Gain_risk.

[0065] For chromosomes at different levels, the copy numbers of different bin groups are calculated to obtain relevant data that can determine the types of chromosome copy number variation at different levels. In a preferred embodiment, the calculation module also includes: a first calculation unit, which is configured to calculate the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio of each region whose copy number is calculated by the CNVkit software in the non-fixed bin group at the chromosome whole and chromosome arm levels respectively; a second calculation unit, which is configured to calculate the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio of the region whose copy number is calculated by the CNVkit software in the second target region and the third target region in the fixed bin group respectively;

[0066] Average copy number = (the sum of the length of the region with calculated copy number by CNVkit software in the subject * the copy number of the corresponding region) / the sum of the lengths of all regions with calculated copy number in the subject;

[0067] The first mutation region ratio = the sum of the lengths of the Loss regions in the subject / the sum of the lengths of all regions with calculated copy numbers in the subject;

[0068] Second mutation region ratio = sum of the lengths of the Loss_risk regions in the object / sum of the lengths of all regions with calculated copy numbers in the object;

[0069] The third mutation region ratio = the sum of the Gain region lengths in the subject / the sum of the lengths of all regions with calculated copy numbers in the subject;

[0070] Fourth, the ratio of mutation regions = the sum of the lengths of the Gain_risk regions in the subject / the sum of the lengths of all regions with calculated copy numbers in the subject;

[0071] The target is: the entire chromosome, a chromosome arm, a second target region or a third target region.

[0072] Using copy number-related data at different chromosome levels (whole, arm) or different target regions (second and third target regions), the copy number variation type is judged, and a relevant scoring model is established to obtain the copy number variation type of the entire chromosome region, and further obtain the tumor molecular typing. In a preferred embodiment, the judgment module includes: a chromosome overall copy number variation type judgment unit, which is configured to use the average copy number at the chromosome overall level, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio to establish a first scoring model to judge the chromosome overall copy number variation type;

[0073] The chromosome arm copy number variation type judgment unit is configured to establish a second scoring model using the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio at the chromosome arm level to judge the chromosome arm copy number variation type;

[0074] a chromosome region copy number variation type determination unit configured to establish a third scoring model based on the copy number variation type of the entire chromosome and the copy number variation type of the chromosome arm, using the average copy number of the second target region, the proportion of the first mutation region, the proportion of the second mutation region, the proportion of the third mutation region, and the proportion of the fourth mutation region, to further determine the copy number variation type of the chromosome region;

[0075] The chromosome band copy number variation type judgment unit is configured to establish a fourth scoring model based on the copy number variation type of the entire chromosome and the copy number variation type of the chromosome arm, using the average copy number of the third target region, the proportion of the first mutation region, the proportion of the second mutation region, the proportion of the third mutation region and the proportion of the fourth mutation region, to further judge the copy number variation type of the chromosome band.

[0076] In a preferred embodiment, the copy number variation type judgment criteria in the chromosome whole copy number variation type judgment unit and the chromosome arm copy number variation type judgment unit include: when the average copy number of the judgment object is ≤ the baseline average copy number * 0.8 and the first mutation region ratio of the judgment object is ≥ 85%, then the copy number variation type of the judgment object is Loss; when the average copy number of the judgment object is ≤ the baseline average copy number * 0.9 and the sum of the first mutation region ratio of the judgment object and the second mutation region ratio of the judgment object is ≥ 85%, then the copy number variation type of the judgment object is Loss_risk; when the average copy number of the judgment object is ≥ the baseline average copy number * 1.2 and the third mutation region ratio of the judgment object is ≥ 85%, then the copy number variation type of the judgment object is Gain; when the average copy number of the judgment object is ≥ the baseline average copy number * 1.1 and the third mutation region ratio of the judgment object and the fourth mutation region ratio of the judgment object are ≥ 85%, then the copy number variation type of the judgment object is Gain_risk; the judgment object is the entire chromosome or the chromosome arm.

[0077] In a preferred embodiment, the judgment criteria of the copy number variation type of the chromosome region in the chromosome region copy number variation type judgment unit include: when the copy number variation type of the chromosome arm in the chromosome region is not amplification Gain, when the average copy number of the second target region is ≤ the baseline average copy number * 0.8 and the first mutation region ratio of the second target region is ≥85%, then the copy number variation type of the chromosome region is deletion Loss; when the copy number variation type of the chromosome arm in the chromosome region is not amplification Gain, when the average copy number of the second target region is ≤ the baseline average copy number * 0.9 and the sum of the first mutation region ratio of the second target region and the second mutation region ratio of the second target region is ≥85%, then the chromosome The copy number variation type of the chromosome arm in the chromosome region is deletion risk Loss_risk; when the copy number variation type of the chromosome arm in the chromosome region is not deletion Loss, when the average copy number of the second target region is ≥ the baseline average copy number * 1.2 and the ratio of the third mutation region in the second target region is ≥ 85%, then the copy number variation type of the chromosome region is amplification Gain; when the copy number variation type of the chromosome arm in the chromosome region is not deletion Loss, when the average copy number of the second target region is ≥ the baseline average copy number * 1.1 and the ratio of the third mutation region in the second target region to the fourth mutation region in the second target region is ≥ 85%, then the copy number variation type of the chromosome region is amplification risk Gain_risk.

[0078] In a preferred embodiment, the judgment criteria of the copy number variation type of the chromosome band in the chromosome band copy number variation type judgment unit include: when the copy number variation type of the chromosome arm of the chromosome band is not amplification Gain, when the average copy number of the third target region is ≤ the baseline average copy number * 0.8 and the first mutation region ratio of the third target region is ≥85%, then the copy number variation type of the chromosome band is deletion Loss; when the copy number variation type of the chromosome arm of the chromosome band is not amplification Gain, when the average copy number of the third target region is ≤ the baseline average copy number * 0.9 and the sum of the first mutation region ratio of the third target region and the second mutation region ratio of the third target region is ≥85%, then the chromosome The copy number variation type of the chromosome band is deletion risk Loss_risk; when the copy number variation type of the chromosome arm of the chromosome band is not deletion Loss, when the average copy number of the third target region is ≥ the baseline average copy number * 1.2 and the proportion of the third mutation region in the third target region is ≥ 85%, then the copy number variation type of the chromosome band is amplification Gain; when the copy number variation type of the chromosome arm of the chromosome band is not deletion Loss, when the average copy number of the third target region is ≥ the baseline average copy number * 1.1 and the proportion of the third mutation region in the third target region and the proportion of the fourth mutation region in the third target region are ≥ 85%, then the copy number variation type of the chromosome band is amplification risk Gain_risk.

[0079] In a third typical embodiment of the present application, a computer-readable storage medium is provided, which includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the above-mentioned method for detecting tumor molecular typing based on low-depth whole genome sequencing.

[0080] In a fourth typical embodiment of the present application, a processor is provided, which is used to run a program, wherein the program runs the above-mentioned method for detecting tumor molecular typing based on low-depth whole genome sequencing.

[0081] The present application is further described in detail below with reference to specific embodiments. These embodiments should not be construed as limiting the scope of protection claimed in this application.

[0082] Example 1

[0083] This application detection method: (such as Figure 1 shown)

[0084] 1) Use fastp software to clean and filter the sWGS data and align the cleaned sequencing data to the human reference genome to generate a BAM format file. The cleaning and filtering process utilizes the test sample's own control sWGS data or negative baseline data.

[0085] 2) The inaccessible portion of the whole genome by next-generation sequencing in the BAM file obtained in 1) is removed, and then the first target region and the anti-target region outside the exon are divided according to the whole exon to obtain the BED format file of the series.

[0086] 3) The inaccessible portion of the whole genome in the BAM file obtained in 1) was removed, and then the second target region and the third target region were continuously split according to fixed lengths of 1M and 10K to obtain the BED format files of this series.

[0087] 4) Use CNVkit software in pairing mode to calculate the copy number of each region in each series of BED files.

[0088] 5) Based on the calculation results of the software, the region where the copy number is calculated is annotated in the reference genome to obtain its specific position on the chromosome arm, region, and band. According to the threshold value (cutoff value) for determining copy number variation, the copy number qualitative results corresponding to the chromosome arm, region, and band (loss, loss risk, gain, and gain risk) are qualitatively determined.

[0089] 6) Calculate the average copy number and the ratio of mutation regions in each region of the non-fixed bin group at the chromosome and chromosome arm level, and the average copy number and the ratio of mutation regions in each region of the fixed bin group. The calculation method is:

[0090] Taking the entire chromosome / chromosome arm / second target region / third target region as the object, the average copy number and mutation region ratio of each region are calculated independently:

[0091] Average copy number = (the sum of the length of the region where the copy number is calculated by the software * the corresponding copy number) / the sum of the lengths of all regions where the copy number is calculated in the object;

[0092] The first mutation region ratio = the sum of the lengths of the Loss regions in the subject / the sum of the lengths of all regions with calculated copy numbers in the subject;

[0093] Second mutation region ratio = sum of the lengths of the Loss_risk regions in the object / sum of the lengths of all regions with calculated copy numbers in the object;

[0094] The third mutation region ratio = the sum of the lengths of the Gain regions in the subject / the sum of the lengths of all regions with calculated copy numbers in the subject;

[0095] Fourth mutation region ratio=the sum of the lengths of the Gain_risk regions in the subject / the sum of the lengths of all regions with calculated copy numbers in the subject.

[0096] 7) The above calculation results are imported into the scoring model, and the copy number variation type results of the corresponding regions at each level are determined by referring to the comparison of the average copy number, the proportion of mutation regions and the threshold matrix, and then the corresponding molecular typing is obtained.

[0097] 1. Determination of copy number variation of the entire chromosome and its arms:

[0098] Clinically, molecular typing in this dimension objectively requires that only a single copy number variation occurs in the vast majority of regions, the average copy number will vary significantly around 2, and the proportion of mutation regions is very high. The various parameters at the chromosome overall and chromosome arm levels are compared using non-fixed bin groups to establish a scoring model.

[0099] Loss of the entire chromosome or chromosome arm: average copy number ≤ baseline average copy number * 0.8 and the proportion of the first mutation region ≥ 85%;

[0100] The risk of deletion of the entire chromosome or chromosome arm is Loss_risk: the average copy number is ≤ the baseline average copy number * 0.9 and the proportion of the first mutation region + the proportion of the second mutation region is ≥ 85%;

[0101] Amplification of the entire chromosome or chromosome arm: average copy number ≥ baseline average copy number * 1.2 and the proportion of the third mutation region ≥ 85%;

[0102] There is a risk of amplification of the entire chromosome or chromosome arm. Gain_risk: average copy number ≥ baseline average copy number * 1.1 and the proportion of the third mutation region + the proportion of the fourth mutation region ≥ 85%.

[0103] 2. Targeting copy number variation in chromosomal regions

[0104] Combined with the copy number variation of the entire chromosome or chromosome arm, the copy number variation position in the chromosome region and band is further confirmed. The variation type of the region and band should be consistent with the chromosome as a whole and the arm.

[0105] Determination of the copy number variation type of the chromosome region:

[0106] Loss: When the chromosome arm is judged not to be a gain, the average copy number of the second target region is ≤ the average copy number of the baseline * 0.8 and the proportion of the first mutation region in the second target region is ≥ 85%;

[0107] Loss_risk: When the chromosome arm is judged not to be Gain, the average copy number of the second target region is ≤ the average copy number of the baseline * 0.9 and the proportion of the first mutation region in the second target region + the proportion of the second mutation region in the second target region is ≥ 85%;

[0108] Gain: When the chromosome arm is judged not to be Loss, the average copy number of the second target region ≥ the baseline average copy number * 1.2 and the proportion of the third mutation region in the second target region ≥ 85%;

[0109] Gain_risk: When the chromosome arm is judged not to be Loss, the average copy number of the second target region ≥ the baseline average copy number * 1.1 and the proportion of the third mutation region of the second target region + the proportion of the fourth mutation region of the second target region ≥ 85%.

[0110] Determination of the copy number variation type of chromosome band:

[0111] Loss: When the chromosome arm is judged not to be a gain, the average copy number of the third target region is ≤ the average copy number of the baseline * 0.8 and the proportion of the first mutation region in the third target region is ≥ 85%;

[0112] Loss_risk: When the chromosome arm is judged not to be Gain, the average copy number of the third target region is ≤ the average copy number of the baseline * 0.9 and the proportion of the first mutation region in the third target region + the proportion of the second mutation region in the third target region is ≥ 85%;

[0113] Gain: When the chromosome arm is judged not to be Loss, the average copy number of the third target region ≥ the baseline average copy number * 1.2 and the proportion of the third mutation region in the third target region ≥ 85%;

[0114] Gain_risk: When the chromosome arm is judged not to be Loss, the average copy number of the third target region ≥ the baseline average copy number * 1.1 and the proportion of the third mutation region in the third target region + the proportion of the fourth mutation region in the third target region ≥ 85%.

[0115] This example uses 12 single tissue samples with a priori standard positive and negative results from the hospital. The types of the samples are shown in Table 1. At the same time, multiple existing schemes are used for cross-validation to prove that the detection method of this application is feasible. The results are shown in Table 2 below.

[0116] Among them, the existing solutions include the Fish method, gene chip and Panel+Marker SNP. The specific methods are as follows:

[0117] The Fish method uses a specific single-stranded nucleic acid of known sequence as a probe, labeled with biotin or fluorescein. Under certain temperature and ion concentration conditions, DNA-DNA in situ hybridization occurs through base pairing, and fluorescence is used for visualization. Ultimately, the DNA is labeled at its original location on cell slides or tissue sections (paraffin or frozen). Because normal chromosomes / genes are diploid, the normal number of detection signals is two. When the number of signals is greater than or less than two (excluding the influence of the section), it can be inferred that a chromosome or gene abnormality (amplification or deletion) has occurred.

[0118] Gene chip: also known as DNA microarray, is a chip that integrates a large number of known sequence probes on the same substrate. Several labeled target nucleotide sequences hybridize with the probes at specific sites on the chip, and then the genetic information in biological cells or tissues is analyzed by detecting the hybridization signal.

[0119] Panel+Marker SNP: Use the second-generation sequencing technology (usually using a hybrid capture platform) to perform high-depth sequencing of the sample, and use probes to comprehensively capture the SNP sites in all exons of the sample-related genes and intron regions of some genes. After standardized library construction, use a second-generation sequencer for precise sequencing. For areas where copy number variations may occur on key chromosomes, more SNP sites to be tested are covered. The results are based on the actual copy number obtained by the corresponding SNP test of each chromosome, combined with the CNV threshold defined by the database, and bioinformatics calculations are performed to evaluate the possible chromosomal heterozygosity loss or homozygous loss of the target chromosome.

[0120] Table 1:

[0121] Sample_ID Sample type Patient 01 Paraffin sections of glioma Patient 02 Paraffin sections of glioma Patient 03 Fresh glioma tissue Patient 04 Fresh glioma tissue Patient 05 Paraffin sections of glioma Patient 06 Paraffin sections of glioma Patient 07 Paraffin sections of brain malignant glioma Patient 08 Secondary glioma paraffin sections Patient 09 Paraffin sections of glioma Patient 10 Paraffin sections of glioma Patient 11 Paraffin sections of glioma Patient 12 Paraffin sections of glioma

[0122] Table 2:

[0123]

[0124]

[0125] Among them, 1p_DEL is a deletion of the short arm of chromosome 1, 19q_DEL is a deletion of the long arm of chromosome 19, chr7_AMP is an overall amplification of chromosome 7, and chr10_DEL is an overall deletion of chromosome 10. The results of various detection schemes "-" mean that the corresponding test was not performed.

[0126] Figure 2The figure shows the variation type of all chromosomes in all 12 samples sequenced by sWGS, which was plotted by the R package of the R language platform based on the calculation results of the scoring model. Sample 1 was a priori negative for 1p19q deletion and positive for Chr7 amplification. Figure 3 The sWGS sequencing results shown clearly show that the results are consistent with the prior knowledge. In addition, amplification of Chr12 and Chr18 chromosomes was detected in the samples.

[0127] For the same sample 1, Figure 4 and Figure 5 As shown in the figure, the Panel+Marker SNP test results showed that 1p19q deletion was positive, which was inconsistent with the prior knowledge. It may be that there will be inaccurate detection during odd-number amplification. Chr7 amplification was positive, which was consistent with the prior knowledge.

[0128] For the same sample 1, Figure 6 As shown, the gene chip test results showed that 1p19q deletion was negative, which was consistent with the prior results. In addition, the amplification of Chr7, Chr12, and Chr18 chromosomes was consistent with the sWGS results.

[0129] For sample 11, Figure 7 As shown in the sWGS sequencing results, it can be clearly seen that 1p19q deletion is negative, Chr7 amplification is negative, and Chr10 deletion is negative, which is consistent with the prior results. Figure 8 As shown, the gene chip test results showed that 1p, 19q, Chr7, and Chr10 were all neutral results, corresponding to negative. It can be concluded that 1p19q deletion was negative, Chr7 amplification was negative, and Chr10 deletion was negative, which was consistent with the prior results. However, 19p and 19q both detected additional small area deletions and amplifications, which had a certain interference with the accuracy.

[0130] From the above description, it can be seen that the above-mentioned embodiments of the present invention achieve the following technical effects: sWGS sequencing data are divided into different block (bin) combinations according to different standards for subsequent copy number analysis, including being divided into non-fixed bin groups based on whether they are exons and being divided into fixed bin groups based on 1M and 10K lengths, and the copy number data of different bin groups are judged to determine the copy number variation types of chromosomes at different hierarchical structural positions, and then the corresponding molecular typing is obtained, thereby achieving more accurate molecular typing detection of the size of tumor chromosomes at each level, and further suitable for the detection of molecular typing of different cancer subspecies.

[0131] In addition, the detection method of the present application can be applied on all high-throughput sequencing platforms, supports single-sample detection (mixed negative samples can be used instead of controls), supports detection and verification of copy number variations in specific large regions, and can be extended to other projects that require chromosome-level copy number detection (molecular typing of other cancer types, large-segment deletions of genetic diseases, etc.).

[0132] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for detecting tumor molecular typing based on low-depth whole genome sequencing, characterized in that: The method comprises: S1, low-depth whole-genome sequencing data are independently divided into non-fixed bin groups and fixed bin groups; S2, using the copy numbers of the non-fixed bin group and the fixed bin group to calculate the average copy number and mutation region ratio of each group; S3, establishing a scoring model based on the average copy number and the ratio of the mutation region to determine the copy number variation type at each structural level of the chromosome, and then deriving the corresponding molecular typing; The non-fixed bin group includes the first target region belonging to the exon and the anti-target region outside the exon; the fixed bin group includes the second target region with a length of 1M and the third target region with a length of 10K obtained by continuous division; The various structural levels of the chromosome include: whole chromosome, chromosome arm, chromosome region and chromosome band; Before calculating the average copy number and mutation region ratio of each group, the method further includes: Annotating the positions of the regions in the chromosomes in the non-fixed bin group and the fixed bin group whose copy numbers are calculated using CNVkit software; Determining the copy number variation type of each of the regions in the non-fixed bin group and the fixed bin group according to a copy number variation threshold; The copy number variation types include: Loss, Loss risk_risk, Gain and Gain risk_risk; Calculating the average copy number and mutation region ratio of each group includes: Under the calculation of the chromosome as a whole and chromosome arm levels, the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio of each region of the copy number are calculated using the CNVkit software in the non-fixed bin group; Calculating the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio of the second target region and the third target region in the fixed bin group, respectively, for the regions where the copy number is calculated by the CNVkit software; The average copy number = (the sum of the length of the region in the subject for which the copy number is calculated by the CNVkit software * the corresponding copy number of the region) / the sum of the lengths of all regions in the subject for which the copy number is calculated; The first mutation region ratio=the sum of the lengths of the Loss regions in the subject / the sum of the lengths of all the regions in the subject whose copy numbers are calculated; The second mutation region ratio=the sum of the lengths of the Loss_risk regions in the object / the sum of the lengths of all the regions in the object whose copy numbers are calculated; The third mutation region ratio=the sum of the Gain region lengths in the subject / the sum of the region lengths of all the regions in the subject to obtain the calculated copy numbers; The fourth mutation region ratio=the sum of the lengths of the Gain_risk regions in the subject / the sum of the lengths of all the regions in the subject to obtain the copy numbers by calculation; Wherein, the object is: the entire chromosome, the chromosome arm, the second target region or the third target region; The S3 includes: Establishing a first scoring model using the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio at the overall chromosome level to determine the copy number variation type of the overall chromosome; Establishing a second scoring model using the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio at the chromosome arm level to determine the copy number variation type of the chromosome arm; Based on the copy number variation type of the entire chromosome and the copy number variation type of the chromosome arm, a third scoring model is established using the average copy number of the second target region, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio, to further determine the copy number variation type of the chromosome region; Based on the copy number variation type of the entire chromosome and the copy number variation type of the chromosome arm, a fourth scoring model is established using the average copy number of the third target region, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio to further determine the copy number variation type of the chromosome band.

2. The method according to claim 1, characterized in that Before performing the division, the inaccessible portion of the second-generation sequencing in the low-depth whole-genome sequencing data is removed.

3. The method according to claim 2, characterized in that The low-depth whole-genome sequencing data is a BAM file obtained by comparing it with the human reference genome after quality control.

4. The method according to claim 1, wherein The copy numbers of the non-fixed bin group and the fixed bin group were calculated using CNVkit software.

5. The method according to claim 1, wherein The copy number variation type of the entire chromosome and the chromosome arm is determined according to the following criteria: When the average copy number of the judgment object is ≤ the baseline average copy number * 0.8 and the first mutation region ratio of the judgment object is ≥ 85%, the copy number variation type of the judgment object is Loss; When the average copy number of the judgment object is ≤ baseline average copy number * 0.9 and the sum of the first mutation region ratio of the judgment object and the second mutation region ratio of the judgment object is ≥ 85%, the copy number variation type of the judgment object is loss risk; When the average copy number of the judgment object is ≥ the baseline average copy number * 1.2 and the proportion of the third mutation region of the judgment object is ≥ 85%, the copy number variation type of the judgment object is amplification Gain; When the average copy number of the judgment object is greater than or equal to the baseline average copy number*1.1 and the ratio of the third mutation region of the judgment object to the ratio of the fourth mutation region of the judgment object is greater than or equal to 85%, the copy number variation type of the judgment object is amplification risk Gain_risk; The judgment object is the entire chromosome or the chromosome arm.

6. The method according to claim 1, characterized in that The copy number variation type of the chromosome region is determined according to the following criteria: When the copy number variation type of the chromosome arm of the chromosome region is not amplification (Gain), when the average copy number of the second target region is ≤ baseline average copy number*0.8 and the ratio of the first mutation region in the second target region is ≥85%, then the copy number variation type of the chromosome region is deletion (Loss); When the copy number variation type of the chromosome arm of the chromosome region is not amplification Gain, when the average copy number of the second target region is ≤ baseline average copy number * 0.9 and the sum of the first mutation region ratio of the second target region and the second mutation region ratio of the second target region is ≥85%, then the copy number variation type of the chromosome region is deletion risk Loss_risk; When the copy number variation type of the chromosome arm in the chromosome region is not Loss, when the average copy number of the second target region is ≥ baseline average copy number * 1.2 and the third mutation region ratio of the second target region is ≥ 85%, the copy number variation type of the chromosome region is Gain; When the copy number variation type of the chromosome arm in the chromosome region is not loss, when the average copy number of the second target region is ≥ baseline average copy number * 1.1 and the ratio of the third mutation region of the second target region to the fourth mutation region of the second target region is ≥ 85%, then the copy number variation type of the chromosome region is amplification risk Gain_risk.

7. The method according to claim 1, characterized in that The copy number variation type of the chromosome band is determined according to the following criteria: When the copy number variation type of the chromosome arm of the chromosome band is not amplification (Gain), when the average copy number of the third target region is ≤ baseline average copy number*0.8 and the first mutation region ratio of the third target region is ≥85%, then the copy number variation type of the chromosome band is deletion (Loss); When the copy number variation type of the chromosome arm of the chromosome band is not amplification Gain, when the average copy number of the third target region is ≤ baseline average copy number * 0.9 and the sum of the first mutation region ratio of the third target region and the second mutation region ratio of the third target region is ≥85%, then the copy number variation type of the chromosome band is deletion risk Loss_risk; When the copy number variation type of the chromosome arm in the chromosome band is not Loss, when the average copy number of the third target region is ≥ baseline average copy number * 1.2 and the third mutation region ratio of the third target region is ≥ 85%, the copy number variation type of the chromosome band is Gain; When the copy number variation type of the chromosome arm of the chromosome band is not loss, when the average copy number of the third target region is ≥ baseline average copy number * 1.1 and the ratio of the third mutation region of the third target region to the ratio of the fourth mutation region of the third target region is ≥85%, then the copy number variation type of the chromosome band is amplification risk Gain_risk.

8. A device for detecting tumor molecular typing based on low-depth whole genome sequencing, characterized in that: The device comprises: The bin division module is configured to independently divide low-depth whole-genome sequencing data into non-fixed bin groups and fixed bin groups; A calculation module is configured to calculate the average copy number and mutation region ratio of each group using the copy numbers of the non-fixed bin group and the fixed bin group; A judgment module is configured to establish a scoring model based on the average copy number and the ratio of the mutation region, determine the copy number variation type of each structural level of the chromosome, and then derive the corresponding molecular typing; The non-fixed bin group includes the first target region belonging to the exon and the anti-target region outside the exon; the fixed bin group includes the second target region with a length of 1M and the third target region with a length of 10K obtained by continuous division; The various structural levels of the chromosome include: whole chromosome, chromosome arm, chromosome region and chromosome band; Before calculating the average copy number and mutation region ratio of each group, the calculation module further includes: an annotation unit configured to annotate the positions of the regions in the chromosomes in the non-fixed bin group and the fixed bin group whose copy numbers are calculated using CNVkit software; a threshold determination unit configured to determine the copy number variation type of each of the regions in the non-fixed bin group and the fixed bin group based on a copy number variation threshold; The copy number variation types include: Loss, Loss risk_risk, Gain and Gain risk_risk; The calculation module also includes: The first calculation unit is configured to calculate the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio of each region of the copy number calculated by the CNVkit software in the non-fixed bin group at the chromosome whole level and the chromosome arm level respectively; a second calculation unit configured to respectively calculate the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio of the second target region and the third target region in the fixed bin group, in the regions where the copy number is calculated by the CNVkit software; The average copy number = (the sum of the length of the region in the subject for which the copy number is calculated by the CNVkit software * the corresponding copy number of the region) / the sum of the lengths of all regions in the subject for which the copy number is calculated; The first mutation region ratio=the sum of the lengths of the Loss regions in the subject / the sum of the lengths of all the regions in the subject whose copy numbers are calculated; The second mutation region ratio=the sum of the lengths of the Loss_risk regions in the object / the sum of the lengths of all the regions in the object whose copy numbers are calculated; The third mutation region ratio=the sum of the Gain region lengths in the subject / the sum of the region lengths of all the regions in the subject to obtain the calculated copy numbers; The fourth mutation region ratio=the sum of the lengths of the Gain_risk regions in the subject / the sum of the lengths of all the regions in the subject to obtain the copy numbers by calculation; Wherein, the object is: the entire chromosome, the chromosome arm, the second target region or the third target region; The judgment module includes: a chromosome overall copy number variation type determination unit, configured to establish a first scoring model using the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio at the chromosome overall level to determine the chromosome overall copy number variation type; a chromosome arm copy number variation type determination unit, configured to establish a second scoring model using the average copy number, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio at the chromosome arm level to determine the copy number variation type of the chromosome arm; a chromosome region copy number variation type determination unit, configured to establish a third scoring model based on the copy number variation type of the entire chromosome and the copy number variation type of the chromosome arm, using the average copy number of the second target region, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio, and the fourth mutation region ratio, to thereby determine the copy number variation type of the chromosome region; The chromosome band copy number variation type judgment unit is configured to establish a fourth scoring model based on the copy number variation type of the entire chromosome and the copy number variation type of the chromosome arm, using the average copy number of the third target region, the first mutation region ratio, the second mutation region ratio, the third mutation region ratio and the fourth mutation region ratio, to further judge the copy number variation type of the chromosome band.

9. The device according to claim 8, characterized in that In the bin division module, before performing the division, the inaccessible portion of the second-generation sequencing in the low-depth whole-genome sequencing data is removed.

10. The device according to claim 9, characterized in that The low-depth whole-genome sequencing data is a BAM file obtained by comparing it with the human reference genome after quality control.

11. The device according to claim 8, characterized in that The calculation module includes: The CNVkit software calculation unit is configured to use the CNVkit software to respectively calculate the copy numbers of the non-fixed bin group and the fixed bin group.

12. The device according to claim 8, characterized in that The judgment criteria for the copy number variation type in the chromosome overall copy number variation type judgment unit and the chromosome arm copy number variation type judgment unit include: When the average copy number of the judgment object is ≤ the baseline average copy number * 0.8 and the first mutation region ratio of the judgment object is ≥ 85%, the copy number variation type of the judgment object is Loss; When the average copy number of the judgment object is ≤ baseline average copy number * 0.9 and the sum of the first mutation region ratio of the judgment object and the second mutation region ratio of the judgment object is ≥ 85%, the copy number variation type of the judgment object is loss risk; When the average copy number of the judgment object is ≥ the baseline average copy number * 1.2 and the proportion of the third mutation region of the judgment object is ≥ 85%, the copy number variation type of the judgment object is amplification Gain; When the average copy number of the judgment object is greater than or equal to the baseline average copy number*1.1 and the ratio of the third mutation region of the judgment object to the ratio of the fourth mutation region of the judgment object is greater than or equal to 85%, the copy number variation type of the judgment object is amplification risk Gain_risk; The judgment object is the entire chromosome or the chromosome arm.

13. The device according to claim 8, characterized in that The judgment criteria for the copy number variation type of the chromosome region in the chromosome region copy number variation type judgment unit include: When the copy number variation type of the chromosome arm of the chromosome region is not amplification (Gain), when the average copy number of the second target region is ≤ baseline average copy number*0.8 and the ratio of the first mutation region in the second target region is ≥85%, then the copy number variation type of the chromosome region is deletion (Loss); When the copy number variation type of the chromosome arm of the chromosome region is not amplification Gain, when the average copy number of the second target region is ≤ baseline average copy number * 0.9 and the sum of the first mutation region ratio of the second target region and the second mutation region ratio of the second target region is ≥85%, then the copy number variation type of the chromosome region is deletion risk Loss_risk; When the copy number variation type of the chromosome arm in the chromosome region is not Loss, when the average copy number of the second target region is ≥ baseline average copy number * 1.2 and the third mutation region ratio of the second target region is ≥ 85%, the copy number variation type of the chromosome region is Gain; When the copy number variation type of the chromosome arm in the chromosome region is not loss, when the average copy number of the second target region is ≥ baseline average copy number * 1.1 and the ratio of the third mutation region of the second target region to the fourth mutation region of the second target region is ≥ 85%, then the copy number variation type of the chromosome region is amplification risk Gain_risk.

14. The device according to claim 8, characterized in that The judgment criteria for the chromosome band copy number variation type in the chromosome band copy number variation type judgment unit include: When the copy number variation type of the chromosome arm of the chromosome band is not amplification (Gain), when the average copy number of the third target region is ≤ baseline average copy number*0.8 and the first mutation region ratio of the third target region is ≥85%, then the copy number variation type of the chromosome band is deletion (Loss); When the copy number variation type of the chromosome arm of the chromosome band is not amplification Gain, when the average copy number of the third target region is ≤ baseline average copy number * 0.9 and the sum of the first mutation region ratio of the third target region and the second mutation region ratio of the third target region is ≥85%, then the copy number variation type of the chromosome band is deletion risk Loss_risk; When the copy number variation type of the chromosome arm in the chromosome band is not Loss, when the average copy number of the third target region is ≥ baseline average copy number * 1.2 and the third mutation region ratio of the third target region is ≥ 85%, the copy number variation type of the chromosome band is Gain; When the copy number variation type of the chromosome arm of the chromosome band is not loss, when the average copy number of the third target region is ≥ baseline average copy number * 1.1 and the ratio of the third mutation region of the third target region to the ratio of the fourth mutation region of the third target region is ≥85%, then the copy number variation type of the chromosome band is amplification risk Gain_risk.

15. A computer-readable storage medium, characterized in that The storage medium includes a stored program, wherein, when the program is running, the device where the storage medium is located is controlled to execute the method for detecting tumor molecular typing based on low-depth whole genome sequencing according to any one of claims 1 to 7.

16. A processor, characterized in that: The processor is used to run a program, wherein the program runs the method for detecting tumor molecular typing based on low-depth whole genome sequencing according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for distinguishing lymphoma molecular subtypes and storage medium

    CN114093421A