Information processing method and device for gene copy number variation detection, equipment and medium

CN121237201BActive Publication Date: 2026-08-11SUZHOU SAIFU MEDICAL LAB CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]目前基因CNV检测已发展出多种技术方法,然而,现有技术存在明显局限性:一方面,单一检测维度难以全面覆盖不同类型的CNV,易因检测灵敏度不足或区间判定逻辑不完整导致漏检、误判;另一方面,不同检测方法的结果缺乏有效的整合与验证机制,当单一方法的判定结果受干扰时,难以区分真实CNV与假阳性信号,最终导致检测结果的可信度与准确性无法满足临床对精准诊断的需求

Benefits of technology

[0051]本申请有益效果为:本申请应用于计算机装置,获取待测生物样本中目标基因的靶向区域遗传特征数据;其中,所述靶向区域遗传特征数据包括区域定位信息、测序比对数据、变异位点信息;基于所述靶向区域遗传特征数据对所述目标基因分别进行基于单倍型的拷贝数变异判定、基于VAF的拷贝数变异判定以及基于测序深度的拷贝数变异判定,以得到单倍型判定结果、VAF判定结果以及测序深度判定结果;根据所述单倍型判定结果和所述VAF判定结果确定所述目标基因中变异扩增子的目标占比,并根据所述目标占比与预设占比阈值之间的大小关系确定所述目标基因的初步变异检测结果;若所述初步变异检测结果与所述测序深度判定结果相匹配,则将所述初步变异检测结果确定为所述目标基因的基因拷贝数的目标变异检测结果。由此可见,通过获取包含区域定位信息、测序比对数据、变异位点信息的靶向区域遗传特征数据,为基因拷贝数变异判定奠定精准基础,具体的,区域定位信息精准划定目标基因扩增子范围,排除无关区域干扰,测序比对数据提供覆盖度量化依据,变异位点信息辅助过滤干扰区域与单倍型分型,解决传统检测数据维度不足问题;同时,采用基于单倍型、VAF、测序深度的三种基因拷贝数变异判定方法并行分析,先根据单倍型与VAF判定结果确定初步变异检测结果,再与测序深度判定结果验证一致性,如此一来,降低假阳性与假阴性,若初步结果与测序深度判定结果矛盾可排除误差,若一致则提升可信度,减少单一方法误差导致的误诊漏诊,最终系统性解决传统基因拷贝数变异检测中区间判定不准、低丰度漏检、结果可信度低等痛点,实现目标基因的基因拷贝数变异的高准确性、高灵敏度、全区域覆盖检测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237201B_ABST
    Figure CN121237201B_ABST
Patent Text Reader

Abstract

This application discloses information processing methods, devices, equipment, and media for gene copy number variation detection, relating to the fields of molecular biology and gene detection technology. The methods include: acquiring genetic characteristic data of the target region of a target gene in a biological sample; performing copy number variation determination on the target gene based on haplotype, VAF, and sequencing depth according to the genetic characteristic data of the target region, to obtain haplotype determination results, VAF determination results, and sequencing depth determination results; determining the target proportion of variant amplicon in the target gene based on the haplotype determination results and VAF determination results; determining the preliminary variation detection result of the target gene based on the target proportion and a preset proportion threshold; and determining the preliminary variation detection result as the target gene copy number variation detection result if the preliminary variation detection result matches the sequencing depth determination result. This application can directly achieve large-fragment single-gene copy number variation detection across three generations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of molecular biology and gene detection technology, and in particular to information processing methods, devices, equipment and media for detecting gene copy number variations. Background Technology

[0002] Copy number variation (CNV) refers to abnormal changes in the copy number of gene segments within a specific range, encompassing deletions, duplications, and other types. It is closely related to the occurrence and development of many human diseases. Therefore, accurate CNV detection of target genes is a crucial step in disease diagnosis, etiological analysis, treatment planning, and prognostic assessment, and is of significant demand in both clinical practice and basic research.

[0003] Currently, various technologies and methods have been developed for gene CNV detection. However, existing technologies have obvious limitations: on the one hand, a single detection dimension is difficult to fully cover different types of CNVs, and it is easy to miss or misjudge due to insufficient detection sensitivity or incomplete interval judgment logic; on the other hand, the results of different detection methods lack an effective integration and verification mechanism. When the judgment result of a single method is interfered with, it is difficult to distinguish between real CNVs and false positive signals, ultimately resulting in the reliability and accuracy of the detection results failing to meet the clinical needs for precise diagnosis.

[0004] In summary, improving the accuracy of gene copy number variation detection is a problem that needs to be solved in this field. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide an information processing method, apparatus, device, and medium for detecting gene copy number variations, thereby improving the accuracy of gene copy number variation detection. The specific solution is as follows:

[0006] In a first aspect, this application discloses an information processing method for detecting gene copy number variations, applied to a computer device, comprising:

[0007] Obtain genetic feature data of the target region of the target gene in the biological sample to be tested; wherein, the genetic feature data of the target region includes regional location information, sequencing alignment data, and mutation site information;

[0008] Based on the genetic feature data of the target region, the target gene is subjected to copy number variation determination based on haplotype, copy number variation determination based on VAF, and copy number variation determination based on sequencing depth, respectively, to obtain haplotype determination results, VAF determination results, and sequencing depth determination results.

[0009] The target proportion of variant amplicon in the target gene is determined based on the haplotype determination result and the VAF determination result, and the preliminary mutation detection result of the target gene is determined based on the relationship between the target proportion and the preset proportion threshold.

[0010] If the preliminary mutation detection result matches the sequencing depth determination result, then the preliminary mutation detection result is determined as the target mutation detection result for the gene copy number of the target gene.

[0011] Optionally, obtain the regional location information and sequencing alignment data of the target gene, including:

[0012] Generate a BED file containing the regional location information of the target gene based on the genomic coordinates of the target gene;

[0013] The sample genome is amplified according to the coordinates of the amplicon region of the target gene as described in the BED file to obtain the target amplicon fragment; wherein, the sample genome is the complete genetic material of the biological sample to be tested and includes the target gene;

[0014] The target amplicon fragment is sequenced to obtain raw sequencing data;

[0015] The raw sequencing data is compared with a reference genome in a standard genome sequence to obtain a BAM file containing the sequencing alignment data.

[0016] Optionally, obtain information on the mutation sites of the target gene, including:

[0017] The BAM file is preprocessed to obtain a preprocessed file;

[0018] The differences between the preprocessed file and the reference genome are analyzed using a target variant detection tool to obtain initial variant results;

[0019] The initial mutation results are filtered to obtain a VCF file containing mutation site information of the target gene.

[0020] Optionally, based on the genetic feature data of the target region, copy number variation determination of the target gene is performed based on haplotype to obtain haplotype determination results, including:

[0021] Based on the genetic feature data of the target region, haplotype classification and haplotype block labeling are performed on the genome sequence fragments to calculate the haplotype ratio of each haplotype within each haplotype block; wherein, the genome sequence fragment is located within the amplicon interval of each non-overlapping region of the target gene;

[0022] If the current amplicon interval meets the preset haplotype deletion condition, a haplotype determination result is generated indicating that the current amplicon interval has a mutation and the mutation type is a deletion type; wherein, the preset haplotype deletion condition is a first preset deletion condition or a second preset deletion condition, the first preset deletion condition is that the haplotype typing result indicates that there is only 1 haplotype in a single amplicon interval, and the second preset deletion condition is that the haplotype typing result indicates that there are 2 haplotypes in a single amplicon interval and the haplotype ratio of each haplotype is less than a first preset threshold or greater than a second preset threshold;

[0023] If the current amplicon interval meets the preset haplotype normal condition, a haplotype determination result is generated indicating that there is no variation in the current amplicon interval; wherein, the preset haplotype normal condition is that the haplotype typing result indicates that there are 2 haplotypes in a single amplicon interval and the haplotype ratio of each haplotype is within a first preset range.

[0024] If the current amplicon interval does not meet the preset haplotype deletion condition and the preset haplotype normal condition, a haplotype determination result is generated indicating that the current amplicon interval has a variation and the variation type is a repeat type.

[0025] Optionally, the target gene is subjected to VAF-based copy number variation determination based on the mutation site information to obtain VAF determination results, including:

[0026] The VAF value of each amplicon region in the target gene is determined based on the mutation site information;

[0027] If the current amplicon interval meets the preset VAF missing condition, a VAF determination result is generated to indicate that the current amplicon interval has a mutation and the mutation type is a missing type; wherein, the preset VAF missing condition is that the VAF value is less than a third preset threshold or greater than a fourth preset threshold;

[0028] If the current amplicon interval meets the preset VAF normal condition, a VAF determination result indicating that the current amplicon interval does not have a variation is generated; wherein, the preset VAF missing condition is that the VAF value is within a second preset range;

[0029] If the current amplicon interval does not meet the preset VAF missing condition and the preset VAF normal condition, a VAF determination result is generated indicating that the current amplicon interval has a variation and the variation type is repeat.

[0030] Optionally, based on the genetic feature data of the target region, copy number variation determination of the target gene is performed based on sequencing depth to obtain sequencing depth determination results, including:

[0031] The target region is determined based on the genetic feature data of the target region; wherein, the target region includes multiple amplicon regions;

[0032] The sequencing depth of the amplicon intervals is statistically analyzed using a preset window size as a sliding window to obtain the average sequencing depth of each amplicon interval, and the sequencing depth ratio of each sliding window is determined based on the average sequencing depth.

[0033] If the current amplicon region meets the preset sequencing depth missing condition, a sequencing depth determination result is generated to indicate that the current amplicon region has a mutation and the mutation type is a deletion type; wherein, the preset sequencing depth missing condition is that the sequencing depth ratio is greater than a fifth preset threshold.

[0034] If the current amplicon region meets the preset normal sequencing depth condition, a sequencing depth determination result is generated indicating that the current amplicon region does not have a variation; wherein, the preset normal sequencing depth condition is that the sequencing depth ratio is not greater than a sixth preset threshold.

[0035] If the current amplicon interval does not meet the preset sequencing depth missing condition and the preset sequencing depth normal condition, a sequencing depth determination result is generated indicating that the current amplicon interval has a variation and the variation type is repetitive.

[0036] Optionally, the step of determining the target proportion of variant amplicon in the target gene based on the haplotype determination result and the VAF determination result, and determining the preliminary variant detection result of the target gene based on the relationship between the target proportion and a preset proportion threshold, includes:

[0037] The amplicon in the haplotype determination result and / or the VAF determination result that indicates a mutation in the target gene and that the mutation type is deletion is identified as the first variant amplicon;

[0038] The amplicon in the haplotype determination result and / or the VAF determination result that indicates a variation in the target gene and that the variation type is repeat is identified as the second variant amplicon;

[0039] The first target proportion of the first variant amplicon and the second target proportion of the second variant amplicon in the target gene are determined based on the number of the first variant amplicon and the number of the second variant amplicon.

[0040] If the proportion of the first target is greater than the preset proportion threshold, a preliminary mutation detection result is generated to indicate that the gene copy number of the target gene has a variation and the variation type is a deletion type.

[0041] If the proportion of the second target is greater than the preset proportion threshold, a preliminary mutation detection result is generated, indicating that the gene copy number of the target gene has a variation and the variation type is a repeat type.

[0042] Secondly, this application discloses an information processing device for detecting gene copy number variations, applied to a computer device, comprising:

[0043] The data acquisition module is used to acquire genetic feature data of the target region of the target gene in the biological sample to be tested; wherein, the genetic feature data of the target region includes regional location information, sequencing alignment data, and mutation site information;

[0044] The determination module is used to determine the copy number variation of the target gene based on haplotype, copy number variation based on VAF, and copy number variation based on sequencing depth based on the genetic feature data of the target region, so as to obtain the haplotype determination result, VAF determination result, and sequencing depth determination result.

[0045] The preliminary result determination module is used to determine the target proportion of variant amplicon in the target gene based on the haplotype determination result and the VAF determination result, and to determine the preliminary variant detection result of the target gene based on the relationship between the target proportion and the preset proportion threshold.

[0046] The mutation result determination module is used to determine the preliminary mutation detection result as the target mutation detection result of the gene copy number of the target gene if the preliminary mutation detection result matches the sequencing depth determination result.

[0047] Thirdly, this application discloses an electronic device, including:

[0048] Memory, used to store computer programs;

[0049] A processor is configured to execute the computer program to implement the steps of the aforementioned disclosed information processing method for detecting gene copy number variations.

[0050] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the aforementioned disclosed information processing method for detecting gene copy number variations.

[0051] The beneficial effects of this application are as follows: This application is applied to a computer device to acquire genetic feature data of a target region of a target gene in a biological sample to be tested; wherein, the genetic feature data of the target region includes regional location information, sequencing alignment data, and mutation site information; based on the genetic feature data of the target region, the target gene is subjected to copy number variation determination based on haplotype, copy number variation determination based on VAF, and copy number variation determination based on sequencing depth to obtain haplotype determination results, VAF determination results, and sequencing depth determination results; the target proportion of variant amplicon in the target gene is determined according to the haplotype determination results and the VAF determination results, and the preliminary variation detection result of the target gene is determined according to the relationship between the target proportion and a preset proportion threshold; if the preliminary variation detection result matches the sequencing depth determination result, the preliminary variation detection result is determined as the target variation detection result of the gene copy number of the target gene. Therefore, by acquiring genetic feature data of the target region containing regional location information, sequencing alignment data, and variant site information, a precise foundation is laid for gene copy number variation determination. Specifically, regional location information accurately delineates the target gene amplicon range and eliminates interference from irrelevant regions; sequencing alignment data provides a quantitative basis for coverage; and variant site information assists in filtering interfering regions and haplotype classification, solving the problem of insufficient data dimensions in traditional detection methods. Simultaneously, three gene copy number variation determination methods based on haplotype, VAF, and sequencing depth are analyzed in parallel. First, preliminary variation detection results are determined based on haplotype and VAF determination results, and then consistency is verified with sequencing depth determination results. This reduces false positives and false negatives. If the preliminary results contradict the sequencing depth determination results, errors can be eliminated; if they are consistent, reliability is improved, reducing misdiagnosis and missed diagnosis caused by errors in a single method. Ultimately, this systematically solves the pain points of inaccurate interval determination, low abundance missed detection, and low result reliability in traditional gene copy number variation detection, achieving high accuracy, high sensitivity, and full-region coverage detection of gene copy number variations in the target gene. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0053] Figure 1 This is a flowchart of an information processing method for detecting gene copy number variations disclosed in this application;

[0054] Figure 2 This is a schematic diagram illustrating a specific full-length gene copy number variation detection method disclosed in this application;

[0055] Figure 3 This is a schematic diagram illustrating a specific haplotype copy number variation determination disclosed in this application;

[0056] Figure 4 This is a schematic diagram illustrating a specific copy number variation determination based on sequencing depth disclosed in this application;

[0057] Figure 5 This is a schematic diagram illustrating a specific gene copy number variation detection method disclosed in this application;

[0058] Figure 6 This is a schematic diagram of the first specific test result disclosed in this application;

[0059] Figure 7 This is a schematic diagram of the second specific test result disclosed in this application;

[0060] Figure 8 This is a schematic diagram of the third specific test result disclosed in this application;

[0061] Figure 9 This is a schematic diagram of an information processing device for detecting gene copy number variations disclosed in this application.

[0062] Figure 10 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0063] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0064] Copy number variation (CNV) refers to abnormal changes in the copy number of gene segments within a specific range of genome size, encompassing deletions, duplications, and other types. It is closely related to the occurrence and development of various human diseases. Therefore, accurate CNV detection of target genes is a crucial step in disease diagnosis, etiological analysis, treatment planning, and prognostic assessment, and is of significant demand in both clinical practice and basic research.

[0065] Currently, various technologies and methods have been developed for gene CNV detection. However, existing technologies have obvious limitations: on the one hand, a single detection dimension is difficult to fully cover different types of CNVs, and it is easy to miss or misjudge due to insufficient detection sensitivity or incomplete interval judgment logic; on the other hand, the results of different detection methods lack an effective integration and verification mechanism. When the judgment result of a single method is interfered with, it is difficult to distinguish between real CNVs and false positive signals, ultimately resulting in the reliability and accuracy of the detection results failing to meet the clinical needs for precise diagnosis.

[0066] Therefore, this application provides a gene copy number variation detection scheme, which improves the accuracy of gene copy number variation detection.

[0067] See Figure 1 As shown in the embodiments of this application, an information processing method for detecting gene copy number variations is disclosed, applied to a computer device, including:

[0068] Step S11: Obtain the genetic feature data of the target region of the target gene in the biological sample to be tested; wherein, the genetic feature data of the target region includes regional location information, sequencing alignment data, and mutation site information.

[0069] Third-generation sequencing (TGS) technologies, such as PacBio SMRT, use single-molecular real-time (SMRT) sequencing to sequence DNA molecules by capturing fluorescence signals from DNA polymerase during DNA replication. Oxford Nanopore nanopore single-molecular sequencing technology sequences DNA molecules by detecting the electrical signals generated as they pass through nanopores. These sequencing processes do not require PCR (Polymerase Chain Reaction) amplification, achieving continuous coverage of reads exceeding 10 kb to several Mb, providing new tools for solving the challenges of genetic disease diagnosis. Compared to traditional short-read next-generation sequencing (NGS), it can directly span entire regions of structural variation or both ends of breakpoints, providing a "single-molecule view" to directly determine the location, size, orientation, and sequence content of variations. It can also span entire repeat units or homologous regions, directly covering complete sequences containing multiple repeat units or unique regions that distinguish true from false genes, enabling precise differentiation of repeat counts and homologous genes. Long reads can cover very long continuous sequences (typically tens of thousands of bp) on a single sequencing reaction, directly determining the linkage (phase) of variations on the same chromosome without additional experiments. This is of great significance for clarifying the pathogenicity of compound heterozygous pathogenic variants, significantly improving diagnostic rates and reducing uncertainty.

[0070] Based on the above advantages, third-generation sequencing technology has been widely used in the fields of genetic disease and cancer screening and research. However, due to the high cost of whole-genome third-generation sequencing, targeted third-generation sequencing (TGS) technology has been promoted more quickly in the actual clinical translation and application of detection technology for complex genes (high homology, tandem duplication, complex structural variations, etc.). This includes Duchenne muscular dystrophy (DMD), spinal muscular atrophy (SMN1), intranuclear inclusion disease and essential tremor (NOTCH2NLC), amyotrophic lateral sclerosis (C9orf72), distal oculopharyngeal myopathy, myotonic dystrophy type 1 (DMPK), late-onset spinocerebellar ataxia 27B, epilepsy (CLN6, SAMD12), Parkinson's disease; thalassemia (HBA1, HBA2); congenital adrenal hyperplasia (CYP21A2); polycystic kidney disease (PKD1), etc. Conventional NGS methods (such as whole-exome sequencing (WES), gene panels, etc.) have significant limitations in detecting the aforementioned gene variations, tandem repeats, and copy number. Targeted third-generation sequencing (TGS) is mainly divided into PCR capture and liquid-phase capture. Among them, PCR amplicon-TGS is a method that directly obtains the full-length amplicon sequence by amplifying the target region using targeted PCR and combining it with third-generation long reads. Compared with capture third-generation technology, it only requires PCR amplification using primers with barcodes to enrich the target gene, eliminating the need for a complicated library construction and hybridization process. It is simple to operate and has low library construction costs.

[0071] Copy number variations (CNVs) are crucial in the diagnosis of genetic diseases. Intragenic insertion and deletion CNVs can be precisely identified by analyzing structural variations in long amplicon third-generation sequencing data. However, this method cannot identify full-length duplications and deletions in single genes. Existing capture NGS technologies primarily use batch controls or samples with similar capture efficiencies to calibrate the coverage of reads in the test samples to improve the identification capability of CNVs within 500kb. The amplification efficiency of long-range LR-PCR (>5kb) is easily affected by primer sequence polymorphism, primer concentration, amplification system, and length, especially when the system contains only a single target gene, making the identification of full-length duplications and deletions of the target gene challenging. Therefore, a scheme for identifying full-length copy number variations (CNVs) of target genes was developed by comprehensively considering the haplotype ratio, variant allele frequency (VAF), and sequencing depth in each amplicon to improve the comprehensiveness of long-read amplicon variation analysis.

[0072] In this embodiment, obtaining the regional location information and sequencing alignment data of the target gene includes: generating a BED file containing the regional location information of the target gene based on the genomic coordinates of the target gene; amplifying the sample genome according to the amplicon region coordinates of the target gene in the BED file to obtain a target amplicon fragment; wherein the sample genome is all the genetic material of the biological sample to be tested and includes the target gene; sequencing the target amplicon fragment to obtain raw sequencing data; and aligning the raw sequencing data with a reference genome in a standard genome sequence to obtain a BAM file containing sequencing alignment data.

[0073] For example Figure 2 The diagram illustrates a specific method for detecting full-length gene copy number variations. First, based on the known full-length coordinates of the target gene in a standard genome sequence (e.g., chromosome number, start and end positions), non-overlapping amplicons covering the entire gene length are designed. The precise genomic coordinates of each amplicons (including chromosome, start site, end site, and unique identifier) ​​are organized in a standardized format to generate a BED (Browser Extensible Data) file containing the target gene's regional location information. Then, specific primers are designed based on the amplicons' coordinates in this BED file, and PCR amplification is performed on the genomic DNA of the sample to be tested (i.e., the sample genome containing the target gene) to obtain each amplicons fragment of the target gene. High-throughput sequencing is performed on the amplification products to obtain raw sequencing data containing the base sequence and quality values ​​of each amplicons fragment. Finally, bioinformatics tools are used to align the raw sequencing data with the reference genome of the corresponding species (e.g., the human hg38 version). After format conversion and sorting, a BAM (Binary Alignment / Map) is generated, recording the alignment position, coverage, and matching quality of the sequencing reads within the target gene's amplicons region. The format (binary alignment / mapping format) file is used to obtain sequencing alignment data.

[0074] In this embodiment, obtaining the mutation site information of the target gene includes: preprocessing the BAM file to obtain a preprocessed file; using a target mutation detection tool to analyze the difference between the preprocessed file and the reference genome to obtain initial mutation results; and filtering the initial mutation results to obtain a VCF file containing the mutation site information of the target gene.

[0075] The BAM file undergoes preprocessing, including using tools to label repetitive reads generated during PCR amplification to avoid interference from repetitive sequences in subsequent analyses. The base quality values ​​of the sequencing reads are recalibrated based on the reference genome to correct base quality deviations caused by sequencing instrument errors, resulting in a preprocessed file. Subsequently, using target variant detection tools suitable for third-generation sequencing data, such as variant identification algorithms optimized for long-read sequences, the sequencing reads of the target gene amplicon region in the preprocessed file are compared with the corresponding sequences in the reference genome. This identifies differential sites such as Single Nucleotide Variations (SNVs) and Small Insertions / Deletions (InDels), generating initial variant results containing information such as variant location, reference base, variant base, and the number of reads supporting the variant. Finally, based on preset quality filtering criteria, such as variant quality value, minimum number of reads supporting the variant, and allele frequency range, the initial variant results are screened to remove low-quality variants, false positive variants, and variants in non-target gene regions. The final result is a VCF (Variant Call) containing only high-confidence variant sites within the target gene amplicon region. The format (Variant Site Call Format) file. The final generated VCF file focuses on high-confidence variants in the target gene amplicon region, providing accurate variant markers for VAF-based copy number variant determination, and also providing key molecular tags for haplotype typing. This solves the copy number variant determination error problem caused by low data quality and insufficient confidence of variant sites in traditional variant detection. Specifically, corresponding BED files are prepared according to the target gene amplicon region. To avoid abnormal sequencing depth determination in subsequent analysis results, non-overlapping region amplicon BED files need to be generated, as shown in Table 1:

[0076] Table 1. Non-overlapping region amplicon BED file format

[0077]

[0078] Prepare the VCF file after single base variation detection, filter and correct the VCF of the target variant. Specifically, revise the haplotypes annotated as "RefCall" and whose VAF (Variant Allele Frequency) variation is between 0.15 and 0.85 to 0 / 1 (originally 0 / 0) for heterozygous site extraction.

[0079] Step S12: Based on the genetic feature data of the target region, the target gene is subjected to copy number variation determination based on haplotype, copy number variation determination based on VAF, and copy number variation determination based on sequencing depth, respectively, to obtain haplotype determination results, VAF determination results, and sequencing depth determination results.

[0080] In a specific embodiment of haplotype copy number variation determination, the target gene is subjected to haplotype-based copy number variation determination based on the genetic feature data of the target region to obtain a haplotype determination result. This includes: performing haplotype typing and haplotype block labeling on the genomic sequence fragment based on the genetic feature data of the target region, and calculating the haplotype ratio of each haplotype within each haplotype block; wherein the genomic sequence fragment is located within an amplicon interval of each non-overlapping region of the target gene; if the current amplicon interval meets a preset haplotype deletion condition, a haplotype determination result is generated indicating that the current amplicon interval has a variation and the variation type is a deletion type; wherein the preset haplotype deletion condition is a first preset deletion condition or a second preset deletion condition, and the first preset deletion condition is a haplotype typing result that characterizes a single The amplicon interval contains only one haplotype. The second preset missing condition is that the haplotype typing result indicates that there are two haplotypes in a single amplicon interval and the haplotype ratio of each haplotype is less than a first preset threshold or greater than a second preset threshold. If the current amplicon interval meets the preset haplotype normal condition, a haplotype determination result indicating that there is no variation in the current amplicon interval is generated. The preset haplotype normal condition is that the haplotype typing result indicates that there are two haplotypes in a single amplicon interval and the haplotype ratio of each haplotype is within a first preset range. If the current amplicon interval does not meet the preset haplotype missing condition and the preset haplotype normal condition, a haplotype determination result indicating that there is variation in the current amplicon interval and the variation type is a repeat type is generated.

[0081] For example Figure 3 The diagram illustrates a specific method for determining copy number variation (CNV) in haplotypes. During this process, the corrected VCF file is combined with the alignment result BAM file for haplotype typing and haplotype block labeling. The number of reads and the ratio of each haplotype within each haplotype block are then calculated. The analysis results are shown in Table 2.

[0082] Table 2. Statistics of Haplotype Blocks

[0083]

[0084] Where Amplicon is the amplicon name, PS is the haplotype block, PS_Start is the start position of the haplotype block, PS_end is the end position of the haplotype block, PS_Length is the length of the haplotype block, HP is the different haplotype numbers in the haplotype block, and Reads_Count is the number of reads supported by each haplotype.

[0085] When Reads Count < 5, the corresponding HP will be removed; each amplicon interval is judged to have full-length duplication based on the number of detected haplotypes and the haplotype ratio.

[0086] If the current amplicon interval meets the preset haplotype deletion condition, a haplotype determination result is generated indicating that the current amplicon interval has a mutation and the mutation type is deletion. Specifically, if a single amplicon interval has only one haplotype, or if there are two haplotypes but the Ratio value is greater than 0.9 or less than 0.1, then the interval is considered to have a deletion (i.e., DEL). That is, if the current amplicon interval has only one haplotype, then the current amplicon interval has a mutation and the mutation type is deletion. If the current amplicon interval has two haplotypes and the haplotype ratio of each haplotype is less than a first preset threshold or greater than a second preset threshold, then the current amplicon interval has a mutation and the mutation type is deletion. The first preset threshold can be 0.9 and the second preset threshold can be 0.1.

[0087] If the current amplicon interval meets the preset haplotype normal conditions, a haplotype determination result indicating that the current amplicon interval does not have variation is generated. Specifically, if there are two haplotypes in the current amplicon interval and the haplotype ratio of each haplotype is within a first preset range, then the current amplicon interval is considered to have no variation (i.e., NORMAL). The first preset range can be 0.425-0.575. That is, if the number of haplotypes in a single amplicon interval is equal to 2 and the ratio value is between 0.425 and 0.575, then the interval is considered to have no CNV variation (i.e., NORMAL).

[0088] If the current amplicon interval does not meet the preset haplotype deletion condition and the preset haplotype normal condition, a haplotype determination result is generated to indicate that the current amplicon interval has a variation and the variation type is a duplication type. In other words, the interval is considered to have CNV duplication (i.e., DUP, duplicate).

[0089] The specific result format is shown in Table 3:

[0090] Table 3 Haplotype Determination Results

[0091]

[0092] Wherein, Segment_Chr is the chromosome name, Segment_Start is the start position of the haplotype block, Segment_end is the end position of the haplotype block, Phase_Block is the name of the haplotype block, Ratio is the support ratio of two haplotype reads (i.e., haplotype ratio), CNV_Type is the CNV determination result, and Amplicon_name is the amplicon name.

[0093] In a specific embodiment of VAF-based copy number variation determination, the target gene is subjected to VAF-based copy number variation determination based on the variation site information to obtain a VAF determination result, including: determining the VAF value of each amplicon interval in the target gene based on the variation site information; if the current amplicon interval meets a preset VAF deletion condition, a VAF determination result is generated indicating that the current amplicon interval has a variation and the variation type is deletion; wherein, the preset VAF deletion condition is that the VAF value is less than a third preset threshold or greater than a fourth preset threshold; if the current amplicon interval meets a preset VAF normal condition, a VAF determination result is generated indicating that the current amplicon interval does not have a variation; wherein, the preset VAF deletion condition is that the VAF value is within a second preset range; if the current amplicon interval does not meet the preset VAF deletion condition and the preset VAF normal condition, a VAF determination result is generated indicating that the current amplicon interval has a variation and the variation type is duplication.

[0094] SNV variant information was extracted from the corrected VCF file, and the format of the extracted result file is shown in Table 4.

[0095] Table 4. Variant Site Information

[0096]

[0097] Where Chr is the chromosome name, Pos is the variant location, Ref is the reference genome base, Alt is the variant base, and AF is the allelic frequency.

[0098] CNV variant type determination is performed by combining the VAF values ​​of all variants in each amplicon. Specifically, the VAF values ​​of variants within each amplicon are analyzed, i.e., the VAF value of each amplicon interval in the target gene is determined based on the variant site information. If the current amplicon interval meets the preset VAF deletion condition, a VAF determination result is generated indicating that the current amplicon interval has a variant and the variant type is deletion. Specifically, if the current amplicon interval VAF value is less than the third preset threshold or greater than the fourth preset threshold, it is considered that the site has been deleted (i.e., DEL). The third preset threshold is 0.9, and the fourth preset threshold is 0.1. If the current amplicon interval meets the preset VAF normal condition, a VAF determination result is generated indicating that the current amplicon interval has no variant. Specifically, if the current amplicon interval VAF value is within the second preset range, it is considered that the site has not been mutated (i.e., NORMAL). The second preset range can be 0.425-0.575. If the current amplicon interval does not meet the preset VAF deletion condition and the preset VAF normal condition, a VAF determination result is generated indicating that the current amplicon interval has a variation and the variation type is duplication, that is, the site is considered to have duplication (i.e., DUP). Finally, the number of bases with the corresponding CNV variation in each amplicon is counted.

[0099] The VAF determination results are shown in Table 5:

[0100] Table 5 VAF Judgment Results

[0101] Where Segment_Chr is the chromosome name, Segment_Start is the amplicon start position, Segment_end is the amplicon end position, Pos is the base variation position, VAF is the allelic frequency, CNV_Type is the CNV determination result, and Amplicon_name is the amplicon name.

[0102] In a specific embodiment of copy number variation determination based on sequencing depth, the target gene is subjected to copy number variation determination based on sequencing depth based on the genetic feature data of the target region to obtain a sequencing depth determination result. This includes: determining the target region of the target gene based on the genetic feature data of the target region; wherein the target region includes multiple amplicon regions; performing sequencing depth statistics on the amplicon regions using a preset window size as a sliding window to obtain the average sequencing depth of each amplicon region, and determining the sequencing depth ratio of each sliding window based on the average sequencing depth; if the current amplicon region meets a preset sequencing depth missing condition, then generating a characterization of the current amplicon region. The sequencing depth determination result indicates that the amplicon interval has a mutation and the mutation type is deletion; wherein, the preset sequencing depth deletion condition is that the sequencing depth ratio is greater than a fifth preset threshold; if the current amplicon interval meets the preset normal sequencing depth condition, then a sequencing depth determination result indicating that the current amplicon interval has no mutation is generated; wherein, the preset normal sequencing depth condition is that the sequencing depth ratio is not greater than a sixth preset threshold; if the current amplicon interval does not meet the preset sequencing depth deletion condition and the preset normal sequencing depth condition, then a sequencing depth determination result indicating that the current amplicon interval has a mutation and the mutation type is repetitive is generated.

[0103] For example Figure 4 The diagram illustrates a specific copy number variation (CNV) determination based on sequencing depth. To exclude cases where a small number of variant sites exist in the sample, leading to abnormal CNV identification results, local CNV determination is performed by combining the sliding window sequencing depth statistics for each amplicon, reducing false positive errors. The target region of the target gene is determined based on the genetic characteristic data of the target region. This target region includes multiple amplicon regions. Sequencing depth statistics are performed on the amplicon regions using a preset window size as the sliding window. The average sequencing depth of each amplicon region is obtained, and the sequencing depth ratio (Log2 Ratio) of each sliding window is determined based on the average sequencing depth. Specifically, for a single amplicon (without overlapping amplicon BED files), sequencing depth statistics are performed using a 100bp sliding window. The specific results are shown in Table 6.

[0104] Table 6. Statistical results of sequencing depth

[0105]

[0106] Where Amplicon is the amplicon name, Start is the start position of the amplicon, End is the end position of the amplicon, Window_Start is the start position of the sliding window, Window_End is the end position of the sliding window, amplicon_avg_depth is the average sequencing depth of the amplicon, and Window_Depth is the sequencing depth of the sliding window.

[0107] The CNV (Cyclic Binary Separation) algorithm is used for CNV (Collapse Novel) determination. The specific process is as follows: 1) If the current amplicon interval meets the preset sequencing depth deletion condition, a sequencing depth determination result is generated indicating that the current amplicon interval has a mutation and the mutation type is deletion. Specifically, if the current amplicon interval sequencing depth ratio (Log2 Ratio) is greater than the fifth preset threshold, then the site is considered to have a deletion (DEL). The fifth preset threshold can be 0.58. 2) If the current amplicon interval meets the preset sequencing depth normal condition, a sequencing depth determination result is generated indicating that the current amplicon interval has no mutation. Specifically, if the current amplicon interval sequencing depth ratio (Log2 Ratio) is not greater than the sixth preset threshold, then the site is considered to have no mutation (NORMAL). The sixth preset threshold can be -1. 3) If the current amplicon interval does not meet the preset sequencing depth deletion condition or the preset sequencing depth normal condition, a sequencing depth determination result is generated indicating that the current amplicon interval has a mutation and the mutation type is repetition, i.e., the CNV type is determined to be NORMAL. See Table 7 for details.

[0108] Table 7 Sequencing Depth Determination Results

[0109] Wherein, Segment_Chr is the chromosome name, Segment_Start is the amplicon start position, Segment_End is the amplicon end position, Phase_Block is the sliding window number, Log2_Ratio is the sequencing depth ratio, CNV_Type is the CNV determination result, and Amplicon_name is the amplicon name.

[0110] Step S13: Determine the target proportion of variant amplicon in the target gene based on the haplotype determination result and the VAF determination result, and determine the preliminary variant detection result of the target gene based on the relationship between the target proportion and the preset proportion threshold.

[0111] In this embodiment, determining the target proportion of variant amplicones in the target gene based on the haplotype determination result and the VAF determination result, and determining the preliminary variant detection result of the target gene based on the relationship between the target proportion and a preset proportion threshold, includes: identifying amplicones in the haplotype determination result and / or the VAF determination result that indicate a mutation in the target gene and that the mutation type is deletion as first variant amplicones; identifying amplicones in the haplotype determination result and / or the VAF determination result that indicate a mutation in the target gene and that the mutation type is duplication as second variant amplicones; determining a first target proportion of the first variant amplicones and a second target proportion of the second variant amplicones in the target gene based on the number of the first variant amplicones and the number of the second variant amplicones; if the first target proportion is greater than the preset proportion threshold, generating a preliminary variant detection result indicating that the gene copy number of the target gene has a mutation and that the mutation type is deletion; if the second target proportion is greater than the preset proportion threshold, generating a preliminary variant detection result indicating that the gene copy number of the target gene has a mutation and that the mutation type is duplication.

[0112] First, based on the haplotype determination results and VAF determination results, the number of amplicon supported by each CNV type is counted. Specifically, the amplicon in the haplotype determination results and / or VAF determination results that indicates a mutation in the target gene and that the mutation type is deletion is identified as the first variant amplicon. For example, if the haplotype determination result indicates that amplicon A1 has a mutation and that the mutation type is deletion, even if the VAF determination result indicates that amplicon A1 does not have a mutation, amplicon A1 is still the first variant amplicon. Similarly, if the haplotype determination result indicates that amplicon A2 has a mutation and that the mutation type is deletion, the VAF determination result also indicates that amplicon A2 has a mutation. If a variant exists and its type is deletion, then amplicon A2 is the first variant amplicon. In other words, if either the haplotype determination result or the VAF determination result indicates that amplicon An has a variant and its type is deletion, then amplicon An is the first variant amplicon. Similarly, amplicons that indicate a variant in the target gene and its type is duplication in either the haplotype determination result or the VAF determination result are identified as the second variant amplicon. That is, if either the haplotype determination result or the VAF determination result indicates that amplicon An has a variant and its type is duplication, then amplicon An is the second variant amplicon.

[0113] Next, calculate the percentage of each type of CNV amplicon. Based on the number of first variant amplicon and the number of second variant amplicon, determine the first target percentage of the first variant amplicon and the second target percentage of the second variant amplicon in the target gene. For example, if there are a total of 10 amplicon in the target gene, with 2 first variant amplicon and 6 second variant amplicon, then the first target percentage is 20% and the second target percentage is 70%.

[0114] Furthermore, the preliminary mutation detection result of the target gene is determined based on the relationship between the target percentage and the preset percentage threshold. Specifically, if the first target percentage is greater than the preset percentage threshold, a preliminary mutation detection result is generated indicating that the gene copy number of the target gene has a mutation and the mutation type is deletion. If the second target percentage is greater than the preset percentage threshold, a preliminary mutation detection result is generated indicating that the gene copy number of the target gene has a mutation and the mutation type is duplication. The preset percentage threshold can be 50%. If the first target percentage is 20% and the second target percentage is 70%, then the second target percentage is greater than the preset percentage threshold, and a preliminary mutation detection result is generated indicating that the gene copy number of the target gene has a mutation and the mutation type is duplication.

[0115] Step S14: If the preliminary mutation detection result matches the sequencing depth determination result, then the preliminary mutation detection result is determined as the target mutation detection result for the gene copy number of the target gene.

[0116] For example Figure 5 The diagram illustrates a specific gene copy number variation detection method. If the preliminary variation detection result matches the sequencing depth determination result, the preliminary variation detection result is determined as the target gene copy number variation detection result. For example, if the preliminary variation detection result indicates that the target gene copy number has a variation and the variation type is repeat, and the sequencing depth determination result also indicates that the target gene copy number has a variation and the variation type is repeat, then the two results match, and the result that the target gene copy number has a variation and the variation type is repeat can be directly taken as the target variation detection result. Conversely, if the preliminary variation detection result does not match the sequencing depth determination result, then the final target variation detection result will be determined manually based on the preliminary variation detection result and the sequencing depth determination result.

[0117] Furthermore, the target variant detection results can be annotated with gene regions, OMIM (Online Mendelian Inheritance in Man) associated diseases, gnomAD database, etc., and the gene (region) can be mapped by combining VAF and sequencing depth.

[0118] The beneficial effects of this application are as follows: This application obtains genetic feature data of the target region of a target gene in a biological sample to be tested; wherein, the genetic feature data of the target region includes regional location information, sequencing alignment data, and mutation site information; based on the genetic feature data of the target region, the target gene is subjected to copy number variation determination based on haplotype, copy number variation determination based on VAF, and copy number variation determination based on sequencing depth to obtain haplotype determination results, VAF determination results, and sequencing depth determination results; the target proportion of variant amplicon in the target gene is determined according to the haplotype determination results and the VAF determination results, and the preliminary variation detection result of the target gene is determined according to the relationship between the target proportion and a preset proportion threshold; if the preliminary variation detection result matches the sequencing depth determination result, the preliminary variation detection result is determined as the target variation detection result of the gene copy number of the target gene. Therefore, by acquiring genetic feature data of the target region containing regional location information, sequencing alignment data, and variant site information, a precise foundation is laid for gene copy number variation determination. Specifically, regional location information accurately delineates the target gene amplicon range and eliminates interference from irrelevant regions; sequencing alignment data provides a quantitative basis for coverage; and variant site information assists in filtering interfering regions and haplotype classification, solving the problem of insufficient data dimensions in traditional detection methods. Simultaneously, three gene copy number variation determination methods based on haplotype, VAF, and sequencing depth are analyzed in parallel. First, preliminary variation detection results are determined based on haplotype and VAF determination results, and then consistency is verified with sequencing depth determination results. This reduces false positives and false negatives. If the preliminary results contradict the sequencing depth determination results, errors can be eliminated; if they are consistent, reliability is improved, reducing misdiagnosis and missed diagnosis caused by errors in a single method. Ultimately, this systematically solves the pain points of inaccurate interval determination, low abundance missed detection, and low result reliability in traditional gene copy number variation detection, achieving high accuracy, high sensitivity, and full-region coverage detection of gene copy number variations in the target gene.

[0119] The following analysis of test samples will be used to illustrate this application.

[0120] 1) Taking the PKD1 gene multiplex amplification system as an example, two PKD1 amplification systems are shown in Tables 8 and 9:

[0121] Table 8. Seven amplicon systems

[0122]

[0123] Table 9. Four amplicon systems

[0124]

[0125] 2) Test sample information: PKD1 positive samples were selected for third-generation amplicon sequencing. Two samples had full-length PKD1 repeats, and one sample had a full-length PKD1 deletion. Specific information is shown in Table 10.

[0126] Table 10 Third-generation amplicon sequencing

[0127]

[0128] 3) Data quality control: The three sequencing data were subjected to quality control analysis and aligned to the hg38 reference genome. The quality control and alignment results were statistically analyzed. The coverage was greater than 99.9%, the sequencing quality value Q20 was greater than 90%, and the average sequencing depth was greater than 200X. The data can be used for subsequent full-length CNV identification. The details are shown in Table 11.

[0129] Table 11 Quality Control and Comparison Results

[0130] Wherein, Sample is the sample number, Raw Bases (Mb) is the number of raw sequencing bases, Q10 (%) is the number of bases with a sequencing quality value greater than 10, Q20 (%) is the number of bases with a sequencing quality value greater than 20, GC (%) is the GC content, Ave_Depth (X) is the average sequencing depth of the target region, Depth>=1X is the proportion of sites with a sequencing depth greater than 1X, Depth>=10X is the proportion of sites with a sequencing depth greater than 10X, and Depth>=50X is the proportion of sites with a sequencing depth greater than 50X.

[0131] 4) Full-length CNV identification results: According to the above analysis process, the results show that samples S1 and S2 were both identified as full-length duplications of PKD1, while sample S3 was identified as full-length missing PKD1.

[0132] 4.1) The identification results of sample S1 are shown in Table 12:

[0133] Table 12. Identification results of S1 sample

[0134] Wherein, Amplicon_name is the amplicon name, CNV_Type is the CNV determination result, Segment_Chr is the chromosome name, Segment_Start is the start position of the amplicon, Segment_End is the end position of the amplicon, Call_Method is the source of the CNV variant identification result, and Vaf_num is the number of SNVs supporting the occurrence of this variant type and the total number of SNVs in this amplicon interval.

[0135] For example Figure 6The first specific detection result diagram shown illustrates the plotting of the target variant detection results. Figure 6 The red scatter dots represent the VAF value of SNV, the black scatter dots represent the sequencing depth ratio Log2 Ratio, and the green dashed line represents the AF baseline value of 0.5.

[0136] 4.2) The identification results of sample S2 are shown in Table 13:

[0137] Table 13. Identification results of S2 sample

[0138] Wherein, Amplicon_name is the amplicon name, CNV_Type is the CNV determination result, Segment_Chr is the chromosome name, Segment_Start is the start position of the amplicon, Segment_End is the end position of the amplicon, Call_Method is the source of the CNV variant identification result, and Vaf_num is the number of SNVs supporting the occurrence of this variant type and the total number of SNVs in this amplicon interval.

[0139] For example Figure 7 The second specific detection result diagram shown illustrates the target variant detection results. Figure 7 The red scatter dots represent the VAF value of SNV, the black scatter dots represent the sequencing depth ratio Log2 Ratio, and the green dashed line represents the AF baseline value of 0.5.

[0140] 4.3) The identification results of sample S3 are shown in Table 14:

[0141] Table 14. Identification Results of Sample S3

[0142] Wherein, Amplicon_name is the amplicon name, CNV_Type is the CNV determination result, Segment_Chr is the chromosome name, Segment_Start is the start position of the amplicon, Segment_End is the end position of the amplicon, Call_Method is the source of the CNV variant identification result, and Vaf_num is the number of SNVs supporting the occurrence of this variant type and the total number of SNVs in this amplicon interval.

[0143] For example Figure 8 The third specific detection result diagram shown illustrates the target variant detection results. Figure 8 The red scatter dots represent the VAF value of SNV, the black scatter dots represent the sequencing depth ratio Log2 Ratio, and the green dashed line represents the AF baseline value of 0.5.

[0144] See Figure 9 As shown in the embodiment of this application, an information processing device for detecting gene copy number variations is disclosed, comprising:

[0145] The data acquisition module 11 is used to acquire the genetic feature data of the target region of the target gene in the biological sample to be tested; wherein, the genetic feature data of the target region includes regional location information, sequencing alignment data, and mutation site information;

[0146] The determination module 12 is used to determine the copy number variation of the target gene based on haplotype, copy number variation based on VAF, and copy number variation based on sequencing depth based on the genetic feature data of the target region, so as to obtain the haplotype determination result, VAF determination result, and sequencing depth determination result.

[0147] The preliminary result determination module 13 is used to determine the target proportion of variant amplicon in the target gene based on the haplotype determination result and the VAF determination result, and to determine the preliminary variant detection result of the target gene based on the relationship between the target proportion and the preset proportion threshold.

[0148] The mutation result determination module 14 is used to determine the preliminary mutation detection result as the target mutation detection result of the gene copy number of the target gene if the preliminary mutation detection result matches the sequencing depth determination result.

[0149] Furthermore, embodiments of this application also provide an electronic device. Figure 10 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0150] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Specifically, it may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the information processing method for detecting gene copy number variations performed by the electronic device disclosed in any of the foregoing embodiments.

[0151] In this embodiment, the power supply 23 is used to provide operating voltage for various hardware devices on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0152] The processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 21 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0153] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored on it include operating system 221, computer program 222 and data 223, etc., and the storage method can be temporary storage or permanent storage.

[0154] The operating system 221 manages and controls the various hardware devices and computer programs 222 on the electronic device to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. The operating system can be Windows, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of performing the information processing method for gene copy number variation detection executed by the electronic device as disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the electronic device from external devices, as well as data collected by its own input / output interface 25.

[0155] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned information processing method for detecting gene copy number variations. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0156] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0157] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software module may be located in random access memory (RAM), memory, read-only memory (ROM), electrically programmable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), register, hard disk, removable disk, CD-ROM (Compact Disc Read-Only Memory), or any other form of storage medium known in the art.

[0158] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0159] The above provides a detailed description of the information processing method, apparatus, device, and medium for detecting gene copy number variations provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only intended to help understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. An information processing method for gene copy number variation detection, characterized by, Applied to computer devices, including: Obtain genetic feature data of the target region of the target gene in the biological sample to be tested; wherein, the genetic feature data of the target region includes regional location information, sequencing alignment data, and mutation site information; Based on the genetic feature data of the target region, the target gene is subjected to copy number variation determination based on haplotype, copy number variation determination based on VAF, and copy number variation determination based on sequencing depth, respectively, to obtain haplotype determination results, VAF determination results, and sequencing depth determination results. The target proportion of variant amplicon in the target gene is determined based on the haplotype determination result and the VAF determination result, and the preliminary mutation detection result of the target gene is determined based on the relationship between the target proportion and the preset proportion threshold. If the preliminary mutation detection result matches the sequencing depth determination result, then the preliminary mutation detection result is determined as the target mutation detection result for the gene copy number of the target gene.

2. The information processing method for gene copy number variation detection according to claim 1, characterized in that, Obtain the regional location information and sequencing alignment data of the target gene, including: Generate a BED file containing the regional location information of the target gene based on the genomic coordinates of the target gene; The sample genome is amplified according to the coordinates of the amplicon region of the target gene as described in the BED file to obtain the target amplicon fragment; wherein, the sample genome is the complete genetic material of the biological sample to be tested and includes the target gene; The target amplicon fragment is sequenced to obtain raw sequencing data; The raw sequencing data is compared with a reference genome in a standard genome sequence to obtain a BAM file containing the sequencing alignment data.

3. The information processing method for gene copy number variation detection according to claim 2, characterized in that, Obtain information on the mutation sites of the target gene, including: The BAM file is preprocessed to obtain a preprocessed file; The differences between the preprocessed file and the reference genome are analyzed using a target variant detection tool to obtain initial variant results; The initial mutation results are filtered to obtain a VCF file containing mutation site information of the target gene.

4. The information processing method for detecting gene copy number variations according to claim 1, characterized in that, Based on the genetic feature data of the target region, the target gene is subjected to copy number variation determination based on haplotype to obtain haplotype determination results, including: Based on the genetic feature data of the target region, haplotype classification and haplotype block labeling are performed on the genome sequence fragments to calculate the haplotype ratio of each haplotype within each haplotype block; wherein, the genome sequence fragment is located within the amplicon interval of each non-overlapping region of the target gene; If the current amplicon interval meets the preset haplotype deletion condition, a haplotype determination result is generated indicating that the current amplicon interval has a mutation and the mutation type is a deletion type; wherein, the preset haplotype deletion condition is a first preset deletion condition or a second preset deletion condition, the first preset deletion condition is that the haplotype typing result indicates that there is only 1 haplotype in a single amplicon interval, and the second preset deletion condition is that the haplotype typing result indicates that there are 2 haplotypes in a single amplicon interval and the haplotype ratio of each haplotype is less than a first preset threshold or greater than a second preset threshold; If the current amplicon interval meets the preset haplotype normal condition, a haplotype determination result is generated indicating that there is no variation in the current amplicon interval; wherein, the preset haplotype normal condition is that the haplotype typing result indicates that there are 2 haplotypes in a single amplicon interval and the haplotype ratio of each haplotype is within a first preset range. If the current amplicon interval does not meet the preset haplotype deletion condition and the preset haplotype normal condition, a haplotype determination result is generated indicating that the current amplicon interval has a variation and the variation type is a repeat type.

5. The information processing method for detecting gene copy number variations according to claim 1, characterized in that, Based on the mutation site information, the target gene is subjected to VAF-based copy number variation determination to obtain VAF determination results, including: The VAF value of each amplicon region in the target gene is determined based on the mutation site information; If the current amplicon interval meets the preset VAF missing condition, a VAF determination result is generated to indicate that the current amplicon interval has a mutation and the mutation type is a missing type; wherein, the preset VAF missing condition is that the VAF value is less than a third preset threshold or greater than a fourth preset threshold; If the current amplicon interval meets the preset VAF normal condition, a VAF determination result indicating that the current amplicon interval does not have a variation is generated; wherein, the preset VAF missing condition is that the VAF value is within a second preset range; If the current amplicon interval does not meet the preset VAF missing condition and the preset VAF normal condition, a VAF determination result is generated indicating that the current amplicon interval has a variation and the variation type is repeat.

6. The information processing method for detecting gene copy number variations according to claim 1, characterized in that, Based on the genetic feature data of the target region, copy number variation of the target gene is determined based on sequencing depth to obtain sequencing depth determination results, including: The target region is determined based on the genetic feature data of the target region; wherein, the target region includes multiple amplicon regions; The sequencing depth of the amplicon intervals is statistically analyzed using a preset window size as a sliding window to obtain the average sequencing depth of each amplicon interval, and the sequencing depth ratio of each sliding window is determined based on the average sequencing depth. If the current amplicon region meets the preset sequencing depth missing condition, a sequencing depth determination result is generated to indicate that the current amplicon region has a mutation and the mutation type is a deletion type; wherein, the preset sequencing depth missing condition is that the sequencing depth ratio is greater than a fifth preset threshold. If the current amplicon region meets the preset normal sequencing depth condition, a sequencing depth determination result is generated indicating that the current amplicon region does not have a variation; wherein, the preset normal sequencing depth condition is that the sequencing depth ratio is not greater than a sixth preset threshold. If the current amplicon interval does not meet the preset sequencing depth missing condition and the preset sequencing depth normal condition, a sequencing depth determination result is generated indicating that the current amplicon interval has a variation and the variation type is repetitive.

7. The information processing method for detecting gene copy number variations according to any one of claims 1 to 6, characterized in that, The step of determining the target proportion of variant amplicon in the target gene based on the haplotype determination result and the VAF determination result, and determining the preliminary variant detection result of the target gene based on the relationship between the target proportion and a preset proportion threshold, includes: The amplicon in the haplotype determination result and / or the VAF determination result that indicates a mutation in the target gene and that the mutation type is deletion is identified as the first variant amplicon; The amplicon in the haplotype determination result and / or the VAF determination result that indicates a variation in the target gene and that the variation type is repeat is identified as the second variant amplicon; The first target proportion of the first variant amplicon and the second target proportion of the second variant amplicon in the target gene are determined based on the number of the first variant amplicon and the number of the second variant amplicon. If the proportion of the first target is greater than the preset proportion threshold, a preliminary mutation detection result is generated to indicate that the gene copy number of the target gene has a variation and the variation type is a deletion type. If the proportion of the second target is greater than the preset proportion threshold, a preliminary mutation detection result is generated, indicating that the gene copy number of the target gene has a variation and the variation type is a repeat type.

8. An information processing device for detecting gene copy number variations, characterized in that, Applied to computer devices, including: The data acquisition module is used to acquire genetic feature data of the target region of the target gene in the biological sample to be tested; wherein, the genetic feature data of the target region includes regional location information, sequencing alignment data, and mutation site information; The determination module is used to determine the copy number variation of the target gene based on haplotype, copy number variation based on VAF, and copy number variation based on sequencing depth based on the genetic feature data of the target region, so as to obtain the haplotype determination result, VAF determination result, and sequencing depth determination result. The preliminary result determination module is used to determine the target proportion of variant amplicon in the target gene based on the haplotype determination result and the VAF determination result, and to determine the preliminary variant detection result of the target gene based on the relationship between the target proportion and the preset proportion threshold. The mutation result determination module is used to determine the preliminary mutation detection result as the target mutation detection result of the gene copy number of the target gene if the preliminary mutation detection result matches the sequencing depth determination result.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the information processing method for detecting gene copy number variations as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the information processing method for detecting gene copy number variations as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Detection method and system for copy number variation in genome

    CN107423534A

  • Detecting cancer mutations and aneuploidy in chromosomal segments

    US20160333416A1