Method, computing device and program product for determining copy number of a gene of interest in a family of homologous genes
By introducing a depth ratio correction mechanism for absolute homology and difference regions into high-throughput sequencing data, the accuracy and stability issues of homology gene family copy number detection were resolved, enabling accurate determination of the copy number of homology gene families.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BERRYGENOMICS CO LTD
- Filing Date
- 2026-03-26
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies suffer from inaccurate copy number calculations due to insufficient or uneven sequencing depth in highly homologous regions when determining the copy number of genes in homologous gene families. This is especially true in high-throughput sequencing data where it is difficult to effectively utilize sequencing signals from highly homologous regions.
By introducing a depth ratio correction mechanism for absolute homologous regions and cross-regional regions, sequencing depth information of highly homologous regions is reused to screen stable homologous regions. Combined with the depth of differential regions, the total copy number of homologous gene families and the copy number of target member genes are determined.
It significantly improves the accuracy and stability of homologous gene copy number detection, is applicable to different sequencing platforms and capture probe systems, and is compatible with whole genome sequencing and hybridization capture sequencing data.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This application relates to bioinformatics processing, and more specifically, to methods, computing devices, computer-readable storage media, and computer program products for determining the copy number of a target member gene in a homologous gene family. Background Technology
[0002] The human genome contains many highly homologous gene regions, such as those associated with spinal muscular atrophy (SMA). SMN1 / SMN2 Genes and thalassemia-related HBA1 / HBA2 Genes, etc. These genes typically have highly consistent nucleotide sequences in most regions, including promoters, introns, and exons, with base differences only at a very few sites.
[0003] Copy number changes in these genes are often closely related to the occurrence, development, and phenotypic severity of specific genetic diseases. For example, in the case of SMA, its pathogenesis mainly involves... SMN1 The deletion of a gene, when an individual's SMN1 SMA can occur when a gene has only one copy or zero copies. SMN2 Gene copy number has been shown to be a significant factor influencing the severity of SMA clinical phenotype. Therefore, accurately distinguishing and determining the copy number of each member gene in a homologous gene family is of great importance for the clinical diagnosis of related diseases and carrier screening.
[0004] Currently, homologous gene copy number detection schemes based on next-generation sequencing (NGS) technology typically first calculate the total copy number of the homologous gene family, then calculate the relative proportion of copy numbers based on the specific differential sites between each member gene, and finally infer the copy number of each member gene.
[0005] However, in practical applications, due to the high sequence consistency among genes in homologous gene families, alignment software often fails to accurately map highly homologous reads to their unique physical origins when performing sequence alignment based on high-throughput sequencing data. According to conventional alignment quality control standards (e.g., requiring an alignment quality value MAPQ ≥ 20), these indistinguishable "multiple alignment sequences" are usually considered low-quality alignments or noise and are filtered out and discarded. This results in the sequencing depth of highly homologous regions being artificially zeroed or exhibiting extremely low depth, making subsequent copy number calculations lack effective data information from highly homologous regions, or forcing them to rely solely on a very small number of specific differential sites for copy number calculations.
[0006] Furthermore, in common scenarios such as hybridization capture sequencing, probe design is often limited by factors such as GC content and repetitive sequences in genomic regions, leading to differences in probe density and capture efficiency across different gene regions. This inherent non-uniformity results in significant inhomogeneity in the sequencing depth distribution across regions, making it difficult to directly apply the uniformity-based depth correction strategies commonly used in whole-genome sequencing (WGS). Forcibly using internal reference genes for standardization often leads to difficulties in threshold setting due to inconsistent probe response characteristics, making it challenging to obtain accurate copy number signals. Simultaneously, differences in GC content, systematic fluctuations in sequencing platforms, batch effects in experimental operations, and variations in sample quality can also cause a mismatch between the depth distribution characteristics of the sample being tested and the external background set, ultimately affecting the accuracy of copy number analysis.
[0007] Therefore, it remains necessary to provide a method that can stably and accurately determine the copy number of highly homologous genes based on high-throughput sequencing data. Summary of the Invention
[0008] This application provides a method, computing device, computer-readable storage medium, and computer program product for determining the copy number of target members in a homologous gene family. It can stably and accurately determine the copy number of highly homologous genes based on high-throughput sequencing data. For regions in a homologous gene family where a large number of base sequences are completely identical and difficult to uniquely locate under conventional sequencing alignment and quality control conditions, unlike conventional procedures that discard multiple alignment sequences (such as sequencing sequences with MAPQ=0), this application's method introduces an absolute homologous region and cross-region depth ratio correction mechanism. This algorithmically reuses the sequencing depth information of these highly homologous regions, making greater use of the sequencing signal contributed by each base in the homologous gene region, significantly improving the accuracy of copy number detection. Validated on tens of thousands of samples, this method demonstrates excellent compatibility and accuracy across different sequencing platforms and capture probe systems, stably and accurately determining the copy number of homologous genes.
[0009] According to a first aspect of this application, a method is provided for determining the copy number of a target member gene in a homologous gene family. The method includes: determining multiple analysis intervals of a reference genome, wherein the analysis intervals within the regions containing each member gene of the homologous gene family correspond to each other based on the alignment results of the reference sequences of each member gene, and the multiple analysis intervals include at least one absolutely homologous interval where the base sequences are completely identical among the reference sequences of each member gene, and at least one differential interval where there are base differences representing specific base differences of the member gene at differential sites among the reference sequences of each member gene; determining the sequencing depth of the test sample and the reference sample in each analysis interval based on the alignment results of sequencing data of the test sample and the reference sample with the retained multiple alignment sequences of the reference genome; screening stable homologous intervals based on the sequencing depth of the reference sample and the test sample in the absolutely homologous intervals to determine the total copy number of the homologous gene family; and determining the copy number of the target member gene based on the sequencing depth of the test sample in the differential intervals supporting each member gene, and the total copy number of the homologous gene family.
[0010] In some embodiments, the absolute homology region does not include regions whose base sequences are completely identical to those of other genomic regions outside the target homologous gene family.
[0011] In some embodiments, the differential sites do not include InDel sites and / or differential sites lacking single-member gene directionality determined based on allele frequencies in population samples.
[0012] In some embodiments, after determining multiple analysis intervals of the reference genome, the method further includes: filtering the analysis intervals according to preset filtering conditions, the preset filtering conditions including: removing analysis intervals whose GC content does not meet a preset GC content range based on the GC content of each analysis interval; and / or removing analysis intervals whose interval length does not meet a preset length range based on the interval length of each analysis interval.
[0013] In some embodiments, the reference sample is a sample whose similarity to the sequencing depth distribution of the sample to be tested in each analysis interval meets a preset similarity condition.
[0014] In some embodiments, determining the sequencing depth of the test sample and the reference sample in each analysis interval includes: determining the original sequencing depth of the test sample and the reference sample in each analysis interval based on the alignment results of the sequencing data of the test sample and the reference sample with the preserved multiple alignment sequence of the reference genome; and performing depth correction on the original sequencing depth to determine the sequencing depth of the test sample and the reference sample in each analysis interval.
[0015] In some embodiments, the depth correction includes GC correction based on the GC content of the analysis interval, and data volume normalization correction based on the total amount of sequencing data.
[0016] In some embodiments, the sequencing data is obtained based on high-throughput sequencing.
[0017] In some embodiments, the sequencing data is obtained based on hybridization capture high-throughput sequencing.
[0018] In some embodiments, before screening stable homologous regions, the method further includes: determining the global baseline depth of the absolute homologous regions based on the central tendency index of the sequencing depth of the reference sample in each absolute homologous region; and, based on the global baseline depth, dividing the absolute homologous regions into subsets of regions at different depth levels, and independently screening stable homologous regions in each subset of regions.
[0019] In some embodiments, screening stable homologous regions based on the sequencing depth of the reference sample and the test sample in the absolute homologous regions includes: determining the degree of fluctuation in the sequencing depth of each absolute homologous region based on the dispersion index of the sequencing depth of the reference sample in each absolute homologous region, and selecting absolute homologous regions whose fluctuation degree meets a first preset stability screening condition as first stable homologous regions; determining a reference depth value for each first stable homologous region based on the sequencing depth of the reference sample and the test sample in each first stable homologous region, calculating a deviation cost function value corresponding to each first stable homologous region, and selecting first stable homologous regions whose deviation cost function value meets a second preset stability screening condition as second stable homologous regions; and determining the predicted total copy number of the homologous gene family based on the sequencing depth of the test sample in each second stable homologous region, calculating a fitting cost function value corresponding to each second stable homologous region, and selecting second stable homologous regions whose fitting cost function value meets a third preset stability screening condition as stable homologous regions.
[0020] In some embodiments, determining the degree of fluctuation in sequencing depth of each absolute homology interval includes: for each of the absolute homology intervals: determining the degree of fluctuation in sequencing depth of the current absolute homology interval based on the dispersion index of sequencing depth of the reference sample in the current absolute homology interval and the deviation of the sequencing depth of the test sample in the current absolute homology interval from the sequencing depth of the reference sample in the current absolute homology interval.
[0021] In some embodiments, determining the reference depth value for each first stable homology interval includes: for each of the first stable homology intervals: based on the sequencing depth of the test sample in each of the remaining first stable homology intervals other than the current first stable homology interval, and the reference depth ratio between the current first stable homology interval and the corresponding remaining first stable homology intervals, determining the reference depth value of the current first stable homology interval, wherein the reference depth ratio is determined based on a central tendency index of the sequencing depth of the reference sample in each first stable homology interval.
[0022] In some embodiments, calculating the deviation cost function value corresponding to each first stable homology interval includes: for each of the first stable homology intervals: determining the deviation cost function value corresponding to the current first stable homology interval based on the statistical distance between the sequencing depth of the test sample in the current first stable homology interval and the reference depth value of the current first stable homology interval, and the degree of deviation of the central tendency index of the sequencing depth of the test sample in the current first stable homology interval and the sequencing depth of the reference sample in the current first stable homology interval.
[0023] In some embodiments, determining the predicted total copy number of the homologous gene family includes: determining the observed total copy number corresponding to each second stable homologous region based on the sequencing depth of the sample to be tested in each second stable homologous region; calculating the total copy number cost function value corresponding to each hypothetical total copy number based on the observed total copy number corresponding to each second stable homologous region and multiple hypothetical total copy numbers; and selecting the hypothetical total copy number corresponding to the minimum total copy number cost function value as the predicted total copy number.
[0024] In some embodiments, calculating the total copy number cost function value corresponding to each hypothesized total copy number includes: for each of the plurality of hypothesized total copy numbers: determining the total copy number cost function value corresponding to the current hypothesized total copy number based on the statistical distance of the total copy number observations corresponding to each second stable homology region relative to the hypothesized total copy number value, and the degree of deviation of the sequencing depth of the test sample in each second stable homology region relative to the sequencing depth of the reference sample in the corresponding second stable homology region, wherein the hypothesized total copy number value is determined based on the current hypothesized total copy number.
[0025] In some embodiments, calculating the fitting cost function value corresponding to each second stable homology interval includes: for each of the second stable homology intervals: determining the total copy number observation corresponding to the current second stable homology interval based on the sequencing depth of the test sample in the current second stable homology interval; and determining the fitting cost function value corresponding to the current second stable homology interval based on the statistical distance of the total copy number observation corresponding to the current second stable homology interval to the predicted total copy number, and the degree of deviation of the central tendency index of the sequencing depth of the test sample in the current second stable homology interval to the sequencing depth of the reference sample in the current second stable homology interval, wherein the predicted total copy number is determined based on the predicted total copy number.
[0026] In some embodiments, determining the total copy number of the homologous gene family includes: correcting the sequencing depth of each stable homologous region based on the sequencing depth of the sample under test in each stable homologous region and a reference depth ratio between each stable homologous region; correcting the sequencing depth of each unstable homologous region based on the corrected depth value of the sample under test in each stable homologous region and a reference depth ratio relative to each unstable homologous region in the absolute homologous region excluding the stable homologous region, wherein the reference depth ratio is determined based on a central tendency index of the sequencing depth of the reference sample in each absolute homologous region; and determining the total copy number of the homologous gene family based on the corrected depth values of the sample under test in each stable homologous region and each unstable homologous region.
[0027] In some embodiments, correcting the sequencing depth of each stable homology region includes: for each of the stable homology regions: determining a reference depth value for the current stable homology region based on the sequencing depth of the sample under test in each of the other stable homology regions excluding the current stable homology region, and the reference depth ratio between the current stable homology region and the corresponding other stable homology regions; and determining a corrected depth value for the current stable homology region based on the sequencing depth of the sample under test in the current stable homology region, the reference depth value, and a preset stable homology region correction weight.
[0028] In some embodiments, correcting the sequencing depth of each unstable homology region includes: for each of the unstable homology regions: determining a reference depth value for the current unstable homology region based on the corrected depth value of the sample to be tested in each stable homology region and the reference depth ratio between the current unstable homology region and the corresponding stable homology region; and determining a corrected depth value for the current unstable homology region based on the sequencing depth of the sample to be tested in the current unstable homology region, the reference depth value, and a preset unstable homology region correction weight.
[0029] In some embodiments, determining the total copy number of the homologous gene family based on the corrected depth values of the test sample in each stable homologous region and each unstable homologous region includes: determining the total copy number of the homologous gene family based on the ratio of the central tendency index of the corrected depth values of the test sample in each absolute homologous region to the sequencing depth of the reference sample in the corresponding absolute homologous region, and the baseline total copy number of the homologous gene family.
[0030] In some embodiments, determining the copy number of the target member gene based on the sequencing depth of each member gene supported by the test sample in the differential interval and the total copy number of the homologous gene family includes: determining the relative proportion of the copy number of the target member gene to the total copy number of the remaining member genes in the homologous gene family, excluding the target member gene, based on the sequencing depth of each member gene supported by the test sample in each differential interval; and determining the copy number of the target member gene based on the total copy number of the homologous gene family and the relative proportion of the copy number of the target member gene to the total copy number of the remaining member genes.
[0031] In some embodiments, determining the relative proportion of the target member gene copy number to the total copy number of the remaining member genes includes: determining the observed relative proportion of the target member gene copy number to the total copy number of the remaining member genes corresponding to each difference interval based on the sequencing depth supported by the sample under test in each difference interval, thereby determining the predicted relative proportion of the target member gene copy number to the total copy number of the remaining member genes; screening core difference intervals based on the observed relative proportions corresponding to each difference interval and the predicted relative proportions; and determining the relative proportion of the target member gene copy number to the total copy number of the remaining member genes based on the observed relative proportions corresponding to each core difference interval.
[0032] In some embodiments, determining the predicted relative proportion of the target member gene copy number relative to the total copy number of the remaining member genes includes: calculating a relative proportion cost function value corresponding to each hypothesis condition based on the observed relative proportion corresponding to each difference interval and multiple hypothesis conditions, wherein each of the multiple hypothesis conditions includes: a hypothetical value of the target member gene copy number determined based on the hypothetical copy number of the target member gene, a hypothetical value of the total copy number of the remaining member genes determined based on the hypothetical total copy number of the remaining member genes, and a hypothetical relative proportion of the target member gene copy number relative to the hypothetical total copy number of the remaining member genes; and selecting the hypothesis condition corresponding to the minimum relative proportion cost function value, and using the hypothetical relative proportion under this hypothesis condition as the predicted relative proportion.
[0033] In some embodiments, the sum of the hypothetical copy number of the target member gene and the hypothetical total copy number of the remaining member genes is within a preset floating range of the total copy number of the homologous gene family.
[0034] In some embodiments, calculating the relative proportion cost function value corresponding to each hypothesis includes: selecting a prior difference interval from the difference intervals based on preset screening conditions; determining the baseline observed value of the target member gene copy number corresponding to the target member gene and the baseline observed value of the total copy number of the remaining member genes corresponding to the remaining member genes based on the sequencing depth supported by the sample to be tested in the prior difference interval; and, for each of the plurality of hypotheses: determining the relative proportion cost function value corresponding to the current hypothesis based on the statistical distance between the baseline observed values corresponding to the target member gene and the remaining member genes and the corresponding hypothesis values under the current hypothesis, and the degree of deviation of the observed relative proportion relationship corresponding to each difference interval from the hypothesis relative proportion relationship under the current hypothesis.
[0035] In some embodiments, screening core difference intervals includes: selecting difference intervals whose statistical distance satisfies preset difference interval screening conditions based on the statistical distance between the observed relative proportion relationship corresponding to each difference interval and the predicted relative proportion relationship, as the core difference intervals.
[0036] In some embodiments, the screening of core difference intervals further includes: removing a predetermined number of difference intervals from the core difference intervals, so that the absolute value of the sum of the residuals of the observed relative proportions relative to the predicted relative proportions for each of the retained core difference intervals is minimized.
[0037] In some embodiments, determining the relative proportion of the target member gene copy number to the total copy number of the remaining member genes based on the observed relative proportions corresponding to each core difference interval includes: determining the relative proportion of the target member gene copy number to the total copy number of the remaining member genes based on the central tendency index of the observed relative proportions corresponding to each core difference interval.
[0038] In some embodiments, determining the copy number of the target member gene based on the total copy number of the homologous gene family and the relative proportion of the copy number of the target member gene to the total copy number of the remaining member genes includes: determining the relative abundance of the target member gene copy number relative to the total copy number of the homologous gene family based on the relative proportion of the copy number of the target member gene to the total copy number of the remaining member genes; and determining the copy number of the target member gene based on the total copy number of the homologous gene family and the relative abundance of the copy number of the target member gene.
[0039] In some embodiments, determining the copy number of the target member gene includes: determining the relative proportion of the copy number of each member gene in the homologous gene family to the total copy number of the remaining member genes in the homologous gene family, based on the sequencing depth of each member gene supported by the test sample in each differential interval; determining the relative abundance of each member gene copy number relative to the total copy number of the homologous gene family, based on the relative proportion of each member gene copy number relative to the total copy number of the remaining member genes in the homologous gene family; determining a normalization factor based on the sum of the relative abundance of the copy numbers of each member gene in the homologous gene family, so as to normalize and correct the relative abundance of the target member gene copy number; and determining the copy number of the target member gene based on the total copy number of the homologous gene family and the corrected relative abundance of the target member gene copy number.
[0040] According to a second aspect of this application, a method for determining SMN1 Genes and SMN2 A method for determining gene copy number. This method includes: identifying multiple analytical regions of a reference genome, wherein... SMN1 Genes and the SMN2 The analysis interval within the region where the gene is located is based on SMN1 Gene reference sequence and SMN2 The alignment results of the gene reference sequences correspond to each other, and the plurality of analytical regions contain at least one base sequence in the... SMN1 Gene reference sequence and the SMN2 Absolutely homologous regions that are completely identical between gene reference sequences, and at least one of the above... SMN1 Gene reference sequence and the SMN2 The presence of differentially expressed sites between gene reference sequences can represent the aforementioned SMN1 Genes and the SMN2The differential intervals of gene-specific base differences; based on the alignment results of the sequencing data of the test sample and the reference sample with the retained multiple alignment sequences of the reference genome, the sequencing depth of the test sample and the reference sample in each analysis interval is determined; based on the sequencing depth of the reference sample and the test sample in the absolute homology interval, stable homology intervals are screened to determine the gene-specific base differences; based on the sequencing depth of the reference sample and the test sample in the absolute homology interval, stable homology intervals are screened to determine the gene-specific base differences. SMN1 Genes and the SMN2 The total copy number of the gene; and, based on the test sample, supporting the statement in the difference interval. SMN1 The sequencing depth of the gene and supporting the above SMN2 The sequencing depth of the gene, and the aforementioned SMN1 Genes and the SMN2 The total copy number of the gene determines the... SMN1 Gene copy number and the above SMN2 Gene copy number.
[0041] In some embodiments, determining the SMN1 Gene copy number and the above SMN2 Before the gene copy number, it also includes: supporting the statement based on the difference intervals of the test sample. SMN1 The sequencing depth of the gene and supporting the above SMN2 The sequencing depth of the gene determines the corresponding first and second difference intervals. SMN1 Gene sequencing depth relative to the above SMN2 The relative proportion of observed gene sequencing depth, where the first difference interval is the... SMN1 Genes and the described SMN2 The differential interval within the regions containing intron 6 and exon 7 of the gene, wherein the second differential interval is the... SMN1 Genes and the described SMN2 The differential interval within the region where intron 7 of the gene is located; and, in response to the observational relative proportions corresponding to each first differential interval and the observational relative proportions corresponding to each second differential interval not satisfying a preset consistency condition, based on the test sample supporting the first differential interval... SMN1 The sequencing depth of the gene and supporting the above SMN2 The sequencing depth of the gene, and the aforementioned SMN1 Genes and the SMN2 The total copy number of the gene determines the... SMN1 Gene copy number and the above SMN2 Gene copy number.
[0042] In some embodiments, determining whether the relative proportion of observations corresponding to each first difference interval and the relative proportion of observations corresponding to each second difference interval meet a preset consistency condition includes: in response to the fact that the proportion of difference intervals in the first difference intervals that meet a first preset relative proportion threshold and the proportion of difference intervals in the second difference intervals that meet a second preset relative proportion threshold both meet a preset proportion threshold, it is determined that the preset consistency condition is not met.
[0043] According to a third aspect of this application, a computing device is also provided. The computing device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, cause the computing device to perform the steps of the methods described in the first and / or second aspects of this application.
[0044] According to a fourth aspect of this application, a computer-readable storage medium is also provided. This computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the methods described in the first and / or second aspects of this application.
[0045] According to a fifth aspect of this application, a computer program product is also provided. This computer program product includes a computer program that, when executed by a processor, implements the steps of the methods described in the first and / or second aspects of this application.
[0046] It should be understood that the summary section is provided to present the selected concepts in a simplified form, which will be further described in the detailed embodiments below. The summary section is not intended to identify key or principal features of this application, nor is it intended to limit the scope of this application. Attached Figure Description
[0047] Figure 1 A schematic diagram of a system for determining the copy number of a target member gene in a homologous gene family, according to an embodiment of this application, is shown.
[0048] Figure 2 A flowchart illustrating a method for determining the copy number of a target member gene in a homologous gene family according to an embodiment of this application is shown.
[0049] Figure 3 A flowchart of a method for screening stable homologous regions according to an embodiment of this application is shown.
[0050] Figure 4 A flowchart of a method for correcting sequencing depth of absolute homology regions according to an embodiment of this application is shown.
[0051] Figure 5 A flowchart is shown of a method for determining the relative proportion of the target member gene copy number to the total copy number of the remaining member genes, according to an embodiment of this application.
[0052] Figure 6 An embodiment of the present application is shown for determining SMN1 Genes and SMN2 A flowchart of the gene copy number method.
[0053] Figure 7 A block diagram schematically illustrates a device suitable for implementing embodiments of this application.
[0054] In the various figures, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation
[0055] Preferred embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While preferred embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments described herein. Rather, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0056] The terms “comprising” or “including” and their variations, as used herein, indicate open inclusion, meaning “including but not limited to”. Unless otherwise stated, the term “or” means “and / or”. The term “based on” means “at least partially based on”. The terms “one example embodiment” and “one embodiment” mean “at least one example embodiment”. The term “another embodiment” means “at least one additional embodiment”. The terms “first,” “second,” etc., may refer to different or the same objects. The terms “not higher than” or “below” mean “lower than or equal to”, and the terms “not lower than” or “above” mean “higher than or equal to”, and are both to be understood as including the stated number.
[0057] As described above, for homologous genes with highly similar sequences (e.g., sequence identity greater than 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%), existing homologous gene copy number detection schemes based on high-throughput sequencing typically rely on high-quality unique alignment sequences for in-depth statistical analysis. This approach suffers from problems such as difficulty in fully utilizing sequencing data from highly homologous regions, difficulty in overcoming sequencing data heterogeneity, and batch effect interference, leading to inaccurate copy number analysis results for homologous genes.
[0058] To at least partially address one or more of the aforementioned problems and other potential issues, an example embodiment of this application proposes a method for determining the copy number of each member gene in a homologous gene family. The method includes: determining multiple analysis intervals of a reference genome, wherein corresponding analysis intervals within the gene regions of each member gene in the homologous gene family are aligned with each other, and the multiple analysis intervals include at least one absolutely homologous interval where the base sequences of each member gene are completely identical, and at least one differential interval containing a base difference site that represents a unique base difference between each member gene; determining the sequencing depth of the test sample in each analysis interval based on sequencing data of the test sample; and selecting historical samples from a historical dataset containing multiple historical sample data whose correlation scores with the test sample meet preset correlation screening criteria as reference samples. Further, the method of this application, based on the sequencing depth of the reference sample and the test sample in the absolutely homologous intervals, screens stable homologous intervals to determine the total copy number of the homologous gene family; and, based on the sequencing depth of the reference sample and the test sample in the differential intervals and the total copy number of the homologous gene family, determines the copy number of each member gene in the homologous gene family.
[0059] Unlike conventional procedures that discard multiple alignment sequences (such as sequencing sequences with MAPQ=0), this application maximizes the mining and utilization of sequencing signals contributed by each base in highly homologous regions. By integrating information on absolute homologous and dissimilar regions, it significantly improves the accuracy and stability of homologous gene copy number detection under high-throughput sequencing conditions. Furthermore, this method is applicable to homologous gene families consisting of two or more paralogous genes with highly similar sequences (e.g., sequence identity greater than 99%), and is compatible with both whole-genome sequencing data and probe-capture-based high-throughput sequencing data.
[0060] Figure 1 A schematic diagram of an exemplary system 100 for implementing a method for determining the copy number of a target member gene in a homologous gene family, according to an embodiment of this application, is shown. Figure 1 As shown, system 100 includes at least a computing device 110. In some embodiments, system 100 may further include a sequencing device 130, a server 140, and a network 150. In some embodiments, the computing device 110, sequencing device 130, and server 140 may communicate and interact with each other via wired or wireless means (e.g., via network 150).
[0061] Regarding the sequencing device 130, it is used, for example, to sequence the sample to be tested and / or the reference sample to generate sequencing data about the nucleic acid sequence information of the corresponding sample. In some embodiments, the sequencing device 130 is, for example, a sequencing device based on high-throughput sequencing technology. For example, the sequencing device 130 includes, but is not limited to, sequencers based on sequencing-by-synthesis (SBS) technology developed by Illumina (e.g., including but not limited to NovaSeq, HiSeq, and MiSeq series sequencers), sequencers based on DNA nanosphere (DNB) technology developed by MGI (e.g., including but not limited to DNBSEQ series sequencers), or sequencers based on semiconductor sequencing technology developed by Thermo Fisher Scientific (e.g., including but not limited to Ion Torrent series sequencers). In some embodiments, the sequencing data can be obtained based on various library construction schemes, such as, but not limited to, whole-genome sequencing (WGS) data, whole-exome sequencing (WES) data, or targeted genome panel sequencing data constructed based on hybridization capture probes. In some embodiments, the sequencing device 130 may also send the generated sequencing data about the corresponding sample to the computing device 110. In some embodiments, the sequencing device 130 may also store the generated sequencing data about the corresponding sample in the server 140 for later use.
[0062] Regarding server 140, it may be used, for example, to store or provide to computing device 110 reference genome sequence information obtained from a public database. The reference genome may be, for example, a human reference genome, which may have different versions, including but not limited to, human reference genome GRCh37 (hg19), GRCh38 (hg38), or the T2T-CHM13 reference sequence released by the Telomere-to-Telomere Consortium. In some embodiments, server 140 may also be used, for example, to store or provide to computing device 110 annotation files that specify the start and end positions of analysis intervals. These annotation files may include, for example, BED files. In some embodiments, server 140 may also be used, for example, to store or provide to computing device 110 reference sample data, which may include, for example, sequencing depth information of samples with known copy number states in each analysis interval, to support computing device 110 in performing copy number analysis algorithms.
[0063] Regarding the computing device 110, it is used, for example, to perform the method provided in this application for determining the copy number of a target member gene in a homologous gene family. In some embodiments, the computing device 110 is used, for example, to determine a plurality of analysis intervals of a reference genome, the plurality of analysis intervals including at least one absolutely homologous interval and at least one differential interval; based on the alignment results of retained multiple alignment sequences, to determine the sequencing depth of the test sample and the reference sample in each analysis interval, and based on the sequencing depth of the reference sample and the test sample in the absolutely homologous interval, to screen stable homologous intervals to determine the total copy number of the homologous gene family; and, based on the sequencing depth of the test sample in the differential interval supporting each member gene, and the total copy number of the homologous gene family, to determine the copy number of the target member gene.
[0064] In some embodiments, the computing device 110 may have one or more processing units, including dedicated processing units such as GPUs, FPGAs, and ASICs, and general-purpose processing units such as CPUs. The computing device 110 may be a single physical device, or an integration of multiple physical servers, an integration of multiple processing units, etc. Additionally, one or more virtual machines may run on each computing device 110. The specific structure of the computing device 110 may be, for example, combined as follows: Figure 7 In some embodiments, the computing device 110 and the server 140 may be integrated together or set up separately.
[0065] In some embodiments, the computing device 110 includes, for example, an analysis interval determination unit 112, a sequencing depth determination unit 114, a total copy number determination unit 116, and a target member gene copy number determination unit 118. The analysis interval determination unit 112, sequencing depth determination unit 114, total copy number determination unit 116, and target member gene copy number determination unit 118 may be configured on one or more computing devices 110.
[0066] Regarding the analysis interval determination unit 112, it is used, for example, to determine multiple analysis intervals of a reference genome. The analysis intervals within the regions containing each member gene of the homologous gene family correspond to each other based on the alignment results of the reference sequences of each member gene, and the multiple analysis intervals include at least one absolutely homologous interval where the base sequence is completely identical among the reference sequences of each member gene, and at least one differential interval where there are differential sites among the reference sequences of each member gene that represent specific base differences of that member gene.
[0067] Regarding the sequencing depth determination unit 114, it is used, for example, to determine the sequencing depth of the test sample and the reference sample in each analysis interval based on the alignment results of the sequencing data of the test sample and the reference sample with the retained multiple alignment sequence of the reference genome.
[0068] Regarding the total copy number determination unit 116, it is used, for example, to screen stable homologous regions based on the sequencing depth of the reference sample and the test sample in the absolute homologous region, so as to determine the total copy number of the homologous gene family.
[0069] Regarding the target member gene copy number determination unit 118, it is used, for example, to determine the target member gene copy number based on the sequencing depth of each member gene supported by the test sample in the difference interval and the total copy number of the homologous gene family.
[0070] The following will combine Figure 2 This application describes a method for determining the copy number of each target member gene in a homologous gene family, according to embodiments of the present application. Figure 2 A flowchart illustrating a method 200 for determining the copy number of each target member gene in a homologous gene family according to an embodiment of this application is shown. It should be understood that method 200 can, for example, be used in... Figure 7 The described device executes at 700 points, and can also be used in Figure 1 The described computing device 110 performs the operation. It should be understood that method 200 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of this application is not limited in this respect.
[0071] At step 202, computing device 110 determines multiple analysis intervals of the reference genome for subsequent statistical analysis of the sequencing depth of the corresponding genomic regions.
[0072] Regarding the reference genome, it may be, for example, the human reference genome, which may have different versions, including but not limited to, the human reference genome GRCh37 (hg19), GRCh38 (hg38), or the T2T-CHM13 reference sequence released by the Telomere-to-Telomere Consortium. In some embodiments, depending on the actual detection target, the reference genome may also be a standard reference genome sequence of other target species. This application does not limit the specific version of the reference genome, and those skilled in the art can choose flexibly according to the actual situation.
[0073] In some embodiments, the computing device 110 determines multiple non-overlapping analysis regions in a reference genome based on a preset analysis region partitioning strategy. The preset analysis region partitioning strategy can be flexibly selected according to the specific experimental design, such as based on the exon / intron boundary structure of genes, the coverage area of hybridization capture probes, or partitioning analysis regions based on a fixed region length. In other embodiments, the computing device 110 can also directly load the defined region partitioning information from a pre-stored annotation file, such as a BED format file.
[0074] Specifically, the analysis intervals within the genomic regions of each member gene in the target homologous gene family must meet constraints based on homology alignment. Specifically, to ensure the comparability of sequencing data of different member genes in the target homologous gene family at homologous sites, the computing device 110 determines the corresponding analysis intervals within the regions of each member gene based on the multiple sequence alignment results between the reference sequences of each member gene. That is, based on the corresponding homologous positions between different member genes, a one-to-one mapping relationship is established between the analysis intervals within the regions of different member genes. Depending on the degree of sequence homology within the analysis intervals, the analysis intervals within the regions of each member gene can be further divided into absolutely homologous intervals and differential intervals.
[0075] The absolute homology region is defined as an analytical region in which the base sequences are completely identical among the reference sequences of each member gene in the homologous gene family. Within this region, since there are no base differences between the reference sequences of each member gene, alignment algorithms typically cannot distinguish sequencing reads from different member genes. Therefore, in the processing mode that preserves multiple alignment sequences, the sequencing depth of the absolute homology region can reflect the total copy number of the target homologous gene family. It should be understood that if differential sites are mixed into the absolute homology region, it may lead to incorrect correction of the copy number calculation of the absolute homology region, thereby affecting the accuracy of the copy number calculation. Therefore, the definition of the absolute homology region needs to exclude differential sites between the reference sequences of each member gene. In some preferred embodiments, the computing device 110 can also exclude absolute homology regions whose base sequences are completely identical to those of other genomic regions outside the target homologous gene family based on the uniqueness of the absolute homology region across the entire genome, in order to avoid systematic depth correction bias caused by the mixing of sequencing reads from non-target regions.
[0076] The difference interval is defined as an analytical interval where base differences exist only at difference sites between the reference sequences of each member gene. The presence of these difference sites allows the computing device 110 to specifically distinguish sequencing reads from different member genes based on specific sequence characteristics, thereby providing a basis for calculating the copy number of each member gene. In some preferred embodiments, to prevent different alignment algorithms from using different alignment rules or causing boundary slippage when processing InDels, thus affecting the accuracy of sequencing depth calculation, the computing device 110 can avoid InDel sites when determining the difference interval. Those skilled in the art will understand that the term "InDel" refers to an insertion or deletion of a nucleotide fragment relative to the reference sequence; such variations often result in gaps during alignment. Furthermore, in some population samples, characteristic bases used to distinguish one member gene may appear in another member gene. Limited by the read length of existing high-throughput sequencing technologies, it is difficult to accurately distinguish sequencing reads from different member genes based on such difference sites that lack single-member gene specificity. Therefore, in some preferred embodiments, in order to ensure that the sequencing depth of the difference interval can truly and stably reflect the relative proportion between each member gene, the computing device 110 can also avoid such difference sites that lack single member gene directionality when determining the difference interval based on the allele frequency of the difference site in the population sample.
[0077] As used herein, the term "homologous gene family" includes two or more member genes that are highly similar in sequence (e.g., sequence identity greater than 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%). In some embodiments, these member genes may be paralogous genes resulting from evolutionary events such as gene duplication, for example... SMN1 Genes and SMN2 Genes. In other embodiments, the homologous gene family may also include functional genes that perform normal physiological functions and corresponding pseudogenes that have lost their function of encoding normal proteins, such as... CYP21A2 Genes and their highly homologous CYP21A1P Pseudogenes. It should be understood that the above sequence identity can be determined by calculating the reference genome sequence of the member gene using any suitable bioinformatics sequence alignment algorithm (such as the local alignment algorithm BLAST, Smith-Waterman algorithm, or the global alignment algorithm Needleman-Wunsch, etc.).
[0078] In some embodiments, after determining multiple analysis intervals of the reference genome, the computing device 110 can also filter the analysis intervals according to preset filtering conditions to remove intervals that are not suitable for sequencing depth analysis.
[0079] For example, excessively high or low GC content in genomic regions can lead to abnormal fluctuations in sequencing depth, such as decreased PCR amplification bias and hybridization capture efficiency. Therefore, in some embodiments, the computing device 110 can eliminate analysis intervals whose GC content does not meet a preset GC content range based on the GC content of each analysis interval. It is understood that the term "GC content" refers to the proportion of guanine (G) and cytosine (C) in the total number of bases in a sequence. Not meeting the preset GC content range can, for example, mean that the GC content of the analysis interval is higher than the upper limit of the preset GC content range or lower than the lower limit of the preset GC content range. For example, in some embodiments, the upper limit of the preset GC content range can be 60%, 65%, 70%, 75%, 80%, or 85%, and the lower limit of the preset GC content range can be 40%, 35%, 30%, 25%, 20%, or 15%.
[0080] For example, the length of the analysis interval is also a key factor affecting the quality of depth statistics. An interval that is too short will result in too few sequencing sequences falling within that interval, leading to significant statistical noise; an interval that is too long may span multiple different functional units or contain regions with different copy number characteristics. Therefore, in some embodiments, the computing device 110 can also eliminate analysis intervals whose lengths do not meet a preset length range based on the interval length of each analysis interval. It is understood that the preset length range can be determined based on sequencing read length, sequencing depth level, or sequencing library type. Not meeting the preset length range can, for example, mean that the length of the analysis interval is less than the lower limit of the preset length range or greater than the upper limit of the preset length range. For example, in some embodiments, the lower limit of the preset length range can be 100 bp, 200 bp, or 500 bp, and the upper limit of the preset length range can be 2 kb, 5 kb, 10 kb, or 50 kb.
[0081] It should be understood that those skilled in the art can flexibly select filtering methods and set filtering conditions and specific parameter ranges according to the specific experimental purpose, the biological characteristics of the target gene and the characteristics of the sequencing platform, etc., and this application does not limit this.
[0082] In step 204, the computing device 110 determines the sequencing depth of the test sample and the reference sample in each analysis interval based on the alignment results of the sequencing data of the test sample and the reference sample with the preserved multiple alignment sequence of the reference genome.
[0083] The sample to be tested usually refers to the target biological sample that needs to be analyzed for copy number, the purpose of which is to determine whether there are copy number variations in the sample in a specific genomic region, especially in a highly homologous gene family region.
[0084] Regarding reference samples, they typically refer to biological samples or standard samples with known copy number states. The sequencing data of the reference samples are used as baseline data in this application to compare with the sequencing depth distribution of the test sample, thereby eliminating systematic noise generated during library construction, probe capture, and sequencing processes, and thus extracting true biological signals. In some embodiments, the computing device 110 may select samples from the same batch of experiments as the test sample as reference samples. In other embodiments, the computing device 110 may also select samples from a preset set of candidate reference samples whose sequencing depth distribution similarity to the test sample in each analysis interval meets a preset similarity condition as reference samples. The similarity can be determined, for example, by calculating the correlation coefficient between the test sample and each candidate reference sample. The correlation coefficient includes, but is not limited to, the Pearson correlation coefficient or the Spearman correlation coefficient, and the preset similarity condition may be, for example, a Pearson correlation coefficient ≥ 0.95. Reference samples that meet the preset similarity condition typically have sequencing background characteristics similar to the test sample, thus more effectively offsetting systematic errors in subsequent steps.
[0085] Regarding sample sequencing data, it includes, for example, sequencing reads of genes belonging to each member of the homologous gene family in the sample, such as, but not limited to, high-throughput sequencing data obtained via a high-throughput sequencing platform. High-throughput sequencing platforms are well-known in the art, and include, but are not limited to, sequencing platforms based on sequencing-by-synthesis technology developed by Illumina, sequencing platforms based on DNA nanosphere technology developed by BGI Genomics, and sequencing platforms based on semiconductor sequencing technology developed by Thermo Fisher Scientific. Specifically, the sample sequencing data includes, for example, whole-genome sequencing data, or more preferably, targeted high-throughput sequencing data obtained based on hybridization capture technology, such as whole-exome sequencing data or targeted genome panel sequencing data designed for specific genes. Since the hybridization capture process is greatly affected by probe hybridization efficiency, GC content of the target region, and sequence homology, it is prone to capture heterogeneity. The method of this application, through subsequent correction and screening processing, can effectively solve the analytical challenges caused by data fluctuations in such targeted sequencing data.
[0086] Before performing in-depth statistical analysis, raw sequencing data typically undergoes preliminary quality control and filtering processes, including but not limited to removing low-quality sequencing sequences, correcting sequencing errors, and removing redundant sequences. Those skilled in the art should understand that the aforementioned quality control and filtering can be implemented using any suitable quality control and filtering tools, and the specific processes and parameters can be flexibly selected based on the sequencing platform, sample characteristics, and experimental objectives.
[0087] A method for determining the sequencing depth of the test sample and the reference sample in each analysis interval includes, for example: a computing device 110 determining the original sequencing depth of the test sample and the reference sample in each analysis interval based on the alignment results of the sequencing data of the test sample and the reference sample with the retained multiple alignment sequence of the reference genome; and performing depth correction on the original sequencing depth to determine the sequencing depth of the test sample and the reference sample in each analysis interval.
[0088] As used herein, the term "sequencing depth" refers to the number of sequencing reads covering a given genomic region. Those skilled in the art will understand that, under ideal high-throughput sequencing conditions free from systematic bias, the sequencing depth of a specific genomic region is proportional to the copy number of the DNA molecules in that region.
[0089] Those skilled in the art will understand that the term "multiple alignment sequence" refers to sequencing reads that, during alignment, can be mapped to multiple different physical locations on the reference genome with similar alignment scores. Because the true origin of such sequencing reads cannot be distinguished, conventional alignment algorithms typically assign them extremely low alignment quality values (usually 0), and often filter out sequencing reads with alignment quality values below a certain threshold (e.g., MAPQ < 20) during subsequent quality control processes to ensure the specificity of the analysis. However, for the homologous gene family of interest in this application, due to the extremely high sequence consistency among member genes, a large number of sequencing reads originating from this region exhibit multiple alignment characteristics. If conventional filtering strategies are used, the sequencing depth data within the homologous region will be artificially zeroed out, resulting in a significant loss of sequencing data from the homologous region.
[0090] Therefore, in order to make full use of sequencing data in homologous regions that are usually unavailable under conventional alignment quality control conditions, the method in this application retains multiple alignment sequences during sequence alignment and includes all sequencing reads that can be mapped to homologous gene family regions in the sequencing depth statistics of the corresponding analysis interval, providing complete data support for subsequent homologous gene family copy number analysis.
[0091] Specifically, computing device 110 can use sequence alignment tools, in a processing mode that preserves multiple alignment sequences (e.g., setting the MAPQ filtering threshold to 0), to map the sequencing sequence of the sample onto a reference genome, and for each analysis interval determined in step 202. The number of sequencing reads falling within the analysis interval is counted using depth statistics tools to determine the original sequencing depth of the sample within that interval. In some embodiments, the sequence alignment tool may be any sequence alignment tool suitable for high-throughput sequencing data, such as, but not limited to, BWA-MEM, Bowtie, or STAR; the depth statistics tool may be any depth statistics tool suitable for high-throughput sequencing data, such as, but not limited to, samtools, mosdepth, or bedtools. The specific implementation of sequence alignment and depth statistics does not constitute a limitation of this application.
[0092] In some embodiments, when determining the original sequencing depth Subsequently, computing device 110 can analyze the original sequencing depth. Depth correction is performed to reduce data fluctuations caused by sequencing bias and differences in sequencing data volume between samples, thereby improving the comparability of sequencing depths between different analysis intervals and between different samples. In some embodiments, the depth correction includes, but is not limited to, GC correction based on GC content and data volume normalization correction based on the total amount of sequencing data. It should be noted that this application does not limit the depth correction to a specific method or statistical model. Those skilled in the art can flexibly choose any correction scheme that can eliminate systematic bias and enhance the comparability of depth data based on the noise characteristics of the actual sequencing platform.
[0093] Regarding GC correction, its purpose is to eliminate or reduce sequencing data fluctuations caused by differences in GC content among DNA fragments during PCR amplification or hybridization capture. The computing device 110 can employ various mathematical models to implement GC correction. For example, in some embodiments, the computing device 110 can use a Loess Regression model to fit the relationship curve between depth and GC content. In other embodiments, the computing device 110 can use the depth distribution of autosomal background analysis intervals with the same or similar GC content within the sample itself as a reference, and perform GC correction on each analysis interval. raw sequencing depth Perform correction to obtain the GC-corrected depth value. .
[0094] Specifically, the computing device 110 can pre-divide the autosomal background analysis intervals across the entire genome into multiple different GC content groups based on the GC content of each analysis interval, according to a preset GC content span (e.g., a GC content span of 1%, 2%, or 5% as a grouping interval). Then, for the analysis intervals to be corrected... The computing device 110 can refer to the following formula for GC correction: in, For analysis interval The original sequencing depth; This is the median sequencing depth of the background analysis region for all autosomes in this sample. For analysis interval Belongs to a specific GC content group; The GC content belongs to the specific GC content group mentioned above. The median sequencing depth of all autosomal background analysis intervals within the range.
[0095] By using the above methods, the sequencing depth affected by GC bias can be brought back to the baseline level of the whole genome, thereby eliminating or reducing the impact of GC bias on subsequent analysis.
[0096] In some embodiments, the computing device 110 may further perform data volume normalization correction based on the total amount of sequencing data after completing GC correction, in order to eliminate the incomparability between different samples caused by differences in the total amount of sequencing data. Data volume normalization correction ensures that the sequencing depth of different samples is on the same order of magnitude.
[0097] Specifically, computing device 110 can refer to the following formula for data volume standardization correction: in, This represents the sequencing depth after GC correction and data volume normalization correction. This is the sum of GC-corrected sequencing depths of the autosomal background analysis regions used for standardization in the current sample (i.e., the total amount of GC-corrected sequencing data). This is an arbitrary preset standardization constant, which can usually be set according to the expected average sequencing depth of the experimental design.
[0098] It should be understood that the term "autosomal background analysis region" refers to a normal diploid region in the genome that is known to not contain copy number variations and does not belong to the target homologous gene family.
[0099] After the above correction process, the computing device 110 finally determined the sequencing depth of the test sample and the reference sample in each analysis interval. This provides normalized input parameters for subsequent copy number analysis.
[0100] Since absolute homologous regions have completely identical base sequences among the reference sequences of each member gene in a homologous gene family, conventional alignment algorithms typically cannot distinguish sequencing reads from different member genes. Therefore, in the multiple alignment sequence preservation processing mode of this application, the sequencing depth of absolute homologous regions can reflect the total copy number of the target homologous gene family in the sample being tested.
[0101] Based on this, in step 206, the computing device 110 screens stable homologous regions based on the sequencing depth of the absolute homologous regions of the reference sample and the test sample to determine the total copy number of the homologous gene family.
[0102] Those skilled in the art should understand that in actual high-throughput sequencing processes, uneven random fluctuations still exist between different analytical intervals due to factors such as differences in probe hybridization efficiency and batch effects. Therefore, in order to eliminate these interfering factors, the method of this application further screens out a set of statistically robust stable homology intervals (denoted as...). This serves as the baseline interval for subsequent calculations of the total copy number.
[0103] The method for screening stable homologous regions includes, for example, the following: A computing device 110 determines the degree of fluctuation in the sequencing depth of each absolute homologous region based on the dispersion index of the sequencing depth of the reference sample in each absolute homologous region; selects absolute homologous regions whose fluctuation satisfies a first preset stability screening condition as first stable homologous regions; determines a reference depth value for each first stable homologous region based on the sequencing depth of the reference sample and the test sample in each first stable homologous region, calculates the deviation cost function value corresponding to each first stable homologous region, selects first stable homologous regions whose deviation cost function value satisfies a second preset stability screening condition as second stable homologous regions; and, based on the sequencing depth of the test sample in each second stable homologous region, determines the predicted total copy number of the homologous gene family, calculates the fitting cost function value corresponding to each second stable homologous region, selects second stable homologous regions whose fitting cost function value satisfies a third preset stability screening condition as the stable homologous regions. The following will combine... Figure 3 The specific method for screening stable homologous intervals is explained in detail in section 300, and will not be repeated here.
[0104] In some implementations, stable homology regions are determined for total copy number analysis. Subsequently, the computing device 110 can determine the total copy number of the target homologous gene family based on the sequencing depth of the reference sample and the sample to be tested in these stable homologous regions. For example, computing device 110 can calculate the total copy number of the target homologous gene family in the sample to be tested using the following formula. : in, For the sample to be tested at the 1st Sequencing depth of stable homologous regions; For reference sample in the first The central tendency index of sequencing depth of a stable homologous region reflects the baseline depth of the region under normal diploid (or known copy number state) conditions. It should be understood that the term "central tendency index" refers to a statistic that can reflect the central position or representativeness of a set of data distributions. It can be, for example, the median, the mean, or the truncated mean. Those skilled in the art can flexibly select appropriate statistics as central tendency indexes according to the characteristics of data distribution. This represents the baseline total copy number of the target homologous gene family under normal conditions (or the known baseline total copy number in the reference sample), for example, for SMN Homologous gene families, typically found in healthy diploid individuals, contain two SMN1 Copy and 2 SMN2 Copy, therefore its It is usually set to 4; Indicates based on the first Sequencing depth data of stable homologous regions were used to estimate the total copy number of the target homologous gene family in the sample to be tested. This represents the selected central tendency operator used to summarize the calculation results of all stable, homogeneous intervals. These include, but are not limited to, central tendency indicators such as the median, mean, and truncated mean. Furthermore, in the calculation... The operator used at that time and They can be the same or different, and those skilled in the art can choose flexibly based on the data characteristics.
[0105] In some preferred embodiments, in order to further approximate the true copy number state of the target homologous gene family in the test sample, stable homologous regions for total copy number analysis are determined. Then, and in calculating the total number of copies Previously, the method in this application also included methods based on the reference sample and the test sample within a stable homology region. The sequencing depth of the sample to be tested is internally corrected for each absolute homology region to eliminate local random noise. For example, the computing device 110 can use the reference depth ratio between each stable homology region and between each stable homology region and each unstable homology region other than the stable homology region in the absolute homology region to perform weighted smoothing correction on the sequencing depth of the sample to be tested in each absolute homology region, thereby reducing the influence of local random noise, and determining the total copy number of the target homology gene family based on the corrected depth of the sample to be tested in each absolute homology region. The internal correction method 400 for sequencing depth of absolute homology regions will be discussed in the following section. Figure 4 Specific details will not be elaborated here.
[0106] Unlike absolute homologous regions, differentially expressed regions contain specific base differences between members of a homologous gene family, allowing sequencing reads to be distinguished by their origin after alignment. Therefore, the sequencing depth of differentially expressed regions can reflect the relative copy number relationships between different gene members in the sample being tested.
[0107] Based on this, in step 208, the computing device 110 supports the sequencing depth of each member gene in the difference interval based on the sample to be tested, and the total copy number of the homologous gene family determined in step 206. Determine the copy number of the target member gene. .
[0108] Specifically, the method by which computing device 110 determines the copy number of the target member gene includes, for example, the following: computing device 110 first determines the copy number of the target member gene based on the difference between the reference sample and the test sample in each interval (denoted as...). Supports sequencing depth of target member genes (denoted as ) ) and the total sequencing depth (denoted as ) of all member genes in the target homologous gene family other than the target member gene. The relative proportion of the target member gene copy number to the total copy number of all other member genes in the target homologous gene family (denoted as ) is determined. Subsequently, based on the total copy number of homologous gene families determined in step 206... And the relative proportion of the target member gene copy number to the total copy number of the other member genes. Determine the copy number of the target member's gene. .
[0109] Regarding the relative proportion of the target member gene copy number to the total copy number of all other member genes in the target homologous gene family, excluding the target member gene. The method, for example, includes: a computing device 110 determining, based on the sequencing depth of each member gene supported in each differential interval of the sample to be tested, the observed relative proportion of the target member gene copy number to the total copy number of the remaining member genes in each differential interval, to determine the predicted relative proportion of the target member gene copy number to the total copy number of the remaining member genes; screening core differential intervals based on the observed relative proportions corresponding to each differential interval and the predicted relative proportions; and determining the relative proportion of the target member gene copy number to the total copy number of the remaining member genes based on the observed relative proportions corresponding to each core differential interval. The following will combine... Figure 5 The specific method for determining the relative proportion of the target member gene copy number to the total copy number of the remaining member genes (500) will not be elaborated here.
[0110] To obtain the relative proportion of the target member gene copy number to the total copy number of the other member genes. Then, the computing device 110 further calculates the relative proportion. Copy number relative abundance, converted to the copy number of the target member gene relative to the total copy number of the target homologous gene family. And combined with the total copy number of homologous gene families determined in step 206 Determine the copy number of the target member's gene. .
[0111] Specifically, the computing device 110 obtains the relative proportion of the target member gene copy number to the total copy number of the remaining member genes. Then, the relative proportions can be expressed by referring to the following mathematical expression. Conversion to copy number relative abundance of target member gene copy number relative to the total copy number of the target homologous gene family : Subsequently, the computing device 110 can combine the total copy number of homologous gene families determined in step 206. The copy number of the target member gene can be determined by referring to the following mathematical expression. : Those skilled in the art will understand that, since the relative proportions of each member gene are determined separately and independently, the sum of the relative abundances of each member gene obtained directly from the above equation (1) may not be strictly equal to 1 due to the influence of sequencing noise or fluctuations in local difference intervals. Therefore, in order to ensure the consistency between the sum of the copy numbers of each member gene in the target homologous gene family and the total copy number of the homologous gene family, in some preferred embodiments, the computing device 110 may also determine a normalization factor based on the sum of the relative abundances of the copy numbers of each member gene in the target homologous gene family, normalize and correct the relative abundances of the copy numbers of the target member gene using the normalization factor, and determine the copy number of the target member gene based on the total copy number of the homologous gene family and the corrected relative abundances of the copy numbers of the target member gene.
[0112] Specifically, computing device 110 targets each member gene in the target homologous gene family. Perform the above method 500 respectively to determine the member genes. Copy number relative to the target homologous gene family excluding member genes The relative proportions of the total copy number of genes of the remaining members. And based on the above equation (1), the member genes are obtained. Copy number relative abundance of the target homologous gene family Based on this, the computing device 110 can determine the normalization factor based on the sum of the relative abundance of the copy numbers of each member gene in the target homologous gene family. : Furthermore, using the aforementioned normalization factor The relative copy number abundance of the target member gene is normalized and corrected to obtain the corrected relative copy number abundance of the target member gene. : Subsequently, the computing device 110 can combine the total copy number of homologous gene families determined in step 206. The copy number of the target member gene can be determined by referring to the following mathematical expression. : The above correction process ensures that the sum of the corrected copy number relative abundance of all member genes is 1, thus eliminating cumulative errors.
[0113] Through the above steps, this application, under the premise of having determined the total copy number of the homologous gene family, utilizes the member gene differentiation ability provided by the difference interval to first determine the relative proportion of the target member gene copy number to the total copy number of the other member genes. Then, under the constraint of the determined total copy number, a unified conversion and normalization correction are performed, which can robustly and accurately determine the target member gene copy number in the homologous gene family.
[0114] In the above-described scheme, this application achieves accurate and robust analysis of copy numbers in homologous gene families, and is particularly suitable for scenarios involving highly homologous genes, multi-copy genes, and complex copy number variations. This scheme can effectively overcome the judgment bias caused by high sequence similarity, accumulated sequencing noise, or unstable reference models in traditional methods.
[0115] The following will combine Figure 3 This application describes a method for screening stable homologous regions according to embodiments of the present application. Figure 3 A flowchart of a method 300 for screening stable homologous regions according to an embodiment of this application is shown. It should be understood that method 300 can, for example, be implemented in... Figure 7 The described device executes at 700 points, and can also be used in Figure 1 The described computing device 110 performs the operation. It should be understood that method 300 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of this application is not limited in this respect.
[0116] To identify analytical intervals within absolute homology intervals that stably reflect the total copy number of homologous gene families, this application provides a multi-round screening method for determining stable homology intervals. This method comprehensively considers the historical stability of reference samples, the depth consistency of test samples, and the degree of fit under the total copy number assumption, thereby gradually eliminating unstable intervals and obtaining a set of stable homology intervals for total copy number calculation.
[0117] Those skilled in the art will understand that due to differences in target probe capture efficiency, local chromatin structure, and regional fluctuations in GC content during high-throughput sequencing, there are inherent differences in the sequencing depth of different absolute homologous regions. Therefore, before screening stable homologous regions, in order to avoid statistical bias when evaluating the stability of absolute homologous regions at different depth levels on the same scale (e.g., misjudging normal depth fluctuations in low capture efficiency regions as unstable regions), thereby affecting the accuracy of subsequent copy number analysis, optionally, in some embodiments, the method of this application further includes step 301 for performing depth-level stratification processing on the absolute homologous regions.
[0118] At step 301, computing device 110 determines the global baseline depth of the absolute homology intervals based on the central tendency index of the sequencing depth of the reference sample in each absolute homology interval; and, based on the global baseline depth, divides the absolute homology intervals into interval subsets of different depth levels, thereby independently performing the subsequent screening steps for stable homology intervals in each interval subset.
[0119] Specifically, for example, firstly, computing device 110 determines the global baseline depth of the absolute homology intervals based on a central tendency index of the sequencing depth of the reference sample in each absolute homology interval. Let... A set of absolutely homologous intervals. For a set of reference samples, for each reference sample In every absolutely homologous region The corrected sequencing depth is denoted as Then the global reference depth It can be represented as: in, This represents the selected operator of central tendency. For example, if the average is used as the indicator of central tendency, then... The arithmetic mean of the depths of all reference samples across all absolutely homogeneous intervals is used. Preferably, in order to eliminate the interference of outlier extreme values on the baseline, the truncated mean can also be used as an indicator of central tendency. For example, the average value can be calculated after removing the highest 10% and lowest 10% of the depth observations in the reference sample set for each absolutely homogeneous interval, thereby ensuring the robustness of the baseline depth.
[0120] Determine the global baseline depth Then, the computing device 110 bases its calculations on the global reference depth. The absolutely homologous intervals are divided into interval subsets of different depth levels. For example, computing device 110 first calculates each absolutely homologous interval. In the reference sample set Central tendency indicator in intervals : Then, the computing device 110 will calculate the intervals. Compared with global baseline depth The comparison is performed, and based on the comparison results, the absolutely homologous regions are divided into at least two depth levels. For example, using a global reference depth... Using the threshold value, the central tendency index of the interval is set. Not less than absolute homology regions Included in high-depth interval set and will Less than absolute homology regions Included in low-depth interval set .
[0121] By independently screening stable homologous intervals in subsets at different depth levels, the method in this application can ensure that subsequent stability evaluations are conducted between intervals with similar statistical distribution characteristics, effectively avoiding the risk of misjudgment due to inherent depth differences between intervals and improving the accuracy of stable homologous interval screening.
[0122] After completing the above-mentioned deep stratification, the computing device 110 independently performs the stable homogeneous interval screening step in each interval subset.
[0123] In step 302, the computing device 110 determines the degree of fluctuation in the sequencing depth of each absolute homology interval based on the dispersion index of the sequencing depth of the reference sample in each absolute homology interval, and selects the absolute homology interval whose fluctuation degree meets the first preset stability screening condition as the first stable homology interval.
[0124] For example, the computing device 110 determines the degree of fluctuation in the sequencing depth of the current absolute homology region based on the dispersion index of the sequencing depth of the reference sample in the current absolute homology region and the deviation of the central tendency index of the sequencing depth of the test sample in the current absolute homology region from the sequencing depth of the reference sample in the current absolute homology region.
[0125] Specifically, computing device 110 calculates the historical coefficient of variation of the reference sample within this interval. and the adaptive weights of the current test sample Adaptive weights The formula for characterizing the deviation of the sequencing depth of the sample in the current absolute homology region from the historical mean is as follows: in, For the sample to be tested in the interval Depth of standardization The historical average depth of the reference sample in this interval.
[0126] Based on this, the degree of fluctuation in this absolutely homogeneous interval is determined. : Among them, the index As a preset parameter, it can be 2, 3 or other suitable values in some embodiments, and those skilled in the art can choose flexibly according to the actual situation.
[0127] Based on the aforementioned volatility levels, intervals with high volatility can be eliminated. For example, a portion of the absolutely homologous intervals with the lowest volatility (e.g., the first third) can be retained as the first stable homologous intervals, denoted as... .
[0128] In step 304, the computing device 110 determines the reference depth value for each first stable homologous region based on the sequencing depth of the reference sample and the test sample in each first stable homologous region, to calculate the deviation cost function value corresponding to each first stable homologous region, and selects the first stable homologous region whose deviation cost function value meets the second preset stability screening condition as the second stable homologous region. This step further evaluates the internal consistency among the first stable homologous regions, aiming to eliminate regions inconsistent with the population trend. The core idea is that if a certain region truly reflects the total copy number change of the homologous gene family, then the sequencing depth of that region should be reasonably predicted by other stable regions through the depth ratio relationship in the reference sample. Therefore, a cross-prediction method can be used to calculate the deviation cost function value for each candidate region.
[0129] Specifically, the computing device 110 uses cross-prediction to calculate the deviation cost function value for each interval. That is, using the set excluding the interval Other intervals Predicting intervals by combining historical depth ratios The depth is calculated, and a weighted sum of the prediction errors is calculated: in, Representative of the interval The interval derived from the observed values The theoretical depth value.
[0130] Subsequently, based on the deviation cost function value, intervals with large deviations are eliminated, and intervals with high consistency (e.g., the top 80%) are retained as the second stable homogeneous intervals, denoted as... .
[0131] In step 306, the computing device 110 determines the observed value corresponding to each second stable homologous region based on the sequencing depth of the sample to be tested in each second stable homologous region, so as to determine the predicted total copy number of the homologous gene family, and calculates the fitting cost function value corresponding to each second stable homologous region. The second stable homologous region whose fitting cost function value satisfies the third preset stability screening condition is selected as the stable homologous region.
[0132] This step, based on the second stable homology interval, determines multiple total copy number hypotheses and evaluates the fit between each hypothesis and the observed values, thereby determining the predicted total copy number. This prediction result is primarily used for subsequent refinement and selection of stable intervals, rather than being directly output as the final total copy number.
[0133] Specifically, computing device 110 sets a total copy number assumption set. For each hypothesis of total copy number k, calculate the total copy number cost function value. : in, This represents the baseline copy number for normal diploids within the homologous gene family.
[0134] Computing device 110 selects to make Minimal assumption The predicted total copy number of homologous gene families .
[0135] After obtaining the predicted total copy number, the computing device 110 uses the predicted total copy number to evaluate the degree of fit of each second stable homologous interval under the prediction result, and further selects the intervals that are most consistent with the prediction result as the core stable homologous intervals.
[0136] Specifically, computing device 110 calculates each interval based on the predicted total copy number. Fitting cost function value : choose The smallest (i.e., the closest to the predicted integer) portion of the interval (e.g., the top 25%) is used as a stable homogeneous interval. The remaining absolutely homologous regions are non-core stable homologous regions. ).
[0137] Through the above-mentioned multi-round screening and correction mechanism, this application can reliably identify stable homologous regions from absolute homologous regions in the context of complex sequencing noise, effectively avoiding the amplified impact of single region anomalies on the determination of total copy number, thereby significantly improving the robustness and accuracy of total copy number analysis of homologous gene families.
[0138] The following will combine Figure 4 A method for correcting sequencing depth for absolute homologous regions according to embodiments of this application is described. Figure 4 A flowchart of a method 400 for correcting sequencing depth of absolute homology regions according to an embodiment of this application is shown. It should be understood that method 400 can, for example, be implemented in... Figure 7 The described device executes at 700 points, and can also be used in Figure 1 The described computing device 110 performs the operation. It should be understood that method 400 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of this application is not limited in this respect.
[0139] Specifically, at step 402, the computing device 110 first determines the reference depth ratio between each absolute homology interval based on the central tendency index of the sequencing depth of the reference sample in each absolute homology interval. Let... Let be the set of all absolutely homologous intervals. For any two intervals in this set... and Its reference depth ratio relationship Characterizes the interval under the reference state and interval The depth distribution ratio between them can be expressed as: in, and The reference sample set is in the interval and interval A measure of central tendency for sequencing depth. Specifically, this can be a statistical measure such as the median, mean, or truncated mean.
[0140] At step 404, the computing device 110 corrects the sequencing depth of each stable homology region based on the sequencing depth of the sample to be tested in each stable homology region and the reference depth ratio between each stable homology region.
[0141] For each absolute homology region, a reference depth ratio between this region and stable homology regions is determined based on reference sample data. This ratio is then combined with the sequencing depth of the sample to be tested within stable homology regions to predict the sequencing depth of the current region. Subsequently, a weighted approach can be used to fuse the original sequencing depth and the predicted depth of the current region to obtain a corrected sequencing depth value. For regions with high stability, such as the stable homology regions, a higher weight can be assigned to the original sequencing depth; for regions with relatively low stability, such as unstable homology regions other than the stable homology regions within the absolute homology regions, more reliance can be placed on the predicted depth.
[0142] For example, in some embodiments, for each of the stable homology intervals: a reference depth value for the current stable homology interval can be determined based on the sequencing depth of the sample under test in each of the other stable homology intervals besides the current stable homology interval, and the reference depth ratio between the current stable homology interval and the corresponding other stable homology intervals; and a corrected depth value for the current stable homology interval can be determined based on the sequencing depth of the sample under test in the current stable homology interval, the reference depth value, and a preset stable homology interval correction weight.
[0143] Specifically, for stable homogeneous regions First, computing device 110 calculates the region based on the sequencing depth of other stable homologous regions. Predicted depth : in, This represents the set of stable, homogeneous intervals obtained through screening. Indicates that the reference sample is in the interval The historical average sequencing depth.
[0144] Then, the corrected sequencing depth is determined by weighting. : in, The preset correction weights for stable homologous regions are used to balance the contributions of the original sequencing depth and the predicted depth. Those skilled in the art can flexibly determine their specific values based on the data characteristics.
[0145] In step 406, the computing device 110 corrects the sequencing depth of each unstable homology interval based on the corrected depth values of the sample to be tested in each stable homology interval and the reference depth ratio of each stable homology interval relative to each unstable homology interval other than the stable homology interval in the absolute homology interval.
[0146] For example, in some embodiments, for each of the unstable homology regions: a reference depth value for the current unstable homology region can be determined based on the corrected depth value of the sample under test in each stable homology region and the reference depth ratio between the current unstable homology region and the corresponding stable homology region; and a corrected depth value for the current unstable homology region can be determined based on the sequencing depth of the sample under test in the current unstable homology region, the reference depth value, and a preset unstable homology region correction weight.
[0147] Specifically, for non-stable homogeneous regions First, computing device 110 is based on other stable homogeneous regions. Corrected sequencing depth Calculate the non-stationary homogeneous region Predicted depth : Then, the corrected sequencing depth is determined by weighting. : in, For the preset correction weights targeting unstable homogeneous regions, those skilled in the art can flexibly determine their specific values based on data characteristics.
[0148] After completing the internal depth correction, at step 408, the computing device 110 can determine the total copy number of the homologous gene family based on the corrected depth values of each stable homologous region and each unstable homologous region.
[0149] Specifically, the computing device 110 determines the total copy number of the homologous gene family based on the ratio of the central tendency index of the corrected sequencing depth of the test sample in each absolute homology region to the sequencing depth of the reference sample in the corresponding absolute homology region, and the baseline total copy number of the homologous gene family. : in, This represents the baseline copy number of a homologous gene family under normal conditions. For interval The corresponding total copy number estimate. In some embodiments, robust statistics such as the median or truncated mean of the total copy number estimates for all absolute homology intervals can also be used to determine the final total copy number. Those skilled in the art can choose flexibly according to the data characteristics.
[0150] Through the above steps, this application can achieve a stable estimate of the total copy number of homologous gene families based on absolute homology intervals; and through mechanisms such as deep stratification and internal correction, it effectively reduces the impact of sequencing noise and systematic bias, obtaining highly reliable total copy number analysis results. The following will combine... Figure 5 This application describes a method for determining the relative proportion of the copy number of a target member gene to the total copy number of the remaining member genes, according to embodiments of the present application. Figure 5 A flowchart of a method 500 for determining the relative proportion of a target member gene copy number to the total copy number of the remaining member genes, according to an embodiment of this application, is shown. It should be understood that method 500 can, for example, be used in... Figure 7 The described device executes at 700 points, and can also be used in Figure 1The described computing device 110 performs the operation. It should be understood that method 500 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of this application is not limited in this respect.
[0151] After determining the total copy number of a homologous gene family, this application provides a method for determining the relative proportions based on difference intervals to further distinguish the copy number distribution among the member genes in the homologous gene family. This method obtains stable and reliable relative proportions of member genes by modeling the relative relationship of sequencing depth within the difference intervals and evaluating the costs under various assumptions.
[0152] In step 502, the computing device 110, based on the sequencing depth of each member gene supported by the reference sample and the test sample in each difference interval, determines the observed relative proportion of the target member gene copy number to the total copy number of the remaining member genes in each difference interval, so as to determine the predicted relative proportion of the target member gene copy number to the total copy number of the remaining member genes.
[0153] Specifically, for the first Let there be intervals of difference: This represents the sequencing depth of the sample under test within this difference interval, aligned to the current member gene; let... The sequencing depth of the sample under test within this difference interval, compared to other member genes, can then be defined as the observed relative proportion. : Specifically, the computing device 110 first determines the anchor value based on the prior difference interval in the difference interval (e.g., key difference sites determined based on clinical phenotype, or pre-screened high stability sites).
[0154] in, and These are the baseline observed copy numbers for the target gene and the background gene set, respectively. This represents the depth of the current member gene at the key site in the sample to be tested. This is the average depth of the current gene at key sites in the reference sample. This represents the depth of the remaining member genes in the sample at key sites. This is the average depth of the remaining genes at key sites in the reference sample. This is the normal baseline copy number.
[0155] To tolerate minor errors that may exist in the total copy number calculation, the computing device 110 constructs a candidate total copy number set (e.g., based on a preset fluctuation range for the total copy number of the homologous gene family). For each candidate total in the set, generate all possible combinations of hypothesis copy numbers.
[0156] Each hypothesis includes: the hypothesis copy number of the current member gene ( ), the sum of the hypothetical copy numbers of the remaining member genes ( ) and the hypothetical relative proportions of the current member genes to the remaining member genes ( ).
[0157] For each hypothesis, calculate the relative proportional cost function value ( This function integrates the deviation of the observed relative proportions in each difference interval from the assumed relative proportions under that assumption; and the difference between the assumed copy number under that assumption and the aforementioned determined baseline observed copy number. in, It is the set of all intervals of difference. This is a proportional consistency weighting term, used to amplify the inconsistency between the observed ratio and the hypothetical ratio: in, For the first The relative proportions of observations in each difference interval (i.e.) ).
[0158] choose The minimum assumptions include the relative proportions of the assumptions, which are used to predict the relative proportions. ).
[0159] In step 504, in order to eliminate interference from individual differential sites (e.g., due to probe nonspecific binding or local sequencing noise), computing device 110 further filters the core differential intervals among the differential intervals based on the observed relative proportion relationship corresponding to each differential interval and the predicted relative proportion relationship.
[0160] In step 506, the computing device 110 determines the relative proportion of the target member gene copy number to the total copy number of the remaining member genes based on the observed relative proportions corresponding to each core difference interval. The aim is to screen out the core difference intervals that best represent the true biological signal from all difference intervals.
[0161] Specifically, computing device 110 first calculates each difference interval. Observed relative proportions and predicted relative proportions The absolute deviation between ( ): in, For the first The relative proportions of observations in each difference interval (i.e.) ), The predicted relative proportion is determined based on the minimum cost function.
[0162] Calculation device 110 based on absolute deviation Sort all the discrepancy intervals and remove a portion of the intervals with the largest deviations. For example, discard the points with the highest deviations, representing the top 25%. The remaining intervals constitute the initial set of discrepancy intervals. .
[0163] In the initial screening set Based on this, we further search for the subset that best matches the overall trend of the predicted relative proportion relationship. Computing device 110 from... Select a predetermined number of intervals (e.g., half the total number of sites) to form a subset, such that the absolute value of the sum of the residuals between the observed relative proportions and the predicted relative proportions of each point in the subset is minimized.
[0164] This subset is the set of core difference intervals. The filtering logic is as follows: in, The number of core intervals (e.g.) Half the size).
[0165] Finally, computing device 110 is based on the core difference interval set. The relative proportions of observations across all intervals are analyzed, and their statistical characteristic values (such as the arithmetic mean) are calculated as the final target-background ratio for the target member gene. ): Through the above steps, this application can stably and accurately determine the relative proportions between members of the homologous gene family under the constraint of the total copy number, effectively filter out random noise in the sequencing data, and ensure that the obtained relative proportions are based on sequencing data from high-quality, highly consistent analysis intervals, thus significantly improving the accuracy of the final copy number determination.
[0166] Spinal muscular atrophy (SMA) is a severe autosomal recessive neuromuscular disease. Two genes in the human genome are associated with the development of SMA: one on the telomere side and one on the telomere side. SMN1 Genes and centromere-side SMN2 The two genes are highly homologous (over 99.9% identity), differing only in a very small number of base sites. This makes it difficult for conventional high-throughput sequencing analysis methods to accurately determine their identities. SMN1 and SMN2 The copy number of the gene. To at least partially address the above problems, this application further provides a method for determining... SMN1 Genes and SMN2 Methods for counting gene copy numbers.
[0167] The following will combine Figure 6 Description of embodiments of the present application for determining SMN1 Genes and SMN2 Methods for counting gene copy numbers. Figure 6 An embodiment of the present application is shown for determining SMN1 Genes and SMN2 A flowchart of method 600 for gene copy number determination. It should be understood that method 600, for example, can be used in... Figure 7 The described device executes at 700 points, and can also be used in Figure 1 The described computing device 110 performs the operation. It should be understood that method 600 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of this application is not limited in this respect.
[0168] At steps 602, 604, and 606, computing device 110 first executes steps 202, 204, and 206 of the aforementioned method 200. Specifically, at step 602, computing device 110 based on SMN1 Gene reference sequence and SMN2 The alignment results of the gene reference sequence, in SMN1 Genes and SMN2 Within the region where the gene is located, corresponding analysis intervals are defined, each analysis interval containing at least one base sequence. SMN1 Gene reference sequence and SMN2 Absolutely homologous regions that are completely identical between gene reference sequences, and at least one in SMN1 Gene reference sequence and SMN2 The presence of differentially expressed sites between gene reference sequences can represent the aforementioned SMN1 Genes and the SMN2The differential intervals of gene-specific base differences; in step 604, the computing device 110 determines the sequencing depth of the test sample and the reference sample in each analysis interval based on the alignment results of the retained multiple alignment sequences; subsequently, in step 606, the computing device 110 screens stable homologous intervals based on the sequencing depth of the reference sample and the test sample in the absolutely homologous intervals to determine the genetic region in the test sample. SMN1 Genes and SMN2 Total number of gene copies.
[0169] In some embodiments, determine the sample to be tested. SMN1 Genes and SMN2 After determining the total copy number of the genes, the computing device 110 can directly proceed to step 614, executing step 208 of the aforementioned method 200. Specifically, the computing device 110 determines the copy number of the SMN1 gene and the copy number of the SMN2 gene based on the sequencing depth supporting the SMN1 gene and the sequencing depth supporting the SMN2 gene in all differential regions of the sample to be tested, and the total copy number of the SMN1 and SMN2 genes in the sample to be tested determined at step 606. This embodiment fully utilizes the sequencing signal contributed by each base in the highly homologous regions of the SMN gene, and can provide robust copy number quantification results in the vast majority of samples without complex structural recombination.
[0170] However, those skilled in the art should understand that the pathogenesis of SMA is mainly caused by the deletion or mutation of key functional regions of the SMN1 gene (especially exon 7). Furthermore, non-allelic homologous recombination frequently occurs between the SMN1 and SMN2 genes, leading to local gene conversion or chimeric recombination. Therefore, to avoid misjudgments caused by local structural recombination of the SMN gene, optionally, in some embodiments, before determining the copy numbers of the SMN1 and SMN2 genes, the method 600 provided in this application further includes a step of determining whether the relative proportions of observations corresponding to different difference intervals meet preset consistency conditions.
[0171] Specifically, in this embodiment, the difference interval is further divided into a first difference interval and a second difference interval. The first difference interval is defined as the difference interval within the regions where intron 6 and exon 7 of the SMN1 and SMN2 genes are located; the second difference interval is defined as the difference interval within the regions where intron 7 of the SMN1 and SMN2 genes are located.
[0172] At step 608, the computing device 110 first supports the test sample in each difference interval. SMN1 Gene sequencing depth and support SMN2 The sequencing depth of the gene is used to determine the corresponding first and second differential intervals. SMN1 Gene sequencing depth relative to SMN2 The relative proportion of gene sequencing depth observations.
[0173] Subsequently, in step 610, the computing device 110 determines whether the observation relative proportion relationship corresponding to each first difference interval and the observation relative proportion relationship corresponding to each second difference interval meet the preset consistency condition.
[0174] Those skilled in the art should understand that if the SMN gene in the sample being tested has not undergone local structural recombination, its SMN1 and SMN2 The DNA molecule should maintain structural continuity across the entire gene span. Therefore, the observed relative ratio presented by the first differential interval (intron 6 / exon 7), within the statistically permissible range of random sequencing noise, should be highly consistent with the ratio presented by the second differential interval (intron 7).
[0175] Conversely, if a large number of observations in the first difference interval (e.g., meeting a preset percentage threshold, such as exceeding 80% of the first difference interval) point to a specific ratio (e.g., meeting a first preset relative ratio threshold, such as...), the relative proportion relationship points to a certain specific ratio. SMN1 : SMN2 The depth ratio is approximately 1:3, while a large number of observations in the second difference interval (e.g., meeting a preset percentage threshold, such as exceeding 80% of the second difference interval) point to a completely different ratio (e.g., meeting a second preset relative ratio threshold, such as...). SMN1 : SMN2 The depth ratio is approximately 2:2. This significant deviation in the relative proportions of this specific region indicates that the genome of the sample under test has undergone local structural reorganization in that region.
[0176] Therefore, the specific logic of the computing device 110 in determining whether the above-mentioned proportional relationship meets the preset consistency condition includes: in response to the fact that the proportion of the observed relative proportional relationship in the first difference interval meets the first preset relative proportional relationship threshold and the proportion of the observed relative proportional relationship in the second difference interval meets the second preset relative proportional relationship threshold both meet the preset proportion threshold, the computing device 110 determines that the preset consistency condition is not met.
[0177] In response to the determination that the above-mentioned preset consistency conditions are met, the computing device 110 will proceed to step 614. In step 614, the computing device 110 determines the total copy number based on the sequencing depth data of all differential regions and the total copy number determined in step 606. SMN1 Gene copy number and SMN2 Gene copy number.
[0178] In response to the determination that the above-mentioned preset consistency conditions are not met, the computing device 110 will proceed to step 612. In step 612, the computing device 110 determines the specific sequencing depth supported by the first differential interval with core diagnostic significance (especially the pathogenic variant region of exon 7), combined with the total copy number determined in step 606 above. SMN1 Gene copy number and the above SMN2 Gene copy number.
[0179] Figure 7 The schematic diagram illustrates an example device 700 suitable for implementing embodiments of this application. For example, device 700 may be as follows: Figure 1 The computing device 110 shown. For example, device 700 may be used to implement execution Figures 2 to 6 The method shown is applicable to devices ranging from 200 to 600.
[0180] like Figure 7 As shown, device 700 includes a central processing unit (i.e., CPU 701), which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (i.e., ROM 702) or loaded from storage unit 708 into random access memory (i.e., RAM 703). RAM 703 may also store various programs and data required for the operation of device 700. CPU 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output interface (i.e., I / O interface 705) is also connected to bus 704.
[0181] Multiple components in device 700 are connected to I / O output interface 705, including: input unit 706, such as keyboard, mouse, microphone, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0182] The various processes and procedures described above, such as methods 200 to 600, may be executed by CPU 701. For example, in some embodiments, methods 200 to 600 may be implemented as computer software programs stored in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by CPU 701, one or more actions of methods 200 to 600 described above may be performed. Alternatively, in other embodiments, CPU 701 may be configured to perform one or more actions of methods 200 to 600 by any other suitable means (e.g., by means of firmware).
[0183] It should be further noted that this application can be a method, apparatus, system, device, computer-readable storage medium, and / or computer program product. The computer program product may include computer-readable program instructions for performing various aspects of this application.
[0184] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0185] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0186] The computer program instructions used to perform the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing the status information of the computer-readable program instructions. These electronic circuits can execute the computer-readable program instructions to implement various aspects of this application.
[0187] Various aspects of this application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0188] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0189] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0190] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0191] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technological improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0192] The above are merely optional embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. Methods for determining the copy number of a target member gene in a homologous gene family, including: Multiple analysis intervals are determined for a reference genome, wherein the analysis intervals within the regions where each member gene in the homologous gene family is located correspond to each other based on the alignment results of the reference sequences of each member gene, and the multiple analysis intervals include at least one absolute homologous interval in which the base sequence is completely identical among the reference sequences of each member gene, and at least one differential interval at the differential site among the reference sequences of each member gene that can represent the base differences specific to the member gene. Based on the alignment results of the sequencing data of the test sample and the reference sample with the preserved multiple alignment sequence of the reference genome, the sequencing depth of the test sample and the reference sample in each analysis interval is determined. Based on the sequencing depth of the reference sample and the test sample within the absolute homology region, stable homology regions are screened to determine the total copy number of the homology gene family; and... Based on the sequencing depth of each member gene supported by the test sample in the difference interval, and the total copy number of the homologous gene family, the copy number of the target member gene is determined; Among them, based on the sequencing depth of the reference sample and the test sample in the absolute homology region, the screening of stable homology regions includes: Based on the dispersion index of sequencing depth of the reference sample in each absolute homology interval, the fluctuation degree of sequencing depth of each absolute homology interval is determined, and the absolute homology interval whose fluctuation degree meets the first preset stability screening condition is selected as the first stable homology interval. Based on the sequencing depth of the reference sample and the test sample in each first stable homology region, a reference depth value for each first stable homology region is determined to calculate the deviation cost function value corresponding to each first stable homology region. The first stable homology region whose deviation cost function value satisfies the second preset stability screening condition is selected as the second stable homology region; and... Based on the sequencing depth of the test sample in each second stable homology interval, the predicted total copy number of the homology gene family is determined, and the fitting cost function value corresponding to each second stable homology interval is calculated. The second stable homology interval whose fitting cost function value satisfies the third preset stability screening condition is selected as the stable homology interval.
2. The method of claim 1, wherein, The absolute homology region does not include regions whose base sequences are completely identical to those of other genomic regions outside the homologous gene family.
3. The method of claim 1, wherein, The differentially expressed sites do not include InDel sites and / or differentially expressed sites that lack individual member gene directionality based on allele frequencies in the population sample.
4. The method of claim 1, wherein, After determining multiple analysis intervals of the reference genome, the method further includes: filtering the analysis intervals according to preset filtering conditions, the preset filtering conditions including: Based on the GC content of each analysis interval, analysis intervals whose GC content does not meet the preset GC content range are removed; and / or, based on the interval length of each analysis interval, analysis intervals whose interval length does not meet the preset length range are removed.
5. The method of claim 1, wherein, The reference sample is a sample whose similarity to the sequencing depth distribution of the sample to be tested in each analysis interval meets the preset similarity condition.
6. The method of claim 1, wherein, Determining the sequencing depth of the test sample and the reference sample in each analysis interval includes: Based on the alignment results between the sequencing data of the test sample and the reference sample and the preserved multiple alignment sequence of the reference genome, the original sequencing depth of the test sample and the reference sample in each analysis interval is determined; and, Depth correction is performed on the original sequencing depth to determine the sequencing depth of the sample to be tested and the reference sample in each analysis interval.
7. The method of claim 6, wherein, The depth correction includes GC correction based on the GC content of the analysis interval, and data volume normalization correction based on the total amount of sequencing data.
8. The method of claim 1, wherein, The sequencing data was obtained based on high-throughput sequencing.
9. The method according to claim 1, wherein, The sequencing data was obtained based on hybridization capture high-throughput sequencing.
10. The method of claim 1, wherein, Before screening for stable homologous regions, the following steps are also included: Based on the central tendency index of sequencing depth of the reference sample in each absolute homology region, the global baseline depth of the absolute homology region is determined; and, Based on the global baseline depth, the absolute homogeneous intervals are divided into interval subsets of different depth levels, and stable homogeneous intervals are independently selected in each interval subset.
11. The method of claim 1, wherein, Determining the degree of variation in sequencing depth for each absolute homology region includes: For each of the absolute homologous intervals: The degree of fluctuation in the sequencing depth of the current absolute homology region is determined based on the dispersion index of the sequencing depth of the reference sample in the current absolute homology region and the deviation of the central tendency index of the sequencing depth of the test sample in the current absolute homology region from the sequencing depth of the reference sample in the current absolute homology region.
12. The method of claim 1, wherein, The reference depth values for each first stable homogeneous region include: For each of the first stable homogeneous regions: Based on the sequencing depth of the test sample in each of the other first stable homology regions besides the current first stable homology region, and the reference depth ratio between the current first stable homology region and the corresponding other first stable homology regions, the reference depth value of the current first stable homology region is determined. The reference depth ratio is determined based on the central tendency index of the sequencing depth of the reference sample in each first stable homology region.
13. The method of claim 1, wherein, The calculation of the deviation cost function value corresponding to each first stable homogeneous interval includes: For each of the first stable homogeneous regions: Based on the statistical distance between the sequencing depth of the test sample in the current first stable homology region and the reference depth value of the current first stable homology region, and the degree of deviation of the central tendency index of the sequencing depth of the test sample in the current first stable homology region and the sequencing depth of the reference sample in the current first stable homology region, the deviation cost function value corresponding to the current first stable homology region is determined.
14. The method of claim 1, wherein, Determining the predicted total copy number of the homologous gene family includes: Based on the sequencing depth of the sample to be tested in each second stable homology region, the total copy number observation value corresponding to each second stable homology region is determined; Based on the observed total copy number corresponding to each second stable homogeneous region, and multiple hypothetical total copy numbers, calculate the total copy number cost function value corresponding to each hypothetical total copy number; and, The hypothetical total copy number corresponding to the minimum total copy number cost function value is selected as the predicted total copy number.
15. The method of claim 14, wherein, The calculation of the total copy number cost function value corresponding to the total copy number of each hypothesis includes: For each of the plurality of hypothetical total copy numbers: Based on the statistical distance between the observed total copy number corresponding to each second stable homology region and the assumed total copy number, and the deviation of the central tendency index of the sequencing depth of the test sample in each second stable homology region from the sequencing depth of the reference sample in the corresponding second stable homology region, the total copy number cost function value corresponding to the current assumed total copy number is determined, and the assumed total copy number value is determined based on the current assumed total copy number.
16. The method according to claim 1, wherein, The calculation of the fitting cost function value corresponding to each second stable homogeneous interval includes: For each of the second stable homogeneous regions: Based on the sequencing depth of the sample to be tested within the current second stable homology region, determine the observed total copy number corresponding to the current second stable homology region; and, Based on the statistical distance between the observed total copy number corresponding to the current second stable homology region and the predicted total copy number, and the deviation of the central tendency index of the sequencing depth of the test sample in the current second stable homology region from the sequencing depth of the reference sample in the current second stable homology region, the fitting cost function value corresponding to the current second stable homology region is determined, and the predicted total copy number is determined based on the predicted total copy number.
17. The method of claim 1, wherein, Determining the total copy number of the homologous gene family includes: Based on the sequencing depth of the test sample in each stable homology region and the reference depth ratio between each stable homology region, the sequencing depth of each stable homology region is corrected. Based on the corrected depth values of the test sample in each stable homology region, and the reference depth ratio of each stable homology region relative to each unstable homology region other than the stable homology regions in the absolute homology regions, the sequencing depth of each unstable homology region is corrected. The reference depth ratio is determined based on the central tendency index of the sequencing depth of the reference sample in each absolute homology region; and, Based on the corrected depth values of the test sample in each stable homology region and each unstable homology region, the total copy number of the homology gene family is determined.
18. The method of claim 17, wherein, Correction of sequencing depth for each stable homology region includes: For each of the stable homogeneous regions: Based on the sequencing depth of the sample under test in all other stable homology regions besides the current stable homology region, and the ratio of the reference depth of the current stable homology region to the corresponding other stable homology regions, the reference depth value of the current stable homology region is determined; and, Based on the sequencing depth of the sample to be tested in the current stable homology region and the reference depth value, as well as the preset stable homology region correction weight, the corrected depth value of the current stable homology region is determined.
19. The method of claim 17, wherein, Correction of sequencing depth for each unstable homology region includes: For each of the aforementioned unstable homogeneous regions: Based on the corrected depth values of the sample under test in each stable homogeneous region, and the ratio of the reference depth of the current unstable homogeneous region to the corresponding stable homogeneous region, the reference depth value of the current unstable homogeneous region is determined; and, Based on the sequencing depth of the sample to be tested in the current unstable homology region and the reference depth value, as well as the preset unstable homology region correction weight, the corrected depth value of the current unstable homology region is determined.
20. The method of claim 17, wherein, Based on the corrected depth values of the test sample in each stable homology region and each unstable homology region, the total copy number of the homologous gene family is determined, including: Based on the ratio of the central tendency index of the sequence depth of the test sample in each absolute homology interval to that of the reference sample in the corresponding absolute homology interval, and the baseline total copy number of the homology gene family, the total copy number of the homology gene family is determined.
21. The method of claim 1, wherein, Based on the sequencing depth of each member gene supported by the test sample in the difference interval, and the total copy number of the homologous gene family, the copy number of the target member gene is determined as follows: Based on the sequencing depth of each member gene supported by the test sample in each differential interval, the relative proportion of the target member gene copy number to the total copy number of the remaining member genes in the homologous gene family, excluding the target member gene, is determined; and, The copy number of the target member gene is determined based on the total copy number of the homologous gene family and the relative proportion of the copy number of the target member gene to the total copy number of the remaining member genes.
22. The method of claim 21, wherein, Determining the relative proportion of the target member gene copy number to the total copy number of the remaining member genes includes: Based on the sequencing depth of each member gene supported by the test sample in each difference interval, the observed relative proportion of the target member gene copy number relative to the total copy number of the remaining member genes in each difference interval is determined, so as to determine the predicted relative proportion of the target member gene copy number relative to the total copy number of the remaining member genes. Based on the observed relative proportions corresponding to each difference interval, and the predicted relative proportions, core difference intervals are selected; and, Based on the observed relative proportions corresponding to each core difference interval, the relative proportion of the target member gene copy number to the total copy number of the remaining member genes is determined.
23. The method according to claim 22, wherein, Determining the predicted relative proportion of the target member gene copy number to the total copy number of the remaining member genes includes: Based on the observed relative proportions corresponding to each difference interval, and multiple assumptions, a relative proportion cost function value corresponding to each assumption is calculated. Each of the multiple assumptions includes: an assumed copy number of the target member gene determined based on the assumed copy number of the target member gene; an assumed copy number of the remaining member genes determined based on the assumed total copy number of the remaining member genes; and an assumed relative proportion between the assumed copy number of the target member gene and the assumed total copy number of the remaining member genes. Select the assumption that the relative proportion cost function value is minimized, and use the assumed relative proportion relationship under this assumption as the predicted relative proportion relationship.
24. The method according to claim 23, wherein, The sum of the hypothetical copy number of the target member gene and the hypothetical total copy number of the remaining member genes is within a preset floating range of the total copy number of the homologous gene family.
25. The method according to claim 23, wherein, Calculating the relative proportional cost function values corresponding to each assumption includes: Based on preset filtering conditions, a priori difference interval is selected from the difference intervals; Based on the sequencing depth of each member gene supported by the test sample within the prior difference interval, the baseline observation of the target member gene copy number corresponding to the target member gene and the baseline observation of the total copy number of the remaining member genes corresponding to the remaining member genes are determined; and, For each of the plurality of assumptions: Based on the statistical distance between the baseline observations corresponding to the target member gene and the remaining member genes and the corresponding hypothetical values under the current hypothetical conditions, and the degree of deviation of the observed relative proportion relationship corresponding to each difference interval from the hypothetical relative proportion relationship under the current hypothetical conditions, the relative proportion cost function value corresponding to the current hypothetical conditions is determined.
26. The method of claim 22, wherein, The core difference intervals to be selected include: Based on the statistical distance between the observed relative proportion relationship corresponding to each difference interval and the predicted relative proportion relationship, the difference intervals whose statistical distance satisfies the preset difference interval screening conditions are selected as the core difference intervals.
27. The method of claim 26, wherein, The core difference intervals also include: A predetermined number of difference intervals are removed from the core difference intervals to minimize the absolute value of the sum of the residuals of the observed relative proportions relative to the predicted relative proportions for each of the retained core difference intervals.
28. The method of claim 22, wherein, Based on the observed relative proportions corresponding to each core difference interval, the relative proportion of the target member gene copy number to the total copy number of the remaining member genes is determined as follows: Based on the central tendency index corresponding to the observed relative proportions of each core difference interval, the relative proportion of the target member gene copy number to the total copy number of the remaining member genes is determined.
29. The method of claim 21, wherein, Based on the total copy number of the homologous gene family and the relative proportion of the target member gene copy number to the total copy number of the remaining member genes, the copy number of the target member gene is determined as follows: Based on the relative proportion of the target member gene copy number to the total copy number of the remaining member genes, the relative abundance of the target member gene copy number relative to the total copy number of the homologous gene family is determined; and, The copy number of the target member gene is determined based on the total copy number of the homologous gene family and the relative abundance of the copy number of the target member gene.
30. The method of claim 21, wherein, Determining the copy number of the target member gene includes: Based on the sequencing depth of each member gene supported by the test sample in each difference interval, the relative proportion of the copy number of each member gene in the homologous gene family to the total copy number of the remaining member genes in the homologous gene family is determined, and based on the relative proportion of the copy number of each member gene to the total copy number of the remaining member genes in the homologous gene family, the relative abundance of the copy number of each member gene to the total copy number of the homologous gene family is determined. Based on the sum of the relative abundance of copy numbers of each member gene in the homologous gene family, a normalization factor is determined to normalize and correct the relative abundance of copy numbers of the target member gene; and, The copy number of the target member gene is determined based on the total copy number of the homologous gene family and the relative abundance of the corrected target member gene copy number.
31. A method for determining SMN1 a gene with SMN2 a gene copy number, comprising: Multiple analytical regions of the reference genome were determined, wherein... SMN1 Genes and the SMN2 The analysis interval within the region where the gene is located is based on SMN1 Gene reference sequence and SMN2 The alignment results of the gene reference sequences correspond to each other, and the plurality of analytical regions contain at least one base sequence in the... SMN1 Gene reference sequence and the SMN2 Absolutely homologous regions that are completely identical between gene reference sequences, and at least one of the above... SMN1 Gene reference sequence and the SMN2 The presence of differentially expressed sites between gene reference sequences can represent the aforementioned SMN1 Genes and the SMN2 The range of differences in gene-specific base combinations; Based on the alignment results of the sequencing data of the test sample and the reference sample with the preserved multiple alignment sequence of the reference genome, the sequencing depth of the test sample and the reference sample in each analysis interval is determined. based on the sequencing depth of the reference sample and the sample to be tested in the absolute homologous interval, screening a stable homologous interval to determine the total copy number of the SMN1 gene and the total copy number of the SMN2 gene; and, Based on the test sample, the difference interval supports the following. SMN1 The sequencing depth of the gene and supporting the above SMN2 The sequencing depth of the gene, and the aforementioned SMN1 Genes and the SMN2 The total copy number of the gene determines the... SMN1 Gene copy number and the above SMN2 Gene copy number.
32. The method of claim 31, wherein, determining the SMN1 gene copy number and the SMN2 gene copy number, further comprising: Based on the test sample in each difference interval, the following supports the... SMN1 The sequencing depth of the gene and supporting the above SMN2 The sequencing depth of the gene determines the corresponding first and second difference intervals. SMN1 Gene sequencing depth relative to the above SMN2 The relative proportion of observed gene sequencing depth, where the first difference interval is the... SMN1 Genes and the described SMN2 The differential interval within the regions containing intron 6 and exon 7 of the gene, wherein the second differential interval is the... SMN1 Genes and the described SMN2 The differential region within the region containing intron 7 of the gene; and, In response to the fact that the relative proportion of observations corresponding to each first difference interval and the relative proportion of observations corresponding to each second difference interval do not meet a preset consistency condition, based on the fact that the test sample supports the first difference interval, SMN1 The sequencing depth of the gene and supporting the above SMN2 The sequencing depth of the gene, and the aforementioned SMN1 Genes and the SMN2 The total copy number of the gene determines the... SMN1 Gene copy number and the above SMN2 Gene copy number.
33. The method of claim 32, wherein, Determining whether the relative proportion of observations corresponding to each first difference interval and the relative proportion of observations corresponding to each second difference interval meet a preset consistency condition includes: If the percentage of the observed relative proportion in the first difference interval that satisfies the first preset relative proportion threshold and the percentage of the observed relative proportion in the second difference interval that satisfies the second preset relative proportion threshold both satisfy the preset percentage threshold, it is determined that the preset consistency condition is not met.
34. A computing device, comprising: include: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, cause the computing device to perform the steps of the method according to any one of claims 1 to 33.
35. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 33.
36. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 33.