Method for establishing sequencing depth prediction model, method for predicting copy number of predetermined region, and application

CN122551880APending Publication Date: 2026-08-11TIANJIN MEDICAL LAB BGI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]然而,这些方法在应用过程中存在诸多不足,第一,对对照样本和数据均一性存在高度依赖,要求对照样本数量充足、无CNV干扰;第二,假基因区域的干扰缺乏有效处理方案,尤其是在目标基因和假基因同源性较高的情况下;第三,靶向捕获测序的特性(探针效率差异、深度波动)未被充分考虑,影响检测方法在实际应用中的适配性与准确性

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551880A_ABST
    Figure CN122551880A_ABST
Patent Text Reader

Abstract

The application provides a sequencing depth prediction model establishment method, a pre-determined region copy number prediction method and application. The foregoing method comprises the following steps: S1, obtaining sequencing depths of multiple pre-determined regions of a training sample, wherein the multiple pre-determined regions are selected from different genes; S2, constructing a virtual target gene sequencing depth of each pre-determined region based on the sequencing depths of the multiple pre-determined regions of the target gene and the pseudo gene in the training sample; S3, performing correlation analysis on the sequencing depths of each pre-determined region of the virtual target gene and all pre-determined regions of the target gene and the pseudo gene, and determining a depth reference region of each pre-determined region of the virtual target gene; and S4, taking the sequencing depth of the reference region as an input feature, taking the sequencing depth of each pre-determined region of the virtual target gene as a label, and training a machine learning model. The foregoing method avoids the dependence on a control sample set, realizes rapid construction of a sequencing depth prediction model, and obtains a model that is not interfered by a pseudo gene in sequencing depth prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics, specifically to a method for establishing a sequencing depth prediction model, a method for predicting the copy number of a predetermined region, and their applications. Background Technology

[0002] Copy number variants (CNVs) refer to increases or decreases in the copy number of large DNA sequences longer than 1 kb in the genome, mainly manifested as deletions and duplications at the submicroscopic level, and are one of the important pathogenic factors of human diseases. Currently, CNV detection methods for targeted capture sequencing data are mainly calculated based on differences in sequencing depth between samples.

[0003] However, these methods have many shortcomings in application. First, they are highly dependent on the homogeneity of control samples and data, requiring a sufficient number of control samples without CNV interference. Second, there is a lack of effective solutions to handle interference from pseudogene regions, especially when the target gene and pseudogene have high homology. Third, the characteristics of targeted capture sequencing (probe efficiency differences, depth fluctuations) have not been fully considered, affecting the adaptability and accuracy of the detection methods in practical applications. Summary of the Invention

[0004] This application aims to at least partially address one of the technical problems in the related art. To this end, one objective of this application is to propose a copy number variation prediction method that does not rely on control samples, avoids false detections or false negatives due to pseudogene interference, and improves the accuracy of CNV detection.

[0005] Specifically, this application proposes the following technical solution:

[0006] In a first aspect of this application, a method for establishing a sequencing depth prediction model is proposed. According to an embodiment of this application, the method includes: S1, obtaining the sequencing depth of multiple predetermined regions in a training sample, wherein the multiple predetermined regions are selected from different genes; S2, constructing the sequencing depth of each predetermined region of a virtual target gene based on the sequencing depth of multiple predetermined regions of a target gene and a pseudogene in the training sample; S3, performing correlation analysis on the sequencing depth of each predetermined region of the virtual target gene with the sequencing depth of all predetermined regions other than the target gene and pseudogenes to determine a reference sequencing depth for each predetermined region of the virtual target gene; S4, training a machine learning model using the reference sequencing depth as input features and the sequencing depth of each predetermined region of the virtual target gene as a label.

[0007] In some examples of this application, the aforementioned method avoids dependence on control sample sets, enables rapid construction of sequencing depth prediction models, and the obtained models are not affected by pseudogenes in sequencing depth prediction.

[0008] In a second aspect of this application, a method for predicting copy number in predetermined regions is proposed. According to embodiments of this application, the method includes: a. obtaining the actual sequencing depth of multiple predetermined regions of a target gene in a sample to be tested; b. constructing a virtual sequencing depth for each predetermined region of a target gene based on the sequencing depth of the multiple predetermined regions of the target gene; c. inputting the sequencing depth of each predetermined region of the virtual target gene into a pre-trained machine learning model to determine the predicted sequencing depth of each predetermined region; d. predicting the copy number of each predetermined region based on the predicted sequencing depth and the actual sequencing depth of each predetermined region; wherein the machine learning model is trained using the method described in the first aspect. In some examples of this application, the aforementioned method is applicable to targeted sequencing, achieving copy number prediction without relying on samples from the same batch or a pre-constructed set of control samples, and effectively avoiding pseudogene interference, significantly improving prediction accuracy and reducing analysis costs.

[0009] In a third aspect of this application, a method for predicting copy number variation is proposed. According to an embodiment of this application, the method includes: determining the copy number variation of each predetermined region based on the difference between the predicted copy number of each predetermined region and a predetermined copy number threshold; wherein the predicted copy number of each predetermined region is obtained by the method described in the second aspect.

[0010] In some examples of this application, the method avoids the dependence on a control sample set found in traditional methods by using other regions of the sample as a reference set for the predetermined region. By employing machine learning models for prediction, it effectively avoids interference with detection results when CNVs are present in other regions of the sample. Even with uneven sequencing depth, this method can still achieve accurate CNV detection.

[0011] In a fourth aspect, this application proposes a sequencing depth prediction model establishment apparatus. According to an embodiment of this application, the apparatus includes: a sequencing depth acquisition unit for acquiring sequencing depths of multiple predetermined regions in a training sample, wherein the multiple predetermined regions are selected from different genes; a virtual gene sequencing depth construction unit for constructing sequencing depths of each predetermined region of a virtual target gene based on the sequencing depths of multiple predetermined regions of a target gene and a pseudogene in the training sample; a reference sequencing depth determination unit for performing correlation analysis between each predetermined region of the virtual target gene and the sequencing depths of all predetermined regions other than the target gene and pseudogenes to determine a reference sequencing depth for each predetermined region of the virtual target gene; and a training unit for training a machine learning model using the reference sequencing depth as input features and the sequencing depths of each predetermined region of the virtual target gene as labels.

[0012] In some examples of this application, the aforementioned device, through its modular design, can efficiently complete the entire process of obtaining sequencing depth, constructing virtual target genes, determining reference sequencing depth, and training models, thus avoiding dependence on control sample sets.

[0013] In a fifth aspect of this application, a predetermined region copy number prediction system is proposed. According to an embodiment of this application, the system includes: a sequencing depth acquisition module for acquiring the actual sequencing depth of multiple predetermined regions of a target gene in a sample to be tested; a virtual gene sequencing depth acquisition module for constructing a virtual sequencing depth for each predetermined region of a target gene based on the sequencing depth of the multiple predetermined regions of the target gene; a predicted sequencing depth determination module for inputting the sequencing depth of each predetermined region of the virtual target gene into a pre-trained machine learning model to determine the predicted sequencing depth of each predetermined region; and a copy number prediction module for predicting the copy number of each predetermined region based on the predicted sequencing depth and the actual sequencing depth of each predetermined region; wherein the machine learning model is trained using the apparatus described in the fourth aspect.

[0014] In some examples of this application, the aforementioned system relies on a pre-trained machine learning model to accurately predict the copy number of each predetermined region of the target gene, avoiding dependence on external control sample sets and reducing the impact of pseudogene interference on the results.

[0015] In a sixth aspect of this application, a copy number variation prediction device is proposed. According to an embodiment of this application, the device includes: a discrimination module, configured to determine the copy number variation of each predetermined region based on the difference between the predicted copy number of each predetermined region and a predetermined copy number threshold; wherein the predicted copy number of each predetermined region is obtained through the system described in the fifth aspect.

[0016] In some examples of this application, the aforementioned device automatically identifies copy number changes in the target region based on the difference between the predicted copy number and a predetermined copy number threshold, providing an intuitive CNV result judgment and improving discrimination efficiency and detection accuracy. Furthermore, this device does not rely on external control samples or uniform sequencing depth, and can achieve stable and reliable copy number variation discrimination under complex sample conditions.

[0017] In a seventh aspect, this application provides a computing device. According to an embodiment of this application, the computing device includes: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program to implement the sequencing depth prediction model establishment method as described in the first aspect, the predetermined region copy number prediction method as described in the second aspect, or the copy number variation prediction method as described in the third aspect.

[0018] In some examples of this application, the aforementioned computing device achieves high efficiency and automation by automatically executing computer instructions to establish sequencing depth prediction models, predict copy number in predetermined regions, or predict copy number variations, thereby improving prediction efficiency and accuracy. Secondly, the instruction-based nature of these methods ensures high consistency and reliability in complex experimental scenarios.

[0019] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 A schematic diagram of the sequencing depth prediction model establishment method provided in the embodiments of this application;

[0022] Figure 2 This is a schematic flowchart of the predetermined region copy number prediction method provided in the embodiments of this application;

[0023] Figure 3 This is a schematic flowchart of the copy number variation prediction method provided in the embodiments of this application;

[0024] Figure 4 A schematic diagram of the sequencing depth prediction model establishment device provided in the embodiments of this application;

[0025] Figure 5A schematic diagram of a predetermined region copy number prediction system provided in an embodiment of this application;

[0026] Figure 6 This is a schematic diagram of the copy number variation prediction device provided in an embodiment of this application;

[0027] Figure 7 A schematic diagram of an electronic device provided in an embodiment of this application;

[0028] Figure 8 A schematic diagram illustrating the correlation distribution between exon prediction results and actual sequencing depth provided in the embodiments of this application;

[0029] Figure 9 A schematic diagram illustrating the performance results of the all-base classifier prediction and different anomaly handling provided in the embodiments of this application;

[0030] Figure 10 This is a schematic diagram of the training sample copy number distribution results provided in the embodiments of this application;

[0031] Figure 11 This is a schematic diagram of the test sample copy number results provided in the embodiments of this application;

[0032] Figure 12 This is a schematic diagram of the corrected copy number results of the test samples provided in the embodiments of this application. Detailed Implementation

[0033] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in sequences other than those illustrated or described herein. In this application, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices. In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0035] This application proposes a method and apparatus for establishing a sequencing depth prediction model, a method and system for predicting copy number in a predetermined region, a method and device for predicting copy number variations, and a computing device. These are described below:

[0036] Methods for establishing sequencing depth prediction models

[0037] In one aspect of this application, a method for establishing a sequencing depth prediction model is proposed, referring to... Figure 1 The method includes:

[0038] S1, obtain the sequencing depth of multiple predetermined regions of the training sample, wherein the multiple predetermined regions are selected from different genes;

[0039] In some examples of this application, the aforementioned training samples may come from public databases or be multiple samples sequenced independently. The data type may be selected from targeted sequencing, whole exome sequencing, or whole genome sequencing.

[0040] In some examples of this application, the aforementioned predetermined region may be selected from exon regions, intron regions, promoter regions, or enhancer regions. In some preferred examples, the aforementioned predetermined region is selected from exon regions.

[0041] Sequencing depth refers to the number of times a target genome or specific region (such as exons or other functional regions) is read during DNA sequencing. In some examples of this application, sequencing depth data for multiple predetermined regions (such as exons) are obtained using commonly used biological software (such as samtools, GATK, bedtools, etc.).

[0042] It is understood that the aforementioned different genes include target genes, pseudogenes, or other genes (non-target genes). It should be noted that the aforementioned pseudogenes are pseudogenes of the target gene.

[0043] S2, based on the sequencing depth of multiple predetermined regions of the target gene and pseudogene in the training sample, construct the sequencing depth of each predetermined region of the virtual target gene;

[0044] Pseudogenes are a class of non-functional genes present in the genome. They are highly similar in sequence to functional genes (true genes), but due to structural defects or mutations, they have lost the ability to encode proteins or perform biological functions. Pseudogenes typically originate from gene duplication, reverse transcription, or other gene rearrangement events. In some examples of this application, the target gene is a functional gene, while the pseudogene is a gene with a sequence highly similar to the target gene but lacking function. Pseudogenes in the human genome have been extensively studied and are included in public databases such as GeneCards (see database link: https: / / www.genecards.org / cgi-bin / carddisp.pl?gene=STRC).

[0045] In this application, a "virtual gene" is a simulated gene constructed through mathematical modeling or statistical methods. Its sequencing depth is calculated by combining the sequencing depths of the target gene and related pseudogenes in a real sample. The virtual gene does not exist directly in the biological genome but is used to represent a theoretical sequencing depth model of the target gene under the influence of pseudogenes.

[0046] In some examples of this application, the sequencing depth of each predetermined region of the aforementioned virtual target gene is represented as follows:

[0047]

[0048] depth v-i The depth represents the sequencing depth of the i-th region of the virtual target gene. real-i The depth represents the sequencing depth of the i-th region of the target gene. pse-i This represents the sequencing depth of the i-th region of the pseudogene. For example, assuming the sequencing depths of multiple predetermined regions of target gene A are [100, 120, 150, 130], and the sequencing depth of pseudogene B is [110, 115, 140, 125], a weighted average can be used to obtain the sequencing depth of each predetermined region of the virtual target gene, [105, 117, 145, 128]. By constructing a virtual gene, the influence of pseudogenes can be better eliminated, thereby obtaining more accurate target gene sequencing data.

[0049] S3, perform correlation analysis on the sequencing depth of each predetermined region of the virtual target gene and all predetermined regions other than the target gene and pseudogene, and determine the reference sequencing depth of each predetermined region of the virtual target gene;

[0050] In some examples of this application, correlation analysis methods (such as Pearson correlation coefficient, Spearman rank correlation, etc.) are used to identify predetermined regions outside the target gene and pseudogene that have a strong linear relationship with the predetermined region of the virtual target gene. The sequencing depth of the predetermined regions outside the target gene and pseudogene is used as a reference value for the sequencing depth of the predetermined region of the virtual target gene.

[0051] In some examples of this application, the sequencing depth of a predetermined region of a virtual target gene with a relevance higher than a predetermined threshold (such as 0.9, 0.8, or 0.7) can be selected as a reference sequencing depth; alternatively, the sequencing depth of a predetermined region of the top few virtual target genes (such as 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100) ranked from high to low relevance can be selected as a reference sequencing depth.

[0052] S4, using the reference sequencing depth as the input feature and the sequencing depth of each predetermined region of the virtual target gene as the label, the machine learning model is trained.

[0053] In some examples of this application, the aforementioned machine learning model is selected from random forest, and the training process uses regression tree as the base learner to perform regression learning based on the sequencing depth of the training samples.

[0054] In some cases, the sequencing depth of a reference region may be abnormally high or low (possibly due to CNV). To address this, the inventors further eliminated abnormal reference sequencing depths. Specifically, this includes: 1) obtaining all base learner prediction results and determining the mean and standard deviation; 2) eliminating abnormal data, including: i) prediction results greater than N times the mean standard deviation; ii) prediction results less than N times the mean standard deviation; where N is selected from integers, such as 1, 2, 3, 4, or 5; 3) analyzing the prediction results after eliminating abnormal data to determine the predicted sequencing depth of the target gene.

[0055] In some examples of this application, the aforementioned analytical processing may be selected from weighted average.

[0056] Using the above method, a model with higher prediction accuracy can be obtained without relying on the same batch or a pre-constructed control sample set, effectively avoiding the impact of pseudogenes on prediction accuracy.

[0057] Predicting method for the number of copies in a predetermined area

[0058] In another aspect of this application, a method for predicting the copy number of a predetermined region is proposed, with reference to... Figure 2 The method includes:

[0059] a. Obtain the true sequencing depth of multiple predetermined regions of the target gene in the sample to be tested;

[0060] The so-called "true sequencing depth" refers to the sequencing depth obtained after sequencing the sample to be tested.

[0061] In some examples of this application, the aforementioned predetermined regions may be selected from exon regions, intron regions, promoter regions, or enhancer regions. In some preferred examples, the aforementioned predetermined regions are selected from exon regions.

[0062] b. Based on the sequencing depth of multiple predetermined regions of the target gene, construct a virtual sequencing depth for each predetermined region of the target gene;

[0063] In some examples of this application, the sequencing depth of each predetermined region of the aforementioned virtual target gene is represented as follows:

[0064]

[0065] depth v The depth represents the sequencing depth of the i-th region of the virtual target gene. real The depth represents the sequencing depth of the i-th region of the target gene. pse-i This indicates the sequencing depth of the i-th region of the pseudogene.

[0066] In this paper, the descriptions of virtual genes and pseudogenes are the same as those for establishing sequencing depth prediction models, and will not be repeated here.

[0067] c. Input the sequencing depth of each predetermined region of the virtual target gene into a pre-trained machine learning model to determine the predicted sequencing depth of each predetermined region; the aforementioned machine learning model is constructed using any of the methods in the first aspect.

[0068] d. Based on the predicted sequencing depth and the actual sequencing depth of each predetermined region, predict the copy number of each predetermined region.

[0069] In some examples of this application, the predicted copy number for each predetermined region is expressed as:

[0070]

[0071] Among them, CopyNum i The depth represents the predicted copy number of the i-th region of the target gene. real-i Depth represents the actual sequencing depth of the i-th region. predict-i This represents the predicted sequencing depth for the i-th region.

[0072] In some examples of this application, if no CNV occurs in the predetermined region of the target gene, the actual sequencing depth should be close to the theoretically predicted depth, and the copy number should be close to 2 (humans are usually diploid).

[0073] If pseudogene interference exists in the target gene, the ratio between the actual sequencing depth and the theoretical predicted depth will change, thus affecting the copy number calculation results. Therefore, it is necessary to correct the predicted copy number for each predetermined region to improve detection accuracy. In some examples of this application, the correction includes:

[0074] 1) Determine the correction coefficient based on the difference in actual sequencing depth between the target gene and the pseudogene in each predetermined region;

[0075] In some examples of this application, the correction coefficient is expressed as:

[0076]

[0077] Among them, Target_Ratio i The depth represents the correction coefficient for the i-th region of the target gene. real-i-j The depth represents the sequencing depth of the j-th differentially expressed site in the i-th region of the target gene. pse-i-j This represents the sequencing depth of the j-th differential base site in the i-th region of the pseudogene.

[0078] For example, suppose gene A has multiple sequencing sites in the third exon region, and the sequencing depth of each site is as follows:

[0079] For position 1: depth A-3-1 =100, depth pseA-3-1 =20;

[0080] For position 2: depth A-3-2 =80, depth pse--3-2 =30;

[0081] For position 3: depth A-3-3 =120, depth pseA-3-3 =40;

[0082] Position 1 ratio: 100 / (100+20)=0.833;

[0083] Position 2 ratio: 80 / (80+30) = 0.727;

[0084] Position 3 ratio: 120 / (120+40)=0.75;

[0085] Target_Ratio3=(0.833+0.727+0.75) / 3=0.77.

[0086] 2) Based on the correction coefficient and the predicted copy number of each predetermined region, determine the copy number of each predetermined region of the target gene.

[0087] In some examples of this application, the copy number of each predetermined region of the target gene is expressed as:

[0088] CopyNum fit-i =CopyNum i *2*Target_Ratio i ;

[0089] Among them, CopyNum fit-i This represents the copy number correction value for the i-th region of the target gene; CopyNum i Target_Ratio represents the predicted copy number of the i-th region of the target gene. i This represents the correction coefficient for the i-th region of the target gene.

[0090] Pseudogenes and target genes may have base differences at certain sites, which can interfere with sequencing depth ratios. By calculating the Target_Ratio correction coefficient, errors caused by pseudogene contamination can be effectively corrected, thus providing more accurate copy number prediction results when target genes and pseudogenes coexist.

[0091] The aforementioned method effectively eliminates pseudogene interference and improves the accuracy of copy number prediction by constructing a virtual target gene sequencing depth. Analysis based on the internal data of the sample to be tested eliminates the need for an additional control sample set, reducing the dependence on experimental conditions and making it suitable for various sample types and complex experimental scenarios. With the help of a pre-trained machine learning model, this method can adapt to uneven sequencing depths, especially performing well in samples with low coverage or high variation. Furthermore, by directly comparing the actual sequencing depth with the predicted sequencing depth, the calculation process is simplified, improving efficiency.

[0092] Copy number variation prediction method

[0093] In another aspect of this application, a method for predicting copy number variations is proposed, referring to... Figure 3 The method includes: e. determining the copy number change of each predetermined region based on the difference between the predicted copy number of each predetermined region and a predetermined copy number threshold; wherein the predicted copy number of each predetermined region is obtained by the predetermined region copy number prediction method of any of the foregoing examples.

[0094] In some examples of this application, if the predicted copy number of the predetermined region is greater than a predetermined copy number threshold, the copy number of the predetermined region increases; if the predicted copy number of the predetermined region is less than the predetermined copy number threshold, the copy number of the predetermined region decreases.

[0095] In some examples of this application, the predetermined copy number threshold is 2.

[0096] Those skilled in the art will understand that in normal diploid organisms (such as humans), each exon generally has two copies (one from each parent's chromosome). Therefore, for a normal sample, the theoretical copy number of each exon is 2. If the detected copy number is greater than 2, it indicates that the copy number in that region may have increased; if the detected copy number is less than 2, it indicates that the copy number in that region has decreased. However, in practical applications, due to human error or systematic error, the detected value may deviate slightly (e.g., 2.1 or 1.8), which may be within the normal range and does not necessarily indicate copy number variation. Therefore, the inventors further use Z-scores to characterize the reliability of copy number prediction.

[0097] In some examples of this application, the Z-score is represented as:

[0098] Z-score = (CopyNum fit-i -average) / std;

[0099] Among them, CopyNum fit-i represents the copy number correction value of the i-th region of the target gene; average represents the mean copy number of the i-th region of the training samples; std represents the standard deviation of the copy number of the i-th region of the training samples.

[0100] In some examples of this application, the Z-score is compared with a standard Z-value to quantify the reliability of the CNV detection results. In this paper, the aforementioned standard Z-value can be obtained based on a cutoff value from a large number of normal and abnormal samples. For example, it can be 3, 4, or 5. In some examples of this application, a standard Z-value of not less than 3 is considered a reliable CNV detection result.

[0101] The aforementioned method quantifies the reliability of copy number variation results by comparing the predicted copy number with a predetermined copy number threshold and combining this with Z-score calculations. This reduces misjudgments caused by systematic errors or human factors, effectively improving detection accuracy. Furthermore, the Z-score, combined with the statistical distribution of a large number of normal and abnormal samples, enhances the method's adaptability under different experimental conditions and sample types, effectively reducing the interference of minor fluctuations in sequencing depth on the results, especially avoiding misjudgments when the copy number is close to the threshold.

[0102] Sequencing depth prediction model building device

[0103] In another aspect of this application, a sequencing depth prediction model establishment device is proposed, referring to... Figure 4 The device includes: a sequencing depth acquisition unit 100, a virtual gene sequencing depth construction unit 200, a reference sequencing depth determination unit 300, and a training unit 400.

[0104] 100 units are used to obtain the sequencing depth of multiple predetermined regions of training samples, wherein the multiple predetermined regions are selected from different genes;

[0105] In some examples of this application, the predetermined region is selected from exon regions, intron regions, promoter regions, or enhancer regions; preferably, the predetermined region is selected from exon regions.

[0106] 200 units, based on the sequencing depth of multiple predetermined regions of the target gene and pseudogene in the training sample, to construct the sequencing depth of each predetermined region of the virtual target gene;

[0107] In some examples of this application, the sequencing depth of each predetermined region of the virtual target gene is represented as:

[0108]

[0109] Depth v The depth represents the sequencing depth of the i-th region of the virtual target gene. real The depth represents the sequencing depth of the i-th region of the target gene. pse-i This indicates the sequencing depth of the i-th region of the pseudogene.

[0110] 300 units were used to perform correlation analysis between the sequencing depth of each predetermined region of the virtual target gene and the sequencing depth of all predetermined regions other than the target gene and pseudogenes, and to determine the reference sequencing depth of each predetermined region of the virtual target gene.

[0111] The machine learning model is trained using 400 units, with the reference sequencing depth as the input feature and the sequencing depth of each predetermined region of the virtual target gene as the label.

[0112] In some examples of this application, the machine learning model is selected from random forest, and the training process uses regression tree as the base learner to perform regression learning based on the sequencing depth of the training samples.

[0113] In some examples of this application, the method further includes: 1) obtaining all base learner prediction results and determining the mean and standard deviation; 2) removing outlier data, the outlier data including: i) prediction results that are greater than N times the mean and standard deviation; ii) prediction results that are less than N times the mean and standard deviation; wherein N is selected from integers, such as 1, 2, 3, 4 or 5; 3) analyzing and processing the prediction results after removing outlier data to determine the predicted sequencing depth of the target gene.

[0114] In some examples of this application, the aforementioned units 100, 200, 300, and 400 are connected. This connection includes both physical and network connections.

[0115] Those skilled in the art will understand that the features and advantages described above for the sequencing depth prediction model establishment method also apply to the above-mentioned device, and will not be repeated here.

[0116] It should be understood that the device embodiments and method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, further details will not be provided here. Specifically, Figure 4 The apparatus shown can execute the above-described method for establishing a sequencing depth prediction model, and the operations and / or functions performed by each unit in the apparatus correspond to those in the method embodiment. For the sake of brevity, they will not be described in detail here.

[0117] Predicted area copy number prediction system

[0118] In another aspect of this application, a predetermined region copy number prediction system is proposed, with reference to... Figure 5 The system includes: a sequencing depth acquisition module 01, a virtual gene sequencing depth acquisition module 02, a predicted sequencing depth determination module 03, and a copy number prediction module 04. Among them,

[0119] Module 01 is used to obtain the actual sequencing depth of multiple predetermined regions of the target gene in the sample to be tested;

[0120] Module 02 is used to construct a virtual sequencing depth for each predetermined region of the target gene based on the sequencing depth of multiple predetermined regions of the target gene;

[0121] Module 03 is used to input the sequencing depth of each predetermined region of the virtual target gene into a pre-trained machine learning model to determine the predicted sequencing depth of each predetermined region; wherein, the machine learning model is trained using the aforementioned sequencing depth prediction model establishment device.

[0122] Module 04 is used to predict the copy number of each predetermined region based on the predicted sequencing depth and the actual sequencing depth of each predetermined region.

[0123] In some examples of this application, the predicted copy number for each predetermined region is expressed as:

[0124]

[0125] Among them, CopyNum i The depth represents the predicted copy number of the i-th region of the target gene. real-i Depth represents the actual sequencing depth of the i-th region. predict-i This represents the predicted sequencing depth for the i-th region.

[0126] In some examples of this application, if the target gene contains pseudogenes, the method further includes: correcting the predicted copy number of each predetermined region.

[0127] In some examples of this application, the correction includes: 1) determining a correction coefficient based on the difference in actual sequencing depth between each predetermined region of the target gene and the pseudogene; 2) determining the copy number of each predetermined region of the target gene based on the correction coefficient and the predicted copy number of each predetermined region.

[0128] In some examples of this application, the correction coefficient is expressed as:

[0129]

[0130] Among them, Target_Ratio i The depth represents the correction coefficient for the i-th region of the target gene. real-i-j The depth represents the sequencing depth of the j-th differentially expressed site in the i-th region of the target gene. pse-i-j This represents the sequencing depth of the j-th differential base site in the i-th region of the pseudogene.

[0131] In some examples of this application, the copy number of each predetermined region of the target gene is expressed as:

[0132] CopyNum fit-i =CopyNum i *2*Target_Ratio i ;

[0133] Among them, CopyNum fit-i This represents the copy number correction value for the i-th region of the target gene; CopyNum i Target_Ratio represents the predicted copy number of the i-th region of the target gene; Target_Ratio represents the correction coefficient for the i-th region of the target gene.

[0134] In some examples of this application, the aforementioned modules 01, 02, 03, and 04 are interconnected. This interconnection includes both physical and network connections.

[0135] Those skilled in the art will understand that the features and advantages described above for the method of predicting the copy number of a predetermined region also apply to the above system, and will not be repeated here.

[0136] It should be understood that the system embodiments and method embodiments can correspond to each other, and similar descriptions can be found in the method embodiments. To avoid repetition, further details are omitted here. Specifically, Figure 5 The system shown can execute the above-described embodiment of the predetermined region copy number prediction method, and the operations and / or functions performed by each module in the system correspond to those in the method embodiment. For the sake of brevity, they will not be described in detail here.

[0137] Copy number variation prediction device

[0138] In another aspect of this application, a copy number variation prediction device is proposed, with reference to Figure 6 The device includes: a discrimination module 05, used to determine the copy number change of each predetermined region based on the difference between the predicted copy number of each predetermined region and the predetermined copy number threshold; wherein the predicted copy number of each predetermined region is obtained through the aforementioned predetermined region copy number prediction system.

[0139] Those skilled in the art will understand that the features and advantages described above for the copy number variation prediction method also apply to the above system, and will not be repeated here.

[0140] It should be understood that the system embodiments and method embodiments can correspond to each other, and similar descriptions can be found in the method embodiments. To avoid repetition, further details are omitted here. Specifically, Figure 6 The system shown can execute the above-described copy number variation prediction method embodiment, and the operations and / or functions performed by each module in the system correspond to those in the method embodiment. For the sake of brevity, they will not be described in detail here.

[0141] computing devices

[0142] In another aspect of this application, a computing device is proposed that enables the execution of the aforementioned sequencing depth prediction model establishment method, predetermined region copy number prediction method, or copy number variation prediction method.

[0143] The term "electronic device" is intended to refer to various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Computing devices can also refer to various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0144] like Figure 7 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 502 or a computer program loaded from storage unit 508 into RAM (Random Access Memory) 503. The RAM 503 can also store various programs and data required for the operation of the device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.

[0145] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0146] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as sequencing depth prediction model building methods, predetermined region copy number prediction methods, or copy number variation prediction methods. For example, in some embodiments, the sequencing depth prediction model building methods, predetermined region copy number prediction methods, or copy number variation prediction methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by computing unit 501, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, computing unit 501 can be configured by any other suitable means (e.g., by means of firmware) to perform the aforementioned sequencing depth prediction model establishment method, predetermined region copy number prediction method, or copy number variation prediction method.

[0147] In this application, the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, a computer-readable medium can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory. The various computer-readable storage media described in this invention can represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" can include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.

[0148] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0149] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0150] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0151] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0152] Embodiments of this application will now be described in more detail, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the invention, and should not be construed as limiting the scope of this application.

[0153] Example 1: Establishment and Application of Sequencing Depth Prediction Model

[0154] This embodiment uses the detection of CNVs in STRC exons with pseudogenes as an example to verify the feasibility and effectiveness of the technical solution of this application.

[0155] I. Preliminary Data Preparation:

[0156] The preliminary data preparation in this embodiment is all conventional technical operation in the field, and only some concepts are illustrated here.

[0157] Bam file: After the data is processed by quality control (e.g., using SOAPNUKE), it is aligned (e.g., using BWA) to a reference genome (e.g., Hg19), and then deduplicated (e.g., using GATK's MarkDuplicates) to produce the Bam file.

[0158] Target region file: A file containing gene exon annotations of the target capture range, recording all exon regions to be analyzed and their corresponding exon numbers.

[0159] II. Model Building Steps:

[0160] 1. First, based on 355 modeling samples, the depth of these samples in each exon region was calculated. x-i Where x represents an exon (e.g., STRC:exon-1), and i represents the i-th sample;

[0161] 2. Calculate the average depth of each exon for STRC (target gene) and STRCP1 (pseudogene) to obtain the depth of each exon for v-STRC (v- indicates virtual gene). v-STRC:exon-1_i =(Depth) STRC:exon-1_i +Depth STRCP1:exon-1_i ) / 2;

[0162] 3. Calculate the correlation coefficients of exon depths for each exon of v-STRC and all non-target regions (including true and false genes) within the capture interval. Then, select the 100 exons with the highest correlation coefficients as feature exons. A partial list of the top 20 exons is shown in Table 1.

[0163] Table 1 Top 20 exons

[0164]

[0165] 4. Ensemble learning is used to model each exon. During the training phase of ensemble learning, random forest is used to learn the parameters of the base learner. The base learner uses a regression tree as input features, which are the depths of the 100 most relevant exons for each exon obtained in step 3 in each sample. The theoretical depth of the target exon is predicted using default prediction logic. The correlation distribution between the predicted depth and the actual depth of each exon is shown below. Figure 8 As shown, the correlation obtained after ensemble learning is significantly higher than that obtained by directly using the average depth of the samples. For example, v-STRC:Exon14 has a correlation of only 0.66 with the average depth of the samples, but a correlation of over 0.9 with the depth predicted by the model.

[0166] 5. To eliminate CNV positivity in the selected feature exons, anomaly handling was performed in the prediction module, and the performance of the adjusted prediction module was confirmed. The specific anomaly handling process was as follows: the mean and standard deviation of the prediction results of all base learners were calculated; anomalies exceeding twice the standard deviation of the mean were removed; and the mean of the remaining base learner prediction results was calculated. The performance of the prediction module was further improved after anomaly handling, with overall correlation as follows: Figure 9 As shown, compared to the model building method without outlier handling, outlier handling of ±5% extreme values ​​(trim±5%) results in a decrease in model prediction performance; handling outliers outside ±2 standard deviations (std) improves model performance and preserves the parameters of the best-performing model.

[0167] III. Determination of Copy Number Threshold

[0168] Calculate the copy number of the target exon using the actual sequencing depth and the model-predicted depth: CopyNum = (Depth)v-STRC:exon-1_i / Depth pre-v-STRC:exon-1_i *2. The copy number distribution of all training samples is as follows: Figure 10 As shown.

[0169] The copy number information for all samples in each exon is statistically analyzed, and the mean and standard deviation of the copy number results for each exon are calculated. The cutoff threshold for copy number loss in each exon is also calculated. down =mean-2*std and amplification copy number threshold: Cutoff up =mean + 2*std.

[0170] IV. CNV Detection

[0171] 1. First, the depth of a positive sample with a complete exon deletion of STRC (*2928, independent of the training sample) is calculated by aligning it with the Bam file. x Where x represents an exon (e.g., STRC:exon-1);

[0172] 2. Calculate the average depth of each exon for STRC and STRCP1 to obtain the depth of each exon for v-STRC (v- indicates virtual gene). v-STRC:exon-1_i =(Depth) STRC:exon-1_i +Depth STRCP1:exon-1_i ) / 2;

[0173] 3. Substitute the characteristic exon depth information corresponding to each exon obtained statistically into the saved model above to calculate the theoretical depth value of each exon. Then, calculate the theoretical copy number of each exon based on the actual sequencing depth and theoretical depth: CopyNum = (Depth... v-STRC:exon-1_i / Depth pre-v-STRC:exon-1_i )*2. And compare the calculated copy number with the threshold obtained during the model building phase.

[0174] The copy number result is as follows Figure 11 As shown, since STRC and STRCP1 have base differences in exons 19, 20, 22-26, the number of true gene bases and pseudogene bases at the differing sites in each exon is counted, and the ratio of the number of true gene bases to the total number of true and pseudogene bases is calculated and denoted as Ratio. site_i Then, the average base ratio at all sites within the exon is used as the copy number ratio of the true gene exon in the virtual machine, denoted as Target_Ratio. The copy number is then corrected using the copy number ratio of the corresponding exon of the true gene. fit=CopyNum * 2 * Target_Ratio. The sample copy number, after being corrected for the ratio of true genes to false genes, is homozygous deletion (0 copies), as shown in the following figure. Figure 12 As shown in (CopyNumFit).

[0175] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0176] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.

Claims

1. A method for establishing a sequencing depth prediction model, characterized in that, include: S1, obtain the sequencing depth of multiple predetermined regions of the training sample, wherein the multiple predetermined regions are selected from different genes; S2, based on the sequencing depth of multiple predetermined regions of the target gene and pseudogene in the training sample, construct a virtual target gene, with each predetermined region having a sequencing depth of; S3, perform correlation analysis on the sequencing depth of each predetermined region of the virtual target gene and all predetermined regions of non-target genes and pseudogenes to determine the reference sequencing depth of each predetermined region of the virtual target gene; S4. Using the reference sequencing depth as the input feature and the sequencing depth of each predetermined region of the virtual target gene as the label, the machine learning model is trained.

2. The method of claim 1, wherein, The predetermined region is selected from exon regions, intron regions, promoter regions, or enhancement regions; Preferably, the predetermined region is selected from the exon region.

3. The method of claim 2, wherein, The machine learning model is selected from random forest, and the training process uses regression tree as the base learner, and performs regression learning based on the sequencing depth of the training samples.

4. The method of claim 1, wherein, The sequencing depth of each predetermined region of the virtual target gene is represented as follows: Depth v-i represents the sequencing depth of the ith region of the virtual target gene; depth real-i represents the sequencing depth of the ith region of the target gene; depth pse-i represents the sequencing depth of the ith region of the pseudo gene.

5. The method according to any one of claims 1 to 4, characterized in that, Further includes: 1) Obtain the prediction results of all base learners and determine the mean and standard deviation; 2) Remove outlier data, including: i, the predicted result that is greater than N times the standard deviation of the mean; ii. The predicted result is less than N times the standard deviation of the mean; Wherein, N is selected from integers, such as 1, 2, 3, 4 or 5; 3) Analyze and process the prediction results after removing abnormal data to determine the predicted sequencing depth of the target gene.

6. A method of predicting copy number in a predetermined region, characterized by, include: a. Obtain the true sequencing depth of multiple predetermined regions of the target gene in the sample to be tested; b. Based on the sequencing depth of multiple predetermined regions of the target gene, construct a virtual sequencing depth for each predetermined region of the target gene; c. Input the sequencing depth of each predetermined region of the virtual target gene into a pre-trained machine learning model to determine the predicted sequencing depth of each predetermined region; d. Based on the predicted sequencing depth and the actual sequencing depth of each predetermined region, predict the copy number of each predetermined region; The machine learning model is obtained by training using the method described in any one of claims 1-5.

7. The method of claim 6, wherein, The predicted copy number for each predetermined region is expressed as follows: CopyNum i represents the predicted copy number of the i-th region of the target gene; depth real-i represents the real sequencing depth of the i-th region; depth predict-i represents the predicted sequencing depth of the i-th region.

8. The method according to claim 6, characterized in that, Further includes: The predicted copy number for each predetermined region is corrected.

9. The method of claim 8, wherein, The correction includes: 1) Determine the correction coefficient based on the difference in actual sequencing depth between the target gene and the pseudogene in each predetermined region; 2) Based on the correction coefficient and the predicted copy number of each predetermined region, determine the copy number of each predetermined region of the target gene.

10. The method of claim 9, wherein, The correction coefficient is expressed as: Among them, Target_Ratio i The depth represents the correction coefficient for the i-th region of the target gene. real-i-j The depth represents the sequencing depth of the j-th differentially expressed site in the i-th region of the target gene. pse-i-j This represents the sequencing depth of the j-th differential base site in the i-th region of the pseudogene.

11. The method of claim 10, wherein, The copy number of each predetermined region of the target gene is expressed as follows: CopyNum fit-i = CopyNum i * 2 * Target_Ratio i ; Among them, CopyNum fit-i This represents the copy number correction value for the i-th region of the target gene; CopyNum i Target_Ratio represents the predicted copy number of the i-th region of the target gene; Target_Ratio represents the correction coefficient for the i-th region of the target gene.

12. A copy number variation prediction method, characterized by, include: Based on the difference between the predicted copy number and the predetermined copy number threshold for each predetermined region, the change in the copy number for each predetermined region is determined; The predicted copy number for each predetermined region is obtained by the method described in any one of claims 6-11.

13. The method of claim 12, wherein, If the predicted copy number of the predetermined region is greater than the predetermined copy number threshold, the copy number of the predetermined region increases; If the predicted copy number of the predetermined region is less than the predetermined copy number threshold, the copy number of the predetermined region is reduced; Optionally, it further includes: using Z-scores to characterize the confidence level of copy number predictions; Optionally, the Z-score is represented as: Z-score = (CopyNum fit-i -average) / std; Among them, CopyNum fit-i represents the copy number correction value of the i-th region of the target gene; average represents the mean copy number of the i-th region of the training samples; std represents the standard deviation of the copy number of the i-th region of the training samples. 14.A sequencing depth prediction model establishment device, characterized in that, include: A sequencing depth acquisition unit is used to acquire the sequencing depth of multiple predetermined regions of a training sample, wherein the multiple predetermined regions are selected from different genes; The virtual gene sequencing depth construction unit constructs the sequencing depth of each predetermined region of the virtual target gene based on the sequencing depth of multiple predetermined regions of the target gene and pseudogenes in the training sample. The reference sequencing depth determination unit performs correlation analysis between the sequencing depth of each predetermined region of the virtual target gene and the sequencing depth of all predetermined regions of non-target genes and pseudogenes to determine the reference sequencing depth of each predetermined region of the virtual target gene. The training unit uses the reference sequencing depth as input features and the sequencing depth of each predetermined region of the virtual target gene as labels to train the machine learning model.

15. A system for predicting copy number in a predetermined region, comprising: include: The sequencing depth acquisition module is used to obtain the actual sequencing depth of multiple predetermined regions of the target gene in the sample to be tested; The virtual gene sequencing depth acquisition module is used to construct the virtual sequencing depth of each predetermined region of the target gene based on the sequencing depth of multiple predetermined regions of the target gene. The predicted sequencing depth determination module is used to input the sequencing depth of each predetermined region of the virtual target gene into a pre-trained machine learning model to determine the predicted sequencing depth of each predetermined region. The copy number prediction module is used to predict the copy number of each predetermined region based on the predicted sequencing depth and the actual sequencing depth of each predetermined region. The machine learning model is obtained by training the apparatus described in claim 14.

16. A copy number variation prediction device, comprising: include: The discrimination module is used to determine the change in the copy number of each predetermined region based on the difference between the predicted copy number and the predetermined copy number threshold of each predetermined region; The predicted copy number for each predetermined region is obtained using the system described in claim 15.

17. A computing device, comprising: include: Processor and memory; The memory is used to store computer programs; The processor is configured to execute the computer program to implement the sequencing depth prediction model establishment method as described in any one of claims 1-5, the predetermined region copy number prediction method as described in any one of claims 6-11, or the copy number variation prediction method as described in claim 12 or 13.