A copy number variation detection method and system based on TOS feature set learning

By extracting multiple features and combining the CatBoost classifier and sliding window strategy, the problems of insufficient accuracy and boundary precision of existing CNV detection methods in low coverage and low tumor purity data are solved, and high-precision CNV detection is achieved.

CN116935960BActive Publication Date: 2025-10-10XIDIAN UNIV HANGZHOU RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310805614.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-03
Publication Date
2025-10-10
Estimated Expiration
2043-07-03

AI Technical Summary

Technical Problem

Existing CNV detection methods based on high-throughput sequencing technology are easily affected by factors such as GC content, base quality, and alignment quality, resulting in decreased detection accuracy. In particular, the performance is unstable in low-coverage and low-tumor purity data, and the CNV boundary accuracy is insufficient.

Method used

Based on the TOS feature ensemble learning method, by extracting RD signal, GC content, base quality, alignment quality and adjacent position correlation features, seven outlier detection algorithms were used to construct TOS features. Combined with the greedy strategy and CatBoost classifier, a sliding small window strategy was designed to accurately locate CNV boundaries.

Benefits of technology

The accuracy and generalization ability of CNV detection are improved, the algorithm complexity is reduced, the CNV boundary error is ensured to be within 100bp, and the detection performance of low coverage and low tumor purity data is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116935960B_ABST
    Figure CN116935960B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of high-throughput sequencing technology of genome sequence determination, and discloses a copy number variation detection method and system based on TOS feature set learning, which comprises five features extraction of copy number variation, TOS feature construction and selection, integrated learning algorithm classification, copy number variation boundary accurate identification and algorithm performance evaluation. The algorithm performance evaluation adopts the copy number variation detection capability of the algorithm under the indexes of recall rate, accuracy and F1-score. The present application firstly selects five basic features reflecting the sequencing characteristics of copy number variation and TOS features, and uses the hypothesis testing method for feature selection, extracts non-isomorphic features representing the distribution of copy number variation, and uses integrated learning to classify whether it is a copy number variation region. In addition, the present application uses a sliding small window to accurately detect the copy number variation boundary, so that the detection result is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of high-throughput sequencing technology for genome sequence determination, and particularly relates to a copy number variation detection method and system based on TOS feature set integrated learning BACKGROUND

[0002] At present, copy number variation (CNV) widely exists in human genome, which refers to the deletion or duplication of genomic sequence at a submicroscopic level, and is proved to be the main cause of most genotypic diseases. Many biological studies have proved that CNV in genome has a huge impact on human. For example, scientists found that people who take starch as staple food usually have more copies of amylase gene than people in regions where starch is not often eaten, which shows that the increase of amylase gene copy number enables people to obtain more nutrients and energy from starch food to adapt to the living environment. When some CNV occurs improperly in genes, it will bring some diseases, such as type I diabetes, schizophrenia, neurodegenerative diseases, congenital abnormalities, cardiovascular diseases, etc. Therefore, accurately detecting the CNV-occurring fragments in genes and determining the variation type are crucial for cancer analysis and targeted drug treatment of cancer.

[0003] With the rapid development of DNA sequencing technology and molecular diagnostic technology, high-throughput sequencing technology can directly show the deletion or amplification of tumor genomic DNA, providing a higher resolution gene map for CNV analysis. At present, the commonly used CNV detection method and system based on high-throughput sequencing technology are mainly based on the read depth (RD) of sequencing data, that is, the RD signal of a certain position in the genome is proportional to the copy number of the position. Therefore, it is crucial to design a reasonable and effective CNV detection model based on high-throughput sequencing data signal.

[0004] RD-based CNV detection methods are often considered an outlier detection problem and are generally divided into two categories: CNV detection methods based on statistical distributions and those based on density or distance. One category typically detects CNVs by fitting a statistical distribution to the RD signal. For example, CNVnator uses the mean-shift method to segment genomic sequences into segments based on the RD signal and then builds a statistical test model for each genomic segment to determine whether a CNV occurs. It is not only applicable to a variety of sequencing data types but also achieves high CNV detection accuracy. ReadDepth uses a negative binomial distribution to approximate an overdispersed Poisson distribution to identify CNVs and employs a circular binary segmentation algorithm to detect CNV variant boundaries. Another category of methods primarily detect CNVs by calculating the density or distance between the RD signal in a region and the RD signal in normal regions. For example, CNV-LOF assigns a local outlier factor to the RD signal of a genomic segment and declares CNVs using a boxplot procedure. However, even slight fluctuations in the RD signal can easily affect the final CNV detection results.

[0005] Through the above analysis, the problems and defects of the existing technology are as follows:

[0006] (1) Distribution-based methods assume that the RD signal of the genome obeys a certain statistical distribution and is easily affected by the distribution of CNVs and other genomic variations, which reduces the performance of CNV detection.

[0007] (2) All these RD-based methods are susceptible to interference in the accuracy of CNV detection because the RD signal is easily affected by some factors, such as GC content, base quality, and alignment quality, which will be amplified in low-coverage sequencing data, especially when the tumor purity is low.

[0008] (3) Factors that affect the RD signal, such as GC content, base quality, and alignment quality, may be inappropriate during data preprocessing or the parameter selection may be affected by the data itself, resulting in a decrease in the performance of the CNV detection method.

[0009] (4) Some existing algorithms have low boundary accuracy and F1 scores when processing CNVs of different lengths. Summary of the Invention

[0010] In response to the problems existing in the prior art, the present invention provides a copy number variation detection method and system based on TOS feature ensemble learning.

[0011] The present invention is achieved by providing a copy number variation detection method based on TOS feature ensemble learning. The copy number variation detection method based on TOS feature ensemble learning cuts the genome sequence into 1000bp fragments based on high-throughput sequencing, and constructs five original features for CNV feature representation, including RD signal, GC content, base quality, alignment quality, and correlation between adjacent positions for each fragment; based on these original features, outlier scores are constructed from different scales using seven different outlier detection methods;

[0012] A greedy strategy is used to combine AUC and discounted precision function to extract CNV features. The CNV detection problem is regarded as a binary classification problem, that is, whether a genomic segment has CNV. Based on the extracted CNV features, a CatBoost classifier is constructed to detect CNV segments. First, the average RD signal of the normal area is determined and it is assumed that the number of copies of this part is 2. Then, based on the fact that the RD signal of a certain segment is proportional to the copy number, the copy number corresponding to the CNV is calculated and the CNV type is determined. When the copy number is greater than 2, it is copy number amplification, and when the copy number is less than 2, it is copy number loss. Finally, a sliding small window strategy is designed to solve the problem of rough CNV boundaries caused by 1000bp cutting of the genomic sequence during the early preprocessing. The error between the CNV boundary and the true variation boundary is controlled within 100bp to ensure the high precision of the CNV boundary.

[0013] Furthermore, the copy number variation detection method based on TOS feature ensemble learning includes the following steps:

[0014] In the first step, the five major feature information of CNV signal, RD value, adjacent base correlation, GC content, base quality, and alignment quality, are extracted in sequence to form the original feature space;

[0015] In the second step, seven outlier detection methods (KNN, K-Median, Avg-kNN, LOF, LoOP, One-Class SVM, and Isolation Forests) were applied to the original feature space to calculate the outlier score, respectively, to obtain the transformed outlier score (TOS) of the original feature space. A greedy strategy was used to filter the TOS features using AUC and discounted precision function, and the filtered TOS features were merged with the original feature space to form the final copy number variation feature space.

[0016] In the third step, the CatBoost algorithm is used to construct a decision tree for the copy number variation feature space obtained above to partition the feature space and determine whether copy number variation has occurred. The average RD values ​​of regions without copy number variation and regions with copy number variation are calculated and compared to determine the copy number variation type and ploidy.

[0017] In the fourth step, the obtained copy number variation may have a fuzzy boundary problem, and the true copy number variation boundary is located near the left and right sides of the fuzzy boundary.

[0018] Furthermore, the TOS feature ensemble learning-based copy number variation detection method extracts five features associated with CNV: the sequenced genomic data is preprocessed using the BWA tool and Samtools, and the genomic sequence is then divided into multiple fragments using a non-overlapping sliding window; then, for each fragment, the five major features of the average read depth, GC content, base quality, alignment quality, and correlation between adjacent base positions within the fragment are calculated to reflect the characteristic distribution of CNV and the noise information in the data preprocessing step;

[0019] The correlation calculation of adjacent base positions: Considering that the strong correlation between adjacent positions in the gene sequence has an important impact on CNV, the definition of base position correlation is described, that is, the absolute value of the difference between the RD value of the current fragment and the RD mean of the w fragments on the left and right:

[0020]

[0021] Among them, i represents the i-th fragment in the genome sequence, j represents the j-th fragment adjacent to i, w represents the number of custom left or right neighbor fragments, r i and c i represent the RD value and correlation coefficient of the i-th segment, respectively.

[0022] Furthermore, the copy number variation detection method based on TOS feature ensemble learning uses a greedy algorithm and adopts AUC and discounted precision function for feature selection, increasing feature diversity by selecting features with high precision and low correlation with the selected features; the selected TOS is regarded as a new feature that enhances the original feature space, and the new feature is combined with the original feature space to obtain a candidate CNV feature set.

[0023] Furthermore, the TOS feature ensemble learning-based copy number variation detection method uses a CatBoost classifier to divide the genome into CNV regions and normal regions. CatBoost constructs new learners using the loss functions of established weak learners and uses them to improve prediction capabilities. First, a simple base model is initialized by constructing a shallow decision tree. Then, a new decision tree is constructed by calculating the residual between the predicted value and the true value of the current model at each iteration. Finally, the previous process is repeated to complete the construction of the entire tree and obtain the CatBoost classifier.

[0024] Determine the CNV variation type and obtain candidate CNV regions: calculate the copy number of the CNV region by comparing the RD signals of the CNV region and the normal region; calculate the average of the RD signals of all normal regions, which indicates a copy number of 2; calculate the copy number of the CNV region based on the RD signal of a certain region in the genome being proportional to the copy number of the region; finally, merge regions with continuous genomic positions based on whether the copy numbers are consistent to obtain candidate CNV regions.

[0025] Furthermore, the copy number variation detection method based on TOS feature ensemble learning establishes a small sliding window strategy to accurately locate CNV boundaries: according to the non-overlapping sliding windows used in preprocessing and feature selection, the boundaries of the candidate CNV region are relatively rough; in the preprocessing and feature selection part, a large window is selected to divide the genome into multiple segments, the RD signal of each segment is calculated and features are constructed for these segments to save computing resources and smooth noise; in this way, the boundaries of the candidate CNV region are delineated by the large window; first, the boundary of the current candidate CNV region is defined as a fuzzy boundary; a small sliding window is designed within the two segments, and the position where the average RD value of the small sliding window changes significantly is found to accurately locate the boundary of the CNV; the average RD signal of the CNV region and the normal region is calculated; then, a small sliding window is moved from the leftmost window binL of the current window to the rightmost window binR of the current window; at the same time, the average RD value of each movement of the small sliding window is calculated as follows:

[0026]

[0027] where R j represents the read depth value of the sliding window, j represents the number of times the sliding window moves, s is the step size of the sliding window movement, w is the length of the sliding window, which is set to 200bp here, m is the starting base position when the sliding window slides for the first time, r i is the read depth value of the i-th base position; the sliding step is set to 100bp;

[0028] Determine whether the algorithm can achieve a high recall rate while maintaining controllable accuracy; evaluate whether the algorithm can detect CNVs with a high F1-score; evaluate whether the algorithm can accurately detect the boundaries of CNV variations; and analyze the computational complexity of the algorithm.

[0029] Another object of the present invention is to provide a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the copy number variation detection method based on TOS feature ensemble learning.

[0030] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to execute the copy number variation detection method based on TOS feature ensemble learning.

[0031] Another object of the present invention is to provide an information data processing terminal, which is used to implement the copy number variation detection method based on TOS feature ensemble learning.

[0032] Another object of the present invention is to provide a copy number variation detection system based on TOS feature ensemble learning that implements the copy number variation detection method based on TOS feature ensemble learning, wherein the copy number variation detection system based on TOS feature ensemble learning comprises:

[0033] The basic feature extraction module is used to sequentially extract the five feature information of CNV signal, namely RD value, adjacent base correlation, GC content, base quality, and alignment quality, to form the original feature space;

[0034] A feature selection module is constructed to calculate the outlier score in the original feature space using seven outlier detection methods: KNN, K-Median, Avg-kNN, LOF, LoOP, One-Class SVM, and Isolation Forests. This method obtains the outlier score TOS after the original feature space is transformed. A greedy strategy is used to filter the TOS features using AUC and discounted precision function. The filtered TOS features are then merged with the original feature space to form the final copy number variation feature space.

[0035] The copy number variation classification module is used to construct a decision tree for the copy number variation feature space obtained above using the CatBoost algorithm to partition the feature space and determine whether copy number variation has occurred; the average RD values ​​of regions without copy number variation and regions with copy number variation are respectively calculated and compared to determine the copy number variation type and ploidy;

[0036] The copy number variation boundary module is used to obtain copy number variations that may have fuzzy boundaries. The true copy number variation boundaries are located near the left and right sides of the fuzzy boundaries.

[0037] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:

[0038] First, in view of the technical problems existing in the above-mentioned prior art and the difficulty of solving the problems, we closely combine the technical solutions to be protected by the present invention and the results and data during the research and development process, and analyze in detail and in depth how the technical solutions of the present invention solve the technical problems, and some creative technical effects brought about after solving the problems. The specific description is as follows: The present invention extracts five major features from the sequencing data, namely RD signal, GC content, base quality, alignment quality and correlation of adjacent base positions in the fragment, reflecting the characteristic distribution of CNV and the noise information in the data preprocessing step. It can solve the noise bias problem in low sequencing coverage data caused by the existing method using only a single RD signal, while avoiding the influence of errors such as GC correction and alignment quality control on CNV characteristic signals.

[0039] This paper proposes a new feature to describe the correlation between adjacent base positions. This feature, combined with the structural nature of CNV, can further enrich the characteristic dimensions of CNV and help improve the accuracy of CNV detection.

[0040] The present invention uses seven different outlier detection algorithms to extract features from the original five major CNV features and constructs 107 TOS features. Compared with the existing single outlier detection algorithm, it enriches the CNV feature space from multiple angles and can improve the detection accuracy and generalization ability of CNV.

[0041] The present invention uses a greedy algorithm to select features from the extracted TOS features based on both AUC and discounted precision function, eliminating redundant features while maintaining CNV detection performance. This feature selection strategy not only reduces computational complexity but also improves CNV detection accuracy.

[0042] This paper treats CNV detection as a binary classification problem (whether a CNV has occurred) and constructs a CatBoost classifier based on the extracted TOS features and the original five features to determine whether a CNV has occurred. This avoids the problem of unstable CNV detection performance caused by distribution-based methods that assume that the genomic RD signal is affected by the distribution of other genomic variants. The CatBoost classifier uses a learning rate of 0.07, a logloss loss function, a tree depth of 8, and 5200 iterations.

[0043] This paper proposes a sliding window strategy to address the rough CNV candidate boundaries caused by using a 1000bp window to segment the genomic sequence during the initial extraction of the five major features, further refining the CNV boundaries. The sliding window length is 200bp, and the step size is 100bp.

[0044] Secondly, from the perspective of the product or as a whole, the technical effects and advantages of the technical solution to be protected by the present application are described as follows: the detection of CNV is regarded as the detection of abnormal values, and a transformed abnormal value scoring and copy number variation based on CatBoost (TOSCB-CNV) method is proposed. When detecting CNV, the present application considers five basic features of RD signal, GC content, correlation of adjacent positions, base quality and alignment quality, instead of only taking RD signal as the original feature, which not only increases the feature space and fully reflects the characteristics of sequencing data, but also avoids the performance degradation of CNV detection caused by improper processing of features such as GC content. Subsequently, based on the original features, different measurement scales of abnormal value detection methods are used to calculate multiple transformed abnormal value scores (TOS features), further enriching the features of CNV. Considering the redundancy of features and the complexity of calculation time, the present application performs AUC and discount accuracy functions to extract TOS features to ensure the accuracy and efficiency of CNV detection. Based on the above extracted TOS features and the above five original features, the present application uses a CatBoost classifier to determine whether CNV occurs. The present application pays more attention to the features of the genomic sequence itself, and can extract effective information even on low-coverage and low-tumor purity data. In addition, the present application also designs a small sliding window strategy to accurately determine the boundary of CNV.

[0045] The present application solves the problem that the existing RD signal-based method is prone to false positives and low F1-score comprehensive performance in low-coverage and low-tumor purity sequencing data; the present application extracts five features of RD signal, GC content, base quality, alignment quality and correlation of adjacent base positions within the segment, reflects the feature distribution of CNV and the noise information in the data preprocessing step, avoids the problem that the existing method is prone to noise bias or improper GC correction when processing RD signal, and thus has low CNV detection accuracy; the present application uses seven abnormal value detection algorithms to construct TOS features from different scales, solves the problem of low CNV detection generalization ability caused by one-sided feature dimension extracted by a single abnormal value detection algorithm; the present application uses a greedy algorithm to screen CNV features based on the extracted TOS features and the comprehensive AUC and discount accuracy functions, which can solve the problem of CNV accuracy decline caused by feature redundancy, and also reduce the complexity of the algorithm; the present application regards the CNV detection problem as a binary classification (whether CNV occurs) problem, and constructs a CatBoost classifier based on the extracted TOS features and the original five features to determine whether CNV occurs, which can avoid the problem of unstable CNV detection performance caused by the assumption that the distribution-based method assumes that the RD signal of the genome is affected by the distribution of other variations of the genome; the present application proposes a small sliding window strategy, which can solve the problem of inaccurate CNV boundary caused by rough division of the genome in the early stage of feature calculation, and can control the error accuracy of the CNV boundary and the real variation boundary within 100 bp.

[0046] Third, as auxiliary evidence for the inventiveness of the claims of the present invention, it is also reflected in the following important aspects:

[0047] The technical solution of the present invention solves a technical problem that people have been eager to solve but have never been able to solve: the RD signal in low-coverage sequencing data with low tumor purity is easily affected by GC content, base quality, alignment quality and other noise, resulting in errors in the statistical distribution fitted or the constructed outlier score of the existing CNV detection method based on the RD signal and poor generalization ability, resulting in high false positives or unsatisfactory comprehensive performance F1-score in the existing CNV detection method based on high-throughput sequencing. This is a technical problem that people have been eager to solve but have never been able to solve.

[0048] The present invention not only uses RD signals, but also extracts GC content, base quality, alignment quality and correlation between adjacent positions to enrich the feature representation of CNV, and at the same time contains rich information of sequencing, which can avoid the poor performance problem in sequencing data with low coverage and low tumor purity when CNV is detected only based on RD signals. On this basis, the present invention uses 7 different outlier detection methods to construct outlier scores from different scales and uses a greedy strategy to integrate AUC and discounted accuracy function to extract CNV features, while ensuring the accuracy of CNV detection. The complexity of the algorithm is reduced, and the outlier score constructed by the single outlier strategy has more generalization ability. In addition, the present invention constructs a CatBoost classifier based on CNV features, treating the CNV detection problem as a two-class problem, rather than realizing CNV detection by fitting statistical distributions or detecting outliers. The present invention has more generalization ability when detecting CNV. At the same time, CatBoost belongs to an integrated learner, which can improve the performance of the single outlier detection algorithm. Finally, the present invention designs a sliding small window strategy to solve the problem of rough CNV boundaries during early processing, and controls the error between the CNV boundary and the true variation boundary within 100bp to ensure the high precision of the CNV boundary. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 Flowchart of a copy number variation detection method based on TOS feature ensemble learning provided by an embodiment of the present invention;

[0050] Figure 2 Schematic diagram of the TOSCatBoost algorithm provided by an embodiment of the present invention;

[0051] Figure 3 This is a detection effect diagram provided by an embodiment of the present invention;

[0052] Figure 4 Schematic diagrams of actual boundary conditions 1 and 2 provided by embodiments of the present invention;

[0053] Figure 5 is a result diagram of different methods provided by the embodiment of the application on a simulation data set with a sequencing coverage of 4X;

[0054] Figure 6 is a result diagram of different methods provided by the embodiment of the application on a simulation data set with a sequencing coverage of 6X. DETAILED DESCRIPTION

[0055] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.

[0056] The present application cuts the genomic sequence into 1000bp fragments based on high-throughput sequencing, and constructs five original features for CNV feature representation for each fragment, including RD signal, GC content, base quality, alignment quality and correlation between adjacent positions. Based on the original features, seven different outlier detection methods are used to construct outlier scores from different scales to further enrich the CNV feature space. A greedy strategy is used to extract CNV features by combining AUC and discount accuracy function, which reduces the complexity of the algorithm while ensuring the accuracy of CNV detection. The CNV detection problem is regarded as a binary classification problem, i.e. whether a genomic fragment has CNV or not. Based on the extracted CNV features, a CatBoost classifier is constructed to realize the detection of CNV fragments. First, the average RD signal of the normal region is determined and the copy number of this part is assumed to be 2. Then, the CNV corresponding copy number is calculated based on the positive proportion of the RD signal of a certain fragment and the copy number, and the CNV type is determined. When the copy number is greater than 2, it is copy number amplification, and when the copy number is less than 2, it is copy number deletion. Finally, a sliding window strategy is designed to solve the rough CNV boundary problem caused by cutting the genomic sequence into 1000bp in the early preprocessing, and the CNV boundary and the real variation boundary error are controlled within 100bp, ensuring the high accuracy of the CNV boundary.

[0057] The present invention includes the extraction of five major features of copy number variation, construction and selection of TOS features, classification of ensemble learning algorithms, precise identification of copy number variation boundaries and performance evaluation of the algorithm. The performance evaluation of the algorithm uses the judgment algorithm's copy number variation detection ability under indicators such as sensitivity, specificity and F1-score. The present invention solves the problem of low accuracy in copy number variation detection under low sequencing coverage and low tumor purity scenarios. The present invention first selects the five basic features and TOS features that reflect the characteristics of copy number variation sequencing, and uses a greedy strategy method to select features, extracts non-isomorphic features that can represent the distribution of copy number variation, and uses ensemble learning to classify whether the copy number variation region is a copy number variation region. In addition, the present invention uses a small sliding window to accurately detect the copy number variation boundary, making the detection results more accurate.

[0058] like Figure 1 As shown, the copy number variation detection method based on TOS feature ensemble learning provided by the embodiment of the present invention includes the following steps:

[0059] S101: Extract five feature information reflecting CNV signals, namely RD value, adjacent base correlation, GC content, base quality, and alignment quality, to form the original feature space;

[0060] S102: Seven outlier detection methods, including KNN, K-Median, Avg-kNN, LOF, LoOP, One-Class SVM, and Isolation Forests, are applied to the original feature space to calculate the outlier score, respectively, and obtain the outlier score TOS after the original feature space is transformed. A greedy strategy is used to filter the TOS features using AUC and discounted precision function, and the filtered TOS features are merged with the original feature space to form the final copy number variation feature space.

[0061] S103: Using the CatBoost algorithm, a decision tree is constructed for the copy number variation feature space obtained above to partition the feature space and determine whether copy number variation occurs; the average RD values ​​of regions without copy number variation and regions with copy number variation are calculated and compared to determine the copy number variation type and ploidy;

[0062] S104: The acquired copy number variation may have a fuzzy boundary problem, and the true copy number variation boundary is located near the left and right sides of the fuzzy boundary. A small sliding window strategy is proposed to search for the location where the RD signal changes significantly in the two small sliding windows on the left and right of the current candidate copy number variation region, so as to control the error between the detected copy number variation boundary and the true copy number variation boundary within 100bp.

[0063] The copy number variation detection method based on TOS feature ensemble learning provided by an embodiment of the present invention includes the following steps:

[0064] Extracting five features associated with CNVs: Sequenced genomic data were preprocessed using BWA tools and Samtools. The genomic sequence was then divided into multiple segments using non-overlapping sliding windows. Five features were then calculated for each segment: average read depth (i.e., RD ​​signal), GC content, base quality, alignment quality, and correlation between adjacent base positions within the segment. These features reflect the characteristic distribution of CNVs and the noise information in the data preprocessing step.

[0065] The correlation calculation of adjacent base positions: Considering the strong correlation between adjacent positions in the gene sequence has an important impact on CNV, for example, when a fragment and its adjacent fragment have the same type of CNV, the RD values ​​of the two fragments are similar; otherwise, the RD values ​​differ greatly. This paper proposes a definition to describe the correlation of base positions, that is, the absolute value of the difference between the RD value of the current fragment and the mean RD value of the w fragments to the left and right:

[0066]

[0067] Among them, i represents the i-th fragment in the genome sequence, j represents the j-th fragment adjacent to i, w represents the number of custom left or right neighbor fragments, r i and c i represent the RD value and correlation coefficient of the i-th segment, respectively.

[0068] Constructing TOS features based on seven outlier detection algorithms: To enrich the CNV feature space, this paper employs seven basic outlier detection functions: KNN, K-Median, Avg-KNN, LOF, Loop, One-class SVM, and Isolation Forests. These functions each calculate outlier scores based on the five original features mentioned above to obtain a transformed outlier score (TOS) feature. The first three methods all calculate outliers based on distance. These three methods differ in the distance measured using different algorithms: KNN, K-Median, and Avg-KNN use mean distance, median distance, and a combination of mean and median distance, respectively. LOF and Loop perform outlier detection based on density, calculating the local density of data points to determine the degree of data anomaly. LOF is highly tolerant to data noise, while Loop has relatively low computational complexity. One-class SVM, based on a kernel function for outlier detection, constructs a hyperplane to distinguish normal from anomalous data. Isolation Forests performs outlier detection based on a tree model, constructing a tree for each data point and then determining the degree of anomaly based on the tree's depth.

[0069] Feature extraction using AUC and discounted precision: Because these TOS features may contain duplication and redundancy, which can affect CNV detection performance, the present invention uses a greedy algorithm and employs AUC and discounted precision for feature selection to control computational complexity and improve prediction accuracy. The AUC is used to select features that contribute to improved CNV detection accuracy. The discounted precision function, which considers the correlation with the selected features, is used to increase feature diversity by selecting features with high precision and low correlation with the selected features. The present invention treats the selected TOS as new features that enhance the original feature space and combines these new features with the original feature space to obtain a set of candidate CNV features.

[0070] Constructing a CatBoost classifier to detect CNVs: The present invention uses a CatBoost classifier to divide the genome into CNV regions and normal regions. CatBoost is a classification algorithm based on gradient boosting decision trees. It constructs new learners through the loss function of established weak learners and uses them to improve predictive ability. Specifically, the present invention first initializes a simple basic model by constructing a shallow decision tree, and then constructs a new decision tree by calculating the residual between the predicted value and the true value of the current model at each iteration. To ensure that the weight of each tree is reasonable, the present invention adjusts each tree through the learning rate and updates the model. Finally, the present invention repeats the previous process to complete the construction of the entire tree and obtain the CatBoost classifier. In summary, the reasons for choosing the CatBoost classifier are as follows: (i) CatBoost is faster to train on large genetic datasets, which greatly saves time. (ii) CatBoost can handle data with unbalanced distribution of normal and abnormal genes and has good performance in detecting CNVs. (iii) CatBoost is highly robust to noise and missing data in genetic data.

[0071] Determine the CNV variant type and obtain candidate CNV regions: The copy number of the CNV region is calculated by comparing the RD signals of the CNV region with those of normal regions. The present invention calculates the average RD signal of all normal regions, which indicates a copy number of 2. The present invention calculates the copy number of the CNV region based on the fact that the RD signal of a region in the genome is proportional to the copy number of that region. Finally, the present invention merges regions with consecutive genomic locations based on whether the copy numbers are consistent to obtain candidate CNV regions.

[0072] Establish a small sliding window strategy to accurately determine CNV boundaries: Based on the non-overlapping sliding windows used in preprocessing and feature selection, the boundaries of the candidate CNV region are relatively rough. In the preprocessing and feature selection part, the present invention selects a large window to divide the genome into multiple segments, calculates the RD signal of each segment and constructs features for these segments to save computing resources and smooth noise. In this way, the boundaries of the candidate CNV region are delineated by the large window. However, the actual boundaries of the CNV region may be located on both sides of the boundary position of the candidate CNV region. In order to refine the boundaries of the candidate CNV region, the present invention designs a left-right sliding small window algorithm. First, the present invention defines the boundary of the current candidate CNV region as a fuzzy boundary. Based on experience, the present invention defines the error of the fuzzy boundary within two segments (left segment: binL and right segment: binR) near the fuzzy boundary. Taking the starting position of CNV as an example, the actual starting position of CNV may be in the standard area (left segment: binL) or the variation area (right segment: binR). The present invention designs a small sliding window within the two segments and searches for the position where the average RD value of the small sliding window changes significantly to accurately locate the boundary of CNV. First, we calculate the average RD signal of the CNV region and the normal area. Then, a small sliding window moves from the left side of binL to the right side of binR. At the same time, the average RD value of each movement of the small sliding window is calculated as follows:

[0073]

[0074] where R j represents the read depth value of the sliding window, j represents the number of times the sliding window moves, s is the step size of the sliding window movement, w is the length of the sliding window, which is set to 200bp here, m is the starting base position when the sliding window slides for the first time, r i is the read depth value of the i-th base position. The sliding step size is set to 100bp.

[0075] As can be seen from the formula, when the small sliding window passes through the area where the CNV occurs, the RD value changes significantly. For example, when the small sliding window moves and does not pass through the left boundary of the true CNV, its average RD value remains stable and consistent. When part of the small sliding window crosses the true CNV boundary, its average RD value will change continuously. After the entire small sliding window passes through the true CNV boundary, its RD value tends to stabilize. Therefore, the present invention determines that the true left boundary of the CNV is the left boundary of the first small sliding window in the steady state. Similarly, the present invention can infer the true right boundary of the CNV based on this method.

[0076] Performance evaluation of the algorithm of the present invention: determine whether the algorithm achieves a high recall rate while maintaining controllable accuracy; evaluate whether the algorithm can detect CNVs with a high F1-score; evaluate whether the algorithm can accurately detect the boundaries of CNV variations; and analyze the computational complexity of the algorithm.

[0077] The copy number variation detection system based on TOS feature ensemble learning provided by an embodiment of the present invention includes:

[0078] The basic feature extraction module is used to sequentially extract the five feature information of CNV signal, namely RD value, adjacent base correlation, GC content, base quality, and alignment quality, to form the original feature space;

[0079] A feature selection module is constructed to calculate the outlier score in the original feature space using seven outlier detection methods: KNN, K-Median, Avg-kNN, LOF, LoOP, One-Class SVM, and Isolation Forests. This method obtains the outlier score TOS after the original feature space is transformed. A greedy strategy is used to filter the TOS features using AUC and discounted precision function. The filtered TOS features are then merged with the original feature space to form the final copy number variation feature space.

[0080] The copy number variation classification module is used to construct a decision tree for the copy number variation feature space obtained above using the CatBoost algorithm to partition the feature space and determine whether copy number variation has occurred; the average RD values ​​of regions without copy number variation and regions with copy number variation are respectively calculated and compared to determine the copy number variation type and ploidy;

[0081] The copy number variation boundary module uses a small sliding window strategy to accurately identify copy number variation boundaries, as the acquired copy number variation boundaries may be fuzzy. The true copy number variation boundaries are located near the left and right sides of the fuzzy boundaries. This approach allows for accurate identification of copy number variation boundaries, keeping the error between the copy number variation boundaries and the true copy number variation boundaries within 100bp.

[0082] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0083] To evaluate the performance of the present invention, it was applied to sequencing data with coverage of 4X and 6X. In each sequencing depth data, the tumor purity was set to 0.2, 0.3, 0.4, and 0.5, respectively. A total of 8 data sets, each with 50 samples, were used to calculate the recall rate and precision of each data set, and the F1-score was also calculated.

[0084] from Figure 5 and Figure 6It can be seen that in the TOSCB-CNV results proposed by the present invention, whether on the 4X or 6X datasets, the recall rate, precision and F1-score all increase with the increase of tumor purity. For example, on a dataset with a coverage depth of 4X, as the tumor purity increases from 0.2 to 0.5, the recall rate of the TOSCB-CNV proposed by the present invention increases from 0.75 to 0.88, and the F1-score increases from 0.55 to 0.7. The recall rate of the present invention is improved by 13% because high-quality information of samples with high tumor purity can be extracted, thereby obtaining more accurate CNV detection. At the same time, the recall rate and F1-score increase with the increase of the coverage depth of different tumor purity datasets. At the same time, the recall rate and F1-score increase with the increase of the coverage depth of different tumor purity datasets. With a fixed tumor purity of 0.3, the recall rates of the 4X and 6X datasets are 0.86 and 0.95, respectively. As coverage depth increased from 4X to 6X, the F1-score of the proposed TOSCB-CNV method increased from 0.68 to 0.75. This phenomenon demonstrates that as sequencing coverage depth increases, sequencing data is less affected by noise, and CNV detection performance improves. For the three metrics in the same dataset scenario, when recall and F1-score are high, precision is slightly lower. This is consistent with the definition of F1-score, which aims to balance precision and recall. The above analysis is consistent with the phenomenon of CNVs in real genomes, indicating that the experimental results are consistent with the facts.

[0085] In order to demonstrate the advantages of the present invention, the present invention is compared with the existing popular methods CNVnator, FreeC, and ReadDepth in the above datasets. Figure 5 and Figure 6The TOSCB-CNV has the highest recall rate in all data set scenarios, indicating that it has a significant advantage in recall rate among all algorithms. In the case of low coverage and tumor purity, TOSCB-CNV has a significant advantage compared with the comparative methods for detecting CNV. For example, when the coverage is 4X and the tumor purity is 0.2, the recall rate of TOSCB-CNV is 0.75. Compared with the highest recall rate of 0.57 among the three comparative methods, it is increased by 18%, which powerfully proves the high recall rate of the method. In the case of high coverage and high tumor purity, the recall rate and F1-score of TOSCB-CNV are significantly higher than FreeC and ReadDepth, but the F1-score is basically consistent with the detection effect of CNVnator. For example, when the coverage is 6X and the tumor purity is 0.5, the recall rate of TOSCB-CNV is as high as 0.94, which is significantly improved compared with FreeC, ReadDepth and CNVnator. In summary, the recall of TOSCB-CNV proposed in the application is relatively stable and less affected by tumor purity. CNVnator ranks second in recall rate, followed by ReadDepth and FreeC. In addition, TOSCB-CNV proposed in the application has the highest F1-score in 8 data sets, indicating that the method has a better balance between recall rate and precision than other compared methods, that is, better overall performance.

[0086] The above merely illustrates the specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any modification, equivalent replacement and improvement within the technical range disclosed by the present application and within the spirit and principle of the present application shall be covered within the protection scope of the present application.

Claims

1. A copy number variation detection method based on TOS feature ensemble learning, characterized in that: The copy number variation detection method based on outlier score TOS feature ensemble learning cuts the genome sequence into 1000bp fragments based on high-throughput sequencing, and constructs five original features for each fragment, including RD signal, GC content, base quality, alignment quality, and correlation between adjacent positions, which can be used to construct CNV feature representation for copy number variation detection. Based on these original features, outlier scores are constructed at different scales using seven different outlier detection methods: KNN, K-Median, Avg-kNN, LOF, LoOP, One-Class SVM, and Isolation Forests. A greedy strategy is used to combine AUC and discounted precision functions to extract CNV features. The CNV detection problem is regarded as a binary classification problem, that is, whether a genomic segment has CNV. Based on the extracted CNV features, a CatBoost classifier is constructed to detect CNV segments. First, the average RD signal of the normal region is determined and the copy number of this part is assumed to be 2. Then, based on the fact that the RD signal of a certain segment is proportional to the copy number, the copy number corresponding to the CNV is calculated and the CNV type is determined. A copy number greater than 2 indicates copy number amplification, and a copy number less than 2 indicates copy number loss. Finally, a sliding small window strategy is designed to solve the problem of rough CNV boundaries caused by 1000bp cutting of the genomic sequence during the early preprocessing. The error between the CNV boundary and the true variation boundary is controlled within 100bp, ensuring the high precision of the CNV boundary. The copy number variation detection method based on TOS feature ensemble learning establishes a small sliding window strategy to accurately locate CNV boundaries: according to the non-overlapping sliding windows used in preprocessing and feature selection, the boundaries of the candidate CNV region are relatively rough; in the preprocessing and feature selection part, a large window is selected to divide the genome into multiple fragments, the RD signal of each fragment is calculated and features are constructed for these fragments to save computing resources and smooth noise; in this way, the boundaries of the candidate CNV region are delineated by the large window; first, the boundary of the current candidate CNV region is defined as a fuzzy boundary; a small sliding window is designed within the two fragments, and the position where the average RD value of the small sliding window changes significantly is found to accurately locate the boundary of the CNV; the average RD signal of the CNV region and the normal region is calculated; then, a small sliding window is moved from the left side of binL to the right side of binR.

2. The copy number variation detection method based on TOS feature ensemble learning according to claim 1, characterized in that: The copy number variation detection method based on TOS feature ensemble learning includes the following steps: In the first step, the five major feature information of CNV signal, RD value, adjacent base correlation, GC content, base quality, and alignment quality, are extracted in sequence to form the original feature space; In the second step, seven outlier detection methods, including KNN, K-Median, Avg-kNN, LOF, LoOP, One-Class SVM, and Isolation Forests, were applied to the original feature space to calculate the outlier scores, respectively, and obtain the outlier score TOS after the original feature space was transformed. A greedy strategy was used to screen the TOS features using AUC and discounted precision function, and the screened TOS features were merged with the original feature space to form the final copy number variation feature space. In the third step, the CatBoost algorithm is used to construct a decision tree for the copy number variation feature space obtained above to partition the feature space and determine whether copy number variation has occurred. The average RD values ​​of regions without copy number variation and regions with copy number variation are calculated and compared to determine the copy number variation type and ploidy. In the fourth step, the copy number variation obtained may have a fuzzy boundary problem, and the true copy number variation boundary is located near the left and right sides of the fuzzy boundary; a small sliding window strategy is proposed to search for the location where the RD signal changes significantly in the two sliding small windows on the left and right of the current candidate copy number variation region, so as to control the error between the detected copy number variation boundary and the true copy number variation boundary within 100bp.

3. The copy number variation detection method based on TOS feature ensemble learning according to claim 2, characterized in that: The TOS feature ensemble learning-based copy number variation detection method extracts five features associated with CNVs: the sequenced genomic data is preprocessed using the BWA tool and Samtools, and the genomic sequence is then divided into multiple segments using a non-overlapping sliding window. Five features are then calculated for each segment, including the average read depth, GC content, base quality, alignment quality, and correlation between adjacent base positions within the segment, reflecting the characteristic distribution of CNVs and the noise information in the data preprocessing step. The correlation calculation of adjacent base positions: Considering that the strong correlation between adjacent positions in the gene sequence has an important impact on CNV, the definition of base position correlation is described, that is, the absolute value of the difference between the RD value of the current fragment and the RD mean of the w fragments on the left and right: , Where i represents the i-th fragment in the genome sequence, j represents the j-th fragment adjacent to i, w represents the number of custom left or right neighbor fragments, r i and c i represent the RD value and correlation coefficient of the i-th segment, respectively.

4. The copy number variation detection method based on TOS feature ensemble learning according to claim 2, characterized in that: The copy number variation detection method based on TOS feature ensemble learning uses a greedy algorithm and adopts AUC and discounted precision function for feature selection, increasing feature diversity by selecting features with high precision and low correlation with the selected features; the selected TOS is regarded as a new feature to enhance the original feature space, and the new feature is combined with the original feature space to obtain a candidate CNV feature set.

5. The copy number variation detection method based on TOS feature ensemble learning according to claim 2, characterized in that: The TOS feature ensemble learning-based copy number variation detection method uses a CatBoost classifier to divide the genome into CNV regions and normal regions. CatBoost constructs new learners using the loss functions of established weak learners and uses them to improve prediction capabilities. First, a simple base model is initialized by constructing a shallow decision tree. Then, a new decision tree is constructed by calculating the residual between the predicted value and the true value at each iteration. Finally, the previous process is repeated to complete the entire tree construction and obtain the CatBoost classifier. Determine the CNV variation type and obtain candidate CNV regions: calculate the copy number of the CNV region by comparing the RD signals of the CNV region and the normal region; calculate the average of the RD signals of all normal regions, which indicates a copy number of 2; calculate the copy number of the CNV region based on the RD signal of a certain region in the genome being proportional to the copy number of the region; finally, merge regions with continuous genomic positions based on whether the copy numbers are consistent to obtain candidate CNV regions.

6. The copy number variation detection method based on TOS feature ensemble learning according to claim 2, characterized in that: The average RD value of each movement of the small sliding window is calculated as follows: , where R j represents the read depth value of the sliding window, j represents the number of times the sliding window moves, s is the step size of the sliding window movement, w is the length of the sliding window, which is set to 200bp here, m is the starting base position when the sliding window slides for the first time, r i is the read depth value of the i-th base position; the sliding step is set to 100bp; Determine whether the algorithm can achieve a high recall rate while maintaining controllable accuracy; evaluate whether the algorithm can detect CNVs with a high F1-score; evaluate whether the algorithm can accurately detect the boundaries of CNV variations; and analyze the computational complexity of the algorithm.

7. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the copy number variation detection method based on TOS feature ensemble learning according to any one of claims 1 to 6.

8. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the copy number variation detection method based on TOS feature ensemble learning according to any one of claims 1 to 6.

9. An information data processing terminal, characterized in that: The information data processing terminal is used to implement the copy number variation detection method based on TOS feature ensemble learning as described in any one of claims 1 to 6.

10. A copy number variation detection system based on TOS feature ensemble learning that implements the copy number variation detection method based on TOS feature ensemble learning according to any one of claims 1 to 6, characterized in that: The copy number variation detection system based on TOS feature ensemble learning includes: The basic feature extraction module is used to sequentially extract the five feature information of CNV signal, namely RD value, adjacent base correlation, GC content, base quality, and alignment quality, to form the original feature space; A feature selection module is constructed to calculate the outlier score in the original feature space using seven outlier detection methods: KNN, K-Median, Avg-kNN, LOF, LoOP, One-Class SVM, and Isolation Forests. This method obtains the outlier score TOS after the original feature space is transformed. A greedy strategy is used to filter the TOS features using AUC and discounted precision function. The filtered TOS features are then merged with the original feature space to form the final copy number variation feature space. The copy number variation classification module is used to construct a decision tree for the copy number variation feature space obtained above using the CatBoost algorithm to partition the feature space and determine whether copy number variation has occurred; the average RD values ​​of regions without copy number variation and regions with copy number variation are respectively calculated and compared to determine the copy number variation type and ploidy; The copy number variation boundary module is used to obtain copy number variations that may have fuzzy boundaries. The true copy number variation boundary is located near the left and right sides of the fuzzy boundary. A small sliding window strategy is proposed to search for locations where significant changes in RD signals occur in two small sliding windows on the left and right sides of the current candidate copy number variation region, thereby controlling the error between the detected copy number variation boundary and the true copy number variation boundary within 100bp.