Classification method and apparatus, electronic device, and storage medium
By extracting and standardizing the number of amplicones from methylation-specific amplification sequencing data, calculating the average entropy, and inputting it into a classification model, the problem of insufficient analysis in existing methylation amplicon sequencing platforms is solved, thereby improving the accuracy and efficiency of tumor detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANCHORDX MEDICAL CO LTD
- Filing Date
- 2022-06-29
- Publication Date
- 2026-04-14
AI Technical Summary
Existing methylation amplicon sequencing platforms lack effective data analysis methods, resulting in insufficient analysis of batch effects and methylation patterns, which affects the accuracy of predictive classification in tumor detection.
The methylation patterns of each target amplification region and the number of amplicones in the internal reference amplification region are extracted from the sequencing sample data of methylation-specific amplification sequencing. After standardization, the average entropy is calculated and input into a trained one-dimensional or multi-dimensional classification model for classification.
It improves the accuracy, sensitivity, and specificity of early tumor screening classification, and enables rapid and accurate classification of sequencing sample data.
Smart Images

Figure CN117393057B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of bioinformatics analysis engineering technology, and in particular to a classification method and apparatus, electronic equipment and storage medium. Background Technology
[0002] DNA methylation is one of the earliest discovered epigenetic modification pathways. Numerous studies have shown that DNA methylation can alter chromatin structure, DNA conformation, DNA stability, and the way DNA interacts with proteins, thereby controlling gene expression. DNA methylation plays a crucial role in normal human development, X chromosome inactivation, aging, and the development of many human diseases (such as developmental malformations, cancer, cardiovascular disease, diabetes, and neuropsychiatric disorders). DNA methylation typically inhibits gene expression, while demethylation induces gene reactivation and expression. Bisulfite can convert unmethylated cytosine (C) to uracil (U), while leaving methylated cytosine unchanged. After DNA is treated with bisulfite, methylation-specific PCR amplification primers can be designed for specific gene sequence regions of interest. Amplification of the target sequence region followed by NGS sequencing can yield single-base methylation pattern resolution. Then, bioinformatics analysis methods are used to quantitatively analyze the methylation signal in the target region to study the correlation between gene methylation and disease. In our case, it is mainly applied to the early detection of tumors, but is not limited to early tumor screening.
[0003] Currently, methylation-specific amplification sequencing (MS / MS) has begun to be applied in scientific research and clinical testing. Compared with single qPCR, this method has many advantages, such as accurate quantification, obtaining the base content of each amplicon (sequencing reads), which is of great significance for achieving single-base resolution of methylation patterns. It also provides a large amount of DNA methylation status information for subsequent bioinformatics analysis. Utilizing and mining this methylation data can improve the accuracy of tumor detection and classification, reduce cases of misdiagnosis and missed diagnosis, and make early tumor screening more widely used in auxiliary clinical diagnosis.
[0004] However, there is currently no effective data analysis method for methylation amplicon sequencing platforms. Most analyses have some problems, such as batch effects and lack of in-depth analysis of methylation patterns. Therefore, how to extract useful information for tumor prediction and classification from massive methylation amplicon sequencing data (sequencing reads) and improve the accuracy of case classification has become a key issue. Summary of the Invention
[0005] According to one aspect of this disclosure, a classification method is provided, the method comprising:
[0006] The number of amplicones for various methylation patterns in each target amplification region and the number of amplicones in each internal reference amplification region are extracted from the sequencing sample data of methylation-specific amplification sequencing. The number of amplicones represents the number of reads corresponding to the reference genome sequence in the amplification region.
[0007] At least one target internal control amplification region is determined from each internal control amplification region, and the number of amplons in each target amplification region is standardized using the number of amplons in the at least one target internal control amplification region to obtain the standardized number of amplons for each target amplification region for various methylation modes.
[0008] The average entropy of each target amplification region is calculated based on the proportion of amplicon numbers for each methylation mode in each target amplification region in each sequencing sample data.
[0009] The average entropy of each target amplification region, the number of standardized amplicones for the preset methylation pattern of each target amplification region, and the number of standardized amplicones for other methylation patterns besides the preset methylation pattern of each target amplification region are respectively input into the corresponding trained single-dimensional classification model, and / or all are input into the trained multi-dimensional classification model to obtain the classification result of the sequencing sample data.
[0010] In one possible implementation, extracting the number of amplicones for various methylation patterns in each target amplification region from the sequencing sample data of methylation-specific amplification sequencing includes:
[0011] The sequencing sample data is filtered to obtain the first set of reads;
[0012] Remove primer dimer sequences formed between primers in the first set of reads to obtain the second set of reads;
[0013] The reference genome sequence is used to match the reads in the second set of reads to determine the third set of reads;
[0014] The number of amplicones for each methylation pattern is determined based on the third set of reads.
[0015] In one possible implementation, the target internal reference amplification region is at least one of the L internal reference amplification regions with the highest gene stability among a plurality of internal reference amplification regions, where L is a positive integer.
[0016] In one possible implementation, determining at least one target internal control amplification region from the various internal control amplification regions includes:
[0017] Determine the average coefficient of variation of each internal control amplification region relative to the other internal control amplification regions;
[0018] Identify at least one of the L internal reference amplification regions with the smallest average coefficient of variation as the target internal reference amplification region.
[0019] In one possible implementation, determining the average coefficient of variation of each internal control amplification region relative to other internal control amplification regions includes:
[0020] Determine the ratio of the number of amplicones in the first internal control amplification region to the number of amplicones in each second internal control amplification region in each sequencing sample data.
[0021] The coefficients of variation of the first internal reference amplification region relative to each second internal reference amplification region are determined based on the ratios.
[0022] The average coefficient of variation of the first internal reference amplification region is determined based on the various coefficients of variation.
[0023] In one possible implementation, the standardization of the amplicon count of each target amplification region using the amplicon count of the at least one target internal reference amplification region includes:
[0024] Determine the geometric mean of the number of amplicones in each target intrinsic reference amplification region in each sample;
[0025] The normalized number of amplicones for each methylation mode in the target internal reference amplification region is obtained by dividing the number of amplicones for each methylation mode in the target internal reference amplification region by the geometric mean.
[0026] In one possible implementation, calculating the average entropy of each target amplification region based on the proportion of amplicon numbers of each methylation mode in each target amplification region in each sequencing sample data includes:
[0027] The average entropy of the target amplification region is calculated using the following formula:
[0028]
[0029] Where Marker_mEntropy represents the average entropy of the target amplification region, N represents the number of methylation patterns in the target amplification region, P_i represents the number of normalized amplicones of the i-th methylation pattern in the target amplification region, and i≤N.
[0030] In one possible implementation, the sum of the percentages of amplicon numbers for various methylation modes in each target amplification region is 1.
[0031] In one possible implementation, the preset methylation mode includes the fullC methylation mode.
[0032] According to one aspect of this disclosure, a sorting apparatus is provided, the apparatus comprising:
[0033] The extraction module is used to extract the number of amplicons of various methylation patterns in each target amplification region and the number of amplicons in each internal reference amplification region from the sequencing sample data of methylation-specific amplification sequencing. The number of amplicons represents the number of reads of the reference genome sequence that are aligned to the amplification region.
[0034] The first determining module is used to determine at least one target internal reference amplification region from each internal reference amplification region, and to standardize the number of amplons in each target amplification region using the number of amplons in the at least one target internal reference amplification region, so as to obtain the standardized number of amplons for each target amplification region for various methylation modes.
[0035] The second determining module is used to calculate the average entropy of each target amplification region based on the proportion of amplicon numbers of each methylation mode in each target amplification region in each sequencing sample data.
[0036] The classification module is used to input the average entropy of each target amplification region, the number of standardized amplicones of the preset methylation pattern of each target amplification region, and the number of standardized amplicones of other methylation patterns besides the preset methylation pattern of each target amplification region into the corresponding trained single-dimensional classification model, and / or input them into the trained multi-dimensional classification model to obtain the classification result of the sequencing sample data.
[0037] According to one aspect of this disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to perform the method described above.
[0038] According to one aspect of this disclosure, a computer-readable storage medium is provided that stores computer program instructions thereon, which, when executed by a processor, implement the above-described method.
[0039] This embodiment of the present disclosure extracts the number of amplicons for various methylation patterns in each target amplification region and the number of amplicons in each internal control amplification region from the sequencing sample data of methylation-specific amplification sequencing. At least one target internal control amplification region is determined from each internal control amplification region, and the number of amplicons in each target amplification region is standardized using the number of amplicons in the at least one target internal control amplification region to obtain the standardized number of amplicons for each methylation pattern in each target amplification region. Based on the proportion of amplicons for each methylation pattern in each target amplification region in each sequencing sample data, the average entropy of each target amplification region is calculated. The average entropy of each target amplification region, the standardized number of amplicons for the preset methylation pattern in each target amplification region, and the standardized number of amplicons for other methylation patterns in each target amplification region are respectively input into the corresponding trained one-dimensional classification model, and / or all are input into the trained multi-dimensional classification model, thereby quickly and accurately obtaining the classification result of the sequencing sample data.
[0040] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the specification, serve to illustrate the technical solutions of this disclosure.
[0042] Figure 1 A flowchart of a classification method according to an embodiment of the present disclosure is shown.
[0043] Figure 2 A flowchart of a classification method according to an embodiment of the present disclosure is shown.
[0044] Figure 3a , Figure 3b A schematic diagram illustrating the stability assessment of the internal reference amplification region according to an embodiment of this disclosure is shown.
[0045] Figure 4 The box plot shows a prediction classification using a unidimensional classification model based on the amplified region average entropy algorithm.
[0046] Figure 5 A box plot is shown for predictive classification using a unidimensional classification model based on a specific methylation pattern.
[0047] Figure 6 Box plots are shown for predictive classification using a unidimensional classification model based on all methylation patterns except for a specific methylation pattern.
[0048] Figure 7 The diagram shows a box plot illustrating classification prediction using a multidimensional classification model.
[0049] Figure 8 A schematic diagram illustrating the AUC comparison analysis of multi-dimensional algorithm modeling and prediction in embodiments of this disclosure is shown.
[0050] Figure 9 A block diagram of a sorting apparatus according to an embodiment of the present disclosure is shown.
[0051] Figure 10 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0052] Figure 11 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation
[0053] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0054] In the description of this disclosure, it should be understood that the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this disclosure and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this disclosure.
[0055] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise expressly specified.
[0056] In this disclosure, unless otherwise expressly specified and limited, the terms "installation," "connection," "linking," "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this disclosure according to the specific circumstances.
[0057] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.
[0058] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0059] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0060] Please see Figure 1 , Figure 1 A flowchart of a classification method according to an embodiment of the present disclosure is shown.
[0061] like Figure 1 As shown, the method includes:
[0062] Step S11: Extract the number of amplicones for various methylation patterns in each target amplification region and the number of amplicones in each internal reference amplification region from the sequencing sample data of methylation-specific amplification sequencing. The number of amplicones represents the number of reads corresponding to the reference genome sequence aligned to the amplification region.
[0063] Step S12: Determine at least one target internal reference amplification region from each internal reference amplification region, and use the number of amplons in the at least one target internal reference amplification region to standardize the number of amplons in each target amplification region to obtain the standardized number of amplons for each methylation mode in each target amplification region.
[0064] Step S13: Calculate the average entropy of each target amplification region based on the proportion of amplicon numbers of each methylation mode in each target amplification region in each sequencing sample data.
[0065] Step S14: The average entropy of each target amplification region, the number of standardized amplicones of the preset methylation mode of each target amplification region, and the number of standardized amplicones of other methylation modes besides the preset methylation mode of each target amplification region are respectively input into the corresponding trained single-dimensional classification model, and / or all are input into the trained multi-dimensional classification model to obtain the classification result of the sequencing sample data.
[0066] This disclosure discloses embodiments that extract the number of amplicons for various methylation patterns in each target amplification region and the number of amplicons in each internal control amplification region from methylation-specific amplification sequencing sample data. At least one target internal control amplification region is determined from each internal control amplification region, and the number of amplicons in each target amplification region is standardized using the number of amplicons in the at least one target internal control amplification region to obtain the standardized number of amplicons for each methylation pattern in each target amplification region. Based on the proportion of amplicons for each methylation pattern in each target amplification region in each sequencing sample data, the average entropy of each target amplification region is calculated. The average entropy of each target amplification region, the standardized number of amplicons for the preset methylation pattern in each target amplification region, and the standardized number of amplicons for other methylation patterns in each target amplification region are input into corresponding trained one-dimensional classification models, and / or all are input into trained multi-dimensional classification models, to quickly and accurately obtain the classification results of the sequencing sample data. For example, this can improve the accuracy, sensitivity, and specificity of early cancer screening classification prediction.
[0067] The classification method of the embodiments of this disclosure can be applied to processing components, or electronic devices such as terminals and servers that include processing components.
[0068] In one example, the processing component includes, but is not limited to, a single processor, discrete components, or a combination of processors and discrete components. The processor may include a controller in an electronic device with instruction execution capabilities. The processor may be implemented in any suitable manner, for example, by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. Of course, the processing component may also be an artificial intelligence processor (IPU) performing artificial intelligence operations. Artificial intelligence operations may include machine learning operations, neuromorphic operations, etc. Machine learning operations include neural network operations, k-means operations, support vector machine operations, etc. The artificial intelligence processor may, for example, include one or a combination of GPU (Graphics Processing Unit), NPU (Neural-Network Processing Unit), DSP (Digital Signal Processing Unit), and FPGA chips. This disclosure does not limit the specific type of processor. Inside the processor, the executable instructions can be executed through hardware circuits such as logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers.
[0069] In one example, a terminal, also known as User Equipment (UE), Mobile Station (MS), or Mobile Terminal (MT), is a device that provides voice and / or data connectivity to a user. Examples include handheld devices and in-vehicle devices with wireless connectivity. Currently, some examples of terminals include: mobile phones, tablets, laptops, PDAs, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, and wireless terminals in the Internet of Vehicles (IoV).
[0070] In one example, the processing component can invoke instructions stored in the storage module to implement the classification method.
[0071] In one example, a storage module may include a computer-readable storage medium, which can be a tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), programmable read-only memory (PROM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage medium as used herein is not to be construed as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0072] Please see Figure 2 , Figure 2 A flowchart of a classification method according to an embodiment of the present disclosure is shown.
[0073] In one possible implementation, step S11, extracting the number of amplicones for various methylation patterns in each target amplification region from the sequencing sample data of methylation-specific amplification sequencing, may include:
[0074] Step S111: Filter the sequencing sample data to obtain the first set of reads;
[0075] This disclosure does not limit the specific operation and implementation of "filtering". Those skilled in the art can choose appropriate technical means to implement it according to actual conditions and needs. In one example, this disclosure can first filter low-quality (e.g., low QC, short length, too many N, etc.) reads by using fastp (or other filtering tools), and then remove the adapter and PolyA / T at both ends of the reads to obtain clean reads sequence (first reads set). Here, fastp is a filtering tool, and reads (read length) refers to the base sequence obtained by sequencing using a sequencer.
[0076] Step S112: Remove the primer dimer sequences formed between primers in the first reads set to obtain the second reads set;
[0077] Primers can be macromolecules with specific nucleotide sequences that are stimulated to synthesize at the beginning of nucleotide polymerization and are linked to reactants by hydrogen bonds.
[0078] Step S113: Use the reference genome sequence to match the reads in the second set of reads to determine the third set of reads;
[0079] This disclosure does not limit the specific implementation of "matching". Those skilled in the art can choose appropriate technical means to implement it according to actual conditions and needs. In one example, this disclosure embodiment can use bismark (or other methylation alignment tools) to align the clean reads in these second reads sets with the corresponding positions of the genome hg19 (reference genome sequence) to obtain all amplicon reads data (third reads set) for each sample. Here, bismark is a methylation alignment tool.
[0080] Step S114: Determine the number of amplicon for each methylation mode based on the third set of reads.
[0081] In one example, embodiments of this disclosure may extract reads aligned to each marker based on the target amplification region (marker) and calculate the total number of amplicon readscount and the number of readscount supporting each methylation pattern.
[0082] This embodiment of the present disclosure filters the sequencing sample data to obtain a first set of reads, removes primer dimer sequences formed between primers in the first set of reads to obtain a second set of reads, matches the reads in the second set of reads using the reference genome sequence to determine a third set of reads, and determines the number of amplicones for each methylation pattern based on the third set of reads, so as to quickly and efficiently extract the number of amplicones for each methylation pattern in each target amplification region from the sequencing sample data of methylation-specific amplification sequencing.
[0083] This disclosure does not limit the specific implementation of step S12, which determines at least one target internal reference amplification region from each internal reference amplification region. Those skilled in the art can choose appropriate technical means to implement it according to actual conditions and needs. For example, in one possible implementation, the target internal reference amplification region is at least one of the L internal reference amplification regions with the highest gene stability among multiple internal reference amplification regions, where L is a positive integer.
[0084] In one possible implementation, such as Figure 2 As shown, step S12, which determines at least one target internal control amplification region from each internal control amplification region, may include:
[0085] Step S121: Determine the average coefficient of variation (CV) of each internal control amplification region relative to other internal control amplification regions;
[0086] Step S122: Determine at least one of the L internal reference amplification regions with the smallest average coefficient of variation as the target internal reference amplification region.
[0087] This disclosure embodiment determines the average coefficient of variation of each internal reference amplification region relative to other internal reference amplification regions, and can determine at least one of the L internal reference amplification regions with the smallest average coefficient of variation as the target internal reference amplification region.
[0088] This disclosure selects several relatively stable internal control amplification regions from multiple batches for normalization of target amplification region markers, thereby reducing bias generated by the internal control amplification regions themselves. Since most internal control genes are stable, the ratios of pairwise comparisons of internal controls are stable across multiple samples, i.e., essentially equal. This disclosure uses the coefficient of variation (CV) to measure the stability of these ratios. Finally, the CV is averaged to determine the average coefficient of variation of each internal control amplification region relative to other internal control amplification regions, serving as an indicator for evaluating the stability of the corresponding internal control marker. An exemplary description follows.
[0089] In one possible implementation, such as Figure 2 As shown, step S121, determining the average coefficient of variation of each internal control amplification region relative to other internal control amplification regions, may include:
[0090] Step S1211: Determine the ratio of the number of amplicon in the first internal reference amplification region to the number of amplicon in each second internal reference amplification region in each sequencing sample data.
[0091] In one example, embodiments of this disclosure can extract a readcounts matrix of amplicon counts for K internal reference amplification regions in M samples based on sequencing sample data from multiple batches. In the readcounts matrix, rows can represent sample IDs, and columns can represent the IDs of the internal reference amplification regions. For example, this readcounts matrix can be labeled as ref_marker_KMmatrix, and the corresponding row vectors in the readcounts matrix can be labeled as KV. i The column vector is labeled MV i , where K, M, and i are all integers greater than 0.
[0092] In one example, embodiments of this disclosure may iterate through and calculate the ratio vector Ratios of pairwise intrinsic parameter amplification region markers, denoted as R. ij As shown in Formula 1.
[0093] R ij =KV i / KV j Formula 1
[0094] In Formula 1, i and j are both integers greater than 0, representing the numbers of the internal reference amplification regions.
[0095] In one possible implementation, such as Figure 2 As shown, step S121, determining the average coefficient of variation of each internal control amplification region relative to other internal control amplification regions, may include:
[0096] Step S1212: Determine the coefficients of variation of the first internal reference amplification region relative to each of the second internal reference amplification regions based on each ratio.
[0097] This disclosure does not limit the specific implementation method for determining the coefficient of variation. Those skilled in the art can select appropriate techniques according to actual conditions and needs to determine the coefficient of variation of the first internal reference amplification region relative to each of the second internal reference amplification regions based on each ratio. In one example, the coefficient of variation CV of the first internal reference amplification region marker i relative to the second internal reference amplification region marker j is calculated and denoted as cv. ij As shown in Formula 2.
[0098] cv ij =CV(R) ij ) Formula 2
[0099] In one possible implementation, such as Figure 2 As shown, step S121, determining the average coefficient of variation of each internal control amplification region relative to other internal control amplification regions, may include:
[0100] Step S1213: Determine the average coefficient of variation of the first internal reference amplification region based on each coefficient of variation.
[0101] In one example, given the coefficients of variation for each internal control amplification region, this embodiment of the present disclosure can determine the average coefficient of variation of the first internal control amplification region based on these coefficients. For instance, the average coefficient of variation Mcv of the internal control amplification region marker i can be determined according to Formula 3. i .
[0102]
[0103] Among them, cv ij J represents the coefficient of variation of the first internal control amplification region marker i relative to the second internal control amplification region marker j, and J represents the total number of internal control amplification regions used.
[0104] Using the above methods, the embodiments of this disclosure can determine the ratio of the number of amplicones in the first internal reference amplification region to the number of amplicones in each second internal reference amplification region in each sequencing sample data, and determine the coefficient of variation of each first internal reference amplification region relative to each second internal reference amplification region based on each ratio, thereby quickly and accurately determining the average coefficient of variation of the first internal reference amplification region based on each coefficient of variation.
[0105] In one example, embodiments of this disclosure can determine the average coefficient of variation of each internal reference amplification region according to the above method, and determine the size of L according to the actual situation and needs, taking at least one of the L internal reference amplification regions with the smallest average coefficient of variation as the target internal reference amplification region.
[0106] Please see Figure 3a , Figure 3b , Figure 3a , Figure 3b A schematic diagram illustrating the stability assessment of the internal reference amplification region according to an embodiment of this disclosure is shown.
[0107] This disclosure uses gastric cancer datasets (e.g., including 40 cases) and colorectal cancer datasets (e.g., including 42 cases) to evaluate the stability of five internal reference markers. Figure 3a This is a schematic diagram of the assessment corresponding to the gastric cancer dataset. Figure 3b This is a schematic diagram of the evaluation corresponding to the colorectal cancer dataset.
[0108] In one example Figure 3a , Figure 3b The horizontal axis represents five internal reference markers to be evaluated, namely ACTB, KKBBD4, LEKHF1, SYT10, and EPHA3 (in other embodiments, other numbers and types of internal reference markers may also be used). The horizontal axis indicates that the stability gradually increases from left to right. The vertical axis is the average value of expression stability, that is, the average value of the coefficient of variation. The larger this value is, the more unstable the internal reference gene is.
[0109] In one example, such as Figure 3a , Figure 3b As shown, EPHA3 and SYT10 are the two most stable internal reference markers. Therefore, the embodiments of this disclosure can identify internal reference marker EPHA3 and / or internal reference marker SYT10 as the target internal reference amplification region.
[0110] In one example, such as Figure 3a , Figure 3b As shown, the stability scores (Mcv) of the five intrinsic parameter markers show a consistent trend across the two datasets, demonstrating the reproducibility of the evaluation method. This means that the scheme for evaluating the stability of the intrinsic parameter amplification region using the average coefficient of variation in this embodiment can be applied to various datasets and has broad adaptability.
[0111] In one example, given the target intrinsic reference amplification region, this embodiment of the present disclosure can perform prediction classification based on the determined target intrinsic reference amplification region. For example, as shown in Tables 1 and 2, this embodiment of the present disclosure compares the prediction classification results of the two determined target intrinsic reference amplification regions and five intrinsic reference amplification regions on the gastric cancer dataset and the colorectal cancer dataset.
[0112] Number of Markers AUC_AllRefMarks AUC_2refMarkers 10 0.75 0.8 20 0.76 0.8 50 0.8 0.8
[0113] Table 1
[0114] Number of Markers AUC_AllRefMarks AUC_2refMarkers 10 0.72873 0.804762 20 0.770159 0.8 50 0.784444 0.809048
[0115] Table 2
[0116] In Tables 1 and 2, the number of markers represents the number of target amplification region markers used in the modeling and prediction. This embodiment uses 10, 20, and 50 markers for modeling, respectively. AUC_AllRefMarks represents the AUC obtained by normalizing the target marker using all intrinsic amplification region markers (such as the aforementioned ACTB, KKBBD4, LEKHF1, SYT10, and EPHA3) and then modeling. AUC stands for Area Undercurve, and it's the area under the ROC curve, a performance metric for the learner; a value closer to 1 is better. AUC_2refMarkers represents the AUC obtained by normalizing the target marker using two target intrinsic markers (EPHA3 and SYT10) selected using the aforementioned method and then modeling. As can be seen from Tables 1 and 2, under various marker counts, the AUC obtained by normalizing the target marker using the two most stable target intrinsic markers selected using the above method is better than the AUC obtained by normalizing using all intrinsic markers.
[0117] In one possible implementation, such as Figure 2 As shown, step S12, which standardizes the number of amplicones in each target amplification region using the number of amplicones in at least one target internal reference amplification region, may include:
[0118] Step S124: Determine the geometric mean of the number of amplicones in each target internal reference amplification region in each sample;
[0119] Step S125: Divide the number of amplicones for each methylation mode in the target internal reference amplification region by the geometric mean to obtain the normalized number of amplicones for each methylation mode in the target internal reference amplification region.
[0120] Using the above method, the embodiments of this disclosure can determine the geometric mean of the number of amplicons in each target internal reference amplification region in each sample, and divide the number of amplicons in each methylation mode of the target internal reference amplification region by the geometric mean to obtain the normalized number of amplicons in each methylation mode of the target internal reference amplification region.
[0121] By standardizing the number of amplicons in each target amplification region using the number of amplicons in at least one target internal reference amplification region, this embodiment of the disclosure can obtain an amplification region methylation pattern matrix with columns representing marker IDs and behaviors representing methylation pattern types. The standardized number of amplicons in the amplification region methylation pattern matrix is a value between 0 and 1 after standardization n. This value represents the methylation degree of the marker. For each target amplification region marker (each column), this value is summed to 1 (the sum of the proportions of amplicon numbers for various methylation patterns in each target amplification region is 1). In one example, the standardized number of amplicons can also represent the frequency of occurrence of the corresponding methylation pattern in each target amplification region marker.
[0122] In one possible implementation, step S13, calculating the average entropy of each target amplification region based on the proportion of amplicon numbers of each methylation mode in each target amplification region of each sequencing sample data, may include:
[0123] The average entropy of the target amplification region is calculated using the following formula 4:
[0124]
[0125] Where Marker_mEntropy represents the average entropy of the target amplification region, N represents the number of methylation modes in the target amplification region, P_i represents the proportion of amplicon numbers in the i-th methylation mode of the target amplification region, and i≤N.
[0126] For example, the proportion of the number of amplicon in the i-th methylation mode of the target amplification region can be the proportion of the number of amplicon in the i-th methylation mode to the total number of amplicon in the target amplification region, that is, the normalized number of amplicon in the i-th methylation mode of the target amplification region.
[0127] In one example, embodiments of this disclosure can first filter out patterns whose read count is less than 1, and then calculate the average entropy of each target amplification region based on the proportion of the number of amplicones of each methylation mode in each target amplification region in each sequencing sample data. In this way, embodiments of this disclosure can remove false positive patterns caused by sequencing errors, thereby improving the accuracy of the calculation.
[0128] In one possible implementation, the preset methylation mode includes the fullC methylation mode.
[0129] The following provides an exemplary description of possible implementations for determining preset methylation modes, including the fullC methylation mode.
[0130] In one example, based on the amplified region methylation pattern matrix (marker pattern matrix) obtained in step S2, this embodiment of the disclosure uses each pattern in each marker as a feature, labeled as marker_pattern. Then, normalization is performed on each marker_pattern, and its value is labeled as marker_pattern_value. This results in a large number of feature values. To avoid overfitting in the modeling, this embodiment of the disclosure uses a feature selection method to screen the optimal combination of marker features and finds that the feature fullC pattern has better predictive ability than other patterns, which is consistent with the biological characteristics of methylation. Methods for selecting suitable patterns may include, for example, the Sequential Floating Forward Selection (SFFS) algorithm. For instance, it can start with an empty set, and in each round, select a subset x from the unselected features so that the evaluation function is optimal after adding subset x. Then, select a subset z from the selected features so that the evaluation function is optimal after removing subset z. Of course, the above description of the SFFS algorithm is exemplary. For a more detailed introduction to the SFFS algorithm, please refer to relevant technologies; it will not be elaborated upon here. Furthermore, this disclosure does not limit the use of other selection algorithms. In other embodiments, other algorithms can also be used to select suitable patterns.
[0131] This disclosure allows for the pre-establishment and training of unidimensional and multidimensional classification models. The average entropy of each target amplification region, the number of standardized amplicones for each target amplification region based on a preset methylation pattern, and the number of standardized amplicones for other methylation patterns besides the preset methylation pattern in each target amplification region are respectively input into the corresponding trained unidimensional classification model, and / or all are input into the trained multidimensional classification model to obtain the classification result of the sequencing sample data. This disclosure does not limit the technical means used for establishing and training unidimensional and multidimensional classification models. Those skilled in the art can use appropriate technical means to model and train according to actual conditions and needs.
[0132] In one possible implementation, the unidimensional classification model of this disclosure may include a unidimensional classification model based on the amplified region average entropy algorithm (marker_mEntropy algorithm, step S13). This model considers the diversity of methylation patterns on an amplified region marker. Shannon entropy is calculated using the frequency of occurrence of each pattern to measure the degree of disorder in marker methylation patterns. In this model, the disorder is higher in cancer patients than in healthy individuals, indicating greater diversity and complexity.
[0133] Please see Figure 4 , Figure 4 The box plot shows a prediction classification using a one-dimensional classification model based on the amplified region average entropy algorithm.
[0134] Figure 4 The black (dark) and gray (light) line areas on the left represent the healthy group and the cancer patient group, respectively, with the vertical axis representing probability. Figure 4 As can be seen, the embodiments of this disclosure can accurately distinguish between healthy people and cancer patients by utilizing the characteristics of the average entropy dimension.
[0135] Please refer to Table 3, which illustrates the performance of predictive classification using a single-dimensional classification model based on the amplified region average entropy algorithm.
[0136]
[0137] Table 3
[0138] In Table 3, Marker_Count represents the number of markers used in machine learning modeling. For example, top_10markers means modeling using the top 10 best markers, and top_20markers means modeling using the top 20 best markers. As shown in Table 3, the validation set AUC is the largest when modeling with top 10 markers, while the test set AUC is the largest when modeling with all markers.
[0139] In Table 3, AUC.test refers to the AUC on the test set; test_SE refers to the classification sensitivity on the test set; test_SP refers to the classification specificity on the test set; test_ACC is the classification accuracy on the test set; and AUC.validation represents the AUC on the validation set. The test set result indicates the model's performance in learning classification, while the validation set result indicates whether the model can generalize. Table 3 shows that the AUC on the test set reaches approximately 0.82, indicating good classification performance, and the AUC on the validation set is also around 0.8, suggesting that the test set result can generalize.
[0140] In one possible implementation, the unidimensional classification model of this disclosure may include a unidimensional classification model based on a specific methylation pattern (such as the FullC pattern). This model only considers reads in which all CpG sites are methylated, and has the characteristic of high specificity, which can reduce false positives.
[0141] Please see Figure 5 , Figure 5 A box plot is shown for predictive classification using a unidimensional classification model based on a specific methylation pattern.
[0142] Figure 5 In the figure, the specific methylation pattern is designated as `fullC_pattern`. The black (dark) and gray (light) lines on the left represent the healthy group and the cancer patient group, respectively, with the vertical axis representing probability. This embodiment utilizes `fullC_pattern` to classify the two groups of samples. Figure 5 It can be seen that the fullC_pattern feature can effectively distinguish between healthy people and cancer patients.
[0143] Please refer to Table 4, which illustrates the performance of predictive classification using a unidimensional classification model based on a specific methylation pattern.
[0144] Marker type Marker_Count AUC.test test_SE test_SP test_ACC AUC.validation fullC_pattern top_10_markers 0.785 0.711 0.872 0.792 0.818 fullC_pattern top_20_markers 0.815 0.757 0.863 0.811 0.774 fullC_pattern top_50_markers 0.843 0.776 0.891 0.833 0.803 fullC_pattern top_5_markers 0.766 0.692 0.886 0.789 0.816 fullC_pattern top_total_markers 0.843 0.772 0.900 0.836 0.815
[0145] Table 4
[0146] In Table 4, `Marker_Count` represents the number of markers used in machine learning modeling. For example, `top_10markers` means modeling using the top 10 best markers, and `top_20markers` means modeling using the top 20 best markers. As shown in Table 4, the validation set AUC is highest when using `top_10markers`, while the test set AUC is highest when using all markers. Here, `AUC.test` refers to the AUC on the test set; `test_SE` refers to the classification sensitivity on the test set; `test_SP` refers to the classification specificity on the test set; `test_ACC` is the classification accuracy on the test set; and `AUC.validation` is the AUC on the validation set. The test set result represents the model's classification performance (good or bad), while the validation set result represents whether the model can generalize. In Table 4, the test set AUC reaches approximately 0.83, indicating good classification performance, and the validation set AUC is around 0.81, suggesting that the test set result can generalize.
[0147] In one possible implementation, the unidimensional classification model of this disclosure may include a unidimensional classification model based on all methylation patterns except for a specific methylation pattern (such as the FullC pattern). This model can increase the sensitivity of model prediction and reduce the judgment of false negatives while increasing the high specificity of fullC. This dimension can be regarded as an auxiliary fullC dimension.
[0148] Please see Figure 6 , Figure 6 Box plots are shown for predictive classification using a unidimensional classification model based on all methylation patterns except for a specific methylation pattern.
[0149] Figure 6 In the diagram, the black (dark) and gray (light) line areas on the left represent the healthy group and the cancer patient group, respectively, with the vertical axis representing probability. This embodiment utilizes a marker_pattern to classify the two groups of samples. Figure 6 This demonstrates that we can effectively distinguish between healthy people and cancer patients using a one-dimensional classification model based on all methylation patterns except for specific methylation patterns.
[0150] Please refer to Table 5, which illustrates the performance of predictive classification using a unidimensional classification model based on a specific methylation pattern.
[0151] Marker type Marker_Count AUC.test test_SE test_SP test_ACC AUC.validation marker_pattern top_10_markers 0.808 0.775 0.820 0.796 0.821 marker_pattern top_20_markers 0.828 0.780 0.854 0.817 0.741 marker_pattern top_50_markers 0.857 0.812 0.882 0.846 0.768 marker_pattern top_5_markers 0.768 0.735 0.806 0.769 0.820 marker_pattern top_total_markers 0.875 0.842 0.881 0.861 0.782
[0152] Table 5
[0153] In Table 5, `Marker_Count` represents the number of markers used in machine learning modeling. For example, `top_10markers` means modeling using the top 10 best markers, and `top_20markers` means modeling using the top 20 best markers. As shown in Table 5, the validation set AUC is highest when using `top_10markers`, while the test set AUC is highest when using all markers. Here, `AUC.test` refers to the AUC on the test set; `test_SE` refers to the classification sensitivity on the test set; `test_SP` refers to the classification specificity on the test set; `test_ACC` is the classification accuracy on the test set; and `AUC.validation` is the AUC on the validation set. The test set result represents the model's classification performance (good or bad), while the validation set result represents whether the model can generalize. Table 5 shows that the test set AUC reaches approximately 0.80, indicating good classification performance, and the validation set AUC is also around 0.80, indicating that the test set result can generalize.
[0154] In one possible implementation, the multidimensional classification model of this disclosure can be based on multiple single-dimensional classification models, such as the aforementioned single-dimensional classification model based on the amplified region average entropy algorithm, the single-dimensional classification model based on a specific methylation pattern (such as the FullC pattern), and the single-dimensional classification model based on all methylation patterns other than the specific methylation pattern (such as the FullC pattern). The features of these three dimensions are complementary. This disclosure embodiment can use SFFS feature selection to select markers with good complementarity and classification performance as the final marker input prediction model (such as random forest) for classification prediction.
[0155] Please see Figure 7 , Figure 7 The box plot shows the prediction classification using a multidimensional classification model.
[0156] The multidimensional classification model combines three dimensions: the amplified region average entropy algorithm, specific methylation patterns, and all methylation patterns except for specific methylation patterns (such as the FullC pattern).
[0157] Figure 7 In the diagram, the black (dark) and gray (light) lines on the left represent the healthy population and cancer patients, respectively, with the vertical axis representing probability. Figure 7 As shown, this embodiment of the disclosure uses merged features to classify two groups of samples, illustrating that we can effectively distinguish between healthy people and cancer patients using merged features.
[0158] Please refer to Table 6, which illustrates the performance of predictive classification using a unidimensional classification model based on a specific methylation pattern.
[0159] Marker type Marker_Count AUC.test test_SE test_SP test_ACC AUC.validation merge3Type top_10_markers 0.815 0.762 0.854 0.809 0.807 merge3Type top_20_markers 0.841 0.797 0.870 0.833 0.813 merge3Type top_50_markers 0.892 0.865 0.887 0.876 0.851 merge3Type top_5_markers 0.787 0.767 0.799 0.782 0.811 merge3Type top_total_markers 0.893 0.868 0.889 0.879 0.861
[0160] Table 6
[0161] In Table 6, `Marker_Count` represents the number of markers used in machine learning modeling. For example, `top_10markers` means modeling using the top 10 best markers, and `top_20markers` means modeling using the top 20 best markers. When modeling with all markers, the AUC on the validation set is the highest, while the AUC on the test set is the highest. `AUC.test` refers to the AUC on the test set; `test_SE` refers to the classification sensitivity on the test set; `test_SP` refers to the classification specificity on the test set; `test_ACC` is the classification accuracy on the test set; and `AUC.validation` is the AUC on the validation set. The test set result represents the model's performance in learning classification, while the validation set result represents whether the model can generalize. In Table 6, the AUC on the test set reaches approximately 0.85, indicating good classification performance, and the AUC on the validation set is around 0.83, indicating that the test set result can generalize.
[0162] Please see Figure 8 , Figure 8 A schematic diagram illustrating the AUC comparison analysis of multi-dimensional algorithm modeling and prediction in embodiments of this disclosure is shown.
[0163] Figure 8 The table shows the results from left to right using marker_mEntropy (a unidimensional classification model based on the amplified region average entropy algorithm), fullC_pattern (a unidimensional classification model based on a specific methylation pattern), marker_pattern (a unidimensional classification model based on all methylation patterns except for a specific methylation pattern (such as the FullC pattern), and merge3Type modeling (a multidimensional classification model) that combines three features. Black represents the AUC on the test set, and gray represents the AUC on the validation set. It can be seen that the modeling with three features has the highest AUC. These three features are complementary, and combining them can improve the performance of prediction and classification.
[0164] The above analysis demonstrates that even with a single-dimensional model, the proposed computational method outperforms conventional analytical methods. Further analysis combining these single-dimensional feature modeling and ensemble modeling for predictive classification of tumors and healthy individuals reveals that multi-dimensional feature ensemble modeling outperforms single-dimensional feature modeling, thus improving classification and prediction performance from a data analysis perspective.
[0165] It should be noted that the dataset used in the above embodiments includes 162 tumor patients, of whom 68 are male, 84 are female, and 10 are of unknown sex; their ages range from 29 to 73 years, with a mean age of 53.4 years; according to pathological stage, they are divided into 13 stage I patients, 14 stage II patients, 39 stage III patients, 14 stage IV patients, and 1 patient of unknown stage. Healthy individuals include 47 males and 34 females; their ages range from 29 to 73 years, with a mean age of 54.1 years. Those skilled in the art can use other datasets for verification.
[0166] Of course, the specific implementation methods of the single-dimensional classification model and the multi-dimensional classification model in this disclosure are not limited. Those skilled in the art can choose appropriate technical means to implement them according to the actual situation and needs.
[0167] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0168] Please see Figure 9 , Figure 9 A block diagram of a sorting apparatus according to an embodiment of the present disclosure is shown.
[0169] like Figure 9 As shown, the device includes:
[0170] Extraction module 10 is used to extract the number of amplicons of various methylation patterns in each target amplification region and the number of amplicons in each internal reference amplification region from the sequencing sample data of methylation-specific amplification sequencing. The number of amplicons represents the number of reads of the reference genome sequence that are aligned to the amplification region.
[0171] The first determining module 20 is used to determine at least one target internal reference amplification region from each internal reference amplification region, and to standardize the number of amplons in each target amplification region using the number of amplons in the at least one target internal reference amplification region, so as to obtain the standardized number of amplons for each target amplification region for various methylation modes.
[0172] The second determining module 30 is used to calculate the average entropy of each target amplification region based on the proportion of the number of amplicones of each methylation mode in each target amplification region in each sequencing sample data.
[0173] The classification module 40 is used to input the average entropy of each target amplification region, the number of standardized amplicones of the preset methylation mode of each target amplification region, and the number of standardized amplicones of other methylation modes besides the preset methylation mode of each target amplification region into the corresponding trained single-dimensional classification model, and / or input them into the trained multi-dimensional classification model to obtain the classification result of the sequencing sample data.
[0174] This embodiment of the present disclosure extracts the number of amplicons for various methylation patterns in each target amplification region and the number of amplicons in each internal control amplification region from the sequencing sample data of methylation-specific amplification sequencing. At least one target internal control amplification region is determined from each internal control amplification region, and the number of amplicons in each target amplification region is standardized using the number of amplicons in the at least one target internal control amplification region to obtain the standardized number of amplicons for each methylation pattern in each target amplification region. Based on the proportion of amplicons for each methylation pattern in each target amplification region in each sequencing sample data, the average entropy of each target amplification region is calculated. The average entropy of each target amplification region, the standardized number of amplicons for the preset methylation pattern in each target amplification region, and the standardized number of amplicons for other methylation patterns in each target amplification region are respectively input into the corresponding trained single-dimensional classification model, and / or all are input into the trained multi-dimensional classification model, thereby quickly and accurately obtaining the classification result of the sequencing sample data.
[0175] In one possible implementation, extracting the number of amplicones for various methylation patterns in each target amplification region from the sequencing sample data of methylation-specific amplification sequencing includes:
[0176] The sequencing sample data is filtered to obtain the first set of reads;
[0177] Remove primer dimer sequences formed between primers in the first set of reads to obtain the second set of reads;
[0178] The reference genome sequence is used to match the reads in the second set of reads to determine the third set of reads;
[0179] The number of amplicones for each methylation pattern is determined based on the third set of reads.
[0180] In one possible implementation, the target internal reference amplification region is at least one of the L internal reference amplification regions with the highest gene stability among a plurality of internal reference amplification regions, where L is a positive integer.
[0181] In one possible implementation, determining at least one target internal control amplification region from the various internal control amplification regions includes:
[0182] Determine the average coefficient of variation of each internal control amplification region relative to the other internal control amplification regions;
[0183] Identify at least one of the L internal reference amplification regions with the smallest average coefficient of variation as the target internal reference amplification region.
[0184] In one possible implementation, determining the average coefficient of variation of each internal control amplification region relative to other internal control amplification regions includes:
[0185] Determine the ratio of the number of amplicones in the first internal control amplification region to the number of amplicones in each second internal control amplification region in each sequencing sample data.
[0186] The coefficients of variation of the first internal reference amplification region relative to each second internal reference amplification region are determined based on the ratios.
[0187] The average coefficient of variation of the first internal reference amplification region is determined based on the various coefficients of variation.
[0188] In one possible implementation, the standardization of the amplicon count of each target amplification region using the amplicon count of the at least one target internal reference amplification region includes:
[0189] Determine the geometric mean of the number of amplicones in each target intrinsic reference amplification region in each sample;
[0190] The normalized number of amplicones for each methylation mode in the target internal reference amplification region is obtained by dividing the number of amplicones for each methylation mode in the target internal reference amplification region by the geometric mean.
[0191] In one possible implementation, calculating the average entropy of each target amplification region based on the proportion of amplicon numbers of each methylation mode in each target amplification region in each sequencing sample data includes:
[0192] The average entropy of the target amplification region is calculated using the following formula:
[0193]
[0194] Where Marker_mEntropy represents the average entropy of the target amplification region, N represents the number of methylation patterns in the target amplification region, P_i represents the number of normalized amplicones of the i-th methylation pattern in the target amplification region, and i≤N.
[0195] In one possible implementation, the sum of the percentages of amplicon numbers for various methylation modes in each target amplification region is 1.
[0196] In one possible implementation, the preset methylation mode includes the fullC methylation mode.
[0197] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0198] This disclosure also proposes a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the above-described method. The computer-readable storage medium may be a non-volatile computer-readable storage medium.
[0199] This disclosure also proposes an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to execute the above-described method.
[0200] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.
[0201] Electronic devices can be provided as terminals, servers, or other forms of devices.
[0202] Please see Figure 10 , Figure 10 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0203] For example, electronic device 800 can be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, and other terminals.
[0204] Reference Figure 10 The electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0205] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0206] Memory 804 is configured to store various types of data to support the operation of electronic device 800. Examples of this data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0207] Power supply component 806 provides power to various components of electronic device 800. Power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.
[0208] Multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0209] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0210] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0211] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 may detect the on / off state of electronic device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include an optical sensor, such as a complementary metal-oxide-semiconductor (CMOS) or charge-coupled device (CCD) image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0212] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, second-generation mobile communication technology (2G), or third-generation mobile communication technology (3G), or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0213] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0214] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 804 including computer program instructions that can be executed by a processor 820 of an electronic device 800 to perform the above-described method.
[0215] Please see Figure 11 , Figure 11 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0216] For example, electronic device 1900 can be provided as a server. (See reference...) Figure 11 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.
[0217] Electronic device 1900 may also include a power supply component 1926 configured to perform power management of electronic device 1900, a wired or wireless network interface 1950 configured to connect electronic device 1900 to a network, and an input / output (I / O) interface 1958. Electronic device 1900 can operate on an operating system stored in memory 1932, such as Microsoft Server operating system (Windows Server). TM Apple's graphical user interface-based operating system (Mac OSX) TM ), a multi-user, multi-process computer operating system (Unix) TM Linux is a free and open-source Unix-like operating system. TM ), the open-source Unix-like operating system (FreeBSD) TM (or similar.)
[0218] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of an electronic device 1900 to perform the above-described method.
[0219] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.
[0220] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, (but not limited to) electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0221] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0222] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0223] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0224] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0225] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0226] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0227] The computer program product can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0228] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A classification method, characterized in that, The method includes: The number of amplicones for various methylation patterns in each target amplification region and the number of amplicones in each internal reference amplification region were extracted from the sequencing sample data of methylation-specific amplification sequencing. The number of amplicones represents the number of reads corresponding to the reference genome sequence in the amplification region. At least one target internal control amplification region is determined from each internal control amplification region, and the number of amplons in each target amplification region is standardized using the number of amplons in the at least one target internal control amplification region to obtain the standardized number of amplons for each target amplification region for various methylation modes. The average entropy of each target amplification region is calculated based on the proportion of amplicon numbers of each methylation mode in each target amplification region in each sequencing sample data. The average entropy of each target amplification region, the number of standardized amplicones for the preset methylation pattern of each target amplification region, and the number of standardized amplicones for other methylation patterns besides the preset methylation pattern of each target amplification region are respectively input into the corresponding trained single-dimensional classification model, and / or all are input into the trained multi-dimensional classification model to obtain the classification result of the sequencing sample data.
2. The method according to claim 1, characterized in that, The extraction of the number of amplicones for various methylation patterns in each target amplification region from the sequencing sample data of methylation-specific amplification sequencing includes: The sequencing sample data is filtered to obtain the first set of reads; Remove primer dimer sequences formed between primers in the first set of reads to obtain the second set of reads; The reference genome sequence is used to match the reads in the second set of reads to determine the third set of reads; The number of amplicones for each methylation pattern is determined based on the third set of reads.
3. The method according to claim 1, characterized in that, The target internal reference amplification region is at least one of the L internal reference amplification regions with the highest gene stability among multiple internal reference amplification regions, where L is a positive integer.
4. The method according to claim 3, characterized in that, Determining at least one target internal control amplification region from each internal control amplification region includes: Determine the average coefficient of variation of each internal control amplification region relative to the other internal control amplification regions; Identify at least one of the L internal reference amplification regions with the smallest average coefficient of variation as the target internal reference amplification region.
5. The method according to claim 4, characterized in that, Determining the average coefficient of variation of each internal control amplification region relative to other internal control amplification regions includes: Determine the ratio of the number of amplicones in the first internal control amplification region to the number of amplicones in each second internal control amplification region in each sequencing sample data. The coefficients of variation of the first internal reference amplification region relative to each second internal reference amplification region are determined based on the ratios. The average coefficient of variation of the first internal reference amplification region is determined based on the various coefficients of variation.
6. The method according to any one of claims 1-5, characterized in that, The standardization of the amplicon count of each target amplification region using the amplicon count of at least one target internal reference amplification region includes: Determine the geometric mean of the number of amplicones in each target intrinsic reference amplification region in each sample; The normalized number of amplicones for each methylation mode in the target internal reference amplification region is obtained by dividing the number of amplicones for each methylation mode in the target internal reference amplification region by the geometric mean.
7. The method according to claim 1, characterized in that, The step of calculating the average entropy of each target amplification region based on the proportion of amplicon numbers for each methylation mode in each target amplification region of each sequencing sample data includes: The average entropy of the target amplification region is calculated using the following formula: Where Marker_mEntropy represents the average entropy of the target amplification region, N represents the number of methylation patterns in the target amplification region, P_i represents the number of normalized amplicones of the i-th methylation pattern in the target amplification region, and i≤N.
8. The method according to claim 1, characterized in that, The sum of the percentages of amplicon numbers for each methylation mode in each target amplification region is 1.
9. The method according to claim 1, characterized in that, The preset methylation mode includes the fullC methylation mode.
10. A sorting device, characterized in that, The device includes: The extraction module is used to extract the number of amplicons of various methylation patterns in each target amplification region and the number of amplicons in each internal reference amplification region from the sequencing sample data of methylation-specific amplification sequencing. The number of amplicons represents the number of reads of the reference genome sequence that are aligned to the amplification region. The first determining module is used to determine at least one target internal reference amplification region from each internal reference amplification region, and to standardize the number of amplons in each target amplification region using the number of amplons in the at least one target internal reference amplification region, so as to obtain the standardized number of amplons for each target amplification region for various methylation modes. The second determining module is used to calculate the average entropy of each target amplification region based on the proportion of amplicon numbers of each methylation mode in each target amplification region in each sequencing sample data. The classification module is used to input the average entropy of each target amplification region, the number of standardized amplicones of the preset methylation pattern of each target amplification region, and the number of standardized amplicones of other methylation patterns besides the preset methylation pattern of each target amplification region into the corresponding trained single-dimensional classification model, and / or input them into the trained multi-dimensional classification model to obtain the classification result of the sequencing sample data.
11. An electronic device, characterized in that, include: processor; A memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to perform the method as claimed in any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, It stores computer program instructions that, when executed by a processor, implement the method as described in any one of claims 1-9.
Citation Information
Patent Citations
Urinary sediment genome DNA classification method and device and application
CN111833965A
Detection kit for TMEM101 gene methylation in human peripheral blood circulating tumor DNA for early screening of endometrial cancer
CN113215264A