A method, device and program product for tumor methylation typing and discovery of new subtypes based on multi-task learning

Through the neural network method based on multitask learning, brain tumor methylation data is processed, and the problem of insufficient brain tumor diagnosis accuracy and new subtype discovery ability in the prior art is solved, and high stability and repeatability diagnosis results are achieved.

CN119646524BActive Publication Date: 2025-06-27BEIJING NEUROSURGICAL INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411909512.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-06-27
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

The prior art has problems with insufficient accuracy, repeatability and new subtype discovery capabilities in the diagnosis of brain tumors, especially in the case of small sample size and poor quality.

Method used

A multi-task learning-based method is adopted to process and classify tumor methylation data through parallel neural networks, and determine whether new subtypes appear in combination with clustering analysis, and improve the adaptability and accuracy of the model through the model update mechanism.

Benefits of technology

It realizes simultaneous classification and new subtype monitoring in brain tumor diagnosis, with good stability and repeatability, and improves the accuracy and reliability of clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119646524B_ABST
    Figure CN119646524B_ABST
Patent Text Reader

Abstract

This application relates to the field of intelligent medicine, and specifically relates to a method, device and program product for tumor methylation typing and discovery of new subtypes based on multi-task learning. It includes: S1, obtaining tumor methylation data; S2, inputting the tumor methylation data into a trained first neural network and a second neural network in parallel, where the first neural network performs feature processing to obtain feature data, and the second neural network performs classification to obtain a classification type; S3, inputting the feature data and the classification type into a trained third neural network to obtain a classification prediction value, and comparing the classification prediction value with a preset threshold. When the classification prediction value is less than the preset threshold interval, it is determined as a new subtype, otherwise it is determined as a known type. This application can detect new subtypes when obtaining the tumor type, and has good clinical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent medicine, and particularly to a method, device, program product and computer-readable storage medium for tumor methylation typing and discovery of new subtypes based on multi-task learning. Background Art

[0002] Central nervous system (CNS) tumors are a highly complex and heterogeneous group of diseases, with the WHO classification defining approximately 100 different types. These tumors vary in malignancy from benign to malignant, and their accurate diagnosis is crucial for the selection of treatment regimens and prognosis assessment of patients. As an important epigenetic marker, DNA methylation can not only reflect the developmental origin of tumors, but also be used to trace the primary sites of highly dedifferentiated metastatic cancers. Even in cases with small sample sizes and poor quality, DNA methylation analysis still has high robustness and reproducibility.

[0003] Currently, the diagnosis of brain tumors mainly relies on the following methods: 1. Traditional diagnostic methods: Histopathological diagnosis based on morphology is the current gold standard, but this method mainly relies on the subjective judgment of pathologists and there are significant inter-observer differences. At the same time, conventional molecular tests such as IDH mutation status, TERT promoter mutation status detection, BRAF mutation status, fluorescence in situ hybridization (FISH) detection, and immunohistochemical detection are also commonly used in clinical practice, but these methods are often difficult to standardize and the result stability is insufficient. 2. Existing machine learning methods: The machine learning methods currently applied to tumor classification mainly include two categories: supervised learning and unsupervised learning. Although supervised learning methods can classify known types relatively accurately, they cannot discover new subtypes; while unsupervised learning methods have the potential to discover new categories, but lack the ability to accurately identify known categories and the result stability is poor. In addition, the existing active learning methods are inefficient in the sample annotation process and require a large amount of expert time for manual annotation. 3. Existing methylation typing methods: The current methylation typing methods mainly focus on the prediction of central nervous system tumor types. These methods usually adopt a single analysis strategy, fail to effectively integrate multi-task learning methods, and do not establish a reasonable new subtype discovery mechanism. At the same time, the existing data processing solutions have obvious deficiencies in terms of automation level, quality control, standardization of result interpretation, etc., and it is difficult to support the verification of results among multiple centers.

[0004] All of the above technical solutions have obvious limitations in practical applications, especially the deficiencies in accuracy, reproducibility, and discovery of new subtypes, resulting in an increase in the uncertainty of clinical diagnosis results and affecting the formulation of treatment plans. Summary of the Invention

[0005] In view of the above problems, the present invention provides a method for tumor methylation typing and discovery of new subtypes based on multi-task learning. A typing method that takes into account both the recognition of known types and the discovery of new subtypes and has good stability and repeatability has important clinical application value. Specifically, it includes: S1. Obtain tumor methylation data; S2. Input the tumor methylation data into a trained first neural network and a second neural network in parallel. The first neural network performs feature processing to obtain feature data, and the second neural network performs classification to obtain classification types; S3. Input the feature data and classification types into a trained third neural network to obtain a classification prediction value, and compare the classification prediction value with a preset threshold. When the classification prediction value is less than the preset threshold, it is determined as a new subtype; otherwise, it is determined as a known type.

[0006] When it is determined as a new subtype based on the classification prediction value, it also includes new subtype confirmation. Input the feature data determined as a new subtype into a fourth neural network for classification to obtain a final classification result. When the final classification result includes a new subtype, annotate the new subtype.

[0007] Optionally, the fourth neural network is a clustering network.

[0008] The first neural network includes N fully connected modules, and the second neural network includes N fully connected modules. Each two fully connected modules are connected by a feature exchange layer. The methylation data alternately passes through the fully connected modules and the feature exchange layer to obtain feature data in the first neural network and classification types in the second neural network.

[0009] Optionally, the feature exchange layer of the first neural network further includes fusing the features of the second neural network; after the first fully connected module of the first neural network outputs features, obtain the output features of the first fully connected module of the second neural network for fusion to obtain fused features, input the fused features into the feature exchange layer of the first neural network for processing to obtain first exchange features, and input the first exchange features into the second fully connected module of the first neural network. The feature exchange layer repeats the above process until the Nth fully connected module.

[0010] Optionally, the feature exchange layer of the second neural network further includes fusing the features of the first neural network; after the first fully connected module of the second neural network outputs features, obtain the output features of the first fully connected module of the first neural network for fusion to obtain fused features, input the fused features into the feature exchange layer of the second neural network for processing to obtain second exchange features, and input the second exchange features into the second fully connected module of the second neural network. The feature exchange layer repeats the above process until the Nth fully connected module.

[0011] Optionally, the fully connected module includes a batch normalization layer and an activation layer.

[0012] Optionally, one or more of the following are adopted as the loss function of the neural network: cross-entropy loss function, hinge loss function, Huber loss function.

[0013] Optionally, the loss function of the first neural network is calculated as follows:

[0014] where x is the target sample, and x i is any one of m noise samples, q θ is the encoder network, p represents the probability distribution, θ represents the parameters of the neural network, and x m represents the m-th noise sample, and ξ represents the target sample.

[0015] Optionally, one or more of the following are adopted for the first neural network model or the second neural network: convolutional neural network, residual network, Transformer.

[0016] The third neural network adopts one or more of the following: GMM, DBSCAN, K-means clustering, hierarchical clustering, Bayesian nonparametric method, infinite Gaussian mixture model.

[0017] The method further includes model update. When a new subtype type appears, after accumulating the sample data of the new subtype type, the first neural network, the second neural network, and the third neural network are updated to obtain the updated first neural network, second neural network, and third neural network. Feature data is obtained based on the updated first neural network, classification type data is obtained based on the updated second neural network, and it is determined whether there is a new subtype based on the updated third neural network.

[0018] Optionally, the neural network is updated by incremental learning for the update method.

[0019] Optionally, the model update further includes model verification. The performance of the updated model is verified through historical data. When the model performance reaches or exceeds the historical model performance, the first neural network, the second neural network, and the third neural network are updated to obtain the updated first neural network, second neural network, and third neural network.

[0020] The method further includes methylation data visualization. The methylation data is input into the first neural network model to obtain feature data. The feature data includes the abscissa and ordinate of the feature data, and two-dimensional visualization data is obtained through visualization processing based on the abscissa and ordinate.

[0021] The abscissa and ordinate represent the distribution positions of the classification types to which the methylation samples belong; the abscissa and ordinate are obtained based on the coordinates of similar samples.

[0022] Optionally, the two-dimensional visualization data includes the distribution regions of each type.

[0023] Optionally, during the training process of the first neural network, the abscissa and ordinate of the methylation data in each type are obtained through the acquired methylation dataset and type labels, and then the relationship between the classification type and the abscissa and ordinate is learned to obtain a mapping relationship. The training is repeated until the loss function remains unchanged, and the trained first neural network is obtained.

[0024] Optionally, when new methylation sample data is added, the two-dimensional abscissa and ordinate of the new methylation sample data are obtained through the mapping relationship in the first neural network.

[0025] The purpose of the present invention is to provide a computer program product, which includes a computer program or instruction. The computer program or instruction is executed by a processor to implement the above-mentioned tumor methylation typing and new subtype discovery method based on multi-task learning.

[0026] The purpose of the present invention is to provide a computer device, which includes a memory, a processor, and a computer program or instruction stored on the memory. The computer program or instruction is executed by the processor to implement the above-mentioned tumor methylation typing and new subtype discovery method based on multi-task learning.

[0027] The purpose of the present invention is to provide a computer-readable storage medium, which stores a computer program or instruction. The computer program or instruction is executed by a processor to implement the above-mentioned tumor methylation typing and new subtype discovery method based on multi-task learning.

[0028] Advantages of the present invention:

[0029] 1. Aiming at the problem of accurately classifying known tumor subtypes while effectively discovering new types, the present invention proposes to use multi-task learning, perform tumor classification and feature processing through parallel neural networks, and then perform clustering analysis based on tumor types and feature data to determine whether new subtypes appear. This solution can perform classification and new subtype monitoring simultaneously, has good stability and repeatability, and has important clinical significance.

[0030] 2. The present invention visualizes the methylation typing, with the same types clustering into a cluster and occupying one position. After adding new samples, not only does the number of classified types change, but also the distribution positions of the types in the visualized data change, which is not conducive to doctors observing the type attribution of new samples. Therefore, the present invention maps the relationship between the methylation data and the two-dimensional data to obtain the coordinates of the new samples in the two-dimensional visualized data, so that the distribution positions of the types remain unchanged, and the classification of new samples only shows a change in quantity, which helps doctors compare and obtain the types of new samples.

[0031] 3. When the present invention detects a new subtype, it annotates the new subtype, reducing the workload in the process of manual sample annotation and improving the usability of the model.

[0032] 4. For the detection of new subtypes, after discrimination by the clustering algorithm, the present invention also adds a new subtype confirmation process, improving the accuracy and reliability of the model in identifying new types.

[0033] 5. The network model of the present invention also includes a model update mechanism. For the newly added subtypes and samples, the model is updated through incremental learning, enabling the model to complete self-optimization and iterative update, having sustainable usability, and meeting all classifications and new subtype detections. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0035] Figure 1 Schematic flow chart of the method for tumor methylation typing and new subtype discovery based on multi-task learning provided by the embodiment of the present invention;

[0036] Figure 2 Schematic diagram of the system for tumor methylation typing and new subtype discovery based on multi-task learning provided by the embodiment of the present invention;

[0037] Figure 3 Schematic diagram of the device for tumor methylation typing and new subtype discovery based on multi-task learning provided by the embodiment of the present invention;

[0038] Figure 4 Network structure diagram of the tumor methylation typing and new subtype discovery based on multi-task learning provided by the embodiment of the present invention;

[0039] Figure 5 Comparison results of TADM2 and other classification algorithms provided by the embodiment of the present invention;

[0040] Figure 6 AUC result comparison between the present invention and other algorithms provided by the embodiments of the present invention;

[0041] Figure 7 Workload comparison between TADM2 annotation and complete dataset annotation provided by the embodiments of the present invention;

[0042] Figure 8 Visualization graph provided by the embodiments of the present invention;

[0043] Figure 9 Comparison results between TAMD2 and other visualization methods provided by the embodiments of the present invention. Detailed implementation manners

[0044] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0045] In some processes described in the specification, claims and above-mentioned drawings of the present invention, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as S101, S102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0046] Figure 1 Schematic diagram of a method for tumor methylation typing and discovery of new subtypes based on multi-task learning provided by the embodiments of the present invention, specifically including:

[0047] S1: Obtain tumor methylation data;

[0048] In one embodiment, the tumor includes one or more of the following: brain tumor, epithelial tumor, mesenchymal tumor, lymphoma, melanoma, breast cancer, lung cancer, colorectal cancer, prostate cancer, ovarian cancer, pancreatic cancer.

[0049] S2: Input the tumor methylation data into a trained first neural network and a second neural network in parallel. The first neural network performs feature processing to obtain feature data, and the second neural network performs classification to obtain classification types;

[0050] In one embodiment, the first neural network includes N fully-connected modules, and the second neural network includes N fully-connected modules. One feature exchange layer is connected between every two fully-connected modules. The methylation data alternately passes through the fully-connected modules and the feature exchange layer to obtain feature data in the first neural network and classification types in the second neural network.

[0051] In one embodiment, the feature exchange layer of the first neural network further includes fusing the features of the second neural network; after the first fully-connected module of the first neural network outputs features, it obtains the output features of the first fully-connected module of the second neural network for fusion to obtain fused features, and inputs the fused features into the feature exchange layer of the first neural network for processing to obtain first exchanged features, and inputs the first exchanged features into the second fully-connected module of the first neural network. The feature exchange layer repeats the above process until the Nth fully-connected module.

[0052] In one embodiment, the feature exchange layer of the second neural network further includes fusing the features of the first neural network; after the first fully-connected module of the second neural network outputs features, it obtains the output features of the first fully-connected module of the first neural network for fusion to obtain fused features, and inputs the fused features into the feature exchange layer of the second neural network for processing to obtain second exchanged features, and inputs the second exchanged features into the second fully-connected module of the second neural network. The feature exchange layer repeats the above process until the Nth fully-connected module.

[0053] In one embodiment, the fully-connected module includes a batch normalization layer and an activation layer.

[0054] In one embodiment, the loss function of the neural network adopts one or more of the following: cross-entropy loss function, hinge loss function, Huber loss function.

[0055] In one embodiment, the loss function of the first neural network is the following calculation method:

[0056] where x is the target sample, x i is any one of m noise samples, q θ is the encoder network, p represents the probability distribution, θ represents the parameters of the neural network, x m represents the mth noise sample, and ξ represents the target sample.

[0057] In one embodiment, the first neural network model or the second neural network adopts one or more of the following: convolutional neural network, residual network, Transformer.

[0058] S3: Input the feature data and classification type into the trained third neural network to obtain a classification prediction value, compare the classification prediction value with a preset threshold. When the classification prediction value is less than the preset threshold range, it is determined as a new subtype; otherwise, it is determined as a known type.

[0059] In one embodiment, when it is determined as a new subtype based on the classification prediction value, it further includes new subtype confirmation. Input the feature data determined as a new subtype into the fourth neural network for classification to obtain the final classification result. When the final classification result includes a new subtype, annotate the new subtype.

[0060] In one embodiment, the fourth neural network is a clustering network.

[0061] In one embodiment, the third neural network adopts one or more of the following: GMM, DBSCAN, K - means clustering, hierarchical clustering, Bayesian non - parametric method, infinite Gaussian mixture model.

[0062] In one embodiment, the method further includes model update. When a new subtype type appears, after accumulating the sample data of the new subtype type, update the first neural network, the second neural network, and the third neural network to obtain the updated first neural network, the second neural network, and the third neural network. Obtain feature data based on the updated first neural network, obtain classification type data based on the updated second neural network, and determine whether there is a new subtype based on the updated third neural network.

[0063] In one embodiment, the update method performs neural network update through incremental learning.

[0064] In one embodiment, the model update further includes model verification. Verify the performance of the updated model through historical data. When the model performance reaches or exceeds the historical model performance, update the first neural network, the second neural network, and the third neural network to obtain the updated first neural network, the second neural network, and the third neural network.

[0065] In a specific embodiment, the present invention proposes a tumor methylation typing model TADM2 (Tumor subtype Annotation and Discovery based on Multi - task learning of DNA Methylation) based on multi - task learning, mainly including: a network architecture, as Figure 4 shown:

[0066] Input layer: Receive methylation data (dimension: 32000);

[0067] Dimensionality reduction network path: 3-layer fully connected layer architecture (1000→500→250→100→2); BatchNorm and ReLU activation functions are used for each layer, and the Dropout rate is 0.2 to prevent overfitting;

[0068] Classification network path: 3-layer fully connected layer architecture (1000→500→250→100→95); The same normalization and activation function settings are used;

[0069] Cross-stitch Unit implementation (feature exchange layer): Cross units are set between every two main layers (fully connected architecture); Implementation formula: Shared Task A = αAA(Task A) + αBA(Task B)

[0070] Shared Task B = αAB(Task A) + αBB(Task B)

[0071] Among them, αAA, αBA, αAB, αBB are learnable parameters;

[0072] Loss function design of the dimensionality reduction network (the first neural network): InfoNCE contrastive learning loss;

[0073] Among them, x is the target sample, x i is m noise samples, q θ is the encoder network.

[0074] Loss function of the classification network (the second neural network), cross-entropy classification loss, is used for the supervised learning branch to calculate the difference between the predicted class and the true label.

[0075] New subtype discovery mechanism:

[0076] 1. Scoring distribution modeling: Use a single-component Gaussian mixture model (GMM) to establish a scoring distribution model for each known category, calculate the probability density of new samples, and convert it into a p-value.

[0077] 2. Adaptive clustering algorithm: Parameter optimization range eps: 0.1 - 2.5, step size 0.1; min_samples: 1 - 9;

[0078] Optimization objective: Maximize the Silhouette Score, clustering condition: at least form more than 2 clusters.

[0079] Self-optimizing system: 1. Data processing flow.

[0080] Input: Methylation data matrix; Preprocessing: Data standardization; Feature extraction: Obtain the low-dimensional representation through a trained model; Classification prediction: Obtain the class prediction result and confidence level.

[0081] 2. New subtype confirmation process.

[0082] Anomaly detection: Screen abnormal samples based on p-value; Cluster analysis: Use optimized DBSCAN parameters for clustering; Manual confirmation: Expert review of the clustering results.

[0083] 3. Model update mechanism.

[0084] Trigger condition: Accumulate a certain number of new annotated samples (new subtype annotated samples); Update method: Incremental learning to maintain model stability; Verification mechanism: Use historical data to verify the performance of the updated model.

[0085] In a specific embodiment, the classification accuracy of the present invention is significantly improved. As Figure 5 shown, compared with the existing common method Random Forest, it is improved by about 5%, and the classification accuracy is also improved compared with other classification methods Xgboost and MLP. It can reliably discover new tumor subtypes, and the AUC value of new subtype discrimination reaches 0.930, as Figure 6 shown. In addition, as Figure 7 shown, using TADM2 to screen the data that needs to be annotated, by comparing the training on the complete dataset and the dataset screened by TADM2, the model prediction accuracy remains unchanged, but the overall required data volume is greatly reduced, reducing the manual annotation workload by about 70%. The model can be continuously optimized and has strong adaptability.

[0086] In an embodiment, the method further includes visualizing the methylation data, inputting the methylation data into the first neural network model to obtain feature data, where the feature data includes the abscissa and ordinate of the feature data, and performing visualization processing based on the abscissa and ordinate to obtain two-dimensional visualization data.

[0087] In an embodiment, the abscissa and ordinate represent the distribution positions of the classification types to which the methylation samples belong, and the abscissa and ordinate are obtained based on the coordinates of similar samples.

[0088] In an embodiment, the two-dimensional visualization data includes the distribution regions of each type.

[0089] In an embodiment, during the training process of the first neural network, self-supervised contrast learning is performed through the obtained methylation dataset to obtain the abscissa and ordinate of each type of methylation data, and the training is repeated until the loss function remains unchanged to obtain the trained first neural network.

[0090] In one embodiment, when adding new methylation sample data, the new methylation sample data is input into a first neural network to obtain two-dimensional abscissa and ordinate, where the two-dimensional abscissa and ordinate of the new sample are obtained based on the coordinates of similar samples.

[0091] In one embodiment, the methylation data is input into a trained first neural network to obtain two-dimensional coordinates of the methylation data. The two-dimensional coordinates are obtained based on the coordinates of similar samples, and the two-dimensional coordinates are visualized, where the visualized data is the distribution area of each type, such as Figure 8 shown.

[0092] When the type of the new sample is a known type, the two-dimensional coordinates of the distribution area of its belonging type are obtained according to the first neural network; when the type of the new sample is a new subtype and the model is updated, the classification type and coordinate mapping relationship in the first neural network are learned and updated.

[0093] In a specific embodiment, the methylation data is input into a dimensionality reduction network (the first neural network) for dimensionality reduction to obtain the abscissa and ordinate of the methylation data. Based on the abscissa and ordinate, the distribution position of the methylation data is obtained and visualized through the two-dimensional coordinate position, such as Figure 8 shown. It can intuitively obtain the distribution positions of various types of methylation data, and when adding new samples, the classification of the new samples can be obtained through the comparison before and after. In addition, compared with the existing clustering type visualization method (although the classification where the new sample is located can be displayed after adding the new sample, the distribution area of each type will change), the two-dimensional coordinates obtained by the dimensionality reduction network of the present invention are visualized and displayed, and the distribution areas of each type remain unchanged after being determined. After adding new samples, they are displayed in the fixed area of the type. Based on this method, doctors can intuitively and simply observe and compare the changes before and after adding new samples. In addition, comparing the TADM2 of the present invention with UMAP and tSNE, as Figure 9 shown, TADM2 is far superior to UMAP and tSNE.

[0094] The disclosed embodiments of the present invention also provide a computer program product or system, including a computer program, which when executed by a processor implements the steps of the above-mentioned tumor methylation typing and new subtype discovery method based on multi-task learning.

[0095] Figure 2 The schematic diagram of the tumor methylation typing and new subtype discovery system based on multi-task learning provided by the embodiments of the present invention specifically includes: an acquisition unit: acquiring tumor methylation data;

[0096] Classification unit: Input the tumor methylation data into the trained first neural network and the second neural network in parallel. The first neural network performs feature processing to obtain feature data, and the second neural network performs classification to obtain classification types.

[0097] Judgment unit: Input the feature data and classification types into the trained third neural network to obtain a classification prediction value. Compare the classification prediction value with a preset threshold. When the classification prediction value is less than the preset threshold interval, it is determined as a new subtype; otherwise, it is determined as a known type.

[0098] Figure 3 The schematic diagram of the device for tumor methylation typing and new subtype discovery based on multi-task learning provided by the embodiments of the present invention specifically includes:

[0099] A memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, any one of the above-mentioned methods for tumor methylation typing and new subtype discovery based on multi-task learning.

[0100] The disclosed embodiments of the present invention also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any one of the above-mentioned methods for tumor methylation typing and new subtype discovery based on multi-task learning.

[0101] The verification results of this verification embodiment show that allocating fixed weights for indications can improve the performance of this method compared to the default settings. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium can include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk, or optical disk, etc.

[0102] Those of ordinary skill in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned medium storage can be read-only memory, magnetic disk, or optical disk, etc.

[0103] The above has introduced in detail a computer device provided by the present invention. For those of ordinary skill in the art, according to the idea of the embodiments of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for tumor methylation typing and new subtype discovery based on multi-task learning, characterized in that: include: S1. Obtain tumor methylation data; S2, inputting the tumor methylation data into a first neural network and a second neural network that are trained in parallel, wherein the first neural network performs feature processing to obtain feature data, and the second neural network performs classification to obtain classification types; The first neural network includes N layers of fully connected modules, and the second neural network includes N layers of fully connected modules. Every two fully connected modules are connected by a feature exchange layer, and the methylation data alternately passes through the fully connected modules and the feature exchange layer; the feature exchange layer of the first neural network also includes a feature of the second neural network; after the first fully connected module of the first neural network outputs the feature, the output feature of the first fully connected module of the second neural network is obtained and fused to obtain a fused feature, the fused feature is input to the feature exchange layer of the first neural network for processing to obtain a first exchange feature, the first exchange feature is input to the second fully connected module of the first neural network, and the feature exchange layer repeats the above process until the Nth fully connected module; The feature exchange layer of the second neural network also includes fusing features of the first neural network; after the first fully connected module of the second neural network outputs features, the output features of the first fully connected module of the first neural network are obtained and fused to obtain fused features, the fused features are input to the feature exchange layer of the second neural network for processing to obtain second exchange features, the second exchange features are input to the second fully connected module of the second neural network, and the feature exchange layer repeats the above process until the Nth fully connected module; Obtain feature data in the first neural network, and obtain classification types in the second neural network; The fully connected module includes a batch normalization layer and an activation layer; S3. Input the feature data and classification type into a trained third neural network to obtain a classification prediction value, and compare the classification prediction value with a preset threshold. When the classification prediction value is less than the preset threshold interval, it is determined to be a new subtype, otherwise it is determined to be a known type.

2. The method for tumor methylation typing and new subtype discovery based on multi-task learning according to claim 1, characterized in that: When a new subtype is determined based on the classification prediction value, it also includes new subtype confirmation, and the feature data determined as the new subtype is input into the fourth neural network for classification to obtain a final classification result. When the final classification result includes a new subtype, the new subtype is annotated.

3. The method for tumor methylation typing and new subtype discovery based on multi-task learning according to claim 2, characterized in that: The fourth neural network is a clustering network.

4. The method for tumor methylation typing and new subtype discovery based on multi-task learning according to claim 1, characterized in that: The loss function of the neural network adopts one or more of the following: cross entropy loss function, hinge loss function, Huber loss function.

5. The method for tumor methylation typing and new subtype discovery based on multi-task learning according to claim 1, characterized in that: The loss function of the first neural network is The calculation method is as follows: Among them, x is the target sample, x i is any one of the m noise samples, q θ is the encoder network, p represents the probability distribution, θ represents the parameters of the neural network, and x m represents the mth noise sample, and ξ represents the target sample.

6. The method for tumor methylation typing and new subtype discovery based on multi-task learning according to claim 1, characterized in that: The first neural network model or the second neural network adopts one or more of the following: convolutional neural network, residual network, Transformer.

7. The method for tumor methylation typing and new subtype discovery based on multi-task learning according to claim 1, characterized in that: The third neural network adopts one or more of the following: GMM, DBSCAN, K-means clustering, hierarchical clustering, Bayesian non-parametric method, infinite Gaussian mixture model.

8. The method for tumor methylation typing and new subtype discovery based on multi-task learning according to claim 1, characterized in that: The method also includes model updating. When a new subtype type appears, the first neural network, the second neural network, and the third neural network are updated after accumulating sample data of the new subtype type to obtain updated first neural network, second neural network, and third neural network, characteristic data is obtained based on the updated first neural network, classification type data is obtained based on the updated second neural network, and whether there is a new subtype is determined based on the updated third neural network.

9. The method for tumor methylation typing and new subtype discovery based on multi-task learning according to claim 8, characterized in that: Neural network updating via incremental learning.

10. The method for tumor methylation typing and new subtype discovery based on multi-task learning according to claim 8, characterized in that: The model update also includes model verification, which verifies the updated model performance through historical data. When the model performance reaches or exceeds the historical model performance, the first neural network, the second neural network, and the third neural network are updated to obtain updated first neural network, second neural network, and third neural network.

11. The method for tumor methylation typing and new subtype discovery based on multi-task learning according to claim 1, characterized in that: The method also includes methylation data visualization, inputting the methylation data into a trained first neural network model to obtain feature data, wherein the feature data includes a horizontal coordinate and a vertical coordinate of the feature data, and performing visualization processing based on the horizontal coordinate and the vertical coordinate to obtain two-dimensional visualization data.

12. The method for tumor methylation typing and new subtype discovery based on multi-task learning according to claim 11, characterized in that: The horizontal and vertical coordinates represent the distribution positions of the classification types to which the methylation samples belong; the horizontal and vertical coordinates are obtained based on the coordinates of similar samples.

13. The method for tumor methylation typing and new subtype discovery based on multi-task learning according to claim 11, characterized in that: The two-dimensional visualization data includes distribution areas of various types.

14. The method for tumor methylation typing and new subtype discovery based on multi-task learning according to claim 11, characterized in that: During the training process, the first neural network performs self-supervised comparative learning on the acquired methylation data set to obtain the horizontal coordinates and vertical coordinates of the methylation data in each type, and repeats the training until the loss function remains unchanged to obtain a trained first neural network.

15. The method for tumor methylation typing and new subtype discovery based on multi-task learning according to claim 11, characterized in that: When newly added methylation sample data is added, the newly added methylation sample data obtains two-dimensional abscissas and ordinates through the first neural network, wherein the two-dimensional abscissas and ordinates of the newly added samples are obtained based on the coordinates of similar samples.

16. A computer program product comprising a computer program or instructions, characterized in that: The computer program or instructions are executed by a processor to implement the tumor methylation typing and new subtype discovery method based on multi-task learning as described in any one of claims 1-15.

17. A computer device comprising a memory, a processor and a computer program or instruction stored in the memory, characterized in that: The computer program or instructions are executed by a processor to implement the tumor methylation typing and new subtype discovery method based on multi-task learning as described in any one of claims 1-15.

18. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: The computer program or instructions are executed by a processor to implement the tumor methylation typing and new subtype discovery method based on multi-task learning as described in any one of claims 1-15.

Citation Information

Patent Citations

  • Primary lesion prediction method and system of unknown primary tumor based on DenseFormer model

    CN117894452A

  • Semi-supervised classification method and system for predicting cancer cfDNA methylation mutation sites

    CN118038974A