Cancer accurate positioning method and system based on multi-omics view missing complementation and dynamic fusion
Through the dynamic fusion strategy of the Transformer encoder-decoder network and Dirichlet distribution, the problem of missing views of cancer multi-omics data is solved, and more efficient and accurate cancer precision positioning is achieved.
Patent Information
- Application Number
- CN202510680308.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-12
AI Technical Summary
In the absence of multi-omics data views of cancer, existing technologies suffer from reduced model performance and inaccurate positioning results. Traditional methods are also difficult to adapt to individual differences and changes in the course of the disease and lack cross-center robustness.
The Transformer encoder-decoder network is used to complete missing views, combined with the dynamic fusion strategy of Dirichlet distribution, and missing views are marked by a mask matrix. The inherent correlation and redundancy of multi-omics data are utilized to dynamically adjust view weights for precise cancer localization.
It improves the accuracy and robustness of cancer precision positioning, can effectively deal with data missing, adapt to individual differences and changes in disease course, and improves the overall performance and diagnostic reliability of the model.
Smart Images

Figure CN120636840A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence, bioinformatics and precision cancer medicine, and in particular to a method and system for precise cancer localization based on multi-omics view defect completion and dynamic fusion. Background Art
[0002] Cancer is a major global health challenge, and its early and precise localization is crucial for improving treatment success and patient survival. However, in real-world clinical practice, due to factors such as sample collection, experimental conditions, or technical limitations, cancer multi-omics data often lack key views, which directly impacts the accuracy of cancer localization. Therefore, utilizing advanced view completion technologies for cancer localization research is particularly important.
[0003] Multi-omics data, including genomics, transcriptomes, and proteomes, provide a wealth of information for a comprehensive understanding of cancer development, progression, and metastasis. Comprehensive analysis of these multi-source data can reveal the complex mechanisms of cancer and provide a strong basis for its precise localization. However, due to technical or sample limitations, these multi-omics data often contain missing views, posing a significant challenge to subsequent analysis and localization. To address this issue, view completion technology has emerged. This technology interactively fills in missing data by leveraging the inherent correlation and redundancy between data, thereby improving the robustness and accuracy of the model. In the precise localization of cancer, view completion technology not only fully utilizes existing data resources but also, to a certain extent, compensates for data missing due to technical limitations, thereby improving the reliability of localization.
[0004] However, current technologies still have shortcomings in view completion. When some key view data is missing, the model's performance often degrades significantly, resulting in inaccurate positioning results. In addition, traditional multi-view fusion methods typically use a fixed weight strategy, which is difficult to adapt to individual differences between patients and the dynamic changes in the disease course, thus limiting further improvements in positioning accuracy. In addition, most studies are only verified on a single dataset and lack robustness in a cross-center, multi-device environment, which limits the practical application range of the system. Summary of the Invention
[0005] In order to improve the accuracy of cancer precise localization results in the case of missing views, the present invention provides a cancer precise localization method and system based on missing complement and dynamic fusion of multi-omics views.
[0006] In a first aspect, the present invention provides a method for precise cancer localization based on deletion completion and dynamic fusion of multi-omics views, comprising:
[0007] Obtain cancer multi-omics view data of the target object and unify the feature dimensions of all the acquired view data;
[0008] Generate a mask matrix to mark the view missing of the target object; wherein one element in the mask matrix corresponds to the missing state of one view feature;
[0009] Inputting all the unified view data and mask matrix into a preset cancer precise positioning model to obtain the cancer type of the target object; wherein the processing process of the preset cancer precise positioning model includes:
[0010] Using the mask matrix to process all view data after unifying the feature dimension, to obtain an original multi-view feature representation of the target object; wherein the original multi-view feature representation carries a view missing flag;
[0011] The original multi-view feature representation of the target object is input into the Transformer encoder-decoder to complete the features at the view missing mark, thereby obtaining a complete multi-view feature representation of the target object;
[0012] Dynamically fuse multi-view features based on the complete multi-view feature representation of the target object to obtain the fused features of the target object;
[0013] The fused features of the target object are input into the classifier to obtain the cancer type of the target object.
[0014] In this example, the Transformer encoder-decoder network is applied to the multi-view missing data completion task, leveraging its powerful sequence modeling capabilities. This network efficiently captures the complex dependencies between views, enabling accurate completion of missing views. This approach not only improves the efficiency of information transfer between views but also enhances the model's robustness to missing data, ensuring the accuracy and efficiency of the completion process.
[0015] In this example, the feature representations of multiple views of the target object are fused and predicted to produce fused features and classification prediction results. This fusion strategy fully exploits the unique contribution of each view to precise cancer localization. This process not only optimizes the interaction and complementarity between views but also enhances the model's ability to process complex data. Through the fused features, the model can make more accurate classification predictions and fully consider the relative importance of each view during the prediction process, thereby achieving more accurate and stable classification results.
[0016] In this example, the complete feature representation of the target object is input into a classifier, which then accurately classifies the fused features to produce the final prediction. This process not only optimizes the model's ability to understand multi-omics data but also ensures that the classifier can effectively and accurately determine cancer type based on the complementarity and information importance of each view, providing efficient and reliable support for cancer diagnosis.
[0017] Furthermore, all view data after unifying the feature dimension are processed using the mask vector to obtain the original multi-view feature representation of the target object, specifically including:
[0018] Set the mask matrix of the target object to mask = {m v |v=1,2,...,V}; where V represents the number of views, m v Represents a vector used to mark the missing state of view v; all view data after setting the unified dimension are X'={X' v |v=1,2,...,V}; where X′ v Represents the feature representation corresponding to view v; then the original multi-view feature representation of the target object is X flattened =cat(X′1⊙m1,X'2⊙m2,...,X' V ⊙m V ); where cat represents concatenation on dimension 1 and ⊙ represents element-wise multiplication.
[0019] Furthermore, the multi-view features are dynamically fused based on the complete multi-view feature representation of the target object to obtain the fused features of the target object, specifically including:
[0020] Input the complete representation of any view v of the object into the classifier to obtain the probability distribution of view v belonging to each category and use it as evidence Use Dirichlet distribution for evidence Modeling is performed to obtain the Dirichlet distribution parameters With evidence The relationship between Thus, the belief quality when any view v belongs to category c is obtained according to the following formula: and the uncertainty u of the classification prediction for that view v v :
[0021]
[0022] Among them, C is the total number of given categories, S v is the Dirichlet distribution intensity, c=1,2,…,C, represents the Dirichlet distribution parameter when view v belongs to the cth category, Evidence that view v belongs to the cth category; v = 1, 2, ..., V;
[0023] The belief quality according to view v and uncertainty u v , calculate the weight w of view v v :
[0024]
[0025] Among them, u v is the normalization factor;
[0026] According to the weight of each view, all views under the target object are fused according to the following formula to obtain the fusion feature h of the target object:
[0027]
[0028] Among them, V represents the number of views, X completed[:,v,:] is the complete multi-view feature representation X of the target object completed The feature representation corresponding to the view v in the figure.
[0029] Furthermore, during the training of the cancer precise localization model, the loss function L total for:
[0030]
[0031] Among them, L is the loss corresponding to the classifier, L completion is the loss corresponding to the Transformer encoder-decoder, L Fusion is the loss corresponding to dynamic fusion, w v is the vth weight vector of the model, ||w v ||2 is the weight vector w v , V is the total number of weight vectors in the model, λ1 and λ2 are the preset coefficients of the corresponding losses, and λ3 is the regularization coefficient.
[0032] In this embodiment, the total loss function is composed of the loss functions of multiple modules in the model, including the loss function for the view missing completion task module, the single-view multi-task loss function for the Transformer encoder-decoder network, and the loss function for the classifier. By organically integrating the loss functions of each module, a multi-level, multi-angle optimization objective is formed, ensuring the balance and coordinated evolution of the model under multiple task objectives.
[0033] Furthermore, the losses corresponding to the classifier and dynamic fusion are both cross-entropy losses, and the losses corresponding to the Transformer encoder-decoder are mean square error losses.
[0034] In a second aspect, the present invention provides a cancer precise localization system based on multi-omics view deletion completion and dynamic fusion, comprising:
[0035] A data acquisition and preprocessing unit, configured to acquire cancer multi-omics view data of a target object and unify the feature dimensions of all acquired view data;
[0036] A data mask generation unit is used to generate a mask matrix to mark the view missing status of the target object; wherein one element in the mask matrix corresponds to marking the missing status of one view feature;
[0037] A cancer precise positioning unit is configured to input all the unified view data and the mask matrix into a preset cancer precise positioning model to obtain the cancer type of the target object; the preset cancer precise positioning model includes a view missing labeling module, a multi-view missing completion module, a dynamic fusion module, and a classification learning module;
[0038] a view missing marking module, configured to process all view data after unifying the feature dimensions using the mask matrix to obtain an original multi-view feature representation of the target object; wherein the original multi-view feature representation carries a view missing mark;
[0039] The multi-view missing completion module is used to input the original multi-view feature representation of the target object into the Transformer encoder-decoder to complete the features at the view missing mark, thereby obtaining a complete multi-view feature representation of the target object;
[0040] A dynamic fusion module is used to dynamically fuse multi-view features based on the complete multi-view feature representation of the target object to obtain the fused features of the target object;
[0041] The classification learning module is used to input the fusion features of the target object into the classifier to obtain the cancer type of the target object.
[0042] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.
[0043] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the first aspect.
[0044] The beneficial effects of the present invention are:
[0045] The cancer precise localization method and system provided by the present invention are designed with missing completion and dynamic fusion strategies, which can fully utilize the multi-omics information in multi-view data, thereby improving the accuracy and reliability of cancer precise localization.
[0046] During training, the multi-omics view training set includes omics data from multiple levels, covering multiple dimensions of cancer characteristics, providing the model with a more comprehensive perspective. Through a multi-view missing completion strategy, the model can effectively address missing data and ensure data integrity and validity. At the same time, a dynamic fusion strategy enables the model to flexibly adjust the weights of each view's data based on the characteristics and relative importance of different views, achieving more accurate cancer localization. This process effectively improves the model's overall performance and ensures the efficiency and accuracy of precise cancer localization. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 A schematic diagram of a process for accurately localizing cancer based on multi-omics view deletion completion and dynamic fusion provided by an embodiment of the present invention;
[0048] Figure 2 A framework diagram of a cancer precision localization model provided by an embodiment of the present invention;
[0049] Figure 3 A schematic diagram of the structure of a cancer precise localization system based on multi-omics view deletion completion and dynamic fusion provided by an embodiment of the present invention;
[0050] Figure 4 This is a structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0051] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0052] Combine Figure 1 and Figure 2 As shown, an embodiment of the present invention provides a method for accurately localizing cancer based on multi-omics view deletion completion and dynamic fusion, comprising the following steps:
[0053] S101: Acquire cancer multi-omics view data of the target object and unify the feature dimensions of all acquired view data;
[0054] Specifically, in multi-view learning, since the data dimensions of different views vary greatly, if the dimensions are not unified first, in the subsequent fusion process, directly fusing the view data of different data dimensions may result in information loss or increased errors. Therefore, it is particularly important to design a feature conversion layer to map the data dimensions of different views to a unified dimensional space. In this embodiment, a linear layer is used as a feature conversion layer to map the data features from different views to the same dimensional space, ensuring the consistency and comparability of the multi-view data in subsequent processing and providing a unified basis for subsequent fusion operations. That is, in this embodiment, after obtaining the cancer multi-omics view data of the target object, each acquired view data will be processed through a linear layer to obtain its potential representation, as shown in formula (1).
[0055] X'={Linear(X v )|v=1,2,...,V} (1)
[0056] Among them, Linear() is a linear layer, whose main function is to transform different views X v The feature dimensions of the model are reduced to a unified dimension.
[0057] S102: Generate a mask vector to mark the view missing status of the target object; wherein one element in the mask vector corresponds to marking the missing status of one view;
[0058] Specifically, in order to effectively simulate the situation of view missing, this embodiment generates a binary mask matrix (mask) to mark which view data is missing. Specifically, each row vector of the mask matrix mask is used to mark the missing state of the corresponding view. Taking a row vector as an example, if a view feature data is missing, the element at the corresponding position in the row vector is 0; on the contrary, when the view feature data exists, the element at the corresponding position is 1. Figure 2 In , the position of “?” is the missing position of view feature data. Finally, the mask matrix is formed as shown in formula (2).
[0059] mask={m v |v=1,2,...,V} (2)
[0060] Among them, V represents the number of views, m v A vector representing missing values for view v.
[0061] S103: Inputting all the unified view data and mask vectors into a preset cancer precise positioning model to obtain the cancer type of the target object; wherein the processing process of the preset cancer precise positioning model includes:
[0062] S1031: Using the mask matrix to process all view data after unifying the feature dimension, obtain the original multi-view feature representation of the target object; wherein the original multi-view feature representation carries a view missing mark, such as Figure 2 The feature matrix of the latent representation stage is shown in .
[0063] Specifically, the mask matrix of the target object is set to mask = {m v |v=1,2,...,V}; all view data after setting unified dimension are X'={X′ v |v=1,2,...,V}; where X′ v represents the feature representation corresponding to view v; then the original multi-view feature representation of the target object is:
[0064] X flattened =cat(X′1⊙m1,X'2⊙m2,...,X' V ⊙m V ) (3)
[0065] Here, cat represents concatenation along dimension 1, and ⊙ represents element-by-element multiplication. Dimension 1 is the feature dimension, and dimension 0 is the batch dimension. That is, in this embodiment, features from multiple views are concatenated along the feature vector of each sample.
[0066] In this embodiment, a binary mask m is randomly set for each view. v , and the m v The missing view is constructed by element-wise multiplication with the corresponding view. This mask m v is a vector where each element corresponds to a view feature. If the view feature is missing, the mask element is 0, otherwise it is 1. In this way, the missing conditions of different views can be simulated, further enhancing the robustness and adaptability of the model.
[0067] S1032: Inputting the original multi-view feature representation of the target object into the Transformer encoder-decoder to complete the features at the view missing mark, thereby obtaining a complete multi-view feature representation of the target object;
[0068] Specifically, the Transformer has excellent sequence modeling capabilities. This embodiment uses the Transformer to fully capture the complex dependencies between views. On this basis, it accurately completes the missing views and reshapes the completed views into the same shape as the original views, achieving complete reconstruction of the information. Through this strategy, the model can not only recover the missing view feature data, but also enhance the mutual correlation and complementarity of information between different views while maintaining the consistency of view features. The view completion process can be expressed as follows:
[0069] X completed =Transformer(X flattened )(4)
[0070] Among them, Transformer() is the Transformer encoder-decoder network, X completed Represents the multi-omics view data completed after the Transformer encoder-decoder, that is, the complete multi-view feature representation of the target object.
[0071] S1033: Dynamically fuse multi-view features based on the complete multi-view feature representation of the target object to obtain a fused feature of the target object;
[0072] Specifically, to better integrate information from different views, a dynamic fusion strategy is employed to deeply fuse and predict features from multiple views of the same object. This fusion strategy enables the model to adaptively weight features from each view, generating a more comprehensive and highly representative fused feature. This process not only optimizes the synergy between features but also enables the generation of multi-task prediction results, improving classification accuracy and diagnostic reliability.
[0073] Specifically, during the training phase, the fusion process includes:
[0074] Input the complete representation of any view v of the object into the classifier to obtain the probability distribution of view v belonging to each category and use it as evidence Use Dirichlet distribution for evidence Modeling is performed to obtain the Dirichlet distribution parameters With evidence The relationship between Thus, the belief quality when any view v belongs to category c is obtained according to the following formula: and the uncertainty u of the classification prediction for that view v v :
[0075]
[0076] Among them, C is the total number of given categories, Sv is the Dirichlet distribution intensity, c=1,2,…,C, represents the Dirichlet distribution parameter when view v belongs to the cth category, Evidence that view v belongs to the cth category; v = 1, 2, ..., V;
[0077] The belief quality according to view v and uncertainty u v , calculate the weight w of view v v :
[0078]
[0079] Among them, u v is a normalization factor used to adjust the scale of the weights;
[0080] According to the weight of each view, all views under the target object are fused according to the following formula to obtain the fusion feature h of the target object:
[0081]
[0082] Among them, V represents the number of views, X completed [ :,v,: ] is the complete multi-view feature representation X of the target object completed The feature representation corresponding to the view v in the figure.
[0083] It is understandable that in the training phase, the goal is to obtain the optimal Dirichlet distribution parameters through training sample objects.
[0084] S1034: Input the fusion features of the target object into the classifier to obtain the cancer type of the target object.
[0085] Specifically, the fused features of the target object are input into the classifier. The classifier can accurately identify and classify the cancer type based on the fused features. The process can be expressed as follows:
[0086] CA=Classification(h,C)(10)
[0087] Among them, Classification1(h,C) is the accurate cancer classifier, and C is the number of categories of different cancer diseases.
[0088] In the cancer precise localization method provided in an embodiment of the present invention, a Transformer encoder-decoder architecture is used to complete missing views. This architecture can effectively utilize known information to accurately predict and complete missing views. Through the powerful modeling capabilities of Transformer, the modality-invariant content across different omics data can be retained as much as possible, that is, those cancer features that remain consistent in different omics views. In addition, by introducing a dynamic fusion strategy based on Dirichlet distribution, information from different omics views can be flexibly integrated. Dirichlet distribution allows each omics view to be assigned a weight, which reflects the relative importance of different views in the classification task. By dynamically adjusting these weights, the information relationship between different views can be better captured, thereby improving the accuracy of classification.
[0089] Therefore, the cancer localization method provided by the embodiments of this invention reduces the loss of valuable task-related information through view gap completion based on the Transformer encoder-decoder, and integrates information from different omics views using a dynamic fusion method based on the Dirichlet allocation. This method can more comprehensively utilize multi-omics data and improve the accuracy of cancer classification.
[0090] In one embodiment, taking breast cancer as an example, the training process of the cancer precise positioning model is described as follows:
[0091] S201: Construct a multi-omics view data training set and unify the feature dimensions of all view data in the training set;
[0092] Specifically, the training set contains cancer multi-omics view data of several subjects. The multi-omics view data of each subject covers different cancer disease signals and fully demonstrates the multidimensional characteristics of cancer. This embodiment takes breast cancer as an example. In the multi-omics view data processing of breast cancer, this embodiment extracts the mRNA data, miRNA data and DNA methylation data in epigenomics of the training subjects as the three views required for the experiment. These three views provide different levels of biological information for the precise positioning of breast cancer, covering both the dynamic changes in gene expression and the deep-level regulation of genetic phenotypes. The cancer precise positioning model can effectively integrate the features of these heterogeneous views, thereby achieving precise positioning of breast cancer. In this embodiment, the classification results of breast cancer are divided into four types: Luminal A, Luminal B, HER-2 positive and triple negative breast cancer.
[0093] S202: For each training subject, the training set is processed according to step S102 in the above embodiment, and the processed training set is input into a preset cancer precise positioning network;
[0094] S203: Designing a loss function and optimizing the parameters of the cancer precise localization network by minimizing the loss function until all view data of the training objects are trained or the set number of iterations is reached;
[0095] Specifically, in the training process of the cancer precise positioning model, the loss function L used in this embodiment is total for:
[0096]
[0097] Among them, L is the loss corresponding to the classifier, L completion is the loss corresponding to the Transformer encoder-decoder, L Fusion is the loss corresponding to dynamic fusion, w v is the vth weight vector of the model (also represents the weight vector corresponding to view v), ||w v ||2 is the weight vector w v The L2 norm of m is the weight vector w v , V is the total number of weight vectors in the model, λ1 and λ2 are the preset coefficients of the corresponding losses, and λ3 is the regularization coefficient; λ1, λ2, and λ3 are used to control the weights of their respective losses.
[0098] Specifically, the losses corresponding to the classifier and dynamic fusion are both cross-entropy losses, and the losses corresponding to the Transformer encoder-decoder are mean square error losses. The formulas are as follows:
[0099]
[0100] Among them, p i,j is the predicted probability that the i-th object (i.e., patient) belongs to the j-th class.
[0101] L=CE(CA,y)(13)
[0102] Among them, CE(·,·) represents the calculation of the cross entropy loss between the cancer accurate classification result and the sample label, CA represents the cancer accurate classification result, and y represents the sample label.
[0103] The classifier's corresponding loss function aims to minimize classification error, ensuring that the model can effectively distinguish different cancer types during multi-task learning, effectively optimizing classification results and ultimately achieving accurate cancer diagnosis. This process not only ensures that the classifier can accurately process complex multi-omics data, but also effectively improves the reliability of diagnostic results, providing strong technical support for clinical applications.
[0104]
[0105] Among them, MSE(·,·) represents the mean square error loss between the view completed by the Transformer encoder-decoder and the original view, L completion represents the sum of all view completion losses.
[0106] In this embodiment, the trained cancer precision localization model can not only achieve efficient breast cancer type classification, but also conduct in-depth analysis of the occurrence and development mechanism of cancer from multiple perspectives, further improving the accuracy of the classification results and the clinical application value.
[0107] It is understood that the above training process is merely an illustrative example using a breast cancer dataset and is not intended to be limiting. The cancer precision localization solution provided by the present invention is applicable to multi-omics datasets focused on highly heterogeneous diseases, particularly the TCGA dataset, which encompasses multidimensional physiological views such as the genome, transcriptome, and proteome, providing valuable information resources for a deeper understanding of the complex mechanisms of cancer.
[0108] Based on the same inventive concept, Figure 3 As shown, an embodiment of the present invention further provides a cancer precise positioning system based on multi-omics view missing completion and dynamic fusion, including a data acquisition and preprocessing unit, a data mask generation unit and a cancer precise positioning unit.
[0109] Among them, the data acquisition and preprocessing unit is used to obtain the cancer multi-omics view data of the target object and unify the feature dimensions of all the acquired view data; the data mask generation unit is used to generate a mask matrix to mark the view missing status of the target object; wherein, one element in the mask matrix corresponds to the missing status of one view feature; the cancer precise positioning unit is used to input all the view data and mask matrix after unification into the preset cancer precise positioning model to obtain the cancer type of the target object.
[0110] Specifically, the preset cancer precision positioning model includes a view missing labeling module, a multi-view missing completion module, a dynamic fusion module and a classification learning module. Among them, the view missing labeling module is used to use the mask matrix to process all view data after the unified feature dimension to obtain the original multi-view feature representation of the target object; wherein the original multi-view feature representation carries a view missing label; the multi-view missing completion module is used to input the original multi-view feature representation of the target object into the Transformer encoder-decoder to complete the features at the view missing label, thereby obtaining a complete multi-view feature representation of the target object; the dynamic fusion module is used to dynamically fuse the multi-view features based on the complete multi-view feature representation of the target object to obtain the fused features of the target object; the classification learning module is used to input the fused features of the target object into the classifier to obtain the cancer type of the target object.
[0111] In the cancer precise localization model designed by the cancer precise localization system provided by the embodiment of the present invention, multiple modules collaborate with each other and significantly improve the accuracy and reliability of the model in cancer precise localization by fully exploring the complementarity and dynamic fusion of multi-omics view data.
[0112] It should be noted that the cancer precise positioning system provided in the embodiments of the present invention is intended to implement the above-mentioned method. Its specific functions can be referred to the above-mentioned method embodiments and will not be described in detail here.
[0113] In order to verify the effectiveness of the solution of the present invention, the present invention also provides the following experimental data.
[0114] The method proposed in this invention (MVMCDF, Multi-view missing completion and dynamic fusion) is compared with existing view missing fusion methods (LHGN, CPM-Net); among them, LHGN explores the complex relationship between samples by constructing a heterogeneous graph and uses graph learning technology to mine the structural information in the latent space to achieve good classification results; CPM-Net constructs a common latent space to process data with arbitrary view missing patterns and uses encoding networks and clustering-like classification loss to mine the structural information in the latent space, thereby achieving effective classification of incomplete multi-view data. Experiments were conducted on each method using the same multi-omics dataset BRCA and OV, while controlling the classifiers of the method of the present invention and the comparison method to be the same. The results are shown in Tables 1 and 2 below. Table 1 shows the experimental results of the precise localization of breast cancer (BRCA) for each model. Table 2 shows the experimental results of the precise localization of ovarian cancer (OV) for each model.
[0115] Table 1 Experimental results of breast cancer (BRCA) precise positioning based on various models
[0116]
[0117] Table 2 Experimental results of precise positioning of ovarian cancer (OV) based on various models
[0118]
[0119] As shown in the table above, the proposed Multi-View Cancer Precision Localization Model (MVMCDF) performs well across various multi-omics cancer datasets. Compared to existing methods, MVMCDF significantly improves cancer classification accuracy. Comparing LHGN with MVMCDF reveals that the addition of dynamic fusion significantly improves cancer classification performance.
[0120] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4As shown, the electronic device may include: a processor (processor) 401, a communication interface (Communications Interface) 402, a memory (memory) 403 and a communication bus 404, wherein the processor 401, the communication interface 402, and the memory 403 communicate with each other through the communication bus 404. The processor 401 can call logic instructions in the memory 403 to execute a cancer precise positioning method based on multi-omics view missing completion and dynamic fusion, the method comprising: obtaining cancer multi-omics view data of a target object and unifying the feature dimensions of all the obtained view data; generating a mask matrix to mark the view missing status of the target object; wherein one element in the mask matrix corresponds to marking the missing status of one view feature; inputting all the view data after the unified dimension and the mask matrix into a preset cancer precise positioning model to obtain the cancer type of the target object; wherein the processing process of the preset cancer precise positioning model comprises: using the mask matrix to process all the view data after the unified feature dimension to obtain an original multi-view feature representation of the target object; wherein the original multi-view feature representation carries a view missing mark; inputting the original multi-view feature representation of the target object into a Transformer encoder-decoder to complete the features at the view missing mark, thereby obtaining a complete multi-view feature representation of the target object; dynamically fusing the multi-view features based on the complete multi-view feature representation of the target object to obtain a fused feature of the target object; and inputting the fused feature of the target object into a classifier to obtain the cancer type of the target object.
[0121] In addition, when the logic instructions in the above-mentioned memory 403 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0122] An embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the cancer precise positioning method based on multi-omics view missing completion and dynamic fusion provided by the above-mentioned method embodiments.
[0123] An embodiment of the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for precise cancer localization based on multi-omics view missing completion and dynamic fusion provided by the above-mentioned method embodiments is implemented.
[0124] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A cancer precise localization method based on multi-omics view deletion completion and dynamic fusion, characterized by: include: Obtain cancer multi-omics view data of the target object and unify the feature dimensions of all the acquired view data; Generate a mask matrix to mark the view missing of the target object; wherein one element in the mask matrix corresponds to the missing state of one view feature; Inputting all the unified view data and mask matrix into a preset cancer precise positioning model to obtain the cancer type of the target object; wherein the processing process of the preset cancer precise positioning model includes: Using the mask matrix to process all view data after unifying the feature dimension, to obtain an original multi-view feature representation of the target object; wherein the original multi-view feature representation carries a view missing flag; The original multi-view feature representation of the target object is input into the Transformer encoder-decoder to complete the features at the view missing mark, thereby obtaining a complete multi-view feature representation of the target object; Dynamically fuse multi-view features based on the complete multi-view feature representation of the target object to obtain the fused features of the target object; The fused features of the target object are input into the classifier to obtain the cancer type of the target object.
2. The method for precise cancer localization based on multi-omics view deletion completion and dynamic fusion according to claim 1, characterized in that: The mask vector is used to process all view data after unifying the feature dimension to obtain the original multi-view feature representation of the target object, specifically including: Set the mask matrix of the target object to mask = {m v |v=1,2,...,V}; where V represents the number of views, m v Represents a vector used to mark the missing state of view v; all view data after setting the unified dimension are X′={X′ v |v=1,2,...,V}; where X′ v Represents the feature representation corresponding to view v; then the original multi-view feature representation of the target object is X flattened =cat(X′1⊙m1,X′2⊙m2,...,X′ V ⊙m V ), where cat represents concatenation on dimension 1 and ⊙ represents element-wise multiplication.
3. The method for precise cancer localization based on multi-omics view deletion completion and dynamic fusion according to claim 1, characterized in that: Based on the complete multi-view feature representation of the target object, the multi-view features are dynamically fused to obtain the fused features of the target object, including: Input the complete representation of any view v of the object into the classifier to obtain the probability distribution of view v belonging to each category and use it as evidence Use Dirichlet distribution for evidence Modeling is performed to obtain the Dirichlet distribution parameters With evidence The relationship between Thus, the belief quality when any view v belongs to category c is obtained according to the following formula: and the uncertainty u of the classification prediction for that view v v : Among them, C is the total number of given categories, S v is the Dirichlet distribution intensity, c=1,2,…,C, represents the Dirichlet distribution parameter when view v belongs to the cth category, Evidence that view v belongs to the cth category; v = 1, 2, ..., V; The belief quality according to view v and uncertainty u v , calculate the weight w of view v v : Among them, u v is the normalization factor; According to the weight of each view, all views under the target object are fused according to the following formula to obtain the fusion feature h of the target object: Among them, V represents the number of views, X completed[:,v,:] is the complete multi-view feature representation X of the target object completed The feature representation corresponding to the view v in the figure.
4. The method for precise cancer localization based on multi-omics view deletion completion and dynamic fusion according to claim 1, characterized in that: During the training of the cancer precise positioning model, the loss function L is used. total for: Among them, L is the loss corresponding to the classifier, L completion is the loss corresponding to the Transformer encoder-decoder, L Fusion is the loss corresponding to dynamic fusion, w v is the vth weight vector of the model, ||w v ||2 is the weight vector w v , V is the total number of weight vectors in the model, λ1 and λ2 are the preset coefficients of the corresponding losses, and λ3 is the regularization coefficient.
5. The method for precise cancer localization based on multi-omics view deletion completion and dynamic fusion according to claim 1, characterized in that: The losses corresponding to the classifier and dynamic fusion are both cross entropy loss, and the loss corresponding to the Transformer encoder-decoder is mean square error loss.
6. Cancer precise positioning system based on multi-omics view deletion completion and dynamic fusion, characterized by: include: A data acquisition and preprocessing unit, configured to acquire cancer multi-omics view data of a target object and unify the feature dimensions of all acquired view data; A data mask generation unit is used to generate a mask matrix to mark the view missing status of the target object; wherein one element in the mask matrix corresponds to marking the missing status of one view feature; A cancer precise positioning unit is configured to input all the unified view data and the mask matrix into a preset cancer precise positioning model to obtain the cancer type of the target object; the preset cancer precise positioning model includes a view missing labeling module, a multi-view missing completion module, a dynamic fusion module, and a classification learning module; a view missing marking module, configured to process all view data after unifying the feature dimensions using the mask matrix to obtain an original multi-view feature representation of the target object; wherein the original multi-view feature representation carries a view missing mark; The multi-view missing completion module is used to input the original multi-view feature representation of the target object into the Transformer encoder-decoder to complete the features at the view missing mark, thereby obtaining a complete multi-view feature representation of the target object; A dynamic fusion module is used to dynamically fuse multi-view features based on the complete multi-view feature representation of the target object to obtain the fused features of the target object; The classification learning module is used to input the fusion features of the target object into the classifier to obtain the cancer type of the target object.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.