Adenocarcinoma image grading prediction method based on multi-modal data fusion and fuzzy measurement

Through multimodal data fusion and fuzzy measurement methods, the problems of incompleteness and uncertainty of data information in traditional adenocarcinoma grading prediction methods are solved, the accuracy and reliability of prediction are improved, and the real-time prediction needs of clinical applications are met.

CN120047391APending Publication Date: 2025-05-27NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510055198.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Traditional adenocarcinoma grading prediction methods rely on single modal data, resulting in incomplete information, limited feature characterization ability, difficult to process heterogeneity and uncertainty of data, and low computational efficiency, making it difficult to meet the needs of clinical applications.

Method used

Adenocarcinoma hierarchical prediction method based on multimodal data fusion and fuzzy metrics is adopted. By fusing CT imaging omics data and DNA methylation data, and combining fuzzy metrics, the uncertainty and fuzzy information in the data are processed to improve the generalization ability and prediction performance of the model.

Benefits of technology

It improves the accuracy and reliability of adenocarcinoma grading prediction, enhances the generalization ability and robustness of the model, optimizes the overall performance of the image grading prediction model, and meets the real-time prediction needs of clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047391A_ABST
    Figure CN120047391A_ABST
Patent Text Reader

Abstract

The invention discloses an adenocarcinoma image grading prediction method based on multi-modal data fusion and fuzzy measurement, and aims to solve the problems of data source singleness and data fuzzy uncertainty in adenocarcinoma grading prediction. According to the method, by fusing CT image omics data and DNA methylation data and combining fuzzy measurement, uncertainty and fuzziness information in the data is better processed, and therefore the generalization ability and robustness of a classification model are improved. According to the method, the overall performance of the cancer grading prediction model is effectively optimized, and the accuracy and usability of the model are improved. According to the feature fusion method, medical data in multiple fields can be utilized for data integration, the problem that data in a certain field is sparse can be relieved, and therefore the coverage rate of a classification model is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical image data processing and prediction, and particularly relates to a method for predicting adenocarcinoma grading based on multi-modal data fusion and fuzzy metrics. Background Art

[0002] Traditional methods for predicting adenocarcinoma grading mainly rely on single-modal data, such as pathological images or gene expression data. This method has significant limitations in information expression and prediction accuracy. To address this challenge, methods for predicting adenocarcinoma grading based on multi-modal data fusion have gradually become a research hotspot. By integrating data from different sources, it is possible to more comprehensively characterize the biological characteristics of tumors and improve the accuracy and reliability of grading prediction.

[0003] Currently, the methods for adenocarcinoma grading based on single-modal data mainly face the following several significant problems. First, single-modal data has problems of incomplete information and limited characterization ability when expressing tumor characteristics, and it is difficult to fully reflect the complex biological mechanisms of adenocarcinoma. Second, there is data heterogeneity between different modal data, and how to effectively fuse this heterogeneous data to improve prediction performance remains a technical challenge. Third, traditional grading methods perform poorly in dealing with uncertain and fuzzy information, and it is difficult to handle the fuzziness and uncertainty of data in clinical practice, resulting in insufficient stability and reliability of prediction results. Finally, the existing methods have low computational efficiency in large-scale data processing and real-time prediction, and it is difficult to meet the needs of clinical applications.

[0004] The method for predicting adenocarcinoma grading based on multi-modal data fusion and fuzzy metrics emerges as the times require, aiming to construct a more comprehensive and refined tumor characterization model by comprehensively using multi-source information of CT radiomics and DNA methylation. Specifically, multi-modal data fusion technology can effectively integrate data from different sources, overcome the challenges brought by data heterogeneity, and improve the generalization ability and prediction performance of the model. At the same time, introducing the fuzzy metric method can better handle the uncertain and fuzzy information in the data and improve the stability and credibility of the prediction results. Although the method for predicting adenocarcinoma grading based on multi-modal data fusion and fuzzy metrics has significant advantages in theory, it still faces many challenges in practical applications: 1. How to design an efficient multi-modal data fusion algorithm to fully explore and utilize the complementary information of each modal data; 2. How to accurately measure and process the fuzzy information in the data to improve the robustness of the model to uncertain data; 3. How to optimize the computational efficiency and model training time on large-scale data sets to meet the needs of clinical real-time prediction; 4. How to ensure the data privacy and security of multi-modal data during the fusion process to ensure the confidentiality of information. Addressing these challenges, the research on the method for predicting adenocarcinoma grading based on multi-modal data fusion and fuzzy metrics not only has important theoretical value but also shows broad prospects in clinical applications. Summary of the Invention

[0005] Object of the Invention: The object of the present invention is to provide an adenocarcinoma image grading prediction method based on multi-modal data fusion and fuzzy measurement. It aims to solve the problems of single data source and data fuzzy uncertainty in adenocarcinoma grading prediction. By fusing CT radiomics data and DNA methylation data and combining fuzzy measurement, it can better process the uncertainty and fuzzy information in the data, thereby improving the generalization ability and robustness of the classification model. This method effectively optimizes the overall performance of the image grading prediction model and improves the accuracy and usability of the model.

[0006] Technical Solution: An adenocarcinoma grading prediction method based on multi-modal data fusion and fuzzy measurement of the present invention includes the following steps:

[0007] S1. Data collection and preprocessing: Collect CT radiomics data and DNA methylation data as the original data set, and preprocess the original data set;

[0008] S2. Data fusion: Standardize the features of each type of data in the data set respectively, adjust the features of each data set according to the number of features of the two types of data to balance the contribution, and then fuse the two types of data features;

[0009] S3. Fuzzy measurement: First, calculate the membership degree of each feature value to the grading label for the data after fusing the two types of data features, and use the Gaussian membership function for membership degree calculation; then normalize the membership degree so that the sum of the membership degrees of each sample on all categories is 1, and finally splice the two membership degree matrices with the fused data to obtain the final data;

[0010] S4. Model training: After obtaining the final fused data containing membership degrees, use lasso regression for feature selection and training of the model;

[0011] S5. Model evaluation: Evaluate the model performance and compare it with other models by calculating the AUC value and drawing the ROC curve.

[0012] Further, step S1 specifically includes the following steps:

[0013] S101: Clean the two datasets of the obtained CT imaging omics data and DNA methylation data respectively, delete the samples without grading labels and delete the target names that appear repeatedly; the information contained in the data is the target ID, CT imaging omics features, DNA methylation features, and target grading labels respectively; the CT imaging omics features include tumor image shape, tumor image first-order features, gray-level co-occurrence matrix features GLCM, gray-level run length matrix features GLRLM, and gray-level zone size matrix features GLSZM, and there are 77 features in total for these five types of features; the DNA methylation features include the end motif features analyzed from plasma cfDNA in the 5mC and 5hmC enrichment regions, including 256 - 4bp and 4098 - 6bp features respectively;

[0014] S102: To exclude the influence of uncertain grading on the data, select and discard those target samples with non-unique target sample grading to ensure the certainty of data labels and improve the generalization ability of the model.

[0015] Further, step S2 specifically includes the following steps:

[0016] S201: Import the CT imaging omics data and DNA methylation to obtain feature data and grading label data, and at the same time delete the target samples with only single-modal data to find common samples;

[0017] S202: Standardize the data features of each modality respectively, so that the data is converted into a distribution with a mean of 0 and a standard deviation of 1 to avoid the algorithm being affected by different feature scales;

[0018] S203: To solve the problem that the different numbers of data features in the two modalities lead to different contributions to the model, calculate the scaling factors of the two types of data respectively according to the number of features to balance the data contributions;

[0019] S204: To increase the diversity of features and enhance the robustness of the model, construct new feature representations for the data after balancing the contributions and perform data splicing and fusion.

[0020] Further, step S3 specifically includes the following steps:

[0021] S301: After fusing the CT imaging omics data and DNA methylation data, calculate the mean and standard deviation of each category;

[0022] S302: Obtain the membership matrix of features by calculating the fuzzy membership degree of each feature value to the grading label, and use Gaussian membership degree to calculate the fuzzy membership degree;

[0023] S303: To ensure that the sum of the membership degrees of each sample on all grading labels is 1, normalize the fuzzy membership degrees calculated using Gaussian membership degrees.

[0024] Further, step S4 specifically includes the following steps:

[0025] S401: Before model training, split the dataset into a training set and a test set. Set the test set to contain 20% of the original dataset. At the same time, set a random seed to ensure the same training set and test set division each time the code is run, guaranteeing the repeatability of the experiment;

[0026] S402: Use the lasso regression model for training and prediction.

[0027] Further, step S5 specifically includes the following steps:

[0028] S501: According to the results of step S4, use two metrics commonly used to measure performance in classification tasks: the ROC curve and the AUC value to evaluate the classification accuracy and overall performance of the model.

[0029] Further, in step S202, the specific method of standardizing the data features of each modality is as follows:

[0030]

[0031] Among them, X is the feature value after standardization, x is the original feature value, μ is the average value of all values of this feature, that is, the mean value, and σ is the standard deviation of all values of this feature;

[0032] In step S203, the specific method of calculating the scaling factors of the two types of data according to the number of features to balance the data contribution is as follows:

[0033]

[0034] Among them, X (i) scaled_new is the adjusted feature matrix, X (i) scaled_old is the original feature matrix after standardization, i ∈ {ct, methylation} represents different datasets, and n i refers to the number of features of the corresponding dataset i, that is, n ct or n methylation ;

[0035] In step S204, the specific method of constructing a new feature representation for the data after balancing the contribution and performing data splicing and fusion is as follows:

[0036]

[0037] where min(X 1 , X 2 ) represents obtaining the minimum value element-wise, max(X 1 , X 2 ) represents obtaining the maximum value element-wise, represents obtaining the average value element-wise, and the semicolon (;) represents concatenating these newly generated feature matrices column-wise to finally form the fused feature matrix X fused , and the operations of obtaining the minimum, maximum, and average values of the features element-wise capture different aspects of the input data while retaining the information of each feature value.

[0038] Furthermore, in step S301, the calculation of the mean and standard deviation of each category is specifically as follows:

[0039]

[0040] where C c represents the set of samples belonging to category c. For each feature i and each category c ∈ {0, 1}, category 0 represents the low-grade category and category 1 represents the high-grade category, x j is the feature value of sample j, and |C c | is the number of samples in the graded category c;

[0041] In step S302, the Gaussian membership degree is specifically as follows:

[0042]

[0043] where g is the membership degree of each feature value, x is the feature value of the sample, μ c is the mean of this feature under a specific category, and σ c is the standard deviation of this feature under a specific category;

[0044] In step S303, the normalization of the fuzzy membership degree calculated using the Gaussian membership degree is specifically as follows:

[0045]

[0046] where g norm,c (x) represents the normalized Gaussian membership degree, and 1e-6 is added to the denominator to avoid a denominator of 0.

[0047] Furthermore, in step 402, the lasso regression model is specifically as follows:

[0048]

[0049] where, represents the optimal parameter estimate value obtained through optimization, yi is the true value of the i-th sample, x ij is the j-th feature value of the i-th sample, β 0 is the bias, which is usually not regularized, β j is the coefficient corresponding to the j-th feature, and λ is the regularization strength parameter, which controls the regularization intensity and λ > 0. Set λ to 0.05 during model training;

[0050] The performance of the model is demonstrated by the ROC curve by plotting the relationship between the true positive rate (TPR) and the false positive rate (FPR).

[0051] Furthermore, the performance of the model demonstrated by the ROC curve by plotting the relationship between the true positive rate (TPR) and the false positive rate (FPR) is specifically as follows:

[0052]

[0053] Among them, TP represents the number of samples that are actually positive and are correctly predicted as positive, FN represents the number of samples that are actually positive but are incorrectly predicted as negative, FP represents the number of samples that are actually negative but are incorrectly predicted as positive, and TN represents the number of samples that are actually negative and are correctly predicted as negative.

[0054] The true positive rate is also called sensitivity or recall. It can represent the proportion of samples that are actually positive and are correctly predicted as positive. In the work, it represents the proportion of all target samples that are actually highly graded and are correctly predicted as highly graded; the false positive rate represents the proportion of samples that are actually negative and are incorrectly predicted as positive. In the work, it represents the proportion of all target samples that are actually lowly graded and are incorrectly predicted as highly graded.

[0055] The AUC value is obtained by calculating the area under the ROC curve. The closer the AUC value is to 1.0, the better the discrimination ability of the model.

[0056] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages:

[0057] (1) Improve the tumor image feature representation ability: The feature fusion method proposed by the present invention can fuse medical features from different sources, including multi-modal information such as CT radiomics features and DNA methylation, thereby improving the accuracy of the classification model.

[0058] (2) Enhance the generalization ability of the model: Introducing the fuzzy metric can better handle the uncertainty and fuzzy information in the data, enhancing the generalization ability and prediction performance of the model.

[0059] (3) Alleviate the problems of incomplete data information and single data source: The feature fusion method of the present invention can use medical data from multiple fields for data integration, which helps to alleviate the problem of sparse data in a certain field, thereby improving the coverage rate of the classification model. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is the system flowchart of the present invention.

[0061] Figure 2 It is the model structure diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0062] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0063] As Figure 2 shown, the model of the present invention constructs a graph structure based on CT imaging omics data and DNA methylation data. First, the CT imaging omics data and DNA methylation data are respectively preprocessed by data cleaning. Then, the model obtains more abundant fused feature information by obtaining the minimum eigenvalue, maximum eigenvalue and average value for data feature fusion. Next, the fuzzy measure of the fused eigenvalue is calculated to solve the data fuzzy problem. These spliced membership feature matrices enrich the feature representation of tumors and the fuzzy relationship with grading, further improving the representation quality and discrimination ability of features. Then, training and testing are carried out through the lasso model to predict the adenocarcinoma grade. Such an architecture design can make full use of different modality medical data information, capture the fuzzy relationship between feature data and cancer grading, and improve the accuracy and generalization ability of the classification model.

[0064] As Figure 1 shown, a method for predicting adenocarcinoma grade based on multi-modal data fusion and fuzzy measure of the present invention includes the following steps:

[0065] S1: Data preprocessing. In order to verify the effectiveness of the adenocarcinoma grade prediction method, the present invention uses CT imaging omics data and DNA methylation data as the original data set and preprocesses the data. The data of the present invention comes from Wuxi People's Hospital and the original 5mC and 5hmC sequencing data of this study have been stored in the National Genomics Data Center (Genome Sequence Archive, GSA).

[0066] S2: Data fusion. First, obtain the common samples in the CT imaging omics data and DNA methylation data. Secondly, standardize the features of each type of data respectively. Then, adjust the features of each data set according to the number of features of the two types of data to balance the contributions. Finally, fuse the features of the two types of data.

[0067] S3: Fuzzy metric: First, calculate the membership degree of each eigenvalue to the classification label for the data after fusing the two data features in the present invention, and use the Gaussian membership function for membership degree calculation. Then, the present invention normalizes the membership degrees so that the sum of the membership degrees of each sample for all categories is 1. Finally, splice the two membership degree matrices with the fused data to obtain the final data.

[0068] S4: Model training: After obtaining the final fused data containing membership degrees, the present invention uses lasso regression for feature selection and training the model.

[0069] S5: Model evaluation: Evaluate the performance of the model in the present invention and compare it with other models by calculating the AUC value and plotting the ROC curve.

[0070] Specifically, S2 is as follows: The rule for the present invention to standardize the features is:

[0071]

[0072] Among them, X is the eigenvalue after standardization, x is the original eigenvalue, μ is the average value of all values of the feature, that is, the mean value, and σ is the standard deviation of all values of the feature.

[0073] When balancing data from different sources, the rule for the present invention to balance the data contribution is:

[0074]

[0075] Among them, X (i) scaled_new is the adjusted feature matrix, X (i) scaled_old is the original feature matrix after standardization, i ∈ {ct, methylation} represents different data sets, and n i refers to the number of features corresponding to the data set i (that is, n ct or n methylation ).

[0076] The rule for the present invention to construct a new feature representation and perform data splicing and fusion is:

[0077]

[0078] Among them, min(X 1 , X 2 ) represents obtaining the minimum value element by element, max(X 1 , X 2 ) represents obtaining the maximum value element by element, represents obtaining the average value element by element, and the semicolon (;) represents splicing these newly generated feature matrices by column, and finally forming the fused feature matrix Xfused The operation of the present invention to obtain the minimum, maximum, and average value of features element by element can capture different aspects of the input data while retaining the information of each feature value, which helps the model better understand the data.

[0079] Specifically, S3 is as follows: The rules for the present invention to calculate the mean and standard deviation are:

[0080]

[0081] where C c represents the set of samples belonging to class c. For each feature i and each class c ∈ {0, 1}, class 0 represents the low-grade class, class 1 represents the high-grade class, x j is the feature value of sample j, and |C c | is the number of samples in the graded class c.

[0082] The rules for the present invention to calculate the membership degree are:

[0083]

[0084] where g is the membership degree of each feature value, x is the feature value of the sample, μ c is the mean of this feature under a specific class, and σ c is the standard deviation of this feature under a specific class.

[0085] The rules for the present invention to normalize the fuzzy membership degree are:

[0086]

[0087] where g norm,c (x) represents the normalized Gaussian membership degree. Adding 1e-6 to the denominator is to avoid the situation of the denominator being zero.

[0088] Specifically, S4 is as follows: The rules for Lasso regression are:

[0089]

[0090] where represents the optimal parameter estimate value obtained through optimization, y i is the true value of the i-th sample, x ij is the j-th feature value of the i-th sample, β 0 is the bias, which is usually not regularized, β j is the coefficient corresponding to the j-th feature, and λ is the regularization strength parameter, which controls the strength of regularization and λ > 0. λ is set to 0.05 during model training.

[0091] The ROC curve shows the performance of the model by plotting the relationship between the True Positive Rate (TPR) and the False Positive Rate (FPR):

[0092]

[0093] Among them, TP represents the number of samples that are actually positive and are correctly predicted as positive, FN represents the number of samples that are actually positive but are wrongly predicted as negative, FP represents the number of samples that are actually negative but are wrongly predicted as positive, and TN represents the number of samples that are actually negative and are correctly predicted as negative.

[0094] The True Positive Rate is also known as sensitivity or recall. It can represent the proportion of samples that are actually positive and are correctly predicted as positive. In the work, it represents the proportion of all target samples that are actually highly graded and are correctly predicted as highly graded; the False Positive Rate represents the proportion of samples that are actually negative and are wrongly predicted as positive. In the work, it represents the proportion of all target samples that are actually lowly graded and are wrongly predicted as highly graded.

[0095] The AUC value is obtained by calculating the area under the ROC curve. The closer the AUC value is to 1.0, the better the discrimination ability of the model.

[0096] The present invention proposes an adenocarcinoma grading prediction method based on multi-modal data fusion and fuzzy metric, which combines technologies such as multi-modal data fusion and fuzzy metric, alleviates the problems inherent in the classification model prediction task in the field of adenocarcinoma grading prediction, such as incomplete data information, limited feature representation ability, and strong data uncertainty, enhances the understanding of the relationship between tumor features and cancer grading by the classification model, and improves the generalization ability of the model and the accuracy of cross-domain grading prediction.

[0097] The specific implementation schemes described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific implementation schemes of the present invention and are not intended to limit the scope of the present invention. Any equivalent changes and modifications made by those skilled in the art without departing from the concept and principles of the present invention shall fall within the scope of protection of the present invention.

Claims

1. A method for predicting adenocarcinoma image grading based on multimodal data fusion and fuzzy measurement, characterized in that: The steps include: S1. Data acquisition and preprocessing: CT imaging data and DNA methylation data were collected as original data sets, and the original data sets were preprocessed; S2, data fusion: standardize the features of each data set separately, adjust the features of each data set according to the number of features of the two data sets to balance the contribution, and then fuse the features of the two data sets; S3, fuzzy measurement: First, the membership of each feature value to the classification label is calculated for the data after the fusion of the two data features, and the Gaussian membership function is used to calculate the membership; then the membership is normalized so that the sum of the membership of each sample in all categories is 1, and finally the two membership matrices are spliced ​​with the fused data to obtain the final data; S4, model training: After obtaining the final fused data containing membership, lasso regression is used for feature selection and model training; S5. Model evaluation: Evaluate model performance and compare with other models by calculating AUC value and drawing ROC curve.

2. The method for predicting adenocarcinoma image grading based on multimodal data fusion and fuzzy measurement according to claim 1, characterized in that: Step S1 specifically includes the following steps: S101: Data cleaning was performed on the two datasets of CT imaging genomics data and DNA methylation data, and samples without classification labels were deleted while repeated target names were deleted; the data contained information including target ID, CT imaging genomics features, DNA methylation features and target classification labels; CT imaging genomics features included tumor image shape, tumor image first-order features, grayscale co-occurrence matrix features GLCM, grayscale run length matrix features GLRLM and grayscale region size matrix features GLSZM, totaling 77 features in these five categories; DNA methylation features included terminal motif features analyzed from plasma cfDNA in 5mC and 5hmC enriched regions, including 256-4bp and 4098-6bp features respectively; S102: In order to eliminate the influence of uncertain classification on the data, target samples whose classification is not unique are discarded to ensure the certainty of data labels and improve the generalization ability of the model.

3. The method for predicting adenocarcinoma image grading based on multimodal data fusion and fuzzy measurement according to claim 1, characterized in that: Step S2 specifically includes the following steps: S201: Import CT radiomics data and DNA methylation data to obtain feature data and hierarchical label data, and delete target samples with only single modality data to find common samples; S202: Standardize the data features of each modality respectively, so that the data is converted into a distribution with a mean of 0 and a standard deviation of 1 to avoid the algorithm being affected by different feature scales; S203: To solve the problem that the two modal data have different contribution to the model due to different feature numbers, the scaling factors of the two types of data are calculated according to the number of features to balance the data contribution; S204: In order to increase the diversity of features and enhance the robustness of the model, a new feature representation is constructed for the data after balancing the contributions and data splicing and fusion are performed.

4. The method for predicting adenocarcinoma image grading based on multimodal data fusion and fuzzy measurement according to claim 1, characterized in that: Step S3 specifically includes the following steps: S301: After fusing the CT radiomics data and DNA methylation data, the mean and standard deviation of each category are calculated; S302: Obtaining a membership matrix of features by calculating the fuzzy membership of each eigenvalue to the classification label, and using Gaussian membership to calculate the fuzzy membership; S303: To ensure that the sum of the membership of each sample on all the classification labels is 1, the fuzzy membership calculated using the Gaussian membership is normalized.

5. The method for predicting adenocarcinoma image grading based on multimodal data fusion and fuzzy measurement according to claim 1, characterized in that: Step S4 specifically includes the following steps: S401: Before model training, the data set is divided into a training set and a test set. The test set is set to contain 20% of the original data set. At the same time, a random seed is set to ensure that the same training set and test set divisions are obtained each time the code is run; S402: Use the lasso regression model for training and prediction.

6. The method for predicting adenocarcinoma image grading based on multimodal data fusion and fuzzy measurement according to claim 1, characterized in that: Step S5 specifically includes the following steps: S501: According to the result of step S4, two indicators commonly used in classification tasks to measure performance are used: ROC curve and AUC value to evaluate the classification accuracy and overall performance of the model.

7. The method for predicting adenocarcinoma image grading based on multimodal data fusion and fuzzy measurement according to claim 3, characterized in that: In step S202, the data features of each modality are standardized respectively as follows: Among them, X is the eigenvalue after standardization, x is the original eigenvalue, μ is the average of all values ​​of the feature, and σ is the standard deviation of all values ​​of the feature; In step S203, the scaling factors of the two types of data are calculated according to the number of features to balance the data contribution, specifically: Among them, X (i) scaled_new is the adjusted feature matrix, X (i) scaled_old is the normalized original feature matrix, i∈{ct,methylation} represents different data sets, n i Refers to the number of features corresponding to data set i, that is, n ct or methylation ; In step S204, a new feature representation is constructed for the data after balanced contribution and data splicing and fusion are performed, specifically: Among them, min(X1,X2) means getting the minimum value element by element, and max(X1,X2) means getting the maximum value element by element. It means to find the average value element by element, and the semicolon (;) means to concatenate these newly generated feature matrices by column, and finally form the fused feature matrix X fused ,The operation of obtaining the minimum, maximum and average of the features element-wise captures different aspects of the input data while preserving the information of each eigenvalue.

8. The method for predicting adenocarcinoma image grading based on multimodal data fusion and fuzzy measurement according to claim 4, characterized in that: In step S301, the mean and standard deviation of each category are calculated as follows: Among them C c represents the set of samples belonging to category c. For each feature i and each category c∈{0,1}, category 0 represents a low-level category, category 1 represents a high-level category, and x j is the eigenvalue of sample j, |C c | is the number of samples in classification category c; In step S302, the Gaussian membership is specifically: Among them, g is the membership degree of each eigenvalue, x is the eigenvalue of the sample, μ c is the mean of the feature in a specific category, σ c is the standard deviation of the feature in a specific category; In step S303, the fuzzy membership calculated using the Gaussian membership is normalized, specifically: Among them, g norm,c (x) represents the normalized Gaussian membership, and 1e-6 is added to the denominator to avoid the denominator being zero.

9. The method for predicting adenocarcinoma image grading based on multimodal data fusion and fuzzy measurement according to claim 5, characterized in that: In step 402, the lasso regression model is specifically: in, represents the best parameter estimate obtained through optimization, y i is the true value of the i-th sample, x ij is the jth eigenvalue of the i-th sample, β0 is the bias, which is usually not regularized, and β j is the coefficient corresponding to the jth feature, and λ is the regularization strength parameter, which controls the strength of regularization and λ>0. In model training, λ is set to 0.05; The ROC curve shows the performance of the model by plotting the relationship between the true positive rate TPR and the false positive rate FPR.

10. The method for predicting adenocarcinoma image grading based on multimodal data fusion and fuzzy measurement according to claim 9, characterized in that: The ROC curve is used to show the performance of the model by plotting the relationship between the true positive rate TPR and the false positive rate FPR: Among them, TP represents the number of actually positive classes that are correctly predicted as positive classes, FN represents the number of actually positive classes that are incorrectly predicted as negative classes, FP represents the number of actually negative classes that are incorrectly predicted as positive classes, and TN represents the number of actually negative classes that are correctly predicted as negative classes.