Machine learning analysis method and system based on prostate tumor multi-dimensional disease data and electronic equipment

By constructing a multimodal fusion model and nonograph, the problems of invasiveness, low efficiency, and low accuracy of traditional prostate cancer diagnosis methods are solved, and efficient and accurate analysis of multidimensional prostate tumor data is achieved.

CN121983286APending Publication Date: 2026-05-05SHANGHAI EAST HOSPITAL EAST HOSPITAL TONGJI UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI EAST HOSPITAL EAST HOSPITAL TONGJI UNIV SCHOOL OF MEDICINE
Filing Date
2026-01-26
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Traditional methods for diagnosing prostate cancer, such as tissue biopsy, are invasive, and manual analysis of routine medical imaging or clinical data is inefficient and inaccurate, making it difficult to efficiently and accurately distinguish the heterogeneity of prostate tumors.

Method used

This study employs a machine learning analysis method based on multidimensional prostate tumor symptom data. By acquiring and preprocessing imaging and clinical data, a radiomics model, a deep learning model, and a clinical prediction model are established. Feature selection and fusion are performed, and a multimodal fusion model is constructed to generate a nomogram, thereby achieving automated analysis of multidimensional prostate tumor symptom data.

Benefits of technology

It improves the efficiency and accuracy of prostate tumor symptom data analysis, automates the analysis of benign and malignant prostate tumors, and enhances the ability to analyze multidimensional prostate tumor symptom data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983286A_ABST
    Figure CN121983286A_ABST
Patent Text Reader

Abstract

The invention provides a machine learning analysis method and system based on prostate tumor multi-dimensional disease data and electronic equipment. The method comprises the following steps: S1, acquiring original image data, original clinical data and pathological results of prostate dominant focus tissues; s2, screening to obtain target radiomics characteristics, and establishing a radiomics model according to the target radiomics characteristics and in combination with the pathological result; establishing a deep learning model according to the original image data in combination with a pathological result; screening to obtain target clinical features, and establishing a clinical prediction model according to the target clinical features and in combination with the pathological result; s3, carrying out fusion analysis on the image omics model, the deep learning model and the clinical prediction model aiming at the prediction result of the prostate tumor multi-dimensional disease data, screening to obtain a target fusion feature, and establishing a multi-modal fusion model according to the target fusion feature in combination with a pathological result; and S4, establishing a two-dimensional column diagram based on the multi-modal fusion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of automated medical data processing technology, and in particular relates to a machine learning analysis method, system and electronic device based on multidimensional prostate tumor disease data. Background Technology

[0002] Prostate cancer (PCa) is one of the most common malignant tumors in men worldwide. With the increasing incidence of PCa, early diagnosis and treatment have become crucial for improving patient prognosis and survival rates. However, traditional diagnostic methods for PCa, such as tissue biopsy, are invasive and have limitations, and cannot fully reflect tumor heterogeneity. Furthermore, conventional manual analysis of medical imaging or clinical data relies on manually designed radiomics features, simple textures, or clinical characteristics, which are difficult for humans to efficiently and accurately distinguish, resulting in low efficiency and accuracy in manual analysis of medical imaging or clinical data. Therefore, there is an urgent need for a solution that can automate and precisely process medical imaging or clinical data. Summary of the Invention

[0003] This invention provides a machine learning analysis method, system, and electronic device based on multidimensional prostate tumor disease data, to solve the technical problems of conventional prostate biopsy being invasive to human tissue and the low efficiency and accuracy of manual analysis of conventional medical imaging data or clinical data.

[0004] To address the above problems, the technical solution of this invention is: a machine learning analysis method based on multidimensional prostate tumor symptom data, comprising the following steps: S1: Obtain the original imaging data and original clinical data of the predominant prostate lesion tissue of the enrolled individuals, perform preprocessing on the original imaging data, and obtain the corresponding pathological results of the predominant prostate lesion tissue of the enrolled individuals in advance; S2: Filter the original image data to obtain target radiomics features, and establish a radiomics model based on the target radiomics features and pathological results; establish a deep learning model based on the original image data and pathological results; filter the original clinical data to obtain target clinical features, and establish a clinical prediction model based on the target clinical features and pathological results. S3: The prediction results of the radiomics model, the deep learning model and the clinical prediction model for multidimensional prostate tumor disease data are fused and analyzed to screen and obtain target fusion features, and a multimodal fusion model is established based on the target fusion features and combined with pathological results. S4: Based on the multimodal fusion model, a two-dimensional nomogram is established. The multimodal fusion model and the nomogram are used to perform analysis and prediction functions on the multidimensional symptom data of prostate tumors of target individuals from different approaches.

[0005] Preferably, locating the dominant prostate lesion tissue of the enrolled individual in S1 includes the following steps: S11: Digital images of multiple prostate lesions in enrolled individuals were obtained by scanning with multi-parameter magnetic resonance imaging (MRI). The largest group of prostate lesions was selected and scored according to the prostate imaging report and data system, and marked as the dominant prostate lesion of the enrolled individual.

[0006] Preferably, the raw image data categories include T2-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient images. Preprocessing of the raw image data is performed in S1, including the following steps: S12: Mark regions of interest (ROIs) in T2-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient images, including 3D and 2D formats. Extract geometric features of prostate tumors based on their 3D shapes from the ROIs in the T2-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient images, respectively. Extract intensity features of prostate tumors based on the voxel intensity distribution within the prostate tumors, respectively. Extract texture features of prostate tumors using gray-level co-occurrence matrix, gray-level run length matrix, gray-level dependency matrix, gray-level region size matrix, or neighborhood gray-level difference matrix, respectively.

[0007] Preferably, in S2, the process of filtering the raw image data to obtain the target radiomics features and filtering the raw clinical data to obtain the target clinical features includes the following steps: S21: All radiomics features were screened using t-tests and / or Mann-Whitney U tests, and radiomics features with p-values ​​<0.05 were retained; The correlation between different radiomics features was calculated using the Spearman rank correlation coefficient. When there is a ρ value ≥ 0.9 between any two sets of radiomics features, one of the redundant radiomics features was deleted. The radiomics features were screened using the LASSO-Cox regression model to obtain the target radiomics features, which were limited to 26 groups. A radiomics feature scoring algorithm was calculated based on the target radiomics features. The radiomics feature scoring algorithm was used to predict the benign or malignant risk value of prostate tumors based on the geometric, intensity, and texture features of the dominant prostate lesion tissue. S22: All clinical features were screened using baseline statistical methods, and clinical features with p-values ​​<0.05 were retained; The target clinical features are obtained by screening clinical features using a logistic regression model.

[0008] Preferably, in S22, clinical characteristics are screened using a logistic regression model, including the following steps: S221: In the Logistic regression model, the independent predictive power of any clinical feature is evaluated, the target clinical feature is selected, and the influence weight distribution of several target clinical features in the optimal predictive variable combination is calculated.

[0009] Preferably, the target clinical features in the optimal combination of predictive variables include the Prostate Imaging Reporting and Data System (PIRD) score, the ratio of free prostate-specific antigen (PSA) to total PSA, and body surface area.

[0010] Preferably, in S2, a radiomics model is established based on the target radiomics features and combined with pathological results; a deep learning model is established based on the original image data and combined with pathological results; and a clinical prediction model is established based on the target clinical features and combined with pathological results, including the following steps: S23: Input the target radiomics features and corresponding pathological results into a logistic regression model, support vector machine model, K-nearest neighbor model, decision tree model, random forest model, XGBoost model or LightGBM model, and train the model based on the 5-fold cross-validation method to establish the radiomics model; S24: Input the two-dimensional format region of interest of the original image data and the corresponding pathological results into the ResNet50 deep learning model for training, and establish the deep learning model; The deep learning features automatically extracted by the deep learning model are filtered using the LASSO-Cox regression model to obtain the target deep learning features. The number of target deep learning features is limited to 33 groups. A deep learning feature scoring algorithm is calculated based on the target deep learning features. The deep learning feature scoring algorithm is used to predict the benign or malignant risk value of prostate tumors based on the target deep learning features of the original image data. S25: Input the target clinical features and corresponding pathological results into a logistic regression model, support vector machine model, K-nearest neighbor model, decision tree model, random forest model, XGBoost model or LightGBM model, and train the model based on the 5-fold cross-validation method to establish the clinical prediction model.

[0011] Preferably, in S3, a multimodal fusion model is established based on the fusion features and combined with pathological results, including the following steps: S31: The radiomics model, the deep learning model and the clinical prediction model respectively output the benign and malignant risk value prediction results for the same prostate tumor multidimensional disease data. The three sets of benign and malignant risk value prediction results are probabilistically fused by weighted average or meta-classifier to generate a comprehensive prediction probability. The LASSO-Cox regression model is used to filter data features from different sources that affect the overall prediction probability to obtain the target fusion features. The number of target fusion features is limited to 64 groups. A comprehensive scoring algorithm is calculated based on the target fusion features. The comprehensive scoring algorithm is used to simultaneously combine the target radiomics features, the target deep learning features, and the target clinical features to predict the benign and malignant risk values ​​of prostate tumors.

[0012] Based on the same concept, the present invention also provides a machine learning analysis system based on multidimensional prostate tumor disease data, for performing the machine learning analysis method based on multidimensional prostate tumor disease data as described in any one of the above, including: The data acquisition module is used to acquire the original imaging data, original clinical data, and corresponding pathological results of the prostate-dominant lesion tissue of the enrolled individuals, and to perform preprocessing on the original imaging data. The feature extraction module is used to extract target radiomics features and target clinical features from the original image data and the original clinical data, respectively. Independent model building module, used to build radiomics models, deep learning models and clinical prediction models; The integrated model building module is used to build a multimodal fusion model based on the radiomics model, the deep learning model, and the clinical prediction model. The nodal plot generation module is used to generate a two-dimensional nodal plot based on the multimodal fusion model.

[0013] Based on the same concept, the present invention also provides an electronic device, including: a memory and a processor; The memory is used to store computer instructions; The processor is configured to invoke computer instructions stored in the memory, causing the electronic device to execute the machine learning analysis method based on multidimensional prostate tumor disease data as described in any of the above.

[0014] Because the present invention adopts the above technical solution, it has the following advantages and positive effects compared with the prior art: This invention provides a machine learning analysis method, system, and electronic device based on multidimensional prostate tumor symptom data. It deeply integrates deep learning models, radiomics features, and clinical indicators to construct a multimodal fusion model and nomogram. Based on machine learning, it processes multidimensional prostate tumor symptom data from multiple dimensions, automating the analysis of benign and malignant prostate tumor symptom data and effectively improving the efficiency and accuracy of prostate tumor symptom data analysis. Attached Figure Description

[0015] Figure 1 The present invention provides a flowchart of a machine learning analysis method based on multidimensional prostate tumor disease data; Figure 2 Example images of mpMRI annotation and ROI cropping provided by this invention: (A) Pathological result is benign; (B) Pathological result is malignant; From left to right are T2WI, ADC, DWI sequences, and the largest ROI cropping image; Figure 3 The performance of the multi-factor clinical model provided by this invention in the training set and test set is as follows: (A) (B) is the full clinical feature model; (C) (D) is the independent target clinical feature model; (E) (F) is the target clinical feature model after determining the distribution weights. Figure 4 The weight distribution of target radiomics features in the radiomics model provided by this invention is as follows: (A) is the weight distribution of target radiomics features based on T2WI sequence in the radiomics model; (B) is the weight distribution of target radiomics features based on ADC sequence in the radiomics model; (C) is the weight distribution of target radiomics features based on DWI sequence in the radiomics model. Figure 5 The Grad-CAM visualization image provided by this invention; Figure 6 The deep learning feature selection and model building process provided by this invention includes: (A) 33 target deep learning features with non-zero regression coefficients selected based on the optimal penalty coefficient λ; (B) the optimal penalty coefficient λ = 0.0168 selected based on the minimum classification error as the target and 10-fold cross-validation; (C) the weight distribution of the 33 target deep learning features in the deep learning model; and (D) the performance of the ResNet50-based deep learning model on the training and validation sets. Figure 7 The present invention provides a nomogram constructed by combining radiomics features, clinical features, and deep learning features; Figure 8 The nomogram provided by this invention provides a predicted score for the test set. Detailed Implementation

[0016] The following detailed description, in conjunction with the accompanying drawings and specific embodiments, provides a machine learning analysis method, system, and electronic device based on multidimensional prostate tumor disease data, as proposed in this invention. The advantages and features of this invention will become more apparent from the following description and claims.

[0017] First Embodiment This embodiment provides a machine learning analysis method based on multidimensional prostate tumor disease data. (See attached...) Figure 1 It includes the following steps: S1: Obtain the original imaging data and original clinical data of the predominant prostate lesion tissue of the enrolled individuals, perform preprocessing on the original imaging data, and obtain the corresponding pathological results of the predominant prostate lesion tissue of the enrolled individuals in advance; S2: Filter raw image data to obtain target radiomics features, and establish a radiomics model based on the target radiomics features and pathological results; establish a deep learning model based on raw image data and pathological results; filter raw clinical data to obtain target clinical features, and establish a clinical prediction model based on the target clinical features and pathological results. S3: The prediction results of radiomics model, deep learning model and clinical prediction model for multidimensional prostate tumor disease data are fused and analyzed, target fusion features are screened and obtained, and a multimodal fusion model is established based on the target fusion features and combined with pathological results. S4: Based on the multimodal fusion model, a two-dimensional nomogram is established. The multimodal fusion model and the nomogram are used to perform analysis and prediction functions on the multidimensional symptom data of prostate tumors of target individuals from different perspectives.

[0018] This embodiment provides a machine learning analysis method based on multidimensional prostate tumor symptom data. It deeply integrates deep learning models, radiomics features, and clinical indicators to construct a multimodal fusion model and nomogram. Based on machine learning, it analyzes multidimensional prostate tumor symptom data from multiple dimensions, realizing automated analysis of benign and malignant prostate tumor symptom data. This can effectively improve the efficiency and accuracy of prostate tumor symptom data analysis.

[0019] The following will provide a more detailed explanation of the specific steps and implementation functions of a machine learning analysis method based on multidimensional prostate tumor disease data provided in this embodiment: First, it should be noted that the multidimensional prostate cancer data used in this embodiment includes both tumor and non-tumor groups. The inclusion criteria for the tumor group are: having undergone multiparametric magnetic resonance imaging (mpMRI) of the prostate, obtaining a clear positive pathological result after ultrasound-guided biopsy / MRI-ultrasound fusion biopsy or radical prostatectomy with a Gleason score, and having well-defined prostate lesions on T2-weighted imaging (T2WI) and diffusion-weighted imaging (DWI) according to the Prostate Imaging Reporting and Data System (PI-RADS). The inclusion criteria for the non-tumor group are: having undergone mpMRI, obtaining a clear negative pathological result after ultrasound-guided biopsy / MRI-ultrasound fusion biopsy or radical prostatectomy with a Gleason score, and having well-defined prostate lesions on T2WI and DWI images according to PI-RADS.

[0020] The tumor group and the non-tumor group will be further divided proportionally to generate training sets and test sets.

[0021] Preferably, in one embodiment, locating the dominant prostate lesion tissue of the enrolled individual in S1 includes the following steps: S11: Digital images of multiple prostate lesions in enrolled individuals were obtained by scanning with multi-parameter magnetic resonance imaging (MRI). The largest group of prostate lesions was selected and scored according to the Prostate Imaging Reporting and Data System (PI-RADS), and marked as the dominant prostate lesion tissue of the enrolled individual.

[0022] Specifically, all enrolled individuals used the same type of MRI scanner to collect images of multiple prostate lesions. The images were saved in DICOM format, normalized to a resampling format, and the resolution was limited to 1mm × 1mm × 1mm. The prostate lesions were scored according to the same PI-RADS standard. When multiple prostate lesions were present, the largest lesion was selected as the dominant prostate lesion and scored accordingly.

[0023] Preferably, in one embodiment, see Figure 2 The raw image data categories include T2-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient images. Preprocessing of the raw image data is performed in S1, including the following steps: S12: Mark regions of interest (ROIs) in T2-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient images, including 3D and 2D formats. Extract geometric features of prostate tumors based on their 3D shapes from the ROIs in the T2-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient images, respectively. Extract intensity features of prostate tumors based on the voxel intensity distribution within the prostate tumors, respectively. Extract texture features of prostate tumors using gray-level co-occurrence matrix, gray-level run length matrix, gray-level dependency matrix, gray-level region size matrix, or neighborhood gray-level difference matrix, respectively.

[0024] Specifically, to ensure the consistency of multi-parameter magnetic resonance imaging (mpMRI) image quality, N4 bias field correction was first performed on all images to be labeled before image annotation. This was used to reduce intensity inhomogeneities in MRI images, thereby improving the accuracy of subsequent image analysis and feature extraction. Subsequently, regions of interest (ROIs) were labeled from T2-weighted imaging (T2WI), diffusion-weighted imaging (DWI), and apparent diffusion coefficient (ADC) images using ITK-SNAP software. Both the original image data files and the labeled ROI files were saved in nii.gz format for subsequent extraction of radiomics features. Furthermore, two-dimensional (2D) ROIs were obtained from the 3D ROIs through layer-by-layer cropping. The 2D ROIs can be used for subsequent deep learning (DL) feature extraction by deep learning models, aiming to provide high-quality input data for deep learning models and thus improve their training effectiveness and prediction performance.

[0025] Radiomic features of prostate tumors include geometric features, intensity features, and texture features. Geometric features describe the three-dimensional shape of the tumor. Intensity features describe the first-order statistical distribution of voxel intensity within the tumor. Texture features describe the pattern or second- and higher-order spatial distribution of intensity. Texture features can be extracted using various methods, including the Gray Level Co-occurrence Matrix (GLCM), Gray Level Run Length Matrix (GLLM), Gray Level Dependence Matrix (GLDM), Gray Level Size Zone Matrix (GLSZM), and Neighborhood Gray Tone Difference Matrix (NGTDM).

[0026] Radiomic features of prostate tumors can reflect multidimensional information about the tumor's morphology and texture, providing multi-faceted data support for the analysis and identification of prostate tumors.

[0027] Preferably, in one embodiment, the process of filtering raw image data to obtain target radiomics features and filtering raw clinical data to obtain target clinical features in S2 includes the following steps: S21: For the original image data, all radiomics features are first screened by t test and / or Mann-Whitney U test, and radiomics features with p value < 0.05 are retained, that is, radiomics features that can significantly distinguish feature differences from a statistical perspective are obtained. Statistical data analysis was performed using R software. Categorical variables were expressed as median (interquartile range) and frequency (%). For continuously distributed variables that conformed to a normal distribution, an independent sample t-test was used for comparison; for continuously distributed variables that did not conform to a normal distribution, the Mann-Whitney U test was used for analysis. The chi-square test (χ²) was used for comparison of categorical variables. 2 The comparison of the area under the curve (AUC) is performed using the Delong test. The Hosmer-Lemeshow test is used to evaluate the consistency between the expected and actual probabilities of the predictive model; a p-value < 0.05 is considered statistically significant.

[0028] Subsequently, the correlation between the remaining different radiomics features was calculated using the Spearman rank correlation coefficient. When there is a correlation coefficient ρ value ≥ 0.9 between any two sets of radiomics features, one set of redundant radiomics features was deleted. That is, for radiomics features with high repetition, redundant radiomics features were deleted while retaining at least one set of radiomics features. The LASSO-Cox regression model was used to screen radiomics features and obtain target radiomics features. Specifically, this embodiment used LASSO-Cox regression analysis (Least Absolute Shrinkage and Selection Operator-Cox regression). Based on minimizing classification error and 10-fold cross-validation, a penalty coefficient λ=0.0168 was obtained. Finally, 26 groups of target radiomics features with predictive ability were selected. Based on the target radiomics features, a radiomics feature scoring algorithm was calculated. The radiomics feature scoring algorithm is used to predict the benign or malignant risk value of prostate tumors based on the geometric, intensity, and texture features of the dominant prostate lesion tissue. The formula for the radiomics feature scoring algorithm is as follows: label = 0.44827586206896547 -0.023190 × original_firstorder_Kurtosis_ADC -0.057150 × original_firstorder_RootMeanSquared_ADC +0.011045× original_firstorder_Skewness_ADC +0.000523 × original_glszm_LargeAreaLowGrayLevelEmphasis_ADC +0.029502 × original_glszm_SmallAreaEmphasis_ADC +0.026008 × original_ngtdm_Complexity_ADC -0.014707× original_ngtdm_Contrast_ADC -0.024497 × original_glcm_Idn_DWI -0.020157 × original_glcm_Imc1_DWI +0.011060 × original_glcm_InverseVariance_DWI +0.049853 × original_gldm_DependenceEntropy_DWI -0.030524 × original_gldm_LargeDependenceLowGrayLevelEmphasis_DWI +0.016802× original_glszm_LargeAreaEmphasis_DWI +0.018230 × original_glszm_SmallAreaLowGrayLevelEmphasis_DWI -0.015088 × original_ngtdm_Busyness_DWI-0.058273 × original_firstorder_Minimum_T2W -0.006463 × original_glcm_ClusterShade_T2W +0.038243 × original_glcm_Correlation_T2W +0.028154 ×original_glcm_Idn_T2W +0.029155 × original_glcm_MCC_T2W +0.002607 ×original_gldm_LargeDependenceHighGrayLevelEmphasis_T2W -0.012518 ×original_glrlm_ShortRunEmphasis_T2W +0.024240 × original_glszm_GrayLevelNonUniformity_T2W -0.001345 × original_glszm_LargeAreaLowGrayLevelEmphasis_T2W -0.051004 × original_glszm_ZonePercentage_T2W -0.072488 × original_shape_Sphericity_T2W. In this example, original_firstorder_Kurtosis_ADC represents the kurtosis of the pixel value distribution extracted from the ADC image, and original_firstorder_RootMeanSquared_ADC represents the root mean square of the voxel values ​​within the lesion ROI calculated in the ADC image. The target radiomics features involved in the above radiomics feature scoring algorithm are all available in existing technologies and can be extracted using open-source radiomics toolkits such as PyRadiomics, IBEX, and LIFEx, and their meanings are clear. This embodiment will not describe them one by one.

[0029] S22: For the original clinical data, all clinical features are first screened using baseline statistical methods, and clinical features with p-values ​​<0.05 are retained, that is, clinical features that can significantly distinguish feature differences from a statistical perspective are obtained; Subsequently, the clinical features were screened using a logistic regression model to obtain the target clinical features.

[0030] Furthermore, in S22, clinical characteristics are screened using a logistic regression model, including the following steps: S221: In the Logistic regression model, the independent predictive power of any clinical feature is evaluated, target clinical features are selected, and the influence weight distribution of several target clinical features in the optimal combination of predictive variables is calculated.

[0031] Among them, the target clinical features in the optimal combination of predictive variables include the Prostate Imaging Reporting and Data System (PI-RADS) score, the ratio of free prostate-specific antigen (fPSA) to total prostate-specific antigen (tPSA), and body surface area (BSA).

[0032] Specifically, in the logistic regression model, analysis of all individual clinical characteristics showed that the PI-RADS score (OR = 1.052, 95% CI: 1.002-1.104, P = 0.008) and the fPSA / tPSA ratio (OR = 0.971, 95% CI: 0.962-0.980, P < 0.001) both exhibited highly significant statistical value, providing key clues about the nature of prostate tumor lesions and independently predicting benignity or malignancy. Meanwhile, BSA (OR = 0.879, 95% CI: 0.806-0.959, P = 0.014) also demonstrated a certain degree of independent predictive ability for benign and malignant prostate tumors. In contrast, other clinical indicators such as body mass index (BMI), prostate-specific antigen density (PSAD), prostate volume (PV), and total prostate-specific antigen (tPSA) showed limited efficacy in independently predicting benign and malignant prostate tumors. Therefore, in this embodiment, only PI-RADS score, fPSA / tPSA ratio and BSA are selected as target clinical features, and the optimal combination of predictive variables is formed by the three groups of target clinical features.

[0033] Subsequently, the influence weight distribution of the target clinical features in the optimal predictor variable combination was calculated. In the multivariate clinical model analysis (an analysis model constructed using only clinical feature data) of the optimal predictor variable combination, the predictive power of the PI-RADS score showed a highly significant enhancement, with its OR value jumping sharply to 10.097, 95% CI ranging from 6.567 to 15.534, and P value less than 0.001. This indicates that the PI-RADS score, in the multivariate model, can exert a powerful independent predictive effect far exceeding that of other clinical features, providing a core basis for the analysis of benign and malignant prostate tumors. The fPSA / tPSA ratio also maintained significant predictive value in the multivariate model (OR = 0.901, 95% CI: 0.867-0.937, P < 0.001). Although its OR value changed compared to its independent predictive analysis, this precisely reflects that in a multivariate environment, the interaction between the fPSA / tPSA ratio and other factors makes a unique contribution to the differentiation of benign and malignant prostate tumors. Although the odds ratio (OR) of BSA decreased significantly to 0.040 (95% CI: 0.007–0.210, P = 0.001) in the multivariate model, it still maintained a certain degree of statistical significance. This indicates that BSA can still play a unique auxiliary role in the prediction of prostate tumors in the multivariate model.

[0034] Therefore, see Figure 3In this embodiment, a three-stage modeling strategy is proposed (full clinical feature model → independent target clinical feature screening → determination of the weight distribution of target clinical features in the Clinic model). This strategy can quickly and accurately identify target clinical features with independent predictive functions, and correct the mutual interference between different independent target clinical features in the final Clinic model to obtain more robust effect estimates.

[0035] Preferably, in one embodiment, in S2, a radiomics model is established based on the target radiomics features and combined with pathological results; a deep learning model is established based on the original image data and combined with pathological results; and a clinical prediction model is established based on the target clinical features and combined with pathological results, including the following steps: S23: Input the target radiomics features and corresponding pathological results into a logistic regression model, support vector machine model, K-nearest neighbor model, decision tree model, random forest model, XGBoost model or LightGBM model, and train the model based on the 5-fold cross-validation method to establish a radiomics model. Specifically, logistic regression, support vector machine, K-nearest neighbor, decision tree, random forest, XGBoost, or LightGBM models were used to train the 26 selected target radiomics features, radiomics feature scoring algorithms, and pathological results. In one embodiment, the radiomics model based on the XGBoost algorithm exhibited relatively superior analytical and predictive performance. The AUC value of the XGBoost model training set reached a high level of 0.965, with a 95% confidence interval of 0.9506-0.9799, indicating that it has high accuracy, high sensitivity, and low false positive rate in the differentiation of benign and malignant prostate tumors, and can provide effective quantitative decision support for prostate tumor analysis.

[0036] It is worth noting that, in this embodiment, the input data for the radiomics feature scoring algorithm and the radiomics model includes the fusion of three types of data: T2-weighted imaging sequences, diffusion-weighted imaging sequences, and apparent diffusion coefficient sequences. (See [link to documentation]). Figure 4Verification showed that radiomics features based on T2WI sequences have a significant weight distribution in radiomics models, especially shape features extracted from T2WI, such as the long and short axis dimensions of nodules and sphericity, which contribute strongly to the analytical and predictive functions of radiomics models. However, the AUC values ​​of radiomics models based on single T2WI sequences are relatively low on both the training and test sets, indicating limited independent predictive performance. Similarly, radiomics features based on single ADC sequences effectively modeled the texture features and water molecule diffusion characteristics of prostate tumors. ADC sequences mainly reflect the water molecule diffusion capacity of the dominant lesion tissue in prostate tumors and have a strong correlation with cell density and malignancy of nodules. However, radiomics models based on single ADC sequences lack sensitivity to low-diffusion or high-density tissues. Finally, radiomics features based on single DWI sequences are highly sensitive to the microstructure and diffusion characteristics of tumor cells, and therefore can reveal the malignant characteristics of the prostate to some extent. However, radiomics models based on single DWI sequences still fail to fully integrate all important imaging information, resulting in predictive performance that does not reach the optimal level.

[0037] Therefore, in this embodiment, radiomics features from multiple sequences are fused, and the complementary biological information provided by different image sequences is used to achieve synergy in machine learning algorithms. This enables the radiomics model to capture the complex features of prostate lesions from multiple dimensions, thereby making more accurate analysis and judgments, improving the accuracy and stability of predicting the benign and malignant nature of prostate tumors, and achieving a deep characterization of the heterogeneity of prostate cancer. For example, prostate cancer exhibits significant intratumoral and intertumoral heterogeneity in its biological behavior. This heterogeneity may show different feature patterns on different image sequences. Some malignant regions may show reduced signal on T2WI, obvious diffusion restriction on ADC, and high signal on DWI; while other regions may only show abnormalities on some sequences. Therefore, radiomics models constructed from single sequences are prone to missing these complex feature combinations. However, multi-sequence fusion radiomics models and radiomics feature scoring algorithms can comprehensively capture these multidimensional information, thereby more accurately reflecting the true biological characteristics of tumors.

[0038] S24: Input the two-dimensional format region of interest of the original image data and the corresponding pathological results into the ResNet50 deep learning model for training, and establish the deep learning model; The deep learning features automatically extracted by the deep learning model were filtered using the LASSO-Cox regression model to obtain target deep learning features. The number of target deep learning features was limited to 33 groups. A deep learning feature scoring algorithm was calculated based on the target deep learning features. The deep learning feature scoring algorithm was used to predict the benign or malignant risk value of prostate tumors based on the target deep learning features of the original image data. Specifically, this embodiment uses the ResNet50 architecture as the foundation of the deep learning model and introduces the Transformer attention mechanism to enhance feature extraction capabilities. During the training of the ResNet50 deep learning model, the input original image data, specifically the two-dimensional ROI image, is preprocessed, resampled to a size of 224×224, and the pixel intensity is normalized. This processed image effectively improves the model's stability and generalization ability to the input data. During training, a Stochastic Gradient Descent (SGD) optimizer is used, which minimizes the loss function by updating network parameters. To prevent overfitting of the deep learning model, L2 regularization and Dropout techniques are combined. L2 regularization limits the model's complexity by penalizing large weight values, while Dropout randomly discards some neurons in the neural network to prevent the network from over-relying on the training data. Simultaneously, data augmentation techniques (including random rotation, flipping, and cropping) are also used to improve the generalization ability of the deep learning model, making it more adaptable to prostates of different shapes and locations.

[0039] To enhance the interpretability of deep learning models, see [link / reference]. Figure 5 This embodiment further employs the Gradient-weighted Class Activation Mapping (Grad-CAM) method. Grad-CAM is a gradient-based backpropagation technique that can display the regions that a neural network focuses on when making predictions in the form of heatmaps. This helps to understand how deep learning models make decisions, provides visualized prediction results from deep learning models, and reveals the basis for the deep learning models' predictions.

[0040] Then see Figure 6 The LASSO-Cox regression model is used to reduce the dimensionality of deep learning features automatically extracted by the deep learning model. LASSO-Cox regression can filter out deep learning features with significant contributions by adjusting the penalty coefficient, and effectively reduce the dimensionality of the feature space. Based on minimizing classification error and 10x cross-validation, the penalty coefficient λ = 0.0168 is obtained. Finally, 33 sets of target deep learning features are selected. A deep learning feature scoring algorithm is then calculated based on these target deep learning features. The formula for the deep learning feature scoring algorithm is as follows: label = 0.4433734939759036 + +0.139451 × DL_0 -0.009806 × DL_1+0.069015 × DL_2 -0.004876 × DL_3 +0.030689 × DL_4 +0.067325 ×DL_5 -0.091648 × DL_6 +0.047811 × DL_7 +0.003616 × DL_9 -0.014095× DL_14 +0.012989 × DL_15 -0.020529 × DL_18 -0.030356 × DL_21 +0.067679 × DL_23 +0.018770 × DL_24 -0.043652 × DL_25 -0.004792 ×DL_26 +0.005207 × DL_27 -0.010566 × DL_30 +0.000205 × DL_33 +0.007424 × DL_37 +0.010104 × DL_38 +0.021877 × DL_40 -0.007622 ×DL_41 +0.040228 × DL_43 -0.024672 × DL_45 -0.009612 × DL_46 +0.022945 × DL_49 +0.003336 × DL_51 -0.015396 × DL_55 +0.046103 ×DL_58 +0.049427 × DL_61 +0.005187 × DL_62 Similarly, the deep learning features involved in the deep learning feature scoring algorithm are automatically generated by the deep learning model. The meaning of the deep learning features will not be explained one by one in this embodiment.

[0041] S25: Input the target clinical features and corresponding pathological results into a logistic regression model, support vector machine model, K-nearest neighbor model, decision tree model, random forest model, XGBoost model, or LightGBM model, and train the model based on the 5-fold cross-validation method to establish a clinical prediction model.

[0042] Specifically, logistic regression, support vector machine, K-nearest neighbor, decision tree, random forest, XGBoost, or LightGBM models are used to train the selected PI-RADS scores, fPSA / tPSA ratios, BSA, and pathological results. In one embodiment, the XGBoost and LightGBM models exhibit relatively superior analytical and predictive performance, efficiently and accurately capturing the complex nonlinear relationships hidden in the data, thereby achieving high prediction accuracy for benign and malignant prostate tumors.

[0043] Preferably, in one embodiment, establishing a multimodal fusion model in S3 based on fusion features and combined with pathological results includes the following steps: S31: The radiomics model, deep learning model and clinical prediction model respectively output the benign and malignant risk value prediction results for the same prostate tumor multidimensional disease data. The three sets of benign and malignant risk value prediction results are probabilistically fused by weighted average or meta-classifier to generate a comprehensive prediction probability. The LASSO-Cox regression model was used to screen data features from different sources that affect the overall prediction probability, and target fusion features were obtained. The number of target fusion features was limited to 64 groups. Based on the target fusion features, an overall scoring algorithm was calculated. The overall scoring algorithm is used to predict the benign and malignant risk values ​​of prostate tumors by simultaneously combining target radiomics features, target deep learning features, and target clinical features.

[0044] Specifically, in this embodiment, independent prediction models based on deep learning, radiomics, and clinical indicators are trained separately to output their respective benign and malignant probability predictions for the same prostate tumor. Then, the probabilities are fused using a weighted average or a meta-classifier (such as logistic regression) to finally generate a comprehensive prediction probability. This achieves the fusion of the radiomics model, the deep learning model, and the clinical prediction model, generating a multimodal fusion model. The multimodal fusion model can retain the independent learning capabilities of each modality model, while optimizing the overall prediction performance through later integration. Furthermore, the comprehensive prediction probability generated by fusion can reduce the impact of dimensional differences and distribution inconsistencies between different modal features on the multimodal fusion model, thereby improving the stability and generalization ability of the multimodal fusion model.

[0045] For example, deep learning models can automatically extract high-level latent features from raw MRI images, capturing complex image patterns that are difficult for the human eye to recognize, but their decision-making process often lacks interpretability. Radiomics models, on the other hand, extract morphological, textural, and higher-order statistical features through quantitative analysis, providing interpretable image markers and helping users understand the biological basis of the multimodal fusion model's discriminative results. Clinical prediction models supplement individual biological background information, enhancing the relevance of the multimodal fusion model to clinical practice. The combination of these three approaches not only improves the discriminative power of the multimodal fusion model but also enhances its clinical applicability, maintaining high predictive accuracy while providing interpretable decision-making support. Furthermore, multimodal fusion models exhibit greater robustness, effectively mitigating the limitations of single data sources. For instance, MRI images may exhibit biases in radiomics feature extraction due to motion artifacts, magnetic field inhomogeneities, or partial volume effects, while deep learning models are highly dependent on image quality, and clinical features may be affected by detection errors or individual differences. However, within a multimodal framework, noise or deficiencies in one modality can be compensated for by information from other modalities. For example, when T2WI images suffer from unreliable radiomics features due to artifacts, deep learning models can still extract valid information from the original pixels, while clinical prediction models can further provide auxiliary analytical support. This redundancy design significantly improves the robustness of multimodal fusion models in real-world clinical scenarios and reduces the risk of misjudgments due to single data quality issues.

[0046] To further optimize feature selection and improve the generalization ability of the multimodal fusion model, this embodiment uses LASSO-Cox regression analysis to reduce the dimensionality of the fused multimodal features, ultimately obtaining a set of sparse but highly discriminative target fusion features. This provides a more reliable and efficient prediction tool for clinical decision-making. The number of target fusion features is 64. A comprehensive scoring algorithm is calculated based on the target fusion features. The formula for the comprehensive scoring algorithm is: label = 0.44827586206896547 -0.022093 × original_firstorder_Kurtosis_ADC -0.081031 × original_firstorder_RootMeanSquared_ADC +0.008828× original_firstorder_TotalEnergy_ADC -0.019541 × original_firstorder_Uniformity_ADC +0.007526 × original_glszm_SmallAreaEmphasis_ADC +0.046048× original_glszm_SmallAreaHighGrayLevelEmphasis_ADC +0.021138 ×original_ngtdm_Complexity_ADC -0.007220 × original_ngtdm_Contrast_ADC +0.027689 × original_shape_Elongation_ADC +0.004576 × original_shape_MinorAxisLength_ADC -0.076844 × original_shape_Sphericity_ADC +0.007765× original_shape_VoxelVolume_ADC -0.050885 × original_firstorder_Kurtosis_DWI -0.001822 × original_glcm_Idn_DWI -0.038504 × original_glcm_Imc1_DWI +0.038880 × original_glcm_InverseVariance_DWI +0.023328 ×original_gldm_DependenceEntropy_DWI -0.010339 × original_gldm_LargeDependenceEmphasis_DWI +0.020862 × original_gldm_LargeDependenceHighGrayLevelEmphasis_DWI -0.008214 × original_gldm_LargeDependenceLowGrayLevelEmphasis_DWI -0.003182 × original_glszm_GrayLevelNonUniformityNormalized_DWI +0.014772 × original_glszm_LargeAreaLowGrayLevelEmphasis_DWI +0.005267× original_glszm_SmallAreaLowGrayLevelEmphasis_DWI -0.019856 × original_ngtdm_Busyness_DWI +0.013573 × original_ngtdm_Complexity_DWI +0.016697 ×original_shape_Elongation_DWI -0.040633 × original_firstorder_Minimum_T2W+0.000083 × original_glcm_InverseVariance_T2W +0.037153 × original_glcm_MCC_T2W +0.007169 × original_glszm_ZoneEntropy_T2W -0.060288 ×original_glszm_ZonePercentage_T2W +0.018353 × original_shape_Elongation_T2W -0.052658 × original_shape_Sphericity_T2W +0.015526 × DL_0 -0.001529 × DL_1 +0.019195 × DL_2 -0.001613 × DL_3 +0.048423 × DL_5-0.060149 × DL_6 +0.022900 × DL_7 +0.001880 × DL_12 +0.022655 ×DL_13 +0.009182 × DL_15 -0.047546 × DL_21 +0.027961 × DL_23 +0.019430 × DL_24 -0.008803 × DL_25 +0.003901 × DL_28 -0.019702 ×DL_30 +0.005428 × DL_33 +0.010581 × DL_34 -0.000528 × DL_35 +0.008313 × DL_38 +0.001425 × DL_40 +0.013926 × DL_43 -0.013959 ×DL_44 -0.013527 × DL_46 +0.017540 × DL_49 +0.001515 × DL_51 -0.002157 × DL_55 +0.044874 × DL_58 +0.010533 × DL_59 +0.016623 ×DL_60 +0.019053 × DL_61. In the comprehensive scoring algorithm, by integrating features from different sources (radiomics features, deep learning features, and clinical features), the multimodal fusion model can fully explore the potential correlations between multidimensional symptoms of prostate tumors, thereby improving the ability to analyze and predict the benign and malignant nature of prostate tumors.

[0047] Preferably, in one embodiment, see Figures 7-8 By combining target radiomics features, target deep learning features, and target clinical features, a two-dimensional nomogram is constructed. This aims to provide users with a more intuitive tool to quickly predict the benign or malignant nature of prostate tumors based on the target individual's radiological and clinical data. The main features of the nomogram include the inputs from radiomics, deep learning, and clinical data, and their weighting in the final prediction. Validation showed that the nomogram achieved an AUC of 0.985 (95% confidence interval: 0.976–0.994) on the training set, far exceeding the predictive performance of traditional single radiomics models, deep learning models, and clinical models. On the test set, the AUC also reached 0.915 (95% confidence interval: 0.862–0.967), further validating the stability and efficiency of the nomogram on different datasets.

[0048] In summary, this embodiment provides a machine learning analysis method based on multidimensional prostate tumor symptom data. It deeply integrates multidimensional deep learning models, radiomics features, and clinical indicators to construct a multimodal fusion model and nomogram. This enables machine learning analysis of multidimensional prostate tumor symptom data and efficient automated prediction of benign and malignant prostate tumors. It solves the technical problems of traditional prostate biopsy, which is invasive and may cause complications such as infection and bleeding in human tissues, and cannot fully reflect the heterogeneity of tumors, as well as the low efficiency and poor accuracy of traditional manual analysis of medical imaging data or clinical data. In clinical practice, this provides a new approach for processing multidimensional prostate tumor symptom data and analyzing and predicting the benign and malignant nature of prostate tumors.

[0049] Second Embodiment Based on the same concept, the present invention also provides a machine learning analysis system based on multidimensional prostate tumor disease data, for performing the machine learning analysis method based on multidimensional prostate tumor disease data as described in any one of the first embodiments, including: The data acquisition module is used to acquire the original imaging data, original clinical data, and corresponding pathological results of the prostate-dominant lesion tissue of the enrolled individuals, and to perform preprocessing on the original imaging data. The feature extraction module is used to extract target radiomics features and target clinical features from raw image data and raw clinical data, respectively. Independent model building module, used to build radiomics models, deep learning models and clinical prediction models; The integrated model building module is used to build a multimodal fusion model based on radiomics models, deep learning models, and clinical prediction models. The nodal plot generation module is used to generate two-dimensional nodal plots based on the multimodal fusion model.

[0050] The functional implementation of each component in the above-mentioned machine learning analysis system based on multidimensional prostate tumor disease data corresponds to the steps in the above-mentioned machine learning analysis method embodiment based on multidimensional prostate tumor disease data. Their functions and implementation processes will not be described in detail here.

[0051] This embodiment also provides an electronic device, including a processor and a memory. The memory stores computer instructions executable by the processor, and the processor invokes the computer instructions stored in the memory to implement the aforementioned machine learning analysis method based on multidimensional prostate tumor disease data. This electronic device can be a server or a terminal device.

[0052] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings, but the present invention is not limited to the above embodiments. Even if various changes are made to the present invention, if these changes fall within the scope of the claims of the present invention and their equivalents, they shall still fall within the protection scope of the present invention.

Claims

1. A machine learning analysis method based on multidimensional prostate tumor disease data, characterized in that, Includes the following steps: S1: Obtain the original imaging data and original clinical data of the predominant prostate lesion tissue of the enrolled individuals, perform preprocessing on the original imaging data, and obtain the corresponding pathological results of the predominant prostate lesion tissue of the enrolled individuals in advance; S2: Filter the original image data to obtain target radiomics features, and establish a radiomics model based on the target radiomics features and pathological results; establish a deep learning model based on the original image data and pathological results; filter the original clinical data to obtain target clinical features, and establish a clinical prediction model based on the target clinical features and pathological results. S3: The prediction results of the radiomics model, the deep learning model and the clinical prediction model for multidimensional prostate tumor disease data are fused and analyzed to screen and obtain target fusion features, and a multimodal fusion model is established based on the target fusion features and combined with pathological results. S4: Based on the multimodal fusion model, a two-dimensional nomogram is established. The multimodal fusion model and the nomogram are used to perform analysis and prediction functions on the multidimensional symptom data of prostate tumors of target individuals from different approaches.

2. The machine learning analysis method based on multidimensional prostate tumor symptom data as described in claim 1, characterized in that, Locating the dominant prostate lesion tissue in the enrolled individuals in S1 includes the following steps: S11: Digital images of multiple prostate lesions in enrolled individuals were obtained by scanning with multi-parameter magnetic resonance imaging (MRI). The largest group of prostate lesions was selected and scored according to the prostate imaging report and data system, and marked as the dominant prostate lesion of the enrolled individual.

3. The machine learning analysis method based on multidimensional prostate tumor symptom data as described in claim 1, characterized in that, The raw image data categories include T2-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient images. Preprocessing of the raw image data is performed in S1, including the following steps: S12: Mark regions of interest (ROIs) in T2-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient images, including 3D and 2D formats. Extract geometric features of prostate tumors based on their 3D shapes from the ROIs in the T2-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient images, respectively. Extract intensity features of prostate tumors based on the voxel intensity distribution within the prostate tumors, respectively. Extract texture features of prostate tumors using gray-level co-occurrence matrix, gray-level run length matrix, gray-level dependency matrix, gray-level region size matrix, or neighborhood gray-level difference matrix, respectively.

4. The machine learning analysis method based on multidimensional prostate tumor symptom data as described in claim 1, characterized in that, In S2, the process of filtering the raw image data to obtain the target radiomics features and filtering the raw clinical data to obtain the target clinical features includes the following steps: S21: All radiomics features were screened using t-tests and / or Mann-Whitney U tests, and radiomics features with p-values ​​<0.05 were retained; The correlation between different radiomics features was calculated using the Spearman rank correlation coefficient. When there is a ρ value ≥ 0.9 between any two sets of radiomics features, one of the redundant radiomics features was deleted. The radiomics features were screened using the LASSO-Cox regression model to obtain the target radiomics features, which were limited to 26 groups. A radiomics feature scoring algorithm was calculated based on the target radiomics features. The radiomics feature scoring algorithm was used to predict the benign or malignant risk value of prostate tumors based on the geometric, intensity, and texture features of the dominant prostate lesion tissue. S22: All clinical features were screened using baseline statistical methods, and clinical features with p-values ​​<0.05 were retained; The target clinical features are obtained by screening clinical features using a logistic regression model.

5. The machine learning analysis method based on multidimensional prostate tumor symptom data as described in claim 4, characterized in that, In S22, clinical characteristics are screened using a logistic regression model, including the following steps: S221: In the Logistic regression model, the independent predictive power of any clinical feature is evaluated, the target clinical feature is selected, and the influence weight distribution of several target clinical features in the optimal predictive variable combination is calculated.

6. The machine learning analysis method based on multidimensional prostate tumor symptom data as described in claim 5, characterized in that, The target clinical features in the optimal combination of predictive variables include the Prostate Imaging Reporting and Data System (PIRD) score, the ratio of free prostate-specific antigen (PSA) to total PSA, and body surface area.

7. The machine learning analysis method based on multidimensional prostate tumor symptom data as described in claim 4, characterized in that, In S2, a radiomics model is established based on the target radiomics features and combined with pathological results; a deep learning model is established based on the original image data and combined with pathological results; and a clinical prediction model is established based on the target clinical features and combined with pathological results. This includes the following steps: S23: Input the target radiomics features and corresponding pathological results into a logistic regression model, support vector machine model, K-nearest neighbor model, decision tree model, random forest model, XGBoost model or LightGBM model, and train the model based on the 5-fold cross-validation method to establish the radiomics model; S24: Input the two-dimensional format region of interest of the original image data and the corresponding pathological results into the ResNet50 deep learning model for training, and establish the deep learning model; The deep learning features automatically extracted by the deep learning model are filtered using the LASSO-Cox regression model to obtain the target deep learning features. The number of target deep learning features is limited to 33 groups. A deep learning feature scoring algorithm is calculated based on the target deep learning features. The deep learning feature scoring algorithm is used to predict the benign or malignant risk value of prostate tumors based on the target deep learning features of the original image data. S25: Input the target clinical features and corresponding pathological results into a logistic regression model, support vector machine model, K-nearest neighbor model, decision tree model, random forest model, XGBoost model or LightGBM model, and train the model based on the 5-fold cross-validation method to establish the clinical prediction model.

8. The machine learning analysis method based on multidimensional prostate tumor symptom data as described in claim 7, characterized in that, In S3, a multimodal fusion model is established based on the fusion features and combined with pathological results, including the following steps: S31: The radiomics model, the deep learning model and the clinical prediction model respectively output the benign and malignant risk value prediction results for the same prostate tumor multidimensional disease data. The three sets of benign and malignant risk value prediction results are probabilistically fused by weighted average or meta-classifier to generate a comprehensive prediction probability. The LASSO-Cox regression model is used to filter data features from different sources that affect the overall prediction probability to obtain the target fusion features. The number of target fusion features is limited to 64 groups. A comprehensive scoring algorithm is calculated based on the target fusion features. The comprehensive scoring algorithm is used to simultaneously combine the target radiomics features, the target deep learning features, and the target clinical features to predict the benign and malignant risk values ​​of prostate tumors.

9. A machine learning analysis system based on multidimensional prostate tumor symptom data, characterized in that, A machine learning analysis method for performing multidimensional prostate tumor disease data as described in any one of claims 1-8, comprising: The data acquisition module is used to acquire the original imaging data, original clinical data, and corresponding pathological results of the prostate-dominant lesion tissue of the enrolled individuals, and to perform preprocessing on the original imaging data. The feature extraction module is used to extract target radiomics features and target clinical features from the original image data and the original clinical data, respectively. Independent model building module, used to build radiomics models, deep learning models and clinical prediction models; The integrated model building module is used to build a multimodal fusion model based on the radiomics model, the deep learning model, and the clinical prediction model. The nodal plot generation module is used to generate a two-dimensional nodal plot based on the multimodal fusion model.

10. An electronic device, characterized in that, include: Memory and processor; The memory is used to store computer instructions; The processor is configured to invoke computer instructions stored in the memory, causing the electronic device to execute the machine learning analysis method based on multidimensional prostate tumor disease data as described in any one of claims 1-8.