Preoperative prediction method for extracapsular extension of prostate cancer based on multi-modal expert fusion
Patent Information
- Application Number
- CN202610652075.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-13
- Publication Date
- 2026-08-28
AI Technical Summary
[0004]然而,在实际临床影像数据中,不同模态对包膜外侵犯预测所提供的信息并不相同,T2加权成像更有利于显示包膜轮廓及局部结构改变,弥散加权成像和表观扩散系数图更有利于反映肿瘤活性及扩散受限特征;同时,不同患者、不同扫描设备和不同扫描条件下,各模态影像还可能存在信噪比、伪影程度、灰度分布和空间一致性方面的差异
[0015] By acquiring multimodal magnetic resonance imaging (MRI) data of the prostate gland of the patient to be predicted, and performing spatial standardization and intensity normalization processing, standardized three-dimensional image volume data is obtained, enabling different modalities to first form a relatively consistent data representation basis. Based on this, convolution calculation and channel attention enhancement processing are performed on the standardized three-dimensional image volume data to extract the spatial distribution features corresponding to each modality, obtaining the original feature vectors corresponding to each modality. This allows the effective spatial features of each modality to be independently extracted and enhanced. Subsequently, linear projection and sequence modeling are performed on the original feature vectors corresponding to each modality, and intermodal information matching is executed. The system employs a homogeneous and attention-based interactive processing method to obtain an enhanced feature sequence, enabling the interaction of relevant information between different modalities under a unified feature representation. Furthermore, based on the original feature vectors corresponding to each modality and the enhanced feature sequence, multiple expert outputs are generated based on different modality combinations. These expert outputs are then dynamically weighted and fused to obtain an adaptive multimodal fusion feature representation, allowing for differentiated expression of the contributions of different modality combinations in the current patient sample. Finally, the multimodal fusion feature representation is subjected to classification mapping and probability normalization to obtain the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted. Therefore, this application does not directly and fixedly stitch together multimodal image data for prediction. Instead, it first reduces the data scale differences between modalities through standardization, then extracts the spatial distribution features of each modality through convolution calculation and channel attention enhancement, then establishes intermodal correlations through linear projection, sequence modeling, and attention interaction, and finally uses expert outputs corresponding to different modal combinations and dynamic weight fusion mechanisms to form an adaptive multimodal fusion feature representation. Therefore, the prediction process can simultaneously preserve the independent image information of each modality and the complementary relationship between modalities, and can adjust the fusion contribution according to the feature performance of different modal combinations, thereby reducing the adverse effects of single modality quality fluctuations or fixed fusion methods on the prediction results, and improving the stability and reliability of the preoperative prediction probability of extracapsular invasion of prostate cancer.
Smart Images

Figure CN122658643A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of classification and prediction technology, and in particular to a preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion. Background Technology
[0002] Prostate cancer is a common disease, and radical prostatectomy is an important treatment for patients with localized prostate cancer. During the surgical planning process, the preoperative assessment of extracapsular invasion directly affects the extent of resection and whether a nerve-sparing surgical strategy is adopted. If preoperative assessment indicates no extracapsular invasion, the surgeon can employ a nerve-sparing approach, provided the surgical criteria are met, to improve postoperative urinary continence and sexual function. If preoperative assessment indicates extracapsular invasion, a larger resection is usually necessary to reduce the risk of positive margins, biochemical recurrence, and local recurrence. Therefore, accurately predicting the extent of extracapsular invasion in prostate cancer preoperatively is a crucial technical aspect of prostate cancer surgical planning.
[0003] In existing technologies, multi-parameter magnetic resonance imaging (MRI) is the primary imaging basis for preoperative staging and assessment of extracapsular invasion in prostate cancer. Clinically, it is usually combined with multimodal imaging such as T2-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient maps for comprehensive judgment. Among them, T2-weighted imaging mainly reflects the anatomical structure and capsule morphology of the prostate, diffusion-weighted imaging mainly reflects the degree of water molecule diffusion restriction in the tissue, and the apparent diffusion coefficient map provides quantitative information on the degree of diffusion. With the development of computer-aided diagnostic technology, deep learning models have also been used to predict extracapsular invasion in prostate cancer. The common practice is to preprocess the multimodal MRI images and then input the images of each modality or the features extracted from the images into a convolutional neural network, a 3D convolutional network, or other classification models. The model then outputs the predicted probability of extracapsular invasion. To utilize multimodal information, existing models often use the method of stitching together images of different modalities as multiple input channels or directly stitching together features of different modalities and then feeding them into the classification layer to achieve fusion.
[0004] However, in actual clinical imaging data, different modalities provide varying information for predicting extracapsular invasion. T2-weighted imaging is more effective at displaying the capsule contour and local structural changes, while diffusion-weighted imaging and apparent diffusion coefficient maps are better at reflecting tumor activity and restricted spread characteristics. Furthermore, different patients, scanning devices, and scanning conditions may exhibit differences in signal-to-noise ratio, artifact levels, grayscale distribution, and spatial consistency across different modalities. Current technologies using fixed splicing or uniform input fusion methods often struggle to distinguish the differences in contribution of different modalities within specific samples, and also find it difficult to reduce the adverse effects of a poor-quality modality on the overall prediction results. This can easily lead to key signs of capsule invasion being interfered with by background noise or low-quality modalities, thus affecting the stability and reliability of the prediction results. Therefore, how to adaptively utilize multimodal imaging information based on the characteristic differences and modal quality differences of multimodal magnetic resonance imaging in different patients during preoperative prediction of extracapsular invasion in prostate cancer, in order to improve the stability and reliability of extracapsular invasion prediction results, has become a problem that needs to be solved. Summary of the Invention
[0005] This application provides a preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion, which can improve the stability and reliability of the prediction results. The technical solution provided in this application is as follows: In a first aspect, this application provides a preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion, the method comprising: Acquire prostate multimodal magnetic resonance imaging data of the patient to be predicted, and perform spatial standardization and intensity normalization processing on the prostate multimodal magnetic resonance imaging data to obtain standardized three-dimensional image volume data. Convolution calculation and channel attention enhancement processing are performed on the standardized three-dimensional image volume data to extract the spatial distribution features corresponding to each modality and obtain the original feature vectors corresponding to each modality. Linear projection and sequence modeling are performed on the original feature vectors corresponding to each modality, and inter-modal information alignment and attention interaction processing are performed to obtain the feature sequence with enhanced interaction. Based on the original feature vectors corresponding to each modality and the feature sequences after interactive enhancement, multiple expert outputs are formed based on different modality combinations. The multiple expert outputs are then dynamically weighted and fused to obtain an adaptive multimodal fusion feature representation. The multimodal fusion feature representation is classified, mapped, and normalized to obtain the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted.
[0006] In one specific implementation scheme, acquiring multimodal magnetic resonance imaging (MRI) data of the prostate of the patient to be predicted, and performing spatial standardization and intensity normalization processing on the multimodal MRI data of the prostate to obtain standardized three-dimensional image volume data, includes: T2-weighted imaging, apparent diffusion coefficient map and diffusion-weighted imaging of the patient to be predicted are obtained, and the three-dimensional image volume data corresponding to the T2-weighted imaging, the apparent diffusion coefficient map and the diffusion-weighted imaging are read respectively. Spatial resampling is performed on the three-dimensional image volume data corresponding to the T2-weighted imaging, the apparent diffusion coefficient map, and the diffusion-weighted imaging to unify the three-dimensional image volume data of the three modalities to a preset spatial resolution size. After spatial resampling, if the number of voxels in any dimension of the 3D image volume data of any modality is greater than the preset spatial resolution size, then the corresponding dimension is cropped; if the number of voxels in any dimension of the 3D image volume data of any modality is less than the preset spatial resolution size, then the corresponding dimension is filled to make the 3D image volume data of the three modalities consistent in spatial size. After spatial standardization, the volumetric data of the 3D image corresponding to each modality is independently subjected to intensity normalization. The voxel intensity values of each modality are scaled to a uniform range or converted into a standard normal distribution form to obtain standardized 3D image volumetric data.
[0007] In one specific implementation scheme, the step of performing convolution calculation and channel attention enhancement processing on the standardized 3D image volume data to extract the spatial distribution features corresponding to each modality and obtain the original feature vector corresponding to each modality includes: The standardized 3D image volume data is divided into T2-weighted imaging 3D image volume and apparent diffusion coefficient according to modality. Figure 3 The volume of the 3D image and the volume of the diffusion-weighted imaging 3D image are respectively input into the 3D convolutional encoder corresponding to each modality. The network structure of each 3D convolutional encoder is the same and the parameters are independent of each other. The three-dimensional image volume of the corresponding modality is calculated by two-stage convolution through each three-dimensional convolution encoder. Each three-dimensional convolution encoder includes two consecutive three-dimensional convolution blocks. Each three-dimensional convolution block includes a three-dimensional convolution layer, a batch normalization layer and a ReLU activation function. The two three-dimensional convolution stages use 32 output channels and 64 output channels respectively. After each 3D convolutional block, channel attention enhancement and 3D max pooling are performed on the corresponding feature map to form a feature map containing information about the corresponding modal space distribution. Global average pooling is performed on the feature maps corresponding to each modality to aggregate the feature maps with three-dimensional spatial dimensions into one-dimensional feature vectors. The one-dimensional feature vectors are then input into a fully connected layer, which projects the features of each modality onto a 128-dimensional feature space to obtain the original feature vectors of T2-weighted imaging, the original feature vectors of the apparent diffusion coefficient map, and the original feature vectors of diffusion-weighted imaging.
[0008] In one specific implementation, the channel attention enhancement process includes: For any modality of convolutional feature map, given the input channel attention processing feature tensor is: ,in, This represents the feature tensor used for input channel attention processing. Indicates batch size. Indicates the number of channels. , , These represent the spatial dimensions of the feature tensor in the depth, height, and width directions, respectively. Global average pooling and global max pooling are performed on the feature tensors of the input channels for attention processing, respectively, to compress the feature responses of each channel in the spatial dimension, resulting in two tensors of shape. Channel descriptor; The two channel descriptors are respectively input into a multilayer perceptron with shared parameters for transformation to obtain a channel attention map. The channel attention map is calculated as follows: ; in, Represents the feature tensor based on input channel attention processing The calculated channel attention map, This represents the Sigmoid activation function. Represents the characteristic tensor Channel descriptors obtained by global average pooling. Represents the characteristic tensor The channel descriptor obtained by performing global max pooling. Let represent the first weight matrix in a multilayer perceptron with shared parameters, and , Let represent the second weight matrix in a multilayer perceptron with shared parameters, and , Indicates the channel reduction ratio; The channel attention map is expanded to a spatial dimension that matches the feature tensor of the input channel attention processing through a broadcast mechanism, and then multiplied element-wise with the feature tensor of the input channel attention processing to obtain the channel-enhanced feature tensor.
[0009] In a specific implementation, the step of performing linear projection and sequence modeling on the original feature vectors corresponding to each modality, and performing inter-modal information alignment and attention interaction processing to obtain the interaction-enhanced feature sequence includes: For each modality, a modality-specific linear projection is performed on the original feature vector to obtain the projection vector corresponding to each modality. The linear projection process is expressed as follows: ; in, Indicates image modality, , This indicates T2-weighted imaging. This represents the apparent diffusion coefficient diagram. This indicates diffusion-weighted imaging. Representing modes The corresponding original feature vector, Representing modes The corresponding linear projection weight matrix, and , Representing modes The corresponding linear projection bias vector, and , Representing modes The corresponding projection vector; The projection vectors of the three modalities are stacked sequentially according to the order of T2-weighted imaging, apparent diffusion coefficient map, and diffusion-weighted imaging to form a modal feature sequence of length 3, represented as: ; in, This represents a modal feature sequence consisting of three modal projection vectors. This represents the projection vector corresponding to T2-weighted imaging. This represents the projection vector corresponding to the apparent diffusion coefficient map. This represents the projection vector corresponding to diffusion-weighted imaging; The modal feature sequence is input into a single-layer Transformer encoder, and intermodal attention interaction is performed through a multi-head self-attention mechanism to obtain an interaction-enhanced feature sequence.
[0010] In a specific feasible implementation, the step of forming multiple expert outputs based on different modal combinations according to the original feature vectors corresponding to each modality and the interactively enhanced feature sequences, and dynamically assigning weights and fusing the multiple expert outputs to obtain an adaptive multimodal fusion feature representation includes: Based on the enhanced feature sequence, the corresponding modal feature representations are extracted in the modal order of T2-weighted imaging, apparent diffusion coefficient map, and diffusion-weighted imaging. Combined with the original feature vectors corresponding to each modality, the modal feature inputs in the hybrid modality expert fusion processing are formed. The modal feature inputs include T2-weighted imaging feature vectors, apparent diffusion coefficient map feature vectors, and diffusion-weighted imaging feature vectors. Multiple expert networks corresponding to different modal subset combinations are constructed. The expert network includes 3 single-modal experts, 3 bimodal experts, and 1 full-modal expert. The single-modal experts process T2-weighted imaging features, apparent diffusion coefficient map features, and diffusion-weighted imaging features, respectively. The bimodal experts process the combined features of T2-weighted imaging and apparent diffusion coefficient map, the combined features of T2-weighted imaging and diffusion-weighted imaging, and the combined features of apparent diffusion coefficient map and diffusion-weighted imaging, respectively. The full-modal expert processes the complete trimodal combined features of T2-weighted imaging, apparent diffusion coefficient map, and diffusion-weighted imaging. The concatenation result of the modal feature vectors in the modal subset corresponding to each expert output is input into the corresponding expert network to obtain each expert output. Each expert network adopts a two-layer lightweight multilayer perceptron, and the two linear layers are connected by the ReLU activation function. The T2-weighted imaging feature vector, the apparent diffusion coefficient map feature vector, and the diffusion-weighted imaging feature vector are concatenated in a preset modal order to obtain a joint multimodal context vector. The joint multimodal context vector is then input into the routing network, and the routing weights corresponding to each expert output are obtained through linear projection and Softmax normalization. The outputs of each expert are weighted and summed according to their corresponding routing weights to obtain an adaptive multimodal fusion feature representation.
[0011] In one specific implementation, the step of performing classification mapping and probability normalization on the multimodal fusion feature representation to obtain the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted includes: The adaptive multimodal fusion feature representation is input into a lightweight fully connected classifier, which includes two linear layers connected by a ReLU activation function and a Dropout layer. The lightweight fully connected classifier is used to classify and map the adaptive multimodal fusion feature representation to obtain classification output values related to extracapsular invasion of prostate cancer. The classification output value is normalized using the Sigmoid activation function to obtain the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted. The predicted probability of extracapsular invasion of prostate cancer is calculated as follows: ; in, This indicates the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted; This represents the Sigmoid activation function; This represents the classification weight parameters in a lightweight fully connected classifier; This represents an adaptive multimodal fusion feature representation; This represents the classification bias parameter in a lightweight fully connected classifier. Based on the predicted probability of extracapsular invasion of prostate cancer and the preset judgment threshold, the preoperative prediction results of the patients to be predicted are formed.
[0012] Secondly, this application provides a preoperative prediction system for extracapsular invasion of prostate cancer based on multimodal expert fusion, employing the following technical solution: A preoperative prediction system for extracapsular invasion of prostate cancer based on multimodal expert fusion includes: The image preprocessing module is used to acquire prostate multimodal magnetic resonance imaging data of the patient to be predicted, and to perform spatial standardization and intensity normalization processing on the prostate multimodal magnetic resonance imaging data to obtain standardized three-dimensional image volume data. The feature extraction module is used to perform convolution calculation and channel attention enhancement processing on the standardized three-dimensional image volume data, extract the spatial distribution features corresponding to each modality, and obtain the original feature vector corresponding to each modality. The feature interaction module is used to perform linear projection and sequence modeling on the original feature vectors corresponding to each modality, perform intermodal information alignment and attention interaction processing, and obtain the feature sequence after interaction enhancement. The expert fusion module is used to generate multiple expert outputs based on the original feature vectors corresponding to each modality and the feature sequences after interactive enhancement, and to dynamically assign weights and perform weighted fusion on the multiple expert outputs to obtain an adaptive multimodal fusion feature representation. The probability prediction module is used to perform classification mapping and probability normalization processing on the multimodal fusion feature representation to obtain the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted.
[0013] Thirdly, this application provides an electronic device, the device including a processor and a memory; the memory stores a program, the program being loaded and executed by the processor to implement a preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion as described in the first aspect.
[0014] Fourthly, this application provides a computer-readable storage medium storing a program that, when executed by a processor, is used to implement a preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion as described in the first aspect.
[0015] By acquiring multimodal magnetic resonance imaging (MRI) data of the prostate gland of the patient to be predicted, and performing spatial standardization and intensity normalization processing, standardized three-dimensional image volume data is obtained, enabling different modalities to first form a relatively consistent data representation basis. Based on this, convolution calculation and channel attention enhancement processing are performed on the standardized three-dimensional image volume data to extract the spatial distribution features corresponding to each modality, obtaining the original feature vectors corresponding to each modality. This allows the effective spatial features of each modality to be independently extracted and enhanced. Subsequently, linear projection and sequence modeling are performed on the original feature vectors corresponding to each modality, and intermodal information matching is executed. The system employs a homogeneous and attention-based interactive processing method to obtain an enhanced feature sequence, enabling the interaction of relevant information between different modalities under a unified feature representation. Furthermore, based on the original feature vectors corresponding to each modality and the enhanced feature sequence, multiple expert outputs are generated based on different modality combinations. These expert outputs are then dynamically weighted and fused to obtain an adaptive multimodal fusion feature representation, allowing for differentiated expression of the contributions of different modality combinations in the current patient sample. Finally, the multimodal fusion feature representation is subjected to classification mapping and probability normalization to obtain the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted. Therefore, this application does not directly and fixedly stitch together multimodal image data for prediction. Instead, it first reduces the data scale differences between modalities through standardization, then extracts the spatial distribution features of each modality through convolution calculation and channel attention enhancement, then establishes intermodal correlations through linear projection, sequence modeling, and attention interaction, and finally uses expert outputs corresponding to different modal combinations and dynamic weight fusion mechanisms to form an adaptive multimodal fusion feature representation. Therefore, the prediction process can simultaneously preserve the independent image information of each modality and the complementary relationship between modalities, and can adjust the fusion contribution according to the feature performance of different modal combinations, thereby reducing the adverse effects of single modality quality fluctuations or fixed fusion methods on the prediction results, and improving the stability and reliability of the preoperative prediction probability of extracapsular invasion of prostate cancer.
[0016] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, the preferred embodiments of this application are described in detail below with reference to the accompanying drawings. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion in the embodiments of this application.
[0018] Figure 2 This is a schematic diagram of the preoperative prediction model for extracapsular invasion of prostate cancer in the embodiments of this application.
[0019] Figure 3 This is a structural block diagram of the preoperative prediction system for extracapsular invasion of prostate cancer based on multimodal expert fusion in the embodiments of this application.
[0020] Figure 4 This is a block diagram of an electronic device for preoperative prediction of extracapsular invasion of prostate cancer based on multimodal expert fusion, as described in this application. Detailed Implementation
[0021] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate this application, but are not intended to limit the scope of this application.
[0022] Optionally, this application uses the preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion provided in various embodiments as an example for application in an electronic device. The electronic device is a terminal or server. The terminal can be a computer, tablet computer, etc. This embodiment does not limit the type of electronic device.
[0023] Reference Figure 1 This is a flowchart illustrating a preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion, provided in one embodiment of this application. The method includes at least the following steps: Step S101: Obtain prostate multimodal magnetic resonance imaging data of the patient to be predicted, and perform spatial standardization and intensity normalization processing on the prostate multimodal magnetic resonance imaging data to obtain standardized three-dimensional image volume data.
[0024] In step S101, multimodal magnetic resonance imaging (MRI) data of the prostate of the patient to be predicted is acquired, and spatial standardization and intensity normalization are performed on the prostate multimodal MRI data to obtain standardized three-dimensional image volume data. This step is used to organize the image data obtained from different MRI sequences of the patient to be predicted into three-dimensional volume data with uniform spatial size and uniform intensity scale. Since different modal images may differ in voxel spacing, spatial size, and grayscale distribution, without uniform processing, the data representation of different modalities of the same patient will be inconsistent, which can easily affect the comparability and consistency of the multimodal image data itself. Therefore, spatial standardization and intensity normalization processing of the prostate multimodal MRI data are necessary.
[0025] Specifically, the multimodal magnetic resonance imaging data of the prostate includes T2-weighted imaging (T2WI), apparent diffusion coefficient mapping (ADC), and diffusion-weighted imaging (DWI). T2-weighted imaging reflects the anatomical structure and capsule morphology of the prostate; diffusion-weighted imaging reflects the degree of water molecule diffusion restriction in the tissue and identifies tumor-active areas; and the apparent diffusion coefficient mapping provides a quantitative diffusion coefficient and helps differentiate between benign and malignant tissues. For the acquired T2-weighted imaging, apparent diffusion coefficient mapping, and diffusion-weighted imaging data, the corresponding three-dimensional image volume data are read, and the three-dimensional image volume data of the three modalities are used as objects for spatial standardization processing.
[0026] Spatial standardization processing includes spatial resampling and cropping / padding. First, spatial resampling is performed on the 3D image volume data corresponding to T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging to unify the volume data of the three modalities to a preset spatial resolution size. This preset spatial resolution size is pre-set according to the data processing specifications, for example, it can be set to 32×128×128 voxels. After spatial resampling, if the number of voxels in a certain dimension of the 3D image volume data of any modality is greater than the preset spatial resolution size, then that dimension is cropped to match the preset spatial resolution size; if the number of voxels in a certain dimension of the 3D image volume data of any modality is less than the preset spatial resolution size, then that dimension is padded to match the preset spatial resolution size. Through the above processing, the 3D image volume data corresponding to T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging maintain consistency in spatial size.
[0027] After spatial standardization, intensity normalization was performed independently on the 3D image volume data corresponding to each modality. Since T2-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient maps have different imaging physical meanings, and the voxel intensity distributions of the three modalities are also different, intensity normalization was performed separately for each modality. Specifically, the voxel intensity values of each modality were scaled to a uniform range or converted to a standard normal distribution to reduce grayscale distribution differences caused by different devices and scanning parameters. After intensity normalization, the 3D image volume data of each modality formed a relatively uniform intensity scale expression.
[0028] Through the above spatial standardization and intensity normalization processes, standardized three-dimensional image volume data are obtained. This standardized three-dimensional image volume data includes the three-dimensional image volume data corresponding to T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging. The three-dimensional image volume data corresponding to the three modalities are consistent in spatial dimensions, and each undergoes intensity scale unification processing, thereby forming a standardized data representation of the multimodal magnetic resonance imaging of the prostate of the patient to be predicted.
[0029] Step S102: Perform convolution calculation and channel attention enhancement processing on the standardized 3D image volume data, extract the spatial distribution features corresponding to each modality, and obtain the original feature vector corresponding to each modality.
[0030] In step S102, convolution calculations and channel attention enhancement are performed on the standardized 3D image volume data to extract the spatial distribution features corresponding to each modality, obtaining the original feature vectors for each modality. This step is used to extract feature information that can characterize the image characteristics of each modality from the 3D image volumes corresponding to T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging, respectively. Since the three modalities reflect different medical meanings—T2-weighted imaging focuses more on the anatomical structure and capsule morphology of the prostate, apparent diffusion coefficient maps focus more on the quantitative expression of diffusion degree, and diffusion-weighted imaging focuses more on diffusion restriction and tumor activity region expression—directly mixing the three modalities can easily weaken the image features of each modality. Therefore, this step performs convolution feature extraction on each modality separately, so that each modality first forms an independent feature expression.
[0031] Specifically, the standardized three-dimensional image volume data obtained in step S101 is divided into T2-weighted imaging three-dimensional image volume and apparent diffusion coefficient according to modality. Figure 3 The three-dimensional image volume and diffusion-weighted imaging volume are calculated and input into the corresponding three-dimensional convolutional encoders for each modality. The three three-dimensional convolutional encoders have the same network structure but independent parameters, and are used to process T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging, respectively. By setting independent parameters, the three-dimensional convolutional encoders for each modality can learn the spatial distribution patterns in each modality, avoiding mutual interference between grayscale representations, structural representations, and diffusion representations of different modalities during the feature extraction stage.
[0032] Each 3D convolutional encoder comprises two consecutive 3D convolutional blocks. Each 3D convolutional block includes a 3D convolutional layer, a batch normalization layer, and a ReLU activation function. The 3D convolutional layer extracts local spatial features in the depth, height, and width directions of the 3D image volume; the batch normalization layer normalizes the feature responses obtained from the 3D convolutional layer to stabilize the feature distribution; and the ReLU activation function performs a nonlinear transformation on the normalized feature responses, enabling the 3D convolutional encoder to express nonlinear image structural differences. The two 3D convolutional stages use 32 and 64 output channels respectively, gradually transforming the encoding process from lower-level local textures, edges, and grayscale changes to higher-level spatial distribution features.
[0033] Following each 3D convolutional block, channel attention enhancement and 3D max pooling are applied. Channel attention enhancement selectively enhances or suppresses channel features based on the information content of different channels in the current feature map, resulting in higher responses for feature channels related to prostate structure, capsule morphology, diffusion-restricted regions, and quantitative diffusion differences, while reducing the influence of noisy or weakly correlated channels on feature representation. 3D max pooling spatially downsamples the feature map after channel attention enhancement, reducing its spatial size while preserving key response features. Thus, the 3D image volume corresponding to each modality, after two stages of convolution, channel attention enhancement, and 3D max pooling, forms a feature map containing the spatial distribution information of that modality.
[0034] After obtaining the feature maps corresponding to each modality, global average pooling is performed on the feature maps of each modality, aggregating the feature maps with three-dimensional spatial dimensions into one-dimensional feature vectors. Global average pooling aggregates the responses of each channel in the spatial dimension, forming a unified response value for each channel, thus compressing the three-dimensional spatial features into a fixed-length vector representation. Subsequently, the one-dimensional feature vectors obtained from global average pooling are input into a fully connected layer. The fully connected layer projects the features of each modality onto a 128-dimensional feature space, yielding the original feature vectors for T2-weighted imaging, the original feature vectors for the apparent diffusion coefficient map, and the original feature vectors for diffusion-weighted imaging.
[0035] The encoding process corresponding to the three modes is represented as follows: ; ; ; in, This represents the normalized 3D image volume corresponding to T2-weighted imaging; This represents the normalized 3D image volume corresponding to the apparent diffusion coefficient map. This represents the normalized 3D image volume corresponding to diffusion-weighted imaging. This represents a 3D convolutional encoder used for processing T2-weighted imaging; This represents a three-dimensional convolutional encoder used to process the apparent diffusion coefficient map; This represents a three-dimensional convolutional encoder used for processing diffusion-weighted imaging; This represents the original feature vector of T2-weighted imaging; This represents the original eigenvector of the apparent diffusion coefficient map; This represents the original feature vector of diffusion-weighted imaging; It represents a 128-dimensional real vector space.
[0036] Through the above processing, T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging are converted into original feature vectors of equal length. Each original feature vector originates from the 3D spatial image information of the corresponding modality and retains the image representation characteristics of that modality, thereby transforming the standardized 3D image volume data into feature representations with uniform structure and modal distinction.
[0037] Step S103: Perform linear projection and sequence modeling on the original feature vectors corresponding to each modality, and perform inter-modal information alignment and attention interaction processing to obtain the feature sequence after interaction enhancement.
[0038] In step S103, linear projection and sequence modeling are performed on the original feature vectors corresponding to each modality, and intermodal information alignment and attention interaction processing are executed to obtain the interactively enhanced feature sequence. This step is used to ensure that the features of each modality have strong channel discrimination ability after feature extraction has been completed for each modality, and then map the features corresponding to T2-weighted imaging, apparent diffusion coefficient map and diffusion-weighted imaging into a unified feature space for intermodal interaction. Since different modalities have different focuses on characterizing extracapsular invasion of prostate cancer, a single modality feature can only reflect part of the imaging manifestation. Therefore, it is necessary to establish the interdependence between different modalities while retaining the features of each modality itself, so that the features corresponding to the three modalities can form a unified expression after interactive enhancement.
[0039] Specifically, before obtaining the original feature vectors corresponding to each modality, the convolutional feature maps corresponding to each modality undergo channel attention enhancement processing. For any convolutional feature map of any modality, given the input feature tensor: ;in, The feature tensor representing the input channel attention processing; Indicates batch size; Indicates the number of channels; , , These represent the spatial dimensions of the feature tensor in the depth, height, and width directions, respectively. Represents the real tensor space of the corresponding dimension.
[0040] For input feature tensor Global average pooling and global max pooling are performed separately to compress the feature response of each channel in the spatial dimension, resulting in two channel descriptors. The channel descriptor obtained by global average pooling reflects the average response level of each channel in the overall spatial range, while the channel descriptor obtained by global max pooling reflects the maximum response level of each channel in the locally salient region. Both channel descriptors have a shape of [missing information]. ,in, Indicates batch size as The number of channels is The space of real matrix numbers.
[0041] After obtaining two channel descriptors, each descriptor is input into a multilayer perceptron with shared parameters for transformation. This multilayer perceptron employs a bottleneck structure, achieved by reducing the ratio. First, reduce the channel dimension and then restore it to reduce the number of parameters and model the dependencies between channels. The channel attention map is calculated as follows: ; in, Represents the input feature tensor The calculated channel attention map; This represents the Sigmoid activation function; Represents the input feature tensor Channel descriptors obtained by global average pooling; Represents the input feature tensor Channel descriptors obtained by performing global max pooling; This represents the first weight matrix in a shared multilayer perceptron. ; This represents the second weight matrix in a shared multilayer perceptron. ; Indicates the channel reduction ratio; express OK The space of real matrices in the column; express OK The space of real matrix columns.
[0042] After obtaining the channel attention map Then, the channel attention map is extended to the input feature tensor via a broadcast mechanism. Matching spatial dimensions and with the input feature tensor Element-wise multiplication yields the enhanced feature tensors for each channel. This process allows the average pooling path to reflect the overall response of a channel, while the max pooling path reflects its significant response. Both types of channel descriptions participate in the attention weight calculation, enhancing channels with high information content and suppressing noisy or weakly correlated channels. Thus, the convolutional feature maps corresponding to each modality undergo adaptive feature selection at the channel level before forming the original feature vectors.
[0043] After convolution, channel attention enhancement, pooling aggregation, and fully connected projection for each modality, the original feature vectors for T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging are obtained. For each modality, the original feature vector is a 128-dimensional vector. To enable interaction between features from different modalities within a unified space, a modality-specific linear projection is performed on the original feature vector for each modality. The linear projection process is expressed as follows: ; in, Indicates image modality, ; This indicates T2-weighted imaging; This represents a diagram showing the apparent diffusion coefficient. This indicates diffusion-weighted imaging; Representing modes The corresponding original feature vector; Representing modes The corresponding linear projection weight matrix, ; Representing modes The corresponding linear projection bias vector, ; Representing modes The corresponding projection vector; Represents a real matrix space with 128 rows and 128 columns; It represents a 128-dimensional real vector space.
[0044] Through the linear projection process described above, the original feature vectors of each modality are mapped to a shared latent semantic space. Since each modality uses an independent projection weight matrix and bias vector, the feature differences between T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging can still be preserved; at the same time, all projection vectors have the same dimension, enabling correlation modeling between different modalities under a unified dimension.
[0045] After obtaining the projection vectors of the three modes, the projection vectors are stacked sequentially in the order of T2-weighted imaging, apparent diffusion coefficient map, and diffusion-weighted imaging to form a modal feature sequence of length 3, represented as: ; in, This represents a sequence of modal features consisting of three modal projection vectors; This represents the projection vector corresponding to T2-weighted imaging; This represents the projection vector corresponding to the apparent diffusion coefficient map; This represents the projection vector corresponding to diffusion-weighted imaging; This represents a 3x128 matrix space of real numbers. Through sequence modeling, the three modes are represented as three feature units in the same sequence, which facilitates the calculation of the correlation between different feature units along the modality dimension.
[0046] After forming the modal feature sequence, the modal feature sequence is input into a single-layer Transformer encoder, and inter-modal attention interaction is performed through a multi-head self-attention mechanism. In one specific implementation, the number of attention heads is 4, and the dimension of each attention head is 32. For each attention head, the query matrix, key matrix, and value matrix are calculated respectively: ; in, Represents the query matrix; Represents the key matrix; Represents a value matrix; Represents a modal feature sequence; This represents the projected weight matrix corresponding to the query matrix; This represents the projected weight matrix corresponding to the key matrix; This represents the projected weight matrix corresponding to the value matrix; , , All belong to , This represents a real matrix space with 128 rows and 32 columns.
[0047] For each attention head, attention weights are calculated based on the correlation between the query matrix and the key matrix. These attention weights are then used to weight the value matrix to obtain the output corresponding to that attention head. The attention calculation process is represented as follows: ; in, Indicates attention head output; This represents the product of the query matrix and the transpose of the key matrix, used to indicate the degree of correlation between modal feature units; The dimension representing each attention head; Indicates the scaling factor; This represents the normalized exponential function, used to convert correlation results into attention weights; Represents a value matrix.
[0048] After each attention head completes its attention calculation, the outputs of all attention heads are concatenated to form a multi-head self-attention output. Subsequently, the multi-head self-attention output is residually connected to the modal feature sequence, and the residual connection result is layer-normalized to obtain the interactively enhanced feature sequence, as shown below: ;in, This represents the feature sequence after interaction enhancement; Presentation layer normalization processing; Represents a modal feature sequence; Represents the modal feature sequence The output obtained after performing multi-head self-attention computation. Residual connections are used to preserve the original information in the projection vectors of each modality, and layer normalization is used to stabilize the feature distribution after interaction.
[0049] Through the above processing, channel attention enhancement first filters and enhances channel features within each modality, giving the original feature vectors of each modality a more explicit discriminative expression. Then, linear projection and sequence modeling map the three modalities to a shared latent semantic space and organize them in sequence. Finally, a multi-head self-attention mechanism performs information interaction based on inter-modal correlations, resulting in an interactively enhanced feature sequence. This interactively enhanced feature sequence retains the feature information of T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging, while also containing the interdependencies among the three modalities.
[0050] Step S104: Based on the original feature vectors corresponding to each modality and the feature sequences after interactive enhancement, multiple expert outputs are formed based on different modality combinations. Dynamic weight allocation and weighted fusion are performed on the multiple expert outputs to obtain an adaptive multimodal fusion feature representation.
[0051] In step S104, based on the original feature vectors corresponding to each modality and the enhanced feature sequences, multiple expert outputs are formed based on different modality combinations. Dynamic weight allocation and weighted fusion are then performed on these multiple expert outputs to obtain an adaptive multimodal fusion feature representation. This step is used to further fuse different modality combinations based on their differences in contribution to extracapsular invasion prediction, building upon the independent feature extraction and intermodal attention interaction of the three modalities. Since the image quality and diagnostic relevance of T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging may differ in different patients, using a fixed single-modality combination or simple splicing is insufficient to adapt to variations in modality contributions across different samples. Therefore, this step uses single-modality, bimodal, and full-modality combinations as different expert inputs and dynamically determines the contribution weight of each expert output based on the joint multimodal context.
[0052] Specifically, based on the enhanced feature sequence obtained in step S103, the corresponding modal feature representations are extracted according to the modal order of T2-weighted imaging, apparent diffusion coefficient map, and diffusion-weighted imaging. These are then combined with the original feature vectors of each modality obtained in step S102 to form the modal feature input in the hybrid modal expert fusion processing. This modal feature input includes the T2-weighted imaging feature vector, the apparent diffusion coefficient map feature vector, and the diffusion-weighted imaging feature vector, denoted as follows: , and .in, This represents the modal feature vector corresponding to T2-weighted imaging. This represents the modal eigenvector corresponding to the apparent diffusion coefficient map. This represents the modal feature vector corresponding to diffusion-weighted imaging. The modal feature vectors mentioned above are consistent with the feature sequence after interactive enhancement in step S103 in terms of modal order, enabling different experts to read the corresponding feature inputs according to the determined modal combination relationship.
[0053] After generating modal feature inputs, multiple expert networks corresponding to different modal subset combinations are constructed. For the three modalities of T2-weighted imaging, apparent diffusion coefficient map, and diffusion-weighted imaging, the expert networks are divided into three categories: the first category consists of three single-modal experts, each handling T2-weighted imaging features, apparent diffusion coefficient map features, and diffusion-weighted imaging features respectively; the second category consists of three bimodal experts, each handling combined features of T2-weighted imaging and apparent diffusion coefficient map, combined features of T2-weighted imaging and diffusion-weighted imaging, and combined features of apparent diffusion coefficient map and diffusion-weighted imaging respectively; the third category consists of one full-modal expert, used to handle the complete trimodal combination features of T2-weighted imaging, apparent diffusion coefficient map, and diffusion-weighted imaging. Therefore, the total number of experts is [number missing]. .
[0054] Each expert network employs a two-layer lightweight multilayer perceptron. The input is the concatenation of feature vectors from each modality subset corresponding to the expert, and the output is the fixed-dimensional expert output. The output of each expert is represented as follows: ; in, Indicates the first Experts from an expert network stated; Indicates the first A two-layer lightweight multilayer perceptron corresponding to an expert network; Indicates feature concatenation operation; , to Indicates the first Feature vectors of each modality in the modality subset corresponding to each expert; , to Indicates the first A modality selected by an expert; Indicates the first Number of modalities corresponding to each expert .when At that time, expert networks process single-modal features; when At that time, the expert network processes the splicing features of the two modalities; when At that time, the expert network processes the spliced features of the three modalities. Each multilayer perceptron consists of two linear layers connected by a ReLU activation function to perform a nonlinear mapping of the feature relationships in the corresponding modal combination.
[0055] After obtaining the outputs from each expert, the routing weights of each expert's output are further calculated based on the joint multimodal context. Specifically, the T2-weighted imaging feature vector, the apparent diffusion coefficient map feature vector, and the diffusion-weighted imaging feature vector are concatenated according to a preset modal order to obtain the joint multimodal context vector: ; in, This represents a joint multimodal context vector formed by concatenating feature vectors from three modalities. , and All are 128-dimensional feature vectors; It represents a 384-dimensional real vector space.
[0056] The joint multimodal context vector is input into the routing network, and after linear projection and Softmax normalization, the routing weights corresponding to the expert outputs are obtained, expressed as follows: ; in, Represents the route weight vector; This represents the linear projection weight matrix in the routing network, with the matrix dimensions matching the joint multimodal context vector and the total number of experts. Represents the bias vector in the routing network; This represents the normalization exponential function, used to convert the linear projection results into weights corresponding to each expert. This represents the total number of experts, and ; , express A 3D real vector space. Routing weight vector. The first in Each component Indicates the first The dynamic contribution weight of each expert output in the current patient sample to be predicted.
[0057] After calculating the outputs of each expert and their corresponding routing weights, a weighted fusion of all expert outputs is performed to obtain an adaptive multimodal fusion feature representation. The weighted fusion process is represented as follows: ; in, This represents an adaptive multimodal fusion feature representation; Indicates the total number of experts; Indicates the first Each expert outputs the corresponding routing weight; Indicates the first Experts from an expert network stated; Indicates the first Each expert outputs a weighted result based on its corresponding routing weight.
[0058] Through the above processing, single-modal experts can retain the independent diagnostic information of T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging, respectively; bimodal experts can express the combined relationship between the two modalities; and multimodal experts can express the complete trimodal joint features. The routing network dynamically assigns weights to the outputs of each expert based on the joint multimodal context, enabling different patient samples to correspond to different modal combination contribution relationships. When the image quality of a certain modality is low or its correlation with the judgment of extracapsular invasion in the current sample is weak, the routing weight of the expert related to that modality can be reduced accordingly; when a certain modal combination can provide stronger discriminative information, the corresponding expert output can receive a higher weight. The resulting adaptive multimodal fusion feature representation can simultaneously contain single-modal information, bimodal complementary information, and multimodal joint information, and can adjust the contribution degree of each type of information according to sample differences.
[0059] Step S105: Perform classification mapping and probability normalization on the multimodal fusion feature representation to obtain the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted.
[0060] In step S105, the multimodal fusion feature representation is subjected to classification mapping and probability normalization to obtain the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted. This step is used to convert the adaptive multimodal fusion feature representation obtained in step S104 into a probabilistic result that can reflect the possibility of extracapsular invasion. Since the multimodal fusion feature representation already contains the fusion information of T2-weighted imaging, apparent diffusion coefficient map and diffusion-weighted imaging under different modal combinations, but this feature representation itself is still a vector expression in the feature space, it cannot be directly used as a clinically understandable prediction result. Therefore, it needs to be converted into a predicted probability of extracapsular invasion of prostate cancer through classification mapping.
[0061] Specifically, the adaptive multimodal fusion feature representation is input into a lightweight fully connected classifier. This lightweight fully connected classifier consists of two linear layers connected by a ReLU activation function and a Dropout layer. The linear layers are used for classification mapping of the multimodal fusion feature representation, the ReLU activation function introduces non-linear expressiveness, and the Dropout layer reduces the instability caused by feature dependency settling during the classification mapping process; the dropout rate of the Dropout layer can be set to 0.5. After processing by the lightweight fully connected classifier, classification output values related to extracapsular invasion of prostate cancer are obtained.
[0062] After obtaining the classification output value, the Sigmoid activation function is used to normalize the probability of the classification output value, mapping it to a probability range between 0 and 1, thus obtaining the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted. This predicted probability represents the likelihood that the patient to be predicted has extracapsular invasion of prostate cancer. The calculation method for the predicted probability of extracapsular invasion of prostate cancer is as follows: ; in, This indicates the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted; This represents the Sigmoid activation function; This represents the classification weight parameters in a lightweight fully connected classifier; This represents the adaptive multimodal fusion feature representation obtained in step S104; This represents the classification bias parameter in a lightweight fully connected classifier. This represents the classification output value obtained after classifying and mapping the adaptive multimodal fusion feature representation.
[0063] After obtaining the predicted probability of extracapsular invasion of prostate cancer, a corresponding preoperative prediction result can be generated based on a preset judgment threshold. The preset judgment threshold can be set to 0.5. At that time, it was determined that the patient to be predicted had extracapsular invasion of prostate cancer; when At that time, it is determined that the patient to be predicted does not have extracapsular invasion of prostate cancer. The preset judgment threshold can also be adjusted according to clinical needs. For example, when more attention is paid to the identification of positive cases, the preset judgment threshold can be lowered to improve the sensitivity of identifying patients with positive extracapsular invasion.
[0064] Through the above processing, the adaptive multimodal fusion feature representation is converted into a predicted probability of extracapsular invasion of prostate cancer, so that the fusion features in multimodal magnetic resonance imaging are expressed in probabilistic form as a preoperative prediction result. This predicted probability can intuitively reflect the likelihood of extracapsular invasion of prostate cancer in the patient to be predicted, and provide a quantitative basis for preoperative assessment of the extracapsular invasion status.
[0065] Furthermore, preferably, in one embodiment, the aforementioned steps S101 to S105 can be implemented using a pre-constructed preoperative prediction model for extracapsular invasion of prostate cancer. (Refer to...) Figure 2 This preoperative prediction model for extracapsular invasion of prostate cancer uses standardized 3D image volume data as input. The standardized 3D image volume data includes the 3D image volumes corresponding to T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging. The preoperative prediction model for extracapsular invasion of prostate cancer includes 3D convolutional encoders corresponding to the three modalities, a cross-modal attention interaction structure, a hybrid modal expert fusion structure, and a classification prediction structure.
[0066] The system comprises three 3D convolutional encoders, each receiving the 3D image volumes corresponding to T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging, respectively, and outputting the original feature vectors for T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging. A cross-modal attention interaction structure is connected to the outputs of the three 3D convolutional encoders, used to perform linear projection, sequence modeling, and attention interaction on the original feature vectors of the three modalities, resulting in an enhanced feature sequence. A hybrid modal expert fusion structure receives the original feature vectors of each modality and the enhanced feature sequence, and generates multiple expert outputs based on different combinations of single-modal, dual-modal, and full-modal approaches. The routing weights corresponding to each expert output are then calculated based on the joint multimodal context, and the multiple expert outputs are weighted and fused to obtain an adaptive multimodal fusion feature representation. A classification prediction structure is connected to the output of the hybrid modal expert fusion structure, used to perform classification mapping and probability normalization on the adaptive multimodal fusion feature representation, obtaining the predicted probability of extracapsular invasion of prostate cancer. Therefore, the preoperative prediction model for extracapsular invasion of prostate cancer forms a continuous computational structure from trimodal image input, modal feature extraction, intermodal interaction, expert adaptive fusion to prediction probability output.
[0067] When training a preoperative prediction model for extracapsular invasion of prostate cancer, binary focus loss was used as a supervisory signal. Because the number of positive and negative samples of extracapsular invasion of prostate cancer in clinical data may be unbalanced, directly using ordinary classification loss for training would cause the model to tend to learn more features of the class with a larger number of samples, thus reducing its ability to identify minority class samples. Binary focus loss, by reducing the influence of easily classified samples on the loss, makes the training process focus more on difficult and minority class samples, thereby improving the model's ability to learn positive samples of extracapsular invasion of prostate cancer.
[0068] Specifically, given the classification value output by the preoperative prediction model for extracapsular invasion of prostate cancer... The predicted probability is obtained through the Sigmoid function. , represented as: ;in, Indicates the predicted probability; This represents the classification value output by the preoperative prediction model for extracapsular invasion of prostate cancer. This represents the Sigmoid function.
[0069] To reduce the risk of overconfidence in model predictions, label smoothing is performed on the true labels before calculating the loss, as shown below: ; in, This indicates the smoothed label; Indicates a real label, and , This indicates a positive sample for extracapsular invasion of prostate cancer. This indicates a negative sample indicating extracapsular invasion of prostate cancer. This represents the smoothing factor, used to control the degree of label smoothing.
[0070] The binary focal loss is expressed as: ; in, Indicates the loss at the binary focus; This represents the category balance factor, used to adjust the relative weights of positive and negative samples in the loss calculation. This represents the focusing parameter, used to reduce the contribution of easily classified samples to the loss; This represents the probability term in the predicted probability that corresponds to the true class.
[0071] in, Represented as: ; in, This represents the predicted probability obtained through the Sigmoid function; This represents the real label. When the real label... hour, Take the probability of a positive prediction. When the real label hour, Take the probability of a negative prediction. In one embodiment, the category balance factor It can be set to 0.6, the focus parameter. It can be set to 1.5, the smoothing factor. It can be set to 0.05.
[0072] During training, the Adam optimizer can be used to update the trainable parameters in the preoperative prediction model for extracapsular invasion of prostate cancer. The initial learning rate can be set to... Weight decay can be set to The batch size can be set to 4. The learning rate scheduling can employ a cosine annealing strategy combined with linear preheating to ensure smoother parameter updates in the early stages of training and facilitate gradual convergence of the training process. The dropout rate in the classification prediction structure can be set to 0.5 to reduce excessive reliance on local features during classification mapping. The training framework can be implemented based on PyTorch and the MONAI medical imaging framework.
[0073] During model validation, five-fold hierarchical cross-validation can be used to evaluate the preoperative prediction model for extracapsular invasion of prostate cancer. Five-fold hierarchical cross-validation divides the samples into five subsets, maintaining a consistent ratio of positive to negative samples for extracapsular invasion within each subset. Each time, one subset is selected as the validation set, and the remaining subsets are used as the training set. This process is repeated five times to obtain the overall evaluation result. Evaluation metrics include the area under the ROC curve, accuracy, sensitivity, and specificity. Specifically, the area under the ROC curve evaluates the model's overall ability to distinguish between positive and negative samples; accuracy evaluates the proportion of predictions consistent with the true labels; sensitivity evaluates the proportion of correctly identified positive samples; and specificity evaluates the proportion of correctly identified negative samples.
[0074] In summary, we first acquire T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging of the patient to be predicted. Then, we perform spatial standardization and intensity normalization on the multimodal MRI data of the prostate to obtain standardized 3D image volume data. Next, we perform convolution calculations and channel attention enhancement on the standardized 3D image volume data to extract the spatial distribution features corresponding to each modality, obtaining the original feature vectors for each modality. Then, we perform linear projection and sequence modeling on the original feature vectors corresponding to each modality, and perform intermodal information alignment and attention interaction processing to obtain an interactively enhanced feature sequence. Following this, based on the original feature vectors corresponding to each modality and the interactively enhanced feature sequence, we generate multiple expert outputs based on different modal combinations (single-modal, bimodal, and full-modal). We then dynamically assign weights and perform weighted fusion on the multiple expert outputs through routing weights to obtain an adaptive multimodal fusion feature representation. Finally, we perform classification mapping and probability normalization on the multimodal fusion feature representation to obtain the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted.
[0075] Analysis based on the above technical solutions reveals that this application first uses spatial standardization and intensity normalization to standardize the spatial size and intensity scale of T2-weighted imaging, apparent diffusion coefficient maps, and diffusion-weighted imaging, thereby reducing data expression differences caused by different scanning conditions. Furthermore, it extracts modal features through parameter-independent convolution calculations and introduces channel attention enhancement within each modality to enhance channels with high information content and suppress noisy or weakly correlated channels. On this basis, through linear projection, sequence modeling, and attention interaction processing, the three modalities establish interdependencies within a shared latent semantic space, forming an interactively enhanced feature sequence containing intermodal correlation information. Subsequently, expert outputs are generated based on single-modality, bimodality, and full-modality combinations, and the routing weights of each expert output are dynamically determined according to the joint multimodal context. This allows the prediction process to adaptively increase the weights of effective modal combinations based on the quality and contribution differences of images from different patient samples, while reducing the impact of low-quality or weakly correlated modal combinations on the fusion results. Therefore, this application can make fuller use of the complementary information between different modalities of images, reduce the interference of background noise and low-quality modalities on the expression of key capsule invasion signs, thereby improving the stability and reliability of preoperative prediction results of extracapsular invasion of prostate cancer.
[0076] Figure 3 This is a structural block diagram of a preoperative prediction system for extracapsular invasion of prostate cancer based on multimodal expert fusion, provided in one embodiment of this application. The system includes at least the following modules: The image preprocessing module is used to acquire prostate multimodal magnetic resonance imaging data of the patient to be predicted, and to perform spatial standardization and intensity normalization processing on the prostate multimodal magnetic resonance imaging data to obtain standardized three-dimensional image volume data. The feature extraction module is used to perform convolution calculations and channel attention enhancement on standardized 3D image volume data to extract the spatial distribution features corresponding to each modality and obtain the original feature vectors corresponding to each modality. The feature interaction module is used to perform linear projection and sequence modeling on the original feature vectors corresponding to each modality, perform intermodal information alignment and attention interaction processing, and obtain the feature sequence after interaction enhancement. The expert fusion module is used to generate multiple expert outputs based on the original feature vectors corresponding to each modality and the feature sequences after interactive enhancement, and to dynamically assign weights and perform weighted fusion on the multiple expert outputs to obtain an adaptive multimodal fusion feature representation. The probability prediction module is used to classify, map, and normalize the multimodal fusion feature representation to obtain the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted.
[0077] For relevant details, please refer to the above method implementation examples.
[0078] Figure 4 This is a block diagram of an electronic device provided in one embodiment of this application. The device includes at least a processor 401 and a memory 402.
[0079] The processor 401 executes computer program instructions stored in the memory 402 to implement the preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion provided in this embodiment. The processor 401 may include one or more processing cores and may be implemented in at least one hardware form selected from CPU, GPU, DSP, FPGA, or AI processor. The GPU or AI processor may be used to perform operations such as three-dimensional image volumetric data processing, convolution calculation, attention calculation, expert fusion calculation, and classification prediction calculation.
[0080] The memory 402 may include one or more computer-readable storage media, which may be non-transitory storage media. The memory 402 may include high-speed random access memory or non-volatile memory, such as a disk storage device or a flash memory device. The memory 402 stores at least one computer program instruction, which, when executed by the processor 401, causes the processor 401 to implement the preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion provided in this embodiment of the application.
[0081] Optionally, this application also provides a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion in the above-described method embodiments.
[0082] Optionally, this application also provides a computer product including a computer-readable storage medium storing a program, which is loaded and executed by a processor to implement the preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion in the above-described method embodiments.
[0083] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion, characterized in that, The method includes: Acquire prostate multimodal magnetic resonance imaging data of the patient to be predicted, and perform spatial standardization and intensity normalization processing on the prostate multimodal magnetic resonance imaging data to obtain standardized three-dimensional image volume data. Convolution calculation and channel attention enhancement processing are performed on the standardized three-dimensional image volume data to extract the spatial distribution features corresponding to each modality and obtain the original feature vectors corresponding to each modality. Linear projection and sequence modeling are performed on the original feature vectors corresponding to each modality, and inter-modal information alignment and attention interaction processing are performed to obtain the feature sequence with enhanced interaction. Based on the original feature vectors corresponding to each modality and the feature sequences after interactive enhancement, multiple expert outputs are formed based on different modality combinations. The multiple expert outputs are then dynamically weighted and fused to obtain an adaptive multimodal fusion feature representation. The multimodal fusion feature representation is classified, mapped, and normalized to obtain the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted.
2. The method for preoperative prediction of extracapsular invasion of prostate cancer based on multimodal expert fusion according to claim 1, characterized in that, The process involves acquiring multimodal magnetic resonance imaging (MRI) data of the prostate gland of the patient to be predicted, and performing spatial standardization and intensity normalization on the MRI data to obtain standardized three-dimensional image volume data, including: T2-weighted imaging, apparent diffusion coefficient map and diffusion-weighted imaging of the patient to be predicted are obtained, and the three-dimensional image volume data corresponding to the T2-weighted imaging, the apparent diffusion coefficient map and the diffusion-weighted imaging are read respectively. Spatial resampling is performed on the three-dimensional image volume data corresponding to the T2-weighted imaging, the apparent diffusion coefficient map, and the diffusion-weighted imaging to unify the three-dimensional image volume data of the three modalities to a preset spatial resolution size. After spatial resampling, if the number of voxels in any dimension of the 3D image volume data of any modality is greater than the preset spatial resolution size, then the corresponding dimension is cropped; if the number of voxels in any dimension of the 3D image volume data of any modality is less than the preset spatial resolution size, then the corresponding dimension is filled to make the 3D image volume data of the three modalities consistent in spatial size. After spatial standardization, the volumetric data of the 3D image corresponding to each modality is independently subjected to intensity normalization. The voxel intensity values of each modality are scaled to a uniform range or converted into a standard normal distribution form to obtain standardized 3D image volumetric data.
3. The method for preoperative prediction of extracapsular invasion of prostate cancer based on multimodal expert fusion according to claim 1, characterized in that, The process involves performing convolution calculations and channel attention enhancement on the standardized 3D image volume data to extract the spatial distribution features corresponding to each modality, resulting in the original feature vectors for each modality. This includes: The standardized three-dimensional image volume data is divided into three-dimensional image volume with T2 weighted imaging, three-dimensional image volume with apparent diffusion coefficient map, and three-dimensional image volume with diffusion weighted imaging according to modality, and then input into the three-dimensional convolutional encoder corresponding to each modality. The network structure of each three-dimensional convolutional encoder is the same and the parameters are independent of each other. The three-dimensional image volume of the corresponding modality is calculated by two-stage convolution through each three-dimensional convolution encoder. Each three-dimensional convolution encoder includes two consecutive three-dimensional convolution blocks. Each three-dimensional convolution block includes a three-dimensional convolution layer, a batch normalization layer and a ReLU activation function. The two three-dimensional convolution stages use 32 output channels and 64 output channels respectively. After each 3D convolutional block, channel attention enhancement and 3D max pooling are performed on the corresponding feature map to form a feature map containing information about the corresponding modal space distribution. Global average pooling is performed on the feature maps corresponding to each modality to aggregate the feature maps with three-dimensional spatial dimensions into one-dimensional feature vectors. The one-dimensional feature vectors are then input into a fully connected layer, which projects the features of each modality onto a 128-dimensional feature space to obtain the original feature vectors of T2-weighted imaging, the original feature vectors of the apparent diffusion coefficient map, and the original feature vectors of diffusion-weighted imaging.
4. The method for preoperative prediction of extracapsular invasion of prostate cancer based on multimodal expert fusion according to claim 3, characterized in that, The channel attention enhancement processing includes: For any modality of convolutional feature map, given the input channel attention processing feature tensor is: ,in, This represents the feature tensor used for input channel attention processing. Indicates batch size, Indicates the number of channels. , , These represent the spatial dimensions of the feature tensor in the depth, height, and width directions, respectively. Global average pooling and global max pooling are performed on the feature tensors of the input channels for attention processing, respectively, to compress the feature responses of each channel in the spatial dimension, resulting in two tensors of shape. Channel descriptor; The two channel descriptors are respectively input into a multilayer perceptron with shared parameters for transformation to obtain a channel attention map. The channel attention map is calculated as follows: ; in, Represents the feature tensor based on input channel attention processing The calculated channel attention map, This represents the Sigmoid activation function. Represents the characteristic tensor Channel descriptors obtained by global average pooling. Represents the characteristic tensor The channel descriptor obtained by performing global max pooling. Let represent the first weight matrix in a multilayer perceptron with shared parameters, and , Let represent the second weight matrix in a multilayer perceptron with shared parameters, and , Indicates the channel reduction ratio; The channel attention map is expanded to a spatial dimension that matches the feature tensor of the input channel attention processing through a broadcast mechanism, and then multiplied element-wise with the feature tensor of the input channel attention processing to obtain the channel-enhanced feature tensor.
5. The method for preoperative prediction of extracapsular invasion of prostate cancer based on multimodal expert fusion according to claim 4, characterized in that, The process involves linear projection and sequence modeling of the original feature vectors corresponding to each modality, followed by inter-modal information alignment and attention interaction processing to obtain an enhanced feature sequence, including: For each modality, a modality-specific linear projection is performed on the original feature vector to obtain the projection vector corresponding to each modality. The linear projection process is expressed as follows: ; in, Indicates image modality, , This indicates T2-weighted imaging. This represents the apparent diffusion coefficient diagram. This indicates diffusion-weighted imaging. Representing modes The corresponding original feature vector, Representing modes The corresponding linear projection weight matrix, and , Representing modes The corresponding linear projection bias vector, and , Representing modes The corresponding projection vector; The projection vectors of the three modalities are stacked sequentially according to the order of T2-weighted imaging, apparent diffusion coefficient map, and diffusion-weighted imaging to form a modal feature sequence of length 3, represented as: ; in, This represents a modal feature sequence consisting of three modal projection vectors. This represents the projection vector corresponding to T2-weighted imaging. This represents the projection vector corresponding to the apparent diffusion coefficient map. This represents the projection vector corresponding to diffusion-weighted imaging; The modal feature sequence is input into a single-layer Transformer encoder, and intermodal attention interaction is performed through a multi-head self-attention mechanism to obtain an interaction-enhanced feature sequence.
6. The method for preoperative prediction of extracapsular invasion of prostate cancer based on multimodal expert fusion according to claim 1, characterized in that, The process involves generating multiple expert outputs based on the original feature vectors corresponding to each modality and the interactively enhanced feature sequences, combining different modalities, and dynamically assigning and weighting these expert outputs to obtain an adaptive multimodal fusion feature representation. This includes: Based on the enhanced feature sequence, the corresponding modal feature representations are extracted in the modal order of T2-weighted imaging, apparent diffusion coefficient map, and diffusion-weighted imaging. Combined with the original feature vectors corresponding to each modality, the modal feature inputs in the hybrid modality expert fusion processing are formed. The modal feature inputs include T2-weighted imaging feature vectors, apparent diffusion coefficient map feature vectors, and diffusion-weighted imaging feature vectors. Multiple expert networks corresponding to different modal subset combinations are constructed. The expert network includes 3 single-modal experts, 3 bimodal experts, and 1 full-modal expert. The single-modal experts process T2-weighted imaging features, apparent diffusion coefficient map features, and diffusion-weighted imaging features, respectively. The bimodal experts process the combined features of T2-weighted imaging and apparent diffusion coefficient map, the combined features of T2-weighted imaging and diffusion-weighted imaging, and the combined features of apparent diffusion coefficient map and diffusion-weighted imaging, respectively. The full-modal expert processes the complete trimodal combined features of T2-weighted imaging, apparent diffusion coefficient map, and diffusion-weighted imaging. The concatenation result of the modal feature vectors in the modal subset corresponding to each expert output is input into the corresponding expert network to obtain each expert output. Each expert network adopts a two-layer lightweight multilayer perceptron, and the two linear layers are connected by the ReLU activation function. The T2-weighted imaging feature vector, the apparent diffusion coefficient map feature vector, and the diffusion-weighted imaging feature vector are concatenated in a preset modal order to obtain a joint multimodal context vector. The joint multimodal context vector is then input into the routing network, and the routing weights corresponding to each expert output are obtained through linear projection and Softmax normalization. The outputs of each expert are weighted and summed according to their corresponding routing weights to obtain an adaptive multimodal fusion feature representation.
7. The method for preoperative prediction of extracapsular invasion of prostate cancer based on multimodal expert fusion according to claim 1, characterized in that, The process of classifying, mapping, and normalizing the multimodal fusion feature representation to obtain the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted includes: The adaptive multimodal fusion feature representation is input into a lightweight fully connected classifier, which includes two linear layers connected by a ReLU activation function and a Dropout layer. The lightweight fully connected classifier is used to classify and map the adaptive multimodal fusion feature representation to obtain classification output values related to extracapsular invasion of prostate cancer. The classification output value is normalized using the Sigmoid activation function to obtain the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted. The predicted probability of extracapsular invasion of prostate cancer is calculated as follows: ; in, This indicates the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted; This represents the Sigmoid activation function; This represents the classification weight parameters in a lightweight fully connected classifier; This represents an adaptive multimodal fusion feature representation; This represents the classification bias parameter in a lightweight fully connected classifier. Based on the predicted probability of extracapsular invasion of prostate cancer and the preset judgment threshold, the preoperative prediction results of the patients to be predicted are formed.
8. A preoperative prediction system for extracapsular invasion of prostate cancer based on multimodal expert fusion, characterized in that, include: The image preprocessing module is used to acquire prostate multimodal magnetic resonance imaging data of the patient to be predicted, and to perform spatial standardization and intensity normalization processing on the prostate multimodal magnetic resonance imaging data to obtain standardized three-dimensional image volume data. The feature extraction module is used to perform convolution calculation and channel attention enhancement processing on the standardized three-dimensional image volume data, extract the spatial distribution features corresponding to each modality, and obtain the original feature vector corresponding to each modality. The feature interaction module is used to perform linear projection and sequence modeling on the original feature vectors corresponding to each modality, perform intermodal information alignment and attention interaction processing, and obtain the feature sequence after interaction enhancement. The expert fusion module is used to generate multiple expert outputs based on the original feature vectors corresponding to each modality and the feature sequences after interactive enhancement, and to dynamically assign weights and perform weighted fusion on the multiple expert outputs to obtain an adaptive multimodal fusion feature representation. The probability prediction module is used to perform classification mapping and probability normalization processing on the multimodal fusion feature representation to obtain the predicted probability of extracapsular invasion of prostate cancer in the patient to be predicted.
9. An electronic device, characterized in that, The device includes a processor and a memory; the memory stores a program that is loaded and executed by the processor to implement a preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a program that, when executed by a processor, is used to implement a preoperative prediction method for extracapsular invasion of prostate cancer based on multimodal expert fusion as described in any one of claims 1 to 7.