Prostate image classification method and device, computer equipment and storage medium
Through the CNN-Transformer network and adaptive fusion module of mixed-perspective learning, combined with cross-sectional and sagittal perspective information, the low contrast problem of prostate cancer diagnosis in existing technologies is solved, achieving more accurate prostate cancer identification and reducing surgical risks.
Patent Information
- Application Number
- CN202510791877.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-12
AI Technical Summary
In the diagnosis of prostate cancer, existing transrectal ultrasound imaging methods have a low contrast problem, making it difficult to accurately identify clinically significant prostate cancer. In addition, existing methods ignore the complementary information of cross-sectional and sagittal perspectives, resulting in unnecessary biopsies and increased surgical risks.
A CNN-Transformer network with mixed-view learning is used to extract features from cross-sectional and sagittal images, and feature enhancement is performed in combination with cross-view attention. A mixed-view adaptive fusion module is used to dynamically aggregate features in channel and spatial dimensions to construct a mixed-view classification network.
It improves the accuracy of prostate cancer identification, reduces unnecessary biopsies, lowers surgical risks, and provides a better reference for medical diagnosis.
Smart Images

Figure CN120635576A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a prostate image classification method, device, computer equipment and storage medium. Background Art
[0002] Prostate cancer is one of the most common cancers and a leading cause of cancer-related mortality in men. Prostate biopsy is the gold standard for diagnosing prostate cancer. Based on the pathological characteristics of prostate tissue, prostate cancer can be diagnosed as clinically significant prostate cancer (csPCa), clinically insignificant prostate cancer (cisPCa), and benign prostatic hyperplasia (BPH). csPCa carries a poor prognosis and requires prompt identification and intervention to ensure timely and effective treatment. Current guidelines recommend multiparametric magnetic resonance imaging as the preferred tool for optimizing prostate biopsy. However, its limited availability and complex procedure limit its widespread use. In contrast, transrectal ultrasound (TRUS) is widely used for biopsy guidance due to its low cost, ease of use, and real-time imaging capabilities. However, the low contrast of TRUS complicates the accurate identification and diagnosis of PCa, potentially leading to more unnecessary biopsies, increased surgical risks, and increased patient suffering and burden. Furthermore, during transrectal ultrasound examinations, physicians often scan from both the transverse and sagittal perspectives, leveraging the complementary information from these two perspectives to further confirm suspicious lesions. Therefore, developing a computer-assisted diagnosis method for csPCa based on multi-view TRUS is of great clinical significance.
[0003] In this regard, the prior art provides several solutions, specifically:
[0004] (1) Automated multiparametric localization of prostate cancer based on B-mode, shear-wave elastography, and contrast-enhanced ultrasound radiomics. This work uses radiomic features from B-mode ultrasound, shear-wave elastography (SWE), and contrast-enhanced ultrasound (DCE-US) to construct a random forest classifier for csPCa detection. The problems with this approach are: the feature extraction method relies on annotating the lesion area in the biopsy pathology image, and the multimodal features need to be manually designed and extracted. In addition, it only focuses on cross-sectional imaging, ignoring the complementary information of the sagittal plane in TRUS scanning.
[0005] (2) Three-dimensional convolutional neural network model to identify clinically significant prostate cancer in transrectal ultrasound videos: a prospective, multi-institutional, diagnostic study. This work uses a prostate mask-guided 3D convolutional neural network to classify csPCa. First, the segmentation network is trained to extract the prostate region to reduce the interference of other surrounding tissues; then, the classification network is trained based on the segmentation results to identify csPCa. Problems with this work: This method relies on high-quality prostate annotation for segmentation network training, but the annotation process of prostate scan videos is time-consuming and labor-intensive, and segmentation errors will affect subsequent classification performance, limiting its practical application. In addition, it only focuses on cross-sectional imaging and ignores the complementary information of the sagittal plane in TRUS scans.
[0006] (3) Towards Multi-modality Fusion and Prototype-based Feature Refinement for Clinically Significant Prostate Cancer Classification in Transrectal Ultrasound: This work proposes a multimodal fusion network that combines B-ultrasound and SWE for csPCa classification and enhances the classification encoder capability through a few-shot segmentation task. Problems with this work include: prototype learning uses masked average pooling to extract feature prototypes, which can easily lead to the loss of small target lesion features during downsampling; and it focuses only on cross-sectional imaging, ignoring the complementary information of the sagittal plane in TRUS scans. Summary of the Invention
[0007] The embodiments of the present invention provide a prostate image classification method, apparatus, computer equipment, and storage medium, aiming to improve the processing effect of medical images.
[0008] In a first aspect, an embodiment of the present invention provides a prostate image classification method, comprising:
[0009] Acquiring a prostate image, wherein the prostate image includes a transverse view image and a sagittal view image;
[0010] A CNN-Transformer network with mixed-view learning is used to extract features from the cross-sectional view image and the sagittal view image respectively, and feature enhancement is performed in combination with cross-view attention to obtain cross-sectional view features and sagittal view features;
[0011] Dynamically fusing the cross-sectional view features and the sagittal view features using a hybrid view adaptive fusion module to obtain a fusion feature;
[0012] Classifying the fused features to obtain a classification detection result of the prostate image, thereby constructing a hybrid view classification network;
[0013] The hybrid view classification network is used to classify the designated prostate image.
[0014] In a second aspect, an embodiment of the present invention provides a prostate image classification device, comprising:
[0015] An image acquisition unit, configured to acquire a prostate image, wherein the prostate image includes a cross-sectional view image and a sagittal view image;
[0016] a feature extraction unit, configured to extract features from the cross-sectional view image and the sagittal view image respectively using a CNN-Transformer network learned through mixed viewpoints, and perform feature enhancement in combination with cross-viewpoint attention, thereby obtaining cross-sectional view features and sagittal view features;
[0017] a feature fusion unit, configured to dynamically fuse the cross-sectional view feature and the sagittal view feature using a hybrid view adaptive fusion module to obtain a fused feature;
[0018] a feature classification unit, configured to classify the fused features to obtain a classification detection result of the prostate image, thereby constructing a hybrid view classification network;
[0019] A classification processing unit is used to perform classification processing on a designated prostate image using the hybrid view classification network.
[0020] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the prostate image classification method as described in the first aspect is implemented.
[0021] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the prostate image classification method as described in the first aspect is implemented.
[0022] An embodiment of the present invention provides a prostate image classification method, apparatus, computer equipment, and storage medium. The method includes: acquiring a prostate image, wherein the prostate image includes a cross-sectional view image and a sagittal view image; using a CNN-Transformer network for hybrid perspective learning to extract features from the cross-sectional view image and the sagittal view image respectively, and combining cross-perspective attention to perform feature enhancement, thereby obtaining cross-sectional view features and sagittal view features; using a hybrid perspective adaptive fusion module to dynamically fuse the cross-sectional view features and sagittal view features to obtain fused features; classifying the fused features to obtain a classification detection result of the prostate image, thereby constructing a hybrid perspective classification network; and using the hybrid perspective classification network to classify a specified prostate image. The embodiment of the present invention adopts a CNN-Transformer hybrid architecture and combines cross-view attention to extract fine-grained local features and model global dependencies for cross-sectional and sagittal view images, respectively, to extract cross-sectional and sagittal view features. Then, the hybrid view adaptive fusion module dynamically aggregates features in the channel and spatial dimensions, thereby optimizing the overall representation capability and improving the processing effect of medical images. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0024] Figure 1 A schematic flow chart of a prostate image classification method provided by an embodiment of the present invention;
[0025] Figure 2 A model architecture diagram of a prostate image classification method provided by an embodiment of the present invention;
[0026] Figure 3 Another model architecture diagram of a prostate image classification method provided by an embodiment of the present invention;
[0027] Figure 4 An example image of a prostate image classification method provided by an embodiment of the present invention;
[0028] Figure 5 A schematic block diagram of a prostate image classification device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0030] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0031] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0032] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0033] See below Figure 1 , an embodiment of the present invention provides a prostate image classification method, specifically comprising: steps S101 to S105.
[0034] Step S101: Acquire a prostate image, wherein the prostate image includes a cross-sectional view image and a sagittal view image;
[0035] Step S102: using a CNN-Transformer network with mixed-view learning to extract features from the cross-sectional view image and the sagittal view image respectively, and combining cross-view attention to perform feature enhancement, thereby obtaining cross-sectional view features and sagittal view features;
[0036] Step S103: dynamically fuse the cross-sectional view feature and the sagittal view feature using a hybrid view adaptive fusion module to obtain a fusion feature;
[0037] Step S104: performing classification processing on the fused features to obtain classification detection results of the prostate image, thereby constructing a hybrid view classification network;
[0038] Step S105: using the hybrid view classification network to classify the designated prostate image.
[0039] In this embodiment, combined with Figure 2 First, the images from the cross-sectional and sagittal perspectives are respectively input into the CNN-Transformer network of hybrid perspective learning for feature extraction. The high-quality image information from the orthogonal perspective is introduced through cross-perspective attention to achieve feature enhancement, thereby extracting features from the two perspectives. Next, the high-dimensional features of the two perspectives are sent to the hybrid perspective adaptive fusion module to achieve dynamic fusion of multi-perspective features. Finally, the fused features are sent to the fully connected layer to obtain the classification results. It can be seen that a hybrid perspective classification network is constructed, and this hybrid perspective classification network can be directly applied in the subsequent image processing process to classify the cross-sectional and sagittal perspective images collected from the specified prostate image.
[0040] This embodiment uses a CNN-Transformer hybrid architecture to extract fine-grained local features and model global dependencies for both cross-sectional and sagittal images. It then uses cross-view attention for feature enhancement to extract cross-sectional and sagittal features. A hybrid-view adaptive fusion module then dynamically aggregates features across channels and spatial dimensions to optimize overall representation capabilities. This embodiment leverages the complementary information between cross-sectional and sagittal views to improve medical image processing and provide a reference for subsequent medical diagnosis.
[0041] In particular, the prostate image classification method provided in this embodiment is suitable for training csPCa classification models in scenarios with classification labels and multi-view image data to assist doctors in reading transrectal ultrasound videos, locating cancerous areas, and providing navigation support for biopsy punctures. Compared with the single-view method, the mixed-view attention network of this embodiment fully exploits the complementary features of images from different perspectives, has better csPCa recognition capabilities, and does not require additional imaging modality information (such as SWE). The network enhances features by extracting complementary information between orthogonal perspectives through cross-view attention. In addition, this embodiment also introduces a mixed-view adaptive fusion module that can dynamically aggregate features in both channel and spatial dimensions to achieve more effective multi-view information integration.
[0042] In practical applications, a data set can be collected for training. Specifically, the data set can include 590 pairs of dual-view (cross-sectional and sagittal) transrectal ultrasound images of the prostate. According to the data set, the appropriate target size is selected to unify the image size, and the data distribution of the image is normalized. Since the sagittal data is a video sequence obtained by rotating the ultrasound probe 120° clockwise from 10 o'clock to 2 o'clock, rather than a 3D volume, in order to enable it to be spatially aligned with the cross-sectional data, this embodiment reconstructs the sagittal data according to the sampling rules. Specifically, each frame of the video sequence is spatially mapped according to its corresponding rotation angle, and reconstructed into 3D volume data by interpolation, such as Figure 3 As shown, the left picture is the transverse section and the right picture is the sagittal section.
[0043] In one embodiment, the CNN-Transformer network using mixed-view learning extracts features from the cross-sectional view image and the sagittal view image respectively, and enhances features by combining cross-view attention, thereby obtaining cross-sectional view features and sagittal view features, including:
[0044] Based on a dual-stream encoder architecture, a convolutional layer and a residual block are used to extract corresponding local features from the cross-sectional view image and the sagittal view image respectively;
[0045] Based on the local features, global modeling is performed using Transformer-based cross-view attention, and high-quality spatial information of the orthogonal view is obtained according to the global modeling and used as a query representation. The query representation is then used to optimize the features of another view to obtain optimized cross-sectional view features and sagittal view features.
[0046] The CNN-Transformer network for hybrid view learning proposed in this embodiment uses a dual-stream encoder to capture view-specific features, thereby facilitating feature learning within and across views. Each branch of the dual-stream encoder consists of an initial layer and four downsampling stages, each of which uses a hybrid CNN-Transformer architecture.
[0047] To achieve a modeling transition from local to global, each stage consists of two residual convolutional blocks followed by a cross-view attention module. The cross-view attention module leverages the attention mechanism from the transformer to build long-range dependencies. Specifically, the residual block extracts fine-grained spatial details and local structure, while the cross-view attention module fuses complementary spatial information from orthogonal views and captures long-range dependencies to optimize features.
[0048] By combining the local feature extraction capabilities of CNN with the global modeling capabilities driven by Transformer, the mixed-view network effectively captures fine-grained detail features and long-range dependencies, and utilizes complementary features from different perspectives to achieve image detection and classification tasks.
[0049] In one embodiment, based on the local features, global modeling is performed using Transformer-based cross-view attention, and high-quality spatial information of orthogonal views is obtained according to the global modeling and used as a query representation. The query representation is then used to optimize features of another view to obtain optimized cross-sectional view features and sagittal view features, including:
[0050] Performing perspective conversion on the local features of the current branch of the dual-stream encoder by dimensional permutation to obtain corresponding cross-sectional conversion features and sagittal plane conversion features;
[0051] A linear layer is used to map the conversion features of the current branch and the orthogonal features from the other branch respectively;
[0052] Based on the mapping results, spatial attention and channel attention are calculated respectively;
[0053] The local features of the cross-sectional view image and the sagittal view image, as well as the calculation results of the spatial attention and the channel attention, are combined and summed to obtain the corresponding cross-sectional view features and sagittal view features.
[0054] The slice thickness results in the spatial resolution along the scanning direction being lower than the in-plane spatial resolution. Although interpolation can solve this imbalance problem, it does not provide any information gain and may introduce additional errors. To solve this problem, this embodiment proposes a cross-view attention module that uses high-quality spatial information from orthogonal views to capture the global dependencies between multi-view images, thereby enhancing feature representation (module structure as shown in Figure 2). Figure 2 Cross-view attention acts as a bridge in the two-stream encoder, enabling the interaction of high-quality spatial information between different views.
[0055] Inspired by the attention mechanism for contextual spatial-channel feature aggregation, the cross-view attention module consists of spatial and channel attention modules, aiming to simultaneously capture spatial relationships and channel dependencies. The cross-view attention module adopts a shared query-key mechanism to improve computational efficiency while maintaining representational richness. The module applies attention across two mutually orthogonal views. Features from one view serve as high-quality queries to refine features from the other view, thereby capturing complementary information, enriching the feature map, and improving the network's multi-view learning capabilities.
[0056] Specifically, the cross - perspective attention module adopts parallel spatial and channel attention modules, and the features are transformed in perspective before applying attention. Taking the cross - sectional perspective features as an example (as shown in part (c) of Figure 2 ), the cross - sectional features are first transposed to the sagittal perspective, and cross - perspective attention is applied to the sagittal plane, using the sagittal - perspective features as high - quality query representations, and vice versa. Here, the cross - sectional features are denoted as F t ∈R B×C×H×W×D , and the sagittal - plane features are denoted as F s ∈R B×C×W×D×H , where B is the batch size, C is the number of channels, H×W×D represents the spatial dimension, the cross - section corresponds to H×W, and the sagittal plane corresponds to W×D. To better operate on features of different perspectives, the imaging axis is merged into the batch dimension, thus generating new cross - sectional features F t ∈R BD×C×H×W and sagittal - plane features F s ∈R BH×C×W×D . The cross - perspective attention first reshapes the features of the cross - sectional perspective into two - dimensional sagittal - plane planar features, denoted as F ts ∈R BH×C×W×D . Then, four different linear layers are used to map F s to the shared query representation Q shared , and map F ts to the shared key representation K shared , the spatial value V spatial and the channel value V channel respectively:
[0057] Q shared =W Q F s , K shared =W K F ts ,
[0058] where, W Q , W K , are the corresponding linear - layer parameters, and the final mapped dimension is HW×C. To reduce the computational burden of spatial attention, K shared and V spatial are projected to a lower - dimensional space using a learnable matrix, from HW×C to P×C, where P << HW. Then the spatial attention is calculated:
[0059]
[0060] where, Q shared , K proj , V projRepresent the shared query representation, the projected shared key representation, and the projected value representation, respectively. s represents a learnable scaling parameter. Channel attention captures the interdependence between feature channels in the channel dimension. This operation is performed between the channel value representation and the channel attention map. The channel attention formula is as follows:
[0061]
[0062] Among them, Q shared , K shared , V channel Represent the shared query, shared key and channel value, γ c Then represents the learnable scaling parameter. Finally, F is summed up t The outputs of the two attention modules are aggregated and the feature representation is further enhanced by the residual block.
[0063] In one embodiment, the hybrid view adaptive fusion module is used to dynamically fuse the cross-sectional view feature and the sagittal view feature to obtain the fused feature, including:
[0064] First, the cross-sectional view feature and the sagittal view feature are connected to obtain a connection feature;
[0065] Secondly, channel attention is used to enhance the discrimination of the connection features to obtain intermediate features;
[0066] Specifically, according to the following formula, the channel attention map is calculated based on the multi-layer perceptron:
[0067] A ch =σ(MLP(AvgPool s (F concat ))+MLP(MaxPool s (F concat )))
[0068] Among them, A ch ∈R C×1×1×1 represents the channel attention map, σ is the sigmoid activation function, MLP represents the shared multi-layer perceptron, AvgPool s with MaxPool s Represents global average pooling and maximum pooling in the spatial dimension respectively;
[0069] Perform element-by-element multiplication of the channel attention map and the splicing feature to obtain the intermediate feature
[0070] Thirdly, spatial attention is used to perform local spatial refinement of the intermediate features in the cross-sectional view and the sagittal view respectively;
[0071] Specifically, the spatial attention map is calculated based on the cross-sectional view feature and the sagittal view feature according to the following formula:
[0072] A sp =σ(Conv(Concat(AvgPool c (F),MaxPool c (F))))
[0073] Among them, A sp ∈R 1×H×W×D Represents the spatial attention map, F represents the cross-sectional view feature or sagittal view feature, Conv represents the 3×3×3 convolution operation, AvgPool c with MaxPool c Represents global average pooling and maximum pooling along the channel dimension respectively;
[0074] Performing view feature recovery processing on the intermediate features to obtain cross-sectional disassembly features and sagittal disassembly features;
[0075] Performing element-by-element multiplication of the spatial attention map with the cross-sectional decomposition feature and the sagittal decomposition feature, respectively, to obtain a local spatial refinement result of the corresponding cross-sectional view and a local spatial refinement result of the sagittal view;
[0076] Finally, the local spatial refinement result of the cross-sectional view and the local spatial refinement result of the sagittal view are spliced into the fusion feature.
[0077] In this embodiment, the hybrid perspective adaptive fusion module is set after the dual-stream encoder to fuse the multi-perspective features extracted by the dual-stream encoder to fully utilize the information from different perspectives and achieve accurate classification and recognition. This module can improve the fusion effect of multi-perspective features by dynamically re-weighting features in both channel and spatial dimensions. Figure 4 As shown in , the module first connects the feature maps of the two perspectives, then applies channel attention to strengthen the discriminative channel, and then applies spatial attention to refine the local spatial information in each perspective. Finally, the fused feature maps are spliced again to form the output result. Here, the feature maps of the two perspectives are set as F t ,F s ∈R C×H×W×D , the channel attention process is defined as follows:
[0078] F concat =Concat(F t ,F s )
[0079] F'=A ch (Fconcat )⊙F concat
[0080] Among them, Concat represents the concatenation operation along the channel dimension, F concat represents the feature obtained by splicing features from different perspectives, ⊙ represents element-by-element multiplication, and F' is the fusion feature improved by channel attention. Channel Attention Map A ch ∈R C ×1×1×1 The calculation formula is as follows:
[0081] A ch =σ(MLP(AvgPool s (F concat ))+MLP(MaxPool s (F concat )))
[0082] Among them, σ is the sigmoid activation function, MLP represents the shared multi-layer perceptron, AvgPool s with MaxPool s Represent the global average pooling and maximum pooling in the spatial dimension respectively. It should be noted that during the multiplication operation, the channel attention map will be broadcasted in the spatial dimension to match the input feature map dimension.
[0083] The spatial attention mechanism in the mixed-view adaptive fusion module operates on each view separately. It uses average and maximum pooling in the channel dimension to highlight significant spatial regions, thereby guiding the network to focus on key anatomical semantic information. First, features from different views are recovered from F':
[0084] F' t =F'[:C],F' s =F'[C:]
[0085] Among them, F' t is the cross-sectional feature extracted from the fusion feature F', s is the sagittal plane feature extracted from the fusion feature F'. Then the spatial attention is calculated and applied to the corresponding perspective:
[0086] F” t =A sp (F' t )⊙F' t ,F” s =A sp (F' s )⊙F' s ,F”=Concat(F” t ,F” s )
[0087] Among them, F” t It's F' t The feature is further refined by spatial attention, F” s It's F' s After the spatial attention is further improved, F' is the final fused feature. Spatial Attention Figure A sp ∈R 1×H×W×D The calculation is as follows:
[0088] A sp =σ(Conv(Concat(AvgPool c (F),MaxPool c (F))))
[0089] Among them, F represents the feature of a certain perspective, Conv represents the 3×3×3 convolution operation, AvgPool c with MaxPool c denote global average pooling and maximum pooling along the channel dimension, respectively.
[0090] By fusing channels and spatial attention mechanisms, the mixed-view adaptive fusion module can adaptively integrate features from different perspectives, thereby effectively aggregating complementary information and improving classification performance.
[0091] In one embodiment, the prostate image classification method further includes:
[0092] According to the following formula, the Focal Loss loss function is used to update the parameters of the mixed view classification network:
[0093] L = -α(1-p) γ log(p)
[0094] Among them, L represents the loss function, α is the class weight, γ is the focus parameter, and p is the model's predicted probability for the positive class.
[0095] This example uses Focal Loss as the loss function for model training. Traditional cross-entropy loss functions perform poorly when dealing with class imbalance because they overly focus on easily categorized samples while ignoring difficult samples that are difficult to distinguish. Focal Loss, on the other hand, can increase its focus on minority samples by adjusting weights.
[0096] In addition, in the actual model training process, it can be implemented based on the PyTorch framework using an NVIDIA V100 GPU with 32G memory. The network is optimized for 200 iterations (epochs) using a stochastic gradient descent optimizer with an initial learning rate of 0.0001 and a (1-epoch / 200) 0.9Polynomial learning rate adjustment strategy for .
[0097] Figure 5 This is a schematic block diagram of a prostate image classification device 500 provided in an embodiment of the present invention. The device 500 includes:
[0098] An image acquisition unit 501 is configured to acquire a prostate image, wherein the prostate image includes a cross-sectional view image and a sagittal view image;
[0099] A feature extraction unit 502 is configured to extract features from the cross-sectional view image and the sagittal view image respectively using a CNN-Transformer network with mixed-view learning, and perform feature enhancement in combination with cross-view attention, thereby obtaining cross-sectional view features and sagittal view features;
[0100] A feature fusion unit 503 is configured to dynamically fuse the cross-sectional view feature and the sagittal view feature using a hybrid view adaptive fusion module to obtain a fused feature;
[0101] A feature classification unit 504 is configured to classify the fused features to obtain a classification detection result of the prostate image, thereby constructing a hybrid view classification network;
[0102] The classification processing unit 505 is used to perform classification processing on the designated prostate image using the hybrid view classification network.
[0103] In one embodiment, the feature extraction unit 502 includes:
[0104] A local extraction unit, configured to extract corresponding local features from the cross-sectional view image and the sagittal view image respectively using a convolutional layer and a residual block based on a dual-stream encoder architecture;
[0105] A global construction unit is used to perform global modeling using Transformer-based cross-view attention, and obtain high-quality spatial information of orthogonal views based on the global modeling, and use it as a query representation. The query representation is then used to optimize the features of another view to obtain optimized cross-sectional and sagittal view features.
[0106] In one embodiment, the global construction unit includes:
[0107] A perspective conversion unit, configured to perform perspective conversion on the local features of the current branch of the dual-stream encoder by dimensional permutation to obtain corresponding cross-sectional conversion features and sagittal plane conversion features;
[0108] A linear mapping unit, configured to map the conversion features of the current branch and the orthogonal features from another branch using a linear layer;
[0109] Attention calculation unit, used to calculate spatial attention and channel attention based on the mapping results;
[0110] The summation calculation unit is used to combine the local features of the cross-sectional perspective image and the sagittal perspective image, as well as the calculation results of the spatial attention and the channel attention, to perform summation calculation to obtain corresponding cross-sectional perspective features and sagittal perspective features.
[0111] In one embodiment, the feature fusion unit 503 includes:
[0112] a feature connection unit, configured to connect the cross-sectional view feature and the sagittal view feature to obtain a connection feature;
[0113] An enhanced discrimination unit, configured to utilize channel attention to perform enhanced discrimination on the connection features to obtain intermediate features;
[0114] A local refinement unit, configured to perform local spatial refinement of the cross-sectional view and the sagittal view on the intermediate features using spatial attention;
[0115] The result splicing unit is used to splice the result of the local space refinement of the cross-sectional view and the result of the local space refinement of the sagittal view into the fusion feature.
[0116] In one embodiment, the enhancement determination unit includes:
[0117] The channel calculation unit is used to calculate the channel attention map based on the multi-layer perceptron according to the following formula:
[0118] A ch =σ(MLP(AvgPool s (F concat ))+MLP(MaxPool s (F concat )))
[0119] Among them, A ch ∈R C×1×1×1 represents the channel attention map, σ is the sigmoid activation function, MLP represents the shared multi-layer perceptron, AvgPool s with MaxPool s Represents global average pooling and maximum pooling in the spatial dimension respectively;
[0120] The first multiplication unit is used to perform element-by-element multiplication of the channel attention map and the splicing feature to obtain the intermediate feature.
[0121] In one embodiment, the local refinement unit includes:
[0122] A spatial calculation unit is configured to calculate a spatial attention map based on the cross-sectional view feature and the sagittal view feature according to the following formula:
[0123] A sp =σ(Conv(Concat(AvgPool c (F),MaxPool c (F))))
[0124] Among them, A sp ∈R 1×H×W×D Represents the spatial attention map, F represents the cross-sectional view feature or sagittal view feature, Conv represents the 3×3×3 convolution operation, AvgPool c with MaxPool c Represents global average pooling and maximum pooling along the channel dimension respectively;
[0125] a restoration and disassembly unit, configured to perform a view feature restoration process on the intermediate features to obtain cross-sectional disassembly features and sagittal disassembly features;
[0126] The second multiplication unit is used to perform element-by-element multiplication of the spatial attention map with the cross-sectional decomposition features and the sagittal decomposition features, respectively, to obtain the corresponding local spatial refinement results of the cross-sectional perspective and the local spatial refinement results of the sagittal perspective.
[0127] In one embodiment, the prostate image classification device 500 further includes:
[0128] A parameter updating unit is used to update the parameters of the hybrid view classification network using the FocalLoss loss function according to the following formula:
[0129] L = -α(1-p) γ log(p)
[0130] Among them, L represents the loss function, α is the class weight, γ is the focus parameter, and p is the model's predicted probability for the positive class.
[0131] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, and will not be repeated here.
[0132] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When executed, the computer program can implement the steps provided in the above embodiments. The storage medium can include a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code.
[0133] The present invention also provides a computer device that may include a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, the steps provided in the above embodiment can be implemented. Of course, the computer device may also include various network interfaces, a power supply, and other components.
[0134] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
[0135] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
Claims
1. A prostate image classification method, characterized in that: include: Acquiring a prostate image, wherein the prostate image includes a transverse view image and a sagittal view image; A CNN-Transformer network with mixed-view learning is used to extract features from the cross-sectional view image and the sagittal view image respectively, and feature enhancement is performed in combination with cross-view attention to obtain cross-sectional view features and sagittal view features; Dynamically fusing the cross-sectional view features and the sagittal view features using a hybrid view adaptive fusion module to obtain a fusion feature; Classifying the fused features to obtain a classification detection result of the prostate image, thereby constructing a hybrid view classification network; The hybrid view classification network is used to classify the designated prostate image.
2. The prostate image classification method according to claim 1, characterized in that: The CNN-Transformer network using mixed-view learning extracts features from the cross-sectional view image and the sagittal view image respectively, and enhances features by combining cross-view attention to obtain cross-sectional view features and sagittal view features, including: Based on a dual-stream encoder architecture, a convolutional layer and a residual block are used to extract corresponding local features from the cross-sectional view image and the sagittal view image respectively; Based on the local features, global modeling is performed using Transformer-based cross-view attention, and high-quality spatial information of the orthogonal view is obtained according to the global modeling and used as a query representation. The query representation is then used to optimize the features of another view to obtain optimized cross-sectional view features and sagittal view features.
3. The prostate image classification method according to claim 2, characterized in that: Based on the local features, global modeling is performed using Transformer-based cross-view attention, and high-quality spatial information of the orthogonal view is obtained according to the global modeling, and used as a query representation. The query representation is then used to optimize the features of another view to obtain optimized cross-sectional view features and sagittal view features, including: Performing perspective conversion on the local features of the current branch of the dual-stream encoder by dimensional permutation to obtain corresponding cross-sectional conversion features and sagittal plane conversion features; A linear layer is used to map the conversion features of the current branch and the orthogonal features from the other branch respectively; Based on the mapping results, spatial attention and channel attention are calculated respectively; The local features of the cross-sectional view image and the sagittal view image, as well as the calculation results of the spatial attention and the channel attention, are combined and summed to obtain the corresponding cross-sectional view features and sagittal view features.
4. The prostate image classification method according to claim 1, characterized in that: The hybrid view adaptive fusion module is used to dynamically fuse the cross-sectional view features and the sagittal view features to obtain fusion features, including: Connecting the cross-sectional view feature and the sagittal view feature to obtain a connection feature; Using channel attention to enhance the discrimination of the connection features to obtain intermediate features; Using spatial attention to perform local spatial refinement of the intermediate features in the cross-sectional view and the sagittal view; The local space refinement result of the cross-sectional view and the local space refinement result of the sagittal view are spliced together to form the fusion feature.
5. The prostate image classification method according to claim 4, characterized in that: The use of channel attention to enhance the discrimination of the connection features to obtain intermediate features includes: According to the following formula, the channel attention map is calculated based on the multi-layer perceptron: A ch =σ(MLP(AvgPool s (F concat ))+MLP(MaxPool s (F concat ))) Among them, A ch ∈R C×1×1×1 represents the channel attention map, σ is the sigmoid activation function, MLP represents the shared multi-layer perceptron, AvgPool s with MaxPool s Represents global average pooling and maximum pooling in the spatial dimension respectively; The channel attention map is element-wise multiplied with the concatenated feature to obtain the intermediate feature.
6. The prostate image classification method according to claim 4, characterized in that: The utilizing spatial attention to perform local spatial refinement of the cross-sectional view and the local spatial refinement of the sagittal view on the intermediate features respectively includes: The spatial attention map is calculated based on the cross-sectional view feature and the sagittal view feature according to the following formula: A sp =σ(Conv(Concat(AvgPool c (F),MaxPool c (F)))) Among them, A sp ∈R 1×H×W×D Represents the spatial attention map, F represents the cross-sectional view feature or sagittal view feature, Conv represents the 3×3×3 convolution operation, AvgPool c with MaxPool c Represents global average pooling and maximum pooling along the channel dimension respectively; Performing view feature recovery processing on the intermediate features to obtain cross-sectional disassembly features and sagittal disassembly features; The spatial attention map is respectively multiplied element-by-element with the cross-sectional decomposition feature and the sagittal decomposition feature to obtain the corresponding local spatial refinement result of the cross-sectional perspective and the local spatial refinement result of the sagittal perspective.
7. The prostate image classification method according to claim 1, characterized in that: Also includes: According to the following formula, the FocalLoss loss function is used to update the parameters of the mixed view classification network: L=-α(1-p) γ log(p) Among them, L represents the loss function, α is the class weight, γ is the focus parameter, and p is the model's predicted probability for the positive class.
8. A prostate image classification device, characterized in that: include: An image acquisition unit, configured to acquire a prostate image, wherein the prostate image includes a cross-sectional view image and a sagittal view image; a feature extraction unit, configured to extract features from the cross-sectional view image and the sagittal view image respectively using a CNN-Transformer network learned through mixed viewpoints, and perform feature enhancement in combination with cross-viewpoint attention, thereby obtaining cross-sectional view features and sagittal view features; a feature fusion unit, configured to dynamically fuse the cross-sectional view feature and the sagittal view feature using a hybrid view adaptive fusion module to obtain a fused feature; a feature classification unit, configured to classify the fused features to obtain a classification detection result of the prostate image, thereby constructing a hybrid view classification network; A classification processing unit is used to perform classification processing on a designated prostate image using the hybrid view classification network.
9. A computer device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the prostate image classification method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the prostate image classification method according to any one of claims 1 to 7 is implemented.