PSMA PET / CT and enhanced MRI multi-modal image registration and fusion method and system

Through the deep learning pyramid registration network model, combined with semantic gating convolution and convolution length and short-time memory module, high-precision registration and fusion of PSMA PET/CT and enhanced MRI is achieved, solving the problems of high misdiagnosis rate and poor robustness in the existing technology, and improving the accuracy of prostate cancer diagnosis.

CN120339343APending Publication Date: 2025-07-18SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510263305.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing prostate PSMA PET/CT and enhanced MRI registration and fusion methods have the problems of high misdiagnosis rate and poor robustness, and it is difficult to achieve high-precision image registration and information fusion, resulting in insufficient diagnostic accuracy of prostate cancer.

Method used

Using a pyramid-type registration network model based on deep learning, combining semantic gating convolution module (SGC), convolutional length and short-term memory module (U-CLSTM) and pyramid structure, through multi-scale and multi-stage registration network design, PSMA PET/CT and enhanced MRI information are integrated to generate high-precision fusion images.

Benefits of technology

It significantly improves the accuracy of prostate cancer diagnosis, reduces misdiagnosis and missed diagnosis, provides scientific treatment plan support, and improves the reliability and accuracy of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339343A_ABST
    Figure CN120339343A_ABST
Patent Text Reader

Abstract

The invention discloses a prostate PSMA PET / CT and enhanced MR multi-modal image registration and fusion method based on deep learning. The method comprises the following steps: preprocessing acquired PSMA PET / CT and enhanced MRI data; a multi-scale multi-stage registration network comprising the semantic gating convolution module and a U-CLSTM module is constructed: the SGC module enhances global perception ability through a V channel, the U-CLSTM module reinforces local feature attention by using a U channel, deformation field iterative optimization is performed in combination with a pyramid structure, and a multi-scale multi-stage registration network comprising the semantic gating convolution module and the U-CLSTM module is constructed; and finally, generating a high-precision registration image through Y channel weighted fusion. According to the method, global information self-adaption and local feature enhancement are combined, the registration precision of a gland anatomical structure and a tumor focus is synchronously improved through a multi-scale pyramid structure, PSMAPT / CT functional metabolism information and MRI anatomical structure information are effectively integrated, the generated three-dimensional fusion image can visually display the spatial position, morphological features and invasion range of a tumor, and the accuracy of the three-dimensional fusion image is improved. A multi-modal image basis with high space consistency is provided for clinical diagnosis, and the diagnosis accuracy of the prostate cancer is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical imaging, and specifically relates to a method and system for multi-modal image registration and fusion of prostate-specific membrane antigen (PSMA) PET / CT and enhanced MRI based on deep learning. Background Art

[0002] The diagnosis and treatment of prostate diseases rely on high-precision medical imaging technologies. Prostate-specific membrane antigen (PSMA) PET / CT and enhanced magnetic resonance imaging (MRI) are two commonly used medical imaging technologies. PSMA PET / CT utilizes the high expression characteristics of prostate-specific membrane antigen (PSMA) in prostate cancer cells, and binds radioactive-labeled PSMA ligands to cancer cells, thereby achieving precise localization of prostate cancer. This method has high sensitivity and specificity, can detect small lesions, and performs excellently especially in evaluating the invasion range of tumors, lymph node metastasis, and recurrence monitoring. However, PSMA PET / CT also has a certain misdiagnosis rate, especially for some PSMA-positive lesions that are not prostate cancer, as well as false-positive results caused by physiological activities or inflammation. Enhanced MRI improves the contrast between prostate tissues by intravenous injection of a contrast agent, thereby more clearly showing the vascular structure and blood flow in the prostate. This method is of great value in judging the blood supply of lesions in the prostate, evaluating tumor activity, and guiding puncture biopsy. However, enhanced MRI also has certain limitations in detecting small lesions and evaluating tumor invasiveness, and the positive predictive value is relatively low.

[0003] Therefore, combining PSMA PET / CT and enhanced MRI can make full use of the advantages of the two technologies and improve the accuracy and reliability of prostate disease diagnosis. However, at present, the registration and fusion of prostate PSMA PET / CT and enhanced MRI are still in their infancy, and direct combination still leads to diagnostic failure. Traditional registration methods such as those based on feature points, feature lines, or feature surfaces can achieve image registration to a certain extent, but due to the complex prostate morphology and large individual differences, it is often difficult to achieve an ideal registration effect. In addition, these methods often ignore the detailed information and local features in the images, resulting in the accuracy and robustness of the registration results needing to be improved.

[0004] Therefore, developing an image registration method based on PSMA PET / CT and enhanced MRI to comprehensively and accurately fuse the information of the two has become a hot and difficult point in current research. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method for registering and fusing prostate PSMA PET / CT and enhanced MRI, realizing the registration of prostate PSMA PET / CT and enhanced MRI medical images, and then generating a fused image containing prostate structural and functional information according to the registration result, so as to improve the diagnostic accuracy of prostate cancer. The invention designs a pyramid-shaped registration network model for prostate PSMA PET / CT and enhanced MRI, utilizes the characteristics of different channel information of PSMA PET / CT images, extracts and processes features at both the global and local scales. At the global scale, the model increases the receptive field to capture the information of the entire gland, and at the local scale, the model strengthens the attention mechanism to focus on tumor lesions, which can effectively integrate the advantageous information of the two images, overcome the limitations of single-modal diagnosis, thus significantly improving the diagnostic accuracy of prostate cancer, reducing the occurrence of misdiagnosis and missed diagnosis, providing strong support for formulating scientific and effective treatment plans clinically, and promoting the development of prostate cancer diagnosis technology.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] First, the present invention provides a method for registering and fusing prostate PSMA PET / CT and enhanced MRI, including the design of a registration model and image fusion based on the registration result. The specific steps are as follows:

[0008] S1: Acquisition and preprocessing of PSMA PET / CT and enhanced MRI;

[0009] S2: Design of a semantic gated convolution module (Semantic Gated Convolution, SGC) based on the V channel of PSMA PET / CT

[0010] to enhance the global perception ability of the registration network model;

[0011] S3: Design of a convolutional long short-term memory module (U-CLSTM) based on the U channel of PSMA PET / CT to improve the local attention of the registration

[0012] network model;

[0013] S4: Estimation of the registration deformation field based on the pyramid structure;

[0014] S5: Design of a multi-scale and multi-stage registration network model based on the Y channel of PSMA PET / CT and enhanced MRI;

[0015] S6: Obtaining the fusion result by weighting PSMA PET / CT and enhanced MRI with the deformation field generated by the registration network model. In step one, the acquisition and preprocessing of PSMA PET / CT and enhanced MRI include:

[0016] First, obtain PSMA PET / CT and enhanced MRI data, and perform image format conversion, image size normalization, image enhancement, and dataset partitioning. Through the following formula, convert the PSMA PET / CT in the form of an RGB color image into an image in YUV format:

[0017] Y = 0.299R + 0.587G + 0.114B

[0018] U = -0.169R - 0.331G + 0.5B

[0019] V = 0.5R - 0.419G - 0.081B

[0020] where the Y channel M y is a grayscale image, and the U and V channels are grayscale images containing chromaticity information. Then, through threshold segmentation of the grayscale of the U and V channel images, generate a U component M u feature map sensitive to tumor region information and a V component M v feature map sensitive to the prostate region.

[0021] In step two, design a semantic gating convolutional module (SGC) based on the V channel of PSMA PET / CT to enhance the global perception ability of the registration network model:

[0022] Step 2.1: Construct a semantic encoding module to extract the semantic information of the prostate gland from the input V channel feature map M v containing the overall prostate region information, reduce the spatial resolution through the pooling layer, retain important global glandular information, and finally encode it into a feature form suitable for subsequent processing.

[0023] Step 2.2: Construct a channel interaction module to project the feature map output by the semantic encoding module into a space matching the convolutional kernel dimension, and realize the interaction and conversion of information between different channels through a grouped linear layer. Finally, adjust the dimension of the feature representation to adapt to the convolutional kernel dimension requirements.

[0024] Step 2.3: Construct a gate decoding module. According to the outputs of the semantic encoding module and the channel interaction module, decode the semantic information into a gate of the same size as the convolutional kernel, and realize the per-element modulation of the convolutional kernel weights, thereby integrating the global glandular semantic information into the convolutional operation.

[0025] In step three, design a convolutional long short-term memory module (U-CLSTM) based on the U channel of PSMA PET / CT to improve the local attention of the registration network model:

[0026] Step 3.1: Randomly initialize the hidden state of the U-CLSTM module, and then input the local information of the U-channel component of the PSMA PET / CT image and the multi-level deformation field into the U-CLSTM module.

[0027] Step 3.2: The input gate i i controls the amount of new data adopted from the input . If the input gate is activated, the input information is accumulated in the storage unit to introduce new data.

[0028] Step 3.3: The forget gate f controlled by the feature map weight i determines the degree of memory retention. If the forget gate is open, the cell states with low weights may be forgotten during this process, and important key information is retained.

[0029] Step 3.4: The output gate o i decides whether to pass all memory cells to the subsequent part or retain the information of the memory cells without update.

[0030] In step four, the registration deformation field estimation based on the pyramid structure:

[0031] Step 4.1: Design the recursive operation. First, combine the fixed feature F i , the moving feature after registration in the previous stage and the feature Z i , and then extract the global feature R i through semantic convolution.

[0032] Step 4.2: The processed feature passes through the deformation estimation module containing the deformable convolutional layer DeformConv and the difference layer Diff to predict the deformation field

[0033] Step 4.3: Input the multi-level deformation field and the U-channel information into the U-CLSTM structure, and output the deformation field with enhanced local attention.

[0034] Step 4.4: Then reduce the number of channels of the feature map R i through the DeConv module, and finally upsample the feature map R i and transfer it to the next layer.

[0035] Step 4.5: Finally, set the recursive module of the pyramid structure, and set different numbers of recursive modules (n i = 2, 2, 3, 1) for each layer structure (i = 1, 2, 3, 4).

[0036] In step five, design of a multi-scale and multi-stage registration network model based on the PSMA PET / CT Y-channel and enhanced MRI: Step 5.1: First, construct an encoder for feature extraction. In the present invention, ResNet is used as the backbone network of the encoder, and a two-stream encoder strategy with non-shared weights is combined to independently encode the features from M y and F.

[0037] Step 5.2: Then, construct a recursive pyramid decoder. The decoder has four decoding layers, and each decoding layer includes a semantic convolution module (SGC), a convolutional long short-term memory module based on the U-channel (U-CLSTM), and a deformation estimation module. After being processed by the above modules, the features output a deformation field In the four decoding layers, the lower decoding layers regress large deformations on low-scale features, and the higher decoding layers regress small errors on large-scale features.

[0038] Step 5.3: Combine the encoder and the decoder to form a registration network model, and then input the processed enhanced MRI, PSMA PET / CT Y-channel M y , U-channel M u and V-channel M v into the registration network model.

[0039] In step six, obtain a fusion result by weighting the PSMA PET / CT and enhanced MRI with the deformation field generated by the registration network model: Apply the registration deformation field generated in the last stage of the registration network model to the PSMA PET / CT image, and perform image weighting on it with the enhanced MRI to generate the final fusion result.

[0040] Second, the present invention also provides a prostate PSMA PET / CT and enhanced MRI registration and fusion system, including: Module 1: A data preprocessing module, which is used to perform image format conversion, image size normalization, image enhancement, and dataset division on the PSMA PET / CT and enhanced MRI. For the processed PSMA PET / CT image, use the conversion formula to convert the RGB-channel image into a YUV channel, and perform threshold segmentation on the U and V channel components to generate a U-component feature map sensitive to tumor region information and a V-component feature map sensitive to the prostate region;

[0041] Module 2: A semantic gating convolution module. Design a semantic gating convolution module (SGC), and respectively construct a semantic encoding module, a channel interaction module, and a gate decoding module;

[0042] Module 3: Convolutional Long Short-Term Memory Module Based on U Channel. Design a Convolutional Long Short-Term Memory Module Based on U Channel (U-CLSTM). The module takes the U channel of the PSMA PET / CT image and the multi-level deformation field as inputs, designs three control gates: input gate, forget gate, and output gate, and generates a local enhanced deformation field.

[0043] Module 4: Multi-Scale Multi-Stage Encoder-Decoder Module. Based on the Y channel of PSMA PET / CT and enhanced MRI, construct a multi-scale multi-stage two-stream encoder and a recursive decoder.

[0044] Module 5: Deformation Estimation Module. Recursively fuse multi-level features based on the traditional pyramid structure, recursively combine the SGC module, U-CLSTM module, and deformation estimation module. After recursion, upsample the features and deformation field and input them to the next layer.

[0045] Module 6: Image Weighting Module. Deform the PSMA PET / CT based on the deformation field at the last stage, weight the deformed image and the enhanced MRI, and output the fusion result.

[0046] Module 1 specifically includes the following: perform image format conversion, image size normalization, image enhancement, and dataset partitioning on PSMA PET / CT and enhanced MRI. Considering the difference that PSMA PET / CT is a three-channel RGB color image while enhanced MRI is a single-channel grayscale image, it is proposed to convert the PSMA PET / CT image to the YUV channel. The Y channel highlights the grayscale information, the U channel captures local lesion features, and the V channel enhances the overall prostate region information, laying a foundation for subsequent feature extraction. Finally, input the processed enhanced MRI, PSMA PET / CT Y channel, U channel, and V channel into the registration network model.

[0047] Module 2 specifically includes the following: separately construct a semantic encoding module, a channel interaction module, and a gate decoding module. The semantic encoding module in Semantic Gated Convolution (SGC) first extracts global context information from the input feature map and encodes it into a latent representation. Then, through the channel interaction module and the gate decoding module, this global context information is mapped to the modulation process of the convolutional kernel. Specifically, the gate G (composed of G1 and G2) generated by the gate decoding module is used to modulate the weights of the convolutional kernel. After passing through the above modules, when the convolutional kernel performs a convolution operation, its weights can be dynamically adjusted according to the global context information. Module 3 specifically includes the following: Based on the concept of Convolutional Long Short-Term Memory Network (CLSTM), the invention abstracts the pyramid hierarchical stage as time and designs a U-CLSTM module that combines the U-channel of PSMA PET / CT images and the multi-level deformation displacement field. The module includes three control gates: an input gate, a forget gate, and an output gate. The input gate controls the amount of new displacement field information adopted and accumulates information when activated; the forget gate is controlled by the U-channel information and determines the degree of memory retention according to the weights; the output gate determines whether the memory unit is passed or retained. The storage unit acts as an accumulator, and the control gates operate on the deformation field. By combining the storage unit and the deformation field, the parameter calculation is reduced, the calculation efficiency is improved, and the attention to the local tumor information area is enhanced, retaining the motion specificity of the lesion.

[0048] Module 4 specifically includes the following: Based on the Y-channel of PSMA PET / CT and MRI, construct a multi-scale multi-stage two-stream encoder and a recursive decoder. The encoder uses a ResNet backbone network to extract multi-level features; the decoder gradually predicts the deformation field through a pyramid structure, combines Semantic Gated Convolution (SGC) and the U-channel-based Convolutional Long Short-Term Memory Network (U-CLSTM), and recursively fuses features at different scales to optimize the registration result.

[0049] Module 5 specifically includes the following: The present invention recursively fuses multi-level features using a traditional pyramid structure, sets a specific number of recursive modules for each layer, and after completion of the recursion, up-samples the features and the deformation field and inputs them to the next layer. The recursion includes performing semantic convolution on the fused features, generating a deformation field through a deformation module, generating a new deformation field and a feature map through a U-CLSTM structure. The deformation estimation module is based on ordinary differential equations and uses a deformation convolutional layer and a difference layer to generate a deformation field, ensuring the differentiability of the deformation field and reducing the irreversible deformation field estimation. This method can effectively fuse features, generate a reversible high-quality deformation field, reduce the proportion of unreasonable deformation fields, and improve the registration accuracy and reliability.

[0050] Module 6 specifically includes the following: Deform PSMA PET / CT based on the deformation field of the last stage, weight the deformed image with enhanced MRI, and output the fusion result.

[0051] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0052] 1. The multi-scale feature fusion and semantic gating convolution (SGC) module effectively solves the common problem of difficult multi-modal feature fusion in the pyramid registration network. The convolution long short-term memory network based on the U-channel (U-CLSTM) and the pyramid-based deformation field estimation enhance the attention to the target area;

[0053] 2. The SGC module and the U-CLSTM module solve the unique challenges of prostate imaging, provide a comprehensive evaluation of the prostate, and also pay attention to local cancer lesions. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a flowchart of the prostate PSMA PET / CT and enhanced MRI registration and fusion method in the embodiment of the present invention

[0055] Figure 2 It is a schematic diagram of the modules of the prostate PSMA PET / CT and enhanced MRI registration and fusion method in the embodiment of the present invention Figure 3 It is a schematic diagram of the enhanced MRI, prostate PSMA PET / CT and each component data in the embodiment of the present invention

[0056] Figure 4 It is a schematic diagram of the prostate PSMA PET / CT and enhanced MRI registration network in the embodiment of the present invention

[0057] Figure 5 It is a schematic diagram of the semantic convolution module in the embodiment of the present invention

[0058] Figure 6 It is a schematic diagram of the convolution long short-term memory module based on the U-channel in the embodiment of the present invention

[0059] Figure 7 It is a schematic diagram of the deformation estimation module in the embodiment of the present invention

[0060] Figure 8 It is the registration result in the embodiment of the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the relevant application scenarios of the present invention.

[0062] The prostate PSMA PET / CT and enhanced MRI registration and fusion method, as Figure 1 shown, specifically includes the following steps:

[0063] S1: Acquisition and preprocessing of PSMA PET / CT and contrast-enhanced MRI;

[0064] S2: Design of a semantic gating convolutional module (SGC) based on the V-channel of PSMA PET / CT to enhance the registration network model

[0065] Global perception ability;

[0066] S3: Design of a convolutional long short-term memory module (U-CLSTM) based on the U-channel of PSMA PET / CT to improve the local attention of the registration

[0067] network model;

[0068] S4: Estimation of the registration deformation field based on a pyramid structure;

[0069] S5: Design of a multi-scale multi-stage registration network model based on the Y-channel of PSMA PET / CT and contrast-enhanced MRI;

[0070] S6: Obtaining the fusion result by weighting PSMA PET / CT and contrast-enhanced MRI with the deformation field generated by the registration network model. To achieve the registration of PSMA PET / CT and contrast-enhanced MRI, the present invention constructs Figure 3 the deep learning network architecture shown below, and the present embodiment will be described in detail below.

[0071] In the prostate cancer diagnosis scenario, contrast-enhanced MRI is usually used as the fixed image because it can more clearly show the anatomical details of the prostate and surrounding tissues. Therefore, PSMA PET / CT and contrast-enhanced MRI are respectively referred to as the moving image M and the fixed image F.

[0072] In step one, acquisition and preprocessing of PSMA PET / CT and contrast-enhanced MRI:

[0073] The imaging data in the present invention is provided by the Department of Urology, Ruijin Hospital, Shanghai Jiao Tong University, and includes a total of 77 cases of PSMA PET / CT images and contrast-enhanced MRI. We set the contrast-enhanced MRI as the fixed image and the PSMA PET / CT as the moving image. For the convenience of training, all images are uniformly sampled to a resolution of 160×192×160, and the gray values are normalized to the interval [0,1]. In addition, to expand the dataset, B-spline non-rigid transformation is used for 4-fold data augmentation, and finally 308 samples are obtained, of which 244 are used for training and 64 are used for testing.

[0074] Since PSMA PET / CT is a three-channel RGB color image while enhanced MRI is a single-channel grayscale image, achieving channel consistency is a prerequisite for accurate registration. YUV channel images store grayscale information in the Y channel and chromaticity information in the U and V channels, which is beneficial for expressing the highlighted information of PSMA PET / CT. To solve the problem of channel inconsistency, we first transform the RGB image channels into YUV channels through the following formula:

[0075] Y = 0.299R + 0.587G + 0.114B

[0076] U = -0.169R - 0.331G + 0.5B

[0077] V = 0.5R - 0.419G - 0.081B

[0078] This is a common method for registering multi-channel functional images with single-channel grayscale images. Specifically, the PSMA PET / CT image M is converted into the YUV space to obtain the Y, U, and V components: M y 、M u 、M v 。The M y of the Y component and the enhanced MRI represent the single-channel grayscale image. For the M u and M v components, as Figure 3 shows, due to the difference in grayscale values, the U component M u that is sensitive to the information of the lesion area is obtained by threshold segmentation, and the V component M v that is sensitive to the prostate area. Finally, the processed enhanced MRI, PSMA PET / CT Y-channel M y 、U-channel M u and V-channel M v are input into the registration network model.

[0079] In step two, based on the semantic gated convolution (SGC) design of the PSMA PET / CT V channel, global perception is added to improve the large-range registration effect:

[0080] Neuroscience research shows that global context information is crucial in visual scene interpretation. Inspired by this, the present invention specifically designs an adaptive convolution modulation strategy, which mainly includes a semantic gated convolution module (SGC), which inputs the semantic information of M v into the module. After extracting the semantic information, the weights of the convolution are modulated using the diversity of information expressed in different channels, greatly improving the ability of the feature map to express global information after convolution. As Figure 5As shown, the SGC module consists of three parts: a semantic encoding module, a channel interaction module, and a gated decoding module.

[0081] Step 2.1: The semantic encoding module first downsamples the input of size c i ×h i ×w i ×d i to c ×h′ i ×w′ i ×d′ i ×d′ i , where c i is the number of channels. Subsequently, it is input into the semantic encoding module containing "fully connected layer, normalization, activation function", and while retaining the channel information, the remaining features are flattened into a one-dimensional vector v i , where v i =k×k×k / 2, k represents the convolutional kernel size. After being processed by this module, the output feature size is c i ×v i .

[0082] Step 2.2: The channel interaction module converts the input feature from the size of c i ×v i to p i ×v i . Specifically, the module includes "grouped fully connected layer, normalization, activation function". Compared with the ordinary linear layer, the grouped linear layer uses the grouping strategy to group the weight matrix, thereby reducing the computational amount and the number of parameters of the model.

[0083] Step 2.3: In the gated decoding module, the input feature of size c i ×v i is converted into a fully connected layer of size c i ×k×k×k, and then the input feature of size p i ×v i is converted into p i ×k×k×k. The two groups of converted features are copied along two directions to obtain a feature of size p i ×c i ×k×k×k. Finally, the Sigmoid function is used to implement gating to control the transmission of information flow and enhance feature selectivity.

[0084] Finally, the output is fused with the convolutional kernel using per-pixel multiplication. Through these steps, the semantic gated convolution module enhances the global feature representation in the fused feature map and effectively guides the subsequent deformation field estimation.

[0085] In Step 3, the accuracy is further improved by adding local attention based on the convolutional long short-term memory (U-CLSTM) module of the PSMA PET / CT U-channel:

[0086] The convolutional long short-term memory network (CLSTM) is a structure designed to capture long-term temporal information through the feedback of forget gates and memory cells. The U-CLSTM module of the present invention is based on the CLSTM concept, abstracting the concept of the pyramid structure hierarchy into temporal information for processing multi-level deformation displacement fields and combining lesion-specific motion information. Specifically, we will local information and multi-level deformation displacement fields into the U-CLSTM module, and randomly initialize the hidden state.

[0087] The key innovation of U-CLSTM lies in the combination of the storage unit and the deformation field. The storage unit, as an accumulator of state information, accesses, writes, and clears the deformation field through multiple parameterized control gates. As Figure 6 shown, the U-CLSTM module includes three gates: the input gate i i controls how much new data from the input is used. If the input gate is activated, its information will be accumulated into the unit. The forget gate f controlled by the i weight information determines how much of the past memory should be retained. If the forget gate is open, the past cell state may be forgotten during this process. The final output gate o i decides whether to pass all memory units to the subsequent part or retain the information of the memory units without updating.

[0088]

[0089]

[0090] H i-1 = o i · tanh(c i )

[0091]

[0092] where σ represents the sigmoid activation function, * represents convolution, · represents the Hadamard product, and W is the weight matrix. This module effectively retains the specificity of lesion motion while reducing unnecessary parameter calculations and greatly improving the computational efficiency.

[0093] In Step 4, the registration deformation field estimation based on the pyramid structure:

[0094] Recursively fuse multi-level features (i = 1, 2, 3, 4) by using the traditional pyramid structure combined with the semantic gating convolution module. Each stage has n i recursive modules (n i = 2, 2, 3, 1). After completing n i recursions at the i-th layer, the fused feature map R i and the deformation field are upsampled and passed to the next layer:

[0095] Z i+1 = transConv(R i )

[0096]

[0097] where transConv represents a set of operations of "3D transposed convolution, instance normalization, Leaky ReLU", and upSamp represents trilinear interpolation for upsampling. After obtaining the upsampled feature map and the deformation field, the following recursion can be performed:

[0098]

[0099] Z i = DeConv(R i )

[0100] where * represents semantic convolution. For (i = 2, 3, 4), DeConv represents the combination of "3D convolution, instance normalization, LeakyReLU", and its role is to reduce the number of channels of the feature map R i and generate the recursive feature Z i . When i = 1, R i = Z i . This design allows the recursive modules to share weights and promotes cross-layer transmission of semantic information.

[0101] As Figure 7 shown, the deformation estimation module uses the deformable convolutional layer DeformConv and the difference layer Diff to generate the deformation field to ensure smooth and reliable transformation. Specifically, DeformConv generates a velocity field, and Diff enforces the diffeomorphic property. For time t, the deformation field is defined by the ordinary differential equation:

[0102]

[0103] where represents the identity transformation at t = 0, represents the target deformation field. Integrate from t = 0 to t = 1 in a scaling and squaring manner to obtain the deformation field In step five, a multi-scale and multi-stage encoder-decoder design based on the PSMA PET / CT Y channel and MRI is used to extract features:

[0104] Step 5.1: As Figure 4 shown, the main function of the encoder is feature extraction. In the present invention, ResNet (residual network) is used as the backbone network of the encoder, and a two-stream encoder strategy with non-shared weights is combined to independently encode the features from M y and F. The encoder consists of four convolutional blocks. Each layer of the residual module applies convolution, instance normalization, and Leaky ReLU to generate the feature map of the next stage. The first encoding layer does not perform downsampling, and the subsequent three layers use convolution with a stride of 2 for downsampling, and then are processed through the residual module of each layer.

[0105] The structure of the decoder is a recursive pyramid with four decoding layers. The task of each decoding layer is to sequentially predict the deformation field The lower decoding layer regresses large deformations at a low scale, and the higher decoding layer regresses small errors at a large scale. Such a structure is beneficial to solving the registration problem of complex deformations and improving the registration accuracy. Specifically, first, the fixed feature F i , the distorted motion feature and the feature z i are fused, and a fused feature map is formed through semantic convolution modulation. Then, the fused feature map is input into the deformation estimation module to predict the deformation field Finally, based on the convolutional long short-term memory network (U-CLSTM) module of the U channel, the estimated deformation field is combined with the information of M u to generate an updated deformation field. This decoding process of multi-scale feature fusion can capture complex deformations, while transmitting the global information of the gland and the local tumor information, ensuring that the key details of the lesion area are retained while taking into account the overall registration effect.

[0106] In step six of the present invention, the fused result is obtained by weighting the PSMA PET / CT and the enhanced MRI based on the deformation field generated by the registration network model.

[0107] Finally, the registration deformation field generated in the last stage of the registration model is applied to the PSMA PET / CT image, and it is image-weighted with the enhanced MRI to generate the final fused result.

[0108] In summary, the method inputs PSMA PET / CT and contrast-enhanced MRI into the registration network model. The model outputs the deformation field from the PSMA PET / CT position to the contrast-enhanced MRI position. The method further outputs the fused image of the two by weighting the shifted PSMA PET / CT and contrast-enhanced MRI. The final registration result is shown in Figure 8 As shown, it can be seen that compared with the large spatial misalignment before registration, the deep learning architecture designed by us can achieve the spatial alignment of the prostate gland and the tumor, thus realizing the non-rigid registration of PSMA PET / CT and contrast-enhanced MRI, greatly improving the diagnostic accuracy of clinicians for prostate cancer.

Claims

1. A method for multi-modal image registration and fusion of prostate PSMAPET / CT and enhanced MRI, characterized in that, Including the following steps: S1: Obtain PSMA PET / CT and enhanced MRI data, and convert the RGB format of the PSMA PET / CT into the YUV format to generate a Y-channel grayscale map, a U-channel feature map, and a V-channel feature map; S2: Based on the V-channel feature map, design a semantic gated convolution module (SGC) to enhance the global prostate contour perception ability of the registration network model; S3: Based on the U-channel feature map, design a convolutional long short-term memory module (U-CLSTM) to improve the local tumor region attention of the registration network model; S4: Recursively fuse multi-scale features based on a pyramid structure, combine the semantic gated convolution module and the convolutional long short-term memory module, and estimate the registration deformation field layer by layer; S5: Based on the Y-channel grayscale map and enhanced MRI, design a multi-scale multi-stage registration network model; S6: Generate a final fused image based on the deformed field weighted registration of the PSMA PET / CT and enhanced MRI generated by the registration network model.

2. A prostate PSMA PET / CT and enhanced MRI multimodal image registration and fusion method according to claim 1, characterized in that Specifically included in step 1: After obtaining the original data of PSMA PET / CT and enhanced MRI, perform image format conversion, image size normalization, image enhancement, and dataset partitioning processing; Convert the RGB channels of the PSMA PET / CT into YUV channels and generate feature maps. Among them, the Y channel retains grayscale information to generate a Y component feature map, and the U channel and V channel respectively generate a U component feature map sensitive to local lesions and a V component feature map sensitive to the overall prostate region through threshold segmentation.

3. A method for prostate PSMA PET / CT and enhanced MRI multimodal image registration and fusion according to claim 1, characterized in that The semantic gated convolution module described in step 2 includes: A semantic encoding module for extracting global semantic information containing prostate glands from the V component feature map and reducing the spatial resolution through a pooling layer to output in a feature form suitable for subsequent processing; A channel interaction module for projecting the feature map output by the semantic encoding module into a space matching the convolutional kernel dimension, realizing the interaction and conversion of information between different channels through a grouped linear layer, and finally adjusting the dimension of the feature representation to adapt to the convolutional kernel dimension requirements; A gate decoding module for dynamically modulating the convolutional kernel weights according to the outputs of the semantic encoding module and the channel interaction module to enhance the global feature expression.

4. A prostate PSMA PET / CT and enhanced MRI multimodal image registration and fusion method according to claim 1, characterized in that, The convolutional long short-term memory network module in step 3 includes: Input gate, which controls the weight information obtained from the deformation field. If the input gate is activated, the weight information is accumulated in the storage unit to introduce new data; ​ A forgetting gate that determines the degree of retaining memory based on the U-channel weights. If the forgetting gate is opened, the unit states with low weights are forgotten during this process, and important key information is retained; An output gate that determines whether to pass all memory units to the subsequent part or retain the information of the memory units without updating.

5. A method for prostate PSMA PET / CT and enhanced MRI multimodal image registration and fusion according to claim 1, characterized in that The deformation field estimation based on the pyramid structure in step 4 includes the following steps: Step 4.1: Design a recursive operation to combine the fixed features, the registered moving features, and the fused features. The combined features are used to extract global features through semantic convolution; Step 4.2: Design a deformation estimation module to predict the deformation field; Step 4.3: Input the deformation field and U-channel information into the U-CLSTM module to output the deformation field with enhanced local attention; Step 4.4: Re - perform feature mapping on the features, upsample the feature map and pass it to the next recursion; Step 4.5: Set a recursive module based on the pyramid structure, and set different numbers of recursive modules for each layer structure.

6. A prostate PSMA PET / CT and enhanced MRI multimodal image registration and fusion method according to claim 1, characterized in that In the multi - scale multi - stage registration network model in step 5, it includes: A two - stream encoder, using ResNet and adopting a two - stream encoder strategy with non - shared weights to separately process the feature extraction of the Y - channel image from PSMA PET / CT and the enhanced MRI image; Recursive pyramid decoder for generating a deformation field from the extracted features It consists of four decoding layers, and each decoding layer includes a semantic convolution module, a convolutional long short-term memory module based on the U-channel, and a deformation estimation module. Combine the encoder and decoder structures to form a registration network model, and input the enhanced MRI, the Y - channel, U - channel, and V - channel of PSMA PET / CT into the registration network model.

7. A method for prostate PSMA PET / CT and enhanced MRI multimodal image registration and fusion according to claim 1, characterized in that, The velocity - displacement field estimation of the pyramid structure in step 6 includes the following steps: Apply the deformation field generated in the last stage of the registration network model to the PSMA PET / CT image, and perform image weighting on it with the enhanced MRI to generate the final fusion result.

8. A method for prostate PSMA PET / CT and enhanced MRI multimodal image registration and fusion, characterized in that, This system includes: A data pre - processing module, used for image format conversion, image size normalization, image enhancement, and dataset division of PSMA PET / CT and enhanced MRI. For the processed PSMA PET / CT image, use the conversion formula to convert the RGB - channel image into the YUV channel, perform threshold segmentation on the U and V channel components to generate a U - component feature map sensitive to tumor region information and a V - component feature map sensitive to the prostate region; A semantic - gated convolution module, design a semantic - gated convolution module (SGC), and respectively construct a semantic encoding module, a channel interaction module, and a gate decoding module; A convolutional long - short - term memory module based on the U - channel, design a convolutional long - short - term memory module based on the U - channel (U - CLSTM). The module inputs the U - channel of the PSMA PET / CT image and the multi - level deformation field, and designs three control gates: an input gate, a forget gate, and an output gate to generate a local enhanced deformation field; A multi - scale multi - stage encoder - decoder module, based on the Y - channel of PSMA PET / CT and the enhanced MRI, constructs a multi - scale multi - stage two - stream encoder and a recursive decoder; A deformation estimation module, recursively combines the SGC module, the U - CLSTM module, and the deformation estimation module based on the traditional pyramid structure to fuse multi - level features. After recursion, upsample the features and the deformation field and input them to the next layer; An image weighting module, deform the PSMA PET / CT based on the deformation field in the last stage, weight the deformed image with the enhanced MRI, and output the fusion result.

Citation Information

Cited By

  • Photo-acoustic-magnetic multi-mode imaging system and method for prostate cancer recognition

    CN120852942A

  • Particle occlusion discrimination and PSD measurement method and system under decoupling neural network

    CN121384732A