Deformable Registration Method and System for Pelvic Images Based on Multi-Scale Dynamic Convolution Fusion
By adopting a multi-scale dynamic convolutional fusion method in the pelvic medical image registration technology, multi-scale features are extracted and space-channel information is fused, and the problems of time-consuming and limited registration accuracy in the existing technology are solved, and the rapid and accurate registration of pelvic CT images is achieved, providing better preconditions for automatic planning of pelvic screws.
Patent Information
- Application Number
- CN202510191748.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The existing pelvic medical image registration technology has problems such as time-consuming, limited registration accuracy and difficulty in capturing large-scale and highly nonlinear deformation characteristics, which affects the real-time and accuracy of automatic planning of pelvic screws.
The deformable registration method of pelvic image based on multi-scale dynamic convolutional fusion is adopted. Multi-scale features are extracted through the DKF-Block module, and the space-channel information fusion is used to achieve rapid and accurate registration of pelvic CT images.
The registration speed and accuracy of pelvic CT images were improved, and the limitations of traditional methods in receptive field range and large displacement deformation feature capture were overcome, providing a solid foundation for subsequent automatic pelvic screw channel planning.
Smart Images

Figure CN119693429B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image registration, and in particular to a deformable registration method and system for pelvic images based on multi-scale dynamic convolution fusion. Background Art
[0002] Pelvic fractures are a serious orthopedic trauma, usually caused by high-energy traumas (such as traffic accidents, falls from heights, or industrial accidents), accounting for about 3% to 8% of all fractures. Such fractures often involve multiple regions of the pelvis, especially the anterior pelvic ring and the posterior bony ring, resulting in vertical instability of the pelvis. The treatment methods for pelvic fractures are divided into two categories according to stability: stable fractures are generally treated conservatively, while unstable fractures require surgical intervention. Surgical treatment usually includes internal plate fixation and percutaneous screw internal fixation. Among them, percutaneous screw internal fixation is widely used in clinical treatment due to its advantages of minimally invasive, rapid recovery, and low surgical cost.
[0003] Percutaneous screw internal fixation surgery for the pelvis usually relies on the experience and manual operation of surgeons. This process is not only time-consuming and laborious, but also lacks standardization, resulting in possible differences in the planning results of different doctors. With the continuous development of medical computer software, methods for pelvic screw channel planning have been widely studied. Existing studies have proposed methods for screw channel planning using medical imaging software. These methods usually require manual extraction and adjustment of the pelvic model, and the screw positions are planned by manually operating the software, which often leads to strong subjectivity in the planning results and it is difficult to obtain the best effect. Recently, some studies have developed pelvic screw channel search algorithms to assist in screw channel planning using computer programs. Although these computer-aided methods effectively reduce subjectivity, they still require a large amount of manual annotation to run the program, thus consuming a large amount of time, which is extremely disadvantageous for emergency surgeries.
[0004] In recent years, significant progress has been made in the research on automatic screw planning using registration technology. By using the Free Form registration algorithm to analyze 523 sub-acetabular channels of the pelvis, an average morphological model of the pelvis was constructed through the registration algorithm. They manually marked the screw entry and exit areas, mapped these areas to the rest of the pelvis, and then obtained the pelvic screw path through full traversal calculation. A screw automatic planning method based on multiple templates was also proposed. This method manually planned the screw trajectories of 40 pelvic templates and used the active shape model (ASM) registration method to map the template paths to the patient's pelvis. It planned a path suitable for a specific patient by optimizing the cumulative function.
[0005] After investigation and analysis, there are currently three problems with the current pelvic medical image registration technology:
[0006] (1) The existing automatic pelvic screw planning technology based on image registration usually relies on an iterative function based on energy optimization to achieve free-form registration. Its main drawback is that the optimization parameters need to be recalibrated every time registration is performed, which is time-consuming and limits the real-time performance of the automatic pelvic screw planning technology.
[0007] (2) Currently, free-form registration methods often rely on a general mathematical model for global deformation when processing pelvic images, and it is difficult to fully characterize the anatomical differences between individual pelves. Since the specific features of each pelvis cannot be effectively identified, the registration accuracy of such methods is susceptible to influence when processing data of different patients, thereby reducing the accuracy and reliability of automatic pelvic screw planning.
[0008] (3) Deep learning registration algorithms, with their learnability and real-time inference speed, have to some extent made up for the deficiencies of free-form registration. However, for specific anatomical structures of the human body such as the pelvis, there are still many challenges in its practical application. Since pelvic registration often involves large-scale and highly non-linear deformations, existing networks are difficult to fully capture and characterize long-displacement feature information, resulting in the registration accuracy and robustness being unable to meet clinical requirements. Moreover, the deformation correlation information of the spatial and channel information of the image is not fully mined, and this kind of information is particularly important for the registration technology. Summary of the Invention
[0009] The purpose of the present invention is to provide a deformable registration method and system for pelvic images based on multi-scale dynamic convolution fusion, accurately register pelvic CT images, and provide a prerequisite for subsequent automatic planning of pelvic screw channels.
[0010] To solve the above technical problems, the present invention provides a deformable registration method for pelvic images based on multi-scale dynamic convolution fusion, including the following steps:
[0011] Obtain pelvic CT images;
[0012] Input the pelvic CT images into the DKF-Block module for feature extraction to obtain large-scale features of the pelvic images;
[0013] Input the large-scale features of the pelvic images into the SCC-Block module for fusion to obtain a fusion result.
[0014] Preferably, inputting the pelvic CT images into the DKF-Block module for feature extraction to obtain large-scale features of the pelvic images specifically includes the following steps:
[0015] Downsample the features of the pelvic CT images to obtain the input features of the DKF-Block ;
[0016] The input features of the DKF-Block After passing through a standard 3×3×3 convolutional layer, a Leaky ReLu activation layer, and an instance normalization layer, intermediate features are obtained. ;
[0017] Intermediate features Enter the DKF-Layer to extract multi-scale features and output the fused features. ;
[0018] Fused features Pass through a standard 3×3×3 convolutional layer and are element-wise added to the original features of the residual branch to obtain the output features. , serving as the large-scale features of the pelvic image.
[0019] Preferably, the intermediate features Enter the DKF-Layer to extract multi-scale features and output the fused features. , specifically including the following steps:
[0020] The DKF-Layer includes a high-resolution branch and a low-resolution branch;
[0021] The high-resolution branch extracts image detail features through a 3×3×3 depthwise separable convolution. ;
[0022] After downsampling, the low-resolution branch respectively uses 5×5×5 and 7×7×7 large-kernel depthwise separable convolutions to capture context information with a larger receptive field in parallel, obtaining the first parallel feature and the first parallel feature ;
[0023] The first parallel feature and the first parallel feature are concatenated along the channel dimension to obtain the concatenated feature;
[0024] The global maximum pooling layer is used to perform global feature aggregation on the concatenated feature, and a dynamic weight is generated through the Sigmoid activation function and ;
[0025] The importance of and is calibrated and to dynamically select and fuse the features obtained from different large-kernel convolutions, obtaining the fused low-resolution features;
[0026] The fused low-resolution features are restored to high resolution through upsampling and added element-wise to obtain the fused features. ;
[0027] 。
[0028] Preferably, the large-scale features of the pelvic image are input into the SCC-Block module for fusion to obtain a fusion result, which specifically includes the following steps:
[0029] The large-scale features of the pelvic image are used as the input features of the lower-level encoder and the features of the same-level encoder. First, the features of the lower-level decoder are upsampled as and then, together with the features of the same-level encoder, are used as for concat operation to obtain the input features of the SCC-Block ;
[0030] The input features of the SCC-Block pass through the core component SCC-Layer to extract the image spatial-channel information and fuse the feature maps, obtaining the output features of the residual branch ;
[0031] The output features of the residual branch successively pass through a standard convolutional layer with a convolutional kernel size of 3, a LeakyReLu activation layer, an instance normalization layer, and another standard convolutional layer with a convolutional kernel size of 3, and are added element-wise to the output features of the residual branch to obtain the output features of the SCC-Block , which are used as the fusion result.
[0032] Preferably, the input features of the SCC-Block pass through the core component SCC-Layer to extract the image spatial-channel information and fuse the feature maps, obtaining the output features of the residual branch , which specifically includes the following steps:
[0033] The SCC-Layer uses the cross-attention mechanism to process the input features of the SCC-Block along the spatial attention branch and the channel attention branch respectively, dynamically generating the spatial attention weight and the channel attention weight ; The cross-attention obtains the output features of the residual branch by introducing the spatial weight into the calculation of the channel weight and integrating the channel weight into the generation of the spatial weight.
[0034] Preferably, the SCC-Layer uses the cross-attention mechanism to process the input features of the SCC-Block along the spatial attention branch and the channel attention branch respectively, dynamically generating the spatial attention weight and channel attention weights ; Cross-attention incorporates spatial weights into the calculation of channel weights and integrates channel weights into the generation of spatial weights, which specifically includes the following steps:
[0035] The channel attention weight branch performs a standard 1×1×1 convolution and element-wise addition on the input feature 、 respectively, then multiplies element-wise with the with reduced number of channels, and uses the SoftMax normalization function to obtain the channel result ;
[0036] The spatial attention weight branch uses a global max pooling layer to highlight the spatial features of the input feature , uses a depthwise separable convolution with a kernel size of 3×3×3 to further extract features, and multiplies element-wise with the with expanded number of channels, and obtains the spatial result through the SoftMax normalization function;
[0037] In the SCC-Layer, the input feature x of the SCC-Block is successively fused with the spatial result 、the channel result to obtain the output feature of the residual branch.
[0038] Preferably, the loss functions of the DKF-Block module and the SCC-Block module include: The first part is the similarity calculation between the deformed moving image and the fixed image; The second part is the spatial regularization of the predicted deformation field to make it as smooth as possible and avoid image folding to the greatest extent;
[0039] ;
[0040] In the formula: represents the image similarity loss function; represents the deformation field regularization; λ is the regularization hyperparameter; represents the moving image; represents the fixed image; represents the regularization term.
[0041] Preferably, the similarity calculation between the deformed moving image and the fixed image specifically includes the following steps:
[0042] Use local normalized cross-correlation as the image similarity loss function ;
[0043] ;
[0044] ;
[0045] In the formula: Ω represents the set of all points in the three-dimensional voxel space; and represents the average voxel value within a local window centered on voxel P with a size of 9×9×9.
[0046] Preferably, spatial regularization is performed on the predicted deformation field, which specifically includes the following steps:
[0047] Use a regularization penalty term to implement a smooth regularization term :
[0048] ;
[0049] In the formula: represents the gradient value of voxel p in the three-dimensional registration space. The greater the gradient, the greater the possibility of non-smoothness at point P.
[0050] The present invention also provides a deformable registration system for pelvic images based on multi-scale dynamic convolution fusion, including:
[0051] An acquisition module for acquiring pelvic CT images;
[0052] A feature extraction module for inputting the pelvic CT image into the DKF-Block module for feature extraction to obtain large-range features of the pelvic image;
[0053] A fusion module for inputting the large-range features of the pelvic image into the SCC-Block module for fusion to obtain a fusion result.
[0054] Compared with the prior art, the beneficial effects of the present invention are:
[0055] The deformable registration algorithm for pelvic images based on multi-scale dynamic convolution fusion proposed by the present invention improves the traditional deep learning registration model, designs a multi-scale dynamic convolution fusion module and a spatial channel coupling module, and realizes fast and accurate registration of pelvic CT images. While effectively solving the problems of long time consumption and limited registration accuracy existing in the traditional free-form registration algorithm, it also overcomes the limitations of the deep learning registration network in the receptive field range and the capture of large displacement deformation features. By capturing and fusing feature information at different scales, this method can better perceive and represent large-range deformations, thus laying a solid foundation for the subsequent automatic planning of pelvic screw channels. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings.
[0057] Figure 1 It is the architecture diagram of DMSC-Reg;
[0058] Figure 2 It is the internal architecture of DKF-Block and DKF-Layer;
[0059] Figure 3 It is the internal architecture of SCC-Block and SCC-Layer. Detailed implementation manners
[0060] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific implementations disclosed below.
[0061] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0062] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0063] The following further describes the present invention in detail with reference to the accompanying drawings:
[0064] In order to better illustrate the technical effects of the present invention, the present invention provides the following specific embodiments to illustrate the above technical processes:
[0065] Embodiment 1. A deformable registration algorithm for pelvic images based on multi-scale dynamic convolution fusion (Dynamic Multi-Scale Convolutional Registration Network, DMSC-Reg), and its improvement points are as follows:
[0066] (1) To solve the problem of the relatively distant correspondence of pelvic CT image features, the encoder part of DMSC-Reg is designed based on multi-scale information fusion. The dynamic convolution fusion module (DynamicKernel Fusion Block, DKF-Block) designed in this chapter is used in the upper half of the encoder to extract large-scale features of pelvic images in the high-resolution dimension, which can retain more detailed information on the association between pelvic images. And the large kernel depthwise separable convolution, downsampling, upsampling, and residual connection processes are introduced. These designs not only deepen the overall structure of the algorithm but also control the computational complexity and memory consumption of the algorithm, minimizing the number of parameters of the registration model as much as possible and facilitating the deployment of the algorithm.
[0067] (2) To solve the problem of insufficient extraction of image spatial and channel feature information, the decoder of DMSC-Reg is designed based on the concept of spatial-channel feature fusion, and the spatial-channel coupling module (Spatial-Channel Coupling Block, SCC-Block) proposed in this chapter is introduced in the decoder layer. This module can fully fuse the spatial and channel information of the decoder features and the upsampled image features, effectively enhancing the deformation and displacement perception ability between pelvic images, thereby further improving the feature expression and decoding performance.
[0068] This invention explores how to accurately register pelvic CT images, providing a prerequisite for the subsequent automatic planning of pelvic screw channels. For this purpose, a deformable registration algorithm for pelvic images based on multi-scale dynamic convolution fusion is proposed, and experiments are set up for detailed analysis to verify the feasibility of the proposed method. The model architecture diagram of this invention is as Figure 1 shown.
[0069] (1) Description of the multi-scale dynamic convolution fusion module (DKF-Block), as Figure 2 shown;
[0070] Due to the large deviation in the shape of the human pelvis, traditional deep learning registration networks are difficult to effectively capture long-range dependencies. Therefore, a broader receptive field is required for feature extraction of pelvic CT images. To solve this problem, this invention proposes the DKF-Block module, Figure 2 which shows the internal architecture of the DKF-Block module. Generally speaking, the DKF-Block module is a residual connection structure that effectively transmits feature information through a deeper network depth. Its core component is the DKF-Layer, which is specifically used to extract long-range feature association information in images.
[0071] Down Sampling consists of two sets of convolutional structures (which can reduce the feature size and expand the dimension). The DKF-Block itself does not change the size and dimension of the feature dimension. Therefore, the change in the feature dimension occurs in Down Sampling, and the feature at this time is the input feature of the DKF-Block. The features of the pelvic CT images will be downsampled first and then input into the DKF-Block.
[0072] Given the input feature of the DKF-Block , first, it passes through a standard convolutional layer of 3×3×3, a Leaky ReLu activation layer (LR), and an instance normalization layer (IN) to obtain an intermediate feature representation. .
[0073] Next, it enters the DKF-Layer, extracts multi-scale features, and outputs the fused features. .
[0074] Finally, it passes through another standard convolutional layer of 3×3×3 and adds the original feature of the residual branch (the input feature of the DKF-Block) element-wise to obtain the output feature. . The input-output feature dimension of the DKF-Block remains unchanged, and the left-branch feature dimension is the same as the right-branch feature dimension.
[0075] The feature calculation formula of the DKF-Block is as follows:
[0076] ;
[0077] ;
[0078] ;
[0079] The DKF-Layer is the key to enabling the DKF-Block to have the ability to extract multi-scale features of images. Its core idea is to extract multi-scale features of images through large-kernel convolutions of different sizes, significantly enhancing the network receptive field. To reduce the number of parameters and computational complexity of the large-kernel convolution, the DKF-Layer uses depthwise separable convolutions to replace the original standard convolutions. Specifically, the DKF-Layer includes two feature branches with different resolutions. The high-resolution branch extracts the detailed features of the image through a 3×3×3 depthwise separable convolution. .
[0080] ;
[0081] After downsampling, the low-resolution branch uses large-kernel depthwise separable convolutions of 5×5×5 and 7×7×7 in parallel to capture context information with a larger receptive field, obtaining and .
[0082] ;
[0083] ;
[0084] Subsequently, these two parallel features are concatenated (Concat) along the channel dimension. Then, a global max pooling layer (MAP) is used to perform global feature aggregation on the concatenated features, and dynamic weights are generated through the Sigmoid activation function and .
[0085] ;
[0086] The weights , can be adaptively adjusted according to the context information of the input features. By calibrating the importance of and , the dynamic selection and fusion of features from different large-kernel convolutions are realized. This mechanism further enhances the ability of multi-scale feature representation, enabling the network to more effectively model long-range dependencies and complex structural information in pelvic CT images. Finally, the fused low-resolution features are restored to high resolution through upsampling (UP) and added element-wise to obtain the final output features . The overall feature calculation formula of the DKF-Layer is as follows:
[0087] ;
[0088] (2) Description of the spatial-channel coupling module (SCC-Block), as shown in Figure 2 ;
[0089] To better adaptively fuse spatial-channel information into multi-scale local feature maps, this paper proposes the SCC-Block Figure 3 shows the internal architecture of the SCC-Block module. Similar to the DKF-Block, the SCC-Block is also a residual connection structure. The difference is that the input of the SCC-Block is the feature obtained by concatenating the same-level encoder feature and the upsampled feature of the lower-level decoder along the channel; the output feature of the DKF-Block There are two directions. One is as the input feature of the lower-level encoder, and the other is the feature of the same-level encoder, which is one of the input features of the SCC-Block. UpSampling consists of two sets of convolutional layers. UpSampling will expand the feature size and reduce the dimension, making it consistent with the dimension of the encoder feature, so that the concat operation can be performed. (For example, if the dimension of the same-level encoder feature is ×D×C, and the feature of the lower-level decoder is / 2×D / 2×2C, its feature dimension becomes ×D×C) after upsampling. First, perform upsampling on the feature of the lower-level decoder, and then perform the concat operation with the feature of the same-level encoder (the blue block on the left in the following figure).
[0090] Given the input feature .
[0091] ;
[0092] First pass through the core component SCC-Layer to extract the image spatial-channel information and fuse the feature maps to obtain the output result . Then sequentially pass through a standard convolutional layer with a kernel size of 3, a LeakyReLu activation layer (LR), an instance normalization layer (IN), and another standard convolutional layer with a kernel size of 3, and add them element-wise to the output feature of the residual branch to obtain the output feature of the SCC-Block. The feature calculation formula of the SCC-Block is as follows:
[0093] ;
[0094] ;
[0095] The core of the SCC-Block lies in adaptively extracting and fusing the spatial information and channel information in the input features through the SCC-Layer. Specifically, the SCC-Layer processes the input features along two branches using the cross-attention mechanism to dynamically generate the spatial attention weight and the channel attention weight . In this process, the cross-attention realizes the deep interaction and optimization of the features in the spatial dimension and channel dimension by introducing the spatial weight into the calculation of the channel weight and integrating the channel weight into the generation of the spatial weight. Finally, this fusion mechanism outputs the optimized feature representation . Through the design of the cross-attention mechanism, the model effectively constructs the deep association of multi-dimensional features and effectively enhances the coupling of spatial and channel information.
[0096] Specifically, for the channel attention weight branch, first, the input features and are respectively subjected to a standard 1×1×1 convolution and element-wise addition, and then multiplied element-wise with with the number of channels reduced, and the channel result is obtained using the SoftMax normalization function.
[0097] ;
[0098] For the spatial attention weight branch, it is obtained by processing x. First, the global max pooling layer (MAP) is used to highlight the spatial features, and the depthwise separable convolution with a convolution kernel size of 3×3×3 is used to further extract features. The subsequent operations are similar to those of the channel weight branch, multiplied element-wise with with the number of channels expanded, and finally the spatial result is obtained through the SoftMax normalization function.
[0099] ;
[0100] In the SCC-Layer, the input feature x is successively fused with and to obtain the output feature of the residual branch. The overall feature calculation formula of the SCC-Layer is as follows:
[0101] ;
[0102] (3) Loss function;
[0103] The overall loss function of the DMSC-Reg network is derived from the energy function of traditional image registration algorithms. The loss function consists of two parts: the first part is the similarity calculation between the deformed moving image and the fixed image, and the second part is the spatial regularization of the predicted deformation field to make it as smooth as possible and avoid image folding to the greatest extent:
[0104] ;
[0105] where represents the image similarity loss function, represents the deformation field regularization, and λ is the regularization hyperparameter.
[0106] The local normalized cross-correlation (LNCC) is used as the image similarity loss function . For the moving image and the fixed image For them, if they are highly similar in the local window, the LNCC value approaches 1; when the gray-scale distributions of the two images are irrelevant or weakly correlated in this window, the LNCC value is close to 0; if there is an inverse gray-scale distribution (strong negative correlation) in the local window, the LNCC value will be close to -1. Since a higher LNCC value represents more similar images, in order to convert the "maximizing similarity" into an optimization goal of "minimizing loss", a negative sign is added to the LNCC during implementation, that is , so as to use it as a loss function. Specifically, the formula of LNCC( , ) is as follows:
[0107] ;
[0108] where Ω represents the set of all points in the three-dimensional voxel space, and represent the average voxel values within a local window centered on voxel P with a size of 9×9×9.
[0109] Minimizing the image similarity loss function Promote the fixed image to approximate the moving image . Therefore, a regularization penalty term is used to implement the smooth regularization term :
[0110] ;
[0111] where represents the gradient value of voxel p in the three-dimensional registration space. The greater the gradient, the greater the possibility of non-smoothness at point P. Therefore, the complete loss function of the DMSC-Reg network is:
[0112] ;
[0113] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules, modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units, modules or components can be combined or integrated into another device, or some features can be ignored or not executed.
[0114] The unit(s) may or may not be physically separated. The component(s) shown as a unit can be a single physical unit or multiple physical units, i.e., they can be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0115] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0116] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), the above-mentioned functions defined in the methods of the present invention are performed. It should be noted that the above-mentioned computer-readable medium in the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above.
[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0118] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims described above.
Claims
1. A deformable pelvic image registration method based on multi-scale dynamic convolution fusion, characterized in that: The following steps are involved: Obtain pelvic CT images; Downsample the pelvic CT image features to obtain the input feature f of DKF-Block; The input feature f of DKF-Block is obtained after the 3×3×3 standard convolution layer, LeakyReLu activation layer, and instance regularization layer to obtain the intermediate feature f in ; DKF-Layer includes high-resolution branch and low-resolution branch; The high-resolution branch extracts image detail features f0 through 3×3×3 depth separation convolution; After downsampling, the low-resolution branch uses 5×5×5 and 7×7×7 large kernel depthwise separable convolutions to capture contextual information with a larger receptive field in parallel, respectively, to obtain the first parallel feature f1 and the first parallel feature f2; Splicing the first parallel feature f1 and the first parallel feature f2 along the channel dimension to obtain a spliced feature; The global maximum pooling layer is used to aggregate the concatenated features globally, and the dynamic weights w1 and w2 are generated through the Sigmoid activation function; The importance of f1 and f2 is calibrated through dynamic weights w1 and w2, so as to dynamically select and fuse the features obtained from different large kernel convolutions to obtain the fused low-resolution features; The fused low-resolution features are restored to high resolution by upsampling and then added element by element to obtain the fused feature f out ; The fused feature f out After the 3×3×3 standard convolution layer and the original features of the residual branch are added element by element, the output features are obtained As a large-scale feature of pelvic images; The large-scale features of the pelvic image As the input feature of the lower encoder and the feature of the same-level encoder, the lower decoder feature is first upsampled to obtain the upsampled feature x2 of the lower decoder, and then concat operation is performed with the feature x1 of the same-level encoder to obtain the input feature x of the SCC-Block; The input feature x of SCC-Block is extracted through the core component SCC-Layer to extract the image space-channel information and fuse the feature map to obtain the output feature x of the residual branch. out ; The output feature x of the residual branch out It passes through a standard convolution layer with a convolution kernel size of 3, a LeakyReLu activation layer, an instance regularization layer, and another standard convolution layer with a convolution kernel size of 3, and is combined with the output feature x of the residual branch. out Add element by element to get the output features of SCC-Block As a result of fusion.
2. The deformable pelvic image registration method based on multi-scale dynamic convolution fusion according to claim 1, characterized in that: The input feature x of SCC-Block is extracted through the core component SCC-Layer to extract the image space-channel information and fuse the feature map to obtain the output feature x of the residual branch. out , specifically including the following steps: SCC-Layer uses the cross attention mechanism to process the input feature x of SCC-Block separately along the spatial attention branch and the channel attention branch, and dynamically generates the spatial attention weight w sp and channel attention weight w ch ; Cross attention introduces spatial weights into the calculation of channel weights, and integrates channel weights into the generation of spatial weights to obtain the output feature x of the residual branch out .
3. The deformable registration method for pelvic images based on multi-scale dynamic convolution fusion according to claim 2, characterized in that: SCC-Layer uses the cross attention mechanism to process the input feature x of SCC-Block separately along the spatial attention branch and the channel attention branch, and dynamically generates the spatial attention weight w sp and channel attention weight w ch ; Cross attention introduces spatial weights into the calculation of channel weights, and integrates channel weights into the generation of spatial weights, specifically including the following steps: The channel attention weight branch performs a 1×1×1 standard convolution and element-wise addition on the input features x1 and x2, and then adds them to the reduced channel number w sp Perform element-wise multiplication and use the SoftMax normalization function to get the channel result w ch ; The spatial attention weight branch uses a global maximum pooling layer to highlight the spatial features of the input feature x, and uses a depthwise separable convolution with a kernel size of 3×3×3 to further extract features, and then adds the dilated channel number w ch Perform element-wise multiplication and obtain the spatial result w through the SoftMax normalization function sp ; In SCC-Layer, the input feature x of SCC-Block is sequentially compared with the spatial result w sp 、Channel result w ch After fusion, the output feature x of the residual branch is obtained out .
4. The deformable registration method for pelvic images based on multi-scale dynamic convolution fusion according to claim 3, characterized in that: The loss functions of the DKF-Block module and the SCC-Block module include: the first part is the similarity calculation between the deformed moving image and the fixed image; the second part is to perform spatial regularization on the predicted deformation field to make it as smooth as possible and to avoid image folding to the greatest extent; Where: represents the image similarity loss function; represents the deformation field regularization; λ is the regularization hyperparameter; I f Represents a moving image; I m represents a fixed image; φ represents a regularization term.
5. The deformable registration method for pelvic images based on multi-scale dynamic convolution fusion according to claim 4, characterized in that: The similarity calculation between the deformed moving image and the fixed image specifically includes the following steps: Using local normalized cross-correlation as image similarity loss function Where: Ω represents the set of all points in the three-dimensional voxel space; and It represents the average voxel value in a local window of size 9×9×9 centered on voxel P.
6. The deformable registration method for pelvic images based on multi-scale dynamic convolution fusion according to claim 5, characterized in that: The predicted deformation field is spatially regularized, which includes the following steps: The regularization penalty term is used to achieve a smooth regularization term φ: Where: It represents the gradient value of voxel p in the 3D registration space. The larger the gradient, the greater the possibility that φ is not smooth at point P.
7. A deformable registration system for pelvic images based on multi-scale dynamic convolution fusion, used to implement the deformable registration method for pelvic images based on multi-scale dynamic convolution fusion as described in any one of claims 1 to 6, characterized in that: include: An acquisition module, used for acquiring a pelvic CT image; A feature extraction module is used to input the pelvic CT image into the DKF-Block module for feature extraction to obtain large-scale features of the pelvic image; The fusion module is used to input the large-scale features of the pelvic image into the SCC-Block module for fusion to obtain the fusion result.
Citation Information
Patent Citations
Pelvis automatic segmentation method and system, electronic equipment and storage medium
CN118411370A