Ablation temperature field prediction method and system with full-scale recalibration and spatial compensation

By employing a full-scale recalibration and spatial compensation method for predicting ablation temperature fields, and utilizing a dual-branch decoder network and physical adapter for feature decoupling and weighted fusion, the prediction error and coordinate offset problems in microwave ablation temperature field prediction are solved, achieving high-precision and efficient temperature field prediction and supporting real-time navigation for clinical microwave ablation surgery.

CN122369966APending Publication Date: 2026-07-10BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNIV OF TECH
Filing Date
2026-04-15
Publication Date
2026-07-10

Smart Images

  • Figure CN122369966A_ABST
    Figure CN122369966A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for predicting ablation temperature fields with full-scale recalibration and spatial compensation, relating to the interdisciplinary fields of biomedical engineering, medical imaging computing, and artificial intelligence. The method involves acquiring a three-dimensional tumor mask, a three-dimensional needle path mask, and an initial specific absorptivity field of the case to be predicted, and stitching them together to obtain a multi-channel input feature map. Feature extraction is performed on the multi-channel input feature map, and semantic feature extraction is performed on the obtained multi-scale encoded feature map to obtain a high-level semantic feature map. Parallel feature extraction is performed using dual-path branching to obtain a first feature map and a second feature map, respectively. Spatially adaptive weighted fusion processing is performed using a gated decision block to obtain a three-dimensional ablation temperature field prediction map. This invention solves the problems of insufficient prediction accuracy, loss of local high-temperature details, poor generalization across cases, and sub-voxel-level spatial positioning bias caused by multi-level sampling in existing microwave ablation temperature field prediction technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary fields of biomedical engineering, medical imaging computing and artificial intelligence, and in particular to a method and system for predicting ablation temperature fields with full-scale recalibration and spatial compensation. Background Technology

[0002] In the clinical treatment of solid tumors, microwave ablation, with its advantages of being minimally invasive and allowing for rapid recovery, has become a first-line radical treatment for small solid tumors (such as hepatocellular carcinoma) measuring 1-3 cm. The core objective of this procedure is to achieve conformal ablation, that is, to completely kill tumor cells while precisely controlling the ablation boundary to protect normal tissue. Therefore, real-time, high-precision prediction of the temperature field in the ablation zone is a crucial prerequisite for ensuring the safety and efficacy of the procedure.

[0003] Current techniques for solving ablation temperature fields are mainly divided into traditional numerical simulation methods, purely data-driven deep learning methods, and physics-guided deep learning methods, but all of them have significant shortcomings in practical applications: (a) Traditional numerical simulation methods (such as the finite element method): have low computational efficiency, and single-case simulation takes a very long time, which cannot meet the needs of real-time navigation during surgery; moreover, the simulation accuracy is highly dependent on the accurate assignment of tissue parameters, which is difficult to cope with the individual heterogeneity and dynamic changes of patient tissue parameters in clinical practice, resulting in large prediction errors.

[0004] (ii) Purely data-driven deep learning methods (such as 3D_UNet): lack physical constraints and have extremely low generalization; multi-level downsampling of traditional convolutional networks can lead to the smooth loss of details of local high temperature gradients near the needle tip, making the prediction of the critical boundary of coagulative necrosis that determines the success or failure of surgery ambiguous; at the same time, the integer step size in the multi-level upsampling process cannot accurately restore the asymmetric spatial size, which can easily produce sub-voxel level coordinate shifts. For small tumors, even slight coordinate shifts can lead to millimeter-level positional deviations.

[0005] (III) Existing physical-guided deep learning methods: Although physical prior information such as specific absorptivity (SAR) is introduced, static physical inputs often cannot adapt to the heterogeneity of different patient tissue parameters and the dynamic changes in the ablation process, resulting in a mismatch between physical priors and deep features; attention mechanisms are mostly deployed only in a single deep layer and cannot achieve full-scale feature recalibration; single-branch network architectures cannot simultaneously take into account the global temperature field distribution and local extremely high temperature gradient features, and model training is easily dominated by the low-temperature region, which accounts for a larger proportion; in addition, skip connections still do not effectively compensate for sub-voxel level coordinate offsets during feature fusion. Summary of the Invention

[0006] To address the aforementioned shortcomings in existing technologies, this invention provides a full-scale recalibration and spatial compensation method and system for predicting ablation temperature fields. This solution addresses the problems of insufficient prediction accuracy, loss of local high-temperature details, poor generalization across cases, and sub-voxel-level spatial positioning bias caused by multi-level sampling in microwave ablation temperature field prediction.

[0007] To achieve the aforementioned objectives, the technical solution adopted by this invention is: a method for predicting ablation temperature fields based on full-scale recalibration and spatial compensation, comprising: S1: Obtain the three-dimensional tumor mask, three-dimensional needle track mask and initial specific absorptivity field of the case to be predicted, and perform splicing to obtain a multi-channel input feature map; S2: Use an encoder network with multiple resolution levels to extract features from the multi-channel input feature map, and within each resolution level, use a three-dimensional channel attention module to perform full-scale channel recalibration on the features of the current level to obtain a multi-scale encoded feature map. S3: Use the bottleneck layer to perform semantic feature extraction on the deepest feature map in the multi-scale encoded feature map to obtain a high-level semantic feature map; S4: Upsample the high-level semantic feature map using a dual-branch decoder network, and perform feature extraction in parallel using dual-path branches to obtain the first feature map and the second feature map respectively. S5: The first feature map and the second feature map are spatially adaptively weighted and fused using a gating decision block to obtain a three-dimensional ablation temperature field prediction map, thus completing the ablation temperature field prediction.

[0008] The beneficial effects of this invention are as follows: This invention provides a method for predicting ablation temperature field with full-scale recalibration and spatial compensation. Through a dual-branch decoder network, feature extraction is performed by decoupling the fully symmetrical main temperature field branch and high-temperature gradient branch. Spatial adaptive weighted fusion is performed by gating decision blocks, which effectively takes into account both the global temperature field distribution and the extraction of local high-temperature gradient features. This solves the problem that traditional single-branch networks are prone to losing high-temperature details and greatly improves the prediction accuracy of the core ablation zone.

[0009] By introducing a physical adapter, the initial specific absorptivity field is dynamically calibrated using a 3D needle path mask as a spatial index, generating an adaptive specific absorptivity field. This eliminates the heterogeneity of tissue parameters among different cases, solves the problem of mismatch between static physical input and dynamic ablation features, and improves the model's generalization ability in different clinical scenarios.

[0010] A three-dimensional channel attention module is deployed in each resolution level of the encoder network to perform full-scale channel recalibration on the features of the current level, dynamically enhance the feature channels that are sensitive to temperature field evolution and suppress noise channels, thereby improving the feature expression efficiency and the prediction accuracy of critical necrosis boundaries.

[0011] In the dual-branch decoder network, spatial dimension alignment and position compensation are performed on the feature maps passed by the skip connections through trilinear interpolation. This forces alignment of the target size and achieves corner point coincidence, effectively eliminating sub-voxel level coordinate offset caused by multi-level sampling and ensuring high-precision positioning of the ablation heat source center.

[0012] The constructed end-to-end prediction model ensures high accuracy while possessing high inference efficiency, avoiding the extremely high time consumption of traditional numerical simulation methods, and can meet the real-time navigation needs in clinical microwave ablation procedures.

[0013] Further, S1 includes: Obtain the three-dimensional tumor mask, three-dimensional needle path mask, and initial specific absorbance field of the case to be predicted; The initial specific absorptivity field is cascaded with the three-dimensional needle mask to obtain cascaded features; The initial specific absorptivity field is dynamically calibrated using a physical adapter. The heat source center position marked by the three-dimensional needle track mask is used as the spatial index. The cascaded features are processed using three-dimensional instance normalization to eliminate the heterogeneity of tissue parameters among different cases and obtain an adaptive specific absorptivity field. The three-dimensional tumor mask, the three-dimensional needle path mask, and the adaptive specific absorptivity field are spliced ​​together to obtain a multi-channel input feature map.

[0014] This preferred solution dynamically calibrates the initial specific absorptivity field using a physical adapter. Using the heat source center marked by the three-dimensional needle path mask as a spatial index, combined with three-dimensional instance normalization processing, it specifically eliminates the individual heterogeneity of tissue electrical and thermal parameters among different cases. This solves the core problem of mismatch between static physical priors and dynamic tissue characteristics of patients in existing physical guidance methods, allowing physical prior information to adapt to individual differences among different patients. It significantly improves the model's generalization ability across cases and clinical scenarios, and avoids temperature field prediction system bias caused by individual differences among patients.

[0015] Further, S2 includes: The multi-channel input feature map is processed by an encoder network containing multiple resolution levels to obtain the current level features within each resolution level; The current layer features are processed using three-dimensional global average pooling, and the global spatial features of each channel are compressed into a channel statistic to obtain a channel descriptor. The channel descriptors are nonlinearly mapped using a series of fully connected layers to obtain the channel weight coefficients for each channel. The channel weight coefficients are multiplied with the current level features to obtain the recalibrated multi-scale encoded feature map.

[0016] This preferred scheme clarifies the specific implementation method of full-scale channel recalibration. It compresses the global spatial features of the channels through three-dimensional global average pooling, and combines this with continuous fully connected layers to complete nonlinear mapping to generate channel weights, ultimately achieving precise channel-level recalibration of features. This mechanism can achieve differentiated feature enhancement at different encoder levels: the shallow layer retains key geometric features of tumor boundaries and needle path locations; the middle layer enhances physical features related to heat conduction gradients; and the deep layer adapts to semantic features that accommodate tissue heterogeneity. This fundamentally avoids the weak and effective features of small tumors (1-3 cm) being submerged by background noise during multi-level downsampling, significantly improving the prediction accuracy of the critical boundary of tumor coagulative necrosis.

[0017] Further, S4 includes: S410: The high-level semantic feature map is upsampled using a dual-branch decoder network with multiple upsampling levels. In each upsampling level, spatial dimension alignment and position compensation are performed on the corresponding upsampling level feature map in the multi-scale encoded feature map using trilinear interpolation to obtain an aligned feature map. S420: After concatenating the aligned feature map with the current upsampled level feature map, feature extraction is performed in parallel using dual-path branches to obtain the first feature map and the second feature map respectively.

[0018] This preferred scheme first performs spatial dimension alignment and position compensation on the feature maps of the encoder's skip connections at each upsampling level of the decoder using trilinear interpolation. Then, it concatenates these features with the upsampling features of the current level and sends them to a dual-path branch for parallel feature extraction. On one hand, spatial alignment eliminates spatial coordinate deviations caused by multi-level upsampling and downsampling in advance, avoiding spatial misalignment during high- and low-resolution feature fusion. On the other hand, the dual-path branch enables parallel decoupling extraction of global temperature field features and local high-temperature features, laying the foundation for accurate learning of subsequent differentiated temperature ranges and further improving the spatial localization accuracy and feature representation efficiency of temperature field prediction.

[0019] Further, S410 includes: The high-level semantic feature map is upsampled using a dual-branch decoder network containing multiple upsampling levels to obtain the current upsampling level feature map for each upsampling level. Obtain the target space size of the feature map at the current upsampled level; The corresponding upsampled layer feature map is processed using trilinear interpolation to force its spatial size to be aligned with the target spatial size. During the interpolation process, the corner alignment parameter is enabled so that the spatial coordinates of the corner pixels of the input and output feature maps coincide, resulting in an aligned feature map.

[0020] This preferred scheme clarifies the specific implementation methods of spatial dimension alignment and position compensation. By using trilinear interpolation, skip connection features are forcibly aligned to the target spatial size of the current level of the decoder. Simultaneously, corner alignment parameters are enabled, ensuring that the spatial coordinates of the corner pixels in the input and output feature maps completely coincide. This fundamentally eliminates the sub-voxel-level coordinate offset generated during traditional multi-level upsampling and downsampling. For small tumors of 1-3 cm with a safety margin of only 5 mm, this scheme can control the spatial positioning deviation of the ablation heat source center to the sub-voxel level, avoiding incomplete tumor ablation or accidental damage to surrounding normal tissues and critical structures due to coordinate offset, directly ensuring the safety and effectiveness of the ablation surgery.

[0021] Further, S420 includes: The aligned feature map is concatenated with the current upsampled level feature map to obtain the concatenated features; the concatenated features include global temperature field distribution features and high temperature gradient features; The main temperature field branch and the high temperature gradient branch, which are completely symmetrical, are used as the dual-path branches to perform feature extraction processing on the spliced ​​features. The main temperature field branch is used to learn the global temperature field distribution features to obtain the first feature map, and the high temperature gradient branch is used to learn the high temperature gradient features near the ablation needle path to obtain the second feature map.

[0022] This preferred approach achieves complete decoupling learning of the global temperature field distribution characteristics and the high temperature gradient characteristics near the needle track through a fully symmetrical main temperature field branch and a high-temperature gradient branch. The main temperature field branch focuses on fitting the overall heat conduction trend across the entire range of 0-120℃, while the high-temperature gradient branch is dedicated to capturing the drastic temperature changes in the critical zone of tumor necrosis at 60-120℃. This solves the problem of traditional single-branch networks being dominated by the lower-temperature zone and neglecting the details of the high-temperature critical zone at the needle tip during training. It effectively preserves the high-temperature details of the ablation core area and significantly improves the prediction accuracy and gold standard fit of the critical boundary of tumor coagulative necrosis.

[0023] Further, S5 includes: The first feature map and the second feature map are concatenated by channels, and then processed using a three-dimensional convolutional layer and activation function to obtain a spatial weight map. Based on the spatial weight map, the first feature map and the second feature map are weighted and fused to obtain a three-dimensional ablation temperature field prediction map, thus completing the ablation temperature field prediction.

[0024] This preferred scheme clarifies the specific implementation of the gating decision block. It generates a spatial weight map through 3D convolution and activation functions, and then adaptively weights and fuses the bi-branch features based on this map. This mechanism automatically adjusts the fusion weights of the bi-branch at each voxel location: in the low-temperature region of normal tissue, it automatically increases the weight of the main temperature field branch to ensure the predictive stability of the global temperature field distribution; in the high-temperature gradient regions near the needle path and in the tumor area, it automatically increases the weight of the high-temperature gradient branch to enhance the detailed prediction of critical necrosis boundaries, ultimately achieving the optimal balance between global temperature field smoothness and local high-temperature accuracy.

[0025] This invention provides a full-scale recalibration and spatial compensation ablation temperature field prediction system, comprising: The input module is used to acquire the three-dimensional tumor mask, three-dimensional needle track mask and initial specific absorptivity field of the case to be predicted, and perform splicing processing to obtain a multi-channel input feature map; The encoding extraction module is used to extract features from the multi-channel input feature map using an encoder network containing multiple resolution levels. Within each resolution level, a three-dimensional channel attention module is used to perform full-scale channel recalibration on the features of the current level to obtain a multi-scale encoded feature map. The semantic extraction module is used to extract semantic features from the deepest feature map of the multi-scale encoded feature map using the bottleneck layer, so as to obtain a high-level semantic feature map. The decoding upsampling module is used to upsample the high-level semantic feature map using a dual-branch decoder network and to perform feature extraction in parallel using dual-path branches to obtain the first feature map and the second feature map, respectively. The fusion prediction module is used to perform spatially adaptive weighted fusion processing on the first feature map and the second feature map using a gated decision block to obtain a three-dimensional ablation temperature field prediction map and complete the ablation temperature field prediction. The output display module is used to visualize the three-dimensional ablation temperature field prediction map, providing navigation support for ablation surgery. Attached Figure Description

[0026] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein: Figure 1 This is a schematic diagram of a full-scale recalibration and spatial compensation ablation temperature field prediction system according to some embodiments of this specification. Figure 2 This is an exemplary flowchart of a method for predicting ablation temperature field with full-scale recalibration and spatial compensation, as shown in some embodiments of this specification. Figure 3 This is an exemplary schematic diagram showing a visual comparison of ablation temperature field prediction results according to some embodiments of this specification. Detailed Implementation

[0027] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0028] Example 1 Figure 1 This is a schematic diagram of a full-scale recalibration and spatial compensation ablation temperature field prediction system according to some embodiments of this specification.

[0029] In some embodiments, a full-scale recalibration and spatial compensation ablation temperature field prediction system may include an input module, an encoding extraction module, a semantic extraction module, a decoding upsampling module, a fusion prediction module, and an output display module.

[0030] The input module is used to acquire the three-dimensional tumor mask, three-dimensional needle track mask and initial specific absorption field of the case to be predicted, and perform splicing processing to obtain a multi-channel input feature map.

[0031] The encoding extraction module is used to extract features from the multi-channel input feature map using an encoder network containing multiple resolution levels. Within each resolution level, a three-dimensional channel attention module is used to perform full-scale channel recalibration on the features of the current level to obtain a multi-scale encoded feature map.

[0032] The semantic extraction module is used to extract semantic features from the deepest feature map in the multi-scale encoded feature map using the bottleneck layer, so as to obtain a high-level semantic feature map.

[0033] The decoding upsampling module is used to upsample the high-level semantic feature map using a dual-branch decoder network and to perform feature extraction in parallel using dual-path branches, thereby obtaining the first feature map and the second feature map respectively.

[0034] The fusion prediction module is used to perform spatially adaptive weighted fusion processing on the first feature map and the second feature map using a gated decision block to obtain a three-dimensional ablation temperature field prediction map, thus completing the ablation temperature field prediction.

[0035] The output display module is used to visualize the three-dimensional ablation temperature field prediction map, providing navigation support for ablation surgery.

[0036] In some embodiments, a full-scale recalibration and spatial compensation ablation temperature field prediction system can be used to perform a full-scale recalibration and spatial compensation ablation temperature field prediction method, including: S1: acquiring the three-dimensional tumor mask, three-dimensional needle track mask, and initial specific absorptivity field of the case to be predicted, and performing splicing processing to obtain a multi-channel input feature map; S2: using an encoder network containing multiple resolution levels to extract features from the multi-channel input feature map, and within each resolution level, using a three-dimensional channel attention module to perform full-scale channel recalibration processing on the current level features to obtain a multi-scale encoded feature map; S3: using a bottleneck layer to perform semantic feature extraction processing on the deepest level feature map in the multi-scale encoded feature map to obtain a high-level semantic feature map; S4: using a dual-branch decoder network to perform upsampling processing on the high-level semantic feature map, and using dual-path branches to perform parallel feature extraction processing to obtain a first feature map and a second feature map respectively; S5: using a gating decision block to perform spatially adaptive weighted fusion processing on the first feature map and the second feature map to obtain a three-dimensional ablation temperature field prediction map, thus completing the ablation temperature field prediction.

[0037] In some embodiments of this specification, the processor utilizes a full-scale recalibration and spatial compensation ablation temperature field prediction system to perform a full-scale recalibration and spatial compensation ablation temperature field prediction method. Through a dual-branch decoder network, feature extraction is decoupled using a completely symmetrical backbone temperature field branch and a high-temperature gradient branch. Spatially adaptive weighted fusion is then performed through a gated decision block, effectively balancing the extraction of global temperature field distribution and local high-temperature gradient features. This solves the problem of traditional single-branch networks easily losing high-temperature details and significantly improves the prediction accuracy of the core ablation region.

[0038] By introducing a physical adapter, the initial specific absorptivity field is dynamically calibrated using a 3D needle path mask as a spatial index, generating an adaptive specific absorptivity field. This eliminates the heterogeneity of tissue parameters among different cases, solves the problem of mismatch between static physical input and dynamic ablation features, and improves the model's generalization ability in different clinical scenarios.

[0039] A three-dimensional channel attention module is deployed in each resolution level of the encoder network to perform full-scale channel recalibration on the features of the current level, dynamically enhance the feature channels that are sensitive to temperature field evolution and suppress noise channels, thereby improving the feature expression efficiency and the prediction accuracy of critical necrosis boundaries.

[0040] In the dual-branch decoder network, spatial dimension alignment and position compensation are performed on the feature maps passed by the skip connections through trilinear interpolation. This forces alignment of the target size and achieves corner point coincidence, effectively eliminating sub-voxel level coordinate offset caused by multi-level sampling and ensuring high-precision positioning of the ablation heat source center.

[0041] The constructed end-to-end prediction model ensures high accuracy while possessing high inference efficiency, avoiding the extremely high time consumption of traditional numerical simulation methods, and can meet the real-time navigation needs in clinical microwave ablation procedures.

[0042] Example 2 Figure 2 This is an exemplary flowchart illustrating a method for predicting ablation temperature fields with full-scale recalibration and spatial compensation, based on some embodiments of this specification. Figure 2 As shown, the process includes the following steps. In some embodiments, the process may be executed by a processor.

[0043] S1: Obtain the three-dimensional tumor mask, three-dimensional needle path mask and initial specific absorption field of the case to be predicted, and perform splicing to obtain a multi-channel input feature map.

[0044] A 3D tumor mask is a three-dimensional binary mask image extracted from a 3D medical image of a case to be predicted, used to identify the tumor region. For example, the data for a 3D tumor mask may include voxel matrix data with a spatial dimension of 128×128×32 and values ​​of only 0 or 1, where 1 represents the tumor region and 0 represents the background.

[0045] In some embodiments, the processor can segment the enhanced abdominal CT image of the case to be predicted by calling a semi-automatic segmentation algorithm in the medical image processing library, and crop out a region of interest of a specific size centered on the ablation needle tip to obtain three-dimensional tumor mask data.

[0046] A 3D needle tract mask is a 3D binary mask image extracted from a 3D medical image of a case to be predicted, used to identify the location of the ablation needle tract and its heat source center. For example, the data for a 3D needle tract mask may include voxel matrix data with the same size as a 3D tumor mask.

[0047] In some embodiments, the processor can use a semi-automatic segmentation algorithm to extract and binarize the needle tip and needle path regions of medical images to obtain three-dimensional needle path mask data.

[0048] The initial specific absorptivity field refers to the initial spatial energy distribution field generated by the microwave ablation antenna within the tissue. For example, data on the initial specific absorptivity field can include the energy absorptivity values ​​at each voxel point.

[0049] In some embodiments, the processor can obtain initial specific absorptivity field data by acquiring ablation power parameters, calculating them based on the radiation characteristics of the clinical microwave ablation antenna and the Pennes biological heat conduction equation, and then performing Z-Score normalization.

[0050] Multi-channel input feature maps refer to the basic feature data input into the encoder network after concatenating multiple input sources. For example, the data of a multi-channel input feature map may include a three-dimensional feature tensor with 3 channels and a spatial size of 128×128×32, with the three channels corresponding to the tumor mask, needle path mask, and adaptive SAR field, respectively.

[0051] In some embodiments, the processor can obtain multi-channel input feature map data by performing tensor stitching operations on the normalized three-dimensional tumor mask, the three-dimensional needle path mask, and the adaptive specific absorptivity field in the channel dimension.

[0052] In some embodiments, the processor can acquire a three-dimensional tumor mask, a three-dimensional needle tract mask, and an initial specific absorption rate field for the case to be predicted; cascade the initial specific absorption rate field with the three-dimensional needle tract mask to obtain cascaded features; perform dynamic calibration processing on the initial specific absorption rate field using a physical adapter, using the heat source center position marked on the three-dimensional needle tract mask as a spatial index, and process the cascaded features using three-dimensional instance normalization to eliminate tissue parameter heterogeneity among different cases to obtain an adaptive specific absorption rate field; and stitch the three-dimensional tumor mask, the three-dimensional needle tract mask, and the adaptive specific absorption rate field together to obtain a multi-channel input feature map.

[0053] Cascaded features refer to multi-channel feature representations formed by stitching together an initial specific absorptivity field with a three-dimensional needle-channel mask along the channel dimension. For example, the data for cascaded features can include three-dimensional feature tensor data with two channels.

[0054] In some embodiments, the processor can obtain cascaded feature data by calling the channel stitching function in the deep learning framework to stitch together the initial specific absorptivity field and the three-dimensional needle mask.

[0055] The physical adapter is a neural network module used to eliminate heterogeneity in tissue parameters among different cases and to perform dynamic calibration on the initial specific absorbance field. This physical adapter consists of a 3D instance normalization (InstanceNorm3d) layer and 1×1×1 3D convolutional layers.

[0056] In some embodiments, the processor can acquire the physical adapter by loading pre-built physical adapter module code and deploying it in memory. The specific processing steps are as follows: First, the physical adapter receives an initial specific absorptivity field and a 3D pin-channel mask as input; the channel stitching unit concatenates the two along the channel dimension. Then, the 3D instance normalization layer uses the heat source center location marked by the pin-channel mask as a spatial index to perform normalization processing on the concatenated data. The calculation formula is: ; in, and These are the mean and variance of the spatial dimension, respectively. and All of these are learnable affine parameters. For a single-input SAR field, To prevent numerically stable terms with a denominator of zero, the normalized features are then input into a 1×1×1 3D convolutional layer for channel dimensionality reduction and feature smoothing, outputting adaptive ratio absorptivity field data after eliminating heterogeneity.

[0057] A spatial index refers to the spatial coordinate reference used to locate the core region of a 3D instance normalization operation. For example, the spatial index data may include the 3D spatial coordinate system data of voxels with a marker value of 1 in a 3D needle mask.

[0058] In some embodiments, the processor can obtain spatial index data by traversing the three-dimensional needle mask matrix and extracting the non-zero voxel coordinates of the needle or heat source center location.

[0059] S2: Use an encoder network with multiple resolution levels to extract features from the multi-channel input feature map, and within each resolution level, use a three-dimensional channel attention module to perform full-scale channel recalibration on the features of the current level to obtain a multi-scale encoded feature map.

[0060] An encoder network is a deep learning backbone feature extraction network used for multi-scale feature extraction and channel recalibration of multi-channel input feature maps. This encoder network consists of three resolution levels (Level 1-Level 3), each level containing a 3D convolutional module, a 3D compression and activation module (SEBlock3D), and a 3D max-pooling layer. The 3D convolutional module consists of two consecutive 3×3×3 3D convolutional layers, a batch normalization layer, and a ReLU activation function.

[0061] In some embodiments, the processor can obtain the encoder network by initializing the network structure and weight parameters of the three resolution levels described above. The specific processing is as follows: In Level 1, the network receives multi-channel input feature maps. First, it extracts shallow spatial features through a 3D convolutional module containing two 3×3×3 convolutional layers. Then, it inputs the features into a three-dimensional channel attention module for dynamic recalibration of the feature channel weights. Finally, it performs downsampling through a 2×2×2 3D max pooling layer to output the Level 1 feature map. This output is directly used as the input for Level 2, and the above convolution extraction-attention recalibration-pooling downsampling processing logic is repeated until Level 3 outputs a multi-scale encoded feature map with the highest level semantics but the lowest spatial resolution.

[0062] In some embodiments, a three-dimensional channel attention module (SEBlock3D) is deployed at each resolution level of the encoder and decoder to address the issue of weak features of small tumors (1-3 cm) being submerged by the background during multi-level downsampling. This full-scale deployment mechanism performs differentiated feature enhancement at different levels.

[0063] In the superficial layer (Level 1), the 3D channel attention module enhances the geometric feature channels representing tumor boundaries and needle path locations, while suppressing background noise channels from normal liver tissue. For small tumors of 1-3 cm, their effective features account for a very small percentage in 3D CT images; this mechanism prevents edge geometric features from being overwhelmed by the large proportion of normal tissue noise in the superficial layer.

[0064] In the middle layer (Level 2), the three-dimensional channel attention module enhances the physical feature channels related to changes in thermal conduction gradients and suppresses irrelevant tissue texture noise channels. This prevents the extremely steep high-temperature gradient features around small tumors from being downsampled and smoothed during transmission, ensuring the prediction accuracy of critical necrosis boundaries.

[0065] At Level 3, the three-dimensional channel attention module enhances high-level semantic feature channels related to tissue heterogeneity, enabling the model to adapt to differences in tissue parameters among different patients and improving the model's cross-case generalization ability.

[0066] In this embodiment, a 3D max pooling layer is set in front of the 3D channel attention module of each level of the encoder. The kernel size is 2×2×2 and the stride is 2, which is used to downsample and reduce the spatial resolution of the feature map.

[0067] The 3D channel attention module refers to the 3D compression and excitation module (SEBlock3D) deployed in each layer of the encoder and decoder to perform channel dynamic recalibration. This module contains a 3D global average pooling layer and two consecutive fully connected layers. For example, the data for the 3D channel attention module may include compression ratio parameters (e.g., r=16) and the mapping weights of the fully connected layers.

[0068] In some embodiments, the processor can obtain a 3D channel attention module by instantiating the SEBlock3D class and setting the corresponding compression ratio parameter. The specific processing is as follows: First, the module receives the feature map of the current layer as input; then, a 3D global average pooling layer compresses the spatial dimension of the input features, outputting a channel descriptor (dimension C×1×1×1) containing global statistics for each channel; subsequently, the channel descriptor passes through a compressed fully connected layer (dimension reduction by compression ratio r, e.g., r=16), ReLU activation, an activated fully connected layer (dimension increase back to C channels), and Sigmoid activation, calculating and outputting the channel weight coefficients with values ​​between 0 and 1; finally, the channel weight coefficients are multiplied element-wise (scaled) with the original input feature map of the current layer, outputting the recalibrated feature data.

[0069] The current level feature refers to the intermediate feature tensor output after processing by the 3D convolution module of the current level within a specific resolution level of the encoder network. For example, in Level 1, the current level feature data may include feature map data with 32 channels and a spatial size of 128×128×32.

[0070] In some embodiments, the processor can obtain the feature data of the current layer by inputting the output of the previous layer into the 3D convolution module of the current layer for forward propagation calculation.

[0071] Multiscale encoded feature maps refer to a collection of feature maps output by an encoder network after progressive downsampling and feature extraction at three different resolution levels. For example, the data of multiscale encoded feature maps may include feature tensor data of 32×64×64×16 output at Level 1, 64×32×32×8 output at Level 2, and 128×16×16×4 output at Level 3.

[0072] In some embodiments, the processor can obtain multi-scale encoded feature map data by saving the outputs of the encoder network before the max pooling layers at each resolution level.

[0073] In some embodiments, the processor may use an encoder network containing multiple resolution levels to perform feature extraction processing on the multi-channel input feature map to obtain the current level features within each resolution level; process the current level features using three-dimensional global average pooling to compress the global spatial features of each channel into a channel statistic to obtain a channel descriptor; perform nonlinear mapping processing on the channel descriptor using consecutive fully connected layers to obtain the channel weight coefficients corresponding to each channel; and perform channel multiplication processing between the channel weight coefficients and the current level features to obtain the recalibrated multi-scale encoded feature map.

[0074] Channel statistics refer to representative values ​​obtained by globally compressing the spatial information of each channel of the input feature map. For example, channel statistics data may include one-dimensional vector data with dimensions C×1×1×1, where C is the number of channels.

[0075] In some embodiments, the processor can obtain channel statistics data by performing a three-dimensional global average pooling (3D) operation on the current level features in three spatial dimensions.

[0076] Channel descriptors are a concrete representation of channel statistics, used to provide a compact description of the global characteristics of each channel. For example, the data for channel descriptors may include C×1×1×1 tensor data consistent with the channel statistics.

[0077] In some embodiments, the processor can obtain channel descriptor data by receiving the output of three-dimensional global average pooling.

[0078] Channel weight coefficients are scalar weights used to measure the sensitivity of each feature channel to the evolution of the temperature field. For example, the data for channel weight coefficients may include a one-dimensional array of C elements with values ​​ranging from 0 to 1 after Sigmoid activation.

[0079] In some embodiments, the processor can obtain channel weight coefficient data by sequentially inputting channel descriptors into two consecutive fully connected layers for nonlinear mapping learning.

[0080] S3: Use the bottleneck layer to perform semantic feature extraction on the deepest feature map in the multi-scale encoded feature map to obtain a high-level semantic feature map.

[0081] The bottleneck layer is a deep network structure that connects the deepest layer of the encoder network to the bottom layer of the decoder network, and is used to extract the highest-level semantic features. This bottleneck layer consists of a 3D convolutional module with 256 output channels and an SEBlock3D module with a compression ratio of 16, connected in series.

[0082] In some embodiments, the processor can obtain the bottleneck layer by initializing the parameters of the deep convolutional block located between the encoder and decoder. The specific processing steps are as follows: First, the bottleneck layer receives the deepest multi-scale encoded feature map output from the Level3 encoder network as input; then, a 3D convolutional module (containing consecutive 3×3×3 convolutions, batch normalization, and ReLU) performs non-linear feature extraction and dimension mapping on the input; next, the extracted features are input to the SEBlock3D module to perform secondary recalibration at the channel level to suppress redundant semantic information; finally, the bottleneck layer outputs high-level semantic feature map data with 256 channels, which will be fed into the auxiliary prediction head and decoder network respectively.

[0083] In some embodiments, the network bottleneck layer derives a specific absorptivity (SAR) auxiliary prediction head. This auxiliary prediction head employs the Softplus activation function, and its calculation formula is as follows: ; in, To assist in predicting the output tensor data of the convolutional layer before the prediction head, is a natural constant. This activation function ensures that the predicted SAR field output value is always greater than 0, satisfying the physical non-negativity constraint of energy conservation.

[0084] The model's total loss function is a weighted average of the primary task's temperature field loss and the auxiliary task's SAR loss, calculated as follows: ; ; ; ; ; in, This represents the total loss function value of the model; The loss function value for the main task of temperature field prediction is a weighted combination of mean squared error (MSE) and mean absolute error (MAE) (e.g., 1:1), and additional loss weights (e.g., 5 times loss weights) are set for the critical zone of tumor necrosis above 60℃. The mean squared error (MSE) is used as the loss function value for SAR prediction in the auxiliary mission. The weighting coefficient for the auxiliary task ranges from 0.2 to 0.5, with 0.3 being the preferred value. The weighting coefficients representing the MSE loss. The weighting coefficients representing the MAE loss are preferably α:β = 1:1. This represents the mean square error loss in temperature field prediction. This represents the average absolute error loss in temperature field prediction. This represents the total number of elements in the 3D feature map. This represents the mean square error loss weighting coefficient for the temperature field prediction of the i-th voxel, when... At ≥ 60℃, =5, when At <60℃, =1, thereby enhancing the learning accuracy of the tumor necrosis critical zone. This represents the predicted temperature value of the i-th voxel. This represents the gold standard temperature value of the i-th voxel. This represents the weighting coefficient for the average absolute error loss in the temperature field prediction of the i-th voxel. This represents the predicted SAR value of the i-th voxel. This represents the gold standard SAR value of the i-th voxel.

[0085] During model training, the AdamW optimizer was used with an initial learning rate of 1e-4 and a weight decay of 1e-5. The learning rate adopted a cosine annealing strategy with a minimum learning rate of 1e-6. A Dropout layer was set during training with a dropout rate of 0.2.

[0086] High-level semantic feature maps refer to feature tensors extracted from the bottleneck layer that contain the richest deep semantic information, such as organizational heterogeneity. For example, the data of a high-level semantic feature map may include feature tensor data with 256 channels and a spatial size of 16×16×4.

[0087] In some embodiments, the processor can obtain high-level semantic feature map data by forward propagating the output feature map of the deepest layer of the encoder network to the bottleneck layer.

[0088] S4: Upsample the high-level semantic feature map using a dual-branch decoder network, and perform feature extraction in parallel using dual-path branches to obtain the first feature map and the second feature map respectively.

[0089] A dual-branch decoder network is a network structure used to restore the feature space resolution and decouple the extraction of features from different temperature ranges. This network contains three upsampling levels that correspond one-to-one with the encoder network. Each upsampling level includes a trilinear interpolation unit and a completely symmetrical main temperature field branch and a high-temperature gradient branch.

[0090] In some embodiments, the processor can obtain a dual-branch decoder network by initializing a decoding structure containing interpolation layers and dual-path convolutional blocks. The specific processing is as follows: In the lowest upsampling layer, the network receives the high-level semantic feature map output from the bottleneck layer. First, it uses a trilinear interpolation unit to double its spatial size to obtain the current upsampling layer feature map. Simultaneously, another trilinear interpolation unit is called to perform spatial alignment and corner compensation on the skip connection features of the corresponding encoder layer, resulting in an aligned feature map. Subsequently, the concatenation unit merges the two in the channel dimension. The merged features are simultaneously split and input to the internal dual-path branches (main branch and high-temperature branch) for parallel processing. After each branch outputs a feature map within a specific temperature range, the processing at that level ends, and these two features are passed to the next shallower upsampling layer, repeating the above logic until the output decoded feature data is restored to the original input size.

[0091] In some embodiments, the spatial dimension alignment and position compensation based on trilinear interpolation in each upsampling layer of the dual-branch decoder network specifically includes: obtaining the target spatial size of the current decoder feature map, performing trilinear interpolation on the feature map of the corresponding layer passed by the encoder through skip connections, and aligning its spatial size to the target size. During the interpolation process, the corner alignment (align_corners=True) parameter is enabled to ensure that the spatial coordinates of the corner pixels of the input and output feature maps completely coincide.

[0092] This mechanism eliminates sub-voxel-level (e.g., 0.5-1 pixel) spatial coordinate offsets generated during multi-level upsampling and downsampling in conventional transposed convolution. For small tumors of 1-3 cm with a safety margin of only 5 mm, sub-voxel-level offsets can lead to a spatial position deviation of 0.5-1 mm, causing the predicted position of the ablation heat source center to deviate from the actual needle path position, resulting in insufficient ablation coverage of the tumor area or damage to surrounding critical anatomical structures. Enabling corner alignment parameters ensures that the heat source center for heat conduction does not experience spatial position deviations during network hierarchical transmission.

[0093] Dual-path branching refers to two structurally identical network branches set up side-by-side in each layer of the decoder. This dual-path branching includes a completely symmetrical backbone temperature field branch and a high-temperature gradient branch. The kernel size (both 3×3×3), number of channels, activation functions, and SEBlock3D module composition of the two branches are completely equivalent.

[0094] In some embodiments, the processor can obtain dual-path branches by defining two parallel 3D convolutional modules and attention layers in the decoder code. The specific processing is as follows: the input ends of the dual-path branches jointly receive the concatenated features from the previous process; within the main temperature field branch, the input features are processed by its unique convolutional kernel and attention weights, focusing on learning the overall heat conduction trend within the 0-120℃ range, outputting a first feature map; simultaneously, within the high-temperature gradient branch, the input features are processed by another independent set of convolutional kernels and attention weights, forced to fit the drastic temperature changes of 60-120℃ near the needle track, outputting a second feature map. The two branches are completely decoupled physically, without intermediate data exchange, until finally outputting two parallel feature map data paths.

[0095] In some embodiments, the dual-path branch comprises a backbone temperature field branch and a high-temperature gradient branch that are identical in structure, kernel size, number of channels, and activation function. This fully symmetrical dual-branch decoupling design addresses the issue of focus separation in single-branch networks during training, which tend to fit large-area background low-temperature regions (37-60°C) while neglecting small, pinpoint-sized high-temperature critical regions (60-120°C).

[0096] The main temperature field branch is used to learn the global temperature field distribution characteristics across the entire range of 0-120℃, capturing the overall heat conduction law and the temperature rise trend of normal tissue; the high temperature gradient branch is dedicated to learning the local high temperature gradient characteristics of 60-120℃, capturing the high temperature details near the needle tip and the temperature boundary of the critical zone of tumor necrosis.

[0097] The gating decision block performs a weighted fusion of the two branches based on the generated spatial weight graph. The formula for calculating the weighted fusion is as follows: ; in, The first feature map output by the main temperature field branch; This is the second feature map output by the high-temperature gradient branch; For spatial weighting, .when When the value approaches 1, the prediction of the voxel's location is dominated by the main temperature field branch; when... When the value approaches 0, the prediction of the voxel position is dominated by the high-temperature gradient branch.

[0098] The first feature map refers to the feature representation of the global temperature field distribution across the entire range of 0-120℃, output after processing by the main temperature field branch. For example, at the top layer of the decoder, the data of the first feature map may include three-dimensional tensor data with 32 channels and a spatial size of 128×128×32.

[0099] In some embodiments, the processor can obtain first feature map data by performing convolution calculations on the main temperature field branch after inputting the spliced ​​features.

[0100] The second feature map refers to the feature representation of a local high temperature gradient of 60-120℃, output after processing by the high temperature gradient branch. For example, the data of the second feature map may include 32-channel three-dimensional tensor data with the exact same size as the first feature map.

[0101] In some embodiments, the processor can obtain second feature map data by synchronously inputting the concatenated features into a high-temperature gradient branch for convolution calculation.

[0102] In some embodiments, the processor may implement S4 based on the following steps.

[0103] S410: The high-level semantic feature map is upsampled using a dual-branch decoder network with multiple upsampling levels. In each upsampling level, spatial dimension alignment and position compensation are performed on the corresponding upsampling level feature map in the multi-scale encoded feature map using trilinear interpolation to obtain an aligned feature map.

[0104] Aligned feature maps refer to feature representations that eliminate sub-voxel-level coordinate offsets by forcibly stretching the spatial dimensions and matching the corners of the skip connection features passed from the encoder. For example, the data in an aligned feature map can include feature tensor data that precisely matches the target size of the current decoder.

[0105] In some embodiments, the processor can obtain aligned feature map data by processing the skip connection feature map by calling a trilinear interpolation function with a corner alignment parameter (align_corners=True).

[0106] In some embodiments, the processor can use a dual-branch decoder network containing multiple upsampling levels to upsample the high-level semantic feature map to obtain the current upsampled level feature map for each upsampled level; obtain the target spatial size of the current upsampled level feature map; process the corresponding upsampled level feature map using trilinear interpolation to force its spatial size to be aligned to the target spatial size, and enable corner alignment parameters during the interpolation process to make the spatial coordinates of the corner pixels of the input and output feature maps coincide, thereby obtaining an aligned feature map.

[0107] The current upsampled layer feature map refers to the feature tensor that the decoder has not yet fused with the skip connection features after completing the basic resolution upsampling at the current layer. For example, the data of the current upsampled layer feature map may include 3D feature data whose size has doubled after upsampling from the output of the previous layer.

[0108] In some embodiments, the processor can obtain the feature map data of the current upsampled level by performing an interpolation amplification operation on the output of the level above the decoder.

[0109] The target spatial size refers to the absolute three-dimensional resolution parameter that the current decoder level needs to recover. For example, the target spatial size data can include specific dimension numeric tuples, such as (32, 32, 8).

[0110] In some embodiments, the processor can obtain target spatial size data by reading the shape attribute of the current upsampled layer feature map.

[0111] S420: After concatenating the aligned feature map with the current upsampled level feature map, feature extraction is performed in parallel using dual-path branches to obtain the first feature map and the second feature map respectively.

[0112] In some embodiments, the processor can concatenate the aligned feature map with the current upsampled level feature map to obtain the concatenated feature map; using the fully symmetrical main temperature field branch and high temperature gradient branch as the dual path branches, the processor performs feature extraction processing on the concatenated feature map; using the main temperature field branch to learn the global temperature field distribution features to obtain the first feature map; and using the high temperature gradient branch to learn the high temperature gradient features near the ablation needle path to obtain the second feature map.

[0113] The concatenated features refer to the hybrid feature tensor obtained by fusing encoder features (with coordinate offset eliminated) with decoder upsampled features. For example, the concatenated feature data may include three-dimensional tensor data of the total number of channels added together in the channel dimension (e.g., 128 + 64 = 192 channels).

[0114] In some embodiments, the processor can combine the aligned feature map with the current upsampled level feature map by performing a channel concat operation to obtain concatenated feature data.

[0115] S5: The first feature map and the second feature map are spatially adaptively weighted and fused using a gating decision block to obtain a three-dimensional ablation temperature field prediction map, thus completing the ablation temperature field prediction.

[0116] A gated decision block is a network module used to perform spatially adaptive nonlinear weighted fusion of the two features output from a dual-path branch. This gated decision block contains a 3×3×3 three-dimensional convolutional layer that compresses the concatenated channels to a single channel, and a sigmoid activation function layer.

[0117] In some embodiments, the processor can obtain a gated decision block by initializing a combination of convolutions and activations at the end of the dual-branch decoder. The specific processing steps are as follows: First, the gated decision block receives the first and second feature maps output from the dual-path branches, and the channel stitching unit stacks them. Next, the stacked features are input to the 3D convolutional layer and the sigmoid activation layer, and a fusion weight (i.e., a spatial weight map g) with a value between 0 and 1 is calculated at each voxel spatial location. Finally, the weighted summation unit performs a fusion calculation based on this weight, and the module ultimately outputs the fused three-dimensional ablation temperature field prediction map data.

[0118] A three-dimensional ablation temperature field prediction map refers to the final output of the entire model, which represents the predicted three-dimensional temperature distribution during the microwave ablation process. For example, the data in a three-dimensional ablation temperature field prediction map may include three-dimensional tensor data with one channel, a spatial size of 128×128×32, and each voxel value mapped to the range of 0-120℃.

[0119] In some embodiments, the processor can obtain three-dimensional ablation temperature field prediction map data by inputting the weighted and fused feature map into the final 1×1×1 three-dimensional convolutional layer for dimensionality reduction output.

[0120] In some embodiments, the processor can perform channel concatenation processing on the first feature map and the second feature map, and process them using a three-dimensional convolutional layer and activation function to obtain a spatial weight map; based on the spatial weight map, the first feature map and the second feature map are weighted and fused to obtain a three-dimensional ablation temperature field prediction map, thus completing the ablation temperature field prediction.

[0121] A 3D convolutional layer is a fundamental deep learning operator used to simultaneously slide and extract stereo features across three spatial dimensions. For example, the data for a 3D convolutional layer may include the kernel size (e.g., 3×3×3), stride (e.g., 1), padding (e.g., 1), and learnable weight tensor data.

[0122] In some embodiments, the processor can obtain a three-dimensional convolutional layer by calling the Conv3d operator at the bottom layer of the deep learning framework and passing in the set parameters.

[0123] A spatial weight map is a three-dimensional probability matrix used to determine the fusion ratio of two features at each voxel location. For example, the data for a spatial weight map may include three-dimensional matrix data with a spatial size of 128×128×32, a single channel, and values ​​ranging from 0 to 1.

[0124] In some embodiments, the processor can obtain spatial weight map data by concatenating the first feature map and the second feature map and inputting them into the convolutional layer and sigmoid activation layer of the gated decision block.

[0125] In some embodiments, the original abdominal CT image is semi-automatically segmented to generate a tumor and needle tract mask; a 3D ROI is obtained by cropping with the ablation needle tip as the center; the mask, SAR field, and temperature field are normalized. The input layer acquires the tumor mask, needle tract mask, and initial SAR field analysis; after processing by the physical adaptation module, the images are stitched and fused before being fed into the feature encoding layer; the feature encoding layer (Level 1-Level 3) extracts high-dimensional medical features; the intermediate prediction layer (bottleneck layer and SAR prediction head) aggregates features and outputs the SAR prediction result; the decoding and fusion layer receives feature information from different levels of the encoder, performs upsampling and fusion, and modulates the weights of the dual-branch output through a gating decision block; the output layer finally outputs the predicted temperature field. Figure 3As shown, the black dashed line marks the critical boundary of tumor coagulative necrosis at 60°C. The prediction results of this application embodiment at this boundary are consistent with the gold standard. Color code explanation: dark blue (0°C) → blue-green (25°C) → yellow-green (50°C) → orange-red (75°C) → bright white (120°C). The black dashed line is the critical boundary of tumor coagulative necrosis at 60°C.

[0126] In some embodiments, the present invention provides a complete process for training and inference of a prediction system based on the above method.

[0127] (1) Implementation environment: The hardware environment is configured with an Intel Xeon Gold 6248R processor, 256GB of memory, and an NVIDIA RTX 4090 24GB graphics processor. The software environment is an Ubuntu 20.04 operating system and a PyTorch 2.0 deep learning framework.

[0128] (2) Input data preprocessing: A medical CT image dataset containing solid tumors of 1-3 cm was acquired. The tumor region and ablation needle tract region were extracted using a semi-automatic segmentation algorithm to generate a binarized three-dimensional tumor mask and needle tract mask. A three-dimensional region of interest of size 128×128×32 was cropped with the tip of the ablation needle tract as the center. At the same time, the gold standard SAR field data was calculated based on the radiation characteristics of the microwave ablation antenna and the Pennes biological heat conduction equation, and the gold standard three-dimensional temperature field data calibrated by needle tract temperature measurement was obtained. The mask data was normalized to [0,1], the SAR field data was Z-Score normalized, and the temperature field data was normalized to [0,1] (corresponding to 0-120℃).

[0129] (3) Network architecture parameter settings: The size of the multi-channel input feature map is 3×128×128×32.

[0130] The encoder module contains three resolution levels (Level 1-Level 3). Level 1 has 32 output channels for the 3D convolutional module, Level 2 has 64, and Level 3 has 128. Each level uses a 3×3×3 3D convolutional kernel with a stride of 1 and padding of 1. Each level is configured with an SEBlock3D compression ratio r=16, a max-pooling kernel size of 2×2×2, and a stride of 2.

[0131] Bottleneck layer: The 3D convolutional module has 256 output channels and is configured with SEBlock3D compression ratio r=16. The derived SAR auxiliary prediction head contains two 3D convolutional layers, with Softplus activation at the end, and an output size of 1×128×128×32.

[0132] The dual-branch decoder module contains three upsampling levels corresponding to the encoder. The bottom layer is upsampled to 128×32×32×8 and concatenated with interpolated Level 2 features (total 192 channels), with both the main branch and the high-temperature branch outputting 128 channels. The middle layer is upsampled to 64×64×64×16 and concatenated with Level 1 features (total 96 channels), with each branch outputting 64 channels. The top layer is upsampled to 32×128×128×32, with each branch outputting 32 channels. Each branch is equipped with a 3D convolutional kernel (3×3×3) and SEBlock3D.

[0133] Gated decision block: Receives two 32-channel features, splices them, and generates a 1×128×128×32 spatial weight map through 3×3×3 convolution and Sigmoid activation. After weighted fusion, it outputs a 1×128×128×32 three-dimensional ablation temperature field prediction map through 1×1×1 convolution.

[0134] (4) Model training strategy: The total loss function weight ratio is 1:1 (MSE+MAE combination), with a 5x loss weight set for regions above 60℃; the auxiliary task weight coefficient λ=0.3. The AdamW optimizer is used with an initial learning rate of 1e-4 and a weight decay of 1e-5; a cosine annealing strategy is adopted. During training, the dropout rate in the convolution module is set to 0.2.

[0135] In some embodiments, the effectiveness of each core module in the method is verified. Using a three-channel 3D U-Net with only an input mask as the baseline model, the core modules of this invention are gradually added for testing. Test results show that when only the SAR dynamic physics adapter module is added to the baseline model, the root mean square error (RMSE) of temperature prediction decreases from 1.0627℃ to 0.8625℃, and the mean absolute error (MAE) decreases from 0.6103℃ to 0.5418℃. When only the full-scale SEBlock3D module is added to the baseline model, the RMSE decreases to 0.9502℃, and the Dice similarity coefficient (Dice50) at the 50℃ necrosis threshold increases to 0.9832. When only the symmetric dual-branch decoder module is added to the baseline model, the RMSE significantly decreases to 0.7517℃, and the MAE decreases to 0.4765℃. When the complete technical solution of this invention is adopted, the RMSE decreases to a minimum of 0.5978℃, the MAE decreases to 0.3394℃, and the Dice50 reaches a maximum of 0.9869. Experiments verified that the physical adapter, full-scale attention recalibration, and dual-branch decoupling architecture significantly improve the high-temperature gradient feature extraction and overall accuracy.

[0136] In some embodiments, the processor can verify the predictive stability of the method across different tumor sizes ranging from 1.0 to 3.0 cm. Test samples were divided into four intervals based on tumor diameter (1.0-1.5 cm, 1.5-2.0 cm, 2.0-2.5 cm, and 2.5-3.0 cm) for independent testing. Test data showed that across all size intervals, the model's Dice50 remained between 0.9862 and 0.9874; the 95% Hausdorff distance (hd95) remained stable between 1.2085 mm and 1.2296 mm; and the predicted temperature MAE was below 0.345 °C. The data demonstrate that the method exhibits no significant performance degradation across the entire 1-3 cm size range, achieving sub-voxel-level accurate localization and high-precision temperature prediction.

[0137] In some embodiments, the method of the present invention is compared with existing mainstream methods (such as the U-Net 2015 baseline and existing physical guidance networks) on the same dataset. Test data shows that compared to the U-Net baseline, the method of the present invention improves Dice50 by 0.88 percentage points (0.9869 vs. 0.9781), reduces RMSE by 43.7% (0.5978 vs. 1.0627), and reduces hd95 by 13.8% (1.2175 vs. 1.4138). In clinical conformity evaluation, the method of the present invention achieves 99.7% tumor region ablation coverage (baseline 90.5%), 98.9% 5mm safe margin coverage (baseline 85.2%), and reduces the proportion of over-ablation of normal tissue to 3.2% (baseline 15.8%). The data demonstrate that the method of the present invention has significant advantages in boundary prediction accuracy and clinical applicability.

[0138] In some embodiments, the processor can verify the adaptability of this method under different ablation powers. Test samples were grouped and tested according to commonly used clinical ablation powers (30W, 40W, 50W, 60W). Test data showed that within the power range of 30W to 60W, the model's Dice50 remained between 0.9863 and 0.9873, RMSE remained stable between 0.5902℃ and 0.6074℃, and MAE remained between 0.3371℃ and 0.3450℃. The data demonstrate that this method can stably adapt to ablation parameter settings in different clinical scenarios.

[0139] It should be noted that different embodiments may produce different beneficial effects. In different embodiments, the beneficial effects may be any one or a combination of the above, or any other possible beneficial effects.

Claims

1. A method for predicting ablation temperature fields with full-scale recalibration and spatial compensation, characterized in that, include: S1: Obtain the three-dimensional tumor mask, three-dimensional needle track mask and initial specific absorptivity field of the case to be predicted, and perform splicing to obtain a multi-channel input feature map; S2: Use an encoder network with multiple resolution levels to extract features from the multi-channel input feature map, and within each resolution level, use a three-dimensional channel attention module to perform full-scale channel recalibration on the features of the current level to obtain a multi-scale encoded feature map. S3: Use the bottleneck layer to perform semantic feature extraction on the deepest feature map in the multi-scale encoded feature map to obtain a high-level semantic feature map; S4: Upsample the high-level semantic feature map using a dual-branch decoder network, and perform feature extraction in parallel using dual-path branches to obtain the first feature map and the second feature map respectively. S5: The first feature map and the second feature map are spatially adaptively weighted and fused using a gating decision block to obtain a three-dimensional ablation temperature field prediction map, thus completing the ablation temperature field prediction.

2. The method for predicting ablation temperature field with full-scale recalibration and spatial compensation according to claim 1, characterized in that, S1 includes: Obtain the three-dimensional tumor mask, three-dimensional needle path mask, and initial specific absorbance field of the case to be predicted; The initial specific absorptivity field is cascaded with the three-dimensional needle mask to obtain cascaded features; The initial specific absorptivity field is dynamically calibrated using a physical adapter. The heat source center position marked by the three-dimensional needle track mask is used as the spatial index. The cascaded features are processed using three-dimensional instance normalization to eliminate the heterogeneity of tissue parameters among different cases and obtain an adaptive specific absorptivity field. The three-dimensional tumor mask, the three-dimensional needle path mask, and the adaptive specific absorptivity field are spliced ​​together to obtain a multi-channel input feature map.

3. The method for predicting ablation temperature field with full-scale recalibration and spatial compensation according to claim 1, characterized in that, S2 includes: The multi-channel input feature map is processed by an encoder network containing multiple resolution levels to obtain the current level features within each resolution level; The current layer features are processed using three-dimensional global average pooling, and the global spatial features of each channel are compressed into a channel statistic to obtain a channel descriptor. The channel descriptors are nonlinearly mapped using a series of fully connected layers to obtain the channel weight coefficients for each channel. The channel weight coefficients are multiplied with the current level features to obtain the recalibrated multi-scale encoded feature map.

4. The method for predicting ablation temperature field with full-scale recalibration and spatial compensation according to claim 1, characterized in that, S4 includes: S410: The high-level semantic feature map is upsampled using a dual-branch decoder network with multiple upsampling levels. In each upsampling level, spatial dimension alignment and position compensation are performed on the corresponding upsampling level feature map in the multi-scale encoded feature map using trilinear interpolation to obtain an aligned feature map. S420: After concatenating the aligned feature map with the current upsampled level feature map, feature extraction is performed in parallel using dual-path branches to obtain the first feature map and the second feature map respectively.

5. The method for predicting ablation temperature field with full-scale recalibration and spatial compensation according to claim 4, characterized in that, S410 includes: The high-level semantic feature map is upsampled using a dual-branch decoder network containing multiple upsampling levels to obtain the current upsampling level feature map for each upsampling level. Obtain the target space size of the feature map at the current upsampled level; The corresponding upsampled layer feature map is processed using trilinear interpolation to force its spatial size to be aligned with the target spatial size. During the interpolation process, the corner alignment parameter is enabled so that the spatial coordinates of the corner pixels of the input and output feature maps coincide, resulting in an aligned feature map.

6. The method for predicting ablation temperature field with full-scale recalibration and spatial compensation according to claim 4, characterized in that, The S420 includes: The aligned feature map is concatenated with the current upsampled level feature map to obtain the concatenated features; the concatenated features include global temperature field distribution features and high temperature gradient features; The main temperature field branch and the high temperature gradient branch, which are completely symmetrical, are used as the dual-path branches to perform feature extraction processing on the spliced ​​features. The main temperature field branch is used to learn the global temperature field distribution features to obtain the first feature map, and the high temperature gradient branch is used to learn the high temperature gradient features near the ablation needle path to obtain the second feature map.

7. The method for predicting ablation temperature field with full-scale recalibration and spatial compensation according to claim 1, characterized in that, S5 includes: The first feature map and the second feature map are concatenated by channels, and then processed using a three-dimensional convolutional layer and activation function to obtain a spatial weight map. Based on the spatial weight map, the first feature map and the second feature map are weighted and fused to obtain a three-dimensional ablation temperature field prediction map, thus completing the ablation temperature field prediction.

8. A full-scale recalibration and spatial compensation ablation temperature field prediction system, used to execute the full-scale recalibration and spatial compensation ablation temperature field prediction method as described in any one of claims 1 to 7, characterized in that, include: The input module is used to acquire the three-dimensional tumor mask, three-dimensional needle track mask and initial specific absorptivity field of the case to be predicted, and perform splicing processing to obtain a multi-channel input feature map; The encoding extraction module is used to extract features from the multi-channel input feature map using an encoder network containing multiple resolution levels. Within each resolution level, a three-dimensional channel attention module is used to perform full-scale channel recalibration on the features of the current level to obtain a multi-scale encoded feature map. The semantic extraction module is used to extract semantic features from the deepest feature map of the multi-scale encoded feature map using the bottleneck layer, so as to obtain a high-level semantic feature map. The decoding upsampling module is used to upsample the high-level semantic feature map using a dual-branch decoder network and to perform feature extraction in parallel using dual-path branches to obtain the first feature map and the second feature map, respectively. The fusion prediction module is used to perform spatially adaptive weighted fusion processing on the first feature map and the second feature map using a gated decision block to obtain a three-dimensional ablation temperature field prediction map and complete the ablation temperature field prediction. The output display module is used to visualize the three-dimensional ablation temperature field prediction map, providing navigation support for ablation surgery.