Microwave thermal therapy ultrasonic noninvasive temperature measurement system based on ThermoSonoNet model
By introducing the ThermoSonoNet model into the ultrasonic image temperature measurement system, using multi-scale feature perception, enhanced visual state space and multi-head attention coding modules, the problem of low temperature measurement accuracy in the existing technology is solved, and higher temperature measurement accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510307843.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-05-02
AI Technical Summary
Existing ultrasonic image temperature measurement methods rely on feature engineering and it is difficult to fully capture temperature-related information in the image, resulting in a low temperature measurement accuracy.
A microwave thermal therapy ultrasonic non-invasive temperature measurement system based on the ThermoSonoNet model is proposed. Through the multi-scale feature perception module, the enhanced visual state space module and the multi-head attention coding module, the multi-scale temperature characteristics in the ultrasound image are extracted and fused to achieve temperature prediction.
It improves the accuracy and robustness of ultrasonic image temperature measurement, enhances the model's ability to capture local details and context information, and improves the stability and classification accuracy of temperature prediction.
Smart Images

Figure CN119908757A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of non-invasive temperature measurement, and in particular to a microwave hyperthermia ultrasound non-invasive temperature measurement system based on a ThermoSonoNet model. Background Art
[0002] As a safe and efficient new tumor treatment method, tumor hyperthermia technology has developed rapidly in recent years. Tumor hyperthermia can be divided into traditional hyperthermia and thermal ablation. Tumor hyperthermia takes advantage of the fact that tumor tissue is more sensitive to temperature than normal tissue, that is, normal human cells will not be damaged at 42.5-43°C, but most tumor cells will be induced to apoptosis at this temperature. At present, the most commonly used clinical technology is lossy temperature measurement technology, that is, placing the temperature probe into the biological tissue for measurement. Although this method can obtain more accurate temperature information, it will cause pain and trauma to the patient and may cause the metastasis of cancer cells. In tumor hyperthermia, the temperature measurement method of ultrasound images is a non-destructive temperature measurement method.
[0003] The temperature measurement methods of ultrasound images mainly include ultrasound backscattering method, ultrasound thermal imaging method and ultrasound temperature measurement method based on deep learning. The basic principle is to use the backscattering characteristics or thermal effects of ultrasound in tissues, combined with image processing technology, to infer the temperature distribution inside the tissue by analyzing the dynamic changes of grayscale values, texture features and other dynamic changes in ultrasound images; Behnia et al. used entropy imaging method to predict temperature by parameter changes of radio frequency time series signals. This method relies on the direct relationship between signal parameters and temperature changes; Pousek et al. measured temperature by evaluating the correlation between texture parameters of ultrasound images, especially the average grayscale level, and invasively measured temperature; the method proposed by Alvarenga's research team evaluates tissue temperature changes based on grayscale changes in ultrasound images. This method has been experimentally verified to be accurate and repeatable for temperature changes within a certain range; Salem's research team performed percutaneous image-guided thermal ablation on 3 patients with acetabular bone metastasis tumors, and used ultrasound images to monitor the temperature of the acetabular articular cartilage during the operation. The results showed that the technical success rate was 100%, so real-time temperature monitoring of the lesion area can be performed. Zhou et al. estimated the temperature by extracting five characteristic parameters of ultrasound images as the input of artificial neural networks, which is a typical application of machine learning; Chen et al. extracted multiple characteristic parameters after preprocessing ultrasound images, and used convolutional neural networks and temperature for multivariate regression analysis to achieve a temperature prediction model with a higher degree of fitting; Cui et al. used multilayer perceptrons for regression analysis and used continuous piecewise function models to fit the extracted characteristic parameters to predict temperature; Guo et al. constructed a random forest model based on the grayscale mean of the image, the grayscale entropy of the grayscale gradient co-occurrence matrix, and the mixed entropy to measure the temperature of tissues during HIFU treatment; compared with the temperature measurement method based on physical properties, this type of method can not only provide temperature information, but also intuitively present the spatial distribution of the temperature field, providing a more comprehensive reference for clinical treatment.
[0004] At present, the research on ultrasound image temperature measurement mainly relies on feature engineering methods, that is, extracting specific features (such as texture features, grayscale distribution, etc.) from ultrasound images through artificially designed rules, and building a temperature prediction model based on these features; however, this method has the following limitations: first, the feature extraction process relies on artificial rule design, and it is difficult to fully capture the temperature-related information in the image; second, the feature extraction and temperature prediction processes are decoupled from each other, resulting in insufficient information utilization; finally, due to the limited feature expression ability, the existing models (such as Alexnet, VGG16, ResNet152, ConvMixer, ConvNext and Mamba models) are applied to ultrasound image temperature measurement, and the temperature measurement accuracy is low. Summary of the invention
[0005] In order to solve the deficiencies in the prior art, the present invention provides a microwave hyperthermia ultrasound non-invasive temperature measurement system based on the ThermoSonoNet model, aiming to improve the prediction accuracy of the model.
[0006] The present invention proposes a microwave hyperthermia ultrasound non-invasive temperature measurement system based on the ThermoSonoNet model, comprising:
[0007] Microwave hyperthermia device;
[0008] A waveguide antenna, the waveguide antenna is connected to the output end of the microwave hyperthermia device, the microwave hyperthermia device feeds energy to the waveguide antenna, and the waveguide antenna generates heat;
[0009] Type B ultrasound equipment, which is used to collect ultrasound image videos;
[0010] The computer is connected to the B-type ultrasound device and is used to receive ultrasound image video collected by the B-type ultrasound device and convert it into ultrasound image data; and input it into the ThermoSonoNet prediction model to perform temperature prediction.
[0011] The computer settings include:
[0012] A conversion module, configured to convert the ultrasound image video into ultrasound image data to be detected;
[0013] The pre-processing module is configured to perform median filtering on the ultrasonic image data to be detected to obtain filtered ultrasonic image data.
[0014] A region of interest positioning module is configured to position the filtered ultrasound image data to obtain image data after the region of interest is positioned;
[0015] A subtraction module is configured to perform a subtraction operation on the image data after the region of interest is located to obtain subtraction image data;
[0016] The temperature prediction module is configured to input the subtraction image data into the ThermoSonoNet prediction model to obtain the predicted temperature.
[0017] The ThermoSonoNet prediction model includes a multi-scale feature perception module, an enhanced visual state space module, a multi-head attention encoding module and a classification output module; the multi-scale feature perception module is used to extract multi-scale features through convolution kernels of different sizes, and use adaptive average pooling, linear transformation and Sigmoid activation function to generate fusion weights to achieve multi-scale feature fusion; the enhanced visual state space module is used to extract local detail features and global features through enhanced convolution branches and visual state space branches respectively, and use the modulated interaction feature aggregation module for deep fusion to enhance the model's ability to represent complex temperature features; the multi-head attention encoding module is used to reorganize and fuse features from three dimensions: channel, space and scale through parallel multi-head attention mechanisms and multi-layer perceptrons, improve feature expression capabilities, and retain original feature information through residual connections; the classification output module is used to map high-dimensional features to temperature labels through global average pooling and fully connected layers to achieve temperature classification output.
[0018] The beneficial effects achieved by the present invention are:
[0019] 1. The present application proposes an ultrasound image processing step for temperature prediction of ultrasound images, including median filtering, region of interest positioning, and subtraction operation. This specific processing sequence improves the accuracy of the prediction.
[0020] 2. Through the multi-scale feature perception module, the multi-scale temperature features in ultrasound images are effectively extracted, enhancing the model's ability to capture local details and contextual information;
[0021] By enhancing the visual state space module, the spatiotemporal continuity of temperature features can be modeled to improve classification accuracy and robustness.
[0022] Through the multi-head attention encoding module, the model's sensitivity to key temperature features is improved and the feature expression ability is enhanced; the overall model structure is efficient and has low computational complexity, making it suitable for practical application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a diagram of the microwave hyperthermia ultrasound non-invasive temperature measurement system of the ThermoSonoNet model of the present invention.
[0024] Figure 2 It is a schematic diagram of the structure of the ThermoSonoNet prediction model of the present invention.
[0025] Figure 3 It is a schematic diagram of the structure of the multi-scale feature perception module of the present invention.
[0026] Figure 4 It is a schematic diagram of the structure of the enhanced visual state space module of the present invention.
[0027] Figure 5 This is a schematic diagram of the structure of the multi-head attention encoding module of the present invention.
[0028] Figure 6 It is a structural schematic diagram of the classification output module of the present invention.
[0029] Figure 7 This is the ultrasonic image after filtering processing of the present invention.
[0030] Figure 8 It is the ultrasound image positioning image of the present invention.
[0031] Fig. 9 It is the contrast image of the subtraction operation of the present invention.
[0032] Fig.10 This is a comparison chart of the effects of different models at different granularities.
[0033] Fig.11 This is a graph of the test results of the ThermoSonoNet prediction model at different granularities.
[0034] Fig.12 This is a comparison chart of the evaluation indicators of the four prediction models, Resnet152, ConvMixer, Mamba and ThermoSonoNet, at three granularities: Coarse, Medium and Fine. DETAILED DESCRIPTION
[0035] To facilitate those skilled in the art to understand the present invention, specific implementations of the present invention are described below with reference to the accompanying drawings.
[0036] like Figure 1 As shown, the present application proposes a microwave hyperthermia ultrasound non-invasive temperature measurement system based on the ThermoSonoNet model, comprising:
[0037] Microwave hyperthermia device, the microwave hyperthermia device used is the MH-IY model of microwave hyperthermia device manufactured by Beijing Muheyu Electronics Co., Ltd., the output power of which is adjustable within the range of 0-60W; at the operating frequency of 2.45GHz, the peak gain reaches 2.2dB;
[0038] The waveguide antenna is connected to the output end of the microwave thermotherapy device. The microwave thermotherapy device feeds energy to the waveguide antenna, and the waveguide antenna generates heat to perform thermal therapy on the tumor. The waveguide antenna uses a compact metamaterial-filled waveguide aperture thermal therapy antenna with a size of 10 mm × 17.4 mm. At an operating frequency of 2.45 GHz, the peak gain reaches 2.2 dB.
[0039] B-type ultrasound equipment, which is used to collect ultrasound video images, has internal parameter settings of 30 frames per second, i.e., 30 images per second. The B-type ultrasound equipment is Mindray M9Vet portable ultrasound diagnostic instrument produced by Mindray. The ultrasound probe used is the linear array L20-5s ultrasound probe that is compatible with the ultrasound measurement system and suitable for small organs, with 192 array elements and a sampling depth of 5 cm.
[0040] The computer is connected to the B-type ultrasound device and is used to receive ultrasound image video collected by the B-type ultrasound device and convert it into ultrasound image data; obtain subtraction image data through preprocessing; and input it into the ThermoSonoNet prediction model to perform temperature prediction.
[0041] The computer settings include:
[0042] A conversion module, configured to convert the ultrasound image video into ultrasound image data to be detected;
[0043] Specifically, the ultrasonic image video is converted into ultrasonic image data to be detected by frame-by-frame extraction.
[0044] The pre-processing module is configured to perform median filtering on the ultrasonic image data to be detected to obtain filtered ultrasonic image data.
[0045] A region of interest positioning module is configured to position the filtered ultrasound image data to obtain image data after the region of interest is positioned;
[0046] The specific method for positioning the filtered ultrasonic image data is as follows: taking the position of the temperature probe as the center, a square area with a side length of 256 pixels is defined as the image after the region of interest is positioned.
[0047] A subtraction module is configured to perform a subtraction operation on the image data after the region of interest is located to obtain subtraction image data;
[0048] The specific method of the subtraction operation is as follows: subtract the image data after the region of interest is located from the image data after the region of interest is located, and obtain the output image by subtracting the corresponding pixels; through the subtraction operation, the contrast of the target tissue is enhanced, and noise and artifacts are eliminated.
[0049] Median filtering is a nonlinear filtering technology, which is mainly used to remove noise in images while retaining the edge information of the image. In ultrasound images, noise usually appears as randomly distributed isolated pixels, which will interfere with subsequent image analysis and processing. It is necessary to remove noise. Noise will affect the accuracy of ROI positioning and may even lead to misjudgment. Median filtering can effectively remove salt and pepper noise, improve the signal-to-noise ratio of the image, and has a significant denoising effect. At the same time, median filtering retains edges: Compared with linear filtering, median filtering can better retain the edge information of the image while removing noise, which is particularly important for subsequent ROI positioning.
[0050] ROI localization is to focus on the target tissue by setting a specific area and ignoring irrelevant background information. In ultrasound images, the position of the temperature probe is usually regarded as the target area. The ROI is located through filtered ultrasound image data to focus on the target: by defining the ROI, the target tissue can be processed intensively to reduce the interference of background noise and irrelevant information. The processing efficiency is improved: reducing the processing range can significantly improve the speed and efficiency of image processing. At the same time, data simplification can be achieved: in the subsequent subtraction operation, ROI localization can simplify the data volume, making the subtraction operation more focused on the target area and improving contrast and clarity.
[0051] The subtraction operation is to enhance the contrast of the target tissue and eliminate noise and artifacts by subtracting the current image from the reference image (usually the first image); Enhance contrast through subtraction operation: The subtraction operation can highlight the difference between the target tissue and the background, making the target clearer; At the same time, artifacts can be eliminated. By subtracting the reference image, artifacts caused by equipment or environmental factors can be eliminated, thereby improving image quality; The image is subjected to median filtering and region of interest positioning before the subtraction operation to ensure that all noise and irrelevant information have been removed, thereby maximizing the contrast of the target tissue.
[0052] The present application proposes an ultrasound image processing step for temperature prediction of ultrasound images, including median filtering, region of interest positioning, and subtraction operation. This specific processing sequence improves the accuracy of prediction.
[0053] First, median filtering is performed. If median filtering is performed after ROI positioning, noise may have interfered with the accuracy of the ROI, resulting in positioning errors. In addition, median filtering requires processing of the entire image, and local filtering within the ROI may not be able to completely remove noise.
[0054] The ROI positioning must be performed before the median filter processing and subtraction operation. If the subtraction operation is performed first, background noise and irrelevant information may be introduced into the subtraction result, reducing the subtraction effect. The ROI positioning ensures that the subtraction operation is only for the target area, thereby improving contrast and clarity.
[0055] The subtraction operation must be performed last: The subtraction operation relies on high-quality input images; if the subtraction operation is performed first, the subsequent filtering and positioning operations may have an adverse effect on the subtraction result, resulting in poor contrast enhancement and noise removal effects.
[0056] The temperature prediction module is configured to input the subtraction image data into the ThermoSonoNet prediction model to obtain the predicted temperature.
[0057] like Figure 2-6 As shown in the figure, the ThermoSonoNet prediction model includes a multi-scale feature perception module, an enhanced visual state space module, a multi-head attention encoding module and a classification output module; the multi-scale feature perception module is used to extract multi-scale features through convolution kernels of different sizes, and use adaptive average pooling, linear transformation and Sigmoid activation function to generate fusion weights to achieve multi-scale feature fusion; the enhanced visual state space module is used to extract local detail features and global features through enhanced convolution branches and visual state space branches respectively, and use the modulated interaction feature aggregation module for deep fusion to enhance the model's ability to represent complex temperature features; the multi-head attention encoding module is used to reorganize and fuse features from three dimensions of channel, space and scale through parallel multi-head attention mechanism and multi-layer perceptron, improve feature expression ability, and retain original feature information through residual connection; the classification output module is used to map high-dimensional features to temperature labels through global average pooling and fully connected layers to achieve temperature classification output.
[0058] The multi-scale feature perception module includes a feature segmentation layer (Split), the output of which is connected to the input of a 3×3 convolutional layer and a 5×5 convolutional layer respectively; the output of the 3×3 convolutional layer and the 5×5 convolutional layer are connected to a batch normalization layer (BN) and a ReLU activation function in turn; the output of the ReLU activation function is connected to an adaptive average pooling layer (AdaptiveAVgPool2d) and a fully connected layer (Linear) in turn through an element-wise addition operation; the output of the fully connected layer (Linear) is processed by a Sigmoid function and a Reshape operation; and the final output is obtained by element-wise multiplication with the output of the ReLU activation function, which is then output to the enhanced visual state space module.
[0059] like Figure 3As shown in the figure, the input features are first split into two branches through the Split operation, and 3×3 and 5×5 convolution kernels are used to extract local detail features and context information respectively; then, the fused features are obtained through batch normalization, ReLU activation function and element addition operation, and the fusion weights are generated by adaptive average pooling, linear transformation and Sigmoid activation function to achieve the fusion of multi-scale features.
[0060] The design of the multi-scale feature perception module fully considers the diversity and complexity of temperature features in ultrasonic microscopy images. Through a 3*3 convolution kernel and a 5*5 convolution kernel, the model's ability to extract multi-scale temperature features is enhanced; this fusion mechanism enables the model to independently determine the contribution of features of different scales according to the characteristics of the input features, thereby providing high-quality initial feature representation for subsequent feature enhancement and classification tasks.
[0061] like Figure 4 As shown in the figure, the enhanced visual state space module includes a feature segmentation layer (Split), the input end of the feature segmentation layer (Split) is connected to the output end of the multi-scale feature perception module; the output end of the feature segmentation layer (Split) is connected to the enhanced convolution branch and the visual state space branch respectively; the enhanced convolution branch includes a depth separable convolution layer (DW Conv), a batch normalization layer (BN), a ReLU activation function, a squeeze excitation module (SE Block), a point convolution layer (PW Conv), a batch normalization layer (BN), a ReLU activation function, and a squeeze excitation module (SE Block) to extract local detail features. The visual state space branch includes a layer normalization layer (LN), an adaptive residual visual state space module (ARVSS Block), and a channel shuffle layer (Channel shuffle); it is used to extract global features; the output ends of the enhanced convolution branch and the visual state space branch are deeply fused through the modulated interactive feature aggregation module (MIFA) to enhance the model's ability to represent complex temperature features; and output to the multi-head attention encoding module.
[0062] The enhanced visual state space module introduces the cascade structure of depthwise separable convolution (DW Conv) and pointwise convolution (PW Conv), as well as the squeeze excitation module (SE Block), combined with batch normalization (BN) and ReLU activation, to reduce the computational complexity while enhancing the local feature expression capability; secondly, through channel shuffle (Channel Shuffle) and adaptive residual visual state space (ARVSS), cross-channel information interaction and feature reuse are strengthened to avoid redundant loss of microtexture information; the squeeze excitation module (SE Block) is an attention mechanism module, which can adaptively adjust channel weights, highlight channels related to temperature features, and effectively enhance the expression of temperature features.
[0063] In the ultrasonic temperature measurement task, the enhanced visual state space module enables the model to process and enhance the temperature features of ultrasound images from two perspectives: local details and global state space. On the one hand, it accurately extracts local features related to temperature, and on the other hand, it fully grasps the overall state of temperature distribution, thereby providing ultrasound images with richer and more discriminative feature representations and improving the performance of the model in this task.
[0064] like Figure 5 As shown in the figure, the multi-head attention encoding module includes a dimensionality transformation layer (Dim Transform), the output end of the dimensionality transformation layer (Dim Transform) is connected to three normalization layers (Layer Norm) respectively; the output ends of the three normalization layers (LayerNorm) are connected to the multi-head attention layer (Multi-head Attention); the output end of the multi-head attention layer (Multi-headAttention) is connected to the input end of the random dropout layer (Dropout), the output of the random dropout layer (Dropout) and the output of the dimensionality transformation layer are element-wise added and input to the layer normalization layer (Layer Norm); the layer normalization layer (Layer The output end of the Norm is connected to the multi-layer perceptron module (MLP) and the random dropout layer (Dropout) in sequence. The output of the random dropout layer (Dropout) and the output of the dimension transformation layer are added element by element again, and finally output to the classification output module; the multi-layer perceptron module (MLP) includes a fully connected layer (Linear), a Gaussian error linear unit activation function layer (GeLU), a random dropout layer (Dropout), a fully connected layer (Linear), and a random dropout layer (Dropout).
[0065] The input features are first adjusted to the form of [B, H+W, C] through a dimensionality transformation module, and then long-distance dependencies are captured through a parallel normalization layer and a multi-head attention mechanism. The features are further transformed nonlinearly through a multi-layer perceptron, and the original feature information is retained through residual connections to enhance the generalization ability of the model.
[0066] The core of the multi-head attention encoding module is to achieve efficient encoding and processing of ultrasound image features through the coordinated operation of multi-head attention mechanism and multi-layer perceptron and other components, providing high-quality feature representation for subsequent image temperature classification.
[0067] Starting from the input stage, the feature X(X∈R B×C×H×W, where B is the batch size, C is the number of channels, H and W are the height and width of the feature map respectively) enter the module in the form of a tensor of [B, C, H, W]. In order to adapt to subsequent operations, the features first enter the dimension transformation layer (Dim Transform) for dimension transformation and become [B, H+W, C]. This operation flattens the spatial dimensions to better interact with information between channels and pixel positions; next, three parallel normalization layers are used to ensure that the feature distribution between different samples and channels is relatively consistent, thereby stabilizing the subsequent model training process; next, the features enter three parallel Layer Norm layers respectively, and the calculation formula is as follows:
[0068]
[0069] Among them, μ and σ 2 are the mean and variance of the features, γ and β are learnable parameters, and ∈ is a small constant for numerical stability. The normalized features are input into the multi-head attention layer for processing to obtain the output feature F attn ; This module effectively captures the long-distance dependencies in ultrasound images through parallel multi-subspace modeling, and its output is regularized by the Dropout layer; then, the features are normalized again by the Layer Norm layer to further stabilize the feature distribution.
[0070] In the multi-layer perceptron part, the features are then input into the MLP (multi-layer perceptron) module for processing; this module implements nonlinear transformation through a multi-layer fully connected structure, thereby effectively modeling the complex relationship between temperature features and other attributes in ultrasound images and enhancing feature extraction capabilities; the output of the MLP module also passes through the Dropout layer to further enhance the generalization ability of the model; finally, the features that have undergone a series of processing are added to the original input features through a residual connection to obtain the final output F. This residual connection design helps to alleviate the gradient vanishing problem in deep neural networks, while retaining the information of the original input features, allowing the model to better integrate feature representations at different levels.
[0071] The classification output module includes an adaptive average pooling layer (AdaptiveAVgPool2d) and a fully connected layer (Linear); the classification output module is used to map high-dimensional features to temperature labels through global average pooling and fully connected layers to achieve temperature classification output.
[0072] Based on the characteristics of ultrasound image temperature measurement, this application proposes an end-to-end deep neural network model ThermoSonoNet for ultrasound image temperature measurement. This model automatically learns temperature-related features directly from the original ultrasound image, overcomes the limitations of feature engineering methods, and significantly improves the accuracy and practicality of ultrasound image temperature measurement;
[0073] Aiming at the multi-scale distribution of features and signal-to-noise ratio fluctuation in ultrasound images, a Multi-scale Feature Perception Module (MFPM) was designed to achieve accurate extraction and fusion of temperature features at different scales, thus improving temperature measurement accuracy and robustness.
[0074] The Enhanced Visual State Space Module (EVSSM) was designed to enhance the spatiotemporal continuity of features through state space modeling methods, thereby improving the stability of temperature prediction and reducing temperature measurement errors.
[0075] The Multiple Attention Encoder Module (MAEM) is introduced to perform feature fusion from three dimensions: channel, space, and scale, to enhance the model’s sensitivity to key temperature features.
[0076] The ThermoSonoNet prediction model is constructed as follows:
[0077] S1, obtain ultrasound image video, and convert the video into ultrasound training image by frame extraction;
[0078] When collecting actual experimental data, this step involves extracting the video frame by frame to obtain 30 pictures per second.
[0079] S2, preprocessing the ultrasonic training image to obtain the ultrasonic training image after filtering;
[0080] The preprocessing method is as follows: in the spatial dimension, the median filtering technology is applied to the trained ultrasound image data; the ultrasound image data after filtering is obtained; Figure 7 A is the original image before filtering, Figure 7 B is the ultrasonic image after filtering.
[0081] We preprocessed the ultrasound images, namely filtering, to reduce the impact of noise interference in B-ultrasound images on the calculation results; in the spatial dimension, we applied the median filtering technology; as a nonlinear signal processing method, median filtering can effectively suppress high-frequency noise in ultrasound images while retaining the edge details of the image and avoiding image distortion.
[0082] S3, positioning the ultrasonic training image after filtering to obtain a training image after the region of interest is positioned;
[0083] The specific operation is: taking the position of the temperature probe as the center, a square area with a side length of 256 pixels is delineated as the training image after the ROI is located.
[0084] To further improve the accuracy of temperature measurement and reduce the scope of image processing, we implemented the localization of the region of interest (ROI). Figure 7 C is the image after the region of interest is located; through subsequent analysis and processing of this specific area, we can more accurately focus on the image information directly related to temperature measurement.
[0085] Figure 8 The left side is the MATLAB interface for the actual positioning operation. Move the mouse on the picture and select the center point in the center of the cross. After selection, a square with 128 on the top, 128 on the bottom, 128 on the left, and 128 on the right will be automatically captured with the center point as the center. Figure 8 As shown on the right, it is a square with a side length of 256 pixels.
[0086] S4, performing a subtraction operation on the training image after the ROI is located to obtain a subtraction training image;
[0087] The specific method of the subtraction operation is as follows: subtract the training image after the region of interest is located from the first training image after the region of interest is located, and obtain the output image by subtracting the corresponding pixels; through the subtraction operation, the contrast of the target tissue is enhanced, and the noise and artifacts are eliminated; Fig. 9 For subtraction operation comparison images, A is the original image and B is the subtracted image.
[0088] S5, using the subtracted training images as samples and the corresponding labeled temperatures as labels, the initial neural network is trained to obtain the ThermoSonoNet prediction model.
[0089] The experimental verification and analysis are as follows;
[0090] Evaluation indicators
[0091] In order to comprehensively evaluate the performance of this model, this study used multiple evaluation indicators for measurement; specifically, accuracy (Acc) was used to reflect the degree of fit and generalization ability of the model on the test set; at the same time, this study also calculated precision (Precision), recall (Recall) and F1 score (F1 Score) to evaluate the accuracy and comprehensiveness of the model in the classification task; among them, precision measures the proportion of samples predicted by the model to be positive that are truly positive, recall reflects the ability of the model to correctly identify all positive samples, and F1 score is the harmonic mean of precision and recall, which is used to comprehensively evaluate the performance of the model; in addition, considering that the data set may be unbalanced, this study also calculated the weighted AUC (Weighted AUC) indicator to more accurately reflect the overall performance of the model in different categories; through the comprehensive analysis of these evaluation indicators, this study can comprehensively and objectively evaluate the performance of this model.
[0092] Training details:
[0093] The operating system of the training model mentioned in this application is Windows 11, the processor is Intel (R) Core (TM) i7-13700H, CPU @ 2.40GHz, the memory is 64GB, the graphics card is NVIDIA GeForce RTX 4060, the graphics card memory is 8GB, the parallel computing environment is CUDA 11.8, the deep learning framework is Pytorch, the programming language is Python, and the development environment is Pycharm; in the hyperparameter setting of the training model, the learning rate of the optimizer is set to 0.001, the Adam optimizer is used, and the total number of training rounds is set to 70.
[0094] Model effect evaluation:
[0095] The ThermoSonoNet prediction model of this application is compared qualitatively and quantitatively with some other commonly used models at three granularities; this application selects six commonly used models: Alexnet, VGG16, ResNet152, ConvMixer, ConvNext and Mamba; the experimental results are as follows Figure 10-11 , as shown in Table 1-Table 3; Fig.10 The comparison of the effects of different models with different granularities is shown, where A is a comparison of coarse-grained classification models; B is a comparison of medium-grained classification models; C is a comparison of fine-grained classification models; Fig.11 The test results of our model (ThermoSonoNet prediction model) at different granularities are shown; A is the performance of the coarse-grained classification model; B is the medium-grained classification performance and C is the fine-grained classification performance.
[0096] Table 1 Comparison of coarse-grained classification models
[0097] Network Model ACC↑ Precision↑ Recall↑ F1 Score↑ Weighted AUC↑ ConvNext 5.09 0.01 0.01 0.01 49.43 VGG16 6.00 0.02 0.06 0.01 50.00 Alexnet 16.35 0.38 6.25 0.71 55.25 Resnet152 97.84 98.02 97.84 98.63 99.38 ConvMixer 97.84 98.02 97.84 98.63 99.77 Mamba 98.61 98.47 98.61 98.70 99.81 Our 99.98 99.54 99.98 99.96 99.92
[0098] Comparison of coarse-grained temperature measurement effects:
[0099] In the comparative experiment of coarse-grained classification models, the performance of each model showed significant differences; the experimental results are shown in Table 1 and Fig.10 (A) shown.
[0100] Traditional models such as ConvNext, VGG16 and Alexnet performed poorly in terms of test accuracy, precision, recall, F1 score and weighted AUC indicators. These models may have difficulty coping with complex ultrasound image classification tasks due to their relatively simple structures and limited feature extraction and expression capabilities. For example, ConvNext and VGG16 are prone to losing detail information when processing high-resolution images, while Alexnet has difficulty capturing high-order features in images due to insufficient network depth. These limitations lead to their performance in classification tasks being far below expectations. In contrast, models such as Resnet152, ConvMixer and Mamba models showed significant advantages; these models significantly improved the classification performance through deep network structure, complex feature extraction mechanism and optimized training strategy; Resnet152 alleviated the gradient vanishing problem by residual connection and was able to train deeper networks; ConvMixer enhanced the ability to extract local features through mixed convolution operations; and the Mamba model introduced an innovative attention mechanism to further improve the efficiency of feature expression; among them, the Mamba model stood out with a test accuracy of 98.61%, becoming the leader among the compared models; it also performed well in indicators such as precision, recall rate and F1 score, showing strong comprehensive performance.
[0101] However, despite the outstanding performance of the Mamba model, the ThermoSonoNet model in this study has surpassed it in all indicators; Fig.11 As shown in (A), the test accuracy of ThermoSonoNet reaches 99.98%, which is 1.37 percentage points higher than that of the Mamba model. This significant improvement is due to the multi-scale feature fusion mechanism and adaptive optimization strategy of ThermoSonoNet, which enables it to capture the key information in the image more accurately. In addition, in terms of precision, recall rate, F1 score and weighted AUC, ThermoSonoNet has also reached a high level, surpassing all the comparison models. For example, in terms of precision, ThermoSonoNet is 1.2 percentage points higher than the Mamba model, and in terms of weighted AUC, the improvement is as high as 1.5 percentage points.
[0102] ThermoSonoNet's outstanding performance fully demonstrates its powerful capabilities in coarse-grained classification tasks; its accurate classification ability, stable performance and high generalization ability not only provide reliable technical support for ultrasonic image analysis, but also lay a solid foundation for the practical application of ultrasonic temperature measurement.
[0103] Table 2 Comparison of particle size classification models
[0104] Network Model ACC↑ Precision↑ Recall↑ F1 Score↑ Weighted AUC↑ ConvNext 3.86 0.01 0.02 0.01 39.99 VGG16 4.00 0.02 3.00 0.02 43.00 Alexnet 5.86 0.12 3.23 0.24 49.66 ConvMixer 87.20 45.86 87.20 30.13 94.52 Resnet152 90.07 91.13 90.07 90.57 95.90 Mamba 90.82 91.58 90.82 90.78 99.42 Our 93.36 94.27 93.36 93.73 99.87
[0105] Comparison of medium-sized temperature measurement effects:
[0106] In the comparison of medium-grained classification models, the performance of each model is significantly different. The experimental results are shown in Table 2 and Fig.10 (B) shown.
[0107] As the classification interval decreases, the changes in ultrasound images become more subtle, which places higher demands on the recognition ability of the model; traditional models such as ConvNext, VGG16 and Alexnet perform poorly in terms of test accuracy, precision, recall, F1 score and weighted AUC. These models have difficulty capturing subtle features in ultrasound images, resulting in limited classification performance.
[0108] In contrast, advanced models such as ConvMixer, Resnet152, and Mamba showed stronger classification capabilities; they successfully captured more detailed information through deeper network structures and complex feature extraction mechanisms, thereby improving classification accuracy; however, despite the excellent performance of the Mamba model, the model in this study still set a new record with better performance. Fig.11 As shown in (B), the test accuracy of the ThermoSonoNet model of this application reaches 93.36%, which is 2.54 percentage points higher than that of the Mamba model. At the same time, in terms of indicators such as precision, recall rate, F1 score and weighted AUC, this model also reaches a relatively high level.
[0109] Table 3 Comparison of fine-grained models
[0110] Network Model ACC↑ Precision↑ Recall↑ F1 Score↑ Weighted AUC↑ ConvNext 1.15 0.01 0.01 0.02 51.43 VGG16 2.00 0.01 1.00 0.03 54.00 Alexnet 2.15 0.03 1.32 0.06 55.04 ConvMixer 66.19 67.62 66.19 65.76 85.69 Resnet152 68.74 68.98 68.47 65.92 89.27 Mamba 75.39 78.20 75.39 76.06 99.76 Our 80.27 80.39 80.27 79.78 99.83
[0111] Comparison of fine-grained temperature measurement effects:
[0112] As the classification interval is reduced to 0.2°C, the changes in the ultrasound image are more subtle than those at the 0.5°C interval, which places higher demands on the recognition ability of each model. The experimental results are shown in Table 3 and Fig.10 (C) shown.
[0113] Traditional models such as ConvNext, VGG16 and Alexnet have extremely low test accuracy when faced with such subtle classification tasks, and are almost unable to effectively distinguish different categories. They also perform poorly in indicators such as precision, recall, F1 score and weighted AUC; these models are difficult to adapt to the high requirements of fine-grained classification tasks and cannot capture the subtle differences in ultrasound images.
[0114] In contrast, advanced models such as ConvMixer, Resnet152 and Mamba have shown stronger adaptability; they successfully capture more detailed information through more complex network structures and sophisticated feature extraction mechanisms, thereby improving classification accuracy; among them, the Mamba model leads with a test accuracy of 75.39%, and all indicators perform well, becoming the best among the compared models; however, despite the excellent performance of the Mamba model, the model of this study still stands out with its more outstanding performance.
[0115] like Fig.11 As shown in (C), the test accuracy of this model reaches 80.27%, which is nearly 5 percentage points higher than that of the Mamba model. This result is particularly outstanding in fine-grained classification tasks. At the same time, this model also reaches a high level in terms of indicators such as precision, recall rate, F1 score and weighted AUC. In particular, the weighted AUC of this model is 99.83%, which fully demonstrates its strong strength and stability in fine-grained classification tasks. This model not only has accurate classification capabilities, but also has a keen ability to capture subtle features and a high degree of generalization ability, which gives it a significant advantage in dealing with complex fine-grained classification tasks.
[0116] Comparison of comprehensive temperature measurement effects of the model:
[0117] Fig.12 The comparison of evaluation indicators of four models, Resnet152, ConvMixer, Mamba and Our model (ThermoSonoNet prediction model), at three granularities of Coarse, Medium and Fine is shown.
[0118] As can be seen from the figure, with the refinement of the granularity, the classification effect of some models has indeed deteriorated, but Our model can still maintain a high classification accuracy at various granularities; specifically, Resnet152 performs well at coarse granularity, but as the granularity increases, its classification effect gradually decreases, especially at fine granularity, the performance deteriorates significantly; the ConvMixer model initially performs well at coarse granularity, but as the granularity increases, its classification effect also gradually decreases, but the decrease is relatively small at fine granularity; the Mamba model has an average classification effect at coarse granularity, improves at medium granularity, but decreases again at fine granularity.
[0119] In contrast, Our model performs best at coarse granularity, and can still maintain high classification accuracy at medium and fine granularity, and is more stable than the other three models; especially at fine granularity, when the classification task is the most complex, the classification effect of Our model is still significantly better than other models; this shows that Our model has high robustness and adaptability in adapting to classification tasks of different granularity, and can maintain stable classification performance under various complexities.
[0120] Model ablation experiment
[0121] In order to verify the effectiveness of each module in the ThermoSonoNet prediction model, this study conducted a series of ablation experiments; in these experiments, this study gradually removed different modules in the model: Multi-scale Feature Perception Module (abbreviated as MFPM), SE module and Multiple Attention Encoder Module (abbreviated as MAEM), and observed the changes in model performance; the results of model ablation experiments at each granularity are shown in Tables 4-6.
[0122] Table 5 Coarse-grained model ablation experiment
[0123] ACC Precision Recall F1 Score Weighted AUC Full Model 99.98 99.54 99.98 99.96 99.92 w / o MFPM 99.01 99.08 99.01 99.21 99.90 w / o SE 99.25 99.23 99.25 99.47 99.91 w / o MAEM 98.56 99.16 98.56 99.35 99.90 w / o MAEM and MFPM 98.84 99.02 98.84 98.83 99.90 w / o MAEM and SE 98.74 98.97 98.74 98.84 99.89 w / o SE and MFPM 98.99 99.24 98.99 98.98 99.91 w / o MAEM, SE and MFPM 98.61 98.47 98.61 98.70 98.81
[0124] Table 6 Ablation experiment of medium-sized model
[0125] ACC Precision Recall F1 Score Weighted AUC Full Model 93.36 94.27 93.36 93.37 99.87 w / o MFPM 92.39 93.02 92.39 92.30 99.63 w / o SE 92.69 92.87 92.69 92.94 99.85 w / o MAEM 91.58 92.39 91.58 92.01 99.52 w / o MAEM and MFPM 91.73 91.76 91.73 92.38 99.50 w / o MAEM and SE 91.36 92.47 91.36 91.06 99.48 w / o SE and MFPM 91.98 92.53 91.98 91.82 99.54 w / o MAEM, SE and MFPM 90.82 91.58 90.82 90.78 99.42
[0126] Table 7 Fine-grained model ablation experiment
[0127] ACC Precision Recall F1 Score Weighted AUC Full Model 80.27 80.39 80.27 79.78 99.83 w / o MFPM 79.31 79.52 79.31 78.32 99.79 w / o SE 78.69 78.19 78.69 77.24 99.82 w / o MAEM 78.60 78.68 78.60 77.18 99.87 w / o MAEM and MFPM 77.78 76.93 77.78 76.23 99.72 w / o MAEM and SE 77.06 77.56 77.06 76.88 99.78 w / o SE and MFPM 76.65 76.80 76.65 75.27 99.77 w / o MAEM, SE and MFPM 75.39 78.20 75.39 76.06 99.76
[0128] In the ablation experiment of the coarse-grained model, we can clearly see the significant impact of each module on the model performance; our model (Our) performs well in all indicators, with a test accuracy of 99.98% and a near-perfect F1 score, which fully proves the rationality and effectiveness of the model design; when we remove the MFPM module, the test accuracy drops by nearly 1 percentage point, and Precision and Recall also decrease, which shows that the MFPM module plays an important role in improving the accuracy and recall of the model; the MFPM module may enhance the model's discriminative ability by capturing specific feature interactions or providing additional contextual information; after removing the SE module, the model performance It also decreased, but the magnitude was slightly smaller than when the MFPM module was removed, which indicates that the SE module also has a positive impact on the model performance. It may improve the generalization ability of the model by adaptively adjusting the importance of features. When the MAEM module is removed, the test accuracy drops to 98.56%, which is the largest drop among all single module removals. This shows that the MAEM module plays a core role in the coarse-grained model, and its powerful sequence modeling ability enables the model to better process input data and capture long-distance dependencies. Further, when the MAEM and MFPM or SE modules are removed simultaneously, the performance decreases more significantly, which shows that there is a synergy between these modules, and they complement each other to jointly improve the overall performance of the model.
[0129] The ablation experiment results of the medium-grained model show that the test accuracy of our model (Our) is 93.36%, which is lower than that of the coarse-grained model, but still maintains a high level; after removing the MFPM module, the test accuracy dropped by about 1 percentage point, and the Precision and Recall also decreased, which shows that the MFPM module is equally important in the medium-grained model, and its effect on improving the model performance has been verified on the medium-grained data; the impact of removing the SE module on the performance is slightly smaller, but still exists, which further proves the effectiveness of the SE module in the model; the removal of the MAEM module causes the test accuracy to drop sharply to 91.58%, which is the most significant performance drop in the medium-grained model, which once again proves the core position of the MAEM module in the model; when multiple modules are removed at the same time, the performance drop is more significant, especially when the MAEM, SE and MFPM modules are removed at the same time, the test accuracy drops to the lowest, which shows that these modules are interdependent in the medium-grained model and jointly maintain the high performance of the model. The synergy between them is crucial to the overall performance of the model.
[0130] The ablation experiment of the fine-grained model shows a similar trend to the coarse-grained and medium-grained models, but the overall performance level is lower. The test accuracy of our model is 80.27%, which is significantly lower than that of the first two granularity models. This may be because fine-grained data is more complex and difficult to capture, and the requirements for the model are higher. After removing the MFPM module, the test accuracy dropped by nearly 1 percentage point, and the Precision and Recall also decreased. This shows that the MFPM module also plays an important role in the fine-grained model. It may enhance the model's discriminative ability by capturing fine-grained feature interactions or providing more refined context information. Removing the SE module has a significant impact on performance. The impact is also quite obvious, with the test accuracy dropping to 78.69%, which further proves the effectiveness of the SE module in the fine-grained model; the removal of the MAEM module causes the test accuracy to drop sharply to 78.60%, which once again verifies the indispensability of the MAEM module in the model, and its powerful sequence modeling capability is essential for processing fine-grained data; when multiple modules are removed at the same time, the performance drops more dramatically, especially when the MAEM, SE and MFPM modules are removed at the same time, the test accuracy drops to a minimum of 75.39%, which shows that these modules collaborate with each other in the fine-grained model and jointly support the performance of the model, and the synergy between them has a crucial impact on the overall performance of the model.
[0131] Overall, no matter it is a coarse-grained, medium-grained or fine-grained model, each module makes an important contribution to the performance of the model, and the interaction between them also greatly affects the overall effect of the model; the MAEM module plays a core role in the models of the three granularities, and its powerful sequence modeling ability enables the model to better process input data and capture long-distance dependencies, thereby improving the performance and generalization ability of the model; the MFPM and SE modules also have a positive impact on the model performance, which may enhance the model's discrimination and generalization capabilities by capturing specific feature interactions, providing additional contextual information or adaptively adjusting the importance of features; when multiple modules are removed at the same time, the performance degradation is more significant, which shows that there is a synergy between these modules, they complement and collaborate with each other, and jointly improve the overall performance of the model.
[0132] Example:
[0133] Taking the use of the microwave hyperthermia ultrasound non-invasive temperature measurement system of the present application to treat breast cancer as an example, first place a B-type ultrasound machine, a microwave hyperthermia machine and a computer, connect the B-type ultrasound machine to the computer, and connect the microwave hyperthermia machine to the antenna; then fix the antenna and the ultrasound probe; the antenna and the ultrasound probe are perpendicular; then place the tumor site of the patient's breast in the central area of the antenna energy radiation and the area directly below the B-ultrasound probe to achieve precise treatment of the tumor area; then turn on the power switch of the microwave hyperthermia machine, set the time to 10 minutes, set the power to 6 watts, turn on the B-ultrasound device and adjust it to the small organ measurement mode, adjust the gain to achieve balance, open the computer microwave hyperthermia ultrasound non-invasive temperature measurement system interface; turn on the microwave hyperthermia machine working switch; dynamically adjust the power and time of the microwave hyperthermia machine according to the temperature value on the computer temperature measurement interface, and control the temperature of the hyperthermia area by controlling the microwave heating power, so as to achieve the maximum effect of treatment; after the treatment is completed, turn off the microwave hyperthermia machine and the B-ultrasound device.
[0134] The above-described embodiments of the present invention do not constitute a limitation on the protection scope of the present invention. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A microwave hyperthermia ultrasound non-invasive temperature measurement system based on the ThermoSonoNet model, characterized in that: include Microwave hyperthermia device; A waveguide antenna, the waveguide antenna is connected to the output end of the microwave hyperthermia device, the microwave hyperthermia device feeds energy to the waveguide antenna, and the waveguide antenna generates heat; Type B ultrasound equipment, which is used to collect ultrasound image videos; The computer is connected to the B-type ultrasound device and is used to receive ultrasound image video collected by the B-type ultrasound device and convert it into ultrasound image data; and input it into the ThermoSonoNet prediction model to perform temperature prediction.
2. According to the microwave hyperthermia ultrasound non-invasive temperature measurement system based on the ThermoSonoNet model of claim 1, it is characterized in that: The computer settings include: A conversion module, configured to convert the ultrasound image video into ultrasound image data to be detected; The pre-processing module is configured to perform median filtering on the ultrasonic image data to be detected to obtain filtered ultrasonic image data. A region of interest positioning module is configured to position the filtered ultrasound image data to obtain image data after the region of interest is positioned; A subtraction module, configured to perform a subtraction operation on the image data after the region of interest is located; obtaining subtraction image data; The temperature prediction module is configured to input the subtraction image data into the ThermoSonoNet prediction model to obtain the predicted temperature.
3. According to the microwave hyperthermia ultrasound non-invasive temperature measurement system based on the ThermoSonoNet model of claim 1, it is characterized in that: The ThermoSonoNet prediction model includes a multi-scale feature perception module, an enhanced visual state space module, a multi-head attention encoding module and a classification output module; the multi-scale feature perception module is used to extract multi-scale features through convolution kernels of different sizes, and use adaptive average pooling, linear transformation and Sigmoid activation function to generate fusion weights to achieve multi-scale feature fusion; The enhanced visual state space module is used to extract local detail features and global features respectively through the enhanced convolution branch and the visual state space branch, and use the modulation interaction feature aggregation module for deep fusion to enhance the model's ability to represent complex temperature features; the multi-head attention encoding module is used to reorganize and fuse features from three dimensions: channel, space, and scale through a parallel multi-head attention mechanism and a multi-layer perceptron to improve feature expression capabilities, and retain the original feature information through residual connections; The classification output module is used to map high-dimensional features to temperature labels through global average pooling and fully connected layers to achieve temperature classification output.
4. The microwave hyperthermia ultrasound non-invasive temperature measurement system based on the ThermoSonoNet model according to claim 3 is characterized in that: The multi-scale feature perception module includes a feature segmentation layer (Split), the output end of the feature segmentation layer (Split) is respectively connected to the input end of the 3×3 convolution layer and the input end of the 5×5 convolution layer; the output ends of the 3×3 convolution layer and the 5×5 convolution layer are sequentially connected to the batch normalization layer (BN) and the ReLU activation function; the output of the ReLU activation function is sequentially connected to the adaptive average pooling layer (AdaptiveAVgPool2d) and the fully connected layer (Linear) through element addition operation; the output of the fully connected layer (Linear) is processed by the Sigmoid function and the Reshape operation; and the final output is obtained by element-by-element multiplication with the output of the ReLU activation function, and output to the enhanced visual state space module.
5. The microwave hyperthermia ultrasound non-invasive temperature measurement system based on the ThermoSonoNet model according to claim 3, characterized in that: The enhanced visual state space module includes a feature segmentation layer (Split), the input end of the feature segmentation layer (Split) is connected to the output end of the multi-scale feature perception module; the output end of the feature segmentation layer (Split) is respectively connected to the enhanced convolution branch and the visual state space branch; the enhanced convolution branch includes a depth separable convolution layer (DW Conv), a batch normalization layer (BN), a ReLU activation function, a squeeze excitation module (SE Block), a point convolution layer (PW Conv), a batch normalization layer (BN), a ReLU activation function, and a squeeze excitation module (SE Block), so as to extract local detail features; the visual state space branch includes a layer normalization layer (LN), an adaptive residual visual state space module (ARVSS Block), and a channel shuffle layer (Channel shuffle); used to extract global features; The outputs of the enhanced convolution branch and the visual state space branch are deeply fused through the modulated interactive feature aggregation module (MIFA) to enhance the model's ability to represent complex temperature features; and the outputs are sent to the multi-head attention encoding module.
6. The microwave hyperthermia ultrasound non-invasive temperature measurement system based on the ThermoSonoNet model according to claim 3, characterized in that: The multi-head attention encoding module includes a dimensionality transformation layer (Dim Transform), the output end of the dimensionality transformation layer (Dim Transform) is connected to three normalization layers (Layer Norm) respectively; the output ends of the three normalization layers (LayerNorm) are connected to the multi-head attention layer (Multi-head Attention); the output end of the multi-head attention layer (Multi-headAttention) is connected to the input end of the random dropout layer (Dropout), the output of the random dropout layer (Dropout) and the output of the dimensionality transformation layer are element-wise added and input to the layer normalization layer (Layer Norm); the output end of the layer normalization layer (Layer Norm) is sequentially connected to a multi-layer perceptron module (MLP) and a random dropout layer (Dropout), the output of the random dropout layer (Dropout) and the output of the dimensionality transformation layer are again element-wise added and finally output to the classification output module.
7. The microwave hyperthermia ultrasound non-invasive temperature measurement system based on the ThermoSonoNet model according to claim 6, characterized in that: The multi-layer perceptron module (MLP) includes a fully connected layer (Linear), a Gaussian error linear unit activation function layer (GeLU), a random dropout layer (Dropout), a fully connected layer (Linear), and a random dropout layer (Dropout).
8. The microwave hyperthermia ultrasound non-invasive temperature measurement system based on the ThermoSonoNet model according to claim 3, characterized in that: The classification output module includes an adaptive average pooling layer (AdaptiveAVgPool2d) and a fully connected layer (Linear); the global average pooling and fully connected layer classification output module is used to map high-dimensional features to temperature labels to achieve temperature classification output.