Thyroid nodule ultrasound image segmentation method fusing factorization VMama and characteristic frequency band separation
Through the method of fusion factorization of VMamba and feature band separation, combined with feature extraction and fusion of FMVSS and DFFT modules, and the edge optimization strategy based on Laplacian, the problems of boundary blurring and feature loss in ultrasonic image segmentation of thyroid nodules are solved, and efficient and accurate image segmentation is achieved.
Patent Information
- Application Number
- CN202510299557.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art has blurred boundary, loss of features and high requirements for image quality in ultrasonic image segmentation of thyroid nodules, making it difficult to achieve efficient segmentation, especially in large-scale image processing.
The segmentation method of fusion factorization VMamba and feature band separation is adopted, multi-level features are extracted through the FMVSS module, and the DFFT module is fusion across domains, and combined with the edge optimization strategy based on the Laplacian operator and the new edge loss function to improve the network segmentation accuracy of the nodule boundaries and detailed areas.
It significantly improves the segmentation accuracy of ultrasound images of thyroid nodules, improves the performance of key segmentation indicators such as IoU and DSC, and provides an efficient and accurate ultrasound image segmentation solution.
Smart Images

Figure CN120235899A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision-assisted medicine, and particularly to a method for segmenting thyroid nodule ultrasound images by fusing factorized VMamba and feature band separation. Background Art
[0002] High-frequency ultrasound detection (HFU), as the preferred detection method for thyroid nodule lesions, can effectively detect up to 50%-68% of thyroid nodules. However, when ultrasonic signals are converted into images, random speckles and artifacts are often generated, resulting in problems such as low-quality, high-noise, and unclear feature contrast in ultrasound images. In addition, nodules themselves have large morphological differences and blurred tissue boundaries. Especially for early nodules, due to their unclear features, it is difficult to identify them with the naked eye, thus increasing the clinical diagnosis difficulty and the risk of misdiagnosis.
[0003] To solve the above problems, Zhou Chunyu et al. proposed a method for segmenting thyroid ultrasound images based on the T-snake model. By introducing adaptive median filtering, speckle noise was effectively removed, and effective extraction of weak edge feature information was achieved. Salles et al. developed a semi-automatic segmentation-assisted diagnostic software, which used the FloodFill algorithm to accurately locate and segment the lesion area in ultrasound images. The above methods can achieve good results in small-scale image processing, but there are still limitations in large-scale image processing, mainly manifested as a high dependence on parameters, high requirements for image quality, and lack of cross-modal applicability, making it difficult to achieve efficient segmentation. Summary of the Invention
[0004] Objective of the present invention: To construct a more refined segmentation model for thyroid nodules in ultrasound images to improve the accuracy of thyroid nodule segmentation and at the same time solve the problems of blurred image boundaries or feature loss in existing models.
[0005] To implement the method described above, the present invention provides a method for segmenting thyroid nodule ultrasound images by fusing factorized VMamba and feature band separation, which includes the following steps:
[0006] Step1: Preprocess the thyroid nodule ultrasound image dataset;
[0007] Step2: Input the preprocessed image into the encoder of the segmentation model FMVM-DFFT for processing: Use multiple FMVSS modules to extract image features at multiple levels, and embed a patch merging module between each layer. During downsampling, feature maps with different resolutions and semantic levels are composed according to the extracted features;
[0008] Step 3: Input the extracted feature maps into the DFFT module. Through the "spatial domain → frequency domain → spatial domain" transformation, cross-domain fusion of the feature maps is performed to extract and integrate the feature information in different domains.
[0009] Step 4: Input the extracted feature information into the decoder of the segmentation model FMVM-DFFT for processing: Use multiple FMVSS modules to generate and fuse feature maps based on feature information at different levels, and embed patch expansion modules between each layer. During the upsampling process, gradually restore the low-resolution feature maps to high-resolution mask images.
[0010] Step 5: Apply an edge optimization strategy based on the Laplacian operator during the training process of inputting training set images into the FMVM-DFFT model for Steps 2 to 4. Calculate the difference between the predicted edge and the actual edge, and act on the next round of the training process in the form of a loss function.
[0011] Step 6: After the training phase ends, retain the weight combinations with the optimal prediction performance in different rounds in Step 5.
[0012] Step 7: Perform Steps 2 to 4 on the best weights in Step 6 to obtain the processed lesion mask image.
[0013] Furthermore, the preprocessing of the thyroid nodule ultrasound image dataset in Step 1 includes: uniformly processing the original images in the dataset to a size of 256×256 pixels; dividing the training set and the test set according to a ratio of 7:3; performing data augmentation operations on the images in the training set by randomly rotating and flipping, and annotating the images.
[0014] Furthermore, the backbone network of the segmentation model FMVM-DFFT in Step 2 is VMamba, and the FMVSS module is a VSS module that combines a factorization machine and external attention.
[0015] Furthermore, the encoder of the segmentation model FMVM-DFFT in Step 2 consists of: an image patch embedding layer → a processing layer composed of two FMVSS modules and a patch merging module → a processing layer composed of two FMVSS modules and a patch merging module → a processing layer composed of three FMVSS modules and a patch merging module → a processing layer composed of two FMVSS modules, which is used to extract hierarchical features of the image data features and capture the global and local information of the image.
[0016] Furthermore, the DFFT module in Step 3 adopts dual-branch fast Fourier transform. First, the spatial domain features are transformed to the frequency domain and decomposed into low-frequency and high-frequency components. Subsequently, the two components are sequentially optimized by pointwise convolution with a size of 1×1 and a channel attention module with a channel reduction ratio of 16 to enhance the feature representation. After recombination in the complex domain, they are restored to the spatial domain through inverse Fourier transform. Finally, they are merged with the retained input features after dynamic convolution with a size of 3×3 and layer normalization processing.
[0017] Furthermore, the decoder of the segmentation model FMVM-DFFT in Step 4 consists of: a processing layer composed of two FMVSS modules → a processing layer composed of a patch expansion module and two FMVSS modules → a processing layer composed of a patch expansion module and one FMVSS module → a final projection layer, which is used to restore the low-resolution feature map to the size and semantic space matching the input data.
[0018] Furthermore, the edge optimization strategy in Step 5 is to use Laplacian to detect the edges of the smoothed image, extract its edge features through convolution operations, and obtain the predicted edge feature map and the actual edge feature map. The two designed edge loss functions and the total loss function are EdgeLoss and BDELoss respectively. Among them, EdgeLoss is the difference between the predicted edge and the actual edge calculated by the absolute error loss of the predicted and actual edge feature maps, and BDELoss is obtained by weighted summation of EdgeLoss and the initial loss function BCEDiceLoss according to the weight.
[0019] Furthermore, the optimal prediction performance weight in Step 6 is determined by comparing the loss values on the validation set, that is, when the validation loss of a certain training epoch is less than the historical minimum loss, the model weight corresponding to this training epoch is the optimal prediction weight.
[0020] Advantages of the present invention: Based on the VMamba architecture, a method for segmenting thyroid nodule ultrasound images that combines factorized VMamba and feature frequency band separation is proposed. By designing the FMVSS module, it can efficiently extract feature information in ultrasound images in multiple dimensions, and adaptively adjust the weights of the fused features, enhancing the ability to capture key information and local details; by designing the DFFT module, the high-frequency and low-frequency features of the image are separated and dynamically decoded, enhancing the model's frequency domain perception ability, and combining the channel attention mechanism to optimize feature selection and fusion, better capturing the detailed information in the image; by designing an edge optimization strategy based on the Laplacian operator, the segmentation accuracy of the network for nodule boundaries and detailed areas is effectively improved. Experimental results show that this method performs well on the key segmentation metrics IoU and DSC, and each sub-module can effectively improve the overall performance, providing an efficient and accurate solution for the ultrasound image segmentation task. Description of the Drawings
[0021] Figure 1 Overall method step diagram;
[0022] Figure 2 Architecture diagram of the segmentation model FMVM-DFFT;
[0023] Figure 3 Architecture diagram of the FMVSS module;
[0024] Figure 4 Architecture diagram of the DFFT module;
[0025] Figure 5 Flowchart of the edge optimization strategy;
[0026] Figure 6 Comparison diagram of using FMVM-DFFT and other segmentation methods on the TN3K dataset;
[0027] Figure 7 Comparison diagram of using FMVM-DFFT and other segmentation methods on the DDTI dataset. Detailed Implementation Manner
[0028] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings.
[0029] The following are the English abbreviation explanations in the present invention: VMamba: Visual State Space Model; VSS module: Visual State Space Module; FMVSS module: Factorized Visual State Space Module; DFFT module: Dual-branch Frequency Band Separation Module; SS2D module: 2D Selective Scanning Module; FFT: Fast Fourier Transform; IFFT: Inverse Fast Fourier Transform; BCEDiceLoss: Loss function combining Dice loss and standard binary cross-entropy; EdgeLoss: Edge loss function; BDELoss: Loss function combining BCEDiceLoss and EdgeLoss; ACC: Accuracy; Recall: Recall rate; Specificity: Specificity; DSC: Similarity coefficient; IoU: Intersection over Union; AUC: Area Under the Curve; GT: Segmentation standard; Relu activation function: Rectified Linear Unit; Sigmoid activation function: Sigmoid function; epoch: Training epoch.
[0030] The present invention uses the latest Visual State Space Model VMamba to propose a thyroid nodule ultrasound image segmentation model FMVM-DFFT that combines factorized VMamba and feature frequency band separation to complete the precise segmentation task of thyroid ultrasound image nodules. The original ultrasound image data is input into the model after preprocessing steps such as label construction, image scaling, and data augmentation. The model combines a factorization machine and external attention to propose a factorized variant FMVSS of VSS, which efficiently extracts information in different dimensions of the input features and adaptively adjusts the feature weights to enhance the ability to capture key information and local details; a DFFT module containing dual-branch fast Fourier transform and dynamic convolution is proposed to perform frequency band dynamic separation and fine extraction of the feature map output by the encoder, thereby improving the network's ability to capture details and macroscopic information, and channel attention is used to adaptively control the weights of each channel; an edge optimization strategy based on the Laplacian operator and a new edge loss function BDELoss are proposed and applied to the training process to further enhance the network's learning ability for the image edge region. After the training process ends, the best weights of the model retained during the training process are used to process the test set images, and the segmented lesion mask images are output to the specified area.
[0031] Generally speaking, this process can be as shown in Figure 1 shown below:
[0032] Step1: Preprocess the thyroid nodule ultrasound image dataset;
[0033] Step 2: Input the preprocessed image into the encoder of the segmentation model FMVM-DFFT for processing: Use multiple FMVSS modules to extract image features at multiple levels, embed a patch merging module between each layer, and form feature maps with different resolutions and semantic levels according to the extracted features during downsampling;
[0034] Step 3: Input the extracted feature maps into the DFFT module, and through the "spatial domain → frequency domain → spatial domain" transformation, perform cross-domain fusion on the feature maps, and extract and integrate the feature information in different domains;
[0035] Step 4: Input the extracted feature information into the decoder of the segmentation model FMVM-DFFT for processing: Use multiple FMVSS modules to generate and fuse feature maps according to different levels of feature information, embed a patch expansion module between each layer, and gradually restore the low-resolution feature maps to high-resolution mask images during upsampling;
[0036] Step 5: Apply an edge optimization strategy based on the Laplacian operator during the training process of inputting the training set images into the FMVM-DFFT model for Steps 2 to 4, calculate the difference between the predicted edge and the actual edge, and act on the next round of training process in the form of a loss function.
[0037] Step 6: After the training phase ends, retain the weight combinations with the optimal prediction performance in different rounds in Step 5.
[0038] Step 7: Perform Steps 2 to 4 on the best weights in Step 6 to obtain the processed lesion mask image.
[0039] As Figure 2 shown, after the data is input into the encoder of the segmentation model FMVM-DFFT, the embedded FMVSS module performs multi-level feature extraction on the input image, captures the global and local information of the image, and generates feature maps with different resolutions and semantic levels.
[0040] As Figure 2 shown, after the data is input into the decoder of the segmentation model FMVM-DFFT, the embedded FMVSS module further refines and fuses the feature information at different frequencies, thereby enhancing the utilization of effective information, and gradually restoring the low-resolution feature maps to high-resolution mask images.
[0041] As Figure 3As shown, the design of FMVSS fully compensates for the deficiencies of VSS in local feature capture and feature interaction capabilities. By adopting a factorization branch containing a factorization machine processing layer, it aims to model the complex interaction relationships between features dimension by dimension. The factorization machine processing layer first maps high-dimensional features to a low-dimensional space to reduce computational complexity; subsequently, it captures second-order interaction relationships through the inner product of latent vectors; finally, it expands the interaction features back to the high-dimensional space to fuse global information, thereby enhancing the feature expression ability and performance of FMVSS. Its core formula is:
[0042]
[0043] where, ω0 represents the global bias, ω i represents the weight of the i-th feature, v i , v j represents the low-dimensional embedding vector of features x i , x j .
[0044] As Figure 3 shown, in FMVSS, the input feature F in ∈R H×W×C is sent to two parallel processing branches after layer normalization: in one branch, the feature expands the number of channels to αC through a linear layer, and then enters the unique SS2D processing module in VMamba after dynamic convolution processing. SS2D adopts a unique four-way scanning strategy and a selective scanning spatial state sequence model, which can efficiently process visual signals and extract features, and obtains feature X1 through normalization; in the second branch, the input feature first passes through the factorization machine processing layer, gradually captures the complex interactions between features through layer-by-layer processing of low-dimensional, medium-dimensional, and high-dimensional, enhances the local feature modeling ability, and then dynamically adjusts the feature weights through an external attention module to further improve the global modeling ability of key features, obtaining feature X2. Finally, X1 and X2 are aggregated through the Hadamard product, and the number of channels is projected back to the original dimension C through a linear layer to generate an output X final with the same shape as the input. This dual-path design combines the global dependence modeling ability of SS2D and the feature interaction advantages of factorization, significantly improving the efficiency and adaptability of the model.
[0045] As Figure 4 shown, the DFFT module focuses on the exploration of feature enhancement mechanisms, realizing the dynamic decomposition and branch processing of high- and low-frequency features. DFFT transforms the input feature F in ∈R H×W×C from the spatial domain to the frequency domain through FFT, so that F in is decomposed into a high-frequency component F high and a low-frequency component F low , and the transformation formula is as follows:
[0046]
[0047] Among them, F represents a feature, (x, y) and (u, v) represent the spatial domain coordinates and the frequency domain coordinates respectively, H and W represent the height and width belonging to the image, (u c , v c ) represents the coordinates of the frequency domain center, R max represents the maximum frequency radius, η represents the proportionality coefficient, and j represents the imaginary unit.
[0048] As Figure 4 shown, the two decomposed components F high and F low enter independent branches respectively and sequentially pass through the pointwise convolution (PWConv) with a size of 1×1 and the channel attention (CA) module with a reduction ratio of 16, and finally obtain the optimized low-frequency feature F low-opt and the optimized high-frequency feature F high-opt . In this process: 1) The low-frequency branch focuses on the macroscopic structure and background information in the image, captures the global semantic features, and strengthens the global context information; 2) The high-frequency branch focuses on the detailed features in the image, captures the local details and edge features, so as to make the representation of the key features more accurate; 3) The channel attention realizes the adaptive adjustment of the channel weights by exploring the mutual dependence of the features between channels. Subsequently, the optimized features F high-opt and F low-opt obtained from the two branches are recombined in the complex domain and restored to the spatial domain through IFFT. The formula for this transformation process is as follows:
[0049] F opt (u, v) = F high-opt (u, v) + j·F low-opt (u, v)
[0050]
[0051] Next, the transformed feature F out is processed by dynamic convolution, which enhances the model's adaptive multi-scale feature extraction ability for spatial domain information and further optimizes the spatial relationship. Finally, after the processed output feature is stabilized by layer normalization, it enters the feed-forward network to achieve deep fusion and non-linear transformation, and the transformed feature is merged with the retained input feature into a feature map and transmitted to the decoder part.
[0052] As Figure 5As shown in the figure, to solve the problem of inaccurate and unrefined edge segmentation, the present invention proposes an edge optimization strategy based on the Laplacian operator and applies it to network training, and designs a new edge loss function BDELoss. During the training process, the specific steps of the edge optimization strategy are as follows: 1) First, perform average pooling on the target ultrasonic image to smooth the image details to optimize the feature extraction effect of subsequent algorithms; 2) Use the Laplacian to detect the edges of the smoothed image, and use convolution operations to extract its edge features to obtain a predicted edge feature map and an actual edge feature map; 3) Use the L1 loss function to calculate the difference between the predicted edge and the actual edge to obtain the edge loss EdgeLoss; 4) Weightedly sum the initial loss function BCEDiceLoss and EdgeLoss according to a certain weight to obtain the total loss BDELoss, which is used to optimize the edge features and overall prediction performance of the model; 5) Finally, the BDELoss calculated during the training process will be used as a loss function to measure the difference between the prediction and the actual value, and further act on the training process, so that the model can focus more on the learning and expression of edge features while optimizing the global features.
[0053] The optimal weight combination in the training stage is determined by comparing the loss values on the validation set. When the validation loss of a certain epoch is less than the recorded minimum loss, the model weights corresponding to this epoch are the optimal weights. Save the model weights at this time as "best.pth", and update the minimum loss value and the corresponding epoch. At the same time, save "latest.pth" containing the states of the model, optimizer, scheduler, etc. at the end of each epoch. Finally, after the training ends, select the "best.pth" weights for testing. After the test stage ends, output the segmented lesion mask image to the specified area.
[0054] To illustrate the effectiveness of the present invention in implementing the above solutions, the following will be described by means of experiments.
[0055] 1. Introduction to evaluation metrics
[0056] To comprehensively evaluate the segmentation performance of the model, this study uses 6 evaluation metrics commonly used in image segmentation tasks: ACC, Recall, Specificity, DSC, IoU, and the area under the curve AUC. The calculation formulas for each metric are as follows:
[0057]
[0058] 2. Ablation experiment to verify the effectiveness of the module
[0059] To evaluate the effectiveness of different components of FMVM-DFFT in improving the overall performance, the present invention conducts ablation experiments on two datasets, TN3K and DDTI. Tables 1-2 show the segmentation performance results of models (v0-v4) with different components on four evaluation metrics, and the inclusion of each component is indicated by "√".
[0060] Table 1 Results of Ablation Experiments on TN3K
[0061]
[0062] Table 2 Results of Ablation Experiments on DDTI
[0063]
[0064] In the ablation experiment, the baseline network is defined as v0. Subsequently, FMVSS, edge optimization, DFFT, and channel attention are successively added to the baseline model, respectively deriving v1, v2, v3, and the proposed v4 model. To comprehensively evaluate the model performance from different perspectives, four widely used metrics in ablation experiments are adopted in this experiment: ACC, DSC, IoU, and AUC. Among them, ACC measures the overall accuracy, DSC and IoU focus on evaluating the segmentation quality, while AUC evaluates the comprehensive performance of the segmentation model at different classification thresholds.
[0065] The analysis of the experimental results in Tables 1-2 shows that with the introduction of different methods, although there are fluctuations in a few metrics, the overall performance of the model shows a steady upward trend. The proposed model reaches the optimal value in all evaluation metrics, and the improvement is most obvious in the DSC and IoU metrics, fully verifying the positive impact of the comprehensive application of multiple methods on the performance of the proposed model.
[0066] 3. Comparative Experiments to Verify the Superiority of the Model
[0067] To verify the superiority of FMVM-DFFT in thyroid nodule segmentation tasks, this experiment compares it with commonly used segmentation models, U-Net, FCN, Segnet, SGUnet, TransUnet, TRFE+, BPAT-UNet, and VM-Unet on six evaluation metrics on the TN3K and DDTI datasets. The quantitative analysis results obtained by different segmentation networks are shown in Tables 3 and 4.
[0068] The experimental results show that for all metrics of FMVM-DFFT on the TN3K and DDTI datasets, except that the Recall of DDTI is slightly lower than that of TRFE+, the others are better than other segmentation models. In particular, it performs excellently in the important metrics DSC and IoU that determine segmentation performance, and also shows an obvious advantage in AUC. In the experimental analysis, it is found that there are obvious differences between FMVM-DFFT and other comparative segmentation models in the three metrics of DSC, IoU and Recall, which determine the accuracy and comprehensiveness of the segmentation results.
[0069] Table 3 Performance comparison of different segmentation models on TN3K
[0070]
[0071] Table 4 Performance comparison of different segmentation models on DDTI
[0072]
[0073] As Figure 6 and Figure 7 shown, these two figures show the effects of FMVM-DFFT and the comparative segmentation models in the image segmentation task, and compare them with the standard. It can be seen from the figures that when processing large and small nodules and double-node images, FMVM-DFFT not only exceeds other comparative networks in the segmentation accuracy of the overall region, but also shows stronger recognition ability in details. Especially in the segmentation of nodule boundaries, FMVM-DFFT can more accurately distinguish nodules from the background, effectively avoiding the problems of blurred or lost edges, and is closest to the standard. In contrast, the segmentation results of other comparative models have problems of inaccurate edge segmentation and target position deviation, and this error is particularly obvious in the processing of small nodules and double-nodes.
[0074] In addition, FMVM-DFFT is compared with the current state-of-the-art model (SOTA) on TN3K and DDTI. Table 5 shows the performance of FMVM-DFFT and SOTA in the key metrics DSC and IoU that measure segmentation quality. The results show that FMVM-DFFT shows better performance compared with SOTA.
[0075] Table 5 Performance comparison of FMVM-DFFT and SOTA
[0076]
[0077] The above are only specific implementation manners of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A thyroid nodule ultrasound image segmentation method integrating factorized VMamba and characteristic frequency band separation, the method comprising the following steps: Step 1: Preprocessing of thyroid nodule ultrasound image dataset; Step 2: Input the preprocessed image into the encoder of the segmentation model FMVM-DFFT for processing: use multiple FMVSS modules to extract image features at multiple levels, and embed a patch merging module between each layer. In the downsampling process, feature maps with different resolutions and semantic levels are formed according to the extracted features; Step 3: Input the extracted feature map into the DFFT module, perform cross-domain fusion on the feature map by transforming from "spatial domain → frequency domain → spatial domain", and extract and integrate feature information in different domains; Step 4: Input the extracted feature information into the decoder of the segmentation model FMVM-DFFT for processing: Use multiple FMVSS modules to generate and fuse feature maps based on feature information at different levels, and embed patch expansion modules between each layer to gradually restore the low-resolution feature map to a high-resolution mask image during the upsampling process; Step 5: Apply the edge optimization strategy based on the Laplacian operator during the training process of the training set image input FMVM-DFFT model from Step 2 to Step 4, calculate the difference between the predicted edge and the actual edge, and use it in the form of a loss function for the next round of training; Step 6: After the training phase, retain the weight combination with the best prediction performance in different rounds in Step 5; Step 7: Process the best weight in Step 6 from Step 2 to Step 4 to obtain the processed lesion mask image.
2. The method for segmenting thyroid nodules ultrasound images by fusing factorized VMamba and characteristic frequency band separation according to claim 1, characterized in that: The preprocessing of the thyroid nodule ultrasound image dataset in Step 1 includes: uniformly resizing the original images in the dataset to 256×256 pixels; dividing the training set and the test set according to a 7:3 ratio; performing data enhancement operations on the images in the training set by randomly rotating and flipping, and annotating the images.
3. The method for segmenting thyroid nodules ultrasound images by fusing factorized VMamba and characteristic frequency band separation according to claim 1, characterized in that: The backbone network of the segmentation model FMVM-DFFT in Step 2 is VMamba, and the FMVSS module is a VSS module that integrates the factorization machine and external attention.
4. The method for segmenting thyroid nodules ultrasound images by fusing factorized VMamba and characteristic frequency band separation according to claim 1, characterized in that: The encoder composition of the segmentation model FMVM-DFFT in the Step 2 is as follows: an image block embedding layer → a processing layer consisting of two FMVSS modules and a patch merging module → a processing layer consisting of two FMVSS modules and a patch merging module → a processing layer consisting of three FMVSS modules and a patch merging module → a processing layer consisting of two FMVSS modules, which is used to extract features of image data layer by layer to capture global and local information of the image.
5. The method for segmenting thyroid nodules ultrasound images by fusing factorized VMamba and characteristic frequency band separation according to claim 1, characterized in that: The DFFT module in Step 3 adopts a dual-branch fast Fourier transform to first transform the spatial domain features into the frequency domain and decompose them into low-frequency and high-frequency components; Subsequently, the two components are optimized in turn by a point-by-point convolution of size 1×1 and a channel attention module with a channel reduction ratio of 16 to enhance the feature representation, and after reorganization in the complex domain, they are restored to the spatial domain by inverse Fourier transform; finally, they are merged with the retained input features by a dynamic convolution of size 3×3 and layer normalization.
6. The method for segmenting thyroid nodules ultrasound images by fusing factorized VMamba and characteristic frequency band separation according to claim 1, characterized in that: The decoder of the segmentation model FMVM-DFFT in the Step 4 is composed of: a processing layer consisting of two FMVSS modules → a processing layer consisting of a patch expansion module and two FMVSS modules → a processing layer consisting of a patch expansion module and a FMVSS module → a final projection layer, which is used to restore the low-resolution feature map to a size and semantic space that matches the input data.
7. The method for segmenting thyroid nodules ultrasound images by fusing factorized VMamba and characteristic frequency band separation according to claim 1, characterized in that: The edge optimization strategy in Step 5 is to use Laplacian to detect the edge of the smoothed image and use convolution operation to extract its edge features. The predicted edge feature map and the actual edge feature map are obtained; the two designed edge loss functions and the total loss function are EdgeLoss and BDELoss, respectively, where EdgeLoss is the difference between the predicted edge and the actual edge calculated by the absolute error loss of the predicted and actual edge feature maps, and BDELoss is obtained by weighted summing EdgeLoss and the initial loss function BCEDiceLoss.
8. The method for segmenting thyroid nodules ultrasound images by fusing factorized VMamba and characteristic frequency band separation according to claim 1, characterized in that: The optimal prediction performance weight in Step 6 is determined by comparing the loss value on the validation set, that is, when the validation loss of a certain training round is less than the historical minimum loss, the model weight corresponding to the training round is the optimal prediction weight.
Citation Information
Cited By
Method and system for predicting residual life of equipment component
CN121412596A
Ultrasonic video target detection method, target detection model training method and electronic equipment
CN121527081A