Multi-scale frequency fusion medical image segmentation system and method based on wavelet transform

Through wavelet transform and multi-scale frequency fusion methods, the deficiencies of feature extraction and fusion in medical image segmentation are solved, and high-precision, low-complexity lesion area segmentation is achieved. It is suitable for diversified medical image segmentation tasks, especially in the electronic health metaverse scenario.

CN120673062APending Publication Date: 2025-09-19QINGDAO UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510750191.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing medical image segmentation methods have insufficient feature extraction capabilities and feature fusion mismatch when dealing with scale diversity, shape irregularities, and differences between images of different modalities, making it difficult to maintain high accuracy and low complexity in diverse clinical applications.

Method used

A multi-scale frequency fusion method based on wavelet transform is adopted. Medical images are decomposed into multi-frequency sub-band feature maps through discrete wavelet transform. Frequency-weighted attention fusion, frequency-aware residual fusion and similar feature alignment modules are combined to achieve adaptive feature extraction and fusion.

Benefits of technology

It significantly improves the accuracy and robustness of complex lesion area segmentation, controls computational complexity, and is suitable for diverse medical image segmentation tasks, especially in electronic health metaverse scenarios with high real-time and interactivity requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673062A_ABST
    Figure CN120673062A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale frequency fusion medical image segmentation method based on wavelet transform, and belongs to the technical field of medical image processing and artificial intelligence. In order to solve the technical problems that in the prior art, precision, detail recovery and efficiency are insufficient in complex medical image segmentation, and multi-scale features are difficult to effectively fuse, the system integrates discrete wavelet transform, frequency weighted attention fusion, frequency perception residual fusion and similar feature alignment units; according to the method, multi-scale frequency decomposition is carried out on an input image, adaptive weighted fusion is carried out on different frequency band features, residual enhancement is carried out in combination with high-resolution original image information, and accurate alignment processing is carried out on multi-scale features, so that high-precision and high-efficiency segmentation of a focus area is realized. The technical scheme of the invention is mainly used for improving the accuracy and robustness of medical image segmentation, and is suitable for an electronic health element universe scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical image processing and artificial intelligence, and specifically relates to a medical image segmentation system that uses wavelet transform to extract and fuse multi-scale frequency features, as well as an efficient and high-precision medical image segmentation method implemented on the system. The method is particularly suitable for accurately identifying and segmenting lesion areas with complex textures and fuzzy boundaries, and has the potential for application in scenarios such as the electronic health metaverse. Background Art

[0002] Medical image segmentation is a key technology in clinical applications such as computer-aided diagnosis, disease monitoring, and surgical planning. Its goal is to accurately identify and delineate anatomical structures or pathological regions of interest from medical images (such as CT, MRI, ultrasound, and X-rays). Traditional medical image segmentation relies primarily on manual labor by clinicians, a time-consuming and inefficient process. Furthermore, segmentation results are susceptible to subjective experience and fatigue, making consistency and reproducibility difficult to ensure.

[0003] With the rapid development of deep learning technology, automatic segmentation methods based on convolutional neural networks (CNNs), especially encoder-decoder architectures represented by U-Net and its various improved networks, have achieved remarkable success in the field of medical image segmentation. These methods extract contextual semantic information from images through step-by-step downsampling and combine shallow high-resolution features with deep semantic features through skip connections, aiming to restore fine segmentation boundaries during the decoding stage. However, CNN-based methods typically use fixed-size convolution kernels, which have limited feature extraction capabilities and adaptability when dealing with the scale diversity, shape irregularities, and significant differences between images of different modalities that are common in medical images. In addition, the common skip connection method directly transfers the high-resolution features of the encoder to the decoder, which may lead to semantic mismatches between features at different levels, affecting the effectiveness of information fusion.

[0004] In recent years, the Transformer architecture has been introduced to the field of medical image segmentation due to its powerful global dependency modeling capabilities, demonstrating promising results. Through its self-attention mechanism, the Transformer can capture long-range pixel-to-pixel relationships within an image, enabling a better understanding of global context. However, Transformer models typically lack the precision of CNNs in capturing local details, and their computational and parameter-intensive nature places high demands on hardware resources, making them challenging to use in clinical scenarios requiring real-time processing or deployment on resource-constrained devices.

[0005] To overcome the shortcomings of the aforementioned single architecture, researchers have begun exploring hybrid models of CNNs and Transformers, as well as various complex network structures designed to enhance multi-scale feature expression, optimize feature fusion strategies, and improve model generalization capabilities. Despite this, existing technologies still have room for improvement in the following areas: 1) How to more effectively separate and utilize information at different scales from medical images to simultaneously account for structural contours and texture details; 2) How to design more intelligent feature fusion mechanisms to adaptively integrate features from different paths, different levels, and different modalities, while suppressing irrelevant or redundant information; and 3) How to control the complexity of the model while ensuring high segmentation accuracy, making it suitable for diverse clinical applications, including emerging e-health and metaverse medical platforms.

[0006] Therefore, developing a segmentation method that can efficiently extract and fuse multi-scale frequency features in medical images while maintaining accurate perception of local details and global structures is of great significance for improving the accuracy and practicality of medical image analysis. Summary of the Invention

[0007] In order to overcome the above problems existing in the prior art, the present invention provides a multi-scale frequency fusion medical image segmentation system based on wavelet transform and a corresponding segmentation method.

[0008] The technical solution adopted by the present invention to solve the technical problem is: a multi-scale frequency fusion medical image segmentation method based on wavelet transform, comprising the following steps:

[0009] Step 1: Perform discrete wavelet transform operation on the input medical image to decompose it into a low-frequency approximate sub-band feature map and three high-frequency detail sub-band feature maps;

[0010] Step 2: Input the low-frequency approximate sub-band feature map and the high-frequency detail sub-band feature map obtained in step 1 into the frequency-weighted attention fusion module, and perform adaptive weighted fusion processing through the frequency-weighted attention fusion module to obtain a first fused feature map;

[0011] Step 3: Inputting the first fusion feature map obtained in step 2 or the combined feature of the first fusion feature map and the input medical image into a multi-stage encoder network, and obtaining the encoding feature map output by the multi-stage encoder network at each corresponding level and the final deep encoding feature map through step-by-step downsampling and feature extraction operations of the multi-stage encoder network;

[0012] Step 4: Inputting the high-resolution original features of the input medical image and the final deep coding feature map obtained in step 3 into a frequency-aware residual fusion module, and obtaining a second fused feature map through the feature fusion operation of the frequency-aware residual fusion module;

[0013] Step 5: Input the second fused feature map obtained in step 4 into a similar feature alignment module, and obtain an aligned decoder input feature map through a spatial alignment operation of the similar feature alignment module;

[0014] Step 6: Input the aligned decoder input feature map obtained in step 5 and the encoded feature map output at each corresponding level of the multi-stage encoder network obtained in step 3 into the multi-stage decoder network, and obtain the final medical image segmentation mask through the step-by-step upsampling and feature fusion operations of the multi-stage decoder network;

[0015] Step 7: Output the final medical image segmentation mask.

[0016] In the above-mentioned multi-scale frequency fusion medical image segmentation method based on wavelet transform, step 2 specifically includes:

[0017] Step 2.1: Concatenate the low-frequency approximate sub-band feature map and the high-frequency detail sub-band feature map obtained in step 1 along the channel dimension to form a multi-channel combined feature tensor;

[0018] Step 2.2: Input the multi-channel combined feature tensor obtained in step 2.1 into the convolutional layer and then pass it through the Sigmoid activation function to learn and generate the attention weight map corresponding to each frequency subband;

[0019] Step 2.3: Perform element-by-element multiplication of each channel of the attention weight map obtained in step 2.2 with the original corresponding frequency sub-band feature map obtained in step 1 to obtain the weighted frequency sub-band feature map;

[0020] Step 2.4: Perform element-by-element summation or concatenation and convolution operations on the weighted frequency sub-band feature maps obtained in step 2.3 to obtain the first fused feature map.

[0021] In the above-mentioned multi-scale frequency fusion medical image segmentation method based on wavelet transform, step 4 specifically includes:

[0022] Step 4.1: Perform a 1x1 convolutional layer channel adjustment operation on the high-resolution original features of the medical image input in step 1 and the final deep coding feature map obtained in step 3, respectively, to obtain adjusted high-resolution features and adjusted deep coding features;

[0023] Step 4.2: The adjusted high-resolution features obtained in step 4.1 are processed by low-pass filtering and high-pass filtering respectively to extract their low-frequency global structure components and high-frequency detail components;

[0024] Step 4.3: Apply content-aware feature reconstruction to the high-frequency detail components obtained in step 4.2 to perform frequency enhancement to obtain enhanced high-resolution high-frequency features;

[0025] Step 4.4: The adjusted deep coding features obtained in step 4.1 are subjected to low-pass filtering to extract their low-frequency components, and are upsampled by applying a content-aware feature reconstruction operation to obtain upsampled low-resolution low-frequency features;

[0026] Step 4.5: Perform a feature fusion operation on the enhanced high-resolution high-frequency features obtained in step 4.3 and the upsampled low-resolution low-frequency features obtained in step 4.4 to obtain the second fused feature map.

[0027] In the above-mentioned multi-scale frequency fusion medical image segmentation method based on wavelet transform, step 5 specifically includes:

[0028] Step 5.1: For each feature point or local area thereof in the second fused feature map obtained in step 4, calculate the similarity score between it and the feature point or local area at the corresponding position in the enhanced high-resolution high-frequency feature obtained in step 4.3, for example, using cosine similarity calculation;

[0029] Step 5.2: Generate a spatial displacement field or sampling grid based on the similarity score obtained in step 5.1;

[0030] Step 5.3: Resample and adjust the second fused feature map obtained in step 4 according to the spatial displacement field or sampling grid obtained in step 5.2 to achieve spatial alignment of features and generate the aligned decoder input feature map.

[0031] The above-mentioned multi-scale frequency fusion medical image segmentation method based on wavelet transform also includes step 8: using the final medical image segmentation mask output in step 7 for three-dimensional reconstruction, visualization display or interactive auxiliary diagnosis in the electronic health metaverse scene.

[0032] A multi-scale frequency fusion medical image segmentation system based on wavelet transform, applied to any of the above-mentioned multi-scale frequency fusion medical image segmentation methods based on wavelet transform, characterized by comprising a multi-stage encoder network, a multi-stage decoder network, a discrete wavelet transform module, a frequency-weighted attention fusion module, a frequency-aware residual fusion module, and a similar feature alignment module;

[0033] The discrete wavelet transform module is configured at the front end of the input path of the multi-stage encoder network or at an early stage thereof, and is used to perform the multi-scale frequency decomposition operation of step 1 of the method on the received original medical image;

[0034] The frequency-weighted attention fusion module is connected to the output end of the discrete wavelet transform module and is used to perform the adaptive weighted fusion processing of step 2 of the method;

[0035] The multi-stage encoder network is used to receive the output of the frequency-weighted attention fusion module and perform the step-by-step downsampling and feature extraction operations of step 3 of the method;

[0036] The frequency-aware residual fusion module, whose input ends are respectively connected to the path for obtaining the high-resolution features of the original medical image and the path for obtaining the final deep encoding features output by the multi-stage encoder network, is used to perform the feature fusion operation in step 4 of the method;

[0037] The similar feature alignment module is connected to the output end of the frequency-aware residual fusion module or its subsequent processing path, and is used to perform the spatial alignment operation in step 5 of the method;

[0038] The multi-stage decoder network is used to combine the output of the similar feature alignment module and the encoding features obtained from the corresponding level of the multi-stage encoder network through the jump connection, perform the feature fusion and reconstruction operation of step 6 of the method, and finally output the segmentation result of the medical image.

[0039] The beneficial effects of the present invention are: a multi-scale frequency fusion medical image segmentation system and method based on wavelet transform of the present invention, by performing explicit multi-scale frequency decomposition of medical images, and combining specific modules such as frequency-weighted attention fusion, frequency-aware residual fusion and similar feature alignment, can more effectively extract and utilize the structural information and detail information of the image, significantly improving the accuracy, robustness and detail recovery ability of the segmentation of complex lesion areas; at the same time, the design of the scheme takes into account the efficiency of feature fusion and the precise alignment of multi-scale information, which helps to control the computational complexity while maintaining high performance, so that it can show superior performance and broad application prospects in a variety of medical image segmentation tasks, especially in emerging medical scenarios such as the electronic health metaverse with high requirements for real-time and interactivity. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0041] Figure 1It is a schematic diagram showing the overall framework of the application of the multi-scale frequency fusion medical image segmentation model (WFANet) based on wavelet transform in the Metaverse medical platform of the present invention.

[0042] Figure 2 Schematic diagram of the overall network architecture of the WFANet model of the present invention, wherein Figure 2 Part (a) shows the adaptive frequency domain decomposition and fusion module (this module implements the discrete wavelet transform DWT and inputs the decomposed multi-band features into the frequency-weighted attention fusion module F-WAF for processing). Figure 2 Part (b) shows the backbone network (encoder-decoder structure). Figure 2 Part (c) illustrates the structure of the frequency-aware residual fusion module (F-ARF).

[0043] Figure 3 Schematic diagram showing the detailed structure of the Similar Feature Alignment module (SFA) in the present invention.

[0044] Figures 4 to 9 These are visual comparison diagrams of the segmentation results of the method of the present invention and other comparison methods on the CVC-ClinicDB dataset, ISIC2017 dataset, BUSI dataset, SZ-CXR dataset, BreaDM dataset and Taidi_7th_B dataset. DETAILED DESCRIPTION

[0045] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] Reference Figures 1 to 3 The present invention discloses a multi-scale frequency fusion medical image segmentation system and method based on wavelet transform, and its core network architecture is named WFANet. The system mainly includes a discrete wavelet transform module, a frequency-weighted attention fusion module (F-WAF), a multi-stage encoder network (part of the backbone network), a frequency-aware residual fusion module (F-ARF), a similar feature alignment module (SFA) and a multi-stage decoder network (another part of the backbone network). These modules work together to implement the medical image segmentation method described in the present invention. Among them, Figure 1 An application scenario of the method of the present invention in the Metaverse medical platform is demonstrated. Figure 2 The overall network architecture of the WFANet model is shown. Figure 2 Part (a) focuses on the collaborative work of discrete wavelet transform (DWT) processing and frequency-weighted attention fusion module (F-WAF). Figure 2 Part (b) shows the encoder-decoder network structure as the backbone. Figure 2Part (c) illustrates the composition of the frequency-aware residual fusion module (F-ARF). Figure 3 The internal structure of the Similarity Feature Alignment module (SFA) is further shown in detail.

[0047] Example 1: Detailed Structure and Working Principle of the WFANet System

[0048] The WFANet system processes input medical images (e.g., Figure 1 The MR angiography images, CT images, etc. in the Metaverse medical platform shown are processed and finally accurate segmentation results (e.g., lesion areas) are output.

[0049] Adaptive frequency domain decomposition and fusion module

[0050] Reference Figure 2 In the adaptive frequency domain decomposition and fusion module shown in (a), the input medical image is first processed by discrete wavelet transform (DWT). DWT is a time-frequency analysis method that uses a mother wavelet function to perform multi-resolution analysis of the signal by adjusting the scale and translation parameters. For a two-dimensional image f(x,y), DWT uses a separable convolution operation by applying a low-pass filter h_L and a high-pass filter h_H to the rows and columns of the image, followed by downsampling. This decomposes the image into four subbands: LL (a low-frequency approximation subband, obtained by row and column low-pass filtering, primarily containing the image's overall contour and smooth regions); LH (a horizontal high-frequency detail subband, obtained by row low-pass filtering and column high-pass filtering, primarily representing horizontal edges and texture information); HL (a vertical high-frequency detail subband, obtained by row high-pass filtering and column low-pass filtering, primarily representing vertical edges and texture information); and HH (a diagonal high-frequency detail subband, obtained by row high-pass filtering and column high-pass filtering, primarily representing diagonal edges, texture, and some noise information). This step effectively separates the structural information and detail information of the image in different frequency channels.

[0051] Subsequently, the four sub-band feature maps of LL, LH, HL, and HH obtained from the aforementioned DWT processing are input into the frequency-weighted attention fusion module (F-WAF). This module aims to adaptively fuse these multi-band information to enhance feature expression. Its specific operations may include:

[0052] (1) Feature concatenation: The feature maps of the four sub-bands LL, LH, HL, and HH are concatenated in the channel dimension to form a multi-channel combined feature tensor, for example, represented as X.

[0053] (2) Attention weight generation: The multi-channel combined feature tensor X obtained in step (1) is input into one or more convolutional layers (for example, the first convolutional layer Conv1 is followed by a ReLU activation function to output an intermediate feature X' with C_inter channels; the second convolutional layer Conv2 processes the intermediate feature X') and finally passes through a Sigmoid activation function to learn and generate a feature tensor corresponding to the number of original sub-bands (for example, 4 channels) and with a spatial size that is the same as the sub-band feature. Figure 1 The attention map of (H x W) is represented as A. Each channel A_k (k=1, 2, 3, 4) of the attention map represents the importance weight of the corresponding k-th frequency subband under the current image content.

[0054] (3) Weighted fusion: Each channel A_k of the attention map A obtained in step (2) is element-wise multiplied with the original corresponding k-th frequency subband feature map to obtain the weighted frequency subband feature maps, such as LL_fused, LH_fused, HL_fused, and HH_fused.

[0055] (4) Fusion output: The weighted frequency sub-band feature maps LL_fused, LH_fused, HL_fused, and HH_fused obtained in step (3) are summed element by element to obtain the final fusion output feature, for example, represented as Y. The fusion output feature Y is then fed into the encoder part of the backbone network.

[0056] In this way, the F-WAF module is able to dynamically adjust the contribution of each frequency component according to the image content, thereby more effectively integrating multi-scale information and highlighting features that are beneficial to the segmentation task.

[0057] backbone network

[0058] Reference Figure 2 In the backbone network shown in (b), the present invention adopts a classic U-shaped encoder-decoder structure.

[0059] Encoder part: via Figure 2 (a) The fused feature Y output by the module after processing is used as the input of the encoder. The encoder usually consists of multiple encoding stages (e.g., Figure 2(b) illustrates the composition of Stage 1 to Stage 4. Each encoding stage contains several convolutional layers (for feature extraction), normalization layers (such as batch normalization, for stable training), and nonlinear activation functions (such as ReLU, for introducing nonlinearity). After each encoding stage, a downsampling operation (such as maximum pooling or strided convolution) is usually performed to gradually reduce the spatial size of the feature map while increasing the number of channels of the feature map accordingly, thereby extracting higher-level and more abstract semantic features. The feature map output by each encoding stage is passed to the corresponding stage of the decoder network through a skip connection (SkipConnection) to retain low-level detail information.

[0060] Decoder part: The decoder also consists of multiple decoding stages, and its structure is usually symmetrical with the encoder in terms of hierarchy. Each decoding stage first gradually restores the spatial size of the feature map through upsampling operations (such as transposed convolution or bilinear interpolation followed by convolution) to match the feature map size of the corresponding stage of the encoder. Then, the upsampled features are fused with the shallow features passed from the corresponding stage of the encoder through jump connections (for example, by splicing and then performing convolution operations). The fused features are then processed through several convolutional layers, normalization layers, and activation functions, and finally a segmentation prediction map with the same resolution as the input image is restored step by step.

[0061] Frequency-aware residual fusion module

[0062] Reference Figure 2 As shown in (c), to better preserve the global spatial structure of the image and enhance the utilization of high-frequency details, the present invention introduces the F-ARF module. The core idea of ​​this module is to fuse the features of the high-resolution original image (e.g., represented by I_H) with the low-resolution features obtained through the encoder path (e.g., represented by I_R).

[0063] Its specific implementation may include:

[0064] (1) Feature compression: The high-resolution original image feature I_H and the low-resolution feature I_R are respectively subjected to channel compression through a 1x1 convolution layer to unify their channel dimensions, thereby obtaining the compressed high-resolution feature I_H' and the compressed low-resolution feature I_R'.

[0065] (2) High-resolution feature processing: The compressed high-resolution feature I_H' obtained in step (1) is subjected to a low-pass filtering process (e.g., a low-pass encoder) to extract its low-frequency information (e.g., represented as I_HL), and is subjected to a high-pass filtering process (e.g., a high-pass encoder) to extract its high-frequency information (e.g., represented as I_HH).

[0066] (3) High-frequency enhancement and low-frequency upsampling: The high-frequency information I_HH extracted in step (2) is subjected to frequency enhancement by applying a content-aware feature recombination (CARAFE) operation (e.g., the upsampling factor is set to 1 for feature enhancement) to obtain enhanced high-resolution high-frequency features (e.g., represented as I_HH'). The compressed low-resolution features I_R' obtained in step (1) (or the low-frequency components obtained by low-pass filtering, e.g., represented as I_RL) are subjected to upsampling by applying a content-aware feature recombination (CARAFE) operation (e.g., the upsampling factor is set to 2 or higher to match the target resolution) to obtain upsampled low-resolution features (e.g., represented as I_RL').

[0067] (4) Feature fusion: The enhanced high-resolution high-frequency features I_HH' obtained in step (3) are fused with the upsampled low-resolution features I_RL' (and optionally, the original high-resolution low-frequency information I_HL obtained in step (2)) (e.g., by element-wise addition or convolution after concatenation) to obtain the final fusion output (e.g., denoted as I_fusion). This fused feature I_fusion is further input to the SFA module or subsequent stages of the decoder.

[0068] Similar feature alignment module

[0069] Reference Figure 3 As shown in FIG, in order to further optimize the spatial correspondence between features from different sources or different resolutions (for example, between the output I_fusion of the aforementioned F-ARF module and a feature representation of a high-resolution path) before fusion, the SFA module is introduced.

[0070] The specific steps may include:

[0071] (1) Similarity calculation: For two input feature maps (for example, one is the fusion feature I_fusion from the F-ARF module, and the other is the feature of the high-resolution path, such as I_HH' processed in the F-ARF module), the feature vectors in the local window are extracted at their corresponding spatial positions, and then the similarity between these local feature vectors is calculated, for example, using cosine similarity to obtain a similarity matrix or similarity score map.

[0072] (2) Spatial offset generation: Based on the similarity matrix or similarity score map calculated in step (1), a spatial offset field is generated. The offset field indicates the optimal spatial displacement of each pixel in the feature map to be aligned relative to the corresponding content of the reference feature map.

[0073] (3) Feature alignment: Using the spatial offset field generated in step (2), the feature map to be aligned (e.g., I_fusion) is sampled, such as by deformable convolution or offset-based interpolation sampling, so that the pixel values ​​of the feature map are recombined according to the offset to obtain an output feature map that is more accurately aligned with the reference feature map in space.

[0074] This module ensures the spatial consistency of features in subsequent fusion operations.

[0075] Example 2: Specific process of WFANet method

[0076] Based on the above system, the medical image segmentation method process of the present invention (as claimed in claim 1) is further described as follows:

[0077] Step 1: Perform a discrete wavelet transform (DWT) operation on an input medical image. For details, refer to the description of DWT processing in Part 1 of Example 1. This step decomposes the input image into a low-frequency approximation subband feature map (abbreviated as feature map L) and three high-frequency detail subband feature maps (abbreviated as feature map H).

[0078] Step 2: Input the feature map L and the feature map H obtained in step 1 into the frequency-weighted attention fusion module (F-WAF). The frequency-weighted attention fusion module performs adaptive weighted fusion processing to obtain a first fused feature map (abbreviated as feature map F1). The specific implementation of this module is as described in the first part of Example 1 regarding the F-WAF module.

[0079] Step 3: Input the feature map F1 obtained in step 2 (or a combination of the feature map F1 and the original features of the medical image input in step 1) into a multi-stage encoder network. Through the multiple encoding stages of the multi-stage encoder network, downsampling and feature extraction operations are performed step by step (see the encoder description in Section 2 of Example 1 for details), and the encoded feature maps output by the multi-stage encoder network at each corresponding level and a final deep encoded feature map (abbreviated as feature map Ed) are obtained.

[0080] Step 4: The high-resolution original features of the medical image input in step 1 (referred to as features Io) and the feature map Ed obtained in step 3 are input to a frequency-aware residual fusion module (F-ARF). Through the feature fusion operation of the frequency-aware residual fusion module (refer to the description in section 3 of Example 1 for details), a second fused feature map (referred to as feature map F2) is obtained.

[0081] Step 5: The feature map F2 obtained in step 4 (and possibly also features from the high-resolution path, such as feature H_H') is input to a similar feature alignment module (SFA). Through the spatial alignment operation of the similar feature alignment module (see the description in section 4 of Example 1 for details), an aligned decoder input feature map (abbreviated as feature map Fa) is obtained.

[0082] Step 6: Input the feature map Fa obtained in step 5 and the encoded feature map obtained from the corresponding level of the multi-stage encoder network in step 3 via skip connections into a multi-stage decoder network. Upsampling and feature fusion operations are performed step by step through multiple decoding stages of the multi-stage decoder network (see the decoder description in Section 2 of Example 1 for details) to obtain the final medical image segmentation mask (abbreviated as mask M).

[0083] Step 7: Output the mask M obtained in step 6. Depending on the specific application, the mask M can be directly used as the segmentation result, or can be output as the final segmentation result after threshold processing and morphological post-processing.

[0084] Step 8 (optional): Use the mask M output from step 7 in the electronic health metaverse scenario (refer to Figure 1 ) for three-dimensional reconstruction, visualization, or interactive auxiliary diagnosis.

[0085] Example 3: Model training and application examples

[0086] When training the WFANet model of the present invention, a combination of multiple loss functions is used for optimization, and the pixel-level binary cross entropy loss (Binary Cross-Entropy Loss, L_CE) and the regional-level Dice loss (DiceLoss, L_DC) are weightedly combined, with a total loss Loss = α*L_CE+β*L_DC, where α and β are preset weight hyperparameters (for example, α = 0.4, β = 0.6) to balance the contributions of different loss functions. The optimizer can use the Adam optimizer and set the initial learning rate (for example, 1e-3 or 1e-4) and the learning rate scheduling strategy (for example, cosine annealing decay). The training process usually sets a fixed number of training rounds (Epochs, for example, 300 rounds), and can be combined with an early stopping mechanism to prevent overfitting when the model performance on the validation set no longer improves.

[0087] The accurate segmentation results obtained by the segmentation method of the present invention can be further applied to the electronic health metaverse scenario, such as referring to Figure 1As shown. Specific applications may include: combining a continuous two-dimensional segmentation mask M with a corresponding sequence of original medical image slices (such as CT, MRI sequences), and generating an accurate three-dimensional digital model of the target area (such as lesions, organs) through a three-dimensional reconstruction algorithm (such as Marching Cubes, etc.). These three-dimensional models can be immersively visualized in the Metaverse platform with the help of virtual reality (VR), augmented reality (AR) or mixed reality (MR) devices, and users can perform interactive operations such as rotation, scaling, translation, and sectioning on the model. This interactive three-dimensional visualization can assist doctors in more intuitive disease analysis, remote collaborative diagnosis, preoperative planning and simulation exercises for complex surgeries, and can also be used in scenarios such as medical teaching, patient science popularization, and rehabilitation training, thereby improving the intelligence level, accessibility, and interactive experience of medical services.

[0088] Example 4: Experimental Verification

[0089] Reference Figures 4 to 9 The WFANet model proposed in the present invention has been extensively experimentally verified on multiple public medical image segmentation datasets, which cover different imaging modalities and segmentation targets, such as CVC-ClinicDB (colorectal polyp video frames), ISIC2017 (dermoscopy images), BUSI (breast ultrasound images), SZ-CXR (chest X-rays), BreaDM (breast MRI images), and Taidi_7th_B (rectal tumor CT images). The performance of the method of the present invention was compared with that of a variety of existing mainstream or advanced medical image segmentation models (such as U-Net, UNet++, Att-UNet, WRANet, DualA-Net, DPMNet, etc.) under the same experimental conditions. The experimental results were quantitatively evaluated using commonly used segmentation evaluation indicators, including Dice similarity coefficient (Dice), intersection over union (IoU), precision (Precision) and recall (Recall). The data show that the method of the present invention has achieved leading or significantly competitive performance in most datasets and evaluation indicators. For example, on the CVC-ClinicDB dataset, the Dice coefficient of WFANet can reach 90.90%; on the ISIC2017 dataset, the Dice coefficient can reach 88.85%; on the BUSI dataset, the Dice coefficient can reach 76.67%. Figures 4 to 9The visual segmentation results shown further intuitively confirm the superiority of the method of the present invention. Compared with the comparative methods, the method of the present invention can more accurately outline the boundaries of the lesion or target area, especially when processing targets with irregular shapes, blurred boundaries, and small sizes. It can better retain detailed information, reduce missed segmentation and over-segmentation, and the segmentation results are more complete and continuous visually, and can effectively suppress noise interference in the background area. These qualitative and quantitative experimental results fully demonstrate the effectiveness, accuracy, and robustness of the method of the present invention in medical image segmentation tasks, as well as its good generalization ability and wide applicability when processing medical images of different modalities and different diseases.

[0090] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A multi-scale frequency fusion medical image segmentation method based on wavelet transform, characterized in that: The method comprises the following steps: Step 1: Perform discrete wavelet transform on the input medical image to decompose it into a low-frequency approximate sub-band feature map and three high-frequency detail sub-band feature maps; Step 2: Input the low-frequency approximate sub-band feature map and the high-frequency detail sub-band feature map obtained in step 1 into the frequency-weighted attention fusion module, and perform adaptive weighted fusion processing through the frequency-weighted attention fusion module to obtain a first fused feature map; Step 3: Inputting the first fusion feature map obtained in step 2 or the combined feature of the first fusion feature map and the input medical image into a multi-stage encoder network, and obtaining the encoding feature map output by the multi-stage encoder network at each corresponding level and the final deep encoding feature map through step-by-step downsampling and feature extraction operations of the multi-stage encoder network; Step 4: Inputting the high-resolution original features of the input medical image and the final deep coding feature map obtained in step 3 into a frequency-aware residual fusion module, and obtaining a second fused feature map through the feature fusion operation of the frequency-aware residual fusion module; Step 5: Input the second fused feature map obtained in step 4 into a similar feature alignment module, and obtain an aligned decoder input feature map through a spatial alignment operation of the similar feature alignment module; Step 6: Input the aligned decoder input feature map obtained in step 5 and the encoded feature map output at each corresponding level of the multi-stage encoder network obtained in step 3 into the multi-stage decoder network, and obtain the final medical image segmentation mask through the step-by-step upsampling and feature fusion operations of the multi-stage decoder network; Step 7: Output the final medical image segmentation mask.

2. The multi-scale frequency fusion medical image segmentation method based on wavelet transform according to claim 1, characterized in that: The step 2 specifically includes: Step 2.1: Concatenate the low-frequency approximate sub-band feature map and the high-frequency detail sub-band feature map obtained in step 1 along the channel dimension to form a multi-channel combined feature tensor; Step 2.2: Input the multi-channel combined feature tensor obtained in step 2.1 into the convolutional layer and then pass it through the Sigmoid activation function to learn and generate the attention weight map corresponding to each frequency subband; Step 2.3: Perform element-by-element multiplication of each channel of the attention weight map obtained in step 2.2 with the original corresponding frequency sub-band feature map obtained in step 1 to obtain the weighted frequency sub-band feature map; Step 2.4: Perform element-by-element summation or concatenation and convolution operations on the weighted frequency sub-band feature maps obtained in step 2.3 to obtain the first fused feature map.

3. The multi-scale frequency fusion medical image segmentation method based on wavelet transform according to claim 1, characterized in that: The step 4 specifically includes: Step 4.1: Perform a 1x1 convolutional layer channel adjustment operation on the high-resolution original features of the medical image input in step 1 and the final deep coding feature map obtained in step 3, respectively, to obtain adjusted high-resolution features and adjusted deep coding features; Step 4.2: The adjusted high-resolution features obtained in step 4.1 are processed by low-pass filtering and high-pass filtering respectively to extract their low-frequency global structure components and high-frequency detail components; Step 4.3: Apply content-aware feature reconstruction to the high-frequency detail components obtained in step 4.2 to perform frequency enhancement to obtain enhanced high-resolution high-frequency features; Step 4.4: The adjusted deep coding features obtained in step 4.1 are subjected to low-pass filtering to extract their low-frequency components, and are upsampled by applying a content-aware feature reconstruction operation to obtain upsampled low-resolution low-frequency features; Step 4.5: Perform a feature fusion operation on the enhanced high-resolution high-frequency features obtained in step 4.3 and the upsampled low-resolution low-frequency features obtained in step 4.4 to obtain the second fused feature map.

4. The multi-scale frequency fusion medical image segmentation method based on wavelet transform according to claim 1, characterized in that: The step 5 specifically includes: Step 5.1: For each feature point or local area thereof in the second fused feature map obtained in step 4, calculate the similarity score between it and the feature point or local area at the corresponding position in the enhanced high-resolution high-frequency feature obtained in step 4.3, for example, using cosine similarity calculation; Step 5.2: Generate a spatial displacement field or sampling grid based on the similarity score obtained in step 5.1; Step 5.3: Resample and adjust the second fused feature map obtained in step 4 according to the spatial displacement field or sampling grid obtained in step 5.2 to achieve spatial alignment of features and generate the aligned decoder input feature map.

5. The multi-scale frequency fusion medical image segmentation method based on wavelet transform according to claim 1, characterized in that: The method also includes step 8: using the final medical image segmentation mask output from step 7 for three-dimensional reconstruction, visualization display or interactive auxiliary diagnosis in the electronic health metaverse scene.

6. A multi-scale frequency fusion medical image segmentation system based on wavelet transform, applied to a multi-scale frequency fusion medical image segmentation method based on wavelet transform according to any one of claims 1 to 5, characterized in that: It includes a multi-stage encoder network, a multi-stage decoder network, a discrete wavelet transform module, a frequency-weighted attention fusion module, a frequency-aware residual fusion module, and a similar feature alignment module; The discrete wavelet transform module is configured at the front end of the input path of the multi-stage encoder network or at an early stage thereof, and is used to perform the multi-scale frequency decomposition operation of step 1 on the received original medical image; The frequency-weighted attention fusion module is connected to the output end of the discrete wavelet transform module and is used to perform the adaptive weighted fusion processing in step 2; The multi-stage encoder network is used to receive the output of the frequency-weighted attention fusion module and perform the step-by-step downsampling and feature extraction operations of step 3; The frequency-aware residual fusion module, whose input ends are respectively connected to the path for obtaining the high-resolution features of the original medical image and the path for obtaining the final deep encoding features output by the multi-stage encoder network, is used to perform the feature fusion operation in step 4; The similar feature alignment module is connected to the output end of the frequency-aware residual fusion module or its subsequent processing path, and is used to perform the spatial alignment operation in step 5; The multi-stage decoder network is used to combine the output of the similar feature alignment module and the encoding features obtained from the corresponding level of the multi-stage encoder network through the jump connection, perform the feature fusion and reconstruction operation of step 6, and finally output the segmentation result of the medical image.

Citation Information

Cited By

  • Biomedical image segmentation method, system, equipment and medium

    CN121259324A