Medical image segmentation method based on hot spot region joint tracking model

By combining the convolutional neural network and Swin Transformer's hot spot area joint tracking model, the problem of insufficient accuracy in liver tumor segmentation is solved, and efficient liver tumor segmentation effect is achieved.

CN120259353APending Publication Date: 2025-07-04CHINA UNIV OF MINING & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510163190.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing medical image segmentation methods have insufficient accuracy when processing complex medical images, especially liver tumor segmentation. Traditional methods rely on expert experience and are inefficient, while deep learning-based methods have limitations when processing global information and small organ segmentation.

Method used

Using a joint tracking model based on hot spot area, combined with a convolutional neural network and a Swin Transformer encoder decoder, high-precision segmentation of liver tumors is achieved through feature extraction, coordinate position coding, hot spot area determination and target association modules, and the hot spot area determination module is used to improve the segmentation effect.

Benefits of technology

The accuracy and efficiency of medical image segmentation were significantly improved, especially in liver tumor segmentation, the segmentation index increased by 4.8% and converged earlier, proving the effectiveness of the hot spot area determination module.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259353A_ABST
    Figure CN120259353A_ABST
Patent Text Reader

Abstract

The invention provides a medical image segmentation method based on a hot spot region joint tracking model, which comprises the following steps of: establishing a hot spot region joint tracking model through a python compiler, comprising a feature extraction module capable of extracting high-level semantic features of an input image and outputting a feature map; the coordinate position coding module can add position information to the feature map; the Swin Transform encoder and decoder module can encode target information in an input image into feature representation of global information, generate a mutual relation between the target information, and perform target prediction and association through a query vector; the hot spot region judgment module can judge a hot spot region in the input image; the target association module can match the detection result of the current frame with the tracking result of the previous frame; and the target segmentation module can segment an image region which is judged to be an abnormal region in the input image according to a matching result of a detection result of the current frame and a tracking result of the previous frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image segmentation, and in particular to a medical image segmentation method based on a hot spot region joint tracking model. Background Art

[0002] Medical image segmentation is a core task in the fields of computer vision and medical image processing. Its purpose is to divide an image according to the characteristics between regions, so as to more accurately locate and analyze the anatomical structure or lesion area of the target region, and plays a crucial role in liver tumor diagnosis, treatment plan formulation, and disease monitoring. However, due to problems such as patient movement, instrument aging, and image noise / artifacts during filming, and the objective problems that liver tumors usually have irregular shapes and blurred boundaries, the accuracy of its medical image segmentation results still faces great challenges in practical applications.

[0003] Current image segmentation methods are mainly divided into two categories: image segmentation based on traditional methods and image segmentation based on deep learning. Traditional methods, such as threshold-based segmentation methods, edge detection-based segmentation methods, region-based segmentation methods, etc., have the advantages of high computational efficiency and simple implementation. However, their segmentation effects highly depend on the professional experience of radiologists. Especially for complex medical images, it is difficult to meet the requirements of automation and high precision. In contrast, image segmentation methods based on deep learning have the advantages of high precision, strong adaptability, and stronger complex image processing capabilities, and have gradually become the mainstream methods for medical image segmentation.

[0004] There are mainly two mainstream models for image segmentation methods based on deep learning: segmentation models based on convolutional neural networks and segmentation models based on Transformers. These two types of methods each have certain limitations, which are specifically as follows:

[0005] ① The segmentation model based on convolutional neural networks is based on the encoder-decoder structure, such as the Unet model. Among them, the encoder extracts the deep semantic features of the image through a series of convolutional and pooling operations, and the decoder restores the detailed information of the image through upsampling operations, and realizes multi-level feature fusion of the image through skip connections. Unet and its variants, such as Unet++, Unet3+, Res-Unet, etc., have achieved remarkable results in medical image segmentation. However, due to the fixed size of the convolutional kernels in the convolutional neural network segmentation model, its receptive field is limited. Therefore, its performance is limited when dealing with global information, fine structures, and edge precise segmentation tasks.

[0006] ②The Transformer model is based on the attention mechanism. SwinUnet is the first pure Transformer segmentation model, which replaces two-dimensional convolutional blocks with Swin Transformer modules to achieve feature representation and global semantic information interaction. In the abdominal multi-organ image and cardiac image segmentation tasks, this model shows good image segmentation performance. In addition, the C2Former model introduces a cross-convolution self-attention mechanism, which enhances the ability to understand image semantic features by modeling long- and short-distance dependencies. However, Transformer-based segmentation models usually focus too much on long-range dependencies and ignore local features, and their segmentation effects on small organ images and precise boundary segmentation are not as good as those of well-designed convolutional neural network segmentation models. Summary of the Invention

[0007] To solve the problems existing in the above-mentioned prior art, the present invention provides a medical image segmentation method based on a hot spot region joint tracking model, and the specific technical solutions are as follows:

[0008] A medical image segmentation method based on a hot spot region joint tracking model includes the following steps:

[0009] S1. Establish an image data set of existing medical images, and the images in the image data set are input images;

[0010] S2. Establish a hot spot region joint tracking model through a python compiler, including a feature extraction module, a coordinate position encoding module, a Swin Transformer encoder-decoder module, a hot spot region determination module, an object association module, and an object segmentation module;

[0011] S21. The feature extraction module can extract high-level semantic features of the input image and output a feature map;

[0012] S22. The coordinate position encoding module can add position information to the feature map of the input image;

[0013] S23. The Swin Transformer encoder-decoder module can encode the target information in the input image into a feature representation of global information, generate the mutual relationship between target information, and perform target prediction and association in the input image through a query vector;

[0014] S24. The hot spot region determination module can determine the hot spot region where the tracking target is located in the input image;

[0015] S25. The target association module can match the detection result of the current frame in the Swin Transformer encoder-decoder module with the tracking result of the previous frame in the hotspot area judgment module;

[0016] S26. The target segmentation module can segment the image area determined as the abnormal area in the input image according to the matching result of the detection result of the current frame in the Swin Transformer encoder-decoder module and the tracking result of the previous frame in the hotspot area judgment module.

[0017] Furthermore, constructing the feature extraction module includes the following specific steps:

[0018] S21-1. Set ResNet50 as the backbone of the convolutional neural network;

[0019] S21-2. Input the input image into the convolutional neural network, and extract the basic features of the input image through preliminary convolutional operations;

[0020] S21-3. Through multi-layer convolutional processing of the convolutional neural network, extract the high-level semantic features of the input image and output a multi-scale feature map, and the expression is:

[0021] F = {F1, F2,..., F n}

[0022] The number of channels of the feature map output by each convolutional layer of the convolutional neural network increases with the increase of the depth of the convolutional neural network, and the expression is:

[0023] F i = Conv(I), i = 1,..., n

[0024] In the formula, I is the input image, and its size is H×W×C in ; F i is the feature map, and its size is H i ×W i ×C i .

[0025] Furthermore, constructing the coordinate position encoding module includes the following specific steps:

[0026] S22-1. Use sine-cosine position encoding to add a pair of position vectors to each pixel of the input image;

[0027] S22-2. Add the feature map and the position encoding to generate an enhanced feature map, where the position encoding is a position information matrix with the same size as the feature map.

[0028] Further, the expression of the position encoding vector for the x-th row and y-th column of the position encoding sequence in sine-cosine encoding is:

[0029]

[0030] In the formula, P i (x, y, 2k) is the position encoding vector of the x-th row of the position encoding P i sequence;

[0031] P i (x, y, 2k + 1) is the position encoding vector of the y-th column of the position encoding P i sequence;

[0032] k is the channel index; C i is the number of channels of the feature map.

[0033] Further, constructing the Swin Transformer encoder-decoder module includes the following specific steps:

[0034] S23-1, perform a window partitioning operation on the enhanced feature map through the following formula:

[0035] W = Split(F' i )

[0036] In the formula, W is the set of partitioned windows, and the size of each window is M×M×C i ; F' i is the enhanced feature map;

[0037] S23-2, perform self-attention operation within the window on the set of partitioned windows, and the expression is:

[0038] W' = Self-Attention(W)

[0039] S23-3, based on the window-based multi-head self-attention mechanism, extract the local and global features of the input image, and the expression is:

[0040]

[0041] In the formula, Q, K, and V are the query, key, and value matrices respectively; d k is the dimension of the key vector;

[0042] S23-4, merge the window features to generate the global feature map, and the expression is:

[0043] F” i = Merge(W')

[0044] In the formula, F”i is the global feature map, and its size is the same as that of the input enhanced feature map F' i identical.

[0045] Furthermore, constructing the hot spot area determination module includes the following specific steps:

[0046] S24-1, perform normalization processing on all pixel points in the input image to obtain the edge distribution of rows or columns. The expression is:

[0047]

[0048] In the formula, P i (x) is the probability distribution of rows; is the probability that each pixel is an outlier; i is the row number of the image pixel; j is the column number of the image pixel;

[0049] S24-2, calculate the KL divergence between the GT distribution and the prediction distribution of columns according to the discrete edge distribution of the input image pixel points. The expression is:

[0050] KL i (GT||Pre) = H i (gt, pre) - H i (gt)

[0051] In the formula, KL i (GT||Pre) is the KL divergence between the GT distribution and the prediction distribution pre; H i (gt, pre) is the prediction distribution on GT; H i (gt) is the actual distribution on GT;

[0052] S24-3, represent the relative entropy loss of the discrete edge distribution of the input image pixel points through the following formula, and complete the determination of the hot spot area where the tracking target is located:

[0053] Loss KL = (KL i (GT||Pre) + KL j (GT||Pre)) / 2

[0054] In the formula, Loss KL is the position loss; KL i (GT||Pre) is the KL divergence loss on rows; KL j (GT||Pre) is the KL divergence loss on columns.

[0055] Furthermore, constructing the target association module includes the following specific steps:

[0056] S25-1. Calculate the matching cost matrix between the detection result of the current frame in the Swin Transformer encoder-decoder module and the tracking result of the previous frame in the hot-spot region judgment module through the following formula:

[0057] C ij = ||Q' i - T t-1,j || + λ·D ij

[0058] In the formula, C ij is the matching cost between the i-th query target and the j-th track target; Q' i is the query embedding of the i-th target; T t-1,j is the track embedding of the j-th track; D ij is the distance between targets; λ is the weight coefficient;

[0059] S25-2. Match the detection result of the current frame in the Swin Transformer encoder-decoder module with the tracking result of the previous frame in the hot-spot region judgment module through the following formula for the matching cost matrix, and output the associated target track:

[0060] T t = Hungarian(C)

[0061] In the formula, T t is the track associated with the current frame.

[0062] Furthermore, the target segmentation module generates a target segmentation mask according to the target track, and the expression is:

[0063] M = SegmentationHead(T t , F” i )

[0064] In the formula, M is the output target segmentation mask; T t is the track associated with the current frame; F” i is the global feature map.

[0065] Furthermore, the medical image segmentation method further includes the following steps:

[0066] S3. Make a training set and a test set to train and test the hot-spot region joint tracking model respectively;

[0067] S4. Obtain the medical image to be segmented, and perform image segmentation based on the trained hot-spot region joint tracking model.

[0068] Furthermore, using the test set to test the hot-spot region joint tracking model includes the following steps:

[0069] S31. After testing the hotspot area joint tracking model with the test set, the tracked image, mAP50 evaluation index, and mAP50-95 evaluation index are obtained, and the qualitative analysis of the hotspot area joint tracking model is completed.

[0070] S32. The image of the segmented abnormal area, as well as three segmentation indexes of MIoU, DICE, and MPA, are obtained through the hotspot area joint tracking model, and the quantitative analysis of the hotspot area joint tracking model is completed.

[0071] Based on the above technical solutions, the present invention has the following beneficial effects:

[0072] In the medical image segmentation method recorded in the present invention, a hotspot area determination module is established. From the verification results of experimental data, it can be known that for the result of BCE Loss after adding the hotspot area determination module, after 70 epochs, the curves of MIOU and DICE coefficients with the hotspot area determination module start to be higher than the curve of BCE and maintain the advantage; for the result of Focal Loss after adding the hotspot area determination module, the average IOU and DICE coefficients are higher than those of Focal Loss after adding MCE Loss after 65 epochs, and after 300 epochs, the final result is increased by 4.8%, and it converges earlier than Focal Loss; after adding the hotspot area determination module, whether CE+DICE or FCL+DICE is adopted, the test results are all improved. Especially when the FCL+DICE loss function is adopted, all three segmentation indexes are increased by more than 4%; the above all fully demonstrate the effectiveness of medical image segmentation based on the hotspot area determination module. Description of the Drawings

[0073] Figure 1 : Schematic diagram of the medical image segmentation process based on the hotspot area joint tracking model;

[0074] Figure 2 : Schematic diagram of the hotspot area determination process;

[0075] Figure 3 : Schematic diagram of the medical image segmentation result after tracking by the hotspot area joint tracking model;

[0076] Figure 4 : Line chart for verifying the effectiveness of BCE loss after adding the hotspot area determination module;

[0077] Figure 5 : Line chart for verifying the effectiveness of FCL loss after adding the hotspot area determination module;

[0078] Figure 6:Schematic diagram of the results of various indicators in the validation set after multiple rounds of learning in the embodiment. Detailed implementation manners

[0079] It should be noted that:

[0080] 1. In the description and claims, certain terms are used to refer to specific components. Those skilled in the art should understand that technicians may use different nouns to refer to the same component. The description and claims of this specification do not use the difference in nouns as a way to distinguish components, but use the difference in the functions of components as the criterion for distinction.

[0081] 2. Unless otherwise defined, the technical terms or scientific terms used in this disclosure should have the ordinary meaning understood by those of ordinary skill in the art to which this disclosure belongs.

[0082] As shown in the attached Figure 1 to the attached Figure 6 This embodiment records a medical image segmentation method based on a hot region joint tracking model, including the following steps.

[0083] S1. Obtain the 3D-IRCADb organ segmentation data set, traverse each slice of the DICOM format file in this data set, save this data set as a png format file, establish an image data set for organ segmentation, and define the images in the image data set as input images I.

[0084] In this step:

[0085] 1. The python compilation language can be used. Through the Pandas library of the python compiler, select an appropriate compiler in PyCharm to make a compilation program, and use this compilation program to traverse each slice of the DICOM format file in the 3D-IRCADb organ segmentation data set.

[0086] 2. The 3D-IRCADb organ segmentation data set can be obtained from a public website, such as the "Extreme City Developer Platform". This data set contains multi-group three-dimensional medical images of anonymous patients and the structures manually segmented by experts. These three-dimensional medical images and their segmentation structures are provided in DICOM and VTK formats.

[0087] S2. Use the python compilation language to establish a hot region joint tracking model through the python compiler, including a liver angiography abnormal feature extraction module, a coordinate position encoding module, a Swin Transformer encoder-decoder module, a hot region determination module, a target association module, and a target segmentation module.

[0088] S21. The liver contrast abnormality feature extraction module is used to extract image features from the input image I. The steps to construct this module are as follows:

[0089] S21-1. Based on the existing convolutional neural network model, set ResNet50 as the backbone of the convolutional neural network model.

[0090] S21-2. Input the input image I into the convolutional neural network model, and extract the basic features of the input image I through preliminary convolutional operations, including but not limited to edges, textures, colors, etc.

[0091] S21-3. Through the processing of multiple convolutional neural networks in the convolutional neural network model, extract the high-level semantic features of the input image I and output a multi-scale feature map.

[0092] By performing deep processing of multiple convolutions on the image basic features in step S21-2 through multiple convolutional neural networks, more abstract and semantic features of the input image I can be obtained. These features can capture more complex patterns or structures, such as the shape and semantics of liver tumors. That is, the high-level semantic features after multiple convolutions, pooling, etc. can reflect the more abstract content of the liver tumor shape in the input image I.

[0093] In the above steps:

[0094] 1. The expression of the multi-scale feature map is:

[0095] F = {F1, F2,..., F n}

[0096] 2. The feature map output by each convolutional layer of the convolutional neural network model is defined as F i , F i . The number of channels of F

[0097] F i = Conv(I), i = 1,..., n

[0098] In the formula, the size of the input image I is H×W×C in ; the size of the feature map F i is H i × W i × C i .

[0099] S22. The coordinate position encoding module is used to add position information to the feature map F i of the input image I. The steps to construct this module are as follows:

[0100] S22-1. Use sine-cosine position encoding to add a pair of position vectors to each pixel of the input image I, and represent the coordinates of each pixel in the input image I through the position vectors, so as to retain the spatial position information of the input image I.

[0101] S22-2. Add the feature map F output by each convolutional layer of the convolutional neural network i to the position encoding P i to generate an enhanced feature map F', i i.e., F' i = F i + P i , where P i is a position information matrix with the same size as F i .

[0102] In the above steps, the sine-cosine encoding for the position encoding vector expression of the x-th row and y-th column of the position encoding P i sequence is:

[0103]

[0104] In the formula,

[0105] P i (x, y, 2k) is the position encoding vector of the x-th row of the position encoding P i sequence;

[0106] P i (x, y, 2k + 1) is the position encoding vector of the y-th column of the position encoding P i sequence;

[0107] k is the channel index;

[0108] C i is the number of channels of the feature map.

[0109] S23. The Swin Transformer encoder-decoder module includes a Swin Transformer encoder and a Swin Transformer decoder. The target in this step refers to specific information in the input image I, such as liver tumor image information.

[0110] The Swin Transformer encoder is used to capture the global context information of the input image I, encode all the target information in the input image I into a feature representation of the global information, generate the mutual relationship between the targets, and provide a basis for subsequent target tracking and association.

[0111] The Swin Transformer decoder makes object predictions and associations in the input image I based on the output of the Swin Transformer encoder through query vectors. Each layer in the Swin Transformer decoder part interacts the query vectors with the global information features output by the Swin Transformer encoder to generate the location, category, and association information of each object.

[0112] Constructing the Swin Transformer encoder-decoder module includes the following steps:

[0113] S23-1, perform a window partitioning operation on the enhanced feature map F' i Extract the local and global features of the input image I based on the window-based multi-head self-attention mechanism, and use the relative entropy loss of the position features to increase the monitoring of the hotspot area.

[0114] For the enhanced feature map F' i The expression for performing the window partitioning operation is:

[0115] W = Split(F' i )

[0116] where W is the set of partitioned windows, and the size of each window is M×M×C i .

[0117] S23-2, perform self-attention operation within the windows on the partitioned window set W, and the expression is:

[0118] W' = Self-Attention(W)

[0119] The expression for the multi-head self-attention mechanism is:

[0120]

[0121] where Q, K, and V are the query, key, and value matrices respectively, which are generated by linear transformation of the partitioned window set W; d k is the key vector dimension for scaling normalization.

[0122] S23-3, merge the window features to generate the global feature map F” i , and the expression is: F” i = Merge(W'), where the output feature F” i has the same size as the input enhanced feature map F' i .

[0123] S24. The hotspot area determination module is used to determine the hotspot area where the tracking target is located. In this embodiment, the tracking target refers to the abnormal area of the input image I, that is, the high-probability position where the abnormality appears.

[0124] The hotspot area determination module determines the location of the hotspot area by calculating the KL divergence between the GT distribution and the predicted distribution of the pixel columns of the input image I, and then performs priority detection.

[0125] Constructing the hotspot area determination module includes the following steps:

[0126] S24-1. Normalize all pixel points in the input image I to obtain the marginal distribution of rows or columns. The expression is:

[0127]

[0128] In the formula,

[0129] P i (x) is the probability distribution of rows;

[0130] is the probability that each pixel is an abnormal point;

[0131] i is the row number of the image pixel;

[0132] j is the column number of the image pixel.

[0133] S24-2. According to the discrete marginal distribution of the pixel points of the input image I, calculate the KL divergence between the GT distribution and the predicted distribution of columns. The expression is:

[0134] KL i (GT||Pre) = H i (gt, pre)-H i (gt)

[0135] In the formula,

[0136] KL i (GT||Pre) is the KL divergence between the GT distribution and the predicted distribution pre;

[0137] H i (gt, pre) is the predicted distribution on GT;

[0138] H i (gt) is the actual distribution on GT.

[0139] S24-3. Represent the relative entropy loss of the discrete marginal distribution of the pixel points of the input image I through the following formula, complete the determination of the hotspot area where the tracking target is located, and improve the attention of the Swin Transformer encoder to the hotspot area. The expression is:

[0140] Loss KL =(KL i (GT||Pre)+KL j (GT||Pre)) / 2

[0141] In the formula,

[0142] Loss KL is the position loss;

[0143] KL i (GT||Pre) is the KL divergence loss on the row;

[0144] KL j (GT||Pre) is the KL divergence loss on the column.

[0145] S25. The target association module is used to match the detection result of the current frame in the Swin Transformer encoder-decoder module with the tracking result of the previous frame in the hot spot area judgment module.

[0146] The target association module uses a cost matching method based on the Hungarian algorithm to match the detection result of the current frame in the Swin Transformer encoder-decoder module with the tracking result of the previous frame in the hot spot area judgment module, ensuring the correct association of the targets determined as tumors in the input image I and ensuring the continuity of the tracked targets in the input image I.

[0147] Constructing the target association module includes the following steps:

[0148] S25-1. Calculate the matching cost matrix of the detection result of the current frame in the Swin Transformer encoder-decoder module and the tracking result of the previous frame in the hot spot area judgment module through the following formula:

[0149] C ij =||Q' i -T t-1,j ||+λ·D ij

[0150] In the formula,

[0151] C ij is the matching cost of the i-th query target and the j-th trajectory target;

[0152] Q' i is the query embedding of the i-th target;

[0153] T t-1,j is the j-th trajectory embedding;

[0154] Dij is the distance between the targets;

[0155] λ is the weight coefficient.

[0156] S25-2, match the cost matrix through the cost matching method of the following Hungarian algorithm, match the detection results of the current frame in the SwinTransformer encoder-decoder module with the tracking results of the previous frame in the hot spot area judgment module, and output the associated target trajectory T t , the expression is:

[0157] T t = Hungarian(C)

[0158] In the formula, T t is the trajectory associated with the current frame.

[0159] S26, the target segmentation module is used to segment the image area determined to be a tumor in the input image I according to the matching result of the detection result of the current frame in the Swin Transformer encoder-decoder module and the tracking result of the previous frame in the hot spot area judgment module, and generate a target segmentation mask according to the target trajectory, the expression is:

[0160] M = SegmentationHead(T t , F” i )

[0161] In the formula,

[0162] M is the output target segmentation mask, and its size is the same as that of the input image I;

[0163] T t is the trajectory associated with the current frame;

[0164] F” i is the feature map providing semantic information.

[0165] S3, divide the 3D-IRCADb organ segmentation dataset obtained in step S1 into a training set and a test set. The training set and the test set can be divided according to a ratio, which will not be elaborated here.

[0166] S4, use the training set to train the hot spot area joint tracking model. The training steps are briefly described as follows:

[0167] S41, preprocess the images in the training set, including but not limited to image normalization and data augmentation;

[0168] S42, select the loss function combination of FCL+MCE+DICE;

[0169] S43. Adjust the learning rate to ensure training convergence and obtain the optimal solution as much as possible.

[0170] S5. Use the test set to test the hotspot region joint tracking model, measure the detection ability and tracking rate of the model for liver tumors in medical images, and compare the test results of the hotspot region loss function in step S24-3 with other loss functions for qualitative and quantitative analysis. Other loss functions include but are not limited to loss functions such as CE, DICE, FCL, CE+DICE, CE+MCE+DICE, FCL+MCE, FCL+MCE+DICE, etc.

[0171] Specifically, it includes the following steps:

[0172] S51. Use the test set to test the hotspot region joint tracking model to obtain the tracked image, mAP50 evaluation index, and mAP50-95 evaluation index, and complete the qualitative analysis of the hotspot region joint tracking model.

[0173] The mAP50 evaluation index measures the average precision of the hotspot region joint tracking model when the IoU threshold is 0.5, ranging from 0 to 1. The closer the value is to 1, the better the detection effect.

[0174] mAP50-95 is a more comprehensive model performance evaluation index, which represents the performance of the model at different IoU levels. mAP50-95 is more challenging than mAP50 because it requires the model to have good performance at a wider range of IoU levels.

[0175] S52. Obtain the segmented liver tumor image and three segmentation indexes of MIoU, DICE, and MPA through the hotspot region joint tracking model to complete the quantitative analysis of the hotspot region joint tracking model.

[0176] MIoU is the overlapping area between the predicted segmentation and the label divided by the union area between the predicted segmentation and the label, ranging from 0 to 1. The closer the value is to 1, the better the segmentation effect.

[0177] DICE, also known as the F1 score, is a set similarity measurement function, which is defined as twice the intersection divided by the pixel sum. This index is very similar to IoU and they are positively correlated, also ranging from 0 to 1. The closer the value is to 1, the better the segmentation effect.

[0178] MPA is the average pixel accuracy, which is used to measure the pixel-level accuracy between the model segmentation result and the true label, also ranging from 0 to 1. The closer the value is to 1, the better the segmentation effect.

[0179] S6. Obtain the medical image to be segmented and perform image segmentation based on the trained hotspot region joint tracking model.Figure 3 It shows the tracking results of the constructed model for liver tumors in medical images and the schematic diagram of the segmented images.

[0180] S7. Conduct an effectiveness verification on the hotspot area determination module (MCE) of the constructed combined tracking model for hotspot areas. For ease of description, the hotspot area determination module is abbreviated as MCE.

[0181] S71. After adding MCE, use the BCE loss function and the FCL loss function to obtain the line chart for verifying the effectiveness of the BCE loss after adding MCE, and the line chart for verifying the effectiveness of the FCL loss after adding MCE.

[0182] Appendix Figure 4 It shows the results of BCE Loss after adding MCE. After 70 epochs, the curves of the MIOU and DICE coefficients of the model with MCE start to be higher than those of BCE and maintain the advantage. This result shows that MCE is slightly different from BCE in the optimization objective, indicating that the position features are different from the positions represented by BCE.

[0183] Appendix Figure 5 It shows the results of Focal Loss after adding MCE. From the curve, it can be seen that the average IOU and DICE coefficients are higher than those of Focal Loss after adding MCE Loss. After 300 epochs, the final result is improved by 4.8%, and it converges earlier than Focal Loss. All these fully demonstrate the effectiveness of MCE.

[0184] S72. Use the 3D-IRCADb organ segmentation dataset to conduct segmentation tests on the combined tracking model for hotspot areas with 7 different loss functions, namely CE, DICE, FCL, CE + DICE, CE + MCE + DICE, FCL + MCE, and FCL + MCE + DICE, to verify the effectiveness of MCE. The comparison results are as shown in the appendix Figure 6 as follows.

[0185] From the appendix Figure 6 it can be seen that after adding MCE, whether using CE + DICE or FCL + DICE, all test results have improved; especially when using the FCL + DICE loss function, all three segmentation indicators have increased by more than 4%. This fully demonstrates the effectiveness of MCE.

[0186] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and all these changes and improvements fall within the scope of the present invention claimed.

Claims

1. A medical image segmentation method based on a joint tracking model of hot regions, characterized in that, Including the following steps: S1. Establish an image dataset of existing medical images, where the images in the image dataset are input images; S2. Establish a hot spot area joint tracking model through a Python compiler, including a feature extraction module, a coordinate position encoding module, a Swin Transformer encoder-decoder module, a hot spot area determination module, a target association module, and a target segmentation module; S21. The feature extraction module can extract high-level semantic features of the input image and output a feature map; S22. The coordinate position encoding module can add position information to the feature map of the input image; S23. The Swin Transformer encoder-decoder module can encode the target information in the input image into a feature representation of global information, generate the mutual relationship between target information, and perform target prediction and association in the input image through a query vector; S24. The hot spot area determination module can determine the hot spot area where the tracking target is located in the input image; S25. The target association module can match the detection result of the current frame in the Swin Transformer encoder-decoder module with the tracking result of the previous frame in the hot spot area judgment module; S26. The target segmentation module can segment the image area determined as an abnormal area in the input image according to the matching result of the detection result of the current frame in the Swin Transformer encoder-decoder module and the tracking result of the previous frame in the hot spot area judgment module.

2. The medical image segmentation method based on the hotspot region joint tracking model according to claim 1, wherein, The specific steps for constructing the feature extraction module include the following: S21-1. Set ResNet50 as the backbone of the convolutional neural network; S21-2. Input the input image into the convolutional neural network, and extract the basic features of the input image through preliminary convolutional operations; S21-3. Through multi-layer convolutional processing of the convolutional neural network, extract the high-level semantic features of the input image and output a multi-scale feature map. The expression of the multi-scale feature map is: F = {F1, F2, ..., F n} The number of channels of the feature map output by each convolutional layer of the convolutional neural network increases with the depth of the convolutional neural network. The expression is: F i = Conv(I), i = 1,..., n Where, I is the input image, and its size is H×W×C in ; F i is the feature map, and its size is H i ×W i ×C i .

3. The medical image segmentation method based on the hotspot region joint tracking model according to claim 1, wherein, The specific steps for constructing the coordinate position encoding module include the following: S22-1. Use sine-cosine position encoding to add a pair of position vectors to each pixel of the input image; S22-2. Add the feature map and the position encoding to generate an enhanced feature map, where the position encoding is a position information matrix with the same size as the feature map.

4. The medical image segmentation method based on the hotspot region joint tracking model according to claim 3, wherein The expression of the position encoding vector for the x-th row and y-th column of the position encoding sequence in sine-cosine encoding is: where P i (x, y, 2k) is the position encoding P i vector of the x-th row of the sequence; P i (x, y, 2k + 1) is the position encoding P i The position encoding vector of the y-th column of the sequence; k is the channel index; C i is the number of channels of the feature map.

5. A medical image segmentation method based on a hotspot region joint tracking model according to claim 3, characterized in that, The specific steps for constructing the Swin Transformer encoder-decoder module include the following: S23-1. Perform a window partitioning operation on the enhanced feature map through the following formula: W = Split(F’ i ) Wherein, W is the set of windows after partitioning, and the size of each window is M×M×C i ; F' i is the enhanced feature map; S23-2. Perform self-attention operation within the window on the partitioned window set. The expression is: W’=Self-Attention(W) S23-3. Based on the window-based multi-head self-attention mechanism, extract the local and global features of the input image, and the expression is: where Q, K, and V are the Query, Key, and Value matrices respectively; d k is the key vector dimension; S23-4. Combine the window features to generate a global feature map, and the expression is: F″ i = Merge(W′) In the formula, F″ i is the global feature map, and its size is the same as that of the input enhanced feature map F′ i identical.

6. A medical image segmentation method based on a hot spot region joint tracking model according to claim 1, characterized in that Constructing the hot spot area determination module includes the following specific steps: S24-1. Normalize all pixel points in the input image to obtain the edge distribution of rows or columns, and the expression is: where P i (x) is the probability distribution of the row; is the probability that each pixel is an outlier; i is the row number of the image pixel; j is the column number of the image pixel; S24-2. According to the discrete edge distribution of the pixel points in the input image, calculate the KL divergence between the GT distribution and the predicted distribution of columns: KL i (GT||Pre) = H i (gt, pre)-H i (gt) where KL i (GT||Pre) is the KL divergence of the GT distribution from the predicted distribution pre; H i (gt, pre) is the predicted distribution on GT; H i (gt) is the actual distribution on GT; S24-3. Represent the relative entropy loss of the discrete edge distribution of the pixel points in the input image through the following formula, and complete the determination of the hot spot area where the tracking target is located: Loss KL =(KL i (GT||Pre)+KL j (GT||Pre)) / 2 where Loss KL is the position loss; KL i (GT||Pre) is the KL divergence loss on the row; KL j (GT||Pre) is the KL divergence loss on the column.

7. A medical image segmentation method based on a hot region joint tracking model according to claim 1, characterized in that Constructing the target association module includes the following specific steps: S25-1. Calculate the matching cost matrix between the detection result of the current frame in the Swin Transformer encoder-decoder module and the tracking result of the previous frame in the hot spot area judgment module through the following formula: C ij = ||Q’ i - T t-1,j || + λ·D ij where C ij is the matching cost between the i-th query target and the j-th trajectory target; Q' i is the query embedding for the i-th target; T t-1,j is the j-th trajectory embedding; D ij is the distance between the targets; λ is the weight coefficient; S25-2. Match the cost matrix through the following formula, match the detection result of the current frame in the Swin Transformer encoder-decoder module with the tracking result of the previous frame in the hot spot area judgment module, and output the associated target trajectory: T t = Hungarian(C) where, T t is the trajectory associated with the current frame.

8. A medical image segmentation method based on a hot region joint tracking model according to claim 7, characterized in that The target segmentation module generates a target segmentation mask according to the target trajectory, and the expression is: M = SegmentationHead(T t , F” i ) where M is the output target segmentation mask; T t is the trajectory associated with the current frame; F” i is the global feature map.

9. A medical image segmentation method based on a hotspot region joint tracking model according to any one of claims 1-8, characterized in that, The medical image segmentation method further includes the following steps: S3. Make a training set and a test set to train and test the hot spot area joint tracking model respectively; S4. Obtain the medical image to be segmented, and perform image segmentation based on the trained hot spot area joint tracking model.

10. The medical image segmentation method based on a hotspot region joint tracking model according to claim 9, wherein, Testing the hot spot area joint tracking model using the test set includes the following steps: S31. After testing the hot spot area joint tracking model using the test set, obtain the tracked image, mAP50 evaluation index, and mAP50-95 evaluation index, and complete the qualitative analysis of the hot spot area joint tracking model; S32. Obtain the image of the segmented abnormal area through the hot spot area joint tracking model, as well as three segmentation indexes of MIoU, DICE, and MPA, and complete the quantitative analysis of the hot spot area joint tracking model.