Kidney tumor image segmentation method and device based on multi-scale dynamic segmentation kernel and storable medium

The problem of effective feature enhancement, localization, and boundary clarity in kidney tumor image segmentation was solved by using a multi-scale dynamic segmentation kernel network, thus achieving accurate segmentation of kidney tumors.

CN117745696BActive Publication Date: 2026-05-15ANHUI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ANHUI UNIV
Filing Date
2023-12-27
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively extract the low-contrast differences between the kidney and kidney tumors, making it difficult to locate kidney tumors in highly similar segmentation backgrounds, and resulting in unclear segmentation boundaries for kidney tumors.

Method used

A multi-scale dynamic segmentation kernel network, including a backbone extraction network, a spatially enhanced receptive field network, an enhanced cross-attention network, and a multi-scale dynamic segmentation kernel network, is used to segment kidney tumor images.

Benefits of technology

It improves the effective feature enhancement of kidney tumor images, accurately locates kidney tumors in highly similar segmentation backgrounds, and enhances the clarity of kidney tumor segmentation boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117745696B_ABST
    Figure CN117745696B_ABST
Patent Text Reader

Abstract

The application discloses a kidney tumor image segmentation method and device based on a multi-scale dynamic segmentation kernel and a storable medium, relates to the technical field of image processing, and comprises the following steps: acquiring a kidney tumor image dataset, and dividing the dataset into a training set and a test set; constructing a multi-scale dynamic segmentation network model, training the multi-scale dynamic segmentation network model by using the training set, and testing the multi-scale dynamic segmentation network model by using the test set after training; and outputting the trained and tested multi-scale dynamic segmentation network model, which is used for subsequent kidney tumor image segmentation. The application solves the problems of effective feature enhancement of a kidney tumor image, positioning of a kidney tumor in a highly similar segmentation background, and the definition of a kidney tumor segmentation boundary.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically to a method, apparatus, and storage medium for kidney tumor image segmentation based on a multi-scale dynamic segmentation kernel. Background Technology

[0002] Currently, medical image segmentation is a technique for pixel-level classification of organs or lesions. However, the low contrast between the kidney and other nearby tissues and organs makes it difficult to accurately extract and identify the details and effective features of the kidney and kidney tumors. Furthermore, kidney tumors vary greatly in size, shape, number, and location within the body, making it challenging to locate them against a highly similar segmentation background. Finally, the presence of residual adipose tissue on kidney tumors, as well as interference from detection equipment and lighting conditions, often results in unclear tumor boundaries on the image, further complicating kidney tumor segmentation.

[0003] However, with the continuous advancement of deep learning technology in recent years, neural network-based segmentation methods have gradually replaced traditional methods in the field of medical image segmentation and achieved remarkable results. Currently, research on renal tumor segmentation based on neural networks still faces the following challenges: effective feature enhancement of renal tumor images, localization of renal tumors in highly similar segmentation backgrounds, and clarity of renal tumor segmentation boundaries.

[0004] Therefore, how to provide a kidney tumor image segmentation method that can solve the above problems is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a method, apparatus and storage medium for kidney tumor image segmentation based on multi-scale dynamic segmentation kernel, which solves the problems of effective feature enhancement of kidney tumor images, localization of kidney tumors in highly similar segmentation backgrounds and clarity of kidney tumor segmentation boundaries.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A kidney tumor image segmentation method based on multi-scale dynamic segmentation kernels includes the following steps:

[0008] Obtain a dataset of kidney tumor images and divide the dataset into training and testing sets;

[0009] A multi-scale dynamic segmentation network model is constructed, and the multi-scale dynamic segmentation network model is trained using the training set. After training, the multi-scale dynamic segmentation network model is tested using the test set.

[0010] The trained and tested multi-scale dynamic segmentation network model is output for subsequent kidney tumor image segmentation.

[0011] Preferably, the multi-scale dynamic segmentation network model includes: a backbone extraction network, a spatially enhanced receptive field network, an enhanced cross-attention network, and a multi-scale dynamic segmentation kernel network.

[0012] Preferably, the backbone extraction network includes multiple pyramid feature structures for extracting backbone features.

[0013] Preferably, the spatially enhanced receptive field network includes: a dilated convolutional layer, an asymmetric convolutional layer, and a self-correcting convolutional layer, used to enhance the backbone features to obtain enhanced spatial features.

[0014] Preferably, the enhanced cross-attention network includes a cross-attention module, a dense context feature module, and a local emphasis submodule, used to process the enhanced spatial features to obtain multi-scale fusion features.

[0015] Preferably, the multi-scale dynamic segmentation kernel network includes multiple dynamic segmentation kernel update sub-modules for updating the multi-scale fusion features to obtain the kidney tumor image segmentation result.

[0016] Preferred options also include:

[0017] Select the average dice similarity coefficient, average intersection-union ratio, average absolute error, recall, precision, and F-value. β Indicators are used to quantitatively evaluate the performance of multi-scale dynamic segmentation network models.

[0018] The present invention also provides a segmentation apparatus for kidney tumor image segmentation based on multi-scale dynamic segmentation kernel as described in any one of the above claims, comprising:

[0019] The acquisition module is used to acquire a kidney tumor image dataset and divide the dataset into a training set and a test set.

[0020] A construction module is used to construct a multi-scale dynamic segmentation network model, train the multi-scale dynamic segmentation network model using the training set, and test the multi-scale dynamic segmentation network model using the test set after training is completed.

[0021] The segmentation module outputs a trained and tested multi-scale dynamic segmentation network model for subsequent kidney tumor image segmentation.

[0022] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the kidney tumor image segmentation method as described in any of the preceding claims.

[0023] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method, device, and storage medium for kidney tumor image segmentation based on multi-scale dynamic segmentation kernels, which has the following advantages:

[0024] (1) This invention utilizes a backbone network to extract more effective backbone features from kidney tumor images, thereby obtaining more local continuity of images and feature maps and more flexible processing of variable resolution inputs.

[0025] (2) This invention simulates the human receptive field by using a spatially enhanced receptive field network (SRF) and further enhances the renal tumor features extracted from the backbone network in the spatial dimension using receptive fields of different sizes.

[0026] (3) This invention utilizes the enhanced cross-attention network (ECA) to generate denser contextual information and fuses it on multi-scale features, which improves the problem of attention distraction and effectively reduces the interference of highly similar segmentation backgrounds on the segmentation results, making the localization of kidney tumors in the segmentation results more accurate.

[0027] (4) This invention introduces a multi-scale dynamic segmentation kernel network (DKS) to improve the clarity of the segmentation boundary of kidney tumors. The segmentation kernel is updated based on the segmentation results with low pixel size, and the edge fineness of the segmentation results is gradually improved.

[0028] (5) The present invention improves the problem of effective feature enhancement of kidney tumor images, the problem of localization of kidney tumors in highly similar segmentation backgrounds, and the problem of clarity of kidney tumor segmentation boundaries. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0030] Figure 1 The overall flowchart of the kidney tumor image segmentation method based on multi-scale dynamic segmentation kernel provided by the present invention;

[0031] Figure 2 The network structure diagram of the multi-scale dynamic segmentation network model provided by this invention;

[0032] Figure 3 A structural diagram of the spatially enhanced receptive field network provided by this invention;

[0033] Figure 4The structural diagram of the enhanced cross-attention network provided by this invention;

[0034] Figure 5 This is a structural diagram of the multi-scale dynamic segmentation kernel network provided by the present invention;

[0035] Figure 6 The diagram shows the results of a qualitative comparison between this embodiment of the invention and other typical model methods.

[0036] Figure 7 The structural principle block diagram of the kidney tumor image segmentation device based on multi-scale dynamic segmentation kernel provided by the present invention is shown. Detailed Implementation

[0037] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0038] See appendix Figure 1 As shown in the figure, this invention discloses a method for kidney tumor image segmentation based on a multi-scale dynamic segmentation kernel, comprising the following steps:

[0039] Obtain a dataset of kidney tumor images and divide the dataset into training and testing sets;

[0040] Construct a multi-scale dynamic segmentation network model, train the multi-scale dynamic segmentation network model using the training set, and test the multi-scale dynamic segmentation network model using the test set after training.

[0041] The trained and tested multi-scale dynamic segmentation network model is output for subsequent kidney tumor image segmentation.

[0042] See appendix Figure 2 As shown, in a specific embodiment, the multi-scale dynamic segmentation network model includes: a backbone extraction network PVTv2, a spatially enhanced receptive field network SRF, an enhanced cross-attention network ECA, and a multi-scale dynamic segmentation kernel network DKS.

[0043] In one specific embodiment, the backbone extraction network PVTv2 includes multiple pyramid feature structures for extracting backbone features.

[0044] Specifically, the backbone extraction network PVTv2 can include four multi-scale pyramid feature structures f1, f2, f3, and f4, with the number of channels for each pyramid feature structure being 64, 128, 320, and 512, respectively. The backbone extraction network PVTv2 has the same linear complexity as CNN, can achieve stronger local continuity, and can handle variable resolution inputs more flexibly. The height and width dimensions of the output feature f1 are one-quarter of the input image, and the height and width dimensions of the output features in subsequent stages are successively reduced to half of the previous stage. As the feature scale decreases and the number of channels increases, f2, f3, and f4 will have richer semantic information.

[0045] See appendix Figure 3 As shown, in a specific embodiment, the Spatial Enhanced Receptive Field Network (SRF) includes: dilated convolutional layers, asymmetric convolutional layers, and self-calibrating convolutional layers, which are used to enhance the backbone features to obtain enhanced spatial features. The dilated convolutions are selected as 3×3 convolutions with an expansion rate of 3, the asymmetric convolutions are 1×3, 3×1, and 3×3 convolutions, and the self-calibrating convolutions fuse information from two different spatial scales through self-calibration operations.

[0046] Specifically, the Spatial Enhanced Receptive Field Network (SRF) simulates human visual perception by using receptive fields of different sizes to obtain stronger multi-scale contextual features. The backbone features extracted by the PVTv2 backbone extraction network are input into three branches: dilated convolution, asymmetric convolution, and self-correcting convolution. This yields multi-scale contextual feature information extracted under different receptive fields. The number of channels in the features remains constant during these three types of convolutions. Then, the features output from the three branches are fused along the channel dimension, resulting in a feature channel count three times that of the input features. Next, a 3×3 convolution reduces the feature channel count back to the original input channel count to reduce subsequent computational resources. Finally, after batch normalization and activation functions, the enhanced spatial features S are obtained after the SRF enhancement. i The specific expression is:

[0047] S i =Concat{Conv at (f i ),Conv sc (f i ),Conv as (f i )} (1)

[0048] In the formula, Concat(·) represents the merge operation, and Conv at (·) denotes dilated convolution, Conv sc (·) denotes self-correcting convolution, Convas (·) indicates asymmetric convolution.

[0049] See appendix Figure 4 As shown, in a specific embodiment, the Enhanced Cross-Attention Network (ECA) includes a cross-attention module, a dense context feature module, and a local emphasis submodule, used to process the enhanced spatial features to obtain multi-scale fused features E. k Furthermore, there can be multiple Enhanced Cross-Attention Networks (ECAs).

[0050] Specifically, the Enhanced Cross-Attention Network (ECA) focuses on denser multi-scale contextual information to effectively reduce interference from segmentation backgrounds highly similar to renal tumors. First, cross-attention is used to acquire dense contextual information of small-sized features in both horizontal and vertical directions, effectively reducing computational resource consumption compared to pixel-by-pixel non-local operations. Then, the dense contextual features are fused with the receptive field enhancement features of this stage along the channel dimension. The fused features enter the local emphasis submodule to prevent attention dispersion or collapse after a series of attention operations, refocusing attention on adjacent features. Simultaneously, the local emphasis module converts the feature channel numbers of the four stages to 64, 128, 256, and 512 respectively. Finally, the upsampling operation unifies the height and width dimensions of the small-sized features in the current stage with the larger-sized features in the previous stage, facilitating subsequent merging operations. The specific expression is as follows:

[0051] E i =L e (Concat{S i Criss(E) i+1 (2)

[0052] L e (f)=up(Relu(Conv(Relu(Conv(f))))) (3)

[0053] In the formula, Concat(·) represents the merge operation, and L e (·) indicates a local emphasis operation, S i The symbol represents spatial augmentation features, Criss(·) represents the cross-attention operation, Relu(·) represents the ReLU activation function operation, Conv(·) represents the 3×3 convolution, and up(·) represents the upsampling operation.

[0054] The multi-scale fusion feature E4, which is the final stage, is generated using the following method:

[0055] E4 = L e (S4) (4)

[0056] See appendix Figure 5 As shown, in a specific embodiment, the multi-scale dynamic segmentation kernel network (DKS) includes multiple dynamic segmentation kernel update submodules for updating multi-scale fusion features to obtain kidney tumor image segmentation results.

[0057] Specifically, the multi-scale dynamic segmentation kernel network (DKS) adopts a method of updating the dynamic segmentation kernel. The segmentation kernel is updated iteratively at each stage. The small-pixel-size predicted segmentation result of the previous stage and the multi-scale fusion features of the current stage are updated by the dynamic segmentation kernel update submodule (DU). The segmentation kernel DK can obtain more accurate location and edge information of the kidney tumor from this. The obtained dynamic segmentation kernel DK is convolved with the multi-scale fusion features of the current stage to obtain a segmentation result with larger pixel size and clearer and more accurate tumor boundaries. Through the gradual update of the dynamic kernel in four stages, the predicted segmentation result steadily approaches the real kidney tumor region.

[0058] Furthermore, the dynamic segmentation kernel update submodule DU integrates multi-scale fusion features E. i The number of channels is unified to 64 using 1×1 convolution, while the small-size prediction segmentation result P from the previous stage is also optimized. i+1 Upsampling was performed to adapt to E. i The size is calculated and then multiplied element-wise, resulting in F. i Compared with the dynamic segmentation kernel K in the previous stage i+1 The input, after linear transformation, is passed to the Gate operation to obtain the weights for the features and the dynamic kernel. and Feature F and dynamic kernel K i+1 The elements are added together according to this weight, and finally a linear transformation is performed to obtain the dynamic segmentation kernel K for the current stage. i The specific expression is:

[0059]

[0060]

[0061]

[0062]

[0063] In the formula, uni(·) represents a 1×1 convolution with 64 output channels, and up(·) represents an upsampling operation. The symbols represent element-wise multiplication, φ1(·), φ2(·), and φ3(·) represent linear transformations, and Ψ1(·) and Ψ2(·) represent linear transformations followed by the addition of the sigmoid function.

[0064] In a specific embodiment, the method further includes: the loss function is designed as a combination of the binary cross-entropy (BCE) loss function and the dice loss function, and the predicted segmentation result map of each dynamic segmentation kernel module adopts a deep supervision method as the optimization objective. The specific expression of the loss function is as follows:

[0065] L = L1 + L2 + L3 + L4 (9)

[0066] L i =L bce (P i ,GT)+L dice (P i ,GT) (10)

[0067] In the formula, L bce L dice Let P represent the binary cross-entropy loss and dice loss, respectively. Let GT represent the truth graph. i This represents the predicted segmentation result at the current size.

[0068] In one specific embodiment, it also includes:

[0069] The following parameters were selected: mean dice similarity coefficient (mDice), mean intersection-union ratio (mIoU), mean absolute error (MAE), recall, precision, and F-squared. β The performance of the multi-scale dynamic segmentation network model is quantitatively evaluated using the above metrics, and the specific expressions for these metrics are as follows:

[0070]

[0071]

[0072]

[0073]

[0074]

[0075]

[0076] In the formula, TP represents a true positive, FP represents a false positive, FN represents a false negative, and n represents the number of test images. and p i This represents the prediction and the corresponding ground truth value of the i-th pixel out of a total of n pixels, where β is set to 2.

[0077] See appendix Figure 7As shown, this embodiment of the invention also provides a segmentation apparatus for a kidney tumor image segmentation method based on a multi-scale dynamic segmentation kernel according to any one of the above embodiments, comprising:

[0078] The acquisition module is used to acquire a kidney tumor image dataset and divide the dataset into a training set and a test set.

[0079] The module is used to build a multi-scale dynamic segmentation network model, train the multi-scale dynamic segmentation network model using the training set, and test the multi-scale dynamic segmentation network model using the test set after training.

[0080] The segmentation module outputs a trained and tested multi-scale dynamic segmentation network model for subsequent kidney tumor image segmentation.

[0081] This invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the kidney tumor image segmentation method as described in any of the above embodiments.

[0082] To evaluate the segmentation performance of the method provided in this embodiment of the invention, a large number of endoscopic images of kidney tumors were screened and annotated to complete the establishment of training and testing datasets, which were named Re-TMRS. The dataset is described as follows:

[0083] The Re-TMRS dataset consists of 2823 endoscopic images of kidney tumors and their corresponding baseline truth kidney tumor masks. The images in the dataset primarily have a resolution of 1920×1080 pixels, with some images at 1440×1080 pixels. To ensure fair comparison, this embodiment of the invention divides the dataset into a training set and a test set according to a certain ratio, used for training and evaluation of this method against other methods. Specifically, the training set contains 2258 kidney tumor images and their corresponding baseline truth kidney tumor masks, while the test set contains the remaining 565 kidney tumor images and their corresponding baseline truth kidney tumor masks.

[0084] In addition, to further analyze the generalization performance of the method provided in this embodiment of the invention, experiments were also conducted on five existing polyp segmentation datasets, including Kvasir-SEG, CVC-ClinicDB, CVC-ColonDB, ETIS, and CVC-T. The datasets are described below:

[0085] Kvasir-SEG contains 1000 polyp images and corresponding annotations; CVC-ClinicDB contains 612 polyp images and corresponding annotations; CVC-ColonDB contains 380 polyp images and corresponding annotations; ETIS contains 192 polyp images and corresponding annotations; CVC-T is a subset of EndoScene, containing 60 polyp images and corresponding annotations.

[0086] To ensure a fair comparison, the generalization performance test of this method uses the same training and test set partitioning as other methods. The training set consists of 900 images from Kvasir-SEG and 550 images from CVC-ClinicDB, totaling 1450 samples. The test set comprises five sets: the remaining 100 images from Kvasir-SEG; the remaining 62 images from CVC-ClinicDB; the complete CVC-ColonDB dataset (380 images); the complete ETIS dataset (196 images); and the CVC-300 dataset (60 images). Specific information on these five polyp datasets and the kidney tumor dataset is shown in Table 1 below.

[0087] Table 1 Dataset Information

[0088]

[0089] This invention employs three experiments to verify the performance of the kidney tumor segmentation method: qualitative and quantitative experiments, and ablation experiments. In these experiments, the model is compared with six different models, including UNet, UNet++, CaraNet, DCRNet, Polyp-PVT, and LDNet. Furthermore, this invention analyzes and compares the performance of each method across three evaluation metrics on a kidney tumor segmentation dataset and a publicly available polyp dataset.

[0090] In the comparative experiment on renal tumor segmentation, the results are shown in the appendix. Figure 6 As shown in Table 2, the quantitative results of the statistical comparison between this model and six different segmentation methods on the kidney tumor dataset are presented. Figure 6 As can be seen from Table 2, compared with other segmentation methods, the model provided by the embodiments of the present invention can more accurately segment the kidney tumor region after image processing.

[0091] Table 2 Comparison of experimental results for different models

[0092]

[0093] To verify the generalization performance of the model, comparative tests were also conducted on the publicly available polyp dataset. Table 3 lists the quantitative results of the model's statistical comparison with six different segmentation methods on the publicly available polyp dataset. The best result for each evaluation metric is highlighted in bold. Compared with the six different segmentation methods, the model proposed in this embodiment also achieved good generalization performance on the publicly available polyp dataset, especially on the more difficult segmentation datasets CVC-ColonDB and ETIS, where the generalization ability of the model provided by this embodiment is significantly improved compared with other models. On CVC-ColonDB, the mDice of the model provided by this embodiment outperforms Polyp-PVT and LDNet by 7.6% and 3.6%, respectively. On ETIS, the mDice of the model provided by this embodiment exceeds Polyp-PVT and LDNet by 4.9% and 3.4%, respectively. On CVC-T, the model provided by this embodiment outperforms Polyp-PVT and LDNet by 2.1% and 0.2%, respectively.

[0094] Table 3 Comparison of experimental results of different methods on public polyp datasets.

[0095]

[0096]

[0097] To validate the effectiveness of each module in the model, ablation experiments were conducted on the spatially enhanced receptive field network (SRF), the enhanced cross-attention network (ECA), and the multi-scale dynamic segmentation kernel network (DKS) on a kidney tumor dataset. The baseline of the model's backbone network consists of PVTv2 and a static segmentation kernel, while the standard model consists of "Baseline + SRF + ECA + DKS". The effectiveness of different modules was evaluated by removing or modifying different modules from the standard model, and the experimental results are shown in Table 4 below.

[0098] Table 4 Ablation experimental results on the kidney tumor dataset.

[0099]

[0100] As shown in Table 4, to quantitatively analyze the effectiveness of the Spatial Enhancement Receptive Field Network (SRF), this embodiment of the invention trained a version lacking the SRF. The SRF was completely removed, and the features extracted by the backbone network were directly fed into the next module. Compared to the standard model, the model without the SRF showed a sharp decline in performance on the kidney tumor dataset, with both mDice and mIou decreasing by 0.5%, demonstrating the excellent enhancement effect of the SRF on the features extracted by the backbone network.

[0101] To quantitatively analyze the effectiveness of the ECA network, this embodiment of the invention trained a version lacking the enhanced cross-attention network (ECA). The ECA was completely removed, and a combination of a 3×3 convolution and upsampling was used for multi-scale fusion. Compared to the standard model, the model without ECA reduced mDice and mIou by 0.2% and 0.3%, respectively, on the kidney tumor dataset.

[0102] To quantitatively analyze the effectiveness of the multi-scale dynamic segmentation kernel network DKS, the entire DKS network was removed and replaced with a static segmentation kernel, i.e., four 1×1 convolutions were used to segment images of four different sizes. Compared with the standard model, the model without DKS reduced mDice and mIou by 0.1% and 0.2%, respectively, on the kidney tumor dataset, and decreased the precision metric by 0.8%.

[0103] In summary, this invention utilizes a backbone network to extract more effective backbone features from kidney tumor images, simulates the human receptive field using a Spatial Enhanced Receptive Field Network (SRF), generates denser contextual information using an Enhanced Cross-Attention Network (ECA) and fuses it across multiple scale features, and introduces a multi-scale dynamic segmentation kernel network (DKS) to improve the clarity of kidney tumor segmentation boundaries, thereby better improving the segmentation performance of kidney tumors.

[0104] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0105] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A kidney tumor image segmentation method based on multi-scale dynamic segmentation kernel, characterized in that, Includes the following steps: Obtain a dataset of kidney tumor images and divide the dataset into training and testing sets; A multi-scale dynamic segmentation network model is constructed, and the multi-scale dynamic segmentation network model is trained using the training set. After training, the multi-scale dynamic segmentation network model is tested using the test set. The multi-scale dynamic segmentation network model includes: a backbone extraction network, a spatially enhanced receptive field network, an enhanced cross-attention network, and a multi-scale dynamic segmentation kernel network. The spatial enhancement receptive field network includes: dilated convolutional layers, asymmetric convolutional layers, and self-calibrating convolutional layers, used to enhance the backbone features to obtain enhanced spatial features; the enhanced cross-attention network includes a cross-attention module, a dense context feature module, and a local emphasis submodule, used to process the enhanced spatial features to obtain multi-scale fusion features; the multi-scale dynamic segmentation kernel network includes multiple dynamic segmentation kernel update submodules, used to update the multi-scale fusion features to obtain the kidney tumor image segmentation result. The multi-scale dynamic segmentation kernel network adopts a dynamic segmentation kernel update method, updating the segmentation kernel at each stage in an iterative manner; The trained and tested multi-scale dynamic segmentation network model is output for subsequent kidney tumor image segmentation.

2. The kidney tumor image segmentation method based on multi-scale dynamic segmentation kernel according to claim 1, characterized in that, The backbone extraction network includes multiple pyramid feature structures for extracting backbone features.

3. The kidney tumor image segmentation method based on multi-scale dynamic segmentation kernel according to claim 1, characterized in that, Also includes: The average dice similarity coefficient, average intersection-union ratio, average absolute error, recall, precision, and Fβ index are selected to quantitatively evaluate the performance of the multi-scale dynamic segmentation network model.

4. A segmentation apparatus for kidney tumor image segmentation based on a multi-scale dynamic segmentation kernel as described in any one of claims 1-3, characterized in that, include: The acquisition module is used to acquire a kidney tumor image dataset and divide the dataset into a training set and a test set. A construction module is used to construct a multi-scale dynamic segmentation network model, train the multi-scale dynamic segmentation network model using the training set, and test the multi-scale dynamic segmentation network model using the test set after training is completed. The segmentation module outputs a trained and tested multi-scale dynamic segmentation network model for subsequent kidney tumor image segmentation.

5. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the kidney tumor image segmentation method as described in any one of claims 1 to 3.