Gastric cancer CT lesion area detection method and device based on ZSNet
By adopting ZSNet-based detection methods and CT-Neck module with ODConv technology in gastric cancer CT image detection, the problems of missed detection, missed detection and inaccurate positioning in micro lesions are solved, and higher detection accuracy and robustness are achieved, providing more effective technical support for clinical diagnosis.
Patent Information
- Application Number
- CN202510507064.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The prior art has problems of missed detection, misdetection and inaccurate positioning when detecting tiny lesions in gastric cancer CT images, which affects the accuracy of clinical diagnosis.
The ZSNet-based CT lesion area detection method is used, combined with the CT-Neck module of ODConv technology, and through multi-scale feature fusion and deep feature refinement, the accuracy and robustness of the detection are improved.
It significantly improves the accuracy and robustness of detection of micro lesions, reduces the risks of missed and misdetection, and improves the comprehensive performance of detection and the accuracy of clinical diagnosis.
Smart Images

Figure CN120047435A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical images, and particularly to a method and device for detecting gastric cancer CT lesion regions based on ZSNet. Background Art
[0002] In recent years, gastric cancer has become one of the highly prevalent malignant tumors globally, and its high incidence and high mortality pose a serious threat to public health. CT imaging technology is widely used in the early diagnosis of gastric cancer due to its high resolution and good display of lesion details. However, CT images of gastric cancer usually have problems such as low contrast, much noise, tiny lesion regions with variable shapes, etc. This makes it easy for traditional image processing methods and conventional object detection models to miss or misdetect when segmenting and locating tiny lesions, thus affecting the accuracy of clinical diagnosis.
[0003] Currently, deep learning technology has achieved remarkable results in the field of object detection. However, when dealing with gastric cancer CT images, due to uneven gray distribution, blurred edges, and interference from complex backgrounds in the images, existing models often struggle to fully extract low-level details and high-level semantic information, resulting in unsatisfactory detection effects for tiny lesions. In particular, the traditional bounding box regression loss has insufficient sensitivity to target center alignment in small object detection, further limiting the performance of the model in detecting gastric cancer CT lesions.
[0004] Some studies have attempted to improve the detection performance of tiny lesions in gastric cancer CT images by introducing technical means such as improved network structures, multi-scale fusion, and dynamic sampling, and certain progress has been made. However, due to the low contrast, much noise, and tiny lesion regions in gastric cancer CT images themselves, these methods still have obvious deficiencies in feature extraction and fusion, and the phenomena of missed detection and misdetection are relatively serious. Their detection accuracy and robustness still cannot meet the strict requirements of clinical early diagnosis and treatment. Therefore, there is an urgent need for a new detection method for the characteristics of gastric cancer CT images that can fully fuse multi-scale features, enhance the capture and accurate positioning of tiny lesions, so as to improve the detection accuracy and robustness and provide more effective technical support for clinical early diagnosis and treatment.
[0005] In summary, although the methods for detecting gastric cancer CT images in the prior art have achieved object detection to a certain extent, there are still problems such as missed detection, misdetection, and inaccurate positioning, which directly affect the clinical diagnosis and treatment effects. Currently, there is an urgent need for a new method to overcome the above deficiencies and provide a more efficient and accurate technical solution for the automatic detection of gastric cancer CT images. Summary of the Invention
[0006] To address the above problems, the present invention proposes a method and device for detecting gastric cancer CT lesion regions based on ZSNet, which adopts an innovative ZSNet structure and a CT-Neck module based on ODConv technology, significantly improving the detection accuracy and robustness for tiny lesions.
[0007] On the one hand, the method for detecting gastric cancer CT lesion regions based on ZSNet is as follows:
[0008] S1. Obtain the grayscale image of the gastric cancer CT lesion region, preprocess the image, and label it to make a dataset.
[0009] S2. Construct a gastric cancer CT object detection model based on the YOLOv8 network; the gastric cancer CT object detection model includes a Backbone part, a Neck part, and a Head part.
[0010] The Backbone part includes a first unit, a second unit, a third unit, and a fourth unit connected in sequence; the second unit and the fourth unit include one or more ZSNet structures based on full-dimensional dynamic convolution; the dataset passes through the first unit, the second unit, the third unit, and the fourth unit in sequence for feature extraction, and the extracted feature representations are respectively output by the second unit and the fourth unit to the Neck part.
[0011] The Neck part performs multi-scale feature fusion on the feature representations input from the Backbone part based on the dysample structure and the ZSNet structure, and then outputs several-dimensional feature maps to the Head part.
[0012] The Head part obtains the dimensional feature maps of the Neck part, and outputs the gastric cancer CT lesion region detection image after feature enhancement by the CBS structure and feature mapping by the Con2d structure.
[0013] S3. Use the dataset to train the gastric cancer CT object detection model to obtain a trained model.
[0014] S4. Use the trained model to detect the gastric cancer CT lesion region.
[0015] Preferably, in the Backbon part, the first unit includes a first CBS structure and a second CBS structure connected in sequence, the second unit includes a first ZSNet structure, a third CBS structure, and a second ZSNet structure connected in sequence, the third unit includes a fourth CBS structure, and the fourth unit includes a third ZSNet structure, a fifth CBS structure, a fourth ZSNet structure, and an SPPF structure connected in sequence.
[0016] Preferably, the Neck part includes a dysample structure, a second dysample structure, a first Concat structure, a second Concat structure, a third Concat structure, a fourth Concat structure, a fifth ZSNet structure, a sixth ZSNet structure, a seventh ZSNet structure, an eighth ZSNet structure, a sixth CBS structure, and a seventh CBS structure; The specific implementation of the Neck part is as follows:
[0017] The feature representation output by the SPPF structure is sampled by the first dysample structure. After feature fusion with the feature representation output by the third ZSNet structure in the first Concat structure, it sequentially passes through the fifth ZSNet structure and the second dysample structure, and is dimensionally concatenated with the feature representation output by the second unit in the second Concate structure. The result after concatenation passes through the sixth ZSNet structure and then outputs the first-dimensional feature map to the Head part;
[0018] The output of the sixth ZSNet structure passes through the sixth CBS structure and is dimensionally concatenated with the output of the fifth ZSNet structure in the third Concat structure. After passing through the seventh ZSNet structure, it outputs the second-dimensional feature map to the Head part;
[0019] The output of the seventh ZSNet structure passes through the seventh CBS structure and is concatenated with the feature representation output by the SPPF structure in the fourth Concat structure. Then, through the eighth ZSNet structure, it outputs the third-dimensional feature map to the Head part.
[0020] Preferably, each ZSNet structure includes a CBS compression module, a channel splitting module, a first CT-Neck module, a second CT-Neck module, a feature fusion module, and a CBS restoration module; The input data of the ZSNet structure undergoes feature transformation and channel compression through the CBS compression module, and then is evenly divided into a first path of features and a second path of features according to the channel dimension by the channel splitting module. The second path of features undergoes enhancement processing through the first CT-Neck module to obtain a third path of features, and the third path of features undergoes enhancement processing through the second CT-Neck module to obtain a fourth path of features; The feature fusion module fuses the first path of features, the second path of features, the third path of features, and the fourth path of features to obtain a unified feature map; The channel restoration module adjusts the number of channels of the unified feature map to be consistent with the number of channels of the subsequent module connected to the ZSNet structure, and the output of the channel restoration module serves as the output of the ZSNet structure.
[0021] Preferably, the first CT-Neck module and the second CT-Neck module adopt the same structure. Each CT-Neck module includes a residual gating function module, a first ODConv module, a second ODConv module, and an addition module. The residual gating function module calculates a gating function value of 1 or 0 based on the input features of the CT-Neck module. The first ODConv module inputs the input features of the CT-Neck module, performs full-scale dynamic convolution, and then outputs to the second ODConv module. The second ODConv module performs full-scale dynamic convolution on the output of the first ODConv module and then outputs to the addition module. When the gating function value is 0, the addition module uses the output of the second ODConv module as the output of the CT-Neck module. When the gating function value is 1, the addition module element-wise adds the input features of the CT-Neck module and the output of the second ODConv module and uses the result as the output of the CT-Neck module.
[0022] Preferably, the processing process of the dysample structure is as follows:
[0023] Perform a linear transformation on the input feature X of the dysample structure to generate a learnable offset.
[0024] Add the regular grid in traditional upsampling to the offset to generate a dynamic sampling grid.
[0025] Use the dynamic sampling grid to perform interpolation sampling on the input feature to obtain the upsampled feature map, as follows:
[0026] ;
[0027] Among them, represents the upsampled feature map, that is, the output of the dysample structure; represents the dynamic sampling grid; represents interpolation sampling; represents the input feature of the dysample structure.
[0028] Preferably, the total loss function of the gastric cancer CT target detection model combines a bounding box regression loss function and a classification loss function, and is expressed as:
[0029] ;
[0030] Among them, represents the total loss function; represents the bounding box regression loss function; represents the classification loss function; and respectively represent the weights of the loss functions.
[0031] Preferably, the bounding box regression loss function includes the WIoU loss and the distribution focal loss; the WIoU loss is expressed as:
[0032] ;
[0033] ;
[0034] ;
[0035] where represents the WIoU loss; represents the exponential weighting factor; represents the CIoU loss; represents the intersection over union; and respectively represent the predicted box and the ground truth box; represents the center coordinate of the predicted box; represents the center coordinate of the ground truth box; and respectively represent the normalized width and height;
[0036] The distribution focal loss is expressed as:
[0037] ;
[0038] where represents the distribution focal loss; and respectively represent the predicted values output by the network; , and respectively represent the ground truth values input to the network, represents the ground truth label; i and i + 1 represent two consecutive adjacent positions in the dataset.
[0039] Preferably, the classification loss function is expressed as:
[0040] ;
[0041] where represents the classification loss function; represents the predicted probability output by the model; represents the ground truth label; represents the balance factor, used to adjust the influence of positive and negative samples; represents the focal factor, used to control the model's attention to difficult samples.
[0042] On the other hand, the gastric cancer CT lesion area detection device based on ZSNet includes the following:
[0043] An image acquisition and preprocessing module, which is used to acquire the grayscale image of the CT lesion area of gastric cancer, preprocess the image, and label and produce a data set;
[0044] A model construction module, which is used to construct a gastric cancer CT target detection model based on the YOLOv8 network; the gastric cancer CT target detection model includes a Backbone part, a Neck part, and a Head part;
[0045] The Backbone part includes a first unit, a second unit, a third unit, and a fourth unit connected in sequence; the second unit and the fourth unit include one or more ZSNet structures based on full-dimensional dynamic convolution; the data set passes through the first unit, the second unit, the third unit, and the fourth unit in sequence for feature extraction, and the second unit and the fourth unit respectively output the extracted feature representations to the Neck part;
[0046] The Neck part performs multi-scale feature fusion on the feature representation input from the Backbone part based on the dysample structure and the ZSNet structure, and then outputs several-dimensional feature maps to the Head part;
[0047] The Head part obtains the dimensional feature maps of the Neck part, and outputs the detection image of the CT lesion area of gastric cancer after feature enhancement by the CBS structure and feature mapping by the Con2d structure;
[0048] A model training module, which is used to train the gastric cancer CT target detection model using the data set to obtain a trained model;
[0049] A gastric cancer CT lesion area detection module, which is used to detect the CT lesion area of gastric cancer using the trained model.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] (1) The present invention innovatively introduces the ZSNet structure, and constructs the ZSNet structure by using adaptive feature reallocation, residual connection, and deep feature refinement technologies. This module can extract more detailed and discriminative features at low computational cost. Even under noise interference and low contrast conditions, it can effectively capture the tiny information of the lesions, thus greatly improving the detection performance;
[0052] (2) The present invention innovatively introduces the CT-Neck module, and adopts the ODConv technology therein. The offset deformable convolution is used to dynamically adjust and deeply fuse the input features, realizing the dual retention of global semantic information and local details, thus significantly enhancing the model's recognition ability for complex lesion areas in gastric cancer CT images and effectively improving the comprehensive detection performance;
[0053] (3) The present invention replaces the traditional bounding box regression loss with the WIoU loss function. By exponentially weighting the center offset between the predicted box and the ground truth box, the constraint on the alignment of the target center is strengthened, the risks of missed detection and false detection are significantly reduced, and the detection accuracy is further improved, especially suitable for dealing with the complex situations in the tiny lesion areas of gastric cancer CT images. Description of the Drawings
[0054] The present invention will be further described in detail below with reference to the drawings;
[0055] Figure 1 It is a flowchart of the method for detecting gastric cancer CT lesion areas based on ZSNet according to an embodiment of the present invention;
[0056] Figure 2 It is a schematic flow diagram of the method for detecting gastric cancer CT lesion areas based on ZSNet according to an embodiment of the present invention;
[0057] Figure 3 It is a structural block diagram of the device for detecting gastric cancer CT lesion areas based on ZSNet according to an embodiment of the present invention. Detailed Embodiments
[0058] The present invention will be further described below through specific embodiments.
[0059] As Figure 1 shown, the method for detecting gastric cancer CT lesion areas based on ZSNet is as follows:
[0060] S1. Obtain the grayscale image of the gastric cancer CT lesion area, preprocess the image and label it to make a data set.
[0061] The preprocessing includes removing the annotation information in the original gastric cancer CT image to obtain the de-annotated image; subsequently, perform a contrast enhancement operation on the de-annotated image. The contrast enhancement can adopt logarithmic transformation, and the calculation formula is as follows: , where is the value of the pixel point of the original image, is the corresponding pixel value after the enhancement process, represents the logarithmic function, and the constant C is used to make the gray dynamic range of the transformed image meet the requirements. Through the above logarithmic transformation, the image contrast can be effectively improved while retaining the image details, providing a higher-quality image input for the subsequent detection and recognition of gastric cancer CT lesion areas.
[0062] S2. Construct a gastric cancer CT target detection model based on the YOLOv8 network.
[0063] As Figure 2As shown in the figure, the Backbone part of the gastric cancer CT target detection model adopts feature extraction modules arranged from top to bottom. The feature extraction module includes four alternating module units arranged from top to bottom. The first unit includes a CBS structure and a CBS structure. The second unit includes a ZSNet structure, a CBS structure, and a ZSNet structure. The third unit includes a CBS structure. The fourth unit includes a ZSNet structure, a CBS structure, a ZSNet structure, and an SPPF structure. The image features extracted by different units are obtained and input into the Neck part, and assist in constructing the loss function of the target detection model, thereby improving the detection ability of the gastric cancer CT lesion area of the overall model and making the model pay more attention to the gastric cancer CT lesion area during the learning process.
[0064] The Neck part adopts a mutually fused FPN network and PANET network, and uses Dysample and ZSNet to improve the FPN network and PANET network to obtain feature maps of multiple dimensions.
[0065] The Head part adopts multiple CBS structures and Conv2d convolutional layers to output images for predicting the gastric cancer CT lesion area from three different dimensions, and assist in constructing the bounding box regression loss function Bbox Loss and the classification loss function Cls Loss. This embodiment uses the combination of ZSNet and YOLOv8 object detection to optimize the problem of false detection and missed detection in gastric cancer CT images, and improve the recognition rate and accuracy of detection.
[0066] In this embodiment, the implementation of the Neck part and the Head part is as follows:
[0067] The output of the SPPF module of the fourth unit alternating module is used as the input of the improved FPN network. This input passes through a dysample structure and a Concat structure in sequence, and performs feature fusion with the output of the ZSNet(40x40x512) structure. Then, after being processed by the ZSNet (to obtain preliminary fusion features) and the dysample structure, it is dimensionally concatenated with the output of the second unit alternating module, and the concatenated result passes through the ZSNet structure in sequence to complete feature fusion. The fusion result of the fused feature map is input into the Head part, and after passing through the CBS structure and the Conv2d structure, it is used as the first-layer output feature map.
[0068] The above fused feature map passes through the CBS structure and then is dimensionally concatenated with the preliminary fusion features again. The concatenated result passes through the ZSNet structure to complete feature fusion to obtain the secondary fusion features. The secondary fusion features are input into the Head part, and after passing through the CBS structure and the Conv2d structure, it is used as the second-layer output feature map;
[0069] After passing through the CBS structure, the secondary fusion features are concatenated in dimension after matching with the feature map output by the SPPF module of the fourth unit's alternating module. The concatenation result then undergoes feature fusion through the ZSNet structure and is input into the Head part. After passing through the CBS structure and the Conv2d structure, it serves as the output feature map of the third layer.
[0070] The implementation of the ZSNet structure in this embodiment is as follows:
[0071] The ZSNet structure receives the input feature map from the previous module. Among them, C represents the number of channels, H represents the height, and W represents the width.
[0072] The channel compression module CBS is adopted, and convolution, batch normalization (BN, Batch Normalization), and activation function are used to perform preliminary transformation and channel compression on the input features. The calculation formula is as follows: , where represents the input using a convolution operation with a kernel size of to compress the number of channels from to , represents batch normalization, represents the activation function, represents the output feature map, with a size of .
[0073] The channel splitting operation Spilt is used to evenly divide the feature map into two parts along the channel dimension and . Among them, and respectively represent the first and second paths of features after splitting, and the size of each path is , providing different sub-features for the subsequent CT-Neck branch.
[0074] The CT-Neck branch module is adopted to further enhance the second path of features after splitting. The CT-Neck branch module uses a gating mechanism to distinguish the processing paths. The process is as follows:
[0075] First, execute the first CT-Neck branch module to perform gating judgment to determine whether to adopt deep processing. Denote the gating function as , if the judgment is True, then ; if it is False, then . When the condition is True, the CT-Neck module enables residual connection, and the input feature is executed twice in sequence After convolution (Omni-Dimensional Dynamic Convolution) and element-wise addition to obtain , the calculation formula is as follows: , when the condition is False, where the CT-Neck module directly performs two convolution operations to obtain features , the calculation formula is as follows: , where represents an Omni-Dimensional Dynamic Convolution with a kernel size of 3×3, which can better adapt to local deformations and capture detailed information of gastric cancer CT.
[0076] Subsequently, the second CT-Neck branch module is executed, and the features output by the first CT-Neck branch module are subjected to the same operations as the first CT-Neck branch module to obtain .
[0077] Finally, after the Spilt operation and CT-Neck module feature extraction, a total of four-way features are output, namely: , , , , and the size of each way is .
[0078] The feature fusion module Concat is used to fuse the four-way enhanced features to form a unified feature map. The calculation formula is as follows: , where represents concatenating the four-way features in the channel dimension, represents the fused feature map, and its size is .
[0079] The channel restoration module CBS is used to again perform convolution, batch normalization (BN), and activation function processing on the fused feature map to restore or adjust the number of channels to be consistent with the requirements of the subsequent module. The calculation formula is as follows: , where represents performing a convolution operation with a kernel size of on the input to restore the number of channels to , represents batch normalization, represents the activation function, represents the output feature map, and its size is .
[0080] In this embodiment, the loss function of the gastric cancer CT target detection model is specifically as follows:
[0081] Construct a bounding box regression loss function (Bbox Loss) to measure the overlap degree between the predicted box and the ground truth box, and prompt the model to accurately regress the target position and size. The bounding box regression loss function includes CIoU loss and DFL loss.
[0082] The calculation formula of CIoU loss is as follows: ;
[0083] Among them, represents the predicted box, represents the ground truth box, represents the intersection over union of the two.
[0084] Use distribution focal loss to measure and optimize the labels, so that the network distribution focuses near the label values. The calculation formula is as follows: ;
[0085] Among them, , is the predicted value output by the network. , , is the ground truth value input to the network.
[0086] Construct a classification loss function (Cls Loss) to measure the difference between the predicted class probability and the true class label, and ensure that the model correctly identifies the target class. The calculation formula of the classification loss function is as follows:
[0087] ;
[0088] Among them, represents the predicted probability output by the model; is the true label; is the balance factor, used to adjust the influence of positive and negative samples, is the focal factor, used to control the model's attention to difficult samples.
[0089] Construct a preliminary total loss function. The calculation formula is as follows:
[0090] ;
[0091] Among them, , respectively represent the weights of the two loss functions.
[0092] In this embodiment, WIoU loss is introduced to replace , by introducing a central offset weighting mechanism, the model's attention to small targets and central alignment problems is enhanced, and the calculation formula is as follows:
[0093] ;
[0094] ;
[0095] ;
[0096] Among them, respectively represent the predicted bounding box and the ground truth bounding box, represents the intersection over union (IoU), and represent the center coordinates of the predicted bounding box and the ground truth bounding box, represents the normalized width and height, represents the exponential weighting factor.
[0097] Finally, the final total loss function is constructed, and the calculation formula is as follows:
[0098] ;
[0099] Among them, and respectively represent the weights of the two loss functions.
[0100] S3. Use the dataset to train the gastric cancer CT target detection model to obtain a trained model.
[0101] S4. Use the trained model to detect the lesion area of gastric cancer CT.
[0102] In this embodiment, an efficient and accurate automatic detection method is designed for the special detection requirements of gastric cancer CT images. Aiming at the characteristics of small volume, low contrast, more noise and complex background in the lesion area of gastric cancer CT images, an improved scheme is proposed, which effectively improves the detection rate and localization accuracy of the lesion area, and provides strong support for clinical early diagnosis and treatment. Specifically, in this embodiment, through multiple improvements based on the YOLOv8 model in the gastric cancer CT image detection scenario, the innovative ZSNet structure and CT-Neck module (combined with ODConv technology) are adopted, and the WIoU loss function is introduced, which significantly improves the detection accuracy and robustness of small lesions. It not only has theoretical innovation, but also shows extremely high clinical value and promotion prospects in practical applications.
[0103] This embodiment makes targeted improvements based on the YOLOv8 model. By optimizing the network structure and multi-scale feature extraction strategy, the model can better fuse high-level semantic information and low-level detail information, thereby significantly improving the detection ability of tiny lesions in complex backgrounds. When processing gastric cancer CT images, the model can not only improve the detection accuracy but also enhance the robustness to image noise and low contrast problems.
[0104] As Figure 3 shown, the present invention also discloses a device for detecting gastric cancer CT lesion regions based on ZSNet, including:
[0105] An image acquisition and preprocessing module 301, which is used to acquire the grayscale image of the gastric cancer CT lesion region, preprocess the image and label it to make a data set.
[0106] A model construction module 302, which is used to construct a gastric cancer CT target detection model based on the YOLOv8 network; the gastric cancer CT target detection model includes a Backbone part, a Neck part, and a Head part.
[0107] The Backbone part includes a first unit, a second unit, a third unit, and a fourth unit connected in sequence; the second unit and the fourth unit include one or more ZSNet structures based on full-dimensional dynamic convolution; the data set passes through the first unit, the second unit, the third unit, and the fourth unit in sequence for feature extraction, and the second unit and the fourth unit respectively output the extracted feature representations to the Neck part;
[0108] The Neck part performs multi-scale feature fusion on the feature representation input from the Backbone part based on the dysample structure and the ZSNet structure, and then outputs several-dimensional feature maps to the Head part;
[0109] The Head part obtains the dimensional feature maps of the Neck part, and outputs the gastric cancer CT lesion region detection image after feature enhancement by the CBS structure and feature mapping by the Con2d structure.
[0110] A model training module 303, which is used to train the gastric cancer CT target detection model using the data set to obtain a trained model.
[0111] A gastric cancer CT lesion region detection module 304, which is used to detect the gastric cancer CT lesion region using the trained model.
[0112] The specific implementation of the device for detecting gastric cancer CT lesion regions based on ZSNet is the same as that of the method for detecting gastric cancer CT lesion regions based on ZSNet, and will not be repeated in this embodiment.
[0113] The above are only the specific embodiments of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantive modification of the present invention using this concept shall fall within the scope of infringement of the protection scope of the present invention.
Claims
1. A gastric cancer CT lesion area detection method based on ZSNet, characterized in that: The steps include: S1, obtaining grayscale images of gastric cancer CT lesion areas, preprocessing and annotating the images to make a data set; S2, constructing a gastric cancer CT target detection model based on the YOLOv8 network; the gastric cancer CT target detection model includes a Backbone part, a Neck part and a Head part; The Backbone part includes a first unit, a second unit, a third unit and a fourth unit connected in sequence; the second unit and the fourth unit include one or more ZSNet structures based on full-dimensional dynamic convolution; the data set is sequentially subjected to feature extraction by the first unit, the second unit, the third unit and the fourth unit, and the extracted feature representations are output by the second unit and the fourth unit respectively to the Neck part; The Neck part performs multi-scale feature fusion on the feature representation of the input Backbone part based on the dysample structure and the ZSNet structure, and then outputs several dimensional feature maps to the Head part; The Head part obtains the dimensional feature map of the Neck part, and outputs the gastric cancer CT lesion area detection image after feature enhancement of the CBS structure and feature mapping of the Con2d structure; S3, using the data set to train the gastric cancer CT target detection model to obtain a trained model; S4, use the trained model to detect gastric cancer CT lesion area.
2. The gastric cancer CT lesion area detection method based on ZSNet according to claim 1 is characterized in that: In the Backbon part, the first unit includes a first CBS structure and a second CBS structure connected in sequence, the second unit includes a first ZSNet structure, a third CBS structure and a second ZSNet structure connected in sequence, the third unit includes a fourth CBS structure, and the fourth unit includes a third ZSNet structure, a fifth CBS structure, a fourth ZSNet structure and an SPPF structure connected in sequence.
3. The gastric cancer CT lesion area detection method based on ZSNet according to claim 2 is characterized in that: The Neck part includes a dysample structure, a second dysample structure, a first Concat structure, a second Concat structure, a third Concat structure, a fourth Concat structure, a fifth ZSNet structure, a sixth ZSNet structure, a seventh ZSNet structure, an eighth ZSNet structure, a sixth CBS structure and a seventh CBS structure; the specific implementation of the Neck part is as follows: The feature representation output by the SPPF structure is sampled by the first dysample structure, and after the feature representation output by the first Concat structure and the third ZSNet structure is feature fused, it passes through the fifth ZSNet structure and the second dysample structure in turn, and is dimensionally spliced with the feature representation output by the second unit in the second Concate structure. The spliced result passes through the sixth ZSNet structure and then outputs the first dimensional feature map to the Head part; The output of the sixth ZSNet structure passes through the sixth CBS structure, and is dimensionally concatenated with the output of the fifth ZSNet structure in the third Concat structure. After passing through the seventh ZSNet structure, the second-dimensional feature map is output to the Head part. The output of the seventh ZSNet structure passes through the seventh CBS structure, is concatenated with the feature representation output by the SPPF structure in the fourth Concat structure, and then passes through the eighth ZSNet structure to output the third dimensional feature map to the Head part.
4. The gastric cancer CT lesion area detection method based on ZSNet according to claim 1, characterized in that: Each of the ZSNet structures includes a CBS compression module, a channel splitting module, a first CT-Neck module, a second CT-Neck module, a feature fusion module and a CBS recovery module; the input data of the ZSNet structure is subjected to feature transformation and channel compression by the CBS compression module, and then is evenly divided into first-path features and second-path features according to the channel dimension by the channel splitting module, the second-path features are enhanced by the first CT-Neck module to obtain third-path features, and the third-path features are enhanced by the second CT-Neck module to obtain fourth-path features; feature The fusion module fuses the first feature, the second feature, the third feature and the fourth feature to obtain a unified feature map; the channel recovery module adjusts the number of channels of the unified feature map to make it consistent with the number of channels of the subsequent modules connected to the ZSNet structure, and the output of the channel recovery module is used as the output of the ZSNet structure.
5. The method for detecting gastric cancer CT lesion area based on ZSNet according to claim 4, characterized in that: The first CT-Neck module and the second CT-Neck module adopt the same structure, and each CT-Neck module includes a residual gating function module, a first ODConv module, a second ODConv module and an addition module; the residual gating function module calculates a gating function value of 1 or 0 based on the input features of the CT-Neck module; the first ODConv module inputs the input features of the CT-Neck module, performs full-scale dynamic convolution and outputs it to the second ODConv module; the second ODConv module performs full-scale dynamic convolution on the output of the first ODConv module and outputs it to the addition module; when the gating function value of the addition module is 0, the output of the second ODConv module is used as the output of the CT-Neck module; when the gating function value of the addition module is 1, the input features of the CT-Neck module and the output of the second ODConv module are element-wise added, and used as the output of the CT-Neck module.
6. The gastric cancer CT lesion area detection method based on ZSNet according to claim 1, characterized in that: The processing of the dysample structure is as follows: Perform a linear transformation on the input feature X of the dysample structure to generate a learnable offset; The regular grid used in traditional upsampling is added with the offset to generate a dynamic sampling grid; The input features are interpolated and sampled using a dynamic sampling grid to obtain the upsampled feature map, as follows: ; in, Represents the upsampled feature map, that is, the output of the dysample structure; represents a dynamic sampling grid; Represents interpolation sampling; Represents the input features of the dysample structure.
7. The gastric cancer CT lesion area detection method based on ZSNet according to claim 1, characterized in that: The total loss function of the gastric cancer CT target detection model combines the bounding box regression loss function and the classification loss function and is expressed as: ; in, represents the total loss function; represents the bounding box regression loss function; represents the classification loss function; and They represent the weights of the loss function respectively.
8. The gastric cancer CT lesion area detection method based on ZSNet according to claim 7, characterized in that: The bounding box regression loss function includes WIoU loss and distribution focus loss; the WIoU loss is expressed as: ; ; ; in, represents WIoU loss; represents the exponential weighting factor; represents CIoU loss; represents intersection and union ratio; and Represent the predicted box and the true box respectively; Represents the center coordinates of the prediction box; Represents the center coordinates of the real box; and Represent the normalized width and height respectively; The distribution focal loss is expressed as: ; in, represents the distribution focal loss; and Respectively represent the predicted values of the network output; , and They represent the true value of the input network, represents the true label; i and i+1 represent two consecutive adjacent positions in the dataset.
9. The gastric cancer CT lesion area detection method based on ZSNet according to claim 7, characterized in that: The classification loss function is expressed as: ; in, represents the classification loss function; Represents the predicted probability of the model output; represents the true label; Represents the balance factor, which is used to adjust the impact of positive and negative samples; Represents the focus factor, which is used to control the model's focus on difficult samples.
10. A gastric cancer CT lesion area detection device based on ZSNet, comprising the following: The image acquisition and preprocessing module is used to acquire the grayscale image of the gastric cancer CT lesion area, preprocess the image and annotate it to make a data set; A model building module, used to build a gastric cancer CT target detection model based on the YOLOv8 network; the gastric cancer CT target detection model includes a Backbone part, a Neck part and a Head part; The Backbone part includes a first unit, a second unit, a third unit and a fourth unit connected in sequence; the second unit and the fourth unit include one or more ZSNet structures based on full-dimensional dynamic convolution; the data set is sequentially subjected to feature extraction by the first unit, the second unit, the third unit and the fourth unit, and the extracted feature representations are output by the second unit and the fourth unit respectively to the Neck part; The Neck part performs multi-scale feature fusion on the feature representation of the input Backbone part based on the dysample structure and the ZSNet structure, and then outputs several dimensional feature maps to the Head part; The Head part obtains the dimensional feature map of the Neck part, and outputs the gastric cancer CT lesion area detection image after feature enhancement of the CBS structure and feature mapping of the Con2d structure; A model training module is used to train a gastric cancer CT target detection model using a data set to obtain a trained model; The gastric cancer CT lesion area detection module is used to detect gastric cancer CT lesion areas using the trained model.
Citation Information
Patent Citations
Rapid behavior detection method based on long-time enhanced feature enhancement and sparse dynamic sampling
CN110688918A
Multi-task histopathological image lesion segmentation method
CN116128832A
System, method, and computer-accessible medium for virtual pancreatography
US20200226748A1
System and method for segmenting medical images
US20240386551A1
A computer-implemented method of enhancing object detection in a digital image of known underlying structure, and corresponding module, data processing apparatus and computer program
US20240404235A1