Wireless capsule endoscope multi-category lesion image detection method based on deep learning

Through the wireless capsule lens multi-category lesion detection method based on deep learning, the Extended-ELAN network and Swin Transformer module optimization model is used to solve the problem of multi-category lesion detection in the digestive tract, efficient and accurate lesion recognition is achieved, and the intelligence level of medical diagnosis is improved.

CN120259797AInactive Publication Date: 2025-07-04CHANGCHUN UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510748023.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing wireless capsule endoscopy technology has difficulty identifying lesions, missed diagnosis and misdiagnosis rates in digestive tract examinations, especially when multiple lesions exist at the same time, and consumes a lot of human resources.

Method used

Using a deep learning-based method, the capsule lens object detection model is designed, and the Extended-ELAN network and the AFPN structure of the asymptotic feature pyramid network is used, and the multi-category lesion detection model is constructed by combining the Swin Transformer module. The model is optimized by the annotation, training and verification sets to achieve accurate detection of multi-category lesions.

Benefits of technology

It improves the detection accuracy and efficiency of multiple categories of the digestive tract lesions, reduces missed diagnosis and misdiagnosis, and improves the intelligence level of medical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259797A_ABST
    Figure CN120259797A_ABST
Patent Text Reader

Abstract

The invention discloses a wireless capsule endoscope multi-class focus image detection method based on deep learning, and belongs to the technical field of WCE multi-class focus image detection of computer vision in deep learning. Through the design scheme, the wireless capsule endoscope multi-category lesion image detection method based on the deep learning model provided by the invention can accurately identify and detect lesion information and position information in the input image, meets the multi-scale detection requirements of multi-category lesion images, and improves the detection efficiency. The innovative technology not only improves the accuracy and efficiency of medical diagnosis, but also injects new vitality to the field of intelligent medical services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of WCE multi-class lesion image detection in computer vision of deep learning, and in particular relates to an object detection method designed for WCE multi-class lesion image detection. Background Art

[0002] Wireless capsule endoscopy (WCE) technology is an important medical technology that can be used for the examination of the digestive tract and other organs. This has enabled the widespread application of wireless capsule endoscopy technology in clinical examinations of the digestive tract since its launch in 2001.

[0003] However, during the use of wireless capsule endoscopy, the built-in miniature camera needs to continuously capture images in the patient's body for 6 to 8 hours. The captured video can be decomposed into a large number of images, approximately 50,000 to 80,000. Moreover, the anatomical structure of the human digestive tract is complex, and lesions in different parts may present different shapes and characteristics.

[0004] Even for trained endoscopists, these images require at least 2 to 3 hours of focused time to complete the meticulous detection work. Identifying diseases not only consumes a large amount of human resources, but also because the lesion images only account for a very small proportion of all WCE images and are extremely easy to be overlooked. In addition, some diseases are difficult to be accurately identified due to their small size or being blocked, which leads to the frequent occurrence of missed diagnoses and misdiagnoses, posing great challenges to medical work.

[0005] Although a large amount of work has been done in the current research on the pathological abnormality detection of WCE images and videos, most of the research focuses on the detection of single or a few types of lesions. However, in real medical practice, there may be multiple different types of lesions in the digestive tract simultaneously, which requires the detection algorithm to be able to identify and process multiple types of lesions simultaneously, further increasing the difficulty of the corresponding work. Therefore, there is an urgent need for a technology in the prior art to solve the above problems. Summary of the Invention

[0006] The technical problem to be solved by the present invention is: to provide a method for solving the problems existing in the background art.

[0007] A method for detecting multi-class lesion images of a wireless capsule endoscope based on deep learning, characterized by comprising the following steps, and the following steps are carried out sequentially; S1: Obtain a multi-class lesion image dataset of a wireless capsule endoscope, annotate the dataset, and divide the dataset into a training set, a validation set, and a test set; S2: Design a capsule endoscope target detection model, using the Extended-ELAN network structure as the feature extraction network and the Asymptotic Feature Pyramid Network (AFPN) as the neck network structure. On this basis, re-optimize the design of the neck network structure, introduce the Swin Transformer module into the neck network structure to enhance the feature extraction and fusion capabilities. In the Head structure, design it as five detection heads to construct a capsule endoscope detection model.

[0008] S3: Bring the training set and validation set obtained in step S1 into the capsule endoscope detection model in step S2 for training and validation, and then use the capsule endoscope detection model with the best performance as the WCE multi-class lesion image detection model. S4: Use the test set obtained in step S1 to test the WCE multi-class lesion image detection model obtained in S3 to obtain a capsule endoscope detection model. S5: Input 28 types of digestive tract WCE lesion images to be detected into the capsule endoscope detection model obtained in step S4. Output the lesion detection result images, including the marked lesion detection boxes, the names of the lesions, and the accuracy rates.

[0009] The above S1 includes the following steps: S1-1: Obtain WCE image data, annotate the WCE image data using the labelImg software, and convert the annotation file into a TXT format file. S1-2: Divide the annotated dataset into a training set, a validation set, and a test set according to the ratio of 60%, 20%, and 20%.

[0010] In the above step S2, the neck network structure uses the Asymptotic Feature Pyramid Network (AFPN). In feature fusion, when the AFPN structure of the Asymptotic Feature Pyramid Network extracts features from bottom to top in the backbone network, it fuses features of different resolutions in stages, gradually integrates low-level feature information into high-level features, realizes the dual fusion of semantic and detailed information, and avoids the information discontinuity between non-adjacent feature layers. At the same time, draw on the adaptive spatial feature fusion technology, dynamically screen information according to feature importance, etc., and optimize the fusion quality. In practical applications, compared with the traditional feature pyramid network, it can improve the detection accuracy and show good performance advantages.

[0011] In the above step S2, the AFPN structure of the Asymptotic Feature Pyramid Network expands the multi-class object detection layer, strengthens the multi-class object feature expression, and uses ASFF to assign different weights to the fusion of different layers when fusing different layers.

[0012] In step S2, a Swin Transformer module is introduced into the neck network structure. Specifically, the 14th, 15th, 16th, and 17th layers in the network structure are designed as Swin Transformer modules. By combining multi-scale feature learning and self-attention mechanism, Swin Transformer can effectively capture multi-scale information and dependencies in images while maintaining efficient computation.

[0013] In step S2, in the Head structure, five detection heads are designed, forming five detection heads of 320*320, 160*160, 80*80, 40*40, and 20*20.

[0014] In step S3, during model training, the input image size is set to 640*640*3, the batch size is set to 64, and the number of training iterations is 300. For the optimizer, the initial learning rate is set to Adjust the learning rate according to the training loss and gradient changes, and update the parameters according to the gradient changes at the same time; after completion, verify the model on the validation set, and use the best-performing capsule endoscope detection model as the WCE multi-class lesion image detection model, and save the model weights as a file.

[0015] In step S4, during model testing, input the test set into the capsule endoscope detection model, and at the same time load the weight file obtained after training in S3 for model testing, output the model test results, evaluate the model results, and obtain the evaluated capsule endoscope detection model.

[0016] Through the above design scheme, the present invention provides a method for detecting multi-class lesion images of wireless capsule endoscopes based on a deep learning model, which can accurately identify and detect the lesion information and location information in the input image, realize the detection of multi-class lesions, and meet the needs of different patients. This innovative technology not only improves the accuracy and efficiency of medical diagnosis, but also injects new vitality into the field of intelligent medical services. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart of the method for detecting multi-class lesion images of wireless capsule endoscopes based on deep learning of the present invention;

[0018] Figure 2 It is a flowchart of the method for designing a capsule endoscope detection model of the present invention; Figure 3 It is the Swin Transformer model structure adopted by the present invention; Figure 4 It is the AFPN network module structure adopted by the present invention; Figure 5This is the structure of the 5 detection heads of the present invention. Specific embodiments

[0019] The present application will be further described with reference to the accompanying drawings: A method for detecting multi-class lesion images of a wireless capsule endoscope based on deep learning, characterized by including the following steps, and the following steps are carried out sequentially; S1: Obtain a multi-class lesion image dataset of a wireless capsule endoscope, label the dataset, and divide the dataset into a training set, a validation set, and a test set; S2: Design a capsule endoscope target detection model, use the Extended-ELAN network structure as the feature extraction network, and use the Asymptotic Feature Pyramid Network AFPN as the neck network structure; on this basis, re-optimize and design the neck network structure, and introduce the Swin Transformer module in the neck network structure to enhance the feature extraction and fusion capabilities; in the Head structure, it is designed as five detection heads to construct a capsule endoscope detection model; S3: Bring the training set and validation set obtained in step S1 into the capsule endoscope detection model in step S2 for training and validation, and then use the best-performing capsule endoscope detection model as the WCE multi-class lesion image detection model; S4: Use the test set obtained in step S1 to test the WCE multi-class lesion image detection model obtained in S3 to obtain a capsule endoscope detection model; S5: Input 28 kinds of digestive tract WCE lesion images to be detected into the capsule endoscope detection model obtained in step S4; output a lesion detection result image, including labeled lesion detection frames, the names of the lesions, and the accuracy rates.

[0020] The said S1 includes the following steps: S1-1, obtain WCE image data, label the WCE image data using the labelImg software, and convert the annotation file into a TXT format file; S1-2, divide the labeled dataset into a training set, a validation set, and a test set according to the ratio of 60%, 20%, and 20%.

[0021] In step S2, the neck network structure adopts the Asymptotic Feature Pyramid Network (AFPN). In feature fusion, when the AFPN structure extracts features from bottom to top in the backbone network, it fuses features of different resolutions in stages, gradually integrates low-level feature information into high-level features, realizes the dual fusion of semantic and detailed information, and avoids the information discontinuity between non-adjacent feature layers. At the same time, referring to the adaptive spatial feature fusion technology, it dynamically filters information based on feature importance, etc., to optimize the fusion quality. In practical applications, compared with the traditional feature pyramid network, it can improve the detection accuracy and show good performance advantages.

[0022] In step S2, the AFPN structure expands the multi-class object detection layer, strengthens the feature expression of multi-class objects, and uses ASFF to assign different weights to the fusion of different layers when fusing different layers.

[0023] In step S2, the Swin Transformer module is introduced into the neck network structure. Specifically, the 14th, 15th, 16th, and 17th layers in the network structure are designed as Swin Transformer modules. By combining multi-scale feature learning and self-attention mechanism, Swin Transformer can effectively capture multi-scale information and dependencies in the image while maintaining efficient computing.

[0024] In step S2, in the Head structure, five detection heads are designed, forming five detection heads of 320*320, 160*160, 80*80, 40*40, and 20*20.

[0025] In step S3, during model training, the input image size is set to 640*640*3, the batch size is set to 64, and the number of training iterations is 300. For the optimizer, the initial learning rate is set to Adjust the learning rate according to the training loss and gradient changes, and update the parameters according to the gradient changes at the same time. After completion, verify the model on the validation set, and use the best-performing capsule endoscope detection model as the WCE multi-class lesion image detection model, and save the model weights as a file.

[0026] In step S4, during model testing, input the test set into the capsule endoscope detection model, and at the same time load the weight file obtained after training in S3 for model testing, output the model test results, evaluate the model results, and obtain the evaluated capsule endoscope detection model. Embodiment

[0027] Figure 1The flowchart of a method for detecting multi-class lesion images of wireless capsule endoscopes based on deep learning according to an embodiment of the present invention is provided. In this embodiment, we collect 28 types of WCE digestive tract lesion images for the training and testing of the capsule endoscope detection model. A method for detecting multi-class lesion images of wireless capsule endoscopes based on a deep learning model according to an embodiment of the present invention is elaborated in detail, as Figure 1 shown, the specific steps are as follows: S1: The lesion images captured by the wireless capsule endoscope using the PillCamSB series are used to construct a WCE multi-class lesion dataset, and the dataset is labeled and divided.

[0028] S1-1. Obtain WCE image data, and use the labelImg software to label the horizontal and vertical coordinates, length, and width of the lesions in each image, and convert the annotation file into a TXT format file.

[0029] S1-2. For the labeled dataset, we divide the dataset into a training set, a validation set, and a test set according to the ratio of 60%, 20%, and 20% for model training and evaluation.

[0030] S2: Use the Extended-ELAN network structure as the feature extraction network, and adopt the asymptotic feature pyramid network as the neck network structure; on this basis, re-optimize the design of the neck network structure, and introduce the Swin Transformer module in the neck network structure to enhance the feature extraction and fusion capabilities; in the Head structure, it is designed as five detection heads to construct a capsule endoscope detection model; as Figure 2 shown, the specific operations are as follows: The capsule endoscope detection model consists of three parts: the Backbone structure, the Neck structure, and the Head structure. The Backbone structure is used to extract the lesion features in the WCE image, the Neck structure is used for the fusion of lesion features, and the Head structure is used for the prediction of WCE lesion images. Among them, Extended-ELAN is used as the Backbone structure, the asymptotic feature pyramid network (AFPN) is used as the Neck structure, and the Head structure is composed of five detection heads.

[0031] In the Backbone structure, the Extended-ELAN network stacks ordinary convolutional layers and E-ELAN layers for feature extraction. The input image is a 640*640px capsule endoscope lesion image with 28 types of lesions, and the output is five feature maps , , , , Among them, the size and number of channels of each feature map can be expressed as: (1) (2) (3) (4) (5) Among them represents a three-dimensional tensor, each element of which is a real number. 320×320 represents the spatial dimension of the feature map, that is, both the height and width of the feature map are 320 pixels, represents the number of channels of the feature map, that is, each spatial position in the feature map has feature channels. Similarly , , , are also the same; in the Neck structure network part, an asymptotic feature pyramid network is used to enhance the fusion of multi-scale features, as shown in Figure 4 ; when the AFPN structure extracts features from bottom to top in the backbone network, it fuses features of different resolutions in stages, integrates low-level feature information into high-level features layer by layer, realizes the dual fusion of semantic and detail information, and avoids the information break between non-adjacent feature layers; at the same time, referring to the adaptive spatial feature fusion technology, it dynamically filters information according to feature importance, etc., and optimizes the fusion quality; in practical applications, compared with the traditional feature pyramid network, it can improve the detection accuracy and show good performance advantages; when , , , is used as the input, first a linear transformation is performed to convert the input feature map X into a query vector Q, a key vector K, and a value vector V that can be queried, and then feature fusion is carried out, (6) (7) (8) Among them, represents the dimension of the feature map X, are respectively the height, width, and number of channels of the feature map, , , are trainable weights. Then the calculation of the self-attention matrix can be expressed as: (9) Among them, represents the dot product between the query vector Q and the transpose of the key vector K, represents the scaling factor, is the dimension of the key vector. The softmax function is used to normalize the attention scores into a probability distribution such that the sum of all scores is 1. A represents the attention matrix, and then the feature map is output. (10) where AV is a matrix multiplication operation, which means multiplying the attention matrix A by the value vector V to obtain the output calculated by the self-attention mechanism. Feature map; fuse the output of the Swin Transformer with the original feature map by combining multi-scale feature learning and self-attention mechanism; the Swin Transformer structure has four stages to obtain feature maps, as Figure 3 shown; each stage contains two parts. Except for the first stage which is composed of a LinearEmbedding and a Swin Transformer Block, the remaining three are all composed of a Patch Merging and a Swin Transformer Block; among them, Patch Merging is similar to a pooling operation but will not cause information loss; after each stage of processing, the resolution becomes half of the original, while the number of channels becomes twice that of before; the Swin Transformer Block is the core of this algorithm and is composed of a window multi-head self-attention layer W-MSA and a shifted window multi-head self-attention layer (SW-MSA); therefore, the number of layers of the Swin Transformer should be an integer multiple of 2, one layer for W-MSA and one layer for SW-MSA; behind this module is a 2-layer multi-layer perceptron MLP with a non-linear rectified linear unit ReLU in between; before and after each MSA module and MLP layer, it is composed of a normalization layer LayerNorm and a residual connection; the window-based W-MSA module and the shifted window-based SW-MSA module are respectively applied to two consecutive transformer blocks; based on this window partitioning mechanism, the process of calculating the feature maps in consecutive Swin Transformer blocks is as follows: (11) (12) (13) (14) where represents the output of the encoder from W-MSA, capturing the relationships between different parts of the sequence. represents the output after the MLP module processes. The main function of this layer is to convert the input data into a more discriminative feature representation, thereby improving the performance of subsequent tasks.

[0032] After being processed by the Swin Transformer module and the SPPF module, a set of features of different scales is generated { In the AFPN structure, the feature fusion process adopts a progressive hierarchical processing strategy. First, the feature set { } features to perform channel dimension reduction operations to generate { } features, and then realize cross-layer feature interaction through the multi-level adaptive spatial feature fusion module (ASSF). First, the low-level features and Input into the ASSF2 module and compare the output with Input them into ASSF3 for secondary fusion. Similarly, after fusion, Input them into ASSF4 for three fusions, and finally combine the output of ASSF4 with Input to ASSF5 for global feature integration. After the fusion of each stage is completed, the feature dimension is restored to the original scale through the channel expansion operation, and finally a set of multi-scale features {P1, P2, P3, P4, P5} is generated and sent to the detection head for detection.

[0033] In order to treat lesions of different sizes, five output heads are designed in the Head part, such as Figure 5 As shown. Multi-scale detection first generates features by the backbone network, which is divided into five stages: C1, C2, C3, C4, and C5. The number indicates the number of times the resolution is halved. For example, C1 represents the feature map output by stage 1, and the resolution is 1 / 2 of the input image. In the backbone network, a new stage 5 is added, which represents the feature map output by stage 5, and the resolution is 1 / 16 of the input image. In the Neck structure, the AFPN structure is adopted, and the output after fusion is P1, P2, P3, P4, and P5, which generate five detection heads of size 320*320 (xxlarge), 160*160 (xlarge), 80*80 (large), 40*40 (medium), and 20*20 (small). Finally, the four detection heads detect the lesion image and generate the detection frame.

[0034] (15) (16) (17) (18) (19) in, It is a detection head (Head) in the network, used to process the final feature maps P1, P2, P3, P4, P5. Each output head uses a classifier to predict the class probabilities , uses a bounding box regressor to predict the bounding boxes , and uses a confidence predictor to predict the confidence .

[0035] In the above predictions, the loss function of the overall model is divided into three parts. First is the bounding box loss: , where is the true bounding box, is the bounding box predicted by the model, is the smooth L1 loss function, used to measure the difference between the predicted bounding box and the true bounding box, is the bounding box loss, used to predict the confidence of the bounding box. Second is the confidence loss: Here, binary cross-entropy is used. Among them, is the true confidence label, 1 indicates the presence of an object, 0 indicates the absence, is the confidence predicted by the model, log is the logarithmic function, used to calculate the binary cross-entropy loss, is the confidence loss, using the binary cross-entropy loss to measure the difference between the predicted confidence and the true confidence label. Finally is the classification loss: , where is the true class label, is the class probability predicted by the model, is the classification loss, and the classification loss uses the cross-entropy loss to measure the difference between the predicted class probability and the true class. The final loss can be expressed as: (20) where is for different sizes, the distribution is small, medium, large, extra-large, the bounding box loss, confidence loss, and classification loss of the lesions, is the weight coefficient, and L is the loss function of the overall model, which combines the bounding box loss, confidence loss, and classification loss of lesions of different sizes, used to optimize the overall performance of the model.

[0036] S3: Train the capsule endoscope detection model. Set the input image size to 640*640*3, the batch size to 16, the number of training iterations to 300. For the optimizer, set the initial learning rate to Adjust the learning rate according to the training loss and gradient change : (21) where β is the attenuation factor and t is the current iteration number. Before updating the parameters each time, a certain number of gradients are accumulated , and then the parameters are updated to reduce the video memory occupation: (22) where is the gradient of the -th mini-batch training, and k is the number of accumulated batches. After the training is completed, the model is verified on the validation set, and the best-performing capsule endoscope detection model is used as the final WCE multi-class lesion image detection model, and the model weights are saved as a file.

[0037] In summary, the purpose of the present invention is to provide a method for detecting multi-class lesion images of wireless capsule endoscopes based on deep learning, and use WCE image data to train the capsule endoscope detection model to solve the problem of detecting 28 types of WCE lesion images and promote the development of intelligent medicine.

[0038] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. In addition, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. A method for detecting multi-class lesion images of wireless capsule endoscopes based on deep learning, characterized in that, Including the following steps, and the following steps are carried out sequentially; S1: Obtain a multi-class lesion image dataset of wireless capsule endoscopy, annotate the dataset, and divide the dataset into a training set, a validation set, and a test set; S2: Design a capsule endoscope object detection model, use the Extended-ELAN network structure as the feature extraction network, and adopt the Asymptotic Feature Pyramid Network AFPN as the neck network structure; on this basis, re-optimize and design the neck network structure, and introduce the Swin Transformer module in the neck network structure to enhance the feature extraction and fusion capabilities; in the Head structure, it is designed as five detection heads to construct a capsule endoscope detection model; S3: Bring the training set and validation set obtained in step S1 into the capsule endoscope detection model in step S2 for training and validation, and then use the best-performing capsule endoscope detection model as the WCE multi-class lesion image detection model; S4: Use the test set obtained in step S1 to test the WCE multi-class lesion image detection model obtained in S3 to obtain a capsule endoscope detection model; S5: Input 28 kinds of digestive tract WCE lesion images to be detected into the capsule endoscope detection model obtained in step S4; output the lesion detection result image, including the marked lesion detection box, the name of the lesion, and the accuracy rate.

2. The method for detecting multi-class lesion images of wireless capsule endoscopes based on deep learning according to claim 1, characterized in that: The said S1 includes the following steps: S1-1, Obtain WCE image data, annotate the WCE image data using the labelImg software, and convert the annotation file into a TXT format file; S1-2, Divide the annotated dataset into a training set, a validation set, and a test set according to the ratio of 60%, 20%, and 20%.

3. The method for detecting multi-class lesion images of wireless capsule endoscopes based on deep learning according to claim 1, characterized in that: In the said step S2, use the Extended-ELAN network structure as the feature extraction network, and the neck network structure is designed as the Asymptotic Feature Pyramid Network AFPN; in feature fusion, when the Asymptotic Feature Pyramid Network AFPN structure extracts features from bottom to top in the backbone network, it fuses different resolution features in stages, gradually integrates low-level feature information into high-level features, realizes the dual fusion of semantic and detail information, and avoids the information discontinuity of non-adjacent feature layers; at the same time, draw on the adaptive spatial feature fusion technology, dynamically screen information according to feature importance, etc., and optimize the fusion quality; in practical applications, compared with the traditional feature pyramid network, it can improve the detection accuracy and show good performance advantages.

4. The method for detecting multi-class lesion images of wireless capsule endoscopes based on deep learning according to claim 1, wherein: In the said step S2, the Asymptotic Feature Pyramid Network AFPN structure expands the multi-class object detection layer, strengthens the multi-class object feature expression, and uses ASFF to assign different weights to the fusion of different layers when fusing different layers.

5. The method for detecting multi-class lesion images of a wireless capsule endoscope based on deep learning according to claim 1, characterized in that: In the said step S2, the Swin Transformer module is introduced into the neck network structure. Specifically, the 14th, 15th, 16th, and 17th layers in the network structure are designed as Swin Transformer modules. By combining multi-scale feature learning and self-attention mechanism, Swin Transformer can effectively capture multi-scale information and dependencies in the image while maintaining high computational efficiency.

6. The method for detecting multi-class lesion images of wireless capsule endoscopes based on deep learning according to claim 1, characterized in that: In the step S2, in the Head structure, it is designed with five detection heads, forming five detection heads of 320*320, 160*160, 80*80, 40*40 and 20*20.

7. The method for detecting multi-class lesion images of wireless capsule endoscopes based on deep learning according to claim 1, characterized in that: In the step S3, when the model is trained, the input image size is set to 640*640*3, the batch size is set to 64, the number of training iterations is 300. For the optimizer, the initial learning rate is set to Adjust the learning rate according to the training loss and gradient changes, and update the parameters according to the gradient changes at the same time; after completion, verify the model on the validation set, use the best-performing capsule endoscope detection model as the WCE multi-class lesion image detection model, and save the model weights as a file.

8. The method for detecting multi-class lesion images of wireless capsule endoscopes based on deep learning according to claim 7, wherein: In the step S4, during model testing, the test set is input into the capsule endoscope detection model, and at the same time, the weight file obtained after training in S3 is loaded for model testing. The model test results are output, and the model results are evaluated to obtain the evaluated capsule endoscope detection model.

Citation Information

Patent Citations

  • Multi-category lesion high-speed detection method and system based on a capsule endoscope image

    CN113222957A

  • Machine room line safety detection method and system based on deep learning

    CN117710795A

  • Mask detection algorithm based on YOLOv7

    CN117726861A

  • Smear detection method and device, electronic equipment, storage medium and program product

    CN118118651A

  • Capsule endoscope image auxiliary diagnosis method based on deep learning

    CN119399679A