Endoscopic Mucosal Dissection Resection Area Identification Method, System, Device and Medium

By constructing a resection area recognition model based on deep learning based TransUNet++ structure, the problem of high difficulty in identifying resectionable areas of the submucosal layer in endoscopic mucosal stripping is solved, and the accurate identification and segmentation of resectionable areas of the submucosal layer is achieved, reducing the risk of surgery.

CN118230120BActive Publication Date: 2025-06-17BEIJING CHAOYANG HOSPITAL CAPITAL MEDICAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410291139.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-14
Publication Date
2025-06-17
Estimated Expiration
2044-03-14

AI Technical Summary

Technical Problem

In endoscopic mucosal dissection, it is difficult to identify the resectable area of ​​the submucosal layer, resulting in an increased risk of perforation and bleeding, and it is difficult for the prior art to assist doctors in accurately identifying resectable areas.

Method used

Using deep learning-based image recognition method, a resection area recognition model with TransUNet++ structure is constructed. By collecting and annotating endoscopic image data, the model is trained to achieve accurate identification and segmentation of resectionable areas in the submucosal layer.

Benefits of technology

Accurate identification and segmentation of the resectable areas of the submucosal layer under digestive endoscopic mucosal dissection has been achieved, reducing the risks of perforation and bleeding, and providing a basis for intelligent identification of resectable areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118230120B_ABST
    Figure CN118230120B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of medical image processing, and discloses a method, system, device and medium for identifying the resection area in endoscopic submucosal dissection, including: collecting the marking data of the resected area in the submucosa of endoscopic submucosal dissection; using the marking data of the resected area in the submucosa of endoscopic submucosal dissection, and adopting a framework based on the TransUNet++ structure to train and obtain a resection area recognition model. The method and system for identifying the resection area in endoscopic submucosal dissection provided by the present invention, by establishing a framework based on the TransUNet++ structure and using the collected marking data of the resected area in the submucosa of endoscopic submucosal dissection to train it, obtain a recognition model for the resected area, realize the accurate recognition and segmentation of the resected area in the submucosa of endoscopic submucosal dissection, and provide a basis for intelligent identification of the resected area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing, and particularly to a method and system for identifying the resectable area of the submucosa in endoscopic submucosal dissection based on deep learning. Background Art

[0002] Endoscopic Submucosal Dissection (ESD) can be used for the resection of early cancers and other large - area lesions in the digestive tract. It is a treatment method that has emerged in recent years and has been quickly recognized and widely applied. Compared with surgical operations, it has the advantages of less trauma, faster recovery, repeatable treatment, maximum retention of the physiological functions of the digestive tract, and fewer sequelae.

[0003] However, the implementation of ESD is highly difficult and requires endoscopists to have very high skills. Especially when dissecting the submucosa, it is necessary to carefully identify the resectable area. Otherwise, it may cause perforation and bleeding. Moreover, problems such as the complex internal environment of the digestive tract, limited vision during ESD surgery, and contamination of the lens by adipose tissue further increase the difficulty of identifying the resectable area during dissection, thus increasing the risk of perforation. Even for some lesions with submucosal fibrosis, only experienced ESD operators can accurately identify the resectable area.

[0004] For the above reasons, there is an urgent need for a means that can accurately identify the resectable area and can be widely used to assist endoscopists in completing ESD surgery and reducing the risks of perforation and bleeding. Moreover, currently, the technology of robotic endoscopic surgery is developing rapidly. If the resectable area can be accurately identified, it will also be possible to perform ESD surgery with a robotic endoscopic surgery system. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the related technologies to some extent.

[0006] One object of the present invention is to provide a method for constructing a recognition model for the resection area in endoscopic submucosal dissection. Based on the image recognition method of deep learning, a recognition model for the resectable area is constructed, providing a basis for intelligent recognition of the resectable area.

[0007] Another object of the present invention is to provide a system for constructing a recognition model for the resection area in endoscopic submucosal dissection.

[0008] Still another object of the present invention is to provide a method for identifying the resection area in endoscopic submucosal dissection, which accurately identifies and segments the resectable area of the submucosa in endoscopic submucosal dissection.

[0009] Another object of the present invention is to provide a system for identifying the resection area in endoscopic submucosal dissection.

[0010] To achieve the above object, on the one hand, the present invention provides a method for constructing a resection area recognition model for endoscopic submucosal dissection, including:

[0011] Collecting the marking data of the resected area of the submucosa in endoscopic submucosal dissection;

[0012] Using the marking data of the resected area of the submucosa in endoscopic submucosal dissection, and adopting a framework based on the TransUNet++ structure to train and obtain a resection area recognition model.

[0013] Further, according to the method for constructing a resection area recognition model of the present invention, the step of using the marking data of the resected area of the submucosa in endoscopic submucosal dissection and adopting a framework based on the TransUNet++ structure to train and obtain a resection area recognition model includes:

[0014] Based on the TransUNet++ structure, establishing a preliminary model;

[0015] According to the marking data of the resected area of the submucosa in endoscopic submucosal dissection, training the preliminary model to obtain a mature resection area recognition model.

[0016] Further, according to the method for constructing a resection area recognition model of the present invention, the step of establishing a preliminary model based on the TransUNet++ structure specifically includes:

[0017] Using a deep residual network as an encoder to perform feature extraction and obtain features of different resolutions;

[0018] Using an MSFM module and a convolutional network as a decoder;

[0019] The encoder and the decoder are densely connected through skip links, and the nodes of the skip links are MSFM modules; MSFM is a multi-scale feature extraction and fusion module;

[0020] The decoder extracts and fuses features of different resolutions through the MSFM module;

[0021] The last stage of the encoder and the first stage of the decoder are linked through a Transformer structure based on a multi-head self-attention layer.

[0022] Further, according to the method for constructing a resection area recognition model of the present invention, the step of collecting the marking data of the resected area of the submucosa in endoscopic submucosal dissection includes:

[0023] Collect endoscopic image data of cases where endoscopic submucosal dissection has been performed in the past;

[0024] The endoscopic image data is handed over to expert doctors to mark the resected area of the submucosa, generating marking data;

[0025] The marking data is divided into a training set, a validation set, and a test set, which are used for training, validating, and testing the resection area recognition model respectively;

[0026] Among them, data augmentation is performed on the training set data to obtain more data.

[0027] Furthermore, according to the method for constructing the resection area recognition model of the present invention, the preliminary model is trained based on the marking data of the resected area of the submucosa in endoscopic submucosal dissection to obtain a mature resection area recognition model, specifically:

[0028] The model is trained using the SGD optimizer. During training, the learning rate is set to 0.005, the momentum is set to 0.9, the weight decay is set to 0.0001, the network batch size is set to 64, and the learning rate is updated automatically according to the following formula:

[0029]

[0030] Where lr0 = 0.005, iter_num is the current iteration step, max_iterations = sample size * Epoch, and Epoch = 200, where Epoch is the total training cycle.

[0031] Furthermore, according to the method for constructing the resection area recognition model of the present invention, binary cross-entropy and Dice loss are used to measure the training effect of the model; the specific expression is:

[0032]

[0033]

[0034] L decoder = λ1L bce + λ2L dice ;

[0035] Where λ1 = λ2 = 0.5; L bce represents cross-entropy; L dice represents the dice coefficient; L decoder represents the overall loss.

[0036] On the other hand, the present invention provides a system for constructing a resection area recognition model for endoscopic submucosal dissection, including:

[0037] A data collection device for collecting marking data of the resectable area of the submucosa in endoscopic submucosal dissection.

[0038] A model training device for using the marking data of the resectable area of the submucosa in endoscopic submucosal dissection and training an excision area recognition model by adopting a framework based on the TransUNet++ structure.

[0039] Preferably, for the construction system of the excision area recognition model according to the present invention, the excision area recognition model includes:

[0040] An encoder module composed of a deep residual network, with a total of four stages, which are sequentially defined as x0, x1, x2, and x3, where x0 is the first stage, the spatial dimension of each stage is 1 / 2 of the previous stage, and inside each stage, after the deep residual network extracts features, the MSFM module is used to further perform multi-scale feature extraction on these features.

[0041] A decoder module, with a total of four stages, which are sequentially defined as d3, d2, d1, and d0. d3 is used as the first stage and corresponds to the deepest stage x3 of the encoder module; each stage includes an MSFM module and a convolution module. The input features are first subjected to multi-scale feature extraction and fusion by the MSFM, and then further feature fusion is performed by the convolution module. Except for the last stage, the outputs of other stages are subjected to upsampling processing with the size doubled.

[0042] A Transformer module for linking the deepest stage x3 of the encoder module and the first stage d3 of the decoder module; the output of the deepest stage x3 of the encoder module is flattened and transmitted to the Transformer module for learning global information and long-tail relationships.

[0043] The encoder module and the decoder module are combined in a TransUNet++ structure to form three-layer skip connections. The stages of the encoder module are used as the head nodes of the three-layer skip connections, and the stages of the decoder module are used as the tail nodes of the three-layer skip connections. The MSFM modules are respectively arranged at the intermediate nodes of the three-layer skip connections; wherein the MSFM module is used for multi-scale feature extraction and fusion.

[0044] Set the jth node in the ith layer of the three-layer skip connection as x_ij, then x_00 is x0, and x_04 is d0. The features of all intermediate nodes of the skip connection and all stages of the decoder module are obtained through the following formula:

[0045] x_01 = MSFM(cat(x0, UP(x1))),

[0046] x_11 = MSFM(cat(x1, UP(x2))),

[0047] x_02 = MSFM(cat(x0, x_01, UP(x_11))),

[0048] x_21 = MSFM(cat(x2, UP(x3))),

[0049] x_12 = MSFM(cat(x1, x_11, UP(x_21))),

[0050] x_03 = MSFM(cat(x0, x_01, x_02, UP(x_12))),

[0051] d3 = x_31 = MSFM(cat(x3, UP(Transformer))),

[0052] d2 = x_22 = MSFM(cat(x2, x_21, UP(x_31))),

[0053] d1 = x_13 = MSFM(cat(x1, x_11, x_12, UP(x_22))),

[0054] d0 = x_04 = MSFM(cat(x0, x_01, x_02, x_03, UP(x_13)));

[0055] Among them, cat is the feature concatenation operation, UP is the upsampling operation that doubles the feature map, and MSFM is the multi-scale feature fusion module.

[0056] Preferably, for the construction system of the resection area recognition model according to the present invention, the four-stage modules of the encoder module are extracted from the deep residual network ResNet50. Among them, each residual convolution block is formed by combining group normalization, rectified linear unit activation, convolution, and identity mapping. The identity mapping connects the input and output, and the output of each residual block is used to construct a nested TransUNet++ structure.

[0057] Preferably, for the construction system of the resection area recognition model according to the present invention, the output of the deepest stage x3 of the encoder module is transformed into a two-dimensional sequence, and this two-dimensional sequence is flattened into one-dimensional vectors of trainable feature vectors Q, V, and K. The Transformer layer is composed of L layers of multi-head self-attention (MSA) and multi-layer perceptron (MLP) blocks. MSA connects multiple SAs, as shown in the formula:

[0058]

[0059] MSA(Q, K, V) = Concate([H1, …, H h );

[0060] where h represents the total number of heads, and Q, K, and V are the query, key, and value respectively.

[0061] Preferably, for the construction system of the resection area recognition model according to the present invention, the MSFM module includes two parts: a convolutional block (CB) and Inception-Resnet. The CB captures information through a 3×3 convolutional layer, batch normalization (BN), and rectified linear unit (ReLU) activation. Inception-Resnet consists of an identity mapping and three convolutional branches b1, b2, and b3. A BN layer follows the convolutional layer of Inception-Resnet. The identity mapping is used to retain the original feature map information to prevent gradient explosion, and residual scaling is performed through a scaling factor. Finally, a ReLU layer is used for activation and output. The MSFM module can be expressed as follows:

[0062] x l = CB(x l );

[0063]

[0064] where, [] represents the concatenation operation, x l and x l+1 represent the input and output of the l-th layer respectively, represents the output of the three branches of the l-th layer of Inception-Resnet, and α represents the scale factor.

[0065] The present invention also provides a method for identifying the resection area of endoscopic submucosal dissection, including:

[0066] Obtaining the endoscopic image of the person to be subjected to endoscopic submucosal dissection;

[0067] Identifying and segmenting the resected area of the submucosa through the resection area recognition model obtained by the above construction method according to the endoscopic image.

[0068] The present invention also provides a system for identifying the resection area of endoscopic submucosal dissection, including:

[0069] An endoscopic imaging module for obtaining the endoscopic image of the person to be subjected to endoscopic submucosal dissection;

[0070] A resection area recognition module for identifying and segmenting the resected area of the submucosa through the resection area recognition model obtained by the above construction method according to the endoscopic image.

[0071] The present invention also provides an electronic device, which is characterized by comprising a processor, a communication interface, a memory and a communication bus;

[0072] Among them, the processor, the communication interface and the memory complete mutual communication through the communication bus;

[0073] The processor is used to call the logical instructions in the memory to execute the above endoscopic submucosal dissection resection area recognition method.

[0074] The present invention also provides a non-transitory computer-readable storage medium, on which computer instructions are stored, and the computer instructions enable the computer to execute the above endoscopic submucosal dissection resection area recognition method.

[0075] The present invention also provides a computer program product, the computer program product includes a computer program, the computer program can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer can execute the above endoscopic submucosal dissection resection area recognition method.

[0076] The endoscopic submucosal dissection resection area recognition method and system provided by the present invention establish a framework based on the TransUNet++ structure, and use the marked data of the resected area of the submucosa of the endoscopic submucosal dissection collected to train it, obtain a recognition model of the resected area, and realize the accurate recognition and segmentation of the resected area of the submucosa of the endoscopic submucosal dissection, providing a basis for intelligent recognition of the resected area. Description of the Drawings

[0077] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0078] Figure 1 It is a framework diagram of the resection area recognition model of the present invention;

[0079] Figure 2 It is a framework diagram of the MSFM module of the present invention;

[0080] Figure 3 It is a framework diagram of the Transformer module of the present invention;

[0081] Figure 4 It is the application of the present invention in recognizing the resected area of the submucosa in endoscopic submucosal dissection. Detailed Embodiments

[0082] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them, and they should not be construed as limitations on the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. In the description of the present invention, it should be understood that the terms used are only for the purpose of description and cannot be construed as indicating or implying relative importance.

[0083] The following describes in conjunction with the accompanying drawings a method and system for constructing an endoscopic submucosal dissection resection area recognition model, as well as a method and system for recognizing an endoscopic submucosal dissection resection area provided by the present invention.

[0084] In the present invention, the endoscopic submucosal dissection may be an endoscopic submucosal dissection for removing digestive tract mucosal lesions in a subject. The endoscope may be a rigid endoscope, a fiber optic endoscope, or an electronic endoscope, and may include a gastroscope, a colonoscope, a small intestine endoscope, an endoscopic ultrasound, etc. according to the location; or a capsule endoscope. The digestive tract mucosal lesion may be cancer, ulcer, inflammation, polyp, or other large-area lesions. The subject may be a mammal suffering from a digestive tract disease, and the mammal includes cattle, horses, dogs, non-human primates, and humans, or various experimental animals, pets, etc.

[0085] First, construct an endoscopic submucosal dissection resection area recognition model, and the construction method is as follows:

[0086] Collect the marking data of the resected area of the submucosa in endoscopic submucosal dissection;

[0087] Using the marking data of the resected area of the submucosa in endoscopic submucosal dissection, adopt a framework based on the TransUNet++ structure to train and obtain a resected area recognition model.

[0088] Collect the marking data of the resected area of the submucosa in endoscopic submucosal dissection, and the specific method includes:

[0089] Collect the endoscopic image data of cases of previous endoscopic submucosal dissections;

[0090] Obtain the marking data of the endoscopic image data, where the resected area of the submucosa in the endoscopic image is marked (which can be completed separately by multiple senior endoscopists);

[0091] Divide the marking data into a training set, a validation set, and a test set according to a ratio of 8:1:1. They are respectively used for the training, validation, and testing of the resection area recognition model.

[0092] Among them, data augmentation is performed on the training set data to obtain more data. The data augmentation methods include random horizontal and vertical flipping, random Gaussian blur, resizing within the range of [0.75, 1.25], and random cropping. All data in the training set are uniformly set to a size of 256*256. The validation set and the test set do not undergo data augmentation.

[0093] In the present invention, several data augmentation strategies are adopted, including random rotation by 20 degrees, random horizontal and vertical flipping, and random Gaussian blur. All images are adjusted to a size of 256*256 to reduce computational complexity and improve training efficiency.

[0094] Using the marked data of the resectable area in the submucosa of endoscopic submucosal dissection, a resection area recognition model is trained using a framework based on the TransUNet++ structure. The specific method includes:

[0095] Based on the TransUNet++ structure, a preliminary model is established;

[0096] According to the marked data of the resectable area in the submucosa of endoscopic submucosal dissection, the preliminary model is trained to obtain a mature resection area recognition model.

[0097] The resection area recognition model includes:

[0098] An encoder module, which is composed of a four-stage deep residual network ResNet50. The four stages are sequentially defined as x0, x1, x2, and x3. x0 is the first stage. Among them, x1, x2, and x3 are three residual convolutional blocks. Each residual convolutional block uses 1*1 and 3*3 as the convolutional kernel sizes, and a stride convolutional layer is used to halve the spatial dimension of the feature map for extracting features of different resolutions; the spatial dimension of each stage of the encoder module is 1 / 2 of the previous stage. Inside each stage, after the features are extracted by the deep residual network, the MSFM module is used to further perform multi-scale feature extraction on these features;

[0099] A decoder module, which has four stages. The four stages are sequentially defined as d3, d2, d1, and d0. d3 is the first stage, corresponding to the deepest stage x3 of the encoder module; each stage includes an MSFM module and a convolutional module. The input features are first subjected to multi-scale feature extraction and fusion by the MSFM, and then further feature fusion is performed by the convolutional module. Except for the last stage, the outputs of other stages are upsampled with the size doubled;

[0100] The Transformer module is used to link the deepest stage x3 of the encoder module with the first stage d3 of the decoder module; the output of the deepest stage x3 of the encoder module is flattened and passed to the Transformer module to learn global information and long-tail relationships;

[0101] The encoder module and the decoder module are combined in a TransUNet++ structure to form three-layer skip connections. Each stage of the encoder module serves as the head node of the three-layer skip connection, and each stage of the decoder module serves as the tail node of the three-layer skip connection. MSFM modules are respectively arranged at the intermediate nodes of the three-layer skip connection; wherein the MSFM module is used for multi-scale feature extraction and fusion;

[0102] Let the i-th and j-th node of the three-layer skip connection be x_ij, then x_00 is x0, x_04 is d0, and all the features of the intermediate nodes of the skip connection and each stage of the decoder module are obtained through the following formula:

[0103] x_01 = MSFM(cat(x0, UP(x1))),

[0104] x_11 = MSFM(cat(x1, UP(x2))),

[0105] x_02 = MSFM(cat(x0, x_01, UP(x_11))),

[0106] x_21 = MSFM(cat(x2, UP(x3))),

[0107] x_12 = MSFM(cat(x1, x_11, UP(x_21))),

[0108] x_03 = MSFM(cat(x0, x_01, x_02, UP(x_12))),

[0109] d3 = x_31 = MSFM(cat(x3, UP(Transformer))),

[0110] d2 = x_22 = MSFM(cat(x2, x_21, UP(x_31))),

[0111] d1 = x_13 = MSFM(cat(x1, x_11, x_12, UP(x_22))),

[0112] d0 = x_04 = MSFM(cat(x0, x_01, x_02, x_03, UP(x_13)));

[0113] Among them, cat is the feature concatenation operation, UP is the upsampling operation that doubles the feature map, and MSFM is the multi-scale feature fusion module.

[0114] It should be noted that nested TransUNet++ is used as the basic structure, and the deep residual network ResNet50 is used as the backbone network for encoder feature extraction. Features are extracted through four stages x0, x1, x2, and x3. The size of the feature map in the next stage is 1 / 2 of that in the previous stage, forming features with different resolutions, and then fused with the decoder through skip connections. Dense connections help reduce the semantic gap between the encoder and the decoder. The deep residual unit makes the deep network easy to train. The skip connections inside the network help propagate gradient information, improving the neural network design by reducing parameters and achieving comparable or better performance in semantic segmentation tasks.

[0115] In one embodiment, each residual convolution module consists of a combination of group normalization (GN), rectified linear unit (ReLU) activation, convolution, and identity mapping. The identity mapping connects the input and output, making the network deeper. The output of each residual block is used to construct the nested TransUNet++ structure.

[0116] It should be noted that the diversity of various submucosal structures (including fat, blood vessels, etc.) and thickness, as well as fibrosis, make the recognition of the cuttable area challenging. The high-resolution features at the shallow level contain boundary information, while the low-resolution features at the deep level contain more context and semantic information. To solve this problem, the present invention introduces a multi-scale feature extraction and fusion (MSFM) module, which uses the Inception-Resnet module to fuse the features of different-scale combinations of convolutional filters to obtain different-scale receptive fields. Our aim is to make the features contain sufficient context and boundary information simultaneously to improve the segmentation accuracy. MSFM acts on the skip link and the decoder to extract and capture features of cuttable areas with different shapes and sizes, improving the model prediction accuracy.

[0117] In one embodiment, the MSFM module consists of two parts: a convolutional block (CB) and Inception-Resnet. The CB captures information through a 3*3 convolutional layer, batch normalization (BN), and ReLU activation. Inception-Resnet consists of an identity mapping and three convolutional branches b1, b2, and b3. A BN layer follows the convolutional layer of Inception-Resnet. The identity mapping is used to retain the original feature map information to prevent gradient explosion, and residual scaling is performed through a scaling factor. Finally, a ReLU layer is used for activation and output. The MSFM module can be expressed as follows:

[0118] xl = CB(x l );

[0119]

[0120] where [ ] represents the concatenation operation, x l and x l+1 represent the input and output of the l-th layer respectively, represents the outputs of the three branches of the l-th layer of Inception-Resnet, and α represents the scaling factor.

[0121] In addition, it should be noted that the convolutional operation of CNN limits the receptive field of CNN through the kernel size. A larger receptive field can be achieved by stacking multiple CNN layers with smaller kernels or using a large kernel, but this will increase the number of parameters. To model the long-tailed dependencies between pixels, we propose to use a Transformer module based on the multi-head self-attention layer between the encoder and the decoder.

[0122] Therefore, in one embodiment, first, the output x3 of the deepest layer stage of the encoder module is converted into a two-dimensional sequence, and this two-dimensional sequence is flattened into one-dimensional vectors of trainable feature vectors Q, V, and K. The Transformer layer consists of L layers of multi-head self-attention (MSA) and multi-layer perceptron (MLP) blocks. MSA connects multiple SAs, as shown in the formula:

[0123]

[0124] MSA(Q, K, V) = Concate([H1,…,H h );

[0125] where h represents the total number of heads, and Q, K, and V are the query, key, and value respectively.

[0126] The construction of the above model is implemented using the PyTorch library. All experiments are conducted on an NVIDIA V100 Tensor Core GPU with 32GB of GPU memory.

[0127] During model training, the SGD optimizer is used for model training. During training, the learning rate is set to 0.005, the momentum is set to 0.9, the weight decay is set to 0.0001, the network batch size is set to 64, and the learning rate is updated automatically according to the following formula:

[0128]

[0129] where lr0 = 0.005, iter_num is the current iteration step, max_iterations = sample size * Epoch, Epoch = 200, and Epoch is the total training cycle.

[0130] Binary cross-entropy and Dice loss are used to measure the training effect of the model; the specific expressions are as follows:

[0131]

[0132]

[0133] L decoder = λ1L bce + λ2L dice ;

[0134] where λ1 = λ2 = 0.5; L bce represents cross-entropy; L dice represents the dice coefficient; L decoder represents the overall loss.

[0135] The model was trained for a total of 200 epochs, using Stochastic Gradient Descent (SGD) with momentum as the optimization algorithm, where the momentum was set to 0.9 and the weight decay was set to 0.0001. The initial learning rate was set to 0.005, and a polynomial learning decay rate schedule was adopted. λ1 and λ2 in the loss function were set to 0.5. Before training, the ImageNet pre-trained weights of the Transformer were loaded, which included 12 layers; the other layers were trained from scratch.

[0136] Based on the resection area recognition model constructed according to the present invention, the present invention can also provide a method for recognizing the resection area of endoscopic submucosal dissection, including:

[0137] Obtaining an endoscopic image of a subject to be subjected to endoscopic submucosal dissection;

[0138] According to the endoscopic image, the resection area recognition model obtained by the above construction method is used to recognize and segment the resected area of the submucosa.

[0139] The recognition result is as Figure 4 shown. Using the resection area recognition model constructed according to the present invention, the resection area of endoscopic submucosal dissection is further recognized, and the recognition result is basically consistent with the real boundary, proving the accuracy of the model constructed in this application and the effectiveness of the recognition method.

[0140] Similarly, a system for recognizing the resection area of endoscopic submucosal dissection can also be provided, including:

[0141] An endoscope imaging module for acquiring endoscope images of a subject to undergo endoscopic submucosal dissection;

[0142] An excision area recognition module for recognizing and segmenting the resected area of the submucosa according to the endoscope image by using the excision area recognition model obtained by the above construction method.

[0143] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the above endoscopic submucosal dissection resection area recognition method, including:

[0144] Acquiring endoscope images of a person to undergo endoscopic submucosal dissection;

[0145] According to the endoscope image, recognizing and segmenting the resected area of the submucosa by using the excision area recognition model obtained by the above construction method.

[0146] On yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the above endoscopic submucosal dissection resection area recognition method, including:

[0147] Acquiring endoscope images of a person to undergo endoscopic submucosal dissection;

[0148] According to the endoscope image, recognizing and segmenting the resected area of the submucosa by using the excision area recognition model obtained by the above construction method.

[0149] On yet another aspect, the present invention also provides an electronic device, which includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory communicate with each other through the communication bus. The processor can call the logical instructions in the memory to execute the above endoscopic submucosal dissection resection area recognition method, including:

[0150] Acquiring endoscope images of a person to undergo endoscopic submucosal dissection;

[0151] According to the endoscope image, recognizing and segmenting the resected area of the submucosa by using the excision area recognition model obtained by the above construction method.

[0152] In addition, when the logical instructions in the above storage can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0153] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0154] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.

[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.

Claims

1. A method for constructing an endoscopic mucosal resection resection area recognition model, characterized in that: include: Data on marking the resectable area of ​​the submucosal layer during endoscopic mucosal dissection were collected; Using the labeling data of the submucosal resectable area of ​​endoscopic mucosal dissection, a resection area recognition model was trained using a framework based on the TransUNet++ structure. Among them, the preliminary model is established based on the TransUNet++ structure, specifically: Use the deep residual network as the encoder to extract features and obtain features of different resolutions; Use MSFM module and convolutional network as decoder; The encoder and decoder are densely connected via skip links, and the nodes of the skip links are MSFM modules; The MSFM module is used to fuse and extract features of different resolutions; The last stage of the encoder is linked to the first stage of the decoder via a Transformer structure based on multi-head self-attention layers; Wherein, the resection area recognition model includes: The encoder module consists of a deep residual network with four stages, which are defined as x0, x1, x2 and x3, respectively. x0 is the first stage. The spatial dimension of each stage is 1 / 2 of the previous stage. Within each stage, after the deep residual network extracts features, the MSFM module is used to further perform multi-scale feature extraction on these features. The decoder module has four stages, which are defined as d3, d2, d1 and d0 in sequence. d3 is the first stage, corresponding to the deepest stage x3 of the encoder module. Each stage includes an MSFM module and a convolution module. The input features are first extracted and fused by the MSFM at multiple scales, and then further fused by the convolution module. Except for the last stage, the outputs of other stages are upsampled to double the size. A Transformer module is used to link the deepest stage x3 of the encoder module with the first stage d3 of the decoder module. The output of the deepest stage x3 of the encoder module is flattened and passed to the Transformer module for learning global information and long-tail relationships. The encoder module and the decoder module are combined with the TransUNet++ structure to form a three-layer skip link, with each stage of the encoder module as the first node of the three-layer skip link, each stage of the decoder module as the last node of the three-layer skip link, and MSFM modules are arranged at the intermediate nodes of the three-layer skip link respectively; wherein the MSFM module is used to fuse the features of different resolutions extracted at each stage; Assume that the jth node of the i-th layer of the three-layer skip link is x_ij, then x_00 is x0, x_04 is d0, and all the features of each intermediate node of the skip link and each stage of the decoder module are obtained by the following formula: x_01 = MSFM(cat(x0, UP(x1)) ), x_11 = MSFM(cat(x1, UP(x2))), x_02 = MSFM(cat(x0,x_01,UP(x_11))), x_21 = MSFM(cat(x2, UP(x3))), x_12 = MSFM(cat(x1,x_11,UP(x_21))), x_03 = MSFM(cat(x0, x_01, x_02, UP(x_12))), d3 = x_31 = MSFM(cat(x3, UP(Transformer))), d2 = x_22 = MSFM(cat(x2, x_21, UP(x_31))), d1 = x_13 =MSFM(cat(x1, x_11, x_12, UP(x_22))), d0 = x_04 = MSFM(cat(x0, x_01, x_02, x_03, UP(x_13))); Among them, cat is the feature concatenation operation, UP is the upsampling operation of doubling the feature map, and MSFM is the multi-scale feature extraction fusion module.

2. The method for constructing an endoscopic mucosal resection resection area identification model according to claim 1, characterized in that: The resectable area labeling data of the submucosal layer of endoscopic mucosal dissection is trained using a framework based on the TransUNet++ structure to obtain a resection area recognition model, including: Based on the TransUNet++ structure, a preliminary model was established; Based on the labeling data of the submucosal resectable area of ​​endoscopic mucosal dissection, a preliminary model was trained to obtain a mature resection area recognition model.

3. The method for constructing an endoscopic mucosal resection resection area identification model according to claim 1, characterized in that: The collection of the marking data of the submucosal resectable area of ​​endoscopic mucosal dissection includes: To collect endoscopic image data of cases that had undergone endoscopic mucosal dissection in the past; obtaining marking data of the endoscopic image data, wherein a submucosal resectable area in the endoscopic image is marked; The labeled data is divided into training set, validation set and test set, which are used for training, validation and testing of the resection area recognition model respectively; Among them, the training set data is enhanced to obtain more data.

4. A system for constructing an endoscopic mucosal resection resection area recognition model, characterized in that: include: A data collection device for collecting data on marking of the resectable area of ​​the submucosal layer during endoscopic mucosal dissection; The model training device is used to train the resection area recognition model using the labeling data of the submucosal resection area of ​​endoscopic mucosal dissection, using a framework based on the TransUNet++ structure; Among them, the preliminary model is established based on the TransUNet++ structure, specifically: Use the deep residual network as the encoder to extract features and obtain features of different resolutions; Use MSFM module and convolutional network as decoder; The encoder and decoder are densely connected via skip links, and the nodes of the skip links are MSFM modules; The MSFM module is used to fuse and extract features of different resolutions; The last stage of the encoder is linked to the first stage of the decoder via a Transformer structure based on multi-head self-attention layers; Wherein, the resection area recognition model includes: The encoder module consists of a deep residual network with four stages, which are defined as x0, x1, x2 and x3, respectively. x0 is the first stage. The spatial dimension of each stage is 1 / 2 of the previous stage. Within each stage, after the deep residual network extracts features, the MSFM module is used to further perform multi-scale feature extraction on these features. The decoder module has four stages, which are defined as d3, d2, d1 and d0 in sequence. d3 is the first stage, corresponding to the deepest stage x3 of the encoder module. Each stage includes an MSFM module and a convolution module. The input features are first extracted and fused by the MSFM at multiple scales, and then further fused by the convolution module. Except for the last stage, the outputs of other stages are upsampled to double the size. A Transformer module is used to link the deepest stage x3 of the encoder module with the first stage d3 of the decoder module. The output of the deepest stage x3 of the encoder module is flattened and passed to the Transformer module for learning global information and long-tail relationships. The encoder module and the decoder module are combined with the TransUNet++ structure to form a three-layer skip link, with each stage of the encoder module as the first node of the three-layer skip link, each stage of the decoder module as the last node of the three-layer skip link, and MSFM modules are arranged at the intermediate nodes of the three-layer skip link respectively; wherein the MSFM module is used to fuse the features of different resolutions extracted at each stage; Assume that the jth node of the i-th layer of the three-layer skip link is x_ij, then x_00 is x0, x_04 is d0, and all the features of each intermediate node of the skip link and each stage of the decoder module are obtained by the following formula: x_01 = MSFM(cat(x0, UP(x1)) ), x_11 = MSFM(cat(x1, UP(x2))), x_02 = MSFM(cat(x0,x_01,UP(x_11))), x_21 = MSFM(cat(x2, UP(x3))), x_12 = MSFM(cat(x1,x_11,UP(x_21))), x_03 = MSFM(cat(x0, x_01, x_02, UP(x_12))), d3 = x_31 = MSFM(cat(x3, UP(Transformer))), d2 = x_22 = MSFM(cat(x2, x_21, UP(x_31))), d1 = x_13 =MSFM(cat(x1, x_11, x_12, UP(x_22))), d0 = x_04 = MSFM(cat(x0, x_01, x_02, x_03, UP(x_13))); Among them, cat is the feature concatenation operation, UP is the upsampling operation of doubling the feature map, and MSFM is the multi-scale feature extraction fusion module.

5. A method for identifying the resection area of ​​endoscopic mucosal resection, characterized in that: include: Obtain endoscopic images of the patient to be subjected to endoscopic mucosal resection; According to the endoscopic image, the resectable area of ​​the submucosal layer is identified and segmented by the resection area recognition model obtained by the construction method according to any one of claims 1 to 3.

6. An endoscopic mucosal resection resection area recognition system, characterized in that: include: An endoscopic camera module, used to obtain an endoscopic image of a person to be subjected to endoscopic mucosal dissection; A resection area recognition module is used to identify and segment the submucosal resectable area based on the endoscopic image using the resection area recognition model obtained by the construction method described in any one of claims 1 to 3.

7. An electronic device, characterized in that: including a processor, a communication interface, a memory and a communication bus; Wherein, the processor, the communication interface, and the memory communicate with each other via a communication bus; The processor is used to call the logic instructions in the memory to execute the endoscopic mucosal resection resection area identification method described in claim 5.

8. A non-transitory computer-readable storage medium having computer instructions stored thereon, the computer instructions enabling a computer to execute the endoscopic submucosal resection resection area identification method of claim 5.

Citation Information

Patent Citations

  • Intracranial aneurysm recognition and detection method, device and system and the computer readable storage medium

    CN113706451A

  • UNet + +-based low-level glioma image segmentation method

    CN114202545A

  • Honeycomb lung lesion image segmentation identification method based on Transform semi-supervised algorithm

    CN117523203A