Enteroscope image detection method and device and electronic equipment

By using a self-supervised learning method of reconstructing a circular template and a two-sample mask in the colonoscopic image, the problem of insufficient classification accuracy of colonoscopic images in the prior art is solved, and efficient classification and auxiliary diagnosis of colonoscopic images are achieved.

CN120339725AActive Publication Date: 2025-07-18ARMY MEDICAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510797133.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-18
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

In the prior art, self-supervised learning methods based on mask reconstruction fail to effectively capture the annular topology and cross-case similarity of colonoscopic images in colonoscopic image classification, resulting in low sensitivity to key identification features extracted by encoder and insufficient classification accuracy.

Method used

The colonoscopy image is divided into multiple annular areas using annular template, combined with the checkerboard template and phase inversion view, a self-supervised learning model is constructed, and forced learning cross-case feature transfer and local-global feature coordination are reconstructed through two-sample masks, and the discriminant feature of feature expression is enhanced by using the radial information leaked by the annular template.

Benefits of technology

The accuracy of colonoscopic image classification was significantly improved, especially in the auxiliary diagnosis of distinguishing ulcerative colitis from Crohn's disease and pathological severity grading, and the sensitivity of the model to subtle pathological changes and classification accuracy were improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339725A_ABST
    Figure CN120339725A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image detection, and provides an enteroscopy image detection method and device and electronic equipment. The enteroscopy image detection method comprises the following steps: inputting an enteroscopy image into a classification model to obtain a classification result; the classification model comprises: a square feature extractor for extracting a plurality of grid feature vectors from an enteroscopy image; the annular template processing module divides the enteroscopy image into a plurality of annular areas based on an annular template, the annular template comprises a plurality of concentric rings, effective rings and invalid rings in the concentric rings are alternately distributed, and one concentric ring area is defined as one annular area; the annular feature extractor is used for acquiring a plurality of annular feature vectors; the encoder is used for encoding the plurality of grid feature vectors, the plurality of annular feature vectors and the classification marks to obtain classification mark feature vectors; and the classifier is used for carrying out classification processing on the classification mark feature vector to obtain a classification result. According to the method, explicit modeling is carried out on the ring topological structure of the enteroscopy image, and the accuracy of enteroscopy image classification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image detection, and in particular, to a colonoscopy image detection method, device, and electronic device. Background Art

[0002] Colonoscopy images generally refer to intestinal images obtained by taking pictures through an intestinal endoscope (also called colonoscope). Through computer vision technology, colonoscopy images can be classified to achieve auxiliary diagnosis of intestinal diseases. For example, ulcerative colitis (UC) and Crohn's disease (CD) can be distinguished. Another example is that the pathological severity of ulcerative colitis (UC) can be graded (four levels: 0, 1, 2, 3). How to accurately extract the key discriminant features of colonoscopy images is the key to accurate classification of colonoscopy images.

[0003] In the related art, a self-supervised learning method based on mask reconstruction (Masked Autoencoder, abbreviated as MAE) is used to train the encoder and decoder so that the encoder can accurately extract the key discriminant features of colonoscopy images. However, in the design of masks in the related art, random block or grid masks are usually adopted, and the inherent characteristics of colonoscopy images are not considered. Through the analysis and summary of a large number of colonoscopy images, the inventor found the following inherent characteristics of colonoscopy images: 1. Depth of field gradient characteristic, that is, the radial characteristic that the proximal mucosa is clear and the distal brightness decays, showing an obvious annular topological structure, which can be observed by referring to Figure 2 the example of the colonoscopy image shown; 2. The lesion-sensitive areas in colonoscopy images (such as Crohn's lesion-sensitive areas) mostly show damage to the annular structure, and it is difficult to capture such morphological clues through existing block or grid masks, such as Figure 9 the sample 1 image in The masks in the related art are difficult to capture the inherent characteristics of colonoscopy images, resulting in low sensitivity of the key discriminant features extracted by the encoder to lesion features, and the classification accuracy still needs to be improved. Summary of the Invention

[0004] This application aims to at least solve the technical problems existing in the prior art, and provides a colonoscopy image detection method, device, and electronic device.

[0005] In a first aspect, the present application provides a colonoscopy image detection method, including: obtaining a colonoscopy image; inputting the colonoscopy image into a classification model to obtain a classification result; wherein, the classification model includes: a square feature extractor for extracting a plurality of grid feature vectors from the colonoscopy image; an annular template processing module for dividing the colonoscopy image into a plurality of annular regions based on an annular template, wherein the annular template includes a plurality of concentric rings, and valid rings and invalid rings are alternately distributed among the plurality of concentric rings, and a concentric ring region is defined as an annular region; an annular feature extractor for obtaining a plurality of annular feature vectors corresponding to the plurality of annular regions; an encoder for encoding the plurality of grid feature vectors, the plurality of annular feature vectors, and a classification label to obtain a classification label feature vector; and a classifier for classifying the classification label feature vector to obtain a classification result.

[0006] In a second aspect, the present application provides a colonoscopy image detection device for implementing the colonoscopy image detection method provided in the first aspect of the present application, including: an image acquisition module for obtaining a colonoscopy image; a classification module for inputting the colonoscopy image into a classification model to obtain a classification result; wherein, the classification model includes: a square feature extractor for extracting a plurality of grid feature vectors from the colonoscopy image; an annular template processing module for dividing the colonoscopy image into a plurality of annular regions based on an annular template, wherein the annular template includes a plurality of concentric rings, and valid rings and invalid rings are alternately distributed among the plurality of concentric rings, and a concentric ring region is defined as an annular region; an annular feature extractor for obtaining a plurality of annular feature vectors corresponding to the plurality of annular regions; an encoder for encoding the plurality of grid feature vectors, the plurality of annular feature vectors, and a classification label to obtain a classification label feature vector; and a classifier for classifying the classification label feature vector to obtain a classification result.

[0007] The beneficial technical effects of the colonoscopy image detection method and device provided by the present application are as follows: Taking into account the inherent characteristics of colonoscopy images, a plurality of annular regions of the colonoscopy image are obtained through an annular template, realizing an explicit modeling of the annular topological structure of the colonoscopy image, and respectively extracting the annular feature vectors corresponding to each annular region. In this way, not only is the extraction of annular feature vectors carried out according to the depth of field gradient, but also the annular feature vectors can focus on the radial structure continuity of the intestinal lumen, enhancing the sensitivity to lesion features such as the destruction of annular folds. At the same time, a plurality of grid feature vectors of the colonoscopy image are extracted to obtain the spatial distribution information of the colonoscopy image, and the encoder is used to encode the plurality of grid feature vectors, the plurality of annular feature vectors, and the classification label, so that the classification label feature vector has significant class discriminability, thereby improving the accuracy of colonoscopy image classification.

[0008] In a third aspect, the present application provides a mask-based self-supervised learning method, including: constructing a set of sample pairs, where one sample pair is composed of randomly extracting two colonoscopy images from a colonoscopy image set; using a checkerboard template to generate a spatially decoupled view and a phase-inverted view for each sample pair, and using the phase-inverted view as the reconstruction target of the sample pair; using an annular template to generate an annular fusion view including multiple annular regions for each sample pair; constructing a self-supervised learning model, and training the self-supervised learning model using the spatially decoupled view, the phase-inverted view, and the annular fusion view of the set of sample pairs, where the self-supervised learning model includes: a square feature extractor for extracting multiple grid feature vectors from the spatially decoupled view of the sample pair; an annular feature extractor for obtaining multiple annular feature vectors corresponding to multiple annular regions of the annular fusion view of the sample pair; an encoder for encoding the multiple grid feature vectors and the multiple annular feature vectors to obtain encoded features; and a decoder for reconstructing the reconstruction target based on the encoded features after adding learnable masked tokens to obtain a reconstructed image.

[0009] The beneficial technical effects of the mask-based self-supervised learning method provided by the present application are as follows: Aiming at the problem that traditional self-supervised algorithms only perform random mask reconstruction on a single colonoscopy image to learn the local feature correlation within a single colonoscopy image, but ignore the high similarity of colonoscopy images across cases, resulting in low class discriminability of the features output by the encoder. The present application introduces a two-sample random mask reconstruction mechanism, randomly selects two colonoscopy images from an unlabeled colonoscopy image set to construct a sample pair, and generates cross-sample synthetic views (spatially decoupled view, phase-inverted view, and annular fusion view), and uses the synthetic views of the sample pair to train the self-supervised learning model, which can achieve cross-case feature transfer and local-global feature collaboration. The decoder utilizes the radial information leaked by the annular template and combines the cross-case knowledge of the two samples of the sample pair to achieve more accurate image reconstruction. Compared with the random mask of the related MAE technology, this method makes the reconstruction process more in line with the anatomical characteristics of the intestinal lumen through the prior constraint of the annular structure, and improves the class discriminability of the features output by the encoder in self-supervised learning.

[0010] Among them, cross-case feature transfer refers to using the cross-case similarity of colonoscopy images (such as limited color distribution, annular topological structure), and forcing the self-supervised learning model to learn the common features (such as the annular continuity of healthy mucosa) and local differences (such as the morphological damage of the lesion area) between cases through two-sample mask reconstruction, significantly improving the sensitivity of the self-supervised learning model to subtle pathological changes; Local-global feature collaboration means obtaining three synthetic views through the masking operations of the checkerboard template and the circular template, enabling the self-supervised learning model to simultaneously reconstruct local regions from different samples, forcing the encoder to establish semantic associations across images (such as the contrast between ulcerous regions and healthy mucosae), and enhancing the discriminability of feature expressions.

[0011] Fourthly, the present application provides a method for training a classification model, including: in the self-supervised pre-training stage, performing the steps of the self-supervised learning method based on masking provided in the third aspect of the present invention; in the supervised fine-tuning stage, performing: using the circular template processing module, the classifier, and the square feature extractor, the circular feature extractor, and the encoder obtained at the end of the self-supervised pre-training stage to construct a classification model; using the colonoscopy image classification sample set to train the classification model.

[0012] The beneficial technical effects of the method for training a classification model provided by the present application are: in addition to having the beneficial effects of the self-supervised learning method based on masking provided in the third aspect, it also has the beneficial technical effects of accurately classifying colonoscopy images and realizing accurate auxiliary diagnosis of intestinal diseases (such as distinguishing ulcerative colitis (UC) and Crohn's disease (CD), such as grading the severity of ulcerative colitis (UC)).

[0013] Fifthly, the present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method provided in the first aspect or the third aspect or the fourth aspect of the present application. Description of the Drawings

[0014] Figure 1 is a schematic flowchart of the colonoscopy image detection method provided in Embodiment 1 of the present invention; Figure 2 is an example of a colonoscopy image; Figure 3 is a schematic structural diagram of the classification model in Embodiment 1 of the present invention; Figure 4 is a schematic structural diagram of the square feature extractor in Embodiment 1 of the present invention; Figure 5 is a schematic structural diagram of the circular feature extractor in Embodiment 1 of the present invention; Figure 6 is a schematic structural diagram of the deep scene processing module in Embodiment 1 of the present invention; Figure 7 is a schematic structural diagram of the encoder in Embodiment 1 of the present invention; Figure 8is a schematic diagram of the structure of a colonoscopy image detection device provided in Example 2 of the present invention; Figure 9 is a process diagram of the self-supervised learning method provided by Embodiment 3 of the present invention; Figure 10 is a schematic diagram of the classification model training method process provided by Example 4 of the present invention; Figure 11 It is a schematic diagram of the structure of an electronic device provided in Example 5 of the present invention. DETAILED DESCRIPTION

[0015] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0016] Example 1 The present embodiment provides a colonoscopy image detection method, and the execution subject of the method includes but is not limited to at least one of the electronic devices such as a server and a terminal that can be configured to execute the method provided in the embodiment of the present application. In other words, the colonoscopy image detection method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.

[0017] In a preferred implementation of this embodiment, the flow chart of the colonoscopy image detection method is as follows: Figure 1 As shown, including: Step A1, obtaining a colonoscopy image.

[0018] In this preferred embodiment, the colonoscopic image is an intestinal image taken by an intestinal endoscope. Figure 2 is an example of a colonoscopy image, whose inner wall wrinkles are in a ring topological structure, and the color distribution is highly limited to a narrow red-yellow-white color range. The execution subject is not limited to obtaining the colonoscopy image to be detected by establishing a communication link with the intestinal endoscope; or, the colonoscopy image to be detected is pre-stored in a memory, and the execution subject reads the colonoscopy image from the memory.

[0019] Step A2: input the colonoscopy image into the classification model to obtain the classification result.

[0020] As shown Figure 3 below, the classification model includes: A square feature extractor for extracting a plurality of grid feature vectors from colonoscopy images; An annular template processing module that divides the colonoscopy image into a plurality of annular regions based on an annular template, where the annular template includes a plurality of concentric rings, and the effective rings and invalid rings in the plurality of concentric rings are alternately distributed. Defining a concentric ring region as an annular region; An annular feature extractor for obtaining a plurality of annular feature vectors corresponding to the plurality of annular regions; An encoder that encodes a plurality of grid feature vectors, a plurality of annular feature vectors, and classification labels to obtain a classification label feature vector; A classifier that classifies the classification label feature vector to obtain a classification result.

[0021] In this preferred embodiment, please refer to Figure 4 shown below. The square feature extractor includes a convolutional layer and a linear mapping module connected in sequence, where the linear mapping module includes one or more cascaded linear mapping layers. The convolutional kernel size of the convolutional layer is , the stride is , and zero-padding convolution operation is performed to divide the colonoscopy image into a plurality of pixel blocks, and the convolution features of each pixel block are obtained. The convolution features of each pixel block are then dimensionally enhanced in the channel dimension through the linear mapping module, so as to obtain the grid feature vector corresponding to each pixel block.

[0022] Exemplarily, the size of the colonoscopy image is 384*384, the convolutional kernel size of the convolutional layer is 32×32, the stride is 32. After performing the convolution operation, the colonoscopy image is divided into 36 16×16 pixel blocks, and the convolution features of each pixel block are obtained. The convolution features of each pixel block are represented by a 256-dimensional vector, and the convolution features of each pixel block use a shared linear mapping module to enhance the channel dimension to 768. Please refer to Figure 4 shown below. In this example, the linear mapping module includes a cascaded first linear mapping layer (linear mapping layer 1) and a second linear mapping layer (linear mapping layer 2), and a Relu activation function is set in the first linear mapping layer. Finally, 36 grid feature vectors are obtained, that is, 36 grid tokens, which can be represented by the token sequence denoted as, denotes the grid feature vector of the th pixel block of the colonoscopy image.

[0023] From Figure 2It can be seen that constrained by the physical imaging mechanism of the annular wide-angle lens, the colonoscopy images exhibit the inherent characteristics of the mapping from a three-dimensional cavity to a two-dimensional plane. First, lens distortion causes barrel distortion at the edges of the colonoscopy images, forming a geometric deformation with a central radial pattern; second, the depth-of-field gradient effect under macro imaging is significant. The proximal mucosa shows clear texture due to sufficient light, while the distal tissue shows brightness attenuation with increasing distance, forming a concentric dark ring feature; more importantly, the annular folds of the healthy intestinal wall form a topological structure with rotational symmetry in the two-dimensional projection, that is, an annular topological structure, while lesions (such as longitudinal ulcers in Crohn's disease) will disrupt this spatial continuity, forming important morphological diagnostic clues, such as Figure 9 the image of sample 1 in

[0024] In this preferred embodiment, the annular template includes a plurality of concentric rings. In the plurality of concentric rings, the effective rings and the invalid rings are alternately distributed. The effective rings represent the parts that need to be extracted in the masking operation, and the invalid rings represent the parts that do not need to be extracted in the masking operation. A concentric ring region is defined as an annular region. The size of the annular template is the same as that of the colonoscopy image. Therefore, the outer edge of the outermost concentric ring is rectangular or square, and the inner edge is circular. The number and width of the concentric rings in the annular template can be adjusted according to the actual size of the colonoscopy image. Figure 5 An example of the annular template is given. In this example, the annular template includes 7 concentric rings, and the white concentric rings are defined as effective rings, and the black concentric rings are defined as invalid rings. From the outside to the inside, they alternate between white and black, and there are 4 effective rings.

[0025] In this preferred embodiment, to quickly extract multiple annular regions of the colonoscopy image, preferably, the annular template processing module performs: Step C21, performing a masking operation on the colonoscopy image using the annular template to obtain multiple first annular regions; specifically, after the masking operation, the first annular regions corresponding to all the effective rings are obtained, that is, multiple first annular regions are obtained.

[0026] Step C22, performing a masking operation on the colonoscopy image using a phase-inverted annular template that forms a radial complementary pair with the annular template to obtain multiple second annular regions; the phase-inverted annular template includes a plurality of concentric rings with the same number of concentric rings as the annular template, and in the plurality of concentric rings, the effective rings and the invalid rings are alternately distributed. A concentric ring region is defined as an annular region. The concentric rings at the corresponding positions of the effective rings of the annular template in the phase-inverted annular template are invalid rings, and the concentric rings at the corresponding positions of the invalid rings of the annular template in the phase-inverted annular template are effective rings. Figure 5A phase-inverted annular template example is given, which includes 7 concentric rings. The white concentric rings are defined as valid rings, and the black concentric rings are defined as invalid rings. From the outside to the inside, it alternates between black and white, and there are 3 valid rings. Through the phase-inverted annular template, the annular regions corresponding to the invalid rings in the annular template can be extracted.

[0027] Step C23, multiple first annular regions and multiple second annular regions form multiple annular regions. In Figure 5 the given example, there are 4 first annular regions and 3 second annular regions, a total of 7 annular regions.

[0028] The radius ranges of the 4 first annular regions are respectively: (central proximal region), (central proximal region), (middle section), (peripheral distal region).

[0029] The radius ranges of the 3 second annular regions are respectively: (central proximal region), (middle section), (peripheral distal region). represents the height or width of the colonoscopy image, and its height is equal to its width.

[0030] In this preferred embodiment, exemplarily, the annular feature extractor includes multiple feature extraction networks corresponding one-to-one to multiple annular regions. Each feature extraction network is used to extract the annular feature vector of its corresponding annular region. The feature extraction network includes a cascaded sparse convolutional layer and an average pooling layer. For each annular region, mask operations are performed on all image regions (invalid regions) in the colonoscopy image except the image region (valid region) corresponding to this annular region, to obtain the mask image corresponding to each annular region. Using the sparse convolutional layer to process the mask image corresponding to each annular region can skip the calculation of the invalid regions and only obtain the image features corresponding to the annular region, and then use the average pooling layer to compress the image features to obtain the annular feature vector corresponding to the annular region.

[0031] In this preferred embodiment, please refer to Figure 7 , the encoder includes a feature splicing module, a position encoding module, and a feature interaction module connected in sequence.

[0032] The feature splicing module is used to splice the multiple grid feature vectors generated by the square feature extractor and the multiple annular feature vectors generated by the annular feature extractor into a mixed vector sequence.

[0033] In the above example, the feature splicing module is used to splice the 36 grid feature vectors generated by the square feature extractor, that is, 36 grid token sequences , and the 7 circular feature vectors generated by the circular feature extractor, i.e., 7 circular token sequences are concatenated into a mixed vector sequence , denotes the circular feature vector corresponding to the th circular region, and d = 768 is the feature dimension (i.e., the number of channels).

[0034] The position encoding module is used to inject a position encoding matrix into the mixed vector sequence. The position encoding matrix includes the position encoding vectors corresponding to each grid feature vector and the position encoding vectors corresponding to each circular feature vector. The position encoding vectors corresponding to the grid feature vectors adopt the standard grid coordinate encoding, and the position encoding vectors corresponding to the circular feature vectors are circular polar coordinate encodings.

[0035] In the above example, the position encoding matrix is represented as , and the matrix the first 36 position encoding vectors adopt the standard grid encoding, and the last 7 position encoding vectors adopt the circular polar coordinate encoding , denotes the normalized value of the circular radius of the circular region, and the circular radius is the average of the minimum radius and the maximum radius of the circular region.

[0036] The feature interaction module adopts a multi-layer Transformer encoder structure. Exemplarily, 12 layers of Transformer encoders can be adopted. In the training of the classification model, the multi-head attention mechanism in the Transformer encoder automatically establishes the mapping relationship between the grid feature vectors, the circular feature vectors, and the classification tokens. Therefore, through the feature interaction module, the accurate classification token feature vectors can be obtained according to the mapping relationship between the grid feature vectors, the circular feature vectors, and the classification tokens obtained by training. In the above example, the classification token feature vector corresponding to the classification token [CLS] is .

[0037] In this preferred embodiment, the classifier is not limited to using a multi-layer perceptron (MLP) or a multi-layer cascaded fully connected structure to implement the classification decision. Preferably, the classifier adopts a two-layer multi-layer perceptron (MLP) to obtain the classification result which is represented as: where, , respectively represent the weight matrix and the bias matrix of the first fully connected layer in the two-layer multi-layer perceptron (MLP), , respectively represent the weight matrix and the bias matrix of the second fully connected layer in the two-layer multi-layer perceptron (MLP), denotes the GELU activation function.

[0038] In an application scenario of this embodiment, the colonoscopy image detection method provided in this embodiment is used for the classification of ulcerative colitis and Crohn's disease, and the classification results include the confidence levels of three categories: ulcerative colitis, Crohn's disease, and normal.

[0039] In another application scenario of this embodiment, the colonoscopy image detection method provided in this embodiment is used for grading the severity level of ulcerative colitis, corresponding to the pathological severity levels of 0-3. The classification results include the confidence levels of four levels: 0, 1, 2, and 3.

[0040] In another preferred implementation manner of this embodiment, please refer to Figure 5 , the annular feature extractor includes: A preliminary feature extraction module that respectively extracts the depth-of-field features of multiple annular regions to obtain multiple annular preliminary features.

[0041] A depth-of-field level processing module that guides the selection of at least one expert model through the annular radius of the annular region, and uses the selected at least one expert model to process the annular preliminary features of the annular region to obtain the annular feature vector corresponding to the annular region.

[0042] In this preferred implementation manner, first, the annular preliminary features of the annular region are extracted. In order to achieve feature extraction at different depth-of-field levels, at least one expert model is adaptively selected according to the annular radius of each annular region, and the selected at least one expert model is used to process the annular preliminary features of the annular region to obtain the annular feature vector corresponding to the annular region, realizing a dynamic expert model selection mechanism. The expert model for the proximal region focuses on the villus microstructure, and the expert model for the distal region models the overall brightness distribution, thereby realizing the differential extraction and fusion of high-resolution texture in the proximal region and global brightness features in the distal region. Improve the accuracy of intestinal disease classification and the grading of the severity of ulcerative colitis.

[0043] In this preferred implementation manner, the expert model is preferably but not limited to using a multi-layer perceptron (MLP) network, such as a two-layer MLP.

[0044] Further preferably, to more accurately extract the annular preliminary features of each annular region in the colonoscopy image, please refer to Figure 5 , the preliminary feature extraction module includes: A first branch that respectively processes multiple first annular regions through a first sparse convolution processing module to obtain multiple first annular preliminary features; A second branch that respectively processes multiple second annular regions through a second sparse convolution processing module to obtain multiple second annular preliminary features; Multiple first annular preliminary features and multiple second annular preliminary features form multiple annular preliminary features.

[0045] In this preferred embodiment, please refer to Figure 5 , in the first branch, the first sparse convolution processing module includes a cascaded first sparse convolutional layer and a first average pooling layer. First, a mask image corresponding to each first annular region is generated. The mask image corresponding to each first annular region is generated based on the colonoscopy image. In this mask image, only the image region corresponding to the first annular region is the valid region, and all image regions other than the image region corresponding to the first annular region are defined as invalid regions, and the invalid regions are masked. Then, the first sparse convolutional layer is used to process the mask image corresponding to each first annular region respectively. Sparse convolution can ignore the masked part and only perform convolution processing on the valid part to obtain convolution features; then, the first average pooling layer is used to compress the convolution features of each first annular region to obtain the first annular preliminary features corresponding to each first annular region. In the above example, the convolution kernel size of the first sparse convolutional layer is 3×3, and the first average pooling layer is preferably a global average pooling layer, and 4 first annular preliminary features with 256 dimensions are obtained.

[0046] In this preferred embodiment, please refer to Figure 5 , in the second branch, the second sparse convolution processing module includes a cascaded second sparse convolutional layer and a second average pooling layer. First, a mask image corresponding to each second annular region is generated. The mask image corresponding to each second annular region is generated based on the colonoscopy image. In this mask image, only the image region corresponding to the second annular region is the valid region, and all image regions other than the image region corresponding to the second annular region are defined as invalid regions, and the invalid regions are masked. Then, the second sparse convolutional layer is used to process the mask image corresponding to each second annular region respectively. Sparse convolution can ignore the masked part and only perform convolution processing on the valid part to obtain convolution features; then, the second average pooling layer is used to compress the convolution features of each second annular region to obtain the second annular preliminary features corresponding to each second annular region. In the above example, the convolution kernel size of the second sparse convolutional layer is 3×3, and the second average pooling layer is preferably a global average pooling layer, and 3 second annular preliminary features with 256 dimensions are obtained. The 4 first annular preliminary features and the 3 second annular preliminary features form 7 annular preliminary features.

[0047] In this preferred embodiment, the depth-of-field features of each annular region (i.e., each annular band), such as proximal texture and distal brightness gradient, are independently extracted by the first sparse convolutional layer and the second sparse convolutional layer to solve the problem of feature distortion caused by lens distortion.

[0048] Further preferably, for adaptively selecting an expert model, please see Figure 6 , the depth-of-field level processing module includes: The expert model library includes multiple expert models, and the multiple expert models respectively correspond to the circular radii of different circular regions; The position encoding unit adds a position encoding vector to the preliminary circular feature of each circular region to obtain a feature position encoding vector, where the position encoding vector is obtained by encoding the circular radius of the circular region; The model selection unit calculates the product vector of the feature position encoding vector of each circular region and the gating weight matrix, and generates an expert weight vector corresponding to the circular region according to the product vector; The expert fusion processing unit selects at least one expert model to process the feature position encoding vector of the circular region according to the expert weight vector corresponding to each circular region, and fuses the processing results of at least one expert model to obtain a circular feature vector corresponding to the circular region.

[0049] In this preferred embodiment, it is defined in the expert model library that expert models, , which is consistent with the number of circular regions. Exemplarily, ; represents the th expert model, which corresponds to the circular radius of a circular region.

[0050] In this preferred embodiment, the position encoding unit is actually a dynamic gating mechanism: for the preliminary circular feature of any circular region, such as for the preliminary circular feature of the circular region , first calculate the position encoding vector of the circular region , is the cosine position encoding of the circular radius of the circular region . Finally, splice after the preliminary circular feature to obtain the feature position encoding vector of the circular region . By injecting the cosine position encoding of the circular radius , the circular radius information is guided to make the model focus on the radial structure continuity of the intestinal lumen, and the sensitivity to lesion features such as the destruction of circular folds is enhanced.

[0051] In this preferred embodiment, in the step where the model selection unit calculates the product vector of the feature position encoding vector of each circular region and the gating weight matrix, for the circular region , the circular region of the feature position encoding vector and the gating weight matrix of the product vector is​ Preferably, the selection model selection unit generates an annular region according to the following formula The corresponding expert weight vector : Wherein, represents the gating weight matrix, , is a learnable matrix, specifically, it can be a linear layer; represents the annular region The characteristic position encoding vector of; represents the largest in the product vector items of elements, and the non-maximum items of elements are set to zero; represents the normalization function. is greater than 1 and less than , preferably, . That is to say, the numerical values of the two largest items in the product vector are retained, and the elements of the remaining non-maximum items are set to zero. In this way, a new vector is obtained, and after being processed by , the obtained expert weight vector is a vector including elements, and the elements are vectors of 0 or 1, . Each element of the expert weight vector corresponds to an expert model, indicating the processing weight of the corresponding expert model for the annular feature vector of the annular region . 1 means processing, and 0 means not processing. The largest in the product vector items have corresponding elements of 1 in , and any other item in the product vector except these items has a corresponding element of 0 in .

[0052] In this preferred embodiment, the expert fusion processing unit selects at least one expert model according to the expert weight vector corresponding to each annular region to process the characteristic position encoding vector of the annular region, and fuses the processing results of at least one expert model to obtain the annular feature vector corresponding to the annular region. Specifically, the processing process of the expert fusion processing unit for the annular region is: ; Wherein, It is a learnable matrix used to map the fusion features of the processing results of at least one expert model to a new feature space, specifically a linear layer. This design aims to achieve feature extraction at different scene depths by guiding the selection of corresponding regional expert models through the annular radius. Represents an annular region The corresponding annular feature vector. Represents an annular region The corresponding expert weight vector In The Element corresponding to the expert model. Represents the th expert model For the annular region The preliminary annular features Processing result.

[0053] Embodiment 2 This embodiment provides a colonoscopy image detection device for implementing the colonoscopy image detection method provided in Embodiment 1. Please refer to Figure 8 As shown, in a preferred implementation manner of this embodiment, the device includes: An image acquisition module for acquiring colonoscopy images; A classification module for inputting the colonoscopy image into a classification model to obtain a classification result.

[0054] The image acquisition module and the classification module correspond one-to-one to steps A1 and A2 in Embodiment 1 and will not be elaborated here.

[0055] Please refer to Figure 3 , the classification model includes: A square feature extractor for extracting multiple grid feature vectors from the colonoscopy image; An annular template processing module for dividing the colonoscopy image into multiple annular regions based on an annular template, where the annular template includes multiple concentric rings, and the effective rings and invalid rings are alternately distributed among the multiple concentric rings. Defining a concentric ring region as an annular region; An annular feature extractor for obtaining multiple annular feature vectors corresponding to the multiple annular regions; An encoder for encoding the multiple grid feature vectors, the multiple annular feature vectors, and the classification markers to obtain a classification marker feature vector; A classifier for classifying the classification marker feature vector to obtain a classification result.

[0056] In this embodiment, the specific structure and working principle of the classification model have been elaborated in detail in Embodiment 1 and will not be elaborated here.

[0057] Embodiment 3 In the related art of using a self-supervised learning method based on masked reconstruction (Masked Autoencoder, abbreviated as MAE) to train an encoder and a decoder so that the encoder can accurately extract the key discriminative features of colonoscopy images, only single colonoscopy images are randomly masked and reconstructed to learn the local feature correlations within a single colonoscopy image, ignoring the high similarity of colonoscopy images across cases. The cross-case similarity of colonoscopy images is manifested in that the color distribution of colonoscopy images is highly restricted within a narrow color gamut of red-yellow-white, mainly reflecting pathological states such as mucosal congestion and ulcers. This color monotonicity results in a very high visual similarity of the normal mucosal regions of different patients, while the key discriminative features only exist in local microstructural changes (such as villus morphology, microvascular pattern, etc.). The single-sample learning mode of the related art makes it difficult for the model to capture the common features (such as the circular folds of healthy mucosa and the narrow color gamut distribution) and subtle pathological differences (such as ulcer morphology and villus structure changes) between different cases, limiting the generalization ability of the model for disease discrimination and grading.

[0058] Based on this, this embodiment provides a masked-based self-supervised learning method to solve the above problems based on the inherent characteristics of colonoscopy images and the high similarity across cases. The masked-based self-supervised learning method provided in this embodiment is used for the pre-training of the square feature extractor, circular feature extractor, and encoder in Embodiment 1. The execution subject of this method includes but is not limited to at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute this method provided in the embodiments of the present application. In other words, the masked-based self-supervised learning method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0059] In this embodiment, the masked-based self-supervised learning method includes: Step B1, constructing a sample pair set, where a sample pair is composed of randomly extracting two colonoscopy images from a set of colonoscopy images.

[0060] In one example, as Figure 9 shown, randomly select two colonoscopy images and from an unlabeled set of colonoscopy images to form a sample pair.

[0061] Step B2: Using the checkerboard template, generate a spatially decoupled view and a phase-inverted view for each sample pair, and use the phase-inverted view as the reconstruction target for the sample pair; using the annular template, generate an annular fusion view including multiple annular regions for each sample pair.

[0062] As Figure 9 the mask template part of, two spatially complementary checkerboard templates are pre-constructed. Exemplarily, construct spatially complementary checkerboard templates and . The checkerboard template : 6×6 checkerboard, starting with black squares, see Figure 9 the first checkerboard template from top to bottom in the mask template part in. The checkerboard template : 6×6 checkerboard, starting with white squares, see Figure 9 the second checkerboard template from top to bottom in the mask template part in. The checkerboard template is the anti-phase checkerboard template of the checkerboard template . Pre-construct the annular template and the phase-inverted annular template that forms a radially complementary pair with the annular template , the templates and see Figure 9 the third and fourth annular templates from top to bottom in the mask template part in, the templates and have been introduced in detail in Embodiment 1 and will not be elaborated here.

[0063] Exemplarily, define three synthetic views as: Spatially decoupled view: Phase-inverted view: Annular fusion view: Among them, represents a per-pixel masking operation, and it can be seen that the three synthetic views simultaneously retain the local features of and in the sample pair and their spatial associations. In the above template design, the grid features of natural images are fused through the checkerboard templates and , and the annular topological features unique to the intestinal lumen are fused using the annular templates and , facilitating the extraction of grid feature vectors and annular feature vectors.

[0064] In this embodiment, using the phase-inverted view As a reconstruction target, the self-supervised learning model only needs to predict the missing small regions based on the local features of adjacent blocks (such as texture, color continuity, etc.), with a lower task complexity and more stable gradient update.

[0065] Step B3: Construct a self-supervised learning model and train the self-supervised learning model using the spatially decoupled view, phase-inverted view, and annular fusion view of the sample pair set.

[0066] Please refer to Figure 9 , the self-supervised learning model includes: A square feature extractor for extracting multiple grid feature vectors from the spatially decoupled view of the sample pair ; Exemplarily, as Figure 9 shown, extract the grid feature vectors of 36 16×16 pixel blocks. The specific structure of the square feature extractor has been described in detail in Embodiment 1 and will not be elaborated here.

[0067] An annular feature extractor for obtaining multiple annular feature vectors corresponding to multiple annular regions of the annular fusion view of the sample pair; Exemplarily, as Figure 9 shown, extract 7 annular feature vectors corresponding to 7 annular regions. The specific structure of the annular feature extractor has been described in detail in Embodiment 1 and will not be elaborated here.

[0068] An encoder for encoding and processing multiple grid feature vectors and multiple annular feature vectors to obtain encoded features. The specific structure of the encoder has been described in detail in Embodiment 1 and will not be elaborated here. Exemplarily, as Figure 9 shown, during the training of each sample pair, the encoded features output by the encoder include 43 tokens.

[0069] A decoder for reconstructing the reconstruction target based on the encoded features after adding learnable mask tokens to obtain a reconstructed image. Exemplarily, as Figure 9 shown, set the mask tokens for the decoder to learn ( ), concatenate the encoded features (43 tokens) output by the encoder with the learnable mask tokens ( ), and input the concatenated tokens into the decoder for decoding processing to obtain the reconstruction target. Specifically, the decoder reconstructs the original information by predicting the pixel values of each mask token (i.e., the mask block). The last layer of the decoder is a linear projection layer, and its output channel number is equal to the number of pixels of each mask block. Then, the output of the decoder is reshaped to construct the reconstructed image.

[0070] In this embodiment, a 12-layer Transformer is used as the encoder, and its multi-head attention mechanism automatically establishes from the colonoscopy image and The pixel-level correspondence between grid tokens and circular tokens is used to achieve cross-case association. Additionally, the information carried by circular tokens can be leaked to grid tokens to reduce the reconstruction difficulty of the decoder.

[0071] In this embodiment, although it is required to reconstruct the complete phase-inverted view , there is knowledge transfer between samples within a sample pair. The high similarity of colonoscopy images enables the model to utilize the common features of colonoscopy images and for knowledge transfer, and the information from the circular fusion view contains information about some of the image patches to be reconstructed, that is, the feature vectors from will leak a part to the decoder for reconstructing the target image . Therefore, in this embodiment, the decoder can obtain a complete reconstructed image.

[0072] In this embodiment, in step B3, the self-supervised learning model is iteratively trained using the spatial decoupled view, phase-inverted view, and circular fusion view of the sample pair set. In each iterative training, the spatial decoupled view and circular fusion view of a sample pair are input into the square feature extractor and circular feature extractor of the self-supervised learning model. The decoder outputs the reconstructed image, the learning loss is calculated based on the reconstructed image and the reconstruction target, and the network parameters of the self-supervised learning model are adjusted according to the learning loss, and then enter the next iterative training until the training stop condition is reached, such as the number of iterative training reaches the preset number or the learning loss is less than the preset loss threshold.

[0073] Learning loss is defined as: Pixel loss: Structure loss: where represents the height of the colonoscopy image, in pixels; represents the width of the colonoscopy image, in pixels; represents the number of channels of the colonoscopy image; represents the reconstructed image at position , channel of the pixel value; represents the reconstructed target phase-inverted view at position , channel of the pixel value; represents the reconstructed image The local mean, which reflects the brightness feature; Indicates the reconstructed target phase-inverted view The local mean of; Indicates the reconstructed image The local variance of; Indicates the reconstructed target phase-inverted view The local variance of; Indicates the reconstructed image And the reconstructed target phase-inverted view The covariance of, which characterizes the structural similarity; The first stability constant The value is 0.0001; The second stability constant The value is 0.0009; Indicates the pixel value dynamic range. When the pixel value dynamic range of the colonoscopy image is 0 - 255, . Indicates the square operation of the L2 norm.

[0074] This embodiment has the following technical effects: (1) Cross-case feature migration: Utilize the cross-case similarity of colonoscopy images (such as limited color distribution, circular topological structure), and through dual-sample mask reconstruction, force the model to learn the common features between cases (such as the circular continuity of healthy mucosa) and local differences (such as morphological destruction in the lesion area), significantly improving the model's sensitivity to subtle pathological changes.

[0075] (2) Local-global feature collaboration: Through pixel-level mask superposition (such as ), the model needs to reconstruct local regions from different samples simultaneously, forcing the encoder to establish cross-image semantic associations (such as the comparison between the ulcer area and healthy mucosa), enhancing the discriminability of feature expression.

[0076] (3) The decoder utilizes the radial information leaked by the circular mask (such as the complementary region in ), combines the cross-case knowledge of the dual samples, and realizes more accurate image reconstruction. Compared with the random mask of MAE, this method makes the reconstruction process more in line with the anatomical characteristics of the intestinal lumen through the prior constraint of the circular structure, improving the feature quality of self-supervised pre-training.

[0077] Example 4 This embodiment provides a method for training a classification model. The execution subject of this method includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute this method provided by the embodiments of the present application. In other words, the method for training a classification model can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0078] Please refer to Figure 10 , the method for training a classification model includes: In the self-supervised pre-training stage, execute the steps of the mask-based self-supervised learning method provided in Embodiment 3; obtain the pre-trained square feature extractor, circular feature extractor, and encoder.

[0079] In the supervised fine-tuning stage, execute: Step C1, as Figure 10 shown, use the circular template processing module, classifier, and the square feature extractor, circular feature extractor, and encoder obtained at the end of the self-supervised pre-training stage to construct a classification model; Step C2, use the colonoscopy image classification sample set to train the classification model.

[0080] In this embodiment, the specific structures of the classification model, circular template processing module, and classifier have been elaborated in detail in Embodiment 1 and will not be repeated here. The square feature extractor, circular feature extractor, and encoder are obtained through self-supervised pre-training.

[0081] In this embodiment, the colonoscopy image classification sample set includes multiple colonoscopy images and the corresponding category labels for each colonoscopy image. According to the different application scenarios of the classification model, the category labels are defined differently. When performing auxiliary diagnosis of ulcerative colitis and Crohn's disease, the category labels can include ulcerative colitis, Crohn's disease, and normal. When performing the grading of the pathological severity of ulcerative colitis, the category labels include four levels: 0, 1, 2, and 3. The input of the supervised fine-tuning stage is a single randomly sampled colonoscopy image. The square feature extractor directly processes the complete colonoscopy image (instead of the checkerboard mosaic image in the self-supervised pre-training stage) and generates 36 grid tokens. The circular feature extractor also receives the complete single colonoscopy image, retains the generation logic of obtaining 7 circular tokens in Embodiment 3, but removes the cross-sample fusion mechanism. The decoder of self-supervised learning is removed in the supervised fine-tuning stage.

[0082] In step C2, a classification model is trained using the colonoscopy image classification sample set. During the training process, the standard cross-entropy loss function can be used to calculate the loss, and the network parameters of the classification model are adjusted according to the loss.

[0083] Example 5 This example provides an electronic device, such as Figure 11 shown, which is a schematic structural diagram of the electronic device for executing the methods provided in Embodiment 1, Embodiment 3, and Embodiment 4 of the present invention. The electronic device may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may further include a computer program stored in the memory 11 and executable on the processor 10, such as the method programs provided in Embodiment 1, Embodiment 3, and Embodiment 4 of the present invention.

[0084] Among them, the processor 10 may be composed of integrated circuits in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions packaged, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as executing the methods provided in Embodiment 1, Embodiment 3, and Embodiment 4 of the present invention), and calling the data stored in the memory 11, to execute various functions of the electronic device and process data.

[0085] The memory 11 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. The memory 11 may be an internal storage unit of the electronic device in some embodiments, such as the mobile hard disk of the electronic device. The memory 11 may also be an external storage device of the electronic device in other embodiments, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 11 may also include both the internal storage unit and the external storage device of the electronic device. The memory 11 can not only be used to store application software installed in the electronic device and various types of data, such as the code of the method programs provided in Embodiment 1, Embodiment 3, and Embodiment 4 of the present invention, but also be used to temporarily store data that has been output or will be output.

[0086] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to implement the connection communication between the memory 11 and at least one processor 10, etc.

[0087] The communication interface 13 is used for the communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface can include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface can be a display, an input unit (such as a keyboard), and optionally, the user interface can also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display can also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device and to display a visual user interface.

[0088] Figure 11 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 11 the shown structure does not constitute a limitation on the electronic device, and it can include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0089] For example, although not shown, the electronic device can also include a power source (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source can also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or an inverter, and a power status indicator. The electronic device can also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0090] It should be understood that the embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0091] Furthermore, if the modules / units integrated in the electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0092] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A colonoscopy image detection method, characterized in that, Including: Obtain colonoscopy images; Input the colonoscopy images into a classification model to obtain a classification result; Among them, the classification model includes: A square feature extractor for extracting multiple grid feature vectors from the colonoscopy images; An annular template processing module that divides the colonoscopy images into multiple annular regions based on an annular template. Among them, the annular template includes multiple concentric rings, and the effective rings and invalid rings in the multiple concentric rings are alternately distributed. Define a concentric ring region as an annular region; An annular feature extractor for obtaining multiple annular feature vectors corresponding to the multiple annular regions; An encoder for encoding the multiple grid feature vectors, the multiple annular feature vectors, and the classification labels to obtain a classification label feature vector; A classifier for classifying the classification label feature vector to obtain a classification result.

2. The colonoscopy image detection method according to claim 1, wherein The annular feature extractor includes: A preliminary feature extraction module for respectively extracting the depth-of-field features of the multiple annular regions to obtain multiple annular preliminary features; A depth-of-field level processing module that guides the selection of at least one expert model through the annular radius of the annular region, and uses the selected at least one expert model to process the annular preliminary features of the annular region to obtain the annular feature vector corresponding to the annular region.

3. The colonoscopy image detection method according to claim 2, characterized in that, The depth-of-field level processing module includes: An expert model library including multiple expert models, and the multiple expert models respectively correspond to the annular radii of different annular regions; A position encoding unit for adding a position encoding vector to the preliminary annular features of each annular region to obtain a feature position encoding vector, where the position encoding vector is obtained by encoding the annular radius of the annular region; A model selection unit for calculating the product vector of the feature position encoding vector of each annular region and the gating weight matrix, and generating an expert weight vector corresponding to the annular region according to the product vector; An expert fusion processing unit for selecting at least one expert model according to the expert weight vector corresponding to each annular region to process the feature position encoding vector of the annular region, and fusing the processing results of the at least one expert model to obtain the annular feature vector corresponding to the annular region.

4. The colonoscopy image detection method according to claim 3, wherein The model selection unit generates an annular region according to the following formula The corresponding expert weight vector : , Among them, represents the gating weight matrix, represents the characteristic position encoding vector of the circular region ; represents keeping the largest terms in the product vector , and setting the elements of the non-largest terms to zero; represents the normalization function.

5. The colonoscopy image detection method according to any one of claims 1-4, characterized in that, The annular template processing module executes: Perform a masking operation on the colonoscopy images using the annular template to obtain multiple first annular regions; Perform a masking operation on the colonoscopy images using a phase-inverted annular template that forms a radial complementary pair with the annular template to obtain multiple second annular regions; The multiple first annular regions and the multiple second annular regions form multiple annular regions.

6. The colonoscopy image detection method according to claim 5, wherein, The preliminary feature extraction module includes: A first branch for respectively processing the multiple first annular regions through a first sparse convolution processing module to obtain multiple first annular preliminary features; A second branch for respectively processing the multiple second annular regions through a second sparse convolution processing module to obtain multiple second annular preliminary features; The multiple first annular preliminary features and the multiple second annular preliminary features form multiple annular preliminary features.

7. An enteroscope image detection device, characterized in that, For implementing the colonoscopy image detection method described in any one of claims 1-6, characterized in that it includes: An image acquisition module for acquiring colonoscopy images; A classification module for inputting the colonoscopy images into a classification model to obtain a classification result; Among them, the classification model includes: A square feature extractor for extracting multiple grid feature vectors from the colonoscopy images; The annular template processing module divides the colonoscopy image into multiple annular regions based on an annular template, where the annular template includes multiple concentric rings, and valid rings and invalid rings are alternately distributed among the multiple concentric rings. A concentric ring region is defined as an annular region; The annular feature extractor obtains multiple annular feature vectors corresponding to the multiple annular regions; The encoder performs encoding processing on multiple grid feature vectors, multiple annular feature vectors, and classification labels to obtain classification label feature vectors; The classifier performs classification processing on the classification label feature vectors to obtain classification results.

8. A self-supervised learning method based on masking, characterized in that, It includes: Construct a sample pair set, where a sample pair is composed of randomly extracting two colonoscopy images from a colonoscopy image set; Use a checkerboard template to generate a spatially decoupled view and a phase-inverted view for each sample pair, and use the phase-inverted view as the reconstruction target of the sample pair; use an annular template to generate an annular fusion view including multiple annular regions for each sample pair; Construct a self-supervised learning model, and use the spatially decoupled view, phase-inverted view, and annular fusion view of the sample pair set to train the self-supervised learning model, where the self-supervised learning model includes: The square feature extractor is used to extract multiple grid feature vectors from the spatially decoupled view of the sample pair; The annular feature extractor obtains multiple annular feature vectors corresponding to the multiple annular regions of the annular fusion view of the sample pair; The encoder performs encoding processing on multiple grid feature vectors and multiple annular feature vectors to obtain encoded features; The decoder reconstructs the reconstruction target based on the encoded features after adding a learnable masked token to obtain a reconstructed image.

9. A method for training a classification model, characterized in that, It includes: The self-supervised pre-training stage performs the steps of the masked-based self-supervised learning method described in claim 8; The supervised fine-tuning stage performs: Use the annular template processing module, classifier, and the square feature extractor, annular feature extractor, and encoder obtained at the end of the self-supervised pre-training stage to form a classification model; Use a colonoscopy image classification sample set to train the classification model.

10. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method described in any one of claims 1-6, 8, 9.

Citation Information

Patent Citations

  • Endoscopic image lesion detection method based on fusion of global and local features

    CN102722735A