Enteroscope navigation method based on entry image features

Through self-supervised learning and supervised fine-tuning based on Vision Transformer, a multi-level feature reconstruction task is constructed, which solves the problem of low accuracy of forward direction prediction and safe distance judgment in colonoscopic navigation, and improves the accuracy and stability of colonoscopic navigation.

CN120431064APending Publication Date: 2025-08-05HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510574972.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing colonoscopic navigation methods have low accuracy in the prediction of forward direction and the determination of safety distances, making it difficult to effectively navigate in complex intestinal environments, resulting in problems such as examination failure, prolonged time and patient discomfort.

Method used

The colonoscopic image feature navigation method based on Vision Transformer is adopted to construct multi-level feature reconstruction tasks through self-supervised learning and supervised fine-tuning, learn intestinal anatomical structure characteristics, and decouple the navigation task into forward direction prediction and safe distance judgment problems, improving navigation accuracy.

Benefits of technology

The accuracy, stability and efficiency of colonoscopic navigation methods in complex intestinal environments are significantly improved, and the risk of navigation misjudgment is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431064A_ABST
    Figure CN120431064A_ABST
Patent Text Reader

Abstract

The invention relates to an enteroscope navigation method based on entry image features, in particular to an enteroscope navigation method based on entry image features. The objective of the invention is to solve the problem of low accuracy of forward direction prediction and safe distance discrimination of an existing enteroscope navigation method. The invention provides a Vision Transform-based enteroscope image direction prediction model, pre-training of unlabeled data is carried out in combination with self-supervised learning, and prediction and visual labeling of the forward direction of an enteroscope are realized through supervised fine tuning. The method comprises the following steps: firstly, carrying out self-supervised learning by utilizing a large number of unmarked enteroscopy images to reconstruct tasks to obtain efficient visual representation, so as to learn intestinal anatomical structure features; subsequently, supervised fine adjustment is performed on a small amount of labeled data, so that the accuracy of prediction of the forward direction of the enteroscope and judgment of the distance between the lens and the intestinal wall is improved. The method is applied to the field of enteroscope navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a colonoscopy navigation method based on enteroscope image features. Background Art

[0002] As the gold standard for the diagnosis and treatment of digestive system diseases, colonoscopy plays an irreplaceable role in clinical scenarios such as early screening for colorectal cancer, evaluation of inflammatory bowel disease, and localization of gastrointestinal bleeding. Statistics from the World Health Organization in 2020 show that colorectal cancer ranks third in the incidence of malignant tumors worldwide, and early detection of adenomatous polyps can significantly reduce the risk of cancer. However, the effectiveness of the implementation of this key diagnostic and treatment technology is highly dependent on the experience level of the operating physician. In the process, there may be clinical risks such as colonoscopy path loss, excessive inflation causing pain to the patient, and even intestinal perforation. [1] ([1] Ye Guoliang, Hu Kefeng. Research progress on the implementation of high-quality colonoscopy [J]. Zhejiang Medicine, 2025, 47(01): 1-8.). Colonoscopy is an important tool for the diagnosis and treatment of gastrointestinal diseases in clinical practice, and plays an irreplaceable role in the early detection of polyps, cancer and other lesions. However, due to the complex intestinal structure, the advancement direction of the colonoscope is difficult to determine, and it is extremely dependent on the operator's experience and technical level. Improper operation may lead to examination failure, prolonged time, and even cause patient discomfort or intestinal wall damage. Therefore, intelligent colonoscopy navigation technology is of great significance for improving the safety and efficiency of the examination. At present, endoscopic-assisted technology based on image processing mainly focuses on tasks such as polyp detection and lesion segmentation, while research on colonoscopy navigation is still in a blank stage. In addition, colonoscopy image data has challenges such as high noise, low contrast and uneven illumination, which further increases the difficulty of intelligent navigation.

[0003] In recent years, deep learning has made breakthrough progress in the field of medical image analysis and has been widely used in tasks such as polyp detection, lesion segmentation, and tumor classification. [2] ([2] Xu Yuqiu, Tang Wentao, Lu Yang, et al. Progress in the application of artificial intelligence in colonoscopy [J]. Modern Digestion and Interventional Diagnosis and Treatment, 2022, 27(04): 403-6+12.) However, there are still few studies on colonoscopy navigation. Existing research mainly focuses on lesion detection, while there is no mature intelligent solution for the key task of predicting the direction of colonoscopy. [3-5]([3] Weng Xuejian, Zheng Endian, Teng Miaomiao. Research on improving the detection rate of colorectal polyps and adenomas by artificial intelligence-assisted colonoscopy [J]. General Practice Clinical and Education, 2024, 22(11): 981-3+7. [4] He Wenqiong, Li Yinpeng. Research progress of artificial intelligence in the diagnosis of colon adenomatous polyps [J]. Chinese Prescription Drugs, 2024, 22(11): 175-8. [5] Miao Yaxuan, Wang Xicheng, Wang Haixing, et al. Application progress of artificial intelligence-assisted system in endoscopic screening of colorectal cancer [J]. Electronic Journal of Comprehensive Oncology Treatment, 2025, 11(01): 10-20+7.). In addition, the high complexity of the intestinal anatomical structure constitutes multiple technical obstacles. First, the fold morphology of the intestinal cavity is dynamically variable, and the anatomical characteristics of different intestinal segments, such as the ascending colon, transverse colon, and descending colon, are significantly different; second, endoscopic images are limited by lighting conditions and are prone to halo effects. Many endoscopic images have mucosal reflections or secretion interference. This makes it difficult for traditional image processing methods based on convolutional neural networks (CNN) to effectively model [6] ([6] Meng Xiangfu, Zhang Zhichao, Yu Chunlin, et al. Colonoscopic polyp image segmentation method combined with boundary self-knowledge distillation [J]. Journal of Image and Graphics, 2025, 30(02): 589-600.).

[0004] As a core method for diagnosing anorectal diseases, colonoscopy is divided into two stages: inserting the endoscope and withdrawing it. In clinical practice, physicians need to use real-time images to determine the direction of the cavity during the insertion of the endoscope, and this process is highly dependent on operational experience. However, due to factors such as the physiological curvature of the intestine, obstruction of mucosal folds, and image artifacts, cavity tracking deviation is very likely to occur, making it difficult to insert the endoscope and even causing complications such as mucosal damage, which seriously affects the examination efficiency and patient experience. [7] ([7] Wang Hao, Liu Hong, Li Guodong. Application of artificial intelligence technology in digestive endoscopy [J]. Journal of Shandong First Medical University (Shandong Academy of Medical Sciences), 2024, 45(11): 700-4.).

[0005] The current development of intelligent colonoscopy equipment provides new ideas for precise navigation, but existing technical solutions still have significant shortcomings.

[0006] Dark area navigation method: This technology infers direction based on the light and dark distribution of a single-frame image. Its limitations are reflected in two aspects: first, the spatial correspondence between dark area features and anatomical structures lacks deep learning support, and it is easy to cause directional misjudgment when there is interference from intestinal fluid reflection or complex folds; second, feature extraction modules based on traditional image processing technologies, such as threshold segmentation and morphological operations, have difficulty modeling the temporal continuity of the intestinal cavity, and their robustness is significantly reduced when the lens is close to the intestinal wall or there is motion blur.

[0007] Morphological analysis method: Although the geometric features of the colon loop are used for direction estimation, it has inherent defects: First, the dynamic deformation characteristics of intestinal tissue limit the generalization ability of the fixed curvature template matching method; second, the feature extraction process is highly dependent on image clarity. When intestinal contents are attached or the camera moves rapidly, traditional edge detection algorithms have difficulty in stably obtaining effective morphological features, causing navigation direction drift.

[0008] Monocular ranging method: This scheme attempts to infer three-dimensional spatial information from two-dimensional images, but it has fundamental flaws: on the one hand, depth estimation based on shadow changes is easily affected by non-uniform lighting in the intestine and is prone to generating erroneous depth values in highly reflective areas of the mucosa; on the other hand, traditional algorithms lack prior learning of intestinal anatomical features and it is difficult to construct reliable spatial mapping in areas with sparse texture, resulting in systematic deviations in direction estimation and low computational efficiency. Summary of the Invention

[0009] The purpose of the present invention is to solve the problems of low accuracy in predicting the direction of advance and judging the safety distance in existing colonoscopy navigation methods, and to propose a colonoscopy navigation method based on the features of the image entering the scope.

[0010] A colonoscopy navigation method based on the features of the image entering the endoscope has the following specific steps:

[0011] Step 1: Acquire unlabeled colonoscopic image data, process the unlabeled colonoscopic image data, and obtain processed unlabeled colonoscopic image data;

[0012] Step 2: inputting the processed unlabeled colonoscopy image data obtained in step 1 into the reconstruction network model, the reconstruction network model outputting the reconstructed colonoscopy image data, and training the reconstruction network model to obtain a trained reconstruction network model;

[0013] Step 3: Obtain a training set of colonoscopy images with labels of forward direction v(sinθ, cosθ) θ is the forward direction angle;

[0014] The colonoscopy image training set with the forward direction v (sinθ, cosθ) label Image samples in Input the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features

[0015] Step 4: Get the safe distance y dis Labeled colonoscopy image training set

[0016] Keep a safe distance dis Labeled colonoscopy image training set Image samples in Input the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features

[0017] Step 5: For the forward direction prediction task: ViT encoder outputs image features Input direction prediction network, direction prediction network outputs forward direction prediction value Obtain a trained direction prediction network;

[0018] Step 6: For the safety distance discrimination task: ViT encoder outputs image features Input the safety distance judgment network, and the safety distance judgment network outputs a binary judgment of whether the distance to the intestinal wall is too close Obtain a trained safety distance discrimination network;

[0019] Step 7: Obtain the colonoscopy image data to be tested, divide the colonoscopy image data to be tested into blocks of size P×P, flatten each block, and output the embedded representation of each block through a linear layer after flattening. The embedded representations of all blocks are input into the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features;

[0020] The ViT encoder outputs image features and inputs them into the trained direction prediction network, which then outputs the predicted value of the colonoscope's forward direction. Prediction value based on the direction of colonoscopy Calculate radian predictions

[0021] The ViT encoder outputs image features and inputs them into the trained safety distance discrimination network. The trained safety distance discrimination network outputs a prediction value of whether the distance to the intestinal wall is too close.

[0022] Preferably, in step 1, unlabeled colonoscopic image data is obtained, and the unlabeled colonoscopic image data is processed to obtain processed unlabeled colonoscopic image data; the specific process is:

[0023] 1) Obtain unlabeled colonoscopic image data, perform image segmentation on the unlabeled colonoscopic image data, and obtain a colonoscopic image containing only the intestinal area;

[0024] 2) Enhance the colonoscopic image containing only the intestinal area to obtain an enhanced colonoscopic image; the specific process is as follows:

[0025] The colonoscopic image containing only the intestinal area was rotated 90 degrees, 180 degrees, and 270 degrees clockwise to obtain the enhanced colonoscopic image;

[0026] 3) Standardize the enhanced colonoscopic image to obtain a standardized colonoscopic image; the specific process is as follows:

[0027] The pixel values of the enhanced colonoscopy image are in the range of [0, 255]. The enhanced colonoscopy image is normalized and the pixel values of the image are adjusted to the range of [0, 1].

[0028] Preferably, in step 2, the processed unlabeled colonoscopy image data obtained in step 1 is input into the reconstruction network model, the reconstruction network model outputs the reconstructed colonoscopy image data, and the reconstruction network model is trained to obtain a trained reconstruction network model; the specific process is:

[0029] 1) The processed unlabeled colonoscopy image data X obtained in step 1 is unlabled The jth sample image x in un Divide into blocks of size P×P, blocks, and the number of channels in each block is C;

[0030] in,

[0031] x un ∈R H×W×C , H, W, C are each sample image x un The height, width and number of channels of , R is a real number;

[0032] P×P represents the height × width of each block;

[0033] 2) Flatten each block;

[0034] After flattening, it is mapped to the D-dimensional feature space through a linear layer as follows:

[0035] z p =W p x p +b p ,p=1,2,…N

[0036] in,

[0037] x p is the pixel vector representation of the p-th block after flattening;

[0038] W p is the weight matrix of the linear layer;

[0039] b p is the bias term;

[0040] z p is the embedded representation of the p-th block output by the linear layer;

[0041] 3) Combine the embedding representations of N blocks into a sequence Z and add classification token z to the sequence Z cls , get the added classification token z cls The sequence Z0 is expressed as:

[0042] Z=[z1,z2,…,z N ]∈R N×D

[0043] Z0=[z cls ,z1,z2,…,z N ]

[0044] Among them, z1 is the embedding representation of the first block, z2 is the embedding representation of the second block, and z N is the embedded representation of the Nth block; R is a real number;

[0045] 4) Introduce position code E:

[0046] E=[e cls ,e1,e2,…,e N ]

[0047] Among them, e cls is the classification token z cls location information;

[0048] e1 is the embedding of the first block, indicating the position information of z1, and e2 is the embedding of the second block, indicating the position information of z2. N is the embedding representation z of the Nth block N location information;

[0049] 5) Based on adding classification token z cls The sequence Z0 and the position code E are used to get the encoder input sequence Z input :

[0050] Z input =Z0+E

[0051] 6) Randomly occlude the input sequence Z input For a part of the block, define an occlusion mask: M={m1,m2,…,m N},m i ∈{0,1}

[0052] Among them, m i =1 means the i-th block is not blocked, m i =0 means the i-th block is occluded;

[0053] 7) From the input sequence Z input Extract the unoccluded blocks to form a sequence Z visible :

[0054] Z visible ={Z input |m i =1},i=1,2,…,N

[0055] 8) Sequence Z of unoccluded blocks visible Input ViT encoder to encode the unoccluded Patch, and ViT encoder outputs the feature representation H of the unoccluded block pretrain ;

[0056] 9) The feature representation H of the unoccluded block output by the ViT encoder pretrain and the position code of the occlusion block E i Passed together into the decoder, the decoder outputs the reconstructed pixel value of the occluded block

[0057] 10) Reconstructed pixel value of the i-th occlusion block based on the decoder output and the original pixel value I of the i-th occlusion block i Calculate the reconstruction loss function L recon ; expressed as:

[0058]

[0059] Among them, N m is the total number of occluded blocks;

[0060] 11) Repeat 1)-10) until the reconstruction loss function converges and obtains the trained reconstruction network model.

[0061] Preferably, in 8), the sequence Z of the unobstructed blocks is visible Input ViT encoder to encode the unoccluded Patch, and ViT encoder outputs the feature representation H of the unoccluded block pretrain ; expressed as:

[0062] H pretrain =f θ (Z visible )

[0063] in,

[0064] f θ For ViT encoder;

[0065] Z visible is a sequence of unoccluded blocks;

[0066] H pretrain Feature representation of unoccluded blocks output by the ViT encoder.

[0067] Preferably, in 9), the feature representation H of the unoccluded block output by the ViT encoder is pretrain and the position code of the occlusion block E i Passed together into the decoder, the decoder outputs the reconstructed pixel value of the occluded block Expressed as:

[0068]

[0069] in,

[0070] g φ For ViT decoder;

[0071] is the reconstructed pixel value of the i-th occlusion block output by the ViT decoder;

[0072] E i Encode the position of the i-th occluded block.

[0073] Preferably, the reconstructed pixel value of the i-th occlusion block output by the decoder in 10) is and the original pixel value I of the i-th occlusion block i Calculate the reconstruction loss function L recon ; expressed as:

[0074]

[0075] Among them, N m is the total number of occluded blocks.

[0076] Preferably, in step 3, a colonoscopy image training set with a label of forward direction v (sinθ, cosθ) is obtained. θ is the forward direction angle;

[0077] The colonoscopy image training set with the forward direction v (sinθ, cosθ) label Image samples in Input the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features Expressed as:

[0078]

[0079] in,

[0080] Represented by image Extracted embedding representations of all blocks;

[0081] The acquisition process is:

[0082] The image sample Divide into blocks of size P×P, flatten each block, and output the embedded representation of each block through the linear layer after flattening. The embedded representation of all blocks consists of

[0083] Preferably, the step 4 obtains the belt safety distance y dis Labeled colonoscopy image training set

[0084] Keep a safe distance dis Labeled colonoscopy image training set Image samples in Input the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features Expressed as:

[0085]

[0086] in,

[0087] Represented by image Extracted embedding representations of all blocks;

[0088] The acquisition process is:

[0089] The image sample Divide into blocks of size P×P, flatten each block, and output the embedded representation of each block through the linear layer after flattening. The embedded representation of all blocks consists of

[0090] Preferably, in step 5, for the forward direction prediction task: the ViT encoder outputs image features Input direction prediction network, direction prediction network outputs forward direction prediction value Get the trained direction prediction network; the specific process is:

[0091] Step 5.1: The direction prediction network includes a fully connected layer FC, a ReLU activation function layer, a Dropout layer, and a fully connected layer FC in sequence;

[0092] Step 52: ViT encoder outputs image features Input direction prediction network, direction prediction network outputs the predicted value of the colonoscope's forward direction

[0093] Step 53: Prediction of the colonoscope's forward direction based on the output of the direction prediction network And the true value v(sinθ,cosθ) is used to calculate the mean square error loss function; it is expressed as:

[0094]

[0095] in,

[0096] N′ should be a colonoscopy image training set with labels of forward direction v(cosθ, sinθ) The total number of image samples in v i′ For the training set The true value of the forward direction of the i′th image sample in ;

[0097] Training set for the direction prediction network output The predicted value of the forward direction of the i′th image sample in ;

[0098] Step 54: Repeat steps 52 to 53 to traverse the colonoscopy image training set with the forward direction v (sinθ, cosθ) label Image samples in the loss function Converge and obtain a trained direction prediction network.

[0099] Preferably, in step 6, for the safety distance discrimination task: the ViT encoder outputs image features Input the safety distance judgment network, and the safety distance judgment network outputs a binary judgment of whether the distance to the intestinal wall is too close Obtain a trained safety distance discrimination network; the specific process is:

[0100] Step 6.1: The safety distance discrimination network includes a fully connected layer FC, a ReLU activation function layer, a Dropout layer, and a fully connected layer FC in sequence;

[0101] Step 62: ViT encoder outputs image features Input the safety distance judgment network, and the safety distance judgment network outputs a binary judgment prediction value of whether the distance to the intestinal wall is too close

[0102] Step 6.3: Binary classification prediction value based on the output of the safety distance judgment network to determine whether the distance to the intestinal wall is too close and the truth value y dis Calculate the cross entropy loss function; expressed as:

[0103]

[0104] in,

[0105] N″ is the safety distance y dis Labeled colonoscopy image training set The total number of image samples in ;

[0106] For safe distance dis Labeled colonoscopy image training set The true value of whether the distance between the i″th sample and the intestinal wall is too close;

[0107] The training set output by the safety distance discrimination network The predicted value of whether the distance between the i″th sample and the intestinal wall is too close;

[0108] Step 64: Repeat steps 62 to 63 to traverse the colonoscopy image training set with safety distance labels Image samples in the loss function Converge and obtain a trained safety distance discrimination network.

[0109] The beneficial effects of the present invention are:

[0110] This paper introduces the Vision Transformer (ViT) architecture into the field of colonoscopy navigation. The structure of ViT is as follows: Figure 2 As shown, the present invention constructs an intelligent navigation framework based on the self-supervised-supervised hybrid learning paradigm based on ViT. In response to the pain point of scarce navigation annotation data for colonoscopy images, the present invention designs a multi-level feature reconstruction task: by randomly masking 75% of the image area and reconstructing the mucosal texture features, the model can learn the anatomical features of different areas of the intestine in an unsupervised state. In the supervised fine-tuning stage, the navigation task is decoupled into a two-classification problem of forward direction prediction (0-2π continuous regression) and judgment of whether it is at a safe distance from the intestinal wall; the present invention improves the accuracy of forward direction prediction and safe distance judgment of the colonoscopy navigation method; experimental verification shows that the model of the present invention is significantly better than the traditional CNN model in both the forward direction prediction task and the safe distance judgment task on the test set.

[0111] The present invention proposes a colonoscopy image direction prediction model based on Vision Transformer (ViT), combines self-supervised learning to perform pre-training on unlabeled data, and realizes the prediction and visual annotation of the colonoscopy direction through supervised fine-tuning. First, the present invention uses a large number of unlabeled colonoscopy images for self-supervised learning to obtain efficient visual representations for reconstruction tasks, thereby learning the anatomical structure characteristics of the intestine. Subsequently, supervised fine-tuning is performed on a small amount of labeled data to improve the accuracy of colonoscopy direction prediction and judgment of the distance between the lens and the intestinal wall. Experimental results show that this method can effectively capture global and local features in complex intestinal environments, and improve the accuracy and stability of colonoscopy direction prediction and judgment of the distance between the lens and the intestinal wall. BRIEF DESCRIPTION OF THE DRAWINGS

[0112] Figure 1This is a diagram of the colonoscopy navigation model based on ViT of the present invention;

[0113] Figure 2 This is the structure diagram of ViT, where the value of L is 24. DETAILED DESCRIPTION

[0114] Specific implementation method 1: Combination Figure 1 This embodiment describes a colonoscopy navigation method based on the image features of the in-scope image. The specific process is as follows:

[0115] Step 1: Acquire unlabeled colonoscopic image data, process the unlabeled colonoscopic image data, and obtain processed unlabeled colonoscopic image data;

[0116] Step 2: inputting the processed unlabeled colonoscopy image data obtained in step 1 into the reconstruction network model, the reconstruction network model outputting the reconstructed colonoscopy image data, and training the reconstruction network model to obtain a trained reconstruction network model;

[0117] Step 3: Obtain a training set of colonoscopy images with labels of forward direction v(sinθ, cosθ) θ is the forward direction angle;

[0118] The origin is the center of the image, the horizontal right side of the origin is 0°, and the angle range is [0°, 360°) and the radian range is [0, 2π);

[0119] The colonoscopy image training set with the forward direction v (sinθ, cosθ) label Image samples in Input the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features

[0120] Step 4: Get the safe distance y dis Labeled colonoscopy image training set

[0121] Keep a safe distance dis Labeled colonoscopy image training set Image samples in Input the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features

[0122] Step 5: For the forward direction prediction task: ViT encoder outputs image features Input direction prediction network, direction prediction network outputs forward direction prediction value Obtain a trained direction prediction network;

[0123] Step 6: For the safety distance discrimination task: ViT encoder outputs image features Input the safety distance judgment network, and the safety distance judgment network outputs a binary judgment of whether the distance to the intestinal wall is too close Obtain a trained safety distance discrimination network;

[0124] Step 7: Obtain the colonoscopy image data to be tested, divide the colonoscopy image data into P×P patches, flatten each patch, and output the embedded representation of each patch through a linear layer. The embedded representations of all patches are input into the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features.

[0125] The ViT encoder outputs image features and inputs them into the trained direction prediction network, which then outputs the predicted value of the colonoscope's forward direction. Prediction value based on the direction of colonoscopy Calculate radian predictions

[0126] The ViT encoder outputs image features and inputs them into the trained safety distance discrimination network. The trained safety distance discrimination network outputs a prediction value of whether the distance to the intestinal wall is too close.

[0127] Specific embodiment 2: This embodiment differs from specific embodiment 1 in that in step 1, unlabeled colonoscopic image data is obtained, and the unlabeled colonoscopic image data is processed to obtain processed unlabeled colonoscopic image data; the specific process is:

[0128] 1) Obtain unlabeled colonoscopic image data, perform image segmentation on the unlabeled colonoscopic image data, and obtain a colonoscopic image containing only the intestinal area;

[0129] 2) Enhance the colonoscopic image containing only the intestinal area to obtain an enhanced colonoscopic image; the specific process is as follows:

[0130] The colonoscopic images containing only the intestinal area were rotated clockwise by 90 degrees, 180 degrees, and 270 degrees (to expand the number of training samples) to obtain enhanced colonoscopic images;

[0131] 3) Standardize the enhanced colonoscopic image to obtain a standardized colonoscopic image; the specific process is as follows:

[0132] The pixel values of the enhanced colonoscopy image are in the range of [0, 255] (any value in the interval [0, 255]). The enhanced colonoscopy image is standardized and the pixel values of the image are adjusted to the range of [0, 1] (any value in the interval [0, 1]).

[0133] Other steps and parameters are the same as those in the first embodiment.

[0134] Specific embodiment three: This embodiment differs from specific embodiments one or two in that, in step two, the processed unlabeled colonoscopy image data obtained in step one is input into the reconstruction network model, the reconstruction network model outputs the reconstructed colonoscopy image data, and the reconstruction network model is trained to obtain a trained reconstruction network model; the specific process is as follows:

[0135] For the unlabeled data X unlabled , pre-trained through the self-supervised task of occlusion reconstruction, and using Visual ViT to learn efficient visual feature representation from a large amount of unlabeled data.

[0136] 1) The processed unlabeled colonoscopy image data X obtained in step 1 is unlabled The jth sample image x in un Divide into P×P blocks (Patch), a total of blocks, and the number of channels in each block is C;

[0137] in,

[0138] x un ∈R H×W×C , H, W, C are each sample image x un The height, width and number of channels of , R is a real number;

[0139] P×P represents the height × width of each block;

[0140] The dimensions of each block before flattening are (16, 16, 3), which are height, width, and number of channels respectively;

[0141] 2) Flatten each block; the dimension of each block after flattening is 768;

[0142] After flattening, it is mapped to a D-dimensional feature space through a linear layer (the dimension of each block is flattened and converted to 1024 dimensions through a linear layer) as follows:

[0143] z p =W p x p +b p ,p=1,2,…N

[0144] in,

[0145] x p is the pixel vector representation of the p-th block after flattening;

[0146] W p is the weight matrix of the linear layer (linear);

[0147] b p is the bias term;

[0148] z p is the embedded representation of the p-th block output by the linear layer;

[0149] 3) Combine the embedding representations of N blocks into a sequence Z and add classification token z to the sequence Z cls , get the added classification token z cls The sequence Z0 is expressed as:

[0150] Z=[z1,z2,…,z N ]∈R N×D

[0151] Z0=[z cls ,z1,z2,…,z N ]

[0152] Among them, z1 is the embedding representation of the first block, z2 is the embedding representation of the second block, and z N is the embedded representation of the Nth block; R is a real number;

[0153] The classification token is a learnable vector inserted at the beginning of the ViT input sequence that represents the global features of the entire image;

[0154] 4) Introduce position code E:

[0155] E=[e cls ,e1,e2,…,e N ]

[0156] Among them, e cls is the classification token z cls location information;

[0157] e1 is the embedding of the first block, indicating the position information of z1, and e2 is the embedding of the second block, indicating the position information of z2. N is the embedding representation z of the Nth block N location information;

[0158] Because the Transformer module in ViT cannot directly perceive the spatial position of patches, position encoding is introduced to provide each patch with its relative or absolute position in the image. By incorporating position information into the patch embedding, the model can understand the positional relationship of each patch within the overall image. The self-attention mechanism can incorporate position encoding when calculating the relationship between different patches, better capturing spatial context.

[0159] 5) Based on adding classification token z cls The sequence Z0 and the position code E are used to get the encoder input sequence Zinput :

[0160] Z input =Z0+E

[0161] 6) During the training process, randomly block the input sequence Z input For a patch in the image, define an occlusion mask:

[0162] M={m1,m2,…,m N},m i ∈{0,1}

[0163] Among them, m i =1 means the i-th block is not blocked, m i =0 means the i-th block is occluded;

[0164] 7) From the input sequence Z input Extract the unoccluded blocks to form a sequence Z visible :

[0165] Z visible ={Z input |m i =1},i=1,2,…,N

[0166] 8) Sequence Z of unoccluded blocks visible Input ViT encoder to encode the unoccluded Patch, and ViT encoder outputs the feature representation H of the unoccluded block pretrain ;

[0167] 9) The feature representation H of the unoccluded block output by the ViT encoder pretrain and the position code of the occlusion block E i Passed together into the decoder, the decoder outputs the reconstructed pixel value of the occluded block

[0168] 10) Reconstructed pixel value of the i-th occlusion block based on the decoder output and the original pixel value I of the i-th occlusion block i Calculate the reconstruction loss function L recon ; expressed as:

[0169] Use Mean Squared Error (MSE) as the reconstruction loss function to calculate the error between the original occluded patch and the reconstructed patch;

[0170]

[0171] Among them, N m is the total number of occluded blocks;

[0172] 11) Repeat 1)-10) to traverse the unlabeled colonoscopy image data X after step 1 unlabled The sample image is trained (traversal stops when the reconstruction loss function converges) until the reconstruction loss function converges and the trained reconstruction network model is obtained.

[0173] Other steps and parameters are the same as those in the first or second embodiment.

[0174] Specific embodiment 4: This embodiment differs from any one of specific embodiments 1 to 3 in that the sequence Z of the unobstructed blocks in 8) is visible Input ViT encoder to encode the unoccluded Patch, and ViT encoder outputs the feature representation H of the unoccluded block pretrain ; expressed as:

[0175] H pretrain =f θ (Z visible )

[0176] in,

[0177] f θ For ViT encoder;

[0178] Z visible is a sequence of unoccluded blocks;

[0179] H pretrain Feature representation of unoccluded blocks output by the ViT encoder.

[0180] The other steps and parameters are the same as those in the first to third embodiments.

[0181] Specific embodiment 5: This embodiment differs from any one of specific embodiments 1 to 4 in that in 9), the feature representation H of the unoccluded block output by the ViT encoder is used. pretrain and the position code of the occlusion block E i Passed together into the decoder, the decoder outputs the reconstructed pixel value of the occluded block Expressed as:

[0182]

[0183] in,

[0184] g φ For ViT decoder;

[0185] is the reconstructed pixel value of the i-th occlusion block output by the ViT decoder;

[0186] E i Encode the position of the i-th occluded block.

[0187] The other steps and parameters are the same as those in the first to fourth embodiments.

[0188] Specific embodiment 6: This embodiment differs from any one of specific embodiments 1 to 5 in that the reconstructed pixel value of the i-th occlusion block output by the decoder in 10) is and the original pixel value I of the i-th occlusion block i Calculate the reconstruction loss function L recon ; expressed as:

[0189] Use Mean Squared Error (MSE) as the reconstruction loss function to calculate the error between the original occluded patch and the reconstructed patch;

[0190]

[0191] Among them, N m is the total number of occluded blocks.

[0192] The other steps and parameters are the same as those in the first to fifth embodiments.

[0193] Specific embodiment seven: This embodiment differs from any one of the specific embodiments one to six in that in step three, a colonoscopy image training set with a label of the forward direction v (sinθ, cosθ) is obtained. θ is the forward direction angle;

[0194] The origin is the center of the image, the horizontal right side of the origin is 0°, and the angle range is [0°, 360°) and the radian range is [0, 2π);

[0195] The colonoscopy image training set with the forward direction v (sinθ, cosθ) label Image samples in Input the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features Expressed as:

[0196]

[0197] in,

[0198] Represented by image Extracted embedding representations of all blocks;

[0199] The acquisition process is:

[0200] The image sample Divide into blocks of size P×P, flatten each block, and output the embedded representation of each block through the linear layer after flattening. The embedded representation of all blocks consists of

[0201] The other steps and parameters are the same as those in the first to sixth embodiments.

[0202] Specific embodiment eight: This embodiment differs from any one of specific embodiments one to seven in that the safety distance y is obtained in step four. dis Labeled colonoscopy image training set

[0203] Keep a safe distance dis Labeled colonoscopy image training set Image samples in Input the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features Expressed as:

[0204]

[0205] in,

[0206] Represented by image Extracted embedding representations of all blocks;

[0207] The acquisition process is:

[0208] The image sample Divide into blocks of size P×P, flatten each block, and output the embedded representation of each block through the linear layer after flattening. The embedded representation of all blocks consists of

[0209] The other steps and parameters are the same as those in the first to seventh embodiments.

[0210] Specific embodiment 9: This embodiment differs from any one of specific embodiments 1 to 8 in that in step 5, for the forward direction prediction task: the ViT encoder outputs the image feature Input direction prediction network, direction prediction network outputs forward direction prediction value Get the trained direction prediction network; the specific process is:

[0211] Step 5.1: The direction prediction network includes a fully connected layer FC, a ReLU activation function layer, a Dropout layer, and a fully connected layer FC in sequence;

[0212] Step 52: ViT encoder outputs image features Input direction prediction network, direction prediction network outputs the predicted value of the colonoscope's forward direction (The last step of the direction prediction network is L2 normalization, so the direction prediction network output value The values have been L2 normalized to ensure that the sum of the squares of sinθ and cosθ is 1);

[0213] Step 53: Prediction of the colonoscope's forward direction based on the output of the direction prediction network And the true value v(sinθ,cosθ) is used to calculate the mean square error (MSE) loss function; it is expressed as:

[0214]

[0215] in,

[0216] N′ should be a colonoscopy image training set with labels of forward direction v(cosθ, sinθ) The total number of image samples in ;

[0217] v i′ For the training set The true value of the forward direction of the i'th image sample in;

[0218] Training set for the direction prediction network output The predicted value of the forward direction of the i′th image sample in ;

[0219] Step 54: Repeat steps 52 to 53 to traverse the colonoscopy image training set with the forward direction v (sinθ, cosθ) label Image samples in the middle (stop traversal when the reconstruction loss function converges) until the loss function Converge and obtain a trained direction prediction network.

[0220] The other steps and parameters are the same as those in Specific Embodiments 1 to 8.

[0221] Specific embodiment 10: The difference between this embodiment and any one of the specific embodiments 1 to 9 is that in step 6, for the task of determining the safe distance: the ViT encoder outputs the image features Input the safety distance judgment network, and the safety distance judgment network outputs a binary judgment of whether the distance to the intestinal wall is too close Obtain a trained safety distance discrimination network; the specific process is:

[0222] Step 6.1: The safety distance discrimination network includes a fully connected layer FC, a ReLU activation function layer, a Dropout layer, and a fully connected layer FC in sequence;

[0223] Step 62: ViT encoder outputs image features Input the safety distance judgment network, and the safety distance judgment network outputs a binary judgment prediction value of whether the distance to the intestinal wall is too close

[0224] Step 6.3: Binary classification prediction value based on the output of the safety distance judgment network to determine whether the distance to the intestinal wall is too close and the truth value y dis Calculate the cross entropy loss function; expressed as:

[0225]

[0226] in,

[0227] N″ is the safety distance y dis Labeled colonoscopy image training set The total number of image samples in ;

[0228] For safe distance dis Labeled colonoscopy image training set The true value of whether the distance between the i″th sample and the intestinal wall is too close;

[0229] The training set output by the safety distance discrimination network The predicted value of whether the distance between the i″th sample and the intestinal wall is too close;

[0230] Step 64: Repeat steps 62 to 63 to traverse the colonoscopy image training set with safety distance labels Image samples in the loss function Convergence (traversal stops when the cross entropy loss function converges) to obtain a trained safety distance discrimination network.

[0231] The other steps and parameters are the same as those in Specific Embodiments 1 to 9.

[0232] In order to make full use of the existing large amount of unlabeled colonoscopy image data and effectively mine the feature information in colonoscopy images, this study obtained an effective image feature extractor based on self-supervised pre-training of massive unlabeled colonoscopy images, and designed a deep neural network module for two downstream tasks combined with the feature extractor, including a direction prediction network and a distance discrimination network. In the next stage, the entire network was fine-tuned in a supervised manner using labeled colonoscopy images. The overall workflow is as follows: Figure 1 shown.

[0233] In the task of predicting the direction of colonoscopy images, the type and quality of data directly affect the performance and generalization ability of the model. In order to make full use of existing data resources, this study divides colonoscopy image data into two categories: labeled data and unlabeled data: (1) Unlabeled data: does not contain labeling information, only the original colonoscopy image. This type of data occupies the majority of the dataset and is suitable for pre-training through self-supervised learning, using ViT to learn efficient visual feature representations from a large amount of unlabeled data. This method can reduce dependence on manual labeling and improve the generalization performance of the model. (2) Labeled data: contains clear labeling information such as the direction of the colonoscope and whether it is too close to the intestinal wall. This type of data is suitable for supervised learning.

[0234] After self-supervised pre-training, the ViT encoder can be considered an efficient feature extractor for colonoscopy images with good transferability. Based on this ViT encoder, this study constructed a heading prediction network and a safety distance discrimination network. Using a supervised learning strategy, the model was gradually fine-tuned using labeled colonoscopy image data, ultimately resulting in a colonoscopy guidance model with navigation capabilities. During the supervised fine-tuning phase, random occlusion of the input image was eliminated to ensure feature integrity. Furthermore, a staged training strategy was employed to train the heading prediction network and the safety distance discrimination network separately at different stages.

[0235] During the training process of the forward direction prediction network, in order to solve the problem of sudden errors between 0 and 2π in the regression of circular direction radians, the present invention adopts a vector decomposition strategy: the radian θ is decomposed into the cosine component cosθ and the sine component sinθ for joint regression. This representation method not only eliminates the problem of periodic mutations in radian values, but also ensures the unitization characteristics of the prediction vector through the L2 norm constraint. Specifically, given a labeled radian ∈ [0, 2π), we convert it into a two-dimensional vector v = (sinθ, cosθ) as a supervision target, and satisfy ||v||2 = 1. Regression under this geometric constraint effectively improves the model's perception of the continuity and stability of direction angles.

[0236] The following examples are used to verify the beneficial effects of the present invention:

[0237] Example 1:

[0238] Experimental results and analysis

[0239] Experimental environment: This experiment is based on the AutoDL platform. The experimental environment is Ubuntu 20.04 operating system, NVIDIA RTX 4090D GPU, 90GB memory, and uses software and libraries such as PyTorch 2.0.0, Python 3.8, and CUDA 11.8.

[0240] Experimental data set: The data used in the present invention is colonoscopy data, collected from the OLYMPUS CF-H290I ultra-high-definition electronic colonoscope in Japan. The data set includes 2510 colonoscopy images in JPG format and a 3-minute and 22-second colonoscopy video in MP4 format. 1200 of the 2510 colonoscopy images are accompanied by label information on the direction of travel, and this part of the data is used for the task of predicting the direction of travel; 1310 images are accompanied by label information on whether the distance to the intestinal wall is too close, and this part of the data is used for the task of judging whether the distance is too close. The present invention extracted 5000 frames from the colonoscopy video as unlabeled pre-training data. In these two tasks, the data set was randomly divided into training set, validation set and test set, with a ratio of 70%:15%:15%. Data augmentation techniques (such as translation, scaling, etc.) were applied to the training set to improve the generalization ability of the model.

[0241] Experimental Setup: The model was trained using the Adam optimizer, with a batch size of 32 and a learning rate of 0.0001. A learning rate decay strategy was used during training. The number of epochs was set to 100, and an early stopping strategy was used. Training was terminated after the model error on the validation set did not decrease after five consecutive epochs. The model was evaluated on the test set, and the accuracy of the heading prediction task was evaluated using the mean absolute error (MAE) and root mean square error (RMSE). The formulas for calculating the mean absolute error (MAE) and root mean square error (RMSE) are:

[0242]

[0243] in, is the predicted radian, θ i is the true radian, and N is the total number of samples.

[0244] The performance of the binary classification task is evaluated using accuracy, precision, recall, and F1 score on the distance judgment task. The calculation formulas for Accuracy, Precision, Recall, and F1 score are as follows:

[0245]

[0246]

[0247] Among them, TP is the number of samples correctly predicted by the model as positive, TN is the number of samples correctly predicted by the model as negative, FP is the number of samples incorrectly predicted by the model as positive, and FN is the number of samples incorrectly predicted by the model as negative.

[0248] Experimental Results: To verify the effectiveness of the proposed navigation model, ViT-NAV, we compared it with traditional convolutional neural network (CNN) methods in the tasks of heading prediction and safety distance determination. The experimental results are shown in Tables 1 and 2. The results are analyzed from the perspectives of regression metrics (RMSE, MAE) and classification metrics (Precision, Recall, F1, Accuracy).

[0249] Table 1 Experimental results of the heading prediction task

[0250]

[0251] Table 2 Experimental results of safety distance discrimination task

[0252]

[0253] In the forward direction prediction task, the ViT-NAV model proposed in the present invention is significantly better than the traditional CNN method in terms of RMSE and MAE, indicating that the model proposed in the present invention has higher accuracy and smaller error in direction prediction.

[0254] In the safety distance discrimination task, the proposed ViT-NAV model surpassed CNN in all four classification metrics. Among them, the precision reached 0.9826, the recall rate was 0.9339, the F1 value was 0.9576, and the accuracy rate reached 0.9583. In comparison, the CNN model's various metrics were significantly lower, indicating that it made more misjudgments and missed judgments when determining whether the camera was too close to the intestinal wall in complex scenarios. The high precision and robustness demonstrated by ViT-NAV demonstrate its stronger generalization capabilities in feature extraction and discrimination.

[0255] In summary, the ViT-NAV model proposed in this paper shows excellent performance in both key tasks, significantly outperforming the traditional CNN method, verifying the effectiveness of the proposed navigation model.

[0256] The present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.

Claims

1. A colonoscopy navigation method based on image features of the in-scope, characterized by: The specific process of the method is: Step 1: Acquire unlabeled colonoscopic image data, process the unlabeled colonoscopic image data, and obtain processed unlabeled colonoscopic image data; Step 2: inputting the processed unlabeled colonoscopy image data obtained in step 1 into the reconstruction network model, the reconstruction network model outputting the reconstructed colonoscopy image data, and training the reconstruction network model to obtain a trained reconstruction network model; Step 3: Obtain a training set of colonoscopy images with labels of forward direction v(sinθ, cosθ) θ is the forward direction angle; The colonoscopy image training set with the forward direction v (sinθ, cosθ) label Image samples in Input the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features Step 4: Get the safe distance y dis Labeled colonoscopy image training set Keep a safe distance dis Labeled colonoscopy image training set Image samples in Input the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features Step 5: For the forward direction prediction task: ViT encoder outputs image features Input direction prediction network, direction prediction network outputs forward direction prediction value Obtain a trained direction prediction network; Step 6: For the safety distance discrimination task: ViT encoder outputs image features Input the safety distance judgment network, and the safety distance judgment network outputs a binary judgment of whether the distance to the intestinal wall is too close Obtain a trained safety distance discrimination network; Step 7: Obtain the colonoscopy image data to be tested, divide the colonoscopy image data to be tested into blocks of size P×P, flatten each block, and output the embedded representation of each block through a linear layer after flattening. The embedded representations of all blocks are input into the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features; The ViT encoder outputs image features and inputs them into the trained direction prediction network, which then outputs the predicted value of the colonoscope's forward direction. Prediction value based on the direction of colonoscopy Calculate radian predictions The ViT encoder outputs image features and inputs them into the trained safety distance discrimination network. The trained safety distance discrimination network outputs a prediction value of whether the distance to the intestinal wall is too close.

2. The colonoscopy navigation method based on the image features of the in-scope according to claim 1, characterized in that: In the step 1, unlabeled colonoscopic image data is obtained, and the unlabeled colonoscopic image data is processed to obtain processed unlabeled colonoscopic image data; the specific process is: 1) Obtain unlabeled colonoscopic image data, perform image segmentation on the unlabeled colonoscopic image data, and obtain a colonoscopic image containing only the intestinal area; 2) Enhance the colonoscopic image containing only the intestinal area to obtain an enhanced colonoscopic image; the specific process is as follows: The colonoscopic image containing only the intestinal area was rotated 90 degrees, 180 degrees, and 270 degrees clockwise to obtain the enhanced colonoscopic image; 3) Standardize the enhanced colonoscopic image to obtain a standardized colonoscopic image; the specific process is as follows: The pixel values of the enhanced colonoscopy image are in the range of [0, 255]. The enhanced colonoscopy image is normalized and the pixel values of the image are adjusted to the range of [0, 1].

3. The colonoscopy navigation method based on the image features of the in-scope according to claim 2, characterized in that: In the second step, the processed unlabeled colonoscopy image data obtained in the first step is input into the reconstruction network model, the reconstruction network model outputs the reconstructed colonoscopy image data, and the reconstruction network model is trained to obtain a trained reconstruction network model; the specific process is: 1) The processed unlabeled colonoscopy image data X obtained in step 1 is unlabled The jth sample image x in un Divide into blocks of size P×P, blocks, and the number of channels in each block is C; in, x un ∈R H×W×C , H, W, C are respectively each sample image x un The height, width and number of channels of , R is a real number; P×P represents the height × width of each block; 2) Flatten each block; After flattening, it is mapped to the D-dimensional feature space through a linear layer as follows: z p =W p x p +b p ,p=1,2,…N in, x p is the pixel vector representation of the p-th block after flattening; W p is the weight matrix of the linear layer; b p is the bias term; z p is the embedded representation of the p-th block output by the linear layer; 3) Combine the embedding representations of N blocks into a sequence Z and add classification token z to the sequence Z cls , get the added classification token z cls The sequence Z0 is expressed as: Z=[z1,z2,…,z N ]∈R N×D Z0=[z cls ,z1,z2,…,z N ] Among them, z1 is the embedding representation of the first block, z2 is the embedding representation of the second block, and z N is the embedded representation of the Nth block; R is a real number; 4) Introduce position code E: E=[e cls ,e1,e2,…,e N ] Among them, e cls is the classification token z cls location information; e1 is the embedding of the first block, indicating the position information of z1, and e2 is the embedding of the second block, indicating the position information of z2. N is the embedding representation z of the Nth block N location information; 5) Based on adding classification token z cls The sequence Z0 and the position code E are used to get the encoder input sequence Z input : WITH input =Z0+E 6) Randomly occlude the input sequence Z input For a part of the block, define an occlusion mask: M = {m1,m2,…,m N },m i ∈{0,1} Among them, m i =1 means the i-th block is not blocked, m i =0 means the i-th block is occluded; 7) From the input sequence Z input Extract the unoccluded blocks to form a sequence Z visible : Z visible ={Z input |m i =1},i=1,2,…,N 8) Sequence Z of unoccluded blocks visible Input ViT encoder to encode the unoccluded Patch, and ViT encoder outputs the feature representation H of the unoccluded block pretrain ; 9) The feature representation H of the unoccluded block output by the ViT encoder pretrain and the position code of the occlusion block E i Passed together into the decoder, the decoder outputs the reconstructed pixel value of the occluded block 10) Reconstructed pixel value of the i-th occlusion block based on the decoder output and the original pixel value I of the i-th occlusion block i Calculate the reconstruction loss function L recon ; expressed as: Among them, N m is the total number of occluded blocks; 11) Repeat 1)-10) until the reconstruction loss function converges and obtains the trained reconstruction network model.

4. The colonoscopy navigation method based on the image features of the in-scope according to claim 3, characterized in that: In 8), the sequence Z of the unobstructed blocks is visible Input ViT encoder to encode the unoccluded Patch, and ViT encoder outputs the feature representation H of the unoccluded block pretrain ; expressed as: H pretrain =f θ (Z visible ) in, f θ For ViT encoder; Z visible is a sequence of unoccluded blocks; H pretrain Feature representation of unoccluded blocks output by the ViT encoder.

5. The colonoscopy navigation method based on the image features of the in-scope according to claim 4 is characterized by: In 9), the feature representation H of the unoccluded block output by the ViT encoder is pretrain and the position code of the occlusion block E i Passed together into the decoder, the decoder outputs the reconstructed pixel value of the occluded block Expressed as: in, g φ For ViT decoder; is the reconstructed pixel value of the i-th occlusion block output by the ViT decoder; E i Encode the position of the i-th occluded block.

6. The colonoscopy navigation method based on the image features of the in-scope according to claim 5, characterized in that: The reconstructed pixel value of the i-th occlusion block output by the decoder in 10) and the original pixel value I of the i-th occlusion block i Calculate the reconstruction loss function L recon ; expressed as: Among them, N m is the total number of occluded blocks.

7. The colonoscopy navigation method based on the image features of the in-scope according to claim 6, characterized in that: In step 3, a colonoscopy image training set with a label of forward direction v (sinθ, cosθ) is obtained. θ is the forward direction angle; The colonoscopy image training set with the forward direction v (sinθ, cosθ) label Image samples in Input the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features Expressed as: in, Represented by image Extracted embedding representations of all blocks; The acquisition process is: The image sample Divide into blocks of size P×P, flatten each block, and output the embedded representation of each block through the linear layer after flattening. The embedded representation of all blocks consists of 8. The colonoscopy navigation method based on the image features of the in-scope according to claim 7, characterized in that: In step 4, the safety distance y is obtained. dis Labeled colonoscopy image training set Keep a safe distance dis Labeled colonoscopy image training set Image samples in Input the ViT encoder in the trained reconstruction network model, and the ViT encoder outputs image features Expressed as: in, Represented by image Extracted embedding representations of all blocks; The acquisition process is: The image sample Divide into blocks of size P×P, flatten each block, and output the embedded representation of each block through the linear layer after flattening. The embedded representation of all blocks consists of 9. The colonoscopy navigation method based on the image features of the in-scope according to claim 8, characterized in that: In step 5, for the forward direction prediction task: the ViT encoder outputs image features Input direction prediction network, direction prediction network outputs forward direction prediction value Get the trained direction prediction network; the specific process is: Step 5.1: The direction prediction network includes a fully connected layer FC, a ReLU activation function layer, a Dropout layer, and a fully connected layer FC in sequence; Step 52: ViT encoder outputs image features Input direction prediction network, direction prediction network outputs the predicted value of the colonoscope's forward direction Step 53: Prediction of the colonoscope's forward direction based on the output of the direction prediction network And the true value v(sinθ,cosθ) is used to calculate the mean square error loss function; it is expressed as: in, N′ should be a colonoscopy image training set with labels of forward direction v(cosθ, sinθ) The total number of image samples in , v i′ For the training set The true value of the forward direction of the i′th image sample in ; Training set for the direction prediction network output The predicted value of the forward direction of the i′th image sample in ; Step 54: Repeat steps 52 to 53 to traverse the colonoscopy image training set with the forward direction v (sinθ, cosθ) label Image samples in the loss function Converge and obtain a trained direction prediction network.

10. The colonoscopy navigation method based on the image features of the in-scope according to claim 9, characterized in that: In step 6, for the safety distance discrimination task: the ViT encoder outputs image features Input the safety distance judgment network, and the safety distance judgment network outputs a binary judgment of whether the distance to the intestinal wall is too close Obtain a trained safety distance discrimination network; the specific process is: Step 6.1: The safety distance discrimination network includes a fully connected layer FC, a ReLU activation function layer, a Dropout layer, and a fully connected layer FC in sequence; Step 62: ViT encoder outputs image features Input the safety distance judgment network, and the safety distance judgment network outputs a binary judgment prediction value of whether the distance to the intestinal wall is too close Step 6.3: Binary classification prediction value based on the output of the safety distance judgment network to determine whether the distance to the intestinal wall is too close and the truth value y dis Calculate the cross entropy loss function; expressed as: in, N″ is the safety distance y dis Labeled colonoscopy image training set The total number of image samples in ; For safe distance dis Labeled colonoscopy image training set The true value of whether the distance between the i″th sample and the intestinal wall is too close; The training set output by the safety distance discrimination network The predicted value of whether the distance between the i″th sample and the intestinal wall is too close; Step 64: Repeat steps 62 to 63 to traverse the colonoscopy image training set with safety distance labels Image samples in the loss function Converge and obtain a trained safety distance discrimination network.