Method for simulating indigo carmine staining of upper digestive tract based on comparative learning

By comparative learning and generation adversarial networks, combined with multi-frame difference hashing and deep learning object detection, the problems of information loss and high model complexity in endoscopic image generation are solved, and clear and fast virtual indigo carmine dyeing is achieved to meet the real-time needs of endoscopic examination.

CN120495436AInactive Publication Date: 2025-08-15徐勤伟
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510411429.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the generated endoscopic images have problems such as missing or blurring information, high model complexity, long training time, and large memory consumption, and the endoscopic doctor cannot provide virtual stained images in time.

Method used

Using a method based on contrast learning, multi-frame difference hashing and deep learning object detection technology, combined with the generation of adversarial network, automatic lesion judgment of endoscopic images and simulated indigo carmine dyeing. Using simulated indigo carmine dyeing generator, characterization network and negative example generator, the generation process is optimized to improve image clarity and model efficiency.

Benefits of technology

It realizes automatic judgment of suspected lesions of endoscopic doctors, and provides clear virtual stained images. The model complexity is low, the training time is short, the memory consumption is small, and the image generation time is short, meeting the real-time needs of endoscopy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495436A_ABST
    Figure CN120495436A_ABST
Patent Text Reader

Abstract

The invention discloses a contrastive learning-based indigo carmine staining simulation method for an upper digestive tract, and relates to the technical field of medical image processing. The invention discloses an upper digestive tract simulated indigo carmine staining method based on comparative learning. The method comprises the steps of image processing, calculation of difference hash codes among continuous N frames of images, calculation of Hamming distances among adjacent N frames of images, judgment of a video motion state, recognition of lesions by using a lesion detection model, starting of a simulated indigo carmine staining model for staining, and outputting of simulated stained images. Based on the multi-frame difference hash and deep learning target detection technology, automatic judgment of suspected lesions observed by endoscopic doctors is achieved, based on the comparative learning and generative adversarial network technology, the lesions simulate indigo carmine staining effect, virtual staining images can be provided for the endoscopic doctors in time, and the efficiency of the endoscopic doctors is improved. And the method has the advantages of clear image, low model complexity, short training time, low memory consumption, short image generation time and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a method for simulating indigo carmine dyeing of the upper digestive tract based on contrast learning. Background Art

[0002] In 2022, China saw approximately 358,700 new cases of gastric cancer (accounting for 7.43% of all malignant tumors) and 260,400 deaths (accounting for 10.11%). The incidence and mortality rates were 25.41 per 100,000 and 18.44 per 100,000, respectively, ranking fifth and third among malignant tumors. Early gastric cancer is often asymptomatic, making it difficult to detect with traditional examinations such as CT and ultrasound. However, gastroscopy, which allows for direct observation of the gastric mucosa and the performance of biopsies, has gradually become the preferred method for gastric cancer screening.

[0003] To better visualize the gastric mucosa during gastroscopy, indigo carmine is often sprayed for staining. Indigo carmine is a non-absorbable dye that deposits in the grooves of the mucosal surface through physical adsorption, enhancing the three-dimensional contrast of otherwise flat or delicate mucosal structures. This allows for better identification of lesion boundaries and assists doctors in differential diagnosis.

[0004] Indigo carmine staining is a widely used adjunct to endoscopic examinations. However, since the dye must be sprayed under the endoscope, it can cause discomfort or allergic reactions in patients. Furthermore, the staining process requires repeated rinsing and observation, which prolongs the examination time and increases medical costs and time for both patients and physicians.

[0005] With the development of artificial intelligence (AI), deep learning techniques are increasingly being applied to virtual staining in digestive endoscopy. Patent CN116309947A uses a recurrent generative adversarial network to simulate indigo carmine staining of endoscopic images, while patent CN118229600A uses a diffusion model to perform virtual indigo carmine staining of ulcerative colitis-related tumors.

[0006] Currently, patent CN116309947A uses a cyclic adversarial generative network to simulate indigo carmine staining of endoscopic images. It uses cycle consistency to solve the problem of data mismatch between white light images and indigo carmine images. Cycle consistency assumes that there is a one-to-one mapping relationship between the input domain image and the output domain image. However, there is more input domain information than output domain information in white light images and indigo carmine objects, resulting in information loss or blurring in the generated images. At the same time, patent CN116309947A requires training two generators and two discriminators, resulting in high model complexity, long training time, and high memory consumption. Patent CN118229600A uses a diffusion model for virtual indigo carmine staining, which has the problem of long image generation time and inability to provide virtual staining images to endoscopists in a timely manner. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for simulating indigo carmine staining of the upper gastrointestinal tract based on contrastive learning, so as to solve the problems of information loss or blurring in the generated images, high model complexity, long training time, large memory consumption, long image generation time, and inability to provide virtual staining images to endoscopists in a timely manner.

[0008] In order to achieve the above object, the present invention provides the following technical solutions:

[0009] A method for simulating indigo carmine staining of the upper digestive tract based on comparative learning comprises the following steps:

[0010] S1. Obtain gastroscopy video, perform image scaling and grayscale processing, extract a single-frame video image, perform horizontal gradient calculation, and calculate the difference hash code between N consecutive frames of images;

[0011] S2. Calculate the Hamming distance between N adjacent frames to determine the video motion state;

[0012] S3, when the Hamming distance between the adjacent N frames of images is less than the set threshold T, the lesion detection model is used to identify the lesion;

[0013] S4. When the lesion detection model detects a suspected lesion, it starts the simulated indigo carmine dyeing model to perform dyeing and outputs a simulated dyeing image.

[0014] The indigo carmine dyeing model consists of a simulated indigo carmine dyeing generator G, a representation network, and a negative example generator;

[0015] Furthermore, the simulated indigo dyeing generator G converts the source domain image X into the target domain image Y, where the source domain image X is a white light image and the target domain image Y is an indigo dyeing image. The simulated indigo dyeing generator G is an encoder-decoder architecture, i.e., Y = G(X) = G dec (G enc (X)) The indigo carmine dyeing generator is based on the residual module generator or the stylegan2 generator structure. The indigo carmine dyeing generator uses adversarial contrast learning and generative adversarial network loss to ensure the authenticity of the generated image and the semantic consistency with the source image;

[0016] The representation network is a two-layer MLP network, denoted as H i (·), i is the index of the i-th layer of the encoder, mapping the features of each layer in the simulated indigo carmine dyeing generator to a high-dimensional representation space for subsequent contrastive learning; extracting features at different layers of the simulated indigo carmine dyeing encoder and is the source image, To generate an image, extract features and The random spatial position sampling in the input is input into the representation network H and L2 normalized to obtain the query sample q and the positive sample k + , the formula is as follows:

[0017]

[0018] The negative example generator generates negative examples based on the source image features The encoder structure in the simulated indigo carmine dyeing generator is a multi-layer MLP network, represented by N i (·), the input is the mean of the source image features And random noise z~N(0,1). Random noise is used to increase the diversity of generated negative samples. The output of the negative example generator after normalization can be formulated as follows:

[0019]

[0020] The negative example generator and the coloring encoder network are optimized alternately, and the objective function is:

[0021]

[0022] In the function, τ represents the temperature factor. Represent the network parameters of the representation network, coloring encoder, and negative example generator respectively; in the alternating optimization of the negative example generator and the coloring encoder network, the negative example generator is optimized to generate the same negative example as the positive example k + Similar difficult negative examples increase contrast loss; color encoder optimization can reduce the query sample q and positive example k + distance, while staying away from The adversarial contrast loss can be expressed as:

[0023] To prevent the negative example generator from generating the same negative examples for different noise inputs, increase the diversity loss:

[0024]

[0025] In the function, z1 and z2 are independently sampled noises. The generated adversarial loss adopts LSGAN, with G(·) and D(·) representing the coloring generator and discriminator respectively. The discriminator structure is PatchGan. The discriminator adversarial loss and the generator adversarial loss are expressed as:

[0026]

[0027] Characterizing Network Loss Coloring Generator Loss Negative Generator Loss It is expressed as follows:

[0028]

[0029] In the function, λ1 and λ2 are loss hyperparameters respectively. In the inference stage, the dye generator is used to convert the white light image into an indigo carmine dyed image.

[0030] Furthermore, image scaling refers to scaling the input image to 8×9 in width and height, thereby preserving the image structure and removing detail information.

[0031] Furthermore, grayscale processing refers to converting the scaled image into a grayscale image.

[0032] Furthermore, the horizontal gradient calculation refers to comparing the grayscale values of two adjacent pixels in each row from left to right. If the previous pixel is greater than the next pixel, it is marked as 1; otherwise, it is marked as 0.

[0033] Furthermore, the difference hash code refers to the 64-bit binary hash code obtained by concatenating the 8-bit code per row in step 3.

[0034] Furthermore, the lesion detection model consists of three parts: Backbone, Neck, and Head.

[0035] Furthermore, Backbone adopts the YOLOv10s structure, which includes Stem, Stage1, Stage2, Stage3, and Stage4 stages. Each stage corresponds to a different input size, output channel, and core module. The input sizes are 640×640×3, 320×320×64, 160×160×128, 80×80×256, and 40×40×512 respectively. The output channels are 64, 128, 256, 512, and 1024 respectively. The core modules are Conv+BN+Si LU, CBS+C2f_3, CBS+C2f_6, SCDown+C2f_6, SCDown+C2fCIB_3+SPPF+PSA, C2f_n is a combination of CBS+split+bottleneck×n+Concat+CBS, SCDown is a spatial channel decoupling downsampling operation, C2fCIB_n is a combination of CBS+split+CIB×n+Concat+CBS, CIB is a compact inverted residual block, SPPF is a fast spatial pyramid pooling, and PSA is a local self-attention module;

[0036] The Backbone network reduces computational costs by separating spatial and channel operations. It uses the CIB structure to reduce redundant computations in the deep layers of the network, optimizes gradient paths, and designs the PSA structure to introduce global modeling capabilities at a very low computational cost, thereby improving network feature extraction capabilities.

[0037] The Neck layer adopts the PAFPN structure and introduces bidirectional path fusion based on FPN, so that each layer in the feature pyramid can fuse semantic and detail information of different resolutions.

[0038] The head layer contains two detection heads, one-to-many detection head and one-to-one detection head. The one-to-one detection head is used for inference prediction. The one-to-many detection head dynamically selects positive samples based on task alignment loss, classification score and IoU. The one-to-one detection head assigns only one optimal prediction box to each real target, avoiding reliance on NMS during inference. The matching strategy uses the same matching metric as the one-to-many detection head to align their supervision. During the training process, the one-to-one detection head uses the supervisory signal of the one-to-many head to improve the performance of the one-to-one detection head.

[0039] The beneficial effects of the present invention are as follows: the present application realizes automatic judgment of suspected lesions observed by endoscopists through multi-frame difference hashing and deep learning target detection technology, and realizes the simulation of indigo carmine dyeing effect of lesions through comparative learning and generative adversarial network technology, and can provide virtual dyeing images to endoscopists in a timely manner. The virtual dyeing images have the advantages of clear virtual dyeing images, low model complexity, short training time, low memory consumption, and short image generation time.

[0040] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is an overall flow chart showing an embodiment of the present invention.

[0042] Figure 2 FIG. 4 is a flow chart of a lesion detection model according to an embodiment of the present invention.

[0043] Figure 3 This is a data diagram of the core modules of Backbone at different stages according to an embodiment of the present invention.

[0044] Figure 4 FIG. 4 is a flow chart of a simulation model for indigo carmine dyeing according to an embodiment of the present invention.

[0045] Figure 5 FIG. 1 is a schematic diagram of an original video image and a simulated indigo carmine dyeing process according to an embodiment of the present invention. DETAILED DESCRIPTION

[0046] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0047] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0048] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0049] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0050] See Figure 1 A preferred embodiment of the present application shows a method for simulating indigo carmine dyeing of the upper digestive tract based on comparative learning, comprising the following steps:

[0051] S1. Obtain gastroscopy video, perform image scaling and grayscale processing, extract a single-frame video image, perform horizontal gradient calculation, and calculate the difference hash code between N consecutive frames of images;

[0052] S2. Calculate the Hamming distance between N adjacent frames to determine the video motion state;

[0053] S3, when the Hamming distance between the adjacent N frames of images is less than the set threshold T, the lesion detection model is used to identify the lesion;

[0054] S4. When the lesion detection model detects a suspected lesion, it starts the simulated indigo carmine dyeing model to perform dyeing and outputs a simulated dyeing image.

[0055] The indigo carmine dyeing model consists of a simulated indigo carmine dyeing generator G, a representation network, and a negative example generator;

[0056] The simulated indigo dyeing generator G converts the source domain image X into the target domain image Y. The source domain image X is a white light image, and the target domain image Y is an indigo dyeing image. The simulated indigo dyeing generator G is an encoder-decoder architecture, that is, Y = G (X) = G dec (G enc (X)) The indigo carmine dyeing generator is based on the residual module generator or the stylegan2 generator structure. The indigo carmine dyeing generator uses adversarial contrast learning and generative adversarial network loss to ensure the authenticity of the generated image and the semantic consistency with the source image;

[0057] The representation network is a two-layer MLP network, denoted as H i (·), i is the index of the i-th layer of the encoder, mapping the features of each layer in the simulated indigo carmine dyeing generator to a high-dimensional representation space for subsequent contrastive learning; extracting features at different layers of the simulated indigo carmine dyeing encoder and is the source image, To generate an image, extract features and The random spatial position sampling in the input is input into the representation network H and L2 normalized to obtain the query sample q and the positive sample k + , the formula is as follows:

[0058]

[0059] The negative example generator generates negative examples based on the source image features The encoder structure in the simulated indigo carmine dyeing generator is a multi-layer MLP network, represented by N i (·), the input is the mean of the source image features And random noise z~N(0,1). Random noise is used to increase the diversity of generated negative samples. The output of the negative example generator after normalization can be formulated as follows:

[0060]

[0061] The negative example generator and the coloring encoder network are optimized alternately, and the objective function is:

[0062]

[0063] In the function, τ represents the temperature factor. Represent the network parameters of the representation network, coloring encoder, and negative example generator respectively; in the alternating optimization of the negative example generator and the coloring encoder network, the negative example generator is optimized to generate the same negative example as the positive example k+ Similar difficult negative examples increase contrast loss; color encoder optimization can reduce the query sample q and positive example k + distance, while staying away from The adversarial contrast loss can be expressed as:

[0064]

[0065] To prevent the negative example generator from generating the same negative examples for different noise inputs, increase the diversity loss:

[0066]

[0067] In the function, z1 and z2 are independently sampled noises. The generated adversarial loss adopts LSGAN, with G(·) and D(·) representing the coloring generator and discriminator respectively. The discriminator structure is PatchGan. The discriminator adversarial loss and the generator adversarial loss are expressed as:

[0068]

[0069] Characterizing Network Loss Coloring Generator Loss Negative Generator Loss It is expressed as follows:

[0070]

[0071] In the function, λ1 and λ2 are loss hyperparameters respectively. In the inference stage, the dye generator is used to convert the white light image into an indigo carmine dyed image.

[0072] Image scaling refers to scaling the input image to 8×9 in width and height, thereby preserving the image structure and removing detail information.

[0073] Grayscale processing refers to converting the scaled image into a grayscale image.

[0074] Horizontal gradient calculation refers to comparing the grayscale values of two adjacent pixels in each row from left to right. If the previous pixel is greater than the next pixel, it is marked as 1; otherwise, it is marked as 0.

[0075] The differential hash code refers to the 64-bit binary hash code obtained by concatenating the 8-bit code per row in step 3.

[0076] The lesion detection model consists of three parts: Backbone, Neck, and Head.

[0077] Backbone adopts YOLOv10s structure, which includes Stem, Stage1, Stage2, Stage3, and Stage4 stages. Each stage corresponds to different input sizes, output channels, and core modules. The input sizes are 640×640×3, 320×320×64, 160×160×128, 80×80×256, and 40×40×512 respectively. The output channels are 64, 128, 256, 512, and 1024 respectively. The core modules are Conv+BN+SiLU, CBS+C2f_3, CBS+C2f_6, SCDown+C2f_6, SCDown+C2fCIB_3+SPPF+PSA, C2f_n is a combination of CBS+split+bottleneck×n+Concat+CBS, SCDown is a spatial channel decoupling downsampling operation, C2fCIB_n is a combination of CBS+split+CIB×n+Concat+CBS, CIB is a compact inverted residual block, SPPF is a fast spatial pyramid pooling, and PSA is a local self-attention module;

[0078] The Backbone network reduces computational costs by separating spatial and channel operations. It uses the CIB structure to reduce redundant computations in the deep layers of the network, optimizes gradient paths, and designs the PSA structure to introduce global modeling capabilities at a very low computational cost, thereby improving network feature extraction capabilities.

[0079] The Neck layer adopts the PAFPN structure and introduces bidirectional path fusion based on FPN, so that each layer in the feature pyramid can fuse semantic and detail information of different resolutions.

[0080] The head layer contains two detection heads, one-to-many detection head and one-to-one detection head. The one-to-one detection head is used for inference prediction. The one-to-many detection head dynamically selects positive samples based on task alignment loss, classification score and IoU. The one-to-one detection head assigns only one optimal prediction box to each real target, avoiding reliance on NMS during inference. The matching strategy uses the same matching metric as the one-to-many detection head to align their supervision. During the training process, the one-to-one detection head uses the supervisory signal of the one-to-many head to improve the performance of the one-to-one detection head.

[0081] The calculation of the Hamming distance is to compare the 64-bit hash codes of two images and count the number of different bits, which is the Hamming distance; the smaller the Hamming distance between two images, the more similar the two images are, and the less shaking the picture in the video.

[0082] To automatically assist doctors with virtual indigo carmine staining and reduce unnecessary manipulation, virtual indigo carmine staining is performed only when the doctor carefully observes the gastric mucosa and a suspected lesion is identified. This is achieved by: when the Hamming distance between any two of N adjacent image frames is less than a threshold T (T is 3 in this solution), indicating that the N adjacent frames are relatively still or slightly shaky, the endoscopist is presumed to be observing the gastric mucosa or freezing the image, and the lesion detection model is activated.

[0083] In summary, the present invention provides a method for simulating indigo carmine staining of the upper gastrointestinal tract based on contrastive learning. The method realizes automatic judgment of suspected lesions observed by endoscopists through multi-frame difference hashing and deep learning target detection technology, and realizes the effect of simulating indigo carmine staining of lesions through contrastive learning and generative adversarial network technology. The method can provide virtual staining images to endoscopists in a timely manner, and has the advantages of clear virtual staining images, low model complexity, short training time, low memory consumption, and short image generation time.

[0084] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A method for simulating indigo carmine dyeing of the upper digestive tract based on comparative learning, characterized in that: The steps include: S1. Obtain gastroscopy video, perform image scaling and grayscale processing, extract a single-frame video image, perform horizontal gradient calculation, and calculate the difference hash code between N consecutive frames of images; S2. Calculate the Hamming distance between N adjacent frames to determine the video motion state; S3, when the Hamming distance between the adjacent N frames of images is less than the set threshold T, the lesion detection model is used to identify the lesion; S4. When the lesion detection model detects a suspected lesion, it starts the simulated indigo carmine dyeing model to perform dyeing and outputs a simulated dyeing image.

2. The method for simulating indigo carmine dyeing of the upper digestive tract based on contrast learning according to claim 1, wherein: The indigo carmine dyeing model consists of a simulated indigo carmine dyeing generator G, a representation network, and a negative example generator; The simulated indigo dyeing generator G converts the source domain image X into the target domain image Y. The source domain image X is a white light image, and the target domain image Y is an indigo dyeing image. The simulated indigo dyeing generator G is an encoder-decoder architecture, that is, Y = G (X) = G dec (G enc (X)) The indigo carmine dyeing generator is based on the residual module generator or the stylegan2 generator structure. The indigo carmine dyeing generator uses adversarial contrast learning and generative adversarial network loss to ensure the authenticity of the generated image and the semantic consistency with the source image; The representation network is a two-layer MLP network, denoted as H i (·), i is the index of the i-th layer of the encoder, mapping the features of each layer in the simulated indigo carmine dyeing generator to a high-dimensional representation space for subsequent contrastive learning; extracting features at different layers of the simulated indigo carmine dyeing encoder and is the source image, To generate an image, extract features and The random spatial position sampling in the input is input into the representation network H and L2 normalized to obtain the query sample q and the positive sample k + , the formula is as follows: The negative example generator generates negative examples based on the source image features The encoder structure in the simulated indigo carmine dyeing generator is a multi-layer MLP network, represented by N i (·), the input is the mean of the source image features And random noise z~N(0,1). Random noise is used to increase the diversity of generated negative samples. The output of the negative example generator after normalization can be formulated as follows: The negative example generator and the coloring encoder network are optimized alternately, and the objective function is: In the function, τ represents the temperature factor. θ G , Represent the network parameters of the representation network, coloring encoder, and negative example generator respectively; in the alternating optimization of the negative example generator and the coloring encoder network, the negative example generator is optimized to generate the same negative example as the positive example k + Similar difficult negative examples increase contrast loss; color encoder optimization can reduce the query sample q and positive example k + distance, while staying away from The adversarial contrast loss can be expressed as: To prevent the negative example generator from generating the same negative examples for different noise inputs, increase the diversity loss: In the function, z1 and z2 are independently sampled noises. The generated adversarial loss adopts LSGAN, with G(·) and D(·) representing the coloring generator and discriminator respectively. The discriminator structure is PatchGan. The discriminator adversarial loss and the generator adversarial loss are expressed as: Characterizing Network Loss Coloring Generator Loss Negative Generator Loss It is expressed as follows: In the function, λ1 and λ2 are loss hyperparameters respectively. In the inference stage, the dye generator is used to convert the white light image into an indigo carmine dyed image.

3. The method for simulating indigo carmine dyeing of the upper digestive tract based on contrast learning according to claim 1, wherein: Image scaling refers to scaling the input image to 8×9 in width and height, thereby preserving the image structure and removing detail information.

4. The method for simulating indigo carmine dyeing of the upper digestive tract based on contrast learning according to claim 1, wherein: Grayscale processing refers to converting the scaled image into a grayscale image.

5. The method for dyeing the upper digestive tract with simulated indigo carmine based on contrast learning according to claim 1, wherein: Horizontal gradient calculation refers to comparing the grayscale values of two adjacent pixels in each row from left to right. If the previous pixel is greater than the next pixel, it is marked as 1; otherwise, it is marked as 0.

6. The method for simulating indigo carmine dyeing of the upper digestive tract based on contrast learning according to claim 1, wherein: The differential hash code refers to the 64-bit binary hash code obtained by concatenating the 8-bit code per row in step 3.

7. The method for simulating indigo carmine dyeing of the upper digestive tract based on contrast learning according to claim 1, wherein: The lesion detection model consists of three parts: Backbone, Neck, and Head.

8. The method for simulating indigo carmine dyeing of the upper digestive tract based on contrast learning according to claim 7, wherein: Backbone adopts YOLOv10s structure, which includes Stem, Stage1, Stage2, Stage3, and Stage4 stages. Each stage corresponds to different input sizes, output channels, and core modules. The input sizes are 640×640×3, 320×320×64, 160×160×128, 80×80×256, and 40×40×512 respectively. The output channels are 64, 128, 256, 512, and 1024 respectively. The core modules are Conv+BN+SiLU, CBS+C2f_3, CBS+C2f_6, SCDown+C2f_6, SCDown+C2fCIB_3+SPPF+PSA, C2f_n is a combination of CBS+split+bottleneck×n+Concat+CBS, SCDown is a spatial channel decoupling downsampling operation, C2fCIB_n is a combination of CBS+split+CIB×n+Concat+CBS, CIB is a compact inverted residual block, SPPF is a fast spatial pyramid pooling, and PSA is a local self-attention module; The Backbone network reduces computational costs by separating spatial and channel operations. It uses the CIB structure to reduce redundant computations in the deep layers of the network, optimizes gradient paths, and designs the PSA structure to introduce global modeling capabilities at a very low computational cost, thereby improving network feature extraction capabilities. The Neck layer adopts the PAFPN structure and introduces bidirectional path fusion based on FPN, so that each layer in the feature pyramid can fuse semantic and detail information of different resolutions. The head layer contains two detection heads, one-to-many detection head and one-to-one detection head. The one-to-one detection head is used for inference prediction. The one-to-many detection head dynamically selects positive samples based on task alignment loss, classification score and IoU. The one-to-one detection head assigns only one optimal prediction box to each real target, avoiding reliance on NMS during inference. The matching strategy uses the same matching metric as the one-to-many detection head to align their supervision. During the training process, the one-to-one detection head uses the supervisory signal of the one-to-many head to improve the performance of the one-to-one detection head.

Citation Information

Patent Citations

  • Method for simulating indigo carmine dyeing based on cyclic generative adversarial network

    CN116309947A

  • Virtual staining method and device for ulcerative colitis related tumors under enteroscope

    CN118229600A