Colorectal lesion auxiliary diagnosis system based on AI and cytoendoscopy

By using AI technology to adjust brightness, denoising and super-resolution reconstruction in the large intestinal endoscopic image processing system, the problem of insufficient image resolution in the prior art is solved, and more accurate lesion area recognition and more efficient diagnosis are achieved.

CN119494838BActive Publication Date: 2025-05-09JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510080867.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-09
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The prior art is difficult to effectively improve the resolution of the endoscopic images of the colon, resulting in limited diagnostic accuracy and reliability.

Method used

Using AI and endoscopy-based large intestinal lesions assisted diagnosis system, the resolution of the image is improved through image brightness adjustment, denoising and super-resolution reconstruction, and the large intestinal state features are extracted from the resolution-optimized images to determine the lesion area.

Benefits of technology

The quality of the endoscopic images of the colon intestine is significantly improved, making the identification of the lesion area more accurate, and improving the accuracy and efficiency of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119494838B_ABST
    Figure CN119494838B_ABST
Patent Text Reader

Abstract

The present application relates to the field of auxiliary diagnosis, and specifically discloses a colorectal lesion auxiliary diagnosis system based on AI and cytoscopy, which first receives images collected by a cytoscopy, then optimizes the brightness of these images, and then denoises the brightness-optimized images to obtain clearer denoised optimized images. Furthermore, the denoised images are super-resolution reconstructed to improve the resolution of the images, and finally, the characteristics of the colorectal state are extracted from the resolution-optimized images, and the lesion area is determined. In this way, the quality of colon endoscopy images can be significantly improved, making the identification of lesion areas more accurate, thereby improving the accuracy and efficiency of diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of auxiliary diagnosis, and more specifically, to an auxiliary diagnosis system for colorectal lesions based on AI and cytoendoscopy. Background Art

[0002] Colorectal lesions refer to various pathological changes occurring in the large intestine (including the cecum, ascending colon, transverse colon, descending colon, sigmoid colon and rectum), which can be benign or malignant.

[0003] With the continuous advancement of medical imaging technology, endoscopic imaging has been able to provide extremely detailed images of the intestine. However, the quality of these images may be affected by many factors, such as the resolution limitation of the imaging device, artifacts introduced during sample preparation, suboptimal lighting conditions, and motion blur. Therefore, improving image resolution is crucial to ensure the accuracy and reliability of diagnosis.

[0004] Traditional resolution enhancement techniques mainly rely on interpolation methods, which attempt to increase the resolution of an image by inserting new pixels between known pixels. However, such interpolation methods often fail to accurately restore the high-frequency information lost in the original image, resulting in a lack of realism and details in the enlarged image. In addition, most traditional methods are general and difficult to optimize for specific types of images or application scenarios, so they perform poorly when faced with diverse and complex real-world images.

[0005] Therefore, an optimized auxiliary diagnosis scheme for colorectal lesions is desired. Summary of the invention

[0006] In order to solve the above technical problems, this application is proposed.

[0007] According to one aspect of the present application, a colorectal lesion auxiliary diagnosis system based on AI and cytoendoscopy is provided, which includes: a colon endoscopy image acquisition module for receiving a colon endoscopy image acquired by a cytoendoscopy; an endoscopic image brightness adjustment module for adjusting the brightness of the colon endoscopy image to obtain a brightness-optimized colon endoscopy image; a colon endoscopy image denoising module for performing image denoising on the brightness-optimized colon endoscopy image to obtain a denoised-optimized colon endoscopy image; a colon endoscopy image resolution optimization module for performing super-resolution reconstruction on the denoised-optimized colon endoscopy image to obtain a resolution-optimized colon endoscopy image; a colon state feature extraction module for extracting colon state image features from the resolution-optimized colon endoscopy image; a lesion area determination module for determining whether there is a lesion area based on the colon state image features;

[0008] Among them, the colonoscopy image resolution optimization module includes: an initial multi-scale feature extraction unit, used to perform multi-scale feature extraction on the denoised and optimized colonoscopy image to obtain initial colonoscopy multi-scale contextual semantic coding features; a preliminary image reconstruction multi-scale feature extraction unit, used to perform preliminary image reconstruction and multi-scale feature extraction on the initial colonoscopy multi-scale contextual semantic coding features to obtain generated colonoscopy multi-scale contextual semantic coding features; a feature combination unit, used to perform a priori guided feature significant combination on the generated colonoscopy multi-scale contextual semantic coding features and the initial colonoscopy multi-scale contextual semantic coding features to obtain colonoscopy multi-step contextual semantic coding features; an image optimization unit, used to reconstruct the colonoscopy multi-step contextual semantic coding features to obtain the resolution optimized colonoscopy image.

[0009] Furthermore, the initial multi-scale feature extraction unit is used to: input the denoised and optimized colon endoscopy image into an image multi-scale feature extractor based on a multi-frequency state space module to obtain an initial colon endoscopy multi-scale contextual semantic coding feature map as the initial colon endoscopy multi-scale contextual semantic coding feature.

[0010] Furthermore, the preliminary image reconstruction multi-scale feature extraction unit comprises:

[0011] A preliminary image reconstruction subunit, used for inputting the initial colon endoscopy multi-scale context semantic encoding feature map into a preliminary reconstruction module based on a generative adversarial network to obtain a preliminary resolution optimized colon endoscopy image;

[0012] The reconstructed image multi-scale extraction subunit is used to input the preliminary resolution optimized colon endoscopy image into the image multi-scale feature extractor based on the multi-frequency state space module to obtain a generated colon endoscopy multi-scale contextual semantic coding feature map as the generated colon endoscopy multi-scale contextual semantic coding feature.

[0013] Furthermore, the feature combining unit comprises:

[0014] A generated feature initial feature reshaping modulation subunit is used to perform feature shape reshaping and prior modulation on the generated colon endoscopy multi-scale context semantic coding feature map and the initial colon endoscopy multi-scale context semantic coding feature map to obtain a priori modulated generated colon endoscopy multi-scale context semantic feature shape reshaping matrix and a priori modulated initial colon endoscopy multi-scale context semantic feature shape reshaping matrix;

[0015] A generated feature initial feature interaction subunit is used to perform attention interaction and shape reshaping on the prior modulated generated colon endoscopy multi-scale contextual semantic feature shape reshaping matrix and the prior modulated initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix to obtain a generated colon endoscopy multi-scale contextual semantic feature interaction encoding feature map and an initial colon endoscopy multi-scale contextual semantic feature interaction encoding feature map;

[0016] A generated feature initial feature fusion subunit is used to fuse the generated colonoscopy multi-scale contextual semantic feature interactive coding feature map and the initial colonoscopy multi-scale contextual semantic feature interactive coding feature map to obtain a colonoscopy multi-step contextual semantic coding feature map as the colonoscopy multi-step contextual semantic coding feature.

[0017] Furthermore, the generating feature initial feature reshaping modulation subunit is used to:

[0018] Performing feature shape reshaping on the generated colon endoscopy multi-scale context semantic encoding feature map and the initial colon endoscopy multi-scale context semantic encoding feature map to obtain a generated colon endoscopy multi-scale context semantic feature shape reshaping matrix and an initial colon endoscopy multi-scale context semantic feature shape reshaping matrix;

[0019] The generated colon endoscopy multi-scale contextual semantic feature shape reshaping matrix and the initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix are input into a priori guided modulation module to obtain the priori modulated generated colon endoscopy multi-scale contextual semantic feature shape reshaping matrix and the priori modulated initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix.

[0020] Furthermore, the generating feature initial feature interaction subunit is used to:

[0021] Input the prior modulated colon endoscopy multi-scale contextual semantic feature shape reshaping matrix and the prior modulated initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix into an attention interaction coding module based on a transformer-like structure to obtain a generated colon endoscopy multi-scale contextual semantic feature interaction coding feature matrix and an initial colon endoscopy multi-scale contextual semantic feature interaction coding feature matrix;

[0022] The generated colonoscopy multi-scale contextual semantic feature interactive encoding feature matrix and the initial colonoscopy multi-scale contextual semantic feature interactive encoding feature matrix are feature reshaped to obtain the generated colonoscopy multi-scale contextual semantic feature interactive encoding feature map and the initial colonoscopy multi-scale contextual semantic feature interactive encoding feature map.

[0023] Furthermore, the generated feature initial feature fusion subunit is used to: calculate the weighted sum of the generated colonoscopy multi-scale contextual semantic feature interactive encoding feature map and the initial colonoscopy multi-scale contextual semantic feature interactive encoding feature map to obtain the colonoscopy multi-step contextual semantic encoding feature map.

[0024] Furthermore, the image optimization unit is used to: input the colon endoscopy multi-step context semantic encoding feature map into a re-reconstruction module based on a generative adversarial network to obtain the resolution-optimized colon endoscopy image.

[0025] Furthermore, the lesion area determination module is used to: use a lesion area discriminator based on a classifier to process the large intestine state image features to obtain a discrimination result indicating whether a lesion area exists.

[0026] Compared with the prior art, the present application provides a colorectal lesion auxiliary diagnosis system based on AI and cell endoscopy, which first receives images collected by cell endoscopy, then optimizes the brightness of these images, and then denoises the brightness-optimized images to obtain clearer denoised optimized images. Further, the denoised images are super-resolution reconstructed to improve the resolution of the images, and finally, the characteristics of the colorectal state are extracted from the resolution-optimized images, and the lesion area is determined. In this way, the quality of colon endoscopy images can be significantly improved, making the identification of lesion areas more accurate, thereby improving the accuracy and efficiency of diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0028] Figure 1 It is a block diagram of a colorectal lesion auxiliary diagnosis system based on AI and cellular endoscopy according to an embodiment of the present application.

[0029] Figure 2 It is a block diagram of a colon endoscopy image resolution optimization module in a colon lesion auxiliary diagnosis system based on AI and cellular endoscopy according to an embodiment of the present application.

[0030] Figure 3 This is a block diagram of a multi-scale feature extraction unit for preliminary image reconstruction in a colorectal lesion auxiliary diagnosis system based on AI and cellular endoscopy according to an embodiment of the present application.

[0031] Figure 4It is a block diagram of the feature combination unit in the colorectal lesion auxiliary diagnosis system based on AI and cellular endoscopy according to an embodiment of the present application. DETAILED DESCRIPTION

[0032] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, but rather these embodiments are provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0033] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0034] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0035] It is worth noting that in this application, all actions to obtain signals, information or data are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.

[0036] In response to the problems in the above-mentioned background technology, the present application proposes a colorectal lesion auxiliary diagnosis system based on AI and cellular endoscopy, which effectively restores image details through the super-resolution reconstruction method of AI technology and can be flexibly applied at different scales, thus providing a more intelligent and efficient solution. Figure 1 FIG. 1 is a block diagram of a colorectal lesion auxiliary diagnosis system based on AI and cell endoscopy according to an embodiment of the present application. Figure 1As shown, the colorectal lesion auxiliary diagnosis system 100 based on AI and cellular endoscopy specifically includes: a colon endoscopy image acquisition module 110, used to receive the colon endoscopy image collected by the cellular endoscopy; an endoscopic image brightness adjustment module 120, used to adjust the brightness of the colon endoscopy image to obtain a brightness-optimized colon endoscopy image; a colon endoscopy image denoising module 130, used to perform image denoising on the brightness-optimized colon endoscopy image to obtain a denoised-optimized colon endoscopy image; a colon endoscopy image resolution optimization module 140, used to perform super-resolution reconstruction on the denoised-optimized colon endoscopy image to obtain a resolution-optimized colon endoscopy image; a colon state feature extraction module 150, used to extract colon state image features from the resolution-optimized colon endoscopy image; a lesion area determination module 160, used to determine whether there is a lesion area through the colon state image features.

[0037] Specifically, the colon endoscopy image acquisition module 110 is used to receive the colon endoscopy image collected by the cell endoscope. It should be understood that the colon endoscopy image is a view of the inside of the colon obtained by a medical device called an endoscope (or endoscope). This image can provide detailed visual information of the surface of the colon mucosa, including its color, texture, structure and other characteristics. Therefore, in order to discover possible lesions, such as inflammation, polyps, tumors, etc., so as to discover and treat them as early as possible, in the technical solution of the present application, the colon endoscopy image collected by the cell endoscope is received.

[0038] Endocytoscopy (EC) is an endoscopic technique with ultra-high resolution and magnification capabilities. It can directly observe the cells and nuclei of the gastrointestinal mucosa in vivo and achieve cellular-level observation, which plays an irreplaceable role in the screening of gastrointestinal intraepithelial neoplasia, the diagnosis of inflammatory bowel disease, and the identification of early cancer. When doctors use endoscopic examinations, they insert a thin and flexible tube into the patient's intestines. The front end of the tube is equipped with a miniature camera and other sensors that can capture detailed visual information of the surface of the large intestinal mucosa, including its color, texture, and structure. These images are not only crucial for discovering possible lesions such as polyps, inflammation, ulcers, tumors, etc., but also help to detect and treat potential problems as early as possible.

[0039] In actual operation, in order to ensure the quality of the image, multiple factors must be considered. First of all, the design and performance of the cytoscopy equipment itself is one of the key factors that determine the image quality. Modern cytoscopy is usually equipped with high-definition cameras, high-brightness light sources and advanced optical lenses to ensure that clear and color-realistic images can be obtained. In addition, some high-end models may also include functions such as autofocus and white balance adjustment to further optimize the imaging effect.

[0040] Once the cytoendoscopy successfully collects image data from the large intestine, the next step is to transmit it to the auxiliary diagnosis system. This stage mainly relies on digital communication technology and computer networks. Generally speaking, the endoscope device is connected to a dedicated computer terminal via a USB interface or wireless connection, which runs a specially designed software program to receive the data stream from the endoscope and convert it into a digital image file in a standard format for storage.

[0041] After receiving the image, the auxiliary diagnosis system will not start processing immediately, but will first conduct a preliminary assessment to determine whether it meets the requirements for the next step of processing. If there are obvious defects in the image, such as severe blur or insufficient field of view, a new image needs to be collected again. Only after confirming that the image quality is qualified will it officially enter the system's image processing process.

[0042] Specifically, the endoscopic image brightness adjustment module 120 is used to adjust the brightness of the colonoscopic image to obtain a brightness-optimized colonoscopic image. Accordingly, considering that the colonoscopic image contains vascular patterns, mucosal surface structures, and pathological characteristics, such as polyps, inflammation, ulcers, tumors, etc., these important pathological information are the key basis for diagnosis. Therefore, in order to improve the quality of the image and enable doctors to see the various features mentioned above more clearly, in the technical solution of the present application, the brightness of the colonoscopic image is adjusted to obtain a brightness-optimized colonoscopic image. In this way, the contrast in the image can be enhanced, and the differences between different tissue levels can be made more obvious, thereby making it easier to detect tiny lesions.

[0043] When the cytoscopy collects the original images, these images may have uneven brightness due to a variety of factors. For example, during the actual inspection process, due to the complex and changeable internal environment of the intestine, the lighting conditions are often difficult to fully control, resulting in some areas being too bright or too dark; in addition, the physiological differences between different patients may also affect the overall brightness distribution of the image. Therefore, in order to make the image more suitable for medical analysis, it is necessary to adjust its brightness appropriately. This process is not just a simple increase or decrease in the overall brightness value, but requires an intelligent method to optimize the brightness level of each pixel according to the specific situation of the image, ensuring that the final image not only has good visual effects, but also retains enough detail information for subsequent processing.

[0044] At this stage, the system will first conduct a comprehensive evaluation of the received colonoscopy images, identify key areas that may affect the diagnosis, and analyze their brightness characteristics. In this way, it can be determined which parts need special attention to avoid losing important pathological information during the adjustment process. Next, the image brightness is optimized using advanced image processing algorithms and techniques, such as adaptive histogram equalization and local contrast enhancement. These technologies can effectively improve the visibility of local areas while maintaining the global consistency of the image, especially in those areas that were originally darker or had lower contrast, making subtle structures more obvious.

[0045] Specifically, adaptive histogram equalization is a commonly used technique that divides the image into several small blocks and then applies a histogram equalization operation to each small block, which can significantly improve the brightness and contrast of a specific location without affecting other areas. This method is very suitable for dealing with local brightness unevenness that may occur in colon endoscopy images because it can provide personalized brightness adjustment solutions for different anatomical structures. In addition, local contrast enhancement technology focuses on enhancing the contrast within a small range in the image, helping to highlight those details that are slightly different from the surrounding background but are of great significance, such as tiny polyps or early cancer lesions.

[0046] It is worth noting that in the process of implementing brightness adjustment, some additional factors need to be considered to ensure that the final image meets both medical standards and the actual needs of doctors. On the one hand, it is necessary to ensure that the adjusted image does not introduce new artifacts or distortion, so as not to mislead doctors to make wrong judgments; on the other hand, it is necessary to try to maintain the authenticity of the original image so that doctors can work in a familiar visual environment. At the same time, considering that there may be equipment differences between different hospitals, the system should also have a certain degree of flexibility and be able to automatically adjust the processing method according to the characteristics of different types of endoscopic images to ensure the best results.

[0047] Specifically, the colonoscopic image denoising module 130 is used to perform image denoising on the brightness-optimized colonoscopic image to obtain a denoised optimized colonoscopic image. It should be understood that although the representation of key features is enhanced in the brightness-optimized colonoscopic image, it may still contain some noise or artifacts, such as random noise, reflected shadows, etc., which are interference factors inevitably introduced during the acquisition process. Based on this, in the technical solution of the present application, the brightness-optimized colonoscopic image is subjected to image denoising to further improve the image quality, remove or reduce the above-mentioned noise and artifacts, thereby ensuring that the image reflects the true anatomical structure and pathological characteristics as accurately as possible, and obtaining a denoised optimized colonoscopic image.

[0048] When the brightness-optimized colon endoscopy image denoising module after brightness adjustment enters the denoising module, the first thing to consider is the various types of noise in the image and its source. In the actual medical examination process, due to the complex and changeable internal environment of the intestine, the lighting conditions are difficult to fully control, the limitations of the acquisition equipment itself, and the patient's body movement and other factors, the obtained images often inevitably contain some random noise, reflection shadows and other interference factors. These noises not only affect the overall visual effect of the image, but more importantly, they may cover up important medical information, resulting in an increased risk of misdiagnosis or missed diagnosis. Therefore, effective measures must be taken to eliminate these adverse effects.

[0049] First, the system will conduct a comprehensive analysis of the input brightness-optimized image, identify key areas that may be contaminated by noise, and assess their severity. This step is crucial because it provides direction and basis for the subsequent specific processing. For example, if some areas are already clear enough, they do not need to be over-processed; while other areas that are blurry or have low contrast require special attention. In this way, local problems can be solved in a targeted manner while ensuring the overall quality of the image.

[0050] Next, according to the pre-set criteria and parameters, a suitable denoising algorithm is selected and applied to the image. Traditional methods such as mean filtering, Gaussian filtering, and bilateral filtering are simple and easy to use, but they have limited performance when dealing with complex scenes. With the development of computer vision and machine learning, convolutional neural networks and other deep learning frameworks can be used to perform more accurate denoising operations.

[0051] In the specific implementation process, the denoising module receives the image from the brightness optimization module and inputs it into the preprocessing pipeline. This pipeline first performs preliminary filtering on the image to remove obvious noise points while retaining important structural information. Then, according to the specific situation of the image, the most appropriate denoising algorithm is dynamically selected or multiple algorithms are combined to work together. Throughout the process, the system monitors the changes in the image in real time and dynamically adjusts the parameters through the feedback mechanism to ensure that the final output image quality is always in the best state.

[0052] Specifically, the colonoscopic image resolution optimization module 140 is used to perform super-resolution reconstruction on the denoised and optimized colonoscopic image to obtain a resolution-optimized colonoscopic image. It should be understood that the image clarity may be reduced due to a variety of factors, including the resolution limit of the imaging device, artifacts in sample preparation, inappropriate lighting conditions, and motion blur. Therefore, enhancing the resolution of the image is critical to ensuring the accuracy and stability of the diagnosis. Accordingly, the technical concept of the present application is to use an image processing and optimization algorithm based on AI and machine learning to perform the initial multi-scale context semantic encoding of the denoised and optimized colonoscopic image, and then perform preliminary image reconstruction and multi-scale context semantic encoding again on the initial colonoscopic multi-scale context semantic encoding features after the initial encoding, so as to intelligently generate the resolution-optimized colonoscopic image based on the prior-guided fusion representation between the colonoscopic multi-scale context semantic encoding features generated after the re-encoding and the initial colonoscopic multi-scale context semantic encoding features. In this way, structural information at different levels can be captured, and important details can be retained while effectively removing noise. Moreover, the use of prior information can ensure that the generated high-resolution images are not only more realistic visually, but also can adapt to the special requirements of different fields, making the feature expression more accurate, which is helpful for subsequent colorectal status analysis and lesion area determination.

[0053] Figure 2 : is a block diagram of a colon endoscopy image resolution optimization module in a colon lesion auxiliary diagnosis system based on AI and cell endoscopy according to an embodiment of the present application. Specifically, Figure 2 As shown, the colonoscopy image resolution optimization module 140 includes: an initial multi-scale feature extraction unit 141, used to perform multi-scale feature extraction on the denoised and optimized colonoscopy image to obtain initial colonoscopy multi-scale contextual semantic coding features; a preliminary image reconstruction multi-scale feature extraction unit 142, used to perform preliminary image reconstruction and multi-scale feature extraction on the initial colonoscopy multi-scale contextual semantic coding features to obtain generated colonoscopy multi-scale contextual semantic coding features; a feature combination unit 143, used to perform a priori guided feature significant combination on the generated colonoscopy multi-scale contextual semantic coding features and the initial colonoscopy multi-scale contextual semantic coding features to obtain colonoscopy multi-step contextual semantic coding features; an image optimization unit 144, used to reconstruct the colonoscopy multi-step contextual semantic coding features to obtain the resolution optimized colonoscopy image.

[0054] In an embodiment of the present application, the initial multi-scale feature extraction unit 141 is used to perform multi-scale feature extraction on the denoised and optimized colon endoscopy image to obtain an initial colon endoscopy multi-scale contextual semantic coding feature. Specifically, in an embodiment of the present application, the initial multi-scale feature extraction unit 141 is used to: input the denoised and optimized colon endoscopy image into an image multi-scale feature extractor based on a multi-frequency state space (Multi-frequency State Space Model, MFSSM) module to obtain an initial colon endoscopy multi-scale contextual semantic coding feature map as the initial colon endoscopy multi-scale contextual semantic coding feature. It should be understood that, considering that the denoised and optimized colon endoscopy image contains information of different scales, from the macroscopic mucosal surface structure to the microscopic cellular level details, there is a contextual association between these different scales of information in space. Therefore, in the technical solution of the present application, the denoised and optimized colon endoscopy image is input into an image multi-scale feature extractor based on a multi-frequency state space module to capture and mine the associations across different spatial scales, and obtain an initial colon endoscopy multi-scale contextual semantic coding feature map. It is understandable that as an advanced image feature encoding module, the core advantage of the multi-frequency state space module is that it can effectively decompose the image at different frequency scales by using the multi-scale analysis capability of frequency domain transformation, and efficiently capture the global features in the image by using the state space model. This capability enables the multi-frequency state space model to perform well in processing complex medical imaging tasks, especially when high precision and multi-scale information are required. Specifically, the multi-frequency state space module can capture high and low frequency information in the image at different frequency domain scales by introducing frequency domain transformation filtering. This means that it can simultaneously focus on the low-frequency global structural information and high-frequency texture details in the image, thereby better understanding the image content. In addition, the multi-frequency state space module effectively reduces the computational complexity and enhances the processing capability of local information by combining the selective scanning mechanism. For example, in colon endoscopy images, it can efficiently identify the subtle differences between polyps and other tissues, even if these differences span different resolution levels, thereby enhancing the understanding of complex image content.

[0055] In particular, in a specific embodiment of the present application, the specific processing process of inputting the denoised and optimized colon endoscopy image into the image multi-scale feature extractor based on the multi-frequency state space module to obtain the initial colon endoscopy multi-scale contextual semantic encoding feature map is as follows;

[0056] In order to achieve multi-scale feature extraction, the multi-frequency state space module performs a series of frequency domain transformations on the denoised and optimized colon endoscopy images to obtain a variety of different frequency domain components from low to high. Each feature map in the frequency domain contains different levels of detail information, from the coarsest overall outline to the finest local texture. To achieve this, the multi-frequency state space module usually uses Haar wavelets to perform frequency domain transformation. Wavelet transform (WT) is an efficient transformation analysis method. It inherits and develops the idea of ​​localization of short-time Fourier transform, while overcoming the shortcomings of window size not changing with frequency. It can provide a "time-frequency" window that changes with frequency, and is an ideal tool for image analysis and processing. In the multi-frequency state space module, the role of wavelet transform is particularly important, because through transformation, the characteristics of certain aspects of the problem can be fully highlighted, and the localization analysis of time (space) frequency can be performed. Through the scaling and translation operation, the image is gradually refined at multiple scales, and finally the time subdivision at high frequency and the frequency subdivision at low frequency are achieved. It can automatically adapt to the requirements of image analysis, so that any details of the image can be focused. Specifically, the wavelet transform in the multi-frequency state space module uses a set of biorthogonal filter groups to filter the image first in the vertical direction and downsample it twice, and then filter it in the horizontal direction and downsample it twice, and then obtain the corresponding different components, namely low-frequency approximate components, high-frequency horizontal detail components, high-frequency vertical detail components, and high-frequency diagonal detail components. The formula is as follows:

[0057] ;

[0058] in, is the input feature map, represents the frequency domain transform operation, is the low-frequency approximation component, which contains the global information of the image, and , , They are high-frequency horizontal detail component, high-frequency vertical detail component, and high-frequency diagonal detail component (containing detail information).

[0059] Next, the multi-frequency state-space module applies the state-space model to extract features from frequency domain components of different scales. The state-space model (SSM), represented by Mamba, has become a powerful alternative to CNN and Transformer architectures due to its efficient long-distance dependency modeling capabilities and low computational complexity, especially showing strong feature extraction capabilities when processing high-resolution medical images. The multi-frequency state-space module uses the selective scanning mechanism and hardware-aware algorithm in Mamba to achieve efficient long-distance dependency modeling and significantly improve training and inference efficiency. The selective scanning mechanism avoids comprehensive scanning of all areas by focusing on important areas in the image, further reducing the computational burden. Specifically, the feature map of each frequency domain component is first input into a linear layer and a convolutional layer to extract the initial local spatial features. Then, in order to further enrich the feature representation, the features of the previous step will be sent to a state space 2D selective scanning module (SS2DSM), which scans from the top left to the bottom right, the bottom left to the top right, the top right to the bottom left, and the bottom right to the top left of the feature map slice. After the scan is completed, the state-space two-dimensional selective scanning module serializes the results of the "four-way" scan, then uses the state-space model for selective scanning, and finally restores and fuses them into a feature map. In this way, the system can extract global structural and semantic information. Then, the feature map is downsampled to reduce the spatial dimension of the feature map, so that the higher-level feature map can cover a larger receptive field. After four iterations of the above operations, the final feature map of each frequency domain component can be obtained. Finally, the multi-frequency state space module will fuse the feature maps at different frequency domain scales to generate the final initial colon endoscopy multi-scale contextual semantic encoding feature map. This process can be implemented in a variety of ways, such as simple element-by-element addition, pixel-by-pixel multiplication, or a more complex attention mechanism to form a comprehensive feature representation, which provides an important foundation for subsequent super-resolution reconstruction.

[0060] In the embodiment of the present application, the preliminary image reconstruction multi-scale feature extraction unit 142 is used to perform preliminary image reconstruction and multi-scale feature extraction on the initial colon endoscopy multi-scale contextual semantic coding features to generate colon endoscopy multi-scale contextual semantic coding features. Figure 3 FIG. 1 is a block diagram of a multi-scale feature extraction unit for preliminary image reconstruction in a colorectal lesion auxiliary diagnosis system based on AI and cell endoscopy according to an embodiment of the present application. Specifically, Figure 3As shown, the preliminary image reconstruction multi-scale feature extraction unit 142 includes: a preliminary image reconstruction subunit 1421, which is used to input the initial colonoscopy multi-scale context semantic coding feature map into the preliminary reconstruction module based on the adversarial generation network to obtain a preliminary resolution optimized colonoscopy image; a reconstructed image multi-scale extraction subunit 1422, which is used to input the preliminary resolution optimized colonoscopy image into the image multi-scale feature extractor based on the multi-frequency state space module to obtain a generated colonoscopy multi-scale context semantic coding feature map as the generated colonoscopy multi-scale context semantic coding feature.

[0061] Specifically, the preliminary image reconstruction subunit 1421 is used to input the initial colonoscopy multi-scale contextual semantic encoding feature map into a preliminary reconstruction module based on a generative adversarial network to obtain a preliminary resolution-optimized colonoscopy image. Accordingly, considering that high-resolution images can provide clearer and more detailed visual information, allowing doctors to more easily identify and evaluate the subtle structures and potential lesions inside the intestine, based on this, in the technical solution of the present application, the initial colonoscopy multi-scale contextual semantic encoding feature map is input into a preliminary reconstruction module based on a generative adversarial network to use the generative adversarial technology without losing the original image information. , To generate a more realistic and detailed preliminary resolution-optimized colonoscopy image.

[0062] It should be understood that the adversarial generative network consists of two main parts: the generator and the discriminator. The task of the generator is to generate realistic and high-resolution images from low-resolution feature maps; while the discriminator is responsible for distinguishing the difference between the generated images and the real high-resolution images, and providing feedback to the generator to help it continuously improve the generation effect. This generation-discrimination cycle iterative mechanism allows the generator to gradually learn how to produce more and more realistic high-resolution images.

[0063] When the initial colonoscopy multi-scale contextual semantic encoding feature map enters the generator, it first goes through multiple levels of deconvolution layers (also called transposed convolution layers), which gradually enlarge the spatial size of the feature map while keeping its semantic information intact. In this process, the generator also introduces residual connections to ensure that low-level detail information can be directly passed to high-level structures to avoid information loss due to excessive abstraction. In addition, in order to enhance the realism of the generated image, the generator may also use an attention mechanism to allow the model to pay more attention to those areas or features that are most important for specific tasks (such as lesion detection). As the feature map is enlarged after a series of deconvolution operations, the generator finally outputs a colonoscopy image with preliminary resolution optimization. Although this image already has a high resolution, its authenticity and detail level need to be further verified. At this point, the discriminator comes into play. The discriminator receives the synthetic image from the generator and the actual high-resolution colonoscopy image collected as input. After multiple layers of convolution operations, it outputs a probability value indicating the possibility of the input image being true or false. Based on this probability value, the system can calculate the gap between the generator and the real data, and adjust the parameters of the generator accordingly so that it can generate more realistic images next time.

[0064] Specifically, the reconstructed image multi-scale extraction subunit 1422 is used to input the preliminary resolution optimized colon endoscopy image into the image multi-scale feature extractor based on the multi-frequency state space module to obtain the generated colon endoscopy multi-scale context semantic coding feature map as the generated colon endoscopy multi-scale context semantic coding feature. It should be understood that, considering that although the adversarial generative network can significantly improve the image resolution, some very subtle structures (such as tiny polyps, early lesions, etc.) may not be completely restored or ignored during the generation process, and high-frequency spatial changes (such as edges and contours) may be lost during the reconstruction process, resulting in the boundaries in the image becoming unclear. Based on this, in the technical solution of the present application, the image multi-scale feature extractor based on the multi-frequency state space module is used again to process the preliminary resolution optimized colon endoscopy image to effectively capture the context information of different scales in the image, maintain good spatial consistency at different scales, and obtain the generated colon endoscopy multi-scale context semantic coding feature map. In this way, it can recover subtle features that may be lost in the initial reconstruction process, helping to more accurately locate lesions or other important structures, so that it can better understand the overall structure of the image and provide richer semantic information.

[0065] In an embodiment of the present application, the feature combination unit 143 is used to perform a priori guided feature significant combination on the generated colon endoscopy multi-scale contextual semantic coding features and the initial colon endoscopy multi-scale contextual semantic coding features to obtain colon endoscopy multi-step contextual semantic coding features. Accordingly, considering that the generated colon endoscopy multi-scale contextual semantic coding features (optimized images after preliminary reconstruction) and the initial colon endoscopy multi-scale contextual semantic coding features (features extracted from the original image) each have unique advantages. The generated feature map may be better in resolution and visual clarity, while the initial feature map retains more original details and structural information. Therefore, in order to comprehensively utilize feature information from two different sources to enhance the quality and diagnostic value of feature representation, in the technical solution of the present application, the generated colon endoscopy multi-scale contextual semantic coding features and the initial colon endoscopy multi-scale contextual semantic coding features are significantly combined by a priori guidance to obtain colon endoscopy multi-step contextual semantic coding features. In particular, this application can use existing medical knowledge or expert experience by introducing prior information to help the model better understand which features are most important for specific tasks (such as lesion detection, classification, etc.), which helps to improve the model's attention to key areas. At the same time, in this process, a converter architecture is introduced to implement a bidirectional attention mechanism, thereby generating a comprehensive feature representation that can more accurately capture the intrinsic structure and semantic information of the data.

[0066] Figure 4 FIG. 1 is a block diagram of a feature combination unit in a colorectal lesion auxiliary diagnosis system based on AI and cell endoscopy according to an embodiment of the present application. Specifically, Figure 4 As shown, the feature combination unit 143 includes: a feature initial feature reshaping modulation subunit 1431, which is used to perform feature shape reshaping and prior modulation on the generated colon endoscopy multi-scale context semantic coding feature map and the initial colon endoscopy multi-scale context semantic coding feature map to obtain a priori modulated colon endoscopy multi-scale context semantic feature shape reshaping matrix and a priori modulated initial colon endoscopy multi-scale context semantic feature shape reshaping matrix; a feature initial feature interaction subunit 1432, which is used to perform feature shape reshaping and prior modulation on the generated colon endoscopy multi-scale context semantic feature shape reshaping matrix and the priori modulated initial colon endoscopy multi-scale context semantic feature shape reshaping matrix. The prior modulated initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix is ​​subjected to attention interaction and shape reshaping to obtain a generated colon endoscopy multi-scale contextual semantic feature interactive coding feature map and an initial colon endoscopy multi-scale contextual semantic feature interactive coding feature map; a generated feature initial feature fusion subunit 1433 is used to fuse the generated colon endoscopy multi-scale contextual semantic feature interactive coding feature map and the initial colon endoscopy multi-scale contextual semantic feature interactive coding feature map to obtain a colon endoscopy multi-step contextual semantic coding feature map as the colon endoscopy multi-step contextual semantic coding feature.

[0067] Specifically, in the embodiment of the present application, the generating feature initial feature reshaping modulation subunit 1431 is used to:

[0068] The generated colon endoscopy multi-scale context semantic encoding feature map and the initial colon endoscopy multi-scale context semantic encoding feature map are subjected to feature shape reshaping to obtain a generated colon endoscopy multi-scale context semantic feature shape reshaping matrix and an initial colon endoscopy multi-scale context semantic feature shape reshaping matrix, namely:

[0069] ;

[0070] ;

[0071] in, For generating a multi-scale contextual semantic encoding feature map of colon endoscopy, is the initial colon endoscopy multi-scale contextual semantic encoding feature map, For the reshape operation, It is to generate the shape reshaping matrix of multi-scale contextual semantic features of colon endoscopy, is the initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix;

[0072] The generated colon endoscopy multi-scale contextual semantic feature shape reshaping matrix and the initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix are input into a priori guided modulation module to obtain the priori modulated generated colon endoscopy multi-scale contextual semantic feature shape reshaping matrix and the priori modulated initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix, that is:

[0073] ;

[0074] ;

[0075] in, It is to generate the shape reshaping matrix of multi-scale contextual semantic features of colon endoscopy, is the initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix, and represents different learnable memory parameter matrices in the a priori guided modulation module, for The transposed matrix of is matrix multiplication, represents the normalization function, It is the prior modulation to generate the multi-scale contextual semantic feature shape reshaping matrix of colon endoscopy, It is the prior modulated initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix.

[0076] It should be understood that the generated colonoscopy multi-scale contextual semantic encoding feature map and the initial colonoscopy multi-scale contextual semantic encoding feature map are feature reshaped to expand each feature matrix along the channel dimension to obtain the generated colonoscopy multi-scale contextual semantic feature shape reshaped matrix (HW×C) and the initial colonoscopy multi-scale contextual semantic feature shape reshaped matrix (HW×C), so that they can have the same spatial dimension, laying the foundation for subsequent processing.

[0077] Accordingly, the two obtained shape reshaping matrices are input into the prior guided modulation module to utilize the prior knowledge in the medical field to adjust the feature representation, so that the model can pay more attention to those areas or features that are critical to specific tasks (such as lesion detection), and enhance the ability to understand and capture key features, so as to obtain the prior modulated generated colon endoscopy multi-scale contextual semantic feature shape reshaping matrix and the prior modulated initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix.

[0078] Specifically, in the embodiment of the present application, the generating feature initial feature interaction subunit 1432 is used to:

[0079] The prior modulated colon endoscopy multi-scale contextual semantic feature shape reshaping matrix and the prior modulated initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix are input into the attention interaction coding module based on the transformer-like structure to obtain the generated colon endoscopy multi-scale contextual semantic feature interaction coding feature matrix and the initial colon endoscopy multi-scale contextual semantic feature interaction coding feature matrix, that is:

[0080] ;

[0081] ;

[0082] in, It is the prior modulation to generate the multi-scale contextual semantic feature shape reshaping matrix of colon endoscopy, is the prior modulated initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix, yes The transposed matrix of yes The transposed matrix of and They are and The scale is the matrix width multiplied by the matrix height, is the normalization function, It is to generate the interactive encoding feature matrix of multi-scale contextual semantic features of colon endoscopy, is the initial colon endoscopy multi-scale contextual semantic feature interaction encoding feature matrix;

[0083] The generated colon endoscopy multi-scale contextual semantic feature interactive coding feature matrix and the initial colon endoscopy multi-scale contextual semantic feature interactive coding feature matrix are reshaped to obtain the generated colon endoscopy multi-scale contextual semantic feature interactive coding feature map and the initial colon endoscopy multi-scale contextual semantic feature interactive coding feature map, that is:

[0084] ;

[0085] ;

[0086] in, It is to generate the interactive encoding feature matrix of multi-scale contextual semantic features of colon endoscopy, is the initial colon endoscopy multi-scale contextual semantic feature interaction encoding feature matrix, It generates a multi-scale contextual semantic feature interactive encoding feature map of colon endoscopy. It is the initial colon endoscopy multi-scale contextual semantic feature interaction encoding feature map.

[0087] It should be understood that the matrix after prior modulation is input into the attention interaction encoding module based on the transformer-like structure to obtain the generated colon endoscopy multi-scale contextual semantic feature interaction encoding feature matrix and the initial colon endoscopy multi-scale contextual semantic feature interaction encoding feature matrix. In particular, the use of a self-attention mechanism similar to that in the transformer structure can enable a deeper interaction between features, thereby better modeling long-distance dependencies and highlighting the most important parts. In addition, this interaction can also help the model learn the potential connections between different features, further improving its expressiveness.

[0088] Subsequently, in order to make the spatial structure of the interactively encoded feature matrix consistent with the original feature map, so as to facilitate the intuitive understanding and analysis of the feature distribution, the interactively encoded features are reshaped to obtain the generated colon endoscopy multi-scale contextual semantic feature interactive encoding feature map and the initial colon endoscopy multi-scale contextual semantic feature interactive encoding feature map.

[0089] Specifically, in the embodiment of the present application, the generated feature initial feature fusion subunit 1433 is used to: calculate the weighted sum of the generated colon endoscopy multi-scale contextual semantic feature interactive coding feature map and the initial colon endoscopy multi-scale contextual semantic feature interactive coding feature map to obtain the colon endoscopy multi-step contextual semantic coding feature map, that is:

[0090] ;

[0091] in, It generates a multi-scale contextual semantic feature interactive encoding feature map of colon endoscopy. is the initial colon endoscopy multi-scale contextual semantic feature interactive encoding feature map, and is the weighted hyperparameter, It is the multi-step contextual semantic encoding feature map of colon endoscopy.

[0092] Finally, the weighted sum of the reshaped features is calculated to give appropriate weights to the generated feature map and the initial feature map and sum them up. This can effectively reduce noise interference and redundant information while maintaining the advantages of both, and ultimately form a more accurate and reliable colon endoscopy multi-step contextual semantic encoding feature map. In particular, this feature map not only retains the key details of the original image, but also combines the advantages of the generated image, while also incorporating the influence of prior knowledge and the correlation significantly enhanced by the attention mechanism, thereby improving the model's recognition and interpretability of lesions or other important structures in colon endoscopy images.

[0093] In an embodiment of the present application, the image optimization unit 144 is used to obtain the resolution-optimized colonoscopy image based on the colonoscopy multi-step contextual semantic coding features. Specifically, in an embodiment of the present application, the image optimization unit 144 is used to: input the colonoscopy multi-step contextual semantic coding feature map into a re-reconstruction module based on a generative adversarial network to obtain the resolution-optimized colonoscopy image. That is, the colonoscopy multi-step contextual semantic coding features obtained by a priori guidance of the generated colonoscopy multi-scale contextual semantic coding feature map and the initial colonoscopy multi-scale contextual semantic coding feature map are generated and processed to intelligently generate the resolution-optimized colonoscopy image. In this way, the quality and resolution of the image can be further improved, so that the final resolution-optimized colonoscopy image can be clearer and more realistic, and retain important medical diagnostic information.

[0094] It should be understood that here, the generated colon endoscopy multi-scale contextual semantic coding feature map and the initial colon endoscopy multi-scale contextual semantic coding feature map respectively represent the multi-scale image semantic coding features of the denoised optimized colon endoscopy image and the multi-scale image semantic coding features of the preliminary resolution optimized colon endoscopy image. When performing feature significant joint perception based on prior guidance, it is considered that the preliminary resolution optimized colon endoscopy image is adversarially generated based on the image semantic coding features of the denoised optimized colon endoscopy image. Therefore, there is a lot of spatial structural redundancy in the generated colon endoscopy multi-scale contextual semantic coding feature map and the initial colon endoscopy multi-scale contextual semantic coding feature map, which will make the colon endoscopy multi-step contextual semantic coding feature map obtained by feature significant joint perception have significant interactive optimization distribution spatial structure differences, affecting the convergence consistency of the adversarial generation network, thereby affecting the image quality of the resolution optimized colon endoscopy image obtained by the re-reconstruction module based on the adversarial generation network.

[0095] Therefore, in one example, the colon endoscopy multi-step context semantic encoding feature map is optimized, and the optimization process includes:

[0096] The absolute value sum and the square root of the square sum of all feature values ​​of the colon endoscopy multi-step context semantic encoding feature map are calculated to obtain the first colon endoscopy multi-step context semantic encoding space structure value and the second colon endoscopy multi-step context semantic encoding space structure value, that is:

[0097] ;

[0098] ;

[0099] in, is the first The eigenvalues ​​at the positions, is the spatial structure value of the first colon endoscopy multi-step contextual semantic encoding, It is the second colon endoscopy multi-step context semantic encoding spatial structure value;

[0100] Determine the total number of feature values ​​of all feature values ​​of the colon endoscopy multi-step context semantic encoding feature map ;

[0101] For each feature value of the colon endoscopy multi-step context semantic coding feature map, the first colon endoscopy multi-step context semantic coding spatial structure value minus the product of the feature value and the total number of feature values ​​is calculated to obtain the first colon endoscopy multi-step context semantic coding long-range dependency value ;

[0102] in, is the total number of eigenvalues ​​of all eigenvalues ​​of the multi-step context semantic encoding feature map of the colon endoscopy, yes The corresponding first colon endoscopy multi-step context semantic encoding long-range dependency value;

[0103] The second colon endoscopy multi-step context semantic coding long-range dependency value is obtained by calculating the square root of the total number of eigenvalues ​​multiplied by the product of the eigenvalues ​​minus the second colon endoscopy multi-step context semantic coding spatial structure value. ;

[0104] in, yes The corresponding second colon endoscopy multi-step context semantic encoding long-range dependency value;

[0105] The index value calculated by taking the first colonoscopy multi-step context semantic encoding long-range dependency value as the exponent of the natural constant and the inverse of the second colonoscopy multi-step context semantic encoding long-range dependency value are weighted summed to obtain the optimized eigenvalue corresponding to each eigenvalue ;

[0106] in, and is the weighting parameter, yes The corresponding optimized eigenvalue;

[0107] The optimized feature values ​​are combined into an optimized colon endoscopy multi-step contextual semantic encoding feature map.

[0108] That is, in view of the possible lack of spatial structure in the high-dimensional space of the feature set of the colon endoscopy multi-step contextual semantic encoding feature map, which causes the weight of the classifier to implicitly infer the spatial structure information based on the features, resulting in inconsistent convergence, a long-distance feature dependency relationship is established based on the overall feature scale of the colon endoscopy multi-step contextual semantic encoding feature map relative to the spatial structure representation of the colon endoscopy multi-step contextual semantic encoding feature map, so as to establish the local connectivity of the features of the colon endoscopy multi-step contextual semantic encoding feature map, and the spatial ambiguous information of the object feature value is captured by predicting the unstructured feature value points of the colon endoscopy multi-step contextual semantic encoding feature map, thereby improving the spatial inductive bias perception ability of the feature set of the colon endoscopy multi-step contextual semantic encoding feature map, improving the convergence consistency of the classifier, and improving the resolution of the colon endoscopy multi-step contextual semantic encoding feature map obtained by the re-reconstruction module based on the generative adversarial network to optimize the image quality of the colon endoscopy image.

[0109] In summary, the colonoscopy image resolution optimization module 140 is explained, which uses an image processing and optimization algorithm based on AI and machine learning to perform the initial multi-scale context semantic coding of the denoised and optimized colonoscopy image. Then, the initial colonoscopy multi-scale context semantic coding features after the initial coding are preliminarily reconstructed and multi-scale context semantic coding is performed again, so as to intelligently generate the resolution-optimized colonoscopy image based on the prior-guided fusion representation between the re-encoded generated colonoscopy multi-scale context semantic coding features and the initial colonoscopy multi-scale context semantic coding features. In this way, structural information at different levels can be captured, and important details can be retained while effectively removing noise. In addition, by using prior information, it can be ensured that the generated high-resolution image is not only more realistic visually, but also can adapt to the special requirements of different fields, making the feature expression more accurate, thereby facilitating subsequent colon status analysis and lesion area determination.

[0110] Specifically, the colon state feature extraction module 150 is used to extract colon state image features from the resolution-optimized colon endoscopy image. It should be understood that the resolution-optimized colon endoscopy image contains pathological information such as polyps, inflammation, ulcers, tumors, etc. Therefore, in order to more accurately analyze and utilize this information, thereby improving the accuracy and efficiency of diagnosis, in the technical solution of the present application, colon state image features are extracted from the resolution-optimized colon endoscopy image. In particular, in a specific embodiment of the present application, the resolution-optimized colon endoscopy image can be processed by using a hole convolutional neural network model to insert a "hole" into the standard convolution kernel to capture and mine key pathological state feature information to obtain colon state image features.

[0111] First of all, considering that the colorectal endoscopy images after resolution optimization already have high clarity and detail retention capabilities, these images provide a solid foundation for feature extraction. In order to make full use of these high-quality images, dilated convolution or expanded convolution methods are usually used. Unlike standard convolution, dilated convolution expands the receptive field without increasing the number of parameters by inserting "holes" or jump sampling in the convolution kernel, thereby capturing contextual information in a wider range without losing spatial resolution. This is especially important for colorectal endoscopy images, because such images often contain rich textures and complex morphological changes, requiring the model to be able to perceive a wider field of view in order to accurately identify various lesion features.

[0112] Next, when building a neural network for feature extraction, the idea of ​​a residual network is often introduced. The residual network solves the problem of difficult training of deep networks by adding cross-layer connections, and helps to keep the information flow path unobstructed. When applied to the task of extracting features from colon endoscopy images, this architecture can effectively prevent the occurrence of gradient vanishing and ensure stable end-to-end learning even when the network is very deep. In addition, the attention mechanism can be embedded in the residual block to allow the model to focus more on areas or features that are critical to specific tasks, such as polyp edges, inflammatory areas, etc.

[0113] Specifically, the lesion area determination module 160 is used to determine whether there is a lesion area through the large intestine status image features. Specifically, in an embodiment of the present application, the lesion area determination module is used to: use a classifier-based lesion area discriminator to process the large intestine status image features to obtain a discrimination result indicating whether there is a lesion area. That is, the classifier can automatically identify areas where lesions may exist based on the input feature map, and output a probability value or a binary label to indicate whether the area is a lesion. In this way, potential lesion areas, especially those early or tiny lesions, can be quickly identified, providing doctors with reliable auxiliary opinions and providing more comprehensive support for clinical decision-making.

[0114] In summary, the colorectal lesion auxiliary diagnosis system 100 based on AI and cytoscopy according to the embodiment of the present application is explained, which first receives images collected by cytoscopy, then optimizes the brightness of these images, and then denoises the brightness-optimized images to obtain clearer denoised optimized images, and further, super-resolution reconstruction is performed on the denoised images to improve the resolution of the images, and finally, the characteristics of the colorectal state are extracted from the resolution-optimized images, and the lesion area is determined. In this way, the quality of colon endoscopy images can be significantly improved, and the identification of lesion areas can be more accurate, thereby improving the accuracy and efficiency of diagnosis.

[0115] As described above, the colorectal lesion auxiliary diagnosis system 100 based on AI and cytoscopy according to the embodiment of the present application can be implemented in various wireless terminals, such as a colorectal lesion auxiliary diagnosis server based on AI and cytoscopy, etc. In one possible implementation, the colorectal lesion auxiliary diagnosis system 100 based on AI and cytoscopy according to the embodiment of the present application can be integrated into the wireless terminal as a software module and / or hardware module. For example, the colorectal lesion auxiliary diagnosis system 100 based on AI and cytoscopy can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the colorectal lesion auxiliary diagnosis system 100 based on AI and cytoscopy can also be one of the many hardware modules of the wireless terminal.

[0116] Alternatively, in another example, the AI ​​and cytoscopy-based colorectal lesion auxiliary diagnosis system 100 and the wireless terminal may also be separate devices, and the AI ​​and cytoscopy-based colorectal lesion auxiliary diagnosis system 100 may be connected to the wireless terminal via a wired and / or wireless network and transmit interactive information in accordance with an agreed data format.

[0117] The above has described various implementations of the present disclosure, and the above description is exemplary and not exhaustive. It is not limited to the disclosed implementations, and many modifications and changes are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used in this article is intended to best explain the principles of each implementation, practical application or improvement of technology in the market, or to enable other ordinary technicians in this technical field to understand the various embodiments disclosed herein.

Claims

1. A colorectal lesion auxiliary diagnosis system based on AI and cytoendoscopy, characterized in that: include: A colon endoscopy image acquisition module, used for receiving colon endoscopy images acquired by a cytoscopy; An endoscopic image brightness adjustment module, used for performing brightness adjustment on the colonoscopic image to obtain a brightness-optimized colonoscopic image; a colonoscopic image denoising module, used for performing image denoising on the brightness-optimized colonoscopic image to obtain a denoised-optimized colonoscopic image; a colonoscopic image resolution optimization module, used for performing super-resolution reconstruction on the denoised-optimized colonoscopic image to obtain a resolution-optimized colonoscopic image; a colon state feature extraction module, used for extracting colon state image features from the resolution-optimized colonoscopic image; A lesion area determination module, used to determine whether there is a lesion area according to the large intestine state image features; Among them, the colon endoscopy image resolution optimization module includes: an initial multi-scale feature extraction unit, which is used to perform multi-scale feature extraction on the denoised and optimized colon endoscopy image to obtain initial colon endoscopy multi-scale contextual semantic coding features; a preliminary image reconstruction multi-scale feature extraction unit, which is used to perform preliminary image reconstruction and multi-scale feature extraction on the initial colon endoscopy multi-scale contextual semantic coding features to obtain generated colon endoscopy multi-scale contextual semantic coding features; a feature combination unit, which is used to perform a priori guided feature significant combination on the generated colon endoscopy multi-scale contextual semantic coding features and the initial colon endoscopy multi-scale contextual semantic coding features to obtain colon endoscopy multi-step contextual semantic coding features; an image optimization unit, which is used to reconstruct the colon endoscopy multi-step contextual semantic coding features to obtain the resolution optimized colon endoscopy image; Wherein, the preliminary image reconstruction multi-scale feature extraction unit comprises: A preliminary image reconstruction subunit, used for inputting the initial colon endoscopy multi-scale context semantic encoding feature map into a preliminary reconstruction module based on a generative adversarial network to obtain a preliminary resolution optimized colon endoscopy image; A reconstructed image multi-scale extraction subunit is used to input the preliminary resolution optimized colon endoscopy image into an image multi-scale feature extractor based on a multi-frequency state space module to obtain a generated colon endoscopy multi-scale contextual semantic coding feature map as the generated colon endoscopy multi-scale contextual semantic coding feature; Among them, the multi-frequency state space module first uses wavelet transform to filter the image in the vertical direction and downsample it twice, and then filters it in the horizontal direction and downsamples it twice to obtain frequency domain components of different scales. Then, the state space model is used to extract features of frequency domain components of different scales to obtain the final feature map of each frequency domain component. Finally, the feature maps at different frequency domain scales are fused to generate the final feature map.

2. The colorectal lesion auxiliary diagnosis system based on AI and cell endoscopy according to claim 1 is characterized in that: The initial multi-scale feature extraction unit is used to: input the denoised and optimized colon endoscopy image into an image multi-scale feature extractor based on a multi-frequency state space module to obtain an initial colon endoscopy multi-scale contextual semantic coding feature map as the initial colon endoscopy multi-scale contextual semantic coding feature.

3. The colorectal lesion auxiliary diagnosis system based on AI and cell endoscopy according to claim 2 is characterized in that: The feature combining unit comprises: A generated feature initial feature reshaping modulation subunit is used to perform feature shape reshaping and prior modulation on the generated colon endoscopy multi-scale context semantic coding feature map and the initial colon endoscopy multi-scale context semantic coding feature map to obtain a priori modulated generated colon endoscopy multi-scale context semantic feature shape reshaping matrix and a priori modulated initial colon endoscopy multi-scale context semantic feature shape reshaping matrix; A generated feature initial feature interaction subunit is used to perform attention interaction and shape reshaping on the prior modulated generated colon endoscopy multi-scale contextual semantic feature shape reshaping matrix and the prior modulated initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix to obtain a generated colon endoscopy multi-scale contextual semantic feature interaction encoding feature map and an initial colon endoscopy multi-scale contextual semantic feature interaction encoding feature map; A generated feature initial feature fusion subunit is used to fuse the generated colonoscopy multi-scale contextual semantic feature interactive coding feature map and the initial colonoscopy multi-scale contextual semantic feature interactive coding feature map to obtain a colonoscopy multi-step contextual semantic coding feature map as the colonoscopy multi-step contextual semantic coding feature.

4. The colorectal lesion auxiliary diagnosis system based on AI and cell endoscopy according to claim 3 is characterized in that: The generating feature initial feature reshaping modulation subunit is used to: Performing feature shape reshaping on the generated colon endoscopy multi-scale context semantic encoding feature map and the initial colon endoscopy multi-scale context semantic encoding feature map to obtain a generated colon endoscopy multi-scale context semantic feature shape reshaping matrix and an initial colon endoscopy multi-scale context semantic feature shape reshaping matrix; The generated colon endoscopy multi-scale contextual semantic feature shape reshaping matrix and the initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix are input into a priori guided modulation module to obtain the priori modulated generated colon endoscopy multi-scale contextual semantic feature shape reshaping matrix and the priori modulated initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix.

5. The colorectal lesion auxiliary diagnosis system based on AI and cell endoscopy according to claim 4 is characterized in that: The generating feature initial feature interaction subunit is used to: Input the prior modulated colon endoscopy multi-scale contextual semantic feature shape reshaping matrix and the prior modulated initial colon endoscopy multi-scale contextual semantic feature shape reshaping matrix into an attention interaction coding module based on a transformer-like structure to obtain a generated colon endoscopy multi-scale contextual semantic feature interaction coding feature matrix and an initial colon endoscopy multi-scale contextual semantic feature interaction coding feature matrix; The generated colonoscopy multi-scale contextual semantic feature interactive encoding feature matrix and the initial colonoscopy multi-scale contextual semantic feature interactive encoding feature matrix are feature reshaped to obtain the generated colonoscopy multi-scale contextual semantic feature interactive encoding feature map and the initial colonoscopy multi-scale contextual semantic feature interactive encoding feature map.

6. The colorectal lesion auxiliary diagnosis system based on AI and cell endoscopy according to claim 5 is characterized in that: The generated feature initial feature fusion subunit is used to calculate the weighted sum of the generated colon endoscopy multi-scale contextual semantic feature interactive coding feature map and the initial colon endoscopy multi-scale contextual semantic feature interactive coding feature map to obtain the colon endoscopy multi-step contextual semantic coding feature map.

7. The colorectal lesion auxiliary diagnosis system based on AI and cell endoscopy according to claim 6 is characterized in that: The image optimization unit is used to: input the colon endoscopy multi-step context semantic encoding feature map into a re-reconstruction module based on a generative adversarial network to obtain the resolution-optimized colon endoscopy image.

8. The colorectal lesion auxiliary diagnosis system based on AI and cell endoscopy according to claim 7 is characterized in that: The lesion area determination module is used to: use a lesion area discriminator based on a classifier to process the large intestine state image features to obtain a discrimination result indicating whether a lesion area exists.

Citation Information

Patent Citations

  • Reference remote sensing image super-resolution reconstruction method based on feature matching

    CN117196944A

  • Arbitrary scale image super-resolution network construction method, system, chip and equipment

    CN117196952A