Image lesion recognition system to assist gastroenterologists in sampling and examination
By designing a multimodal AI-fusion image lesion recognition system, the problems of strong subjectivity and easy missed lesions in imaging diagnosis of digestive system diseases are solved, efficient and accurate diagnosis of digestive tract diseases are achieved, and sampling operations and data management are optimized.
Patent Information
- Application Number
- CN202510339506.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The prior art has problems such as strong subjectivity in the imaging diagnosis of digestive system diseases, easy to miss diagnosis of small lesions, insufficient data integration and real-time decision-making support, difficulty in dynamic integration of multimodal imaging analysis results, and lack of intelligent recommendation capabilities for individualized cases.
An image lesion recognition system assisting gastroenterologists in sampling examination was designed, including image acquisition module, preprocessing module, AI lesion recognition module, feature analysis module, sampling recommendation module, data storage module, real-time interaction module and system control center. The system integrates white light, narrowband light and fluorescence endoscopic images through multimodal large models, combines superpixel segmentation and three-dimensional reconstruction technology to achieve accurate identification and quantitative analysis of lesions, and dynamically generates sampling solutions through reinforcement learning strategies and biomechanical simulation.
It significantly improves the objectivity and accuracy of diagnosis, optimizes the accuracy and efficiency of doctors' operations, ensures the security of patient privacy, and promotes the optimal allocation of medical resources.
Smart Images

Figure CN119850632B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image lesion recognition, and in particular to an image lesion recognition system that assists gastroenterologists in sampling and examination. Background Art
[0002] The imaging diagnosis of digestive system diseases mainly relies on technologies such as endoscopy, CT, and MRI, but these methods often require doctors to manually analyze massive amounts of imaging data, which is highly subjective and easy to miss tiny lesions. For example, traditional endoscopic examinations rely on the doctor's experience to determine the location of lesions. The recognition rate for early lesions or lesions with blurred boundaries (such as flat polyps and superficial ulcers) is limited, and the selection of sampling locations lacks quantitative basis, which may lead to repeated sampling or missed detection. In addition, in the existing imaging diagnosis process, data integration and real-time decision support are insufficient, the analysis results of different modal images (such as white light and narrow-band light imaging) are difficult to dynamically fuse, and there is a lack of intelligent recommendation capabilities for individualized cases.
[0003] The main defects of existing technologies are reflected in the deficiencies in intelligence and systematization. First, traditional lesion identification mostly relies on single image modality analysis. For example, endoscopic ultrasound or CT alone is difficult to integrate multispectral imaging features (such as molecular information of fluorescence imaging and three-dimensional morphology of structured light projection), resulting in incomplete extraction of lesion features. Secondly, existing auxiliary systems generally lack real-time interaction and dynamic guidance capabilities. Doctors need to frequently switch between observation interfaces and operating tools during operations. The application of augmented reality technologies such as AR navigation and tactile feedback has not yet been popularized, which increases the complexity of operations and learning costs. In addition, there are shortcomings in data management and model optimization: most systems use centralized data storage, which faces the risk of privacy leakage when collaborating across institutions, and model updates rely on local data. It is difficult to achieve multi-center knowledge sharing through federated learning, which limits the generalization ability of diagnostic models.
[0004] The deeper contradiction lies in technology integration and clinical adaptability. Existing image analysis tools are often isolated from the diagnosis and treatment process. For example, endoscopic image processing is separated from the electronic medical record system, and the diagnostic strategy cannot be dynamically adjusted in combination with the patient's medical history. At the hardware level, the computing unit of traditional equipment is fixed, and it is difficult to support real-time 4K image processing and lightweight deployment. It is difficult to balance battery life and performance in mobile diagnosis and treatment scenarios. These defects together make it difficult for existing technologies to meet the core requirements of efficient, accurate and standardized operations in the early diagnosis of gastrointestinal diseases. There is an urgent need to achieve breakthroughs through multimodal AI fusion, edge computing optimization and human-computer collaborative interaction. Summary of the invention
[0005] In order to overcome the shortcomings and deficiencies of the prior art, the present invention provides an image lesion recognition system for assisting gastroenterologists in sampling and examination.
[0006] The image lesion recognition system that assists gastroenterologists in sampling and examination includes: image acquisition module, preprocessing module, AI lesion recognition module, feature analysis module, sampling recommendation module, data storage module, real-time interaction module and system control center;
[0007] The image acquisition module acquires digestive tract endoscopic image data in real time through an endoscopic device, including white light imaging, narrow-band light imaging and fluorescence imaging multimodal images, and transmits them to the preprocessing module;
[0008] The preprocessing module includes an image enhancement unit, a noise suppression unit and an image segmentation unit, wherein the noise suppression unit adopts a non-local mean denoising algorithm combined with an adaptive median filter, and the image segmentation unit realizes a preliminary separation of the digestive tract mucosal area and the lesion based on a U-Net++ network;
[0009] The AI lesion recognition module is driven by a multimodal large model, including a visual Transformer backbone network and an adaptive feature fusion module. The backbone network extracts global image features through a pre-trained ViT-Huge model, and the fusion module combines local features of the convolutional neural network to output category labels and lesion bounding box coordinates for polyps, ulcers, and bleeding points.
[0010] The feature analysis module performs multi-dimensional quantitative analysis on the identified lesions, including a morphological parameter calculation unit based on superpixel segmentation, an HSV color space histogram analysis unit, and a gray-level co-occurrence matrix texture feature extraction unit, and generates a comprehensive report including lesion size, shape irregularity, color heterogeneity, and texture complexity.
[0011] The sampling recommendation module integrates a reinforcement learning strategy, dynamically generates the optimal sampling location and quantity through a probabilistic graph model according to the lesion type, ESD / EMR treatment guidelines and historical sampling success rate data, and outputs a visual heat map superimposed on the endoscopic image;
[0012] The data storage module uses blockchain encryption technology to store original images, lesion annotation data, feature analysis results and sampling records, and supports multi-condition retrieval and comparative analysis based on case ID;
[0013] The real-time interaction module includes an AR display unit and a tactile feedback unit. The AR display unit projects a three-dimensional reconstruction model of the lesion and a sampling path guide through a head-mounted device. The tactile feedback unit generates a vibration prompt when the biopsy forceps approaches the recommended sampling position.
[0014] The system control center coordinates the data flows of each module through edge computing nodes, and deploys a federated learning framework to achieve collaborative optimization of models among multiple medical institutions.
[0015] Furthermore, the focus identification module also includes:
[0016] The dynamic weight adjustment submodule automatically adjusts the attention weights of different feature layers according to the classification results of the digestive tract anatomical parts. The lesion detection in the esophagus, gastric body, and duodenum corresponds to different feature fusion strategies respectively.
[0017] The online incremental learning submodule uses doctors’ revised annotation data to update model parameters in real time and compresses the model through knowledge distillation technology to adapt to embedded device deployment.
[0018] Furthermore, the feature analysis module further includes:
[0019] The 3D reconstruction unit generates a point cloud model of the lesion based on binocular endoscopic images or structured light projection data, and calculates the volume, surface area, and invasion depth parameters;
[0020] The malignancy risk prediction unit inputs the morphological features and the patient's electronic medical record data into the XGBoost classifier and outputs the early cancer probability score.
[0021] Furthermore, the AR display unit of the real-time interaction module implements:
[0022] Multi-view collaborative display function, supporting picture-in-picture display of the main operation screen and the local magnified screen of the lesion;
[0023] Dynamic navigation line generation function adjusts the display density and transparency of the sampling path in real time according to the endoscope advancement speed.
[0024] Furthermore, the sampling recommendation module comprises:
[0025] Biomechanical simulation unit, based on finite element analysis, predicts tissue deformation at different sampling depths to avoid areas with dense blood vessels;
[0026] The historical case matching unit retrieves the sampling schemes and postoperative pathological results of similar lesions through the graph database.
[0027] Furthermore, the image enhancement unit of the preprocessing module adopts:
[0028] Adaptive gamma correction algorithm dynamically adjusts the contrast enhancement curve according to different imaging modes;
[0029] The low-quality image restoration function implemented by the generative adversarial network eliminates motion blur and mirror reflection artifacts.
[0030] Furthermore, the system control center also integrates:
[0031] Dynamic resource allocation unit, automatically switches model accuracy mode according to GPU memory occupancy;
[0032] The multimodal data synchronization unit ensures the alignment of timestamps of the endoscopic video stream, patient vital signs data, and lesion annotation information.
[0033] Furthermore, the data storage module further comprises:
[0034] The privacy protection unit uses homomorphic encryption technology to process metadata containing patient identity information;
[0035] Intelligent compression unit, which performs JPEG2000 lossless compression on medical images while retaining the original bit depth of the lesion area.
[0036] Furthermore, the system also includes:
[0037] The disinfection monitoring module tracks the disinfection status of endoscopic instruments through radio frequency identification technology and displays a warning of the sterilization validity period on the interface;
[0038] The power consumption optimization module uses neural architecture search technology to generate a lightweight inference model on the device side, supporting 8 hours of continuous operation under battery power.
[0039] Furthermore, the hardware architecture of the system includes:
[0040] Removable computing unit, quickly replace and upgrade GPU accelerator card through PCIe interface;
[0041] Redundant power supply module, adopting dual-circuit power supply design and supercapacitor instantaneous power-off protection circuit;
[0042] Modular mechanical structure, connecting the endoscope host and auxiliary equipment through a magnetic quick-release interface. Beneficial Effects
[0043] The present invention proposes an image lesion recognition system to assist gastroenterologists in sampling and examination. The system integrates intelligence with clinical medicine in depth, realizing intelligent innovation of the entire process of gastrointestinal disease diagnosis and treatment. Based on a multimodal large model, the system integrates white light, narrow-band light and fluorescent endoscopic images, accurately identifies lesions such as polyps and ulcers in real time and marks their boundaries. It combines superpixel segmentation and three-dimensional reconstruction technology to quantitatively analyze the morphology, infiltration depth and texture characteristics of the lesions, significantly improving the objectivity of diagnosis. The system relies on reinforcement learning strategies and biomechanical simulation to dynamically generate sampling schemes that adapt to clinical guidelines. It uses augmented reality technology to spatially align the three-dimensional model of the lesion with the real-time endoscopic image, and guides the operation path with tactile feedback, greatly optimizing the accuracy and efficiency of the doctor's operation. At the data management level, the federated learning framework and blockchain encryption technology build a secure channel for cross-institutional collaboration, ensuring patient privacy while promoting continuous optimization of the model, while intelligent compression and multimodal synchronization technology take into account storage efficiency and data integrity. The system also innovatively integrates a full-process quality control system, from multispectral imaging enhancement, artifact repair pretreatment to instrument disinfection status tracking, to comprehensively reduce operational risks. The hardware design adopts a modular architecture, supports flexible upgrades of high-performance computing units and lightweight model deployment, and adapts to real-time processing requirements in mobile diagnosis and treatment scenarios. Through the deep collaboration of AI technology and clinical experience, the system not only improves the detection ability of early gastrointestinal lesions, but also reduces unnecessary repeated sampling operations, providing patients with a safer and more accurate diagnosis and treatment experience, while promoting the optimal allocation of medical resources and establishing an innovative paradigm for the intelligent diagnosis and treatment of gastrointestinal diseases. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0045] Figure 1 It is a module composition diagram of the system of the present invention. DETAILED DESCRIPTION
[0046] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application may be combined with each other. The present application is further described in detail below in conjunction with the drawings and specific embodiments.
[0047] like Figure 1 As shown, the present embodiment provides an image lesion recognition system for assisting gastroenterologists in sampling and examination, the system comprising: an image acquisition module, a preprocessing module, an AI lesion recognition module, a feature analysis module, a sampling recommendation module, a data storage module, a real-time interaction module and a system control center.
[0048] The image acquisition module serves as the data input end of the system, and uses a multispectral endoscope imaging system (Olympus EVISX1 or similar equipment) for image capture. The module integrates three sets of high-sensitivity CMOS sensors, corresponding to white light (400-700nm), narrow-band light (415nm / 540nm blue-green band) and fluorescence imaging (excitation wavelength 635nm) modes. Through the combination of fiber optic bundle transmission and beam splitter prism, multi-modal synchronous acquisition at 30 frames per second is achieved. The core hardware includes: 1. Miniaturized LED array light source, supporting millisecond-level mode switching; 2. Embedded ISP processor performs Bayer interpolation and white balance correction; 3. USB3.2Gen2 interface transmits 12-bit RAW format data.
[0049] Multimodal imaging breaks through the single-modality limitations of traditional endoscopy. White light imaging provides anatomical information, narrow-band light enhances vascular contrast (NBI technology improves the clarity of capillaries on the mucosal surface by 47%), and fluorescence imaging identifies early cancerous tissues through indocyanine green (ICG) labeling. Trimodal data fusion provides complementary features for subsequent AI analysis.
[0050] At the hardware level, a beam splitter is used to decompose the incident light to different sensors, and a time synchronization controller is used to ensure that the spatial registration error of multimodal images is less than 0.1mm. At the software level, a device driver protocol is developed to support DICOM standard metadata embedding, including key information such as shooting timestamp, endoscope insertion depth (obtained through electromagnetic positioning sensor), and light source parameters.
[0051] Preprocessing module, which contains three core units: Image enhancement unit: An improved adaptive gamma correction algorithm is used to establish brightness-contrast mapping tables for different imaging modes (WLI / NBI / FI). For low-light areas, S-curve enhancement (γ=0.5-1.2) is applied, and piecewise linear compression is used for highlight areas. Motion blur repair is achieved through the GAN network (the generator uses the U-Net structure and the discriminator is PatchGAN). The training data set contains 5,000 sets of blur-clear image pairs, and the loss function combines L1 reconstruction loss and perceptual loss (VGG16 feature matching). Noise suppression unit: A hybrid denoising strategy is adopted. The first stage uses the non-local mean (NLM) algorithm, the search window is 21×21, the similarity block is 7×7, and the h parameter is adaptively adjusted according to the noise level (based on image gradient histogram estimation). The second stage uses adaptive median filtering, dynamically adjusts the filter window size (3×3 to 7×7), effectively removes salt and pepper noise while retaining edge details. Actual measurements show that this solution improves the PSNR index by 5.2dB compared with the single filtering method. Image segmentation unit: mucosal area segmentation is achieved based on the U-Net++ architecture, the network input size is 512×512, the encoder uses the ResNet-34 pre-trained model, and the jump connection introduces the attention gate mechanism. The training data is enhanced by the MICCAIEndoVis2018 dataset, and the generalization is improved by random elastic deformation and mirror reflection simulation. The segmentation result is output as a binary mask, and the Dice coefficient reaches 0.923.
[0052] The preprocessing module increases the signal-to-noise ratio of the original image to above 35dB, and the segmentation accuracy meets the requirements of lesion localization, providing standardized input for subsequent AI analysis.
[0053] AI lesion recognition module, this module builds a multimodal large model architecture:
[0054] Visual Transformer backbone network: The ViT-Huge model (632M parameters) is used. The input image is divided into 16×16 blocks, and the global context features are extracted through a 32-layer Transformer encoder. Pre-training uses a self-supervised approach and is optimized on 5 million endoscopic images through contrastive learning (SimCLR framework).
[0055] Adaptive feature fusion module: insert convolutional feature extraction branches at the 8th / 16th / 24th layer of ViT, and use depthwise separable convolution (3×3 kernel) to capture local texture. Dynamically fuse global and local features through the gated attention mechanism, and the calculation formula is: F fused =α·F global +(1-α)·F local , where α is generated by the fully connected network according to the entropy value of the feature map.
[0056] Dynamic weight adjustment submodule: Based on the output of the anatomical site classifier (ResNet-18), the weights of features at different levels are adjusted. For example, esophageal lesions focus on high-frequency features (high-level weight 0.7), while gastric lesions need to combine medium- and low-frequency information (middle-level weight 0.6). This is achieved by modifying the multi-head attention query vector of the Transformer layer.
[0057] Online incremental learning submodule: Deploy the knowledge distillation framework, the teacher model (original ViT-Huge) guides the student model (MobileNetV3), and the detection accuracy is maintained through KL divergence loss (mAP drops <2%). Incremental data is encrypted and transmitted using the federated learning framework, and the elastic weight solidification (EWC) algorithm is used to update parameters to prevent catastrophic forgetting.
[0058] This module achieved 92.4% mAP on the Kvasir-SEG test set, an improvement of 11.6% compared to the single CNN model, and supports simultaneous detection of three types of lesions: polyps (sensitivity 94.2%), ulcers (89.7%), and bleeding spots (86.3%).
[0059] Feature analysis module, which includes a morphological parameter calculation unit: the SLIC superpixel segmentation algorithm is used to generate over-segmented regions, and the precise lesion contour is obtained by regional merging. 12 indicators are calculated, including area (pixel count × 0.01mm² / px), perimeter (chain code method), circularity (4πA / P²), and fractal dimension (box counting method).
[0060] Color heterogeneity analysis unit: Convert the image to HSV space, perform histogram equalization on the H channel, and calculate the entropy, skewness, and kurtosis of the 36-bin histogram. Calculate the local standard deviation map for the S and V channels, and use the Moran's I index to evaluate the color space autocorrelation.
[0061] Texture feature extraction unit: When constructing the gray-level co-occurrence matrix (GLCM), select d=1 pixel, angle 0° / 45° / 90° / 135°, and extract four features: contrast, energy, homogeneity, and correlation. Calculate the LBP (local binary pattern) histogram in a 5×5 sliding window, and use the χ² distance to evaluate the texture difference.
[0062] 3D reconstruction unit: The binocular vision module generates point clouds through disparity maps and uses Poisson surface reconstruction algorithm to generate 3D models. The volume calculation uses voxel projection method with an accuracy of 0.1mm³. The infiltration depth is obtained by fitting the maximum penetration distance in the normal direction of the mucosal surface.
[0063] Malignant risk prediction unit: Building an XGBoost classifier (n estimators =500, max depth=6), the input includes 32-dimensional image features and 8-dimensional clinical data (age, CEA level, etc.). SHAP value is used to perform feature importance analysis and output the probability of canceration and confidence interval.
[0064] This module generates a comprehensive report containing 68 quantitative indicators, with an AUC for malignancy prediction of 0.891, an improvement of 27% over traditional doctor's visual assessment.
[0065] In the clinical diagnosis of gastroenterology, accurate sampling is the key to obtaining accurate pathological results. Different types of lesions have different degrees, locations and characteristics. Traditional sampling methods often rely on the doctor's experience, which is subjective and blind, and may lead to insufficient or inaccurate sampling, thus affecting the accuracy of diagnosis and the formulation of subsequent treatment plans. The emergence of the sampling recommendation module aims to use advanced technical means, combined with a large amount of clinical data and scientific algorithms, to provide doctors with objective and scientific sampling suggestions, improve the success rate and efficiency of sampling, reduce the occurrence of misdiagnosis and missed diagnosis, and buy precious time for patients' treatment.
[0066] Reinforcement learning strategy parameters: The core of reinforcement learning is that the agent continuously learns the optimal strategy through interaction with the environment. In the sampling recommendation module, the agent's reward function needs to be designed based on a variety of factors. For example, a higher positive reward is given for successfully obtaining representative tissue samples, while a negative reward is given for behaviors that cannot be clearly diagnosed due to improper sampling locations or insufficient quantities. The learning rate is usually set between 0.01-0.1. This parameter controls the step size of the agent's strategy update in each iteration. Too large a learning rate may cause the strategy to be unstable, while too small a learning rate will make the learning process too slow. The discount factor is generally between 0.9-0.99. It is used to balance the importance of current rewards and future rewards. A higher discount factor means more emphasis on long-term rewards.
[0067] Probabilistic graph model parameters: The probabilistic graph model is used to infer the optimal sampling location and quantity based on the lesion type, ESD / EMR treatment guidelines, and historical sampling success rate data. When constructing a probabilistic graph model, it is necessary to determine the conditional probability distribution between nodes. For example, for different types of lesions, the probability distribution of their occurrence in different parts of the digestive tract is different, and these probability distributions can be obtained through a large amount of clinical data statistics. At the same time, it is also necessary to consider the requirements of different treatment guidelines for sampling locations and quantities, and incorporate this information into the probabilistic graph model to improve the accuracy of recommendations.
[0068] First, the system collects a large amount of historical case data, including lesion type, location, sampling scheme, and corresponding pathological results. These data are used to train reinforcement learning agents and build probabilistic graph models. In practical applications, when new lesion images and related information are input, the probabilistic graph model will preliminarily infer the possible sampling locations and quantity ranges based on existing knowledge and data. Then, the reinforcement learning agent selects an action (i.e., specific sampling location and quantity recommendations) based on the current state of the environment (i.e., lesion information and preliminary inferences from the probabilistic graph model), interacts with the environment, and adjusts its strategy based on the feedback rewards. Through continuous iterative learning, the agent gradually finds the optimal sampling recommendation plan.
[0069] The biomechanical simulation unit predicts tissue deformation at different sampling depths based on the finite element analysis method. Before performing finite element analysis, a mechanical model of the tissue needs to be established, which usually involves measuring and modeling the material properties of the tissue (such as elastic modulus, Poisson's ratio, etc.). For digestive tract tissue, the elastic modulus is generally between 1-10kPa, and the Poisson's ratio is about 0.4-0.5. By inputting these parameters into the finite element analysis software, the force applied by the biopsy forceps to the tissue at different depths is simulated to obtain the deformation of the tissue. At the same time, the system will combine the vascular distribution map to identify areas with dense blood vessels, and automatically avoid these areas when recommending sampling locations to reduce the occurrence of complications such as bleeding.
[0070] The historical case matching unit uses a graph database to store and retrieve sampling plans and postoperative pathological results of similar lesions. The nodes in the graph database can represent entities such as cases, lesion characteristics, sampling plans, and edges represent the relationship between them. When encountering a new case, the system will search the graph database based on the characteristics of the current lesion (such as size, shape, location, etc.) to find similar historical cases. By comparing the sampling plans and pathological results of these similar cases, more accurate sampling recommendations are provided for the current case. For example, if it is found that multiple similar cases have obtained clear diagnostic results after sampling at a specific location, the system will use that location as one of the recommended sampling locations.
[0071] In the medical field, the secure storage and effective management of data are of vital importance. The image lesion recognition system of the Department of Gastroenterology generates a large amount of sensitive data, including the original images of patients, lesion annotation data, feature analysis results, and sampling records. These data are not only of great significance for the diagnosis and treatment of current cases, but can also be used for medical research and clinical experience summary. The data storage module uses advanced technical means to ensure the security, integrity and traceability of data, while supporting efficient data retrieval and comparative analysis, providing strong data support for doctors and researchers.
[0072] Blockchain encryption parameters: Blockchain uses hash algorithms to encrypt data, and commonly used hash algorithms include SHA-256. The length of the hash value is 256 bits. It is unique and irreversible, and can ensure that the data is not tampered with during transmission and storage. At the same time, the consensus mechanism of the blockchain (such as proof of work) needs to set certain difficulty parameters to ensure the security and stability of the blockchain network. The difficulty parameter determines how many hash operations miners need to perform to find a valid block. The higher the difficulty, the stronger the security of the network, but it will also increase the time for transaction confirmation.
[0073] Homomorphic encryption parameters: Homomorphic encryption technology is used to process metadata containing patient identity information. When selecting a homomorphic encryption algorithm, it is necessary to consider the security and computational efficiency of the algorithm. For example, the Paillier homomorphic encryption algorithm is a commonly used additive homomorphic encryption algorithm. Its key length is usually 2048 bits, which can support addition operations on encrypted data while ensuring data privacy.
[0074] JPEG2000 lossless compression parameters: When performing JPEG2000 lossless compression on medical images, some parameters need to be set to control the compression ratio and image quality. For example, the compression ratio can be adjusted according to actual needs, generally between 2:1-10:1. At the same time, the original bit depth of the lesion area needs to be preserved to ensure that important information is not lost in subsequent analysis and diagnosis.
[0075] The system packages data such as original images, lesion annotation data, feature analysis results, and sampling records into data blocks and calculates the hash value of each data block. Each data block contains the hash value of the previous data block, forming a chain structure. When new data needs to be stored, the system will add it to the end of the blockchain and ensure the consistency and security of the data through a consensus mechanism (such as proof of work). Only authorized nodes can read and write to the blockchain, thereby ensuring the privacy and security of the data.
[0076] The privacy protection unit uses homomorphic encryption technology to encrypt metadata containing patient identity information. During the data collection phase, the system encrypts the patient's identity information and then stores the encrypted data in the blockchain. When the data needs to be used, only authorized personnel can decrypt it using the corresponding private key. Homomorphic encryption technology allows specific calculations to be performed on encrypted data without first decrypting the data, thereby ensuring data availability while protecting the patient's privacy to the greatest extent possible.
[0077] The intelligent compression unit performs JPEG2000 lossless compression on medical images. During the compression process, the system first analyzes the image and identifies the lesion area. Then, according to the importance and characteristics of the lesion area, the compression parameters are adjusted to achieve a higher compression ratio while ensuring the original bit depth of the lesion area. The compressed image data will be stored in the blockchain, while retaining the metadata of the original image for subsequent retrieval and analysis.
[0078] The data storage module supports multi-condition retrieval and comparative analysis based on case ID. Users can enter multiple conditions such as case ID, lesion type, sampling time, etc., and the system will quickly locate the data that meets the conditions in the blockchain. At the same time, the system can also conduct comparative analysis on the data of different cases, such as comparing the characteristic differences of different types of lesions, the effects of different sampling schemes, etc. Through the visual interface, users can intuitively view the comparative analysis results, providing strong support for clinical decision-making and research.
[0079] The real-time interactive module provides doctors with an intuitive and convenient way of operation, enhancing the interactivity between doctors and the system. During the sampling and examination process of gastroenterology, doctors need to understand the three-dimensional structure, location and recommended sampling path of the lesion in real time in order to operate accurately. Through the AR display unit and tactile feedback unit, doctors can perceive the lesion information more intuitively, improve the accuracy and safety of the operation, and reduce the operation time and the pain of the patient.
[0080] AR display unit parameters: The head-mounted device of the AR display unit needs to have the characteristics of high resolution and low latency to provide clear and smooth image display. Generally speaking, the resolution of the head-mounted device should not be less than 2K (2560×1440), and the delay time should be controlled within 20 milliseconds. At the same time, the field of view of the AR display should meet the doctor's operating needs, usually between 60°-90°. In the multi-view collaborative display function, the ratio of the main operation screen and the local magnified screen of the lesion can be adjusted according to the doctor's habits. Generally, the main operation screen accounts for 70%-80%, and the local magnified screen of the lesion accounts for 20%-30%.
[0081] Tactile feedback unit parameters: The vibration intensity and frequency of the tactile feedback unit need to be adjusted according to the actual situation. The vibration intensity is generally between 0.5-2.0g, and the frequency is between 100-300Hz. When the biopsy forceps approaches the recommended sampling position, the vibration intensity and frequency will gradually increase to attract the doctor's attention.
[0082] The AR display unit realizes multi-view collaborative display function through a head-mounted device. The system will render the main operation screen and the local magnified screen of the lesion separately, and project them into the doctor's field of view through the optical system. During the rendering process, the system will adjust the display ratio and position of the screen in real time according to the doctor's perspective and operation requirements. For example, when the doctor needs to observe the lesion in more detail, he can enlarge the local screen of the lesion through gestures or voice commands. At the same time, the system will ensure the synchronization of information between the main operation screen and the local magnified screen of the lesion to avoid screen delays or misalignment.
[0083] The dynamic navigation line generation function adjusts the display density and transparency of the sampling path in real time according to the advancement speed of the endoscope. The system monitors the advancement speed of the endoscope in real time through sensors. When the advancement speed is faster, the display density of the navigation line will decrease and the transparency will increase to avoid excessive information interfering with the doctor's vision; when the advancement speed is slower, the display density of the navigation line will increase and the transparency will decrease so that the doctor can see the sampling path more clearly. The color and style of the navigation line can be set according to the doctor's preferences. Generally, bright colors (such as red and green) and simple line styles are used to improve recognition.
[0084] The system generates a 3D reconstruction model of the lesion based on binocular endoscopic images or structured light projection data. When generating the 3D reconstruction model, steps such as feature extraction, matching, and triangulation are required to ensure the accuracy and precision of the model. Then, the generated 3D reconstruction model and sampling path guidance are projected into the field of view of the headset. By observing the projection image in the headset, the doctor can intuitively understand the 3D structure of the lesion and the recommended sampling path, so as to operate more accurately.
[0085] The tactile feedback unit monitors the distance between the biopsy forceps and the recommended sampling position in real time through the sensor installed on the biopsy forceps. When the biopsy forceps approaches the recommended sampling position, the sensor transmits a signal to the control system, which adjusts the vibration intensity and frequency of the vibration motor according to the preset parameters. The vibration motor provides tactile feedback to the doctor by generating vibrations of different intensities and frequencies, so that the doctor can understand the position of the biopsy forceps in time without relying on vision, thereby improving the accuracy and safety of sampling.
[0086] The system control center is the core hub of the entire image lesion recognition system, responsible for coordinating the data flow between modules to ensure the efficient operation of the system. At the same time, by deploying the federated learning framework, the system control center can achieve collaborative optimization of models among multiple medical institutions, make full use of the data resources of various medical institutions, improve the accuracy and generalization ability of the model, and provide stronger support for the diagnosis and treatment of gastroenterology.
[0087] Edge computing node parameters: Edge computing nodes need to have certain computing power and storage capacity to meet the needs of data processing and caching. Generally speaking, the CPU frequency of edge computing nodes should not be less than 2.0GHz, the memory capacity should not be less than 8GB, and the storage capacity should not be less than 128GB. At the same time, the network bandwidth of edge computing nodes should meet the requirements of data transmission, generally above 100Mbps.
[0088] Federated learning framework parameters: In the federated learning framework, some parameters need to be set to control the model training process. For example, the learning rate is usually set between 0.001-0.01, which controls the step size of the model to update the parameters in each iteration. The number of global iterations is adjusted according to the size of the dataset and the complexity of the model, generally between 10-100 times. At the same time, an aggregation strategy, such as the FedAvg algorithm, needs to be set to aggregate and optimize the model parameters uploaded by various medical institutions.
[0089] GPU memory usage threshold: The resource dynamic allocation unit automatically switches the model precision mode according to the GPU memory usage. Generally speaking, when the GPU memory usage exceeds 80%, the system automatically switches to low-precision mode; when the memory usage is less than 30%, the system switches to high-precision mode.
[0090] The system control center coordinates the data flow of each module through the edge computing node. The raw data obtained by the image acquisition module is first transmitted to the edge computing node, which performs preliminary processing on the data, such as data cleaning and format conversion. Then, the processed data is distributed according to the needs of each module to ensure that the data can be transmitted to the corresponding module in a timely and accurate manner for further analysis and processing. At the same time, the edge computing node can also cache the data to reduce the delay of data transmission and improve the response speed of the system.
[0091] Under the federated learning framework, each medical institution uses its own data set to train the model locally. After the training is completed, each medical institution uploads the trained model parameters to the system control center. The system control center uses aggregation strategies (such as the FedAvg algorithm) to aggregate and optimize these parameters to generate a better global model. Then, this global model is distributed to each medical institution, and each medical institution continues to train the global model using its local data set. Through continuous iterative training, the performance of the model continues to improve, while protecting the data privacy of each medical institution.
[0092] The resource dynamic allocation unit monitors the GPU memory usage in real time. When the GPU memory usage exceeds the preset threshold, the system will automatically switch to low-precision mode to reduce the model's usage of video memory. In low-precision mode, the calculation accuracy of the model will be reduced, but the calculation speed can be significantly improved and the video memory usage can be reduced. When the video memory usage is lower than another preset threshold, the system will switch to high-precision mode to improve the accuracy of the model. This dynamic resource allocation method can fully utilize the computing power of the GPU while ensuring system performance.
[0093] The multimodal data synchronization unit ensures that the timestamps of the endoscopic video stream, the patient's vital signs data, and the lesion annotation information are aligned. The system adds timestamp information to each data sample during the data acquisition phase, and then synchronizes the data of different modalities through a timestamp matching algorithm during the data processing process. For example, when a frame image in the endoscopic video stream corresponds to a moment in the patient's vital signs data, the system will associate the two data samples to ensure that when the doctor views the data, he can accurately understand the time relationship between different data and improve the accuracy of diagnosis.
[0094] In gastroenterology examinations, the disinfection of endoscopic instruments is a key step in preventing cross infection. Endoscopic instruments that are not thoroughly disinfected may carry various pathogens, such as bacteria and viruses, which pose serious health risks to patients. The disinfection monitoring module tracks the disinfection status of endoscopic instruments in real time and displays a sterilization validity period warning on the interface, helping medical staff to understand the disinfection status of the instruments in a timely manner, ensuring that the endoscopic instruments used meet hygiene standards and protect patient safety.
[0095] RFID tag parameters: RFID tags need to have a certain storage capacity and read-write distance. Generally speaking, the storage capacity of RFID tags should not be less than 128 bytes to store basic information of the device and disinfection records. The read-write distance is selected according to the actual application scenario, generally between 1-5 meters. At the same time, the operating frequency of the RFID tag also needs to be selected according to the specific situation. Common operating frequencies are 13.56MHz and 860-960MHz.
[0096] Sterilization validity period threshold: The system will set the sterilization validity period threshold according to different disinfection methods and instrument types. For example, for endoscopic instruments sterilized by high temperature and high pressure, the sterilization validity period is generally 7-14 days; for instruments sterilized by immersion in chemical disinfectants, the sterilization validity period may be shorter, generally 24-48 hours.
[0097] The disinfection monitoring module tracks the disinfection status of endoscopic instruments through radio frequency identification (RFID) technology. Each endoscopic instrument is equipped with an RFID tag, which stores the basic information of the instrument (such as model, number, etc.) and disinfection records. When the instrument is disinfected, the disinfection equipment will write the disinfection information (such as disinfection time, disinfection method, etc.) into the RFID tag through the RFID reader. At the same time, the system will read the information in the tag in real time through the RFID reader installed in the ward or operating room to update the disinfection status of the instrument. When the sterilization validity period of the instrument is approaching, the system will display a warning message on the interface to remind medical staff to disinfect or replace the instrument in time.
[0098] In practical applications, the power consumption of the system is an important factor to consider. If the system power consumption is too high, it will not only increase the cost of use, but may also cause serious heating of the equipment, short battery life and other problems, affecting the normal use of the system. The power consumption optimization module uses advanced technical means to reduce the power consumption of the system and improve the endurance of the equipment, so that the system can work continuously when powered by batteries to meet the needs of clinical use.
[0099] Neural architecture search parameters: Neural architecture search (NAS) technology requires setting some parameters to control the search process. For example, the size of the search space determines the range of neural network architectures that can be searched. Generally, the search space can be defined by setting parameters such as the number of network layers and the size of the convolution kernel. The choice of search algorithm will also affect the search efficiency and results. Common search algorithms include random search, genetic algorithm, and reinforcement learning algorithm. The number of search iterations is adjusted according to the size of the search space and computing resources, generally between 100-1000 times.
[0100] Battery life target: The goal of the power optimization module is to support 8 hours of continuous operation under battery power. In order to achieve this goal, it is necessary to optimize the power consumption of various components of the system, including the processor, memory, display, etc.
[0101] The power optimization module uses neural architecture search (NAS) technology to generate a lightweight inference model on the device side. First, a search space is defined, containing various possible neural network architectures. Then, a search algorithm is used to search in the search space, and each architecture is evaluated according to preset evaluation indicators (such as accuracy, computational complexity, etc.). During the search process, the parameters of the architecture are continuously adjusted until the neural network architecture with the best performance and the lowest complexity is found. The generated lightweight inference model has low computational complexity and power consumption while ensuring a certain degree of accuracy. At the same time, the system will also optimize the power consumption of other components, such as using low-power processors, optimizing memory management, etc., to further reduce the power consumption of the system and achieve the goal of 8 hours of continuous operation under battery power.
[0102] Reasonable hardware architecture design has an important impact on the performance, scalability and stability of the system. The design of detachable computing units, redundant power modules and modular mechanical structures can improve the efficiency of system upgrades and maintenance, ensure that the system can operate stably under various conditions, and provide a solid hardware foundation for the normal operation of the image lesion recognition system.
[0103] PCIe interface parameters: The detachable computing unit can quickly replace and upgrade the GPU accelerator card through the PCIe interface. The version and bandwidth of the PCIe interface will affect the performance of the GPU accelerator card. Currently, the common PCIe interface versions are PCIe3.0 and PCIe4.0. The bandwidth of PCIe3.0 is 8GT / s, and the bandwidth of PCIe4.0 is 16GT / s. Choosing the appropriate PCIe interface version can give full play to the performance of the GPU accelerator card.
[0104] Redundant power supply module parameters: The redundant power supply module adopts a dual-power supply design and a supercapacitor instantaneous power-off protection circuit. The power of the dual power supply needs to be selected according to the total power demand of the system. Generally speaking, the power of each power supply should not be less than 50% of the total power of the system. The capacity and withstand voltage of the supercapacitor also need to be selected according to the needs of the system to ensure that sufficient power support can be provided to the system when the power supply is suddenly interrupted.
[0105] Magnetic quick-release interface parameters: The modular mechanical structure connects the endoscope host and auxiliary equipment through a magnetic quick-release interface. The suction force of the magnetic quick-release interface needs to be moderate, both to ensure the stability of the connection and to facilitate disassembly. Generally speaking, the suction force of the magnetic quick-release interface is between 10-50N.
[0106] The detachable computing unit is connected to the GPU accelerator card through the PCIe interface. When the GPU accelerator card needs to be upgraded, the user only needs to turn off the system power, open the chassis, unplug the old GPU accelerator card from the PCIe interface, and then insert the new GPU accelerator card into the PCIe interface and secure it. The PCIe interface has the characteristics of high bandwidth and low latency, which can ensure high-speed data transmission between the GPU accelerator card and other components of the system. This detachable design allows the system to be continuously upgraded with the development of technology and improve the computing performance of the system.
[0107] The redundant power supply module adopts a dual-power supply design, and each power supply can independently power the system. When one power supply fails, the other power supply will automatically take over the power supply task to ensure the normal operation of the system. At the same time, the supercapacitor instantaneous power failure protection circuit plays a role when the power is suddenly interrupted. The supercapacitor will release the stored energy in an instant to provide short-term power support for the system, so that the system has enough time to save data and shut down safely, avoiding data loss and system damage.
[0108] The modular mechanical structure connects the endoscope host and the auxiliary device through a magnetic quick-release interface. When assembling the device, the user only needs to align the magnetic quick-release interfaces of the endoscope host and the auxiliary device to achieve a quick connection. The magnetic quick-release interface has the functions of automatic alignment and adsorption to ensure the accuracy and stability of the connection. When the device needs to be maintained or replaced, the user only needs to gently pull the magnetic quick-release interface to disassemble the device, which is convenient and quick. This modular design improves the maintenance efficiency and scalability of the equipment.
[0109] In summary, the various modules and hardware architectures of the image lesion recognition system that assists gastroenterologists in sampling and examination work together to provide comprehensive, efficient and accurate support for the clinical diagnosis and treatment of gastroenterology. By continuously optimizing and improving these modules and architectures, it is expected that the performance and application value of the system will be further improved, bringing better medical experience to patients.
[0110] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. An image lesion recognition system for assisting gastroenterologists in sampling and examination, characterized in that: The system includes: image acquisition module, preprocessing module, AI lesion recognition module, feature analysis module, sampling recommendation module, data storage module, real-time interaction module and system control center; The image acquisition module acquires digestive tract endoscopic image data in real time through an endoscopic device, including white light imaging, narrow-band light imaging and fluorescence imaging multimodal images, and transmits them to the preprocessing module; The preprocessing module includes an image enhancement unit, a noise suppression unit and an image segmentation unit, wherein the noise suppression unit adopts a non-local mean denoising algorithm combined with an adaptive median filter, and the image segmentation unit realizes a preliminary separation of the digestive tract mucosal area and the lesion based on a U-Net++ network; The AI lesion recognition module is driven by a multimodal large model, including a visual Transformer backbone network and an adaptive feature fusion module. The backbone network extracts global image features through a pre-trained ViT-Huge model, and the fusion module combines local features of the convolutional neural network to output category labels and lesion bounding box coordinates for polyps, ulcers, and bleeding points. The feature analysis module performs multi-dimensional quantitative analysis on the identified lesions, including a morphological parameter calculation unit based on superpixel segmentation, an HSV color space histogram analysis unit, and a gray-level co-occurrence matrix texture feature extraction unit, and generates a comprehensive report including lesion size, shape irregularity, color heterogeneity, and texture complexity.
2. The image lesion recognition system for assisting gastroenterologists in sampling and examination according to claim 1 is characterized in that: The sampling recommendation module integrates a reinforcement learning strategy, dynamically generates the optimal sampling location and quantity through a probabilistic graph model according to the lesion type, ESD / EMR treatment guidelines and historical sampling success rate data, and outputs a visual heat map superimposed on the endoscopic image; The data storage module uses blockchain encryption technology to store original images, lesion annotation data, feature analysis results and sampling records, and supports multi-condition retrieval and comparative analysis based on case ID; The real-time interaction module includes an AR display unit and a tactile feedback unit. The AR display unit projects a three-dimensional reconstruction model of the lesion and a sampling path guide through a head-mounted device. The tactile feedback unit generates a vibration prompt when the biopsy forceps approaches the recommended sampling position. The system control center coordinates the data flows of each module through edge computing nodes, and deploys a federated learning framework to achieve collaborative optimization of models among multiple medical institutions; The AI lesion recognition module also includes: The dynamic weight adjustment submodule automatically adjusts the attention weights of different feature layers according to the classification results of the digestive tract anatomical parts. The lesion detection in the esophagus, gastric body, and duodenum corresponds to different feature fusion strategies respectively. The online incremental learning submodule uses doctors’ revised annotation data to update model parameters in real time and compresses the model through knowledge distillation technology to adapt to embedded device deployment.
3. The image lesion recognition system for assisting gastroenterologists in sampling and examination according to claim 1 is characterized in that: The feature analysis module further comprises: The 3D reconstruction unit generates a point cloud model of the lesion based on binocular endoscopic images or structured light projection data, and calculates the volume, surface area, and invasion depth parameters; The malignancy risk prediction unit inputs the morphological features and the patient's electronic medical record data into the XGBoost classifier and outputs the early cancer probability score.
4. The image lesion recognition system for assisting gastroenterologists in sampling and examination according to claim 1, characterized in that: The AR display unit of the real-time interaction module realizes: Multi-view collaborative display function, supporting picture-in-picture display of the main operation screen and the local magnified screen of the lesion; Dynamic navigation line generation function adjusts the display density and transparency of the sampling path in real time according to the endoscope advancement speed.
5. The image lesion recognition system for assisting gastroenterologists in sampling and examination according to claim 1, characterized in that: The sampling recommendation module comprises: Biomechanical simulation unit, based on finite element analysis, predicts tissue deformation at different sampling depths to avoid areas with dense blood vessels; The historical case matching unit retrieves the sampling schemes and postoperative pathological results of similar lesions through the graph database.
6. The image lesion recognition system for assisting gastroenterologists in sampling and examination according to claim 1, characterized in that: The image enhancement unit of the preprocessing module adopts: Adaptive gamma correction algorithm dynamically adjusts the contrast enhancement curve according to different imaging modes; The low-quality image restoration function implemented by the generative adversarial network eliminates motion blur and mirror reflection artifacts.
7. The image lesion recognition system for assisting gastroenterologists in sampling and examination according to claim 1 is characterized in that: The system control center also integrates: Dynamic resource allocation unit, automatically switches model accuracy mode according to GPU memory occupancy; The multimodal data synchronization unit ensures the alignment of timestamps of the endoscopic video stream, patient vital signs data, and lesion annotation information.
8. The image lesion recognition system for assisting gastroenterologists in sampling and examination according to claim 1, characterized in that: The data storage module further comprises: The privacy protection unit uses homomorphic encryption technology to process metadata containing patient identity information; Intelligent compression unit, which performs JPEG2000 lossless compression on medical images while retaining the original bit depth of the lesion area.
9. The image lesion recognition system for assisting gastroenterologists in sampling and examination according to claim 1, characterized in that: The system also includes: The disinfection monitoring module tracks the disinfection status of endoscopic instruments through radio frequency identification technology and displays a warning of the sterilization validity period on the interface; The power consumption optimization module uses neural architecture search technology to generate a lightweight inference model on the device side, supporting 8 hours of continuous operation under battery power.
10. The image lesion recognition system for assisting gastroenterologists in sampling and examination according to claim 1, characterized in that: The hardware architecture of the system includes: Removable computing unit, quickly replace and upgrade GPU accelerator card through PCIe interface; Redundant power supply module, adopting dual-circuit power supply design and supercapacitor instantaneous power-off protection circuit; Modular mechanical structure, connecting the endoscope host and auxiliary equipment through a magnetic quick-release interface.
Citation Information
Patent Citations
Intelligent gastrointestinal endoscope operation auxiliary system and method
CN119028538A
Intelligent medical film reading method, apparatus, and device, and storage medium
WO2023178972A1