Multifunctional computer-aided gastroscopy system and method adopting an optimized integrated AI solution

By designing a multifunctional computer-assisted gastroscopy system optimized by integrated AI solutions, the gastroscopy problem that is difficult to achieve low latency and high accuracy in the prior art is solved, and the effect of efficient detection and classification under different hardware configurations is achieved.

CN115004316BActive Publication Date: 2025-05-27HONG KONG APPLIED SCI & TECH RES INST
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202280001595.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-04-25
Filing Date
2022-05-05
Publication Date
2025-05-27
Estimated Expiration
2042-05-05

AI Technical Summary

Technical Problem

It is difficult to develop a multifunctional gastroscopy system that can achieve low latency and high accuracy. The system needs to process 4K video streams, detect different lesions, cancers, Helicobacter pylori (HP) infections, and operate under different hardware configurations.

Method used

A multifunctional computer-assisted gastroscopy system optimized by integrated AI schemes is designed, including an AI image processing system, through collaborative execution of multiple architectural-level modules of image quality assessment, lesion detection, cancer identification, HP classification and site recognition, using neural models to extract and share information, improve classification accuracy, and optimize subtask allocation through coprocessors to reduce response delays.

Benefits of technology

It realizes low-latency response and high-precision detection and classification, meets the detection and classification performance standards of medical professionals, and runs on different hardware platforms, supporting offline and online processing modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115004316B_ABST
    Figure CN115004316B_ABST
Patent Text Reader

Abstract

A multifunctional computer-assisted gastroscopy system with an optimized integrated AI solution is disclosed. The system utilizes multiple deep learning neural models to achieve low latency and high performance requirements for multiple tasks. The optimization is performed at three levels: architectural, modular, and functional. At the architectural level, the model is designed in such a way that it can simultaneously complete HP infection classification and detection of several lesions with a single inference, thereby reducing the computational cost. At the modular level, the site recognition model is optimized using temporal information as a sub-model of HP infection classification. It not only improves the performance of HP infection classification, but also plays an important role in lesion detection and program state determination. At the functional level, inference latency is minimized through configuration and resource-aware optimization. Also at the functional level, preprocessing is accelerated through image resizing parallelization and unified preprocessing.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority application

[0002] This application claims priority to U.S. Non-Provisional Application No. 17 / 660,442, filed on April 25, 2022, which is incorporated herein by reference. Technical Field

[0003] The present invention generally relates to a computer-assisted gastroscopy system, and more particularly to a multifunctional computer-assisted gastroscopy system optimized using its integrated AI solution. Background Art

[0004] It has recently been reported that deep learning-based technologies are very beneficial in the field of endoscopy. According to some trials, the detection rate can be improved by about 50%, while the costs associated with endoscopy can be reduced by 7-20%. Although several commercial products with AI functions have emerged in the endoscopy industry, many of them still have various limitations. There are still many challenges in developing a multifunctional gastroscopy system that can achieve low latency and high accuracy, process 4K video streams from the latest gastroscopic instruments, and can simultaneously detect different lesions, cancers, Helicobacter pylori (HP) infections and can run on different hardware configurations. Therefore, one object of the present invention is to develop a multifunctional, tightly integrated gastroscopy system optimized using AI solutions. Summary of the invention

[0005] In view of the above background, the present invention provides a multifunctional computer-assisted gastroscopy system optimized by its integrated AI solution.

[0006] Therefore, an exemplary embodiment of the present invention provides a computer-assisted gastroscopy system, comprising: a central processing unit coupled to a memory storing an executable software program, wherein the software program comprises: an AI image processing system for analyzing a gastric image sequence obtained from a gastroscopy instrument. The AI ​​image processing system comprises at least three architecture-level modules that collaboratively perform image quality assessment, lesion detection and cancer identification, HP classification, and site identification, wherein at least one of the modules comprises one or more neural models, each neural model extracting different but related information from the gastric image sequence and sharing the information extracted from the gastric image sequence with other modules; and at least one of the neural models fuses HP infection features with site information extracted from other neural models to improve the classification accuracy of the computer-assisted gastroscopy system.

[0007] Another exemplary embodiment of the present invention provides a method for processing a gastric image sequence by a computer-assisted gastroscopy system, comprising: acquiring the gastric image sequence from a gastroscopy instrument; analyzing the gastric image sequence by an AI image processing system including at least three architecture-level modules that collaboratively perform image quality assessment, lesion detection and cancer identification, HP classification, and lesion site recognition, wherein at least one of the modules includes one or more neural models, each neural model extracts different but related information from the gastroscopy image sequence and shares the information extracted from the gastric image sequence with other modules; and at least one of the neural models extracts HP infection features extracted from other neural network models. The computer-assisted gastroscopy system is designed to integrate feature and location information to improve the classification accuracy of the computer-assisted gastroscopy system; each of the at least three modules creates a subtask list for execution by the computer-assisted gastroscopy system; and when the computer-assisted gastroscopy system also includes at least one coprocessor, and a software program executed at a central processor unit of the computer-assisted gastroscopy system wisely allocates subtasks to the at least one coprocessor according to the pre-assigned priority of the subtasks and the capacity and capability of each of the coprocessors, the response delay to the user is reduced, so that the computer-assisted gastroscopy system can achieve the requirements of low-latency response and high-precision detection and classification.

[0008] The above exemplary embodiments have benefits and advantages over conventional techniques. For example, the disclosed computer-assisted gastroscopy system can not only meet all detection and classification performance criteria proposed by medical professionals, but also can run on different hardware platforms with various computing capabilities to achieve low-latency response. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other features, advantages and aspects of the embodiments of the present invention will become more apparent through the following detailed description with reference to the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0010] Figure 1 A block diagram of a computer-assisted gastroscopy system according to an embodiment of the present disclosure is shown;

[0011] Figure 2 Some embodiments of the present disclosure are shown Figure 1 Block diagram of AI functions and image processing modules in the computer-assisted gastroscopy system;

[0012] Figure 3 The training process of various neural models according to some embodiments of the present disclosure is shown;

[0013] Figure 4shows an HP feature table for HP classification according to one embodiment of the present disclosure;

[0014] Figure 5 The composite neural model of the third module and how to train it according to one embodiment of the present disclosure are shown;

[0015] Figure 6 shows a region label for stomach region identification according to an embodiment of the present disclosure;

[0016] Figure 7 illustrates various components of a part recognition model in one embodiment of the present invention and how to train it;

[0017] Figure 8 A method of creating a pseudo video sequence in an exemplary embodiment of the present disclosure is illustrated;

[0018] Fig. 9 An exemplary example of training a neural model for part recognition according to an embodiment of the present disclosure is shown.

[0019] Fig.10 Exemplary experimental results on cumulative position change and frame index according to one embodiment of the present disclosure are shown.

[0020] Fig.11 Two different pre-processing methods according to some embodiments of the present disclosure are shown.

[0021] Fig.12 An example of a modified neural architecture according to an exemplary embodiment of the present disclosure is shown.

[0022] Fig.13 A flow chart for preparing an image for resizing according to one embodiment of the present disclosure is shown.

[0023] Fig.14 A method of performing parallel resizing on a stomach image according to one embodiment of the present disclosure is shown.

[0024] Fig.15 It illustrates how a computer-assisted gastroscopy system according to one embodiment of the present disclosure can detect available hardware configurations and fully utilize their computing capabilities.

[0025] Fig.16 The dynamic batch inference process according to one embodiment of the present disclosure is shown.

[0026] Fig.17 A neural model prioritization according to one embodiment of the present disclosure is shown.

[0027] Fig.18A timing diagram of a bounded-length double-ended queue operation according to one embodiment of the present disclosure is illustrated.

[0028] Fig.19 The figure shows the neural reasoning execution flow in the online operation mode according to one embodiment of the present disclosure.

[0029] Fig. 20 It is shown how model selection is performed in an exemplary embodiment of the present disclosure.

[0030] Fig.21 A hardware schematic diagram of a computer-assisted gastroscopy system in an exemplary embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0031] As used herein and in the claims, "comprising" means including the following elements but not excluding other elements. The term "based on" should be understood as "based at least in part on". The terms "one example embodiment" and "example embodiment" should be understood as "at least one example embodiment". The term "another embodiment" should be understood as "at least one other embodiment".

[0032] As used herein and in the claims, a "module" by itself generally refers to a main software component in the software, unless otherwise specified.

[0033] As used herein and in the claims, a "neural model" refers to a neural network with a pre-specified neural architecture. A "neural architecture" refers to a specific configuration of interconnections between nodes at different layers of a neural network.

[0034] As used herein and in the claims, a "tensor" refers to a multidimensional mathematical object. For example, a [256x256x3] tensor represents a three-dimensional array whose first and second dimensions are 256 and the third dimension is 3. The last dimension is also called the "channel" of the tensor.

[0035] As used herein and in the claims, an "image" refers to a digital image having a plurality of pixels arranged in a two-dimensional array having a certain height and width. A "video" is a series of images arranged in a certain order. An image in a video is also referred to as a "frame". An image with a label representing a specific attribute of the image is called a "tagged image", and a "pseudo-video" is a collection of tagged images arranged in a specific order to mimic a video. In this specification, the terms "video" and "image sequence" are used interchangeably, both referring to an ordered sequence of images.

[0036] The present invention proposes a computer-assisted gastroscopy system. Gastroscopy is a medical procedure that involves inserting a thin, flexible tube called an endoscope through the patient's mouth to examine the internal conditions of the esophagus, stomach, and duodenum. The tip of the flexible tube is equipped with a camera and a light source. The camera captures video throughout the gastroscopy procedure, which is then reviewed by a medical expert or a computer-assisted system to check for any abnormal growths or lesions within the gastrointestinal tract.

[0037] Stomach videos can check whether the patient is infected with Helicobacter pylori (HP) virus or whether there are polyps, ulcers or cancerous tumors in the stomach. HP infection is the main cause of HP infection-related gastritis, and if it is not paid attention to, it may eventually develop into gastric cancer. Therefore, early diagnosis of HP infection is important. In addition to HP infection, stomach videos can also review other types of abnormal growth or lesions in the stomach. Therefore, it is advantageous to develop a computer-assisted gastroscopy system that can help medical professionals analyze stomach videos and report the results. In recent years, AI technology based on deep learning technology has been proven to be successful in various medical image analysis applications. In the present invention, a neural algorithm based on deep learning is developed to help endoscopists perform routine screening for HP infection and other gastric diseases to improve diagnostic efficiency and accuracy.

[0038] There are several challenges in developing such a system. First, the system must meet a set of performance standards specified by the medical community. These requirements are usually specified as a set of performance targets, including but not limited to disease classification accuracy, detection sensitivity and specificity, etc. Second, the system should preferably be able to process the gastric video stream in real time so that medical professionals can view the results while performing the gastroscopy procedure. Third, the system needs to perform multiple diagnostic tasks on the same video stream simultaneously, including but not limited to HP infection detection, lesion detection, organ site identification, and tumor classification. Each of these tasks may require a dedicated deep learning neural network to analyze the same gastric video. However, these deep learning neural networks are very computationally demanding - both in terms of computing speed and memory requirements. On the other hand, recent endoscopes may be equipped with very high-resolution cameras. While this provides better image clarity to the end user, it also requires more computing power to process these images. For this purpose, the computer-assisted system needs to be equipped with additional computing hardware, otherwise it cannot meet the real-time requirements. Nevertheless, it is hoped that the system can support different computing hardware configurations with different computing capabilities. For low-end configurations, the system may not be able to provide real-time response. But offline processing may still be useful in some applications. High-end hardware configurations with additional coprocessors can certainly reduce system response time, but the deployment cost is high. Therefore, developing a system that can provide low latency and high performance with minimal additional hardware accelerators is a great challenge.

[0039] To overcome all the above challenges, the computer-assisted gastroscopy system was optimized at three levels: architectural, modular, and functional. At the architectural level, the model was designed in such a way that it can complete HP infection classification and detect a few lesions in one inference to reduce the computational cost. At the modular level, the site recognition model, as a sub-model of HP infection classification, was optimized using temporal information. It not only improves the performance of HP infection classification, but also plays an important role in lesion detection and program state determination. At the functional level, inference latency was minimized through configuration and resource-aware optimization. Also at the functional level, preprocessing was accelerated through image resizing parallelization and unified preprocessing.

[0040] Numbered Examples

[0041] Group 1

[0042] 1. A computer-assisted gastroscopy system, comprising:

[0043] A central processing unit coupled to a memory storing an executable software program, wherein the software program includes:

[0044] An AI image processing system for analyzing a gastric image sequence obtained from a gastroscopy instrument, wherein the AI ​​image processing system includes at least three architecture-level modules that collaboratively perform image quality assessment, lesion detection, cancer identification, HP classification, and site recognition, wherein at least one of the modules includes one or more neural models that extract different but related information from the gastric image sequence and share the information extracted from the gastric image sequence with other modules; and at least one of the neural models fuses HP infection features with site information extracted from other neural models to improve the classification accuracy of the computer-assisted gastroscopy system.

[0045] 2. The system of embodiment 1, wherein the at least three modules include:

[0046] a first module for image quality control to filter out unqualified images in the stomach image sequence;

[0047] a second module for lesion detection, cancer identification, and lesion tracking; and

[0048] The third module is used for classification of HP infection and site identification,

[0049] Each of these modules includes one or more neural models.

[0050] 3. The system of embodiment 2, wherein the third module further comprises a composite neural model, wherein the composite neural model comprises:

[0051] a first neural model that takes the stomach image sequence as input, performs HP feature extraction, and outputs a first number of feature channels;

[0052] a second neural model, also taking the stomach image sequence as input, performing location feature extraction including spatiotemporal location information of the gastrointestinal tract, and outputting a second number of feature channels; and

[0053] A third neural model takes the concatenation of the first number of feature channels and the second number of feature channels as input and generates a third number of class labels, each of which indicates an HP infection feature.

[0054] 4. A system according to Example 3, wherein the first neural model generates a first tensor of sixty-four channel elements; the second neural model generates a second tensor of twelve channel elements, each of the twelve channel elements corresponds to a site classification label; and the third neural model takes the cascade of the first tensor and the second tensor as input, and outputs nine element classification labels corresponding to nine of the HP infection features.

[0055] 5. The system according to Example 3 further includes a unified preprocessing module, wherein the unified preprocessing module takes the stomach image sequence as input and generates a unified tensor as output for each image in the stomach image sequence, and the unified tensor is fed to the neural models of the first module, the second module, and the third module of the AI ​​image processing system.

[0056] 6. A system according to embodiment 5, wherein the neural network architectures of the first module, the second module and the third module are adjusted so that the output tensors of the neural model remain the same, as if each of the neural network architectures uses a different preprocessing module specifically designed for the neural network architecture.

[0057] 7. A system according to Example 5, wherein the neural network architectures of the first module, the second module, and the third module are adjusted so that the performance of each neural model is not reduced.

[0058] 8. The system of embodiment 5, wherein if the height or width of an image in the stomach image sequence entering the unified pre-processing module is greater than a threshold, a parallelized resizing process is called to resize the image, wherein the parallelized resizing process comprises the following steps:

[0059] If the height is odd, pad the original image with a row of zeros, and if the width is odd, pad the original image with a column of zeros;

[0060] Divide the filled image into four quadrants;

[0061] resizing each quadrant in parallel to produce four resized quadrants; and

[0062] Stitch the four resized quadrants together to get a uniform resized image.

[0063] 9. The system of embodiment 2, wherein

[0064] The neural model of the first module is trained using an image quality dataset to generate a complete image quality assessment neural model;

[0065] The one or more neural models of the second module are trained using the lesion dataset to produce a complete lesion detection neural model; and

[0066] The one or more neural models of the third module are trained using the stomach site dataset and the Helicobacter pylori dataset to produce a complete HP plus site neural model.

[0067] 10. The system according to Example 9 also includes a model pruning and quantization module, wherein the complete image quality assessment neural model, the complete lesion detection neural model and the complete HP plus part neural model are optimized by pruning layer connections and quantizing connection weights to generate an optimized image quality assessment neural model, an optimized lesion detection neural model and an optimized HP plus part neural model, respectively.

[0068] 11. A system according to embodiment 10, wherein the computer-assisted gastroscopy system further comprises at least one coprocessor, and the software program executed at the central processing unit wisely assigns the subtasks initiated by each of the modules to the at least one coprocessor according to the pre-assigned priority of the subtasks and the capability of each of the coprocessors, so that the computer-assisted gastroscopy system can achieve the requirements of low-latency response and high-precision detection and classification.

[0069] 12. The system according to embodiment 11, wherein, when the computer-assisted gastroscopy system is equipped with the at least one coprocessor, the computer-assisted gastroscopy system is capable of operating in an offline processing mode and an online processing mode.

[0070] 13. The system of embodiment 12, wherein when the computer-assisted gastroscopy system is configured to operate in the offline processing mode, the computer-assisted gastroscopy system configures each of the at least one coprocessor to operate a dynamic batch processing process, the dynamic batch processing process comprising the following steps:

[0071] loading at least one complete neural model into the coprocessor; and

[0072] A batch of stomach images is loaded into the coprocessor, wherein the batch size is dynamically determined based on available computing power and resources in the coprocessor.

[0073] 14. The system of embodiment 12, wherein when the computer-assisted gastroscopy system is equipped with the at least one coprocessor and is configured to operate in an online processing mode, the computer-assisted gastroscopy system executes a delay control program, the delay control program comprising the following:

[0074] predetermining a loading priority of each optimized neural model according to the computational resource requirements of each optimized neural model;

[0075] Loading one or more optimized neural models into each of the coprocessors according to the loading priority and the hardware configuration of the coprocessors;

[0076] Establishing a task queue of fixed length for each coprocessor, for the central processing unit to send subtasks to the task queue for execution by the coprocessor; and

[0077] Each of the at least one coprocessor is enabled to operate in parallel, wherein each time the coprocessor becomes idle, the coprocessor takes a subtask from a task queue associated with the coprocessor and starts executing the subtask each time the coprocessor becomes idle.

[0078] 15. The system of embodiment 14, wherein when the computer-assisted gastroscopy system obtains the stomach image, the computer-assisted gastroscopy system executes a resource-aware reasoning procedure, the resource-aware reasoning procedure comprising the following steps:

[0079] Selecting a coprocessor with the shortest work queue among the at least one coprocessor as a designated coprocessor;

[0080] loading one or more of the optimized neural models to the designated coprocessor based on neural model priorities and resource availability at the designated coprocessor;

[0081] Repeating the selecting and loading steps so that as many optimized neural models as possible are loaded into one or more of the coprocessors;

[0082] enabling each of the neural models to start inference in parallel at each of the at least one coprocessor;

[0083] When the inference task of the neural model is split to run on more than one of the coprocessors and an intermediate result generated by one of the coprocessors needs to be shared with other coprocessors, performing inter-coprocessor communication between the coprocessors; and

[0084] The inference results of each of the neural models are collected and reported back to the computer-assisted gastroscopy system.

[0085] 16. A method for processing a sequence of stomach images by a computer-assisted gastroscopy system, comprising:

[0086] Acquiring the stomach image sequence from a gastroscopy instrument;

[0087] analyzing the stomach image sequence by an AI image processing system including at least three architecture-level modules that collaboratively perform image quality assessment, lesion detection, cancer identification, HP classification, and lesion site recognition,

[0088] Wherein, at least one of the modules includes one or more neural models, which extract different but related information from the gastroscopic image sequence and share the information extracted from the gastroscopic image sequence with other modules; and at least one of the neural models fuses HP infection features and location information extracted from other neural network models to improve the classification accuracy of the computer-assisted gastroscopy system;

[0089] Creating a subtask list by each of the at least three modules for execution by the computer-assisted gastroscopy system; and

[0090] reducing delayed responses to a user when the computer-assisted gastroscopy system further comprises at least one co-processor and a software program executed at a central processor unit of the computer-assisted gastroscopy system judiciously assigns subtasks to the at least one co-processor according to pre-assigned priorities of the subtasks and the capacity and capability of each of the co-processors,

[0091] This enables the computer-assisted gastroscopy system to achieve the requirements of low-latency response and high-precision detection and classification.

[0092] 17. The method of embodiment 16, wherein the analyzing step further comprises the following steps:

[0093] Filter out unqualified images in the stomach image sequence by using the neural model in the first module;

[0094] performing lesion detection, cancer identification, and lesion tracking by at least one of the neural models in the second module; and

[0095] HP infection is classified and the gastrointestinal site is identified by at least one of the neural models in the third module.

[0096] 18. The method of embodiment 17, wherein the steps of classifying and identifying further comprise the steps of:

[0097] outputting a first number of feature channels by a first neural model that takes the stomach image sequence as input and performs HP feature extraction;

[0098] Outputting a second number of feature channels through a second neural model that also takes the stomach image sequence as input and performs site feature extraction including spatiotemporal site information of the gastrointestinal tract; and

[0099] generating a third number of class labels by a third neural model having as input a cascade of the first number of feature channels and the second number of feature channels,

[0100] Wherein, each of the class labels indicates HP infection characteristics.

[0101] 19. The method of embodiment 18, wherein the second neural model is a composite neural model comprising a part feature extractor model, an LSTM model, and a final neural model, and the composite neural model is trained according to the following steps:

[0102] creating cumulative batches of stomach images from a complete set of labeled stomach images stored in a stomach region dataset;

[0103] training an auxiliary neural network using the accumulated batch of stomach images;

[0104] Copying the entire auxiliary neural network and using the entire auxiliary neural network as a feature extractor of the composite neural model, wherein the connection weights of the feature extractor are not modified during a subsequent training process;

[0105] creating a pseudo video using the complete set of labeled stomach images, and collecting a predefined number of labeled stomach images from the pseudo video to form a batch unit;

[0106] grouping a predefined number of batch units together as a batch group, and sending the batch group to the feature extractor to generate a feature tensor of the batch group;

[0107] Feeding the feature tensor to the LSTM model, the LSTM model generates an intermediate tensor;

[0108] The intermediate tensor is sent to a final neural model including at least one fully connected layer, wherein an output tensor of the final neural model includes a part vector, each element of the part vector is a part label and represents a stomach part position, and the part vector becomes a feature channel of the second neural model.

[0109] 20. The method of embodiment 19, wherein the pseudo video is created according to the following steps:

[0110] Sorting the labeled stomach images in the stomach dataset in ascending order according to the index of the part label to obtain a sorted list of stomach images;

[0111] Creating a predetermined number of random generators, each of the random generators generating a random number within a predetermined random range;

[0112] selecting one of the random generators to generate an initial random number and using the initial random number to select a stomach image from the ranked list as an anchor image;

[0113] collecting a set of random numbers from each of the plurality of random generators, wherein a total number of the collected random numbers is a predefined number, the predefined number being a batch size of the batch unit;

[0114] converting the set of random numbers into a list of indices offset by the index of the anchor image; and

[0115] Stomach images are selected from the sorted list using the offset index list to form the batch unit.

[0116] System Architecture

[0117] Reference now Figure 1 , discloses a computer-assisted gastroscopy system 100. The system includes an image capture card 103 that captures a series of digital images from an endoscope device 101. In one embodiment, the endoscope device 101 includes an image camera and a light source that can be inserted into an internal organ of a human body 130. When the endoscope device 101 is used to examine the esophagus and stomach, it is generally referred to as a gastroscope instrument.

[0118] The image capture card generates a video image sequence 104, which is sent to the endoscopic image analysis module 120 for further analysis. In one embodiment, the endoscopic image analysis module 120 includes a unified pre-processing module 107, which receives the image sequence 104 and resizes each image to a standard size before feeding them to the AI ​​function and image processing module 108. This module performs the main analysis to detect any abnormalities on the image sequence 104, which will be discussed in detail later. The results of this analysis are sent to the post-processing module 109. After post-processing, the results are sent to the database 112, the program status module 110, and the display 102, where the results are displayed to the end user along with the image sequence 104.

[0119] In one embodiment, the procedure status module 110 captures all information of the gastroscopic analysis procedure, including the patient's information, the start and end of the gastroscopic procedure, and the post-processing results, and feeds them to the case-level analysis module 111 to create a diagnostic case for the patient. The diagnostic case is stored in the database 112 together with the results from the post-processing module 109 and the image sequence 104.

[0120] In one embodiment, the AI ​​function and image processing module 108 employ one or more neural models 106 for image processing and analysis, and they are very computationally intensive. Therefore, the endoscopic image analysis module 120 is designed to utilize all hardware resources available in the system to speed up the response. In particular, the computer-assisted gastroscopy system is developed to run on a variety of hardware platforms, with or without a graphics processing unit (GPU) coprocessor (one or more). The configuration and resource optimization module 105 considers the hardware information 116 and attempts to configure and schedule multiple tasks generated by the neural module 106 so that the AI ​​function and image processing module 108 can execute these tasks in parallel to reduce turnaround time.

[0121] To this end, the neural model needs to be trained before any abnormalities in the gastric image sequence can be classified and detected. The model training and testing module 114 uses the endoscopy data set 115 to train each neural model so that the overall system performance of the computer-assisted gastroscopy system 100 can meet the target requirements specified by medical professionals. After training, the fully trained neural model can be further optimized by the model optimization module 113 to produce an optimized neural model.

[0122] Figure 1 The AI ​​function and image processing module 108 Figure 2. The module 220 also includes three sub-modules, namely, a first module 221 for image quality control, a second module 222 for detecting lesions with clear boundaries, and a third module 223 for detecting lesions without clear boundaries. All three modules use one or more neural models to process stomach images. For clarity, we represent the neural model used in the first module as NN1, the neural model used in the second module as NN2, and the neural model used in the third module as NN3. As described below, both the second module and the third module use more than one neural model for analysis, and we use sub-indexes to represent them. For example, NN2.1 and NN2.2 represent two neural models used in the second module, and NN3.1 and NN3.2 represent the neural models used in the third module.

[0123] The first module 221 also includes a region of interest (ROI) module 201 and an image quality assessment neural model NNl (200) to perform image quality assessment. Its purpose is to filter out unqualified images in the stomach image sequence, thereby saving time for unnecessary reasoning on these bad images. The second module 222 uses a lesion detection neural model NN2.1 (202) to detect lesions with clear boundaries based on object recognition technology. In addition to lesion detection, it also performs cancer identification. The second module 222 also includes a lesion tracking neural model NN2.2 (204) for lesion tracking and key frame selection. The third module 223 uses multiple neural models to detect HP infection and other types of lesions without clear boundaries. It consists of a feature extraction neural model NN3.1 (205) and a site recognition neural model NN3.2 (207). The outputs of these two neural models are cascaded and sent to a merged neural model NN3.3 (206). The site information extracted from the site recognition neural model 207 is not only useful for the third module, but can also be combined with the results of the second module to perform image-level analysis 208.

[0124] The detection results of all three modules will be sent to the case-level analysis module, which will further generate key frame selection diagnosis report 209, cancer risk assessment 210 and HP infection degree analysis 211.

[0125] A neural network is a computational model inspired by biology. It consists of multiple nodes (or neurons) arranged in two or more layers. There are basically three types of layers - input layer, zero or more hidden layers, and output layer. Nodes in the input layer receive sensory input data. This can be in the form of real vectors, two-dimensional matrices (such as pixel values ​​of an image), or even higher-dimensional data structures. The output layer provides the inference results of the neural network. When the neural network performs classification, each output node represents a class. The hidden layer is the layer between the input layer and the output layer. Each node in a hidden layer or output layer is connected to the nodes in the previous hidden layer or input layer; each connection is associated with a real value, called the connection weight. In operation, each node first calculates a weighted sum based on these weight values ​​and the output values ​​of the nodes in the previous layer to which it is connected, and then executes a function to obtain an activation value. This activation value is the output value of the node, which will be sent to those nodes in the next layer to which the node is connected. The function can be the Softmax function, which is usually used in the last layer, the linear rectifier function (ReLU), or just the average or maximum value of the activation values ​​of the nodes in the previous layer to which the node is connected. This forms a common neural network architecture.

[0126] While a hidden node in a hidden or output layer can be fully connected to all nodes in its previous hidden or input layer, other neural architectures employ only partial connections. The former case is called a fully connected (FC) layer. In the popular convolutional neural networks (CNNs), which are often used to process two-dimensional digital images, the nodes in a hidden layer are connected only to a small square grid of nodes in the previous layer. Typically, the grid size is 3x3, 5x5, or 7x7. The weight values ​​of these connections specify the pattern that this hidden node is looking for, and are often called the filter or kernel of this hidden node. There may be more than one node in a hidden layer connected to the same grid; but each of these nodes has a different filter, so they are looking for different patterns from the same grid. These nodes can be stacked on top of each other so that the hidden layer can be concisely represented as a tensor of three numbers - two numbers representing the grid size mentioned above, and the third number representing the number of nodes stacked together.

[0127] CNN neural models can have up to a hundred or more hidden layers. There are usually two types of hidden layers - convolutional layers and pooling layers. In a convolutional layer, each node is looking for a specific pattern as mentioned earlier, while in a pooling layer, all weights in the kernel are the same and the function of the node is either taking the average of all grid values ​​(AvgPooling) or the maximum of them (MaxPooling). Usually, pooling layers are interspersed between one or more convolutional layers. Another important parameter that defines a CNN is the stride, which specifies the number of pixel shifts of the grid between adjacent nodes in the hidden layer.

[0128] Another neural architecture is the recurrent neural network (RNN). In this network architecture, hidden nodes are not only connected to nodes in the previous hidden or input layer, but also to their own outputs. As a result, RNNs are able to remember their own processing state and are able to capture sequential dependencies in sequence prediction problems, such as time series analysis. A specific category of RNNs is the long short-term memory (LSTM) neural model. A common LSTM cell consists of a cell, an input gate, an output gate, and a forget gate. The block remembers values ​​over arbitrary time intervals, and the three gates regulate the flow of information into and out of the cell. It is also possible to combine CNNs and LSTMs together. This is particularly useful for processing videos, as CNNs can be used to extract spatial information, while LSTMs can be used to collect temporal information.

[0129] There are many existing CNN models that developers can use. Some examples are ResNet, CenterNet, and Xception. Using these existing neural models can speed up development.

[0130] Before a neural network can be deployed, it must be trained. Training is the modification of the connection weights of all nodes. In one embodiment, the first step that needs to be done is to collect a data set for training the neural network. For each sample in the data set, a label needs to be assigned to the sample. For classification tasks, this becomes the class label of the sample. In some cases, an expert may be required to examine the sample and assign a class label to it. This can be a tedious task when the number of samples is large. In one embodiment, training the neural network includes presenting the sample to the input layer and comparing the output nodes of the output layer with the class label. Through the loss function, the difference between the assigned class label and the activation value of the output node is used to calculate the loss value, and the training algorithm (such as the back propagation algorithm) is called to adjust the weights of the connection so that after multiple training iterations, the overall loss is reduced to a local minimum. After training, the neural network will have higher discrimination and the classification accuracy will be much higher than random selection.

[0131] In one embodiment, the output nodes of one neural network can be fed to the input nodes of a subsequent neural network. Alternatively, the output nodes of one neural network can be connected to the output nodes of a second neural network and then fed to the input nodes of a third neural network. Therefore, cascading one or more neural networks together in this way creates a larger neural network. We refer to the resulting neural network as a composite neural model. In general, we use symbols such as NN3 to represent composite neural models, and use sub-indexes such as NN3.1, NN3.2,.. to represent individual neural model components.

[0132] Figure 3The training process of various neural models according to one embodiment of the present invention is shown. First, an endoscope raw data set 300 is collected. The raw data set may include multiple different data sets. Through the labeling process 301, each sample in the raw data set 300 is labeled. The image quality data set 302 is used to train the neural model NN1 of the first module using the model training #1 program 307. After training, a complete image quality assessment neural model 315 is obtained. Similarly, the stomach site data set 303 and the Helicobacter pylori data set 304 are used to train multiple neural models 308, 309 and 310 as components of the composite neural model NN3 of the third module using the model training #2 program 311. This produces a complete HP plus site neural model 316. The details of how to train each sub-model will be discussed later. In the same way, the lesion data set 305 is used to train the neural models NN2.1 and NN2.2 of the second module using the model training #3 program 314 to obtain a complete lesion detection neural model. Although the respective complete neural models produce the best detection and classification performance, they may require too many computing resources to run and are not suitable for online processing. Therefore, the three complete neural models are sent to the model pruning and quantization module 321 to prune the layer connections and quantize the connection weights. Then, it generates an optimized image quality assessment neural model 318, an optimized HP plus part model 319, and an optimized lesion detection model 320, respectively. The optimized neural models speed up the inference speed while keeping the performance close to the respective complete neural models. Table 1 shows the actual performance of the optimized neural models. As shown in this table, the target sensitivity and specificity values ​​are specified by medical professionals. The test values ​​are the respective performance values ​​of these optimized neural models. Table 1 shows that for the three classification and detection tasks, the optimized neural models exceed the target values ​​with a large margin.

[0133] Model Sensitivity (test / target) Specificity (test / target) HP infection classification 0.87 / 0.8 0.95 / 0.8 Lesion detection 0.82 / 0.8 0.90 / 0.9 Cancer Detection 0.79 / 0.65 0.92 / Nil

[0134] Table 1. Sensitivity and specificity performance of the optimized neural model for user-specified target values ​​in three separate classification and detection tasks

[0135] In this disclosure, we also refer to the NN1 neural model as the image quality assessment neural model; the NN2 neural model as the lesion detection neural model, and the NN3 neural model as the HP plus site neural model. When these terms are used, the context will specify whether these symbols refer to the complete neural model or the optimized neural model.

[0136] Classification of HP infections and lesions

[0137] We now discuss in detail how the third module performs classification of HP infections and lesions. This module is specifically designed to detect and classify lesions without clear boundaries, as the object detection model NN2 is not suitable for this task.

[0138] According to the Kyoto Classification of Gastritis, Helicobacter pylori (HP) infection is divided into three stages: (i) non-gastritis or HP(-): the gastric mucosa is not infected with H. pylori, (ii) active gastritis or HP(+): the gastric mucosa is currently infected with H. pylori, and (iii) inactive gastritis or past HP(+): the gastric mucosa was previously infected. When HP infection occurs, gastric images show certain symptom patterns. These patterns are used as features of the HP infection classifier to determine which of the three above categories the gastric image belongs to. Figure 4 Table 400 in Figure 400 shows a list of these features. The Kyoto classification highlights six symptom patterns, which are indicated in the first 6 rows of Table 400 and grouped under reference label 401. In addition to extracting these 6 features for HP detection, the feature set is also expanded to include features useful in detecting and classifying other types of lesions that also do not have clear boundaries. In this regard, features L6 and L7 (403) are used to classify lesions that are not related to HP infection, while features L5 (402) and L8 (404) are used to classify HP infection and other types of lesions that do not have clear boundaries.

[0139] Therefore, for the HP classifier, a total of 9 features (L0 to L8) are first extracted from the stomach image. The details on how these 9 values ​​are obtained will be discussed in subsequent sections. Once obtained, they are sent to the tree classifier to determine which of the three HP classes mentioned above the stomach image belongs to.

[0140] Figure 5 The main components of the third module, which performs HP infection feature extraction and site identification, are reviewed. Neural network NN3 (620) is a composite neural model that combines three neural models, namely NN3.1 (621), NN3.2 (622) and NN3.3 (623). NN3.1 (621) is designed to extract features of HP infection and other lesion types without clear boundaries, and it consists of HP and lesion feature extractors 605. Its output layer is a vector of 2048 elements or 2048x1 channels (606). The vector is compressed to a vector of 64 elements (607) in the next layer. In one embodiment, the connection between layer 606 and layer 607 is a fully connected network. In another embodiment, a 1x1 convolution kernel is used.

[0141] NN3.2 (622) is a special neural network designed for part recognition. It consists of a part feature extractor 608, whose output layer is a 2048x1 element vector 609, which is further compressed to a 12-element vector in the next layer 610. The output node of layer 610 is cascaded with the output node of layer 607 in the cascade gate 611 to become the input node 612 of the merged neural network NN3.3 (623). Therefore, the input size of NN3.3 (623) is (64+12=76) elements. The input layer 612 is fully connected to the output layer 613, which is a vector of 9 elements. These 9 elements correspond to Figure 4 The 9 features L0 to L8 in Table 400.

[0142] In operation, a stomach image 601 is sent to the HP feature extractor 605 and the site feature extractor 608. Neural networks NN3.1 (621) and NN3.2 (622) process the same image simultaneously; NN3.1 (621) focuses on extracting information about HP infection and other types of lesions without clear boundaries, while NN3.2 (622) performs site recognition. Site information provides additional information for HP detection, so by cascading the outputs of NN3.1 (621) and NN3.2 (622), the merge neural network NN3.3 (623) combines the two information sources together, thereby improving the sensitivity of HP infection detection.

[0143] In one embodiment, HP feature extractor 605 and part feature extractor 608 are both multi-layer CNN neural models. In another embodiment, they are Xception CNN models.

[0144] These neural models are trained as follows. In one embodiment, NN3.2 (622) is pre-trained. The entire neural model is copied from the stomach location recognition model 602. On the other hand, NN3.1 (621) and NN3.2 (623) are trained with the connection weights in NN3.2 (622) fixed. The two neural models are trained by first presenting the image 601 to the input layer of NN3.1 (621) and NN3.2 (622). The HP label 600 provides a label corresponding to the image and is sent to the loss function evaluator 614. The loss function evaluator 614 compares the output of NN3.3 (623) with the HP label 600 and uses a predefined loss function to evaluate the loss value. The training program adjusts all connection weights in NN3.1 (621) and NN3.3 (623) to minimize this loss value, while the connection weights of NN3.2 (622) remain unchanged. In one embodiment, binary cross entropy (BCE) with a Logits loss function is used in this implementation.

[0145] It was observed that the composite neural model NN3 was able to learn the correlation between HP infection features and gastric site locations. Therefore, the site-assisted neural model NN3 was able to improve the sensitivity of HP infection detection by 5.6% in one experiment while keeping the specificity unchanged.

[0146] The location information from NN3.2 (622) not only helps with HP infection classification, but can also be used to post-process lesion detection results and determine the surgical procedure status of the gastric examination process. In one embodiment, the location information can be used to: (1) determine the endoscope start and stop operation procedures, and (2) reduce false positives in HP infection classification and lesion detection. For example, if the system determines that the operation is in an end state, the system can ignore the results of AI reasoning to reduce false positives in HP classification and lesion detection. As another example, if the endoscope is located in the duodenum inferred from the location information, the system may ignore any HP features in the model to reduce false positives. Therefore, the output of NN3.2 (622) is directed to the lesion detection neural model (603), which is the neural model in NN2 of the second module; and also directed to the program status module 604.

[0147] As previously mentioned, location information is useful to many modules in the system. The stomach location recognition model 602 is a neural model trained to recognize location in stomach images. In one embodiment, the location is displayed in Figure 6 Each entry in this table is associated with a location in the upper gastrointestinal tract. The stomach location dataset contains a large number of images annotated with location labels, such as Figure 6 As shown, it is used to train a neural model for part recognition.

[0148] Figure 7 An alternative embodiment of a part recognition model and how to train the alternative part recognition neural network in one embodiment of the invention is shown. The first step is to train an auxiliary neural network 701 that acts as a part feature extractor. This is itself a CNN. In another embodiment, it uses the Xception neural model. The output of this CNN is a vector of 2048 channels, which are connected to a classifier neural network 702. Neural networks 701 and 702 are trained together to recognize such as Figure 6 The region labels are shown. A cumulative batch of stomach images 700 from the stomach region dataset is used to train this composite neural model.

[0149] After this training step, the auxiliary neural network 701 is retained and copied to the composite neural model 720. The composite neural model 720 includes a part feature extractor model 706, an LSTM model 707, and a final neural model 710. The part feature extractor 706 is a direct copy of the auxiliary neural network 701, which extracts part features, and its connection weights are not updated when the other two neural models are trained. In order to train the LSTM model 707 and the final neural model 710, a series of batches of units are first prepared. In one embodiment, a pseudo video simulating an actual stomach video image sequence is first created based on a labeled stomach image extracted from a stomach part dataset. (For the generation of the pseudo video image sequence, please refer to the following paragraphs and Figure 8 ). From this pseudo video, five consecutive images are grouped together to form a batch unit. In one embodiment, five stomach images are grouped along the time step direction 703 to form a batch unit. Then 32 of these batch units are grouped together to form a batch group along the batch direction 704. During training, the part feature extractor 706 generates a feature vector of 2048 channels for each image. Therefore, for one batch unit, the part feature extractor 706 produces an output tensor of [5x1x2048] elements, and for a batch group with 32 batch units, it produces an output tensor of [5x32x2048] elements. This is fed to the LSTM model 707, which captures the temporal information from the stomach image sequence and outputs a tensor of [5x32x64] elements. This output tensor is used as the input to the final neural model 710. This neural model has a first fully connected layer 708 coupled to another fully connected layer 709. The output of this neural model, which is also the output of the composite neural model 720, is a tensor of [5x32x12] elements. The last dimension is the number of part labels, and each element in this dimension corresponds to a part label, such as Figure 6 As shown in the table.

[0150] It should be understood that Figure 7 Only an exemplary method of training the composite neural model 720 is shown. In this example, the labeled stomach images are arranged into batch units and then arranged into batch groups with sequential indices in preparation for training. The batch size used is 32 and the number of batch units in the batch group, i.e., 5, is used here only for illustration. Other values ​​can be used.

[0151] It should be understood that the LSTM model 707 is able to capture time information. Therefore, the training data needs to show the time information captured by the LSTM. In actual gastroscopy, the video captured during the entire endoscopic examination contains time information. However, in order to use the video to train the LSTM model 707, a part label needs to be assigned to each frame of the video, such as Figure 6This has to be done by an expert in the field and labeling is very time consuming and tedious. On the other hand, a stomach region dataset is available - every image in this dataset has been labeled. However, these are still images that do not carry any temporal information. Therefore, a method to create pseudo videos from these still images is developed and disclosed as follows.

[0152] Reference now Figure 8 , an exemplary embodiment is given to illustrate the method of creating a pseudo video sequence. In this example, we assume that there are only two labels S1 and S2 in the stomach image dataset, and we want to create two batch units. The first step is to sort all the labeled images in the stomach area dataset so that images with the same label are grouped together and the groups are sorted in the order of accent. For example, images with label S1 are grouped together as the first label group 800, while images with label S2 are the second label group 801, which is appended to the end of the first label group 800 to form a long chain of sorted labeled images. Next, an image from the first label group 800 (in this case, Img#2 804) is randomly selected and inserted as the first labeled image in the first batch unit 802. Then, using this image as an anchor point, other images are selected and inserted into the first batch unit 802. To this end, four random numbers are generated. Note that when generating random numbers, the random range can be adjusted. When the range is small, the probability of selecting samples near the anchor image is high. Therefore, the selected image may have the same label as the anchor image. On the other hand, when the random range is high, images from different label groups can be selected. Using four random numbers generated from four different random ranges, a sequence of images can be generated whose labels are mostly the same as the labels of the anchor image, but also contain images with different labels randomly scattered in the first batch unit 802. In a similar manner, the second batch unit 803 is constructed. In this example, the anchor image is Img#7 805. Note that in this example, Img#7 805 is closer to the end of the first label group than Img#2 804. Therefore, the second batch unit 803 has more images selected from the second batch unit 801 than the first batch unit 802. In this way, the sequential arrangement of the synthetic images in the two batch units becomes a pseudo video that can simulate a video image sequence showing part boundary transitions. In other words, the temporal information at the part transitions is artificially created from the still labeled images, and then such part transition information can be used to train the LSTM model.

[0153] It should be understood that Figure 8 Only an exemplary method of generating pseudo video image sequences for training is shown. The same method can be generalized to accommodate more than two part labels and generate more than two batch units.

[0154] In one embodiment, both the part feature extractors 701 and 706 use the Xception neural model. When the part feature extractor 706 captures the spatial information of the stomach image, the LSTM model 707 captures the temporal aspect of the stomach image, thus extracting spatial and temporal information from the stomach image sequence together with the composite neural model 720.

[0155] In one embodiment, Figure 7 The execution time of the part feature extractor 706 in the batch group is the longest. It needs to extract features for each image in the batch unit. To speed up this process, each batch unit in the batch group can be executed in parallel in a system equipped with one or more coprocessors (such as GPU cards) and the results are placed in a result queue. However, each subtask may be completed at different times, which may result in Fig. 9 The out-of-sequence problem shown. Fig. 9 , the multi-processing feature extraction step 901 creates subtasks for each batch unit from the batch group sequence 900, and these subtasks are distributed to one or more coprocessors so that they are executed in parallel. The results are placed in the result queue 902. However, some subtasks may be completed earlier than other subtasks, so when they are inserted into the result queue, they may not be in the expected sequential order. In addition, there may be cases where a subtask may not produce any results. In both cases, it will cause the batch units in the result queue 902 to be not arranged in their expected sequential order, so they need to be rearranged in step 903 before sending them to the LTSM model for training. After rearrangement, they can be sent to the LTSM model and LTSM model reasoning can be performed. The output of the LTSM reasoning is the prediction of the part label for the i-th sequence order in step 905, and it can be inserted back into the result queue.

[0156] The composite neural model 720 can be deployed in an actual application environment. In actual deployment, a stomach image sequence is presented as the input of the composite neural model 720, rather than a set of continuous batch units. Essentially, this is equivalent to setting the batch size to 1. Although the LSTM model can produce outputs of 5 sequence steps, only the latest time step is used. Therefore, the output of the composite neural model 720 is a vector of 12 elements, each element corresponding to a part label, such as Figure 6 shown.

[0157] With this arrangement, the composite neural model 720 can replace Figure 52(622) in the Xception part feature extractor 706. Since the composite neural model 720 utilizes the LSTM model 707 to capture temporal information, it further helps the HP infection classifier improve its accuracy. Table 2 below shows the improvement in accuracy. As shown in the table, adding the LSTM model 707 to the Xception part feature extractor 706 improves the accuracy of HP infection classification from 94.6% to 97.4%; an improvement of 2.8%.

[0158] Model Type Accuracy Xception only 94.6% Xception+LTSM 97.4%

[0159] Table 2. Accuracy improvement after incorporating the LTSM model

[0160] In addition to improving HP detection accuracy, the composite neural model 720 also reduces the number of false positives when detecting site transitions in stomach videos. Fig.10 The results of one experiment are shown. In the figure, the horizontal axis represents the frame index in the video sequence, and the vertical axis is the accumulated position change or part transition. The two curves in the figure show the difference in the accumulated position change in a typical stomach video. Curve 1001 is obtained without using the LSTM 707 model, while curve 1002 is obtained by using LSTM to capture the temporal part transition information. The figure clearly shows that with the help of the LSTM 707 model, the number of erroneous positions of the part transition is greatly reduced.

[0161] Preprocessing

[0162] Preprocessing is the step of preparing the stomach images into a form that can be processed by various neural models. Nowadays, gastroscope instruments come from different manufacturers and different brands. Each of them can use different cameras with different pixel resolutions. This may range from less than 384x384 pixels to more than 3840x2180 pixels. Most are color cameras, so the total number of pixels is three times the resolution. On the other hand, each neural model employed in this system requires input image dimensions of a specific size. For example, the Xception model processes color images of 299x299x3, CenterNet requires 384x384x3 input, and the ResNet model processes 224x224x3 input. Therefore, the incoming stomach images of different resolutions need to be converted into a format suitable for each neural model; this is the task of preprocessing.

[0163] Reference now Fig.11And as mentioned above, the computer-assisted gastroscopy system uses the three neural models mentioned to process gastric images. The traditional endoscopic image analysis scheme 1100 will adopt three different pre-processing modules 1102, 1103 and 1104 to process the input gastric image 1101. For pre-processing module #1 1102, it converts the input image into a tensor 1105 of [299x299x3] dimensions to be fed to the HP plus part neural model 1108. Similarly, pre-processing module #2 1103 produces a tensor of [384x384x3] ​​dimensions 1106 to meet the input requirements of CenterNet used in the lesion detection neural model 1109. Finally, pre-processing module #3 1104 generates a tensor of [224x224x3] for the image quality model 1110. The outputs of the three neural models will be sent to the post-processing module 1111 for further processing.

[0164] This conventional approach is obviously inefficient as it requires three separate preprocessing modules. While each of these modules produces a unique output image dimension specified by its recipient, much of the processing in these individual preprocessing modules is identical, so it is best to merge them to save computational time and resources. Therefore, a unified preprocessing endoscopic image analysis scheme 1113 was developed, whereby it employs a unified preprocessing module 1115 to process the gastric image 1114 and produce an output tensor 1116 of [384x384x3] ​​dimensions. This same tensor is used as input to the individual neural models 1108, 1109, and 1110. The outputs of these three neural models will be sent to the same post-processing module 1111 for further processing.

[0165] Since the output tensor dimension 1116 of the unified pre-processing unit 1115 may be different from what is required by the subsequent neural model, the neural model may need to adjust its internal architecture parameters so that it can adapt to the output tensor dimension 1116. Once the internal architecture parameters are changed, the neural model may produce an output tensor dimension that is different from the original tensor dimension. Most importantly, the performance may be reduced. Therefore, it is desirable to keep the output tensor of the neural model the same and its performance not degraded so as not to affect any downstream processing. Therefore, the neural architecture parameters need to be changed.

[0166] Here is an exemplary embodiment to illustrate how to adjust the neural architecture parameters to adapt to the change in the input tensor dimension while keeping the performance and output tensor dimension the same. In this embodiment, the Xception neural model is used to illustrate the basic concepts. The Xception neural model is a CNN model used to extract part recognition features in this application. This 71-layer neural model is organized into three main parts, namely the inlet flow, the intermediate flow, and the outlet flow. In one embodiment, only the outlet flow part is modified to meet the above criteria. In the original Xception architecture, the input tensor is [299x299x3] and the output tensor is a vector of 2048x1 dimensions. The intermediate input tensor of the outlet flow part of the Xception model is [19x19x728]. If the input tensor is changed to [384x384x3] ​​to accommodate the output tensor 1116 of the preprocessing unit 1115, the input tensor of the outlet flow part becomes [24x24x728] because the first two dimensions of the input tensor at the inlet flow part are larger. Therefore, some parameters in the outlet flow part are modified to keep the performance and output tensor the same.

[0167] Reference now Fig.12 , reference label 1200 shows the original outlet flow architecture of the Xception model, while reference label 1216 shows the modified outlet flow architecture that produces the same output vector as required. As mentioned earlier, when the input tensor of the inlet flow becomes [384x384x3], the input tensor 1218 of the outlet flow layer becomes [24x24x728]. In the original outlet flow model 1200, it produces a feature map 1213 of [12x12x728] tensor dimension. Note that a larger size feature map is undesirable because it loses more information after global pooling. By changing the block 1201 that performs a 1x1 convolution with a stride of 2x2 to a block 1217 that performs a 1x1 convolution with a stride of 3x3, the modified outlet flow 1216 produces a feature map 1229 of [8x8x728] tensor dimension. This effectively reduces the information loss.

[0168] The above examples illustrate the idea of ​​adjusting the neural architecture to adapt to different input tensor sizes. It should be understood that for different neural architectures, different parameters may need to be changed to achieve the desired results, but based on the teachings disclosed in this specification, those skilled in the art will be able to apply the same ideas to solve their specific problems.

[0169] With Fig.11 The unified preprocessing 1113 achieves a faster preprocessing time compared to the conventional preprocessing method shown in 1100. The following table compares the preprocessing time required for the conventional method and the unified preprocessing method.

[0170]

[0171] Table 3. Comparison of preprocessing time for separate preprocessing and unified preprocessing

[0172] Referring to the table, it shows that the unified preprocessing method can reduce the preprocessing time for both image sizes by about one third. The experiment was conducted on a computer equipped with an Intel Core i7-7800X CPU running at a clock frequency of 3.5GHz.

[0173] As mentioned above, different models and brands of gastroscopy instruments use different cameras to produce gastric images. The resolutions of these cameras vary greatly - from low resolution 342x372 pixels to ultra-high resolution 3840x2160 pixels. In the unified preprocessing, the standardized output is a tensor 1116 of [384x384x3] ​​dimensions. Therefore, regardless of the camera resolution, the unified preprocessing module 1115 resizes the input gastric image to the standardized output dimension.

[0174] Fig.13 A flow chart showing how this is done is shown. Referring to the figure, if the input stomach image 1300 has a height or width below 384 pixels, it follows path 1305 to direct the image to undergo a single pass resize 1307, which resizes the image to 384x384 pixels.

[0175] If the height or width of the stomach image 1300 is higher than 384, it takes path 1301. Then it further checks whether the height or width of the stomach image 1300 is an odd number. If yes, path 1302 is taken and the stomach image is padded with a row or a column of zeros so that the resulting image has an even number of rows and columns. The resulting image is then sent to the parallelized resizing module 1304 for resizing. If both the height and width of the original stomach image 1300 are even numbers, path 1306 is taken and the image is directly sent to the parallelized resizing module 1304. After resizing, an image 1308 of 384×384 pixels is obtained.

[0176] Fig.14 Details on how to perform Fig.13 1409. The parallelized resizing module 1304 in FIG. 1400. The input image is first divided into four quadrants 1400, 1401, 1402, and 1403. Each quadrant is sent to the resizing process 1404, 1405, 1406, and 1407, respectively. The results of the four resized quadrants are then stitched together at 1408 to produce a resized image 1409. In this method, overlapping cropping is avoided, thereby saving computation time.

[0177] If the computer-assisted gastroscopy system is equipped with a coprocessor such as a graphics processing unit (GPU) card, the resizing process can be executed in parallel to speed up the resizing process. In this case, the main CPU of the computer-assisted gastroscopy system can assign tasks initiated by the resizing processes 1404, 1405, 1406, and 1407 to the coprocessor when the coprocessor is idle, so that these processes can run in parallel.

[0178] The following table shows the speedup performance between a single process resizing and four processes working in parallel.

[0179]

[0180] Table 4. Speedup performance between single process resizing and four processes working in parallel

[0181] Referring to the table, it clearly shows that using parallelized resizing can roughly cut the resizing time in half. If the stomach image size is less than 384x384 in either dimension, there is no need for parallelized resizing as the single-process system can resize the image in less than half a millisecond. The experiment was conducted on a computer equipped with an Intel Core i9-9900K CPU running at a clock frequency of 3.6GHz.

[0182] Hardware Configuration

[0183] Deep neural models require a lot of computing resources to run. It is advantageous to explore parallel processing techniques to reduce the running time. Nowadays, it is easy to add one or more coprocessors, such as graphics processing unit (GPU) cards, to a personal computer to increase the overall computing power. The development of computer-assisted gastroscopy systems is to take advantage of this trend and be able to use all available hardware facilities to speed up the response.

[0184] Fig.15 The following describes how the computer-assisted gastroscopy system detects the available hardware configuration and makes full use of its computing power. In one embodiment, it performs configuration-aware optimization as follows: When the system is running, it first executes step 1501 to read the hardware configuration of the computer-assisted gastroscopy system hardware platform 1500. If the hardware platform 1500 does not have a coprocessor such as a GPU card, the system must run in offline mode. Therefore, path 1502 is taken. The system will load all complete neural models to the computer in step 1503 and start using the main CPU power for reasoning in 1504.

[0185] If the hardware platform includes one or more coprocessors or GPU cards, path 1505 is taken and the user can choose to run the system in offline mode or online mode. For offline processing, path 1506 is selected. The system then loads the complete neural model to the coprocessor in preparation for the subsequent reasoning process. At this point, the system checks the hardware configuration and capabilities of each GPU card in step 1508. If the GPU card has a mid-to-high-end configuration in terms of its memory size and processing speed, the system will execute the dynamic batch reasoning program in step 1509. Otherwise, it will skip this step. The details of dynamic batch reasoning will be discussed later. At this point, the system further checks whether the hardware platform 1500 has a single GPU or multiple GPUs. If the former, path 1510 is taken, and the system uses a single GPU to perform all neural reasoning in step 1511. If the system has multiple GPUs, path 1514 is taken, and another reasoning program is called using multiple GPUs in step 1515.

[0186] If the system is chosen to run in online mode, path 1512 is taken and the system will perform a latency control step 1513. This step primarily attempts to manage response latency to an acceptable level so that the user does not have to wait a long time to get the results. The details of this step will be discussed later. Also in this step, the optimized neural model is loaded to the GPU instead of the complete neural model because the optimized neural model is able to run faster. After this step, the system will join a path to check whether the hardware platform 1500 has a single GPU or multiple GPUs, and proceed accordingly as described above.

[0187] As mentioned above, when the system runs in offline mode and one or more GPUs have sufficient resources, the system performs dynamic batching. In dynamic batching, the system dynamically changes the batch size based on the computing power and memory availability of the GPU card. For example, when the GPU memory is large enough, loading more images in batches can speed up inference. Fig.16 The relationship between batch size and GPU computing power and available resources is illustrated.

[0188] When the system is running in online mode, the response time needs to be optimized. In one embodiment, the system performs latency control. The system explores every method of parallel processing to perform tasks such as unified preprocessing and / or neural inference in parallel. Note that in this case, the optimized neural models are loaded to the GPU because they can be executed faster than the corresponding full models. To do this, different neural models are prioritized based on their computational power requirements. Fig.171701 illustrates this priority arrangement based on task importance in the present application. In the figure, the horizontal axis represents the computing power of the GPU, and the vertical axis represents various neural models according to priority. As shown in the figure, NN1 is the least important task, so it has the lowest priority, while NN2.1 is the most important task and is assigned the highest priority. Therefore, at any given time, the system uses this table to perform adaptive model loading. If the computing power of the GPU is low or the available computing resources are small, only the NN2.1 model is loaded, as shown in 1701. On the other hand, when the computing power is high or the available computing resources are abundant, all five neural models can be loaded to the CPU, as shown in reference label 1702. When more than one GPU is available, the system will load some neural models onto one GPU and the rest onto other GPUs to balance the workload to minimize the inference time.

[0189] In another embodiment, the system assigns a bounded-length task queue to each GPU and fills the queue with tasks. Fig.18 A timing diagram of the operation of a bounded length double-ended queue in an exemplary embodiment is shown. Fig.18. The horizontal axis refers to the timeline. The main CPU of the system begins to issue tasks to the GPU. In the (i-1)th time slot, the main CPU issues a task, as shown in time step 1811. At this time, the bounded queue is empty, so the task is placed in the highest time slot of the queue, as shown in reference label 1804, and is ready to enter the GPU for execution. However, the GPU is still busy performing inference on the task entered in the (i-2)th time slot, as shown in reference label 1800, so the task needs to wait. In the i-th time slot, the CPU issues another task, as shown in time step 1812, which is also placed in the bounded queue. In this example, the size of the bounded queue is two time slots, so the bounded queue is full, as shown in reference label 1805. Later, the GPU has completed executing the task entered in the (i-2)th time slot, as shown in reference label 1800; therefore, the (i-1)th task at reference label 1805 is popped from the bounded queue and placed in the GPU for execution, as shown in reference label 1801. Afterwards, the bounded queue has one entry, namely the task that entered in the i-th time slot, as shown in reference label 1806. At the (i+1)-th time slot 1813, the main CPU issues a new task and enters the bounded queue, as shown in reference label 1807. At the (i+2)-th time slot 1814, the main CPU again issues another task. However, the bounded queue is full, but the GPU is still reasoning about the (i-1)-th task, as shown in reference label 1801. At this time, the system needs to discard the i-th task from the bounded queue, as shown in reference label 1802. Afterwards, the bounded queue is now occupied by the tasks entered in the (i+1)-th time slot and the (i+2)-th time slot, as shown in reference label 1808. Just after this, the GPU completes reasoning on the (i-1)-th task, as shown in reference label 1801. Therefore, the (i+1)-th task is popped out of the bounded queue and enters the GPU for reasoning, as shown in reference label 1803. At this point, the bounded queue has only one entry, as shown by reference label 1809. In the (i+3)th time slot, the CPU issues another task, as shown by time slot 1815. This task enters the bounded queue 1810; and the process repeats itself.

[0190] The timing diagram shows that when the bounded queue is full and the GPU is still busy inferring the previous task, a new task request entering the bounded queue will cause the top task in the queue to be discarded. This mechanism is to prevent the occurrence of large system delays in online mode. Therefore, this mechanism is required in online mode. The resulting frame drops will be supplemented in offline mode; however, this mechanism is not required in offline mode because large system delays are not important in this mode.

[0191] pass Fig.17 The priority arrangement shown and Fig.18The implementation of the bounded queue scheme shown and other delay control methods such as parallel computing enable the system to utilize all GPU hardware resources to perform inference in parallel at runtime to minimize the delay time. The following table shows the performance improvement of the above delay control scheme.

[0192]

[0193] Table 5. Performance improvements of the delay control scheme.

[0194] This table summarizes the experiments performed on various GPU hardware cards with different memory resources.

[0195] The improvement is measured against a target total latency value of less than 120 milliseconds. This value is specified by the end user. The second column shows the optimized latency in terms of average and maximum values ​​using the above latency control scheme; the third column shows the percentage of improvement. As can be seen in this table, the improvement ranges from just over 30% to as high as 49%.

[0196] Fig.15 We show how a computer-assisted gastroscopy system configures hardware resources to prepare for subsequent reasoning in online mode. When presented with a stomach image, the system performs various reasoning tasks based on this configuration. Fig.19 The actual execution flow of inference in online operation mode is shown.

[0197] refer to Fig.19 , the main reasoning process 1900 starts when the stomach image 1903 is received. The system first selects the GPU with the shortest work queue at step 1904. If no GPU is available, the path of 1905 is taken and the image is discarded at step 1906. Then go to result management 1911. If no GPU is available, path 1907 is taken. Then one or more neural models are selected at step 1908 using the adaptive model selection scheme mentioned in the previous paragraph. If no neural model is selected for reasoning, path 1910 is taken to enter the result management step 1911.

[0198] If a neural model is selected, path 1909 is taken and the GPU will use the neural model to start reasoning. If the selected neural model is not a part recognition neural model, path 1913 is taken and the GPU will continue its reasoning task in step 1912. The result is sent to the result manager 1911. Otherwise, path 1914 is taken and the part feature map of the image is extracted in step 1915. This process is computationally intensive. If the system has a single GPU, path 1918 is taken to start reasoning using the GPU in step 1919. Otherwise, path 1916 is taken. Since there is more than one GPU available, the entire part recognition task is divided into subtasks, and each GPU processes one subtask in parallel. Each subtask will produce intermediate results that need to be shared with other GPUs. Therefore, in step 1917, inter-device communication between GPUs is required. Regardless of whether the system is a single GPU or multiple GPUs, the reasoning results will be sent to the result management 1911.

[0199] Fig. 20 1 shows how model selection is performed in an exemplary embodiment. In this embodiment, there are two GPUs, namely GPU-0 and GPU-1. Each GPU is capable of executing neural models NN1, NN2.1, NN2.2, NN3.1, and NN3.2. However, these neural models are assigned to the two GPUs differently based on the available computing resources in GPU-0 and GPU-1 at the time of model selection. In the figure, the shaded area indicates that the GPU has a lot of computing resources available for neural model inference, while the white area indicates that the GPU is already occupied or has less computing resources available. Fig.17 In the adaptive model loading table shown, NN2.1 is the highest priority task, followed by NN3.1, and NN1 is the lowest priority task. Fig. 20 In row 2001 of , the available computing resources of both GPUs are very low. In this example, the first two highest priority tasks NN2.1 and NN3.1 are assigned to GPU-0 for execution, while NN1, NN2.2 and NN3.2 are sent to GPU-1.

[0200] In row 2002, GPU-1 has higher compute resources to choose from. In this case, GPU-0 uses NN1 and NN3.2, while GPU-1 uses NN2.1, NN2.2, and NN3.1. Similarly, row 3 is the opposite of row 2 in terms of GPU resource availability. So in this case, GPU-0 uses NN2.1, NN2.2, and NN3.1, while GPU-1 uses NN1 and NN3.2.

[0201] for Fig. 20In row 2004, there are a lot of computing resources available for both GPUs, and these tasks can be assigned to either GPU at will. In this example, GPU-0 uses NN1, NN2.1, and NN2.2, while GPU-1 uses NN3.1 and NN3.2.

[0202] The following table shows the comparison of inference speed (in frames per second) between the two systems. In this experiment, the CPU is an Intel Core i7-7800x running at a clock frequency of 3.5GHz; the GPU card is a dual-GPU GeForce GTX 1080Ti. This experiment aims to compare the inference speed (in frames per second (FPS)) of processing stomach videos when using one of the GPUs (single-process case) and when using two GPUs (multi-process case). The table shows that for the multi-process system, it can process 51 frames of video images per second, while the single-process system can only process 19 frames.

[0203]

[0204] Table 6. Comparison of inference speed between single process and multi-process

[0205] The systems and methods of the present invention may be implemented in the form of a software application that runs on a computerized system. In addition, portions of the method may be performed on one such computerized system, while other portions are performed on one or more other such computerized systems. Examples of computerized systems include mainframes, personal computers, handheld computers, servers, etc. The software application may be stored on a recording medium that is locally accessible to the computer system and can be accessed through a hardwired or wireless connection to a network, such as a local area network or the Internet.

[0206] The computerized system may include, for example, a processor, a random access memory (RAM), a printer interface, a display unit, a local area network (LAN) data transfer controller, a LAN interface, a network controller, an internal bus, and one or more input devices, such as a keyboard, a mouse, etc. The computerized system may be connected to a data storage device.

[0207] Fig.21 is a schematic diagram of a computerized system 2100 for a multifunctional computer-assisted gastroscopy system according to an embodiment of the present invention, which is composed of hardware and software components that can be used to implement the embodiment of the present invention.

[0208] The hardware components in this embodiment also include a processor 2105, a memory 2111 and multiple interfaces. It may optionally include one or more coprocessors 2110 to accelerate calculations. Multiple components in the computerized system 2100 are connected to an I / O interface 2120, including an input unit 2112, an output unit 2113, a storage unit 2114 and a communication unit 2115, and the communication unit 2115 includes but is not limited to a network card, a modem, a radio communication transceiver, etc. In another embodiment, the present disclosure can also be deployed in a distributed computing environment, which includes more than one computerized system 2100 connected together through one or more network interfaces in the communication unit 2115. The network interface may include one or more of the Internet, an intranet, an extranet, a cellular network, a local area network (LAN), a home area network (HAN), a metropolitan area network (MAN), a wide area network (WAN), a Bluetooth network, a public network, and a private network.

[0209] The processor 2105 may be a central processing unit (CPU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc., and is used to control the overall operation of a memory (e.g., a random access memory (RAM) for temporary data storage, and / or a read-only memory (ROM) for permanent data storage, and firmware). One or more processors may communicate with each other and with the memory and perform operations and tasks that implement one or more blocks of the flowcharts discussed herein.

[0210] Similarly, the coprocessor 2110 may be a graphics processing unit (GPU) card, which includes its own processing unit and random access memory (RAM); or it may be other hardware circuits that accelerate mathematical calculations, such as a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc. One or more coprocessors 2110 may communicate with the processor 2105 and access the memory 2111. They may also communicate with each other and with the memory 2111 and perform operations and tasks that implement one or more blocks of the flowcharts discussed herein.

[0211] The memory 2111 stores, for example, applications, data, programs, algorithms (including software that implements or assists in implementing the example embodiments) and other data. The memory 2111 may include dynamic or static random access memory (DRAM or SRAM) or read-only memory, such as erasable and programmable read-only memory (EPROM), electrically erasable and programmable read-only memory (EEPROM) and flash memory, as well as other memory technologies, alone or in combination.

[0212] Storage devices 2114 typically include persistent memory, such as magnetic disks, such as fixed disks and removable disks; other magnetic media, including magnetic tape; optical media such as compact disks (CDs) or digital versatile disks (DVDs), and semiconductor storage devices such as flash memory cards, solid-state drives, EPROMs, EEPROMS or other storage technologies, used alone or in combination. Note that the instructions of the above software can be provided on a computer-readable or machine-readable storage medium, or alternatively, can be provided on multiple computer-readable or machine-readable storage media distributed in a large system that may have multiple nodes. Such computer-readable or machine-readable media are considered to be part of an article (or article of manufacture). An article or article of manufacture can refer to any manufactured single component or multiple components.

[0213] Input unit 2112 is an interface component that connects computerized system 2100 to a data input device such as a keyboard, a keypad, a pen-based device, a mouse or other pointing device, a voice input device, a scanner or other input technology. According to an embodiment of the present invention, input unit 1812 may include a gastroscopy instrument including an image camera that can be inserted into the gastrointestinal tract. Output unit 2113 is an interface component that computerized system 2100 uses to send data to an output device such as a CRT or flat panel display, a printer, a voice output device, a speaker or other output technology. Communication unit 2115 can generally include a serial or parallel interface and a USB (Universal Serial Bus) interface, as well as other interface technologies. Communication unit 2115 can also enable computerized system 2100 to exchange information with an external data processing device through a data communication network such as a personal area network (PAN), a local area network (LAN), a wide area network (WAN), the Internet and other data communication network architectures. Communication unit 2115 can include an Ethernet interface, a wireless LAN interface device, a Bluetooth interface device and other networking devices alone or in combination.

[0214] The processor 2105 is capable of executing software program instructions stored in the memory 2111. The software program also includes an operating system, and application software such as an endoscope image analysis module. The operating system is responsible for managing all hardware resources and arranging execution priorities for all tasks and processes.

[0215] The blocks and / or methods discussed herein can be performed and / or constructed by a user, a user agent (including a machine learning agent and an intelligent user agent), a software application, an electronic device, a computer, firmware, hardware, a process, a computer system, and / or an intelligent personal assistant. In addition, the blocks and / or methods discussed herein can be automatically executed with or without instructions from a user.

[0216] Those skilled in the art will appreciate that the division between hardware and software is a conceptual division for ease of understanding and is somewhat arbitrary. In addition, it should be appreciated that a peripheral device installed in one computer may be integrated into a host in another computer. In addition, application software systems may be executed in a distributed computing environment. Software programs and their associated databases may be stored in separate file servers or database servers and transferred to a local host for execution. Therefore, if Fig.21 The computerized system 2100 shown is an exemplary embodiment of how the present invention may be implemented. Those skilled in the art will appreciate that alternative embodiments may be suitable for implementing the present invention.

[0217] Therefore, the exemplary embodiments of the present invention are fully described. Although the description relates to specific embodiments, it will be clear to those skilled in the art that the present invention can be practiced by changes in these specific details. Therefore, the present invention should not be interpreted as being limited to the embodiments set forth herein.

[0218] The methods discussed in different figures can be added to or exchanged with the methods in other figures. In addition, specific numerical data values ​​(e.g., specific quantities, numbers, categories, etc.) or other specific information should be interpreted as illustrative for discussing example embodiments. Providing such specific information is not intended to limit the example embodiments.

Claims

1. A computer-aided gastroscopy system, comprising: a central processing unit coupled to a memory storing an executable software program, wherein the software program includes: an AI image processing system for analyzing a sequence of gastric images obtained from a gastroscopy instrument, wherein the AI image processing system includes at least three architecture-level modules that cooperate to perform image quality assessment, lesion detection, cancer identification, HP classification, and location recognition, wherein at least one of the modules includes one or more neural models that extract different but related information from the sequence of gastric images and share the information extracted from the sequence of gastric images with other modules; and at least one of the neural models fuses HP infection features and location information extracted from other neural models to improve the classification accuracy of the computer-aided gastroscopy system, wherein the at least three architecture-level modules include: a first module for image quality control to filter out unqualified images in the sequence of gastric images; a second module for lesion detection, cancer identification, and lesion tracking; and a third module for classifying the HP infection features and location recognition, wherein each of these modules includes one or more neural models, wherein the third module further includes a composite neural model, and the composite neural model includes: a first neural model that takes the sequence of gastric images as input, performs HP feature extraction, and outputs a first number of feature channels; a second neural model that also takes the sequence of gastric images as input, performs location feature extraction including spatio-temporal location information of the gastrointestinal tract, and outputs a second number of feature channels; and a third neural model that takes the concatenation of the first number of feature channels and the second number of feature channels as input and generates a third number of class labels, each of the class labels indicating HP infection features, wherein the second neural model is a composite neural model including an Xception CNN neural model combined with an LSTM model.

2. The system according to claim 1, wherein, the first neural model generates a first tensor of sixty-four-channel elements; the second neural model generates a second tensor of twelve-channel elements, each of the twelve-channel elements corresponding to a location classification label; and the third neural model takes the concatenation of the first tensor and the second tensor as input and outputs nine element classification labels corresponding to the nine HP infection features.

3. The system according to claim 1, further comprising a unified preprocessing module, wherein, the unified preprocessing module takes the sequence of gastric images as input and generates a unified tensor as output for each image in the sequence of gastric images, and the unified tensor is fed to the neural models of the first module, the second module, and the third module of the AI image processing system.

4. The system according to claim 3, wherein, The neural network architectures of the first module, the second module, and the third module are adjusted such that the output tensors of the neural model remain the same as if each neural network architecture used a different preprocessing module designed specifically for the neural network architecture.

5. The system according to claim 3, wherein, the neural network architectures of the first module, the second module, and the third module are adjusted such that the performance of each neural model is not degraded.

6. The system according to claim 3, wherein, if the height or width of an image in the gastric image sequence entering the unified preprocessing module is greater than a threshold, a parallel resizing process is invoked to resize the image, wherein the parallel resizing process comprises the following steps: if the height is odd, the original image is padded with a row of zeros, and if the width is odd, the original image is padded with a column of zeros; the padded image is divided into four quadrants; each quadrant is resized in parallel to produce four resized quadrants; and the four resized quadrants are stitched together to obtain a uniformly resized image.

7. The system according to claim 1, wherein the neural model of the first module is trained using an image quality data set to produce a complete image quality assessment neural model; one or more neural models of the second module are trained using a lesion data set to produce a complete lesion detection neural model; and one or more neural models of the third module are trained using a gastric region data set and a Helicobacter pylori data set to produce a complete HP plus region neural model.

8. The system according to claim 7, further comprising a model pruning and quantization module, wherein, the complete image quality assessment neural model, the complete lesion detection neural model, and the complete HP plus region neural model are optimized by pruning layer connections and quantizing connection weights to respectively produce an optimized image quality assessment neural model, an optimized lesion detection neural model, and an optimized HP plus region neural model.

9. The system according to claim 8, wherein, the computer-aided gastroscopy system further comprises at least one co-processor, and the software program executed at the central processing unit wisely assigns subtasks initiated by each of the modules to the at least one co-processor according to the pre-assigned priorities of the subtasks and the capabilities of each co-processor, such that the computer-aided gastroscopy system can meet the requirements of low-latency response and high-precision detection and classification.

10. The system according to claim 9, wherein, when the computer-aided gastroscopy system is equipped with the at least one co-processor, the computer-aided gastroscopy system can operate in an offline processing mode and an online processing mode.

11. The system according to claim 10, wherein, When the computer-aided gastroscopy system is set to operate in the offline processing mode, the computer-aided gastroscopy system configures each of the at least one coprocessor to run a dynamic batch process, and the dynamic batch process includes the following steps: Loading at least one complete neural model into the coprocessor; and Loading a batch of gastric images into the coprocessor, wherein the batch size is dynamically determined based on the available computing power and resources in the coprocessor.

12. The system according to claim 10, wherein, When the computer-aided gastroscopy system is equipped with the at least one coprocessor and is set to operate in the online processing mode, the computer-aided gastroscopy system executes a delay control program, and the delay control program includes the following steps: Pre-determining the loading priority of each optimized neural model according to the computing resource requirements of each optimized neural model; Loading one or more of the optimized neural models into each coprocessor according to the loading priority and the hardware configuration of the coprocessor; Establishing a task queue with a fixed length for each coprocessor for the central processing unit to issue subtasks to the task queue for the coprocessor to execute; and Enabling each of the at least one coprocessor to operate in parallel, wherein whenever the coprocessor becomes idle, the coprocessor fetches subtasks from the task queue associated with the coprocessor and starts to execute the subtasks whenever the coprocessor becomes idle.

13. The system according to claim 12, wherein, When the computer-aided gastroscopy system obtains gastric images, the computer-aided gastroscopy system executes a resource-aware inference program, and the resource-aware inference program includes the following steps: Selecting the coprocessor with the shortest work queue among the at least one coprocessor as the designated coprocessor; Loading one or more of the optimized neural models into the designated coprocessor based on the neural model priority and the resource availability at the designated coprocessor; Repeating the selection and loading steps so that as many optimized neural models as possible are loaded into one or more of the coprocessors; Enabling each of the neural models to start inference in parallel at each of the at least one coprocessor; Performing inter-coprocessor communication between the coprocessors when the inference task of the neural model is split to run on more than one coprocessor and the intermediate result generated by one of the coprocessors needs to be shared with other coprocessors; and Collecting the inference results of each of the neural models and reporting them back to the computer-aided gastroscopy system.

14. A method for processing a sequence of gastric images by a computer-aided gastroscopy system, comprising: Obtaining the sequence of gastric images from a gastroscopy instrument; Analyzing the sequence of gastric images by an AI image processing system including at least three architecture-level modules that collaboratively perform image quality assessment, lesion detection, cancer identification, HP classification, and lesion site identification At least one of the modules includes one or more neural models that extract different but related information from the gastroscopy image sequence and share the information extracted from the gastroscopy image sequence with other modules; and at least one of the neural models fuses the HP infection features and location information extracted from other neural network models to improve the classification accuracy of the computer-aided gastroscopy examination system; Create a list of subtasks for the computer-aided gastroscopy examination system to execute by each of the at least three architecture-level modules; and When the computer-aided gastroscopy examination system further includes at least one coprocessor and the software program executed at the central processing unit of the computer-aided gastroscopy examination system wisely assigns subtasks to the at least one coprocessor according to the pre-assigned priorities of the subtasks and the capacity and capabilities of each coprocessor, the latency response to the user is reduced, enabling the computer-aided gastroscopy examination system to meet the requirements of achieving low-latency response and high-precision detection and classification, wherein, the analysis step further includes the following steps: Filter out unqualified images in the gastric image sequence through the neural model in the first module; Perform lesion detection, cancer identification, and lesion tracking through at least one of the neural models in the second module; and Classify HP infection and identify gastrointestinal locations through at least one of the neural models in the third module, The classification and identification step further includes the following steps: Output a first number of feature channels through a first neural model that takes the gastric image sequence as input and performs HP feature extraction; Output a second number of feature channels through a second neural model that also takes the gastric image sequence as input and performs location feature extraction including spatio-temporal location information of the gastrointestinal tract; and Generate a third number of class labels through a third neural model that takes the concatenation of the first number of feature channels and the second number of feature channels as input, wherein each class label indicates HP infection features, wherein the second neural model is a composite neural model including an Xception CNN neural model combined with an LSTM model.

15. The method according to claim 14, wherein, The second neural model is a composite neural model including an Xception CNN neural model, an LSTM model, and a final neural model, and the composite neural model is trained according to the following steps: Create an accumulated batch of gastric images from the complete set of labeled gastric images stored in the gastric location dataset; Train an auxiliary neural network using the accumulated batch of gastric images; Copy the entire auxiliary neural network and use the entire auxiliary neural network as the feature extractor of the composite neural model, wherein the connection weights of the feature extractor are not modified during subsequent training; Create a pseudo-video using the complete set of labeled gastric images and collect a predefined number of labeled gastric images from the pseudo-video to form a batch unit; Combine a predefined number of batch units together as a batch group, and send the batch group to the feature extractor to generate a feature tensor of the batch group; Feed the feature tensor into the LSTM model, and the LSTM model generates an intermediate tensor; Send the intermediate tensor to a final neural model including at least one fully connected layer, wherein an output tensor of the final neural model includes a part vector, each element in the part vector is a part label and represents a stomach part position, and the part vector becomes a feature channel of the second neural model.

16. The method according to claim 15, wherein, the pseudo-video is created according to the following steps: Ascendingly sort the labeled stomach images in the stomach dataset according to the indexes of the part labels to obtain a sorted list of stomach images; Create a predefined number of random generators, each of the random generators generates a random number within a predefined random range; Select one of the random generators to generate an initial random number and use the initial random number to select a stomach image from the sorted list as an anchor image; Collect a set of random numbers from each of the multiple random generators, wherein the total number of the collected random numbers is a predefined number, and the predefined number is the batch size of the batch units; Convert the set of random numbers into a list of indexes offset by the indexes of the anchor image; and Use the offset list of indexes to select stomach images from the sorted list to form the batch units.

Citation Information

Patent Citations

  • Method and apparatus with high conductance components for chamber cleaning

    US20220349050A1

  • Abnormity detection used for medical samples under various kinds of setting

    CN107735838A

  • Medical image segmentation based on hybrid context CNN model

    CN110945564A