Helicobacter pylori stomach video full-automatic intelligent analysis system and marking method thereof
The fully automated intelligent video analysis system for Helicobacter pylori gastric infection utilizes a multi-instance learning mechanism to train a deep learning model, enabling fully automated and real-time diagnosis of Helicobacter pylori infection. This solves the problems of high data annotation costs and low automation in existing technologies, thereby improving diagnostic efficiency.
Patent Information
- Application Number
- CN202111338011.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-12
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2041-11-12
AI Technical Summary
Existing Helicobacter pylori identification technologies suffer from high data annotation costs, low levels of automated reasoning, and poor robustness to video data, increasing the workload of physicians.
The fully automated intelligent video analysis system for Helicobacter pylori infection in the stomach is adopted, which includes modules for image acquisition, preprocessing, effective frame screening, anatomical location localization, and deep learning. The deep learning model is trained using a multi-instance learning mechanism to achieve real-time automated diagnosis of Helicobacter pylori infection and reduce manual operation by physicians.
It enables fully automated, real-time diagnosis of Helicobacter pylori infection, reducing the workload of physicians, improving diagnostic efficiency, and reducing judgment errors caused by the intensity of their work.
Smart Images

Figure CN114359131B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical data mining, specifically relating to a fully automated intelligent analysis system for Helicobacter pylori gastric video and its labeling method. Background Technology
[0002] Helicobacter pylori, or H. pylori, is a bacterium that causes progressive damage to the gastric mucosa. It plays a pathogenic role in several important diseases, including duodenal ulcers, gastric ulcers, gastric adenocarcinoma, and gastric mucosa-associated lymphoid tissue (MALT) lymphoma. Almost all patients with gastric cancer have a history of H. pylori infection, and timely detection and eradication of H. pylori infection can effectively reduce the risk of gastric cancer. Furthermore, according to an international risk assessment analysis, even after H. pylori eradication, the risk of developing gastric cancer remains higher than in people who have never had H. pylori. Therefore, regular endoscopic examination is still recommended after H. pylori eradication surgery. Thus, endoscopic examination of H. pylori infection status is becoming increasingly important.
[0003] Experienced gastroenterologists can determine Helicobacter pylori infection status based on endoscopic image features. According to the Kyoto Global Consensus Guidelines for Helicobacter pylori Gastritis, signs of Helicobacter pylori infection include diffuse redness, mucosal swelling, enlarged and serpentine folds, cloudy mucus, punctate redness, and goosebumps. Signs of non-infection with Helicobacter pylori include regularly arranged clusters of small veins. After eradication, signs often include map-like redness and atrophy. The endoscopic signs vary depending on the degree of infection. Visual assessment alone requires a high level of experience, and junior physicians often struggle to correctly identify the infection.
[0004] Patent CN112651375A discloses a Helicobacter pylori gastric image recognition and classification system based on a deep learning model. The system includes an image acquisition device and a computing host. The computing host includes an image preprocessing module and a deep learning model module, which are connected sequentially. The image acquisition device acquires image data of the gastric region to be detected. The image preprocessing module preprocesses the images. The deep learning model module extracts features from the images and then classifies them based on the extracted features. This system can achieve two functions: 1. Extracting features from endoscopic images using a deep learning model to determine the image category, classifying images into three categories: infected, cured after infection, and uninfected. 2. Extracting features from endoscopic images using a deep learning model to identify the location of the acquired gastric region and classifying the image based on the identified location to detect any missed locations and avoid endoscopic detection errors.
[0005] However, this identification and classification system is trained and tested based on static images, resulting in poor robustness to video data. Constructing the training set requires physicians to select specific video frames with signs of infection from each gastroscopy video, which is very time-consuming and labor-intensive. During testing, physicians need to input specific images into the system via a foot switch for analysis, increasing the workload of routine endoscopy examinations. Therefore, it is essential to develop a fully automated intelligent auxiliary analysis system for Helicobacter pylori gastric video based on video training, which requires no physician intervention during testing. This system can quickly and in real-time determine the status of Helicobacter pylori infection, reduce the need for biopsies, and provide better auxiliary diagnostic and reference value for endoscopic examinations by junior physicians. Summary of the Invention
[0006] The purpose of this invention is to solve the problems of high data annotation costs and low automation of reasoning in existing endoscopic Helicobacter pylori identification technologies, and to provide a fully automated intelligent analysis system for Helicobacter pylori gastric video.
[0007] Another objective of this invention is to provide a fully automated real-time labeling method for Helicobacter pylori.
[0008] The technical solution adopted in this invention is:
[0009] A fully automated intelligent video analysis system for Helicobacter pylori in the stomach includes:
[0010] An image acquisition module is used to acquire an input video stream of conventional white light endoscopy of the stomach from an endoscopic imaging system;
[0011] An image preprocessing module is used to preprocess images from a video stream;
[0012] The effective frame module is used to classify and filter images in the video stream to obtain effective frames with clear visual field.
[0013] Anatomical location positioning module, which is used to identify the specific location of the endoscope in the stomach;
[0014] A deep learning module is used to predict the probability of Helicobacter pylori infection from images at different anatomical locations. Then, at each anatomical location, a prediction result is generated using a multi-example voting mechanism. Finally, the prediction results from multiple anatomical locations are combined and voted again to obtain the real-time Helicobacter pylori infection probability during gastroscopy.
[0015] A marker display module is used to display anatomical location information and the probability of Helicobacter pylori infection via video.
[0016] Furthermore, the preprocessing of the image preprocessing module is to process the image into the input form required by the deep learning module. The processing includes one or more of image normalization, invalid pixel cropping, and image scaling.
[0017] Furthermore, the operation of the effective frame module includes: using deep learning technology to classify video frames into four categories: external frames, narrowband endoscopy frames, invalid examination frames, and effective examination frames, thereby filtering out interfering images and retaining clear endoscopic examination images.
[0018] Furthermore, the anatomical location positioning module uses deep learning technology to identify the specific location of the endoscope in the stomach, including one or more of the following locations: gastric fundus, gastric body, gastric angle, and gastric antrum.
[0019] Furthermore, the deep learning module is used for Helicobacter pylori infection prediction. Its training process is based on a multi-instance learning mechanism and video-level labeling. It includes a model training unit and a model testing unit. The model training unit is used to build a deep learning model, and the model testing unit is used to predict the Helicobacter pylori infection rate in gastroscopy videos in real time.
[0020] Furthermore, the model training unit is performed through the following steps:
[0021] The S200 collects a large number of gastroscopy videos, categorized by the type of examination. 14 C-urea breath test and pathological tissue section examination results serve as video-level labels for Helicobacter pylori infection;
[0022] S201 filters out clear single-frame images from the video using the effective frame module, and then performs data augmentation.
[0023] S202 divides the single-frame image into different anatomical locations using the anatomical location positioning module;
[0024] S203 is based on a multi-instance learning mechanism, which eliminates the need for doctors to screen specific infection frame images. It only uses video-level Helicobacter pylori infection labels to train a deep learning model for predicting the probability of Helicobacter pylori infection in a single frame image at each anatomical location.
[0025] Furthermore, the model testing unit is performed through the following steps:
[0026] The S300 uses the valid frame module to filter and retain images with clear visual information from the input video stream.
[0027] S301 uses the anatomical location localization module to obtain the anatomical location category of each frame of image;
[0028] S302 first obtains the probability of Helicobacter pylori infection for each frame of image using a Helicobacter pylori infection prediction model corresponding to the anatomical location. Then, it generates the prediction result for each anatomical location through a multi-example voting mechanism. Finally, it votes again on the infection probabilities of all anatomical locations to obtain the patient-level Helicobacter pylori infection probability.
[0029] Furthermore, the operation of the marker display module includes: refreshing and displaying the infection probability of the current frame, the overall infection probability of the current video, the maximum infection probability at each anatomical location, and the corresponding image, based on the output information of the deep learning module.
[0030] Compared with existing technologies, the fully automated intelligent video analysis system for Helicobacter pylori in the stomach provided by this invention has the following advantages:
[0031] 1) Fully automated real-time assessment of Helicobacter pylori infection status during gastroscopy using a deep learning module;
[0032] 2) The prediction model is trained based on video-level labels and adopts a multi-instance learning mechanism. The construction of training data for the deep learning module does not require physicians to manually screen and label Helicobacter pylori infection labels for single-frame images. Only the video-level infection labels for each examination video need to be obtained based on pathology and breath results.
[0033] 3) This invention eliminates the need for doctors to step on a pedal and input specific images for Helicobacter pylori infection identification during testing. The entire process requires no human intervention, achieving full automation and high intelligence.
[0034] 4) The marking and display module refreshes and displays the anatomical location information and the probability of Helicobacter pylori infection on the monitor, and refreshes and prompts the unexamined anatomical locations according to the queue of examined locations. This can help alleviate the high-intensity and long-term work of doctors in reading images, avoid subjective judgment errors caused by work intensity and working time, reduce the workload of doctors, and improve the efficiency of medical diagnosis.
[0035] A fully automated real-time labeling method for Helicobacter pylori, which employs the intelligent analysis system described above, includes the following specific steps:
[0036] S100 obtains the routine white light endoscopy video stream of the patient's stomach through the image acquisition module and the endoscopy imaging system equipment, and refreshes and displays it on the marker display module;
[0037] S101 preprocesses the video stream using an image preprocessing module;
[0038] S102 classifies and filters the video stream through the effective frame module to obtain effective frame images;
[0039] S103 uses the anatomical location positioning module to mark and display the anatomical location type of the current frame, refreshes and displays the coverage of the currently checked locations, and prompts the remaining unchecked locations.
[0040] S104 uses a deep learning module to predict the probability of Helicobacter pylori infection in all valid frame images.
[0041] Based on the prediction results of the deep learning module, S105 refreshes and displays the infection probability of the current frame, the overall infection probability of the current video, the maximum infection probability at each dissection location, and the corresponding image.
[0042] The fully automated real-time labeling method for Helicobacter pylori provided by this invention employs the aforementioned intelligent analysis system. Since the beneficial effects of the intelligent analysis system have been described in detail above, they will not be repeated here. Attached Figure Description
[0043] Figure 1 This is a schematic diagram of a fully automated intelligent video analysis system for Helicobacter pylori in the stomach provided in Embodiment 1 of the present invention;
[0044] Figure 2 This is a structural diagram of the model training unit of a fully automated intelligent video analysis system for Helicobacter pylori in the stomach provided in Embodiment 1 of the present invention;
[0045] Figure 3 This is a structural diagram of a model test unit for a fully automated intelligent video analysis system for Helicobacter pylori in the stomach provided in Embodiment 1 of the present invention;
[0046] Figure 4 This is a schematic diagram of the display results of the marking and display module in a fully automated intelligent video analysis system for Helicobacter pylori in the stomach provided in Embodiment 1 of the present invention;
[0047] Figure 5 The flowchart is provided for a fully automated real-time labeling method for Helicobacter pylori according to Embodiment 2 of the present invention. Detailed Implementation
[0048] In the description of this invention, it should be understood that the terms "one end", "the other end", "outer side", "upper", "inner side", "horizontal", "coaxial", "center", "end", "length", "outer end", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing this invention and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.
[0049] The invention will now be further described with reference to the accompanying drawings.
[0050] Example 1
[0051] Figure 1 The diagram illustrates the principle of a fully automated intelligent video analysis system for Helicobacter pylori gastric images according to this embodiment. The intelligent analysis system runs on a computer system and includes an image acquisition module, an image preprocessing module, an effective frame module, an anatomical location positioning module, a deep learning module, and a marker display module, which are connected in sequence.
[0052] The image acquisition module acquires the input gastric conventional white light endoscopy video stream in real time through the endoscopy imaging system.
[0053] The image preprocessing module is used to process the images in the video stream into the input format required by the deep learning model. The image enhancement processing includes: image normalization, invalid pixel cropping, and image scaling pixel cropping.
[0054] The effective frame module uses deep learning technology to classify video frames into four categories: external frames, narrowband endoscope frames, invalid examination frames, and effective examination frames. This filters out non-white light images, motion blur, instrument manipulation, bloodstains, and other interfering images, while retaining clear endoscopic examination images.
[0055] The anatomical location module is used to obtain the stomach location information of the current frame, including the four main gastric anatomical locations: fundus, body, antrum, and pylorus.
[0056] The deep learning module is used to predict Helicobacter pylori infection status in endoscopic videos in real time, and includes a model training unit and a model testing unit.
[0057] Continue reading Figure 2 Explanation of the model training unit: Since Helicobacter pylori infection can occur in various anatomical locations of the stomach, and the anatomical structure and signs of Helicobacter pylori infection differ between these locations, directly training a deep learning model for each anatomical location together would significantly interfere with the network's convergence direction due to the structural information of the anatomical location. This would prevent the effective learning of Helicobacter pylori infection information. Therefore, in this embodiment, to decouple anatomical location information from Helicobacter pylori infection information, separate multi-instance learning-based Helicobacter pylori infection prediction models are trained for each different anatomical location. The specific process is as follows:
[0058] The S200 collects a large number of gastroscopy videos, categorized by the type of examination. 14 C-urea breath test and pathological tissue section examination results serve as video-level labels for Helicobacter pylori infection;
[0059] S201 uses the effective frame module of this embodiment to filter images with clear visual field in each video, and then performs data augmentation by translation transformation, mirror flipping or random cropping to expand the dataset;
[0060] S202 uses the anatomical location positioning module of this embodiment to divide the above image into different anatomical locations, including four main anatomical locations: gastric fundus, gastric body, gastric angle, and gastric antrum.
[0061] S203 employs a weakly supervised mechanism based on multi-instance learning. For videos containing Helicobacter pylori infection, the images with the highest predicted probabilities should possess morphological features associated with Helicobacter pylori infection, such as mucus and redness, while images with lower predicted probabilities may not necessarily exhibit infection features. Conversely, for negative videos, none of the images should contain infection features; that is, the images with the highest predicted probabilities should not contain Helicobacter pylori infection features. Based on this premise, this invention uses a multi-instance learning mechanism. In each iteration of network training, this invention employs a two-step training method:
[0062] Step 1:
[0063] For each anatomical location, a deep convolutional network model is used to calculate the infection probability of all video frames for each patient, and then the images are sorted to select the top-10 images with the highest infection probability for each anatomical location for each patient.
[0064] Step 2:
[0065] The corresponding video tags are used as pseudo-tags for the selected top-10 images. The loss of each image is calculated, and backpropagation is performed to update the weight parameters of the convolutional network model.
[0066] Repeat the above steps until the network converges.
[0067] The convolutional network model used in this invention is based on ResNet18, and the weights are initialized based on a pre-trained model of Image-Net.
[0068] The construction of training datasets for conventional endoscopic Helicobacter pylori identification systems often requires physicians to invest a significant amount of time and effort in selecting images with signs of Helicobacter pylori infection. However, this embodiment, based on a multi-instance learning mechanism, uses only the results obtained through breath and pathology examination reports as labels for each video to train the deep learning model. No annotation work is required from physicians, which can make full use of the hospital's past unlabeled historical data and has the advantage of large-scale promotion.
[0069] In addition, to mitigate the interference of class imbalance, in each iteration of the model, this invention reads several images from the positive image data loader and the negative image loader respectively, forming a batch of training data with a balanced positive and negative ratio, so that the model achieves a better balance between prediction specificity and sensitivity, and has better robustness.
[0070] Continue reading Figure 3 Description of the model prediction unit: It is responsible for real-time, fully automated prediction of Helicobacter pylori infection during gastroscopy. Typical endoscopic Helicobacter pylori identification systems often require physicians to input specific frames into the analysis system via foot pedals or similar operations. This embodiment, however, fully utilizes a weakly supervised multi-instance learning mechanism. During testing, the top-10 video frames with the highest predicted probabilities for each anatomical location are first voted on to obtain the anatomical location-level prediction probability. Then, the infection probabilities for each anatomical location are voted on again to obtain the video-level prediction probability. This eliminates the need for physician input of specific frames, exhibiting fully automated characteristics. Its main process is as follows:
[0071] The S300 uses the valid frame module to filter and retain images with clear visual information from the input video stream.
[0072] S301 determines the anatomical location category of each frame image through the anatomical location positioning module;
[0073] S302 uses the Helicobacter pylori infection prediction model for the corresponding anatomical location to obtain the Helicobacter pylori infection probability of each frame image, and performs cumulative voting on the top-10 images with the highest infection probability at each anatomical location to obtain the cumulative average infection probability at the current time for each anatomical location.
[0074] S303 votes on the infection probability of all anatomical locations to obtain the cumulative average Helicobacter pylori infection probability at the current moment of the gastroscopy video.
[0075] Compared with similar endoscopic Helicobacter pylori infection analysis systems, this embodiment has the following advantages.
[0076] 1. In this embodiment, the Helicobacter pylori infection prediction model is based on a deep learning model with a multi-instance learning mechanism. Its training set does not require doctors to manually select and label single-frame static images. It only needs to obtain the patient's Helicobacter pylori infection results from pathology and breath reports as the corresponding video-level annotations, which greatly reduces the doctor's annotation workload. This is different from any previous endoscopy-based Helicobacter pylori infection prediction model. It belongs to a weakly supervised learning model and has good adaptability and generalization for large-scale video datasets.
[0077] 2. Multi-level voting mechanism: In this embodiment, Helicobacter pylori infection prediction is divided into three levels: image level, anatomical location level, and patient level. The multi-level voting mechanism has very good stability and robustness and can effectively avoid interference from outliers and extreme values.
[0078] 3. The overall analysis and prediction process is fully automated, without requiring doctors to select specific images (similar systems often require doctors to press a foot switch to input a specific image for model analysis). For full video stream analysis, it can predict Helicobacter pylori infection without increasing the workload of endoscopists.
[0079] Example 2
[0080] Please see Figure 5 A fully automated real-time labeling method for Helicobacter pylori, employing the intelligent analysis system described above, includes the following specific steps:
[0081] S100 obtains the routine white light endoscopy video stream of the patient's stomach through the image acquisition module and the endoscopy imaging system equipment, and refreshes and displays it on the marker display module;
[0082] S101 preprocesses the video stream using an image preprocessing module;
[0083] S102 classifies and filters the video stream through the effective frame module to obtain effective frame images;
[0084] S103 uses the anatomical location positioning module to mark and display the anatomical location type of the current frame, refreshes and displays the coverage of the currently checked locations, and prompts the remaining unchecked locations.
[0085] S104 uses a deep learning module to predict the probability of Helicobacter pylori infection in all valid frame images.
[0086] Based on the prediction results of the deep learning module, S105 refreshes and displays the infection probability of the current frame, the overall infection probability of the current video, the maximum infection probability at each dissection location, and the corresponding image.
[0087] Furthermore, in S100, the obtained routine white light endoscopy video stream of the patient's stomach is refreshed and displayed on the left half of the screen for the physician to observe;
[0088] Furthermore, in S105, while obtaining the Helicobacter pylori infection probability of the current image, it is added to the probability value sequence of the corresponding anatomical location. The top-10 Helicobacter pylori infection probabilities of the current anatomical location are calculated and updated, and a vote is taken to obtain the infection probability of that anatomical location. Then, the Helicobacter pylori infection probabilities of the examined locations are voted again to obtain the cumulative average Helicobacter pylori infection probability of the current video. At the same time, the frame image with the highest infection probability under each anatomical location is compared and updated. Finally, this information and images are updated and displayed in the lower right of the screen to remind the doctor of the current Helicobacter pylori infection status.
[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A full-automatic intelligent analysis system for Helicobacter pylori in stomach by video, characterized in that, Comprise: An image acquisition module for acquiring an input regular white light gastroscope video stream from an endoscope image system; An image preprocessing module for preprocessing images of the video stream; An effective frame module for classifying and screening video stream images, obtaining effective frames with clear field of view, the operation of the effective frame module comprising: using deep learning technology to classify video frames into four categories: extracorporeal frame, narrow band endoscope frame, invalid examination frame, and effective examination frame, thereby filtering out interfering images and retaining endoscopic examination images with clear field of view; An anatomical position positioning module for identifying the specific position of the endoscope currently located in the stomach; A deep learning module for predicting the probability of Helicobacter pylori infection for pictures of different anatomical positions, then generating the prediction result of each anatomical position by a voting mechanism, and finally combining the prediction results of multiple anatomical positions to obtain the real-time Helicobacter pylori infection probability of gastroscope by voting again, the deep learning module for Helicobacter pylori infection prediction, the training process of which is based on a multiple-instance learning mechanism and video-level label, which includes a model training unit and a model testing unit, the model training unit is used to establish a deep learning model, and the model testing unit is used to predict the Helicobacter pylori infection rate; and A mark display module for displaying anatomical position information and Helicobacter pylori video infection probability, the operation of the mark display module comprising: refreshing the infection probability of the current frame, the overall infection probability of the current video, the maximum infection probability of each anatomical position and the picture corresponding thereto according to the output information of the deep learning module; The model training unit is performed by the following steps: S200 Collect a large number of gastroscopy videos, corresponding to each examination 14 C Urea breath test and pathological tissue section examination results as video-level H. pylori infection labels; S201 screening single-frame images with clear field of view in the video by the effective frame module, and then performing data augmentation; S202 dividing the single-frame images into different anatomical positions by the anatomical position positioning module; S203 training a deep learning model for each anatomical position based on a multiple-instance learning mechanism; The model testing unit is performed by the following steps: S300 filtering images with unclear field of view using the effective frame module to retain pictures with clear field of view for the input video stream; S301 obtaining the anatomical position category of each frame image using the anatomical position positioning module; S302 obtaining the Helicobacter pylori infection probability of each frame picture using the Helicobacter pylori infection prediction model corresponding to the anatomical position, then generating the prediction result of each anatomical position by a voting mechanism, and finally voting again all the infection probabilities of the anatomical positions to obtain the Helicobacter pylori infection probability at the patient level.
2. The full-automatic intelligent analysis system for H. pylori stomach video according to claim 1, characterized in that, The preprocessing of the image preprocessing module is to process the images into the input form required by the deep learning module, which includes one or more of image normalization, invalid pixel cropping, and image scaling.
3. The full-automatic intelligent analysis system for H. pylori stomach video according to claim 1, characterized in that, The anatomical position positioning module uses deep learning technology to identify the specific position of the endoscope currently located in the stomach.
4. A fully automated real-time labeling method for H. pylori, characterized by, The intelligent analysis system is implemented by using any one of the intelligent analysis systems in claims 1-3, and the specific steps include: S100 obtains a conventional white light endoscope video stream of a stomach of a subject through an image acquisition module and an endoscope image system device and refreshes and displays on a marker display module; S101 pre-processes the video stream through an image pre-processing module; S102 classifies and filters the video stream through an effective frame module to obtain effective frame images; S103 marks and displays an anatomical position type of a current frame through an anatomical position positioning module, refreshes and displays a coverage rate of a current checked position, and prompts a remaining unchecked position; S104 predicts a Helicobacter pylori infection probability of all effective frame images through a deep learning module; S105 refreshes and displays an infection probability of the current frame, an overall infection probability of the current video, a maximum infection probability of each anatomical position, and a picture corresponding to the maximum infection probability according to a prediction result of the deep learning module.
Citation Information
Patent Citations
Digestive endoscopy image abnormal feature real-time labeling system and method
CN108852268A
Helicobacter pylori stomach image recognition and classification system based on deep learning model
CN112651375A