A method and system for detecting lost frames in gastroscopy vision

By preprocessing and segmenting the gastroscopy pictures, the classification model is used to detect visual field loss and analyze the causes, and the examination time and noise problems caused by visual field loss in gastroscopy are solved, achieving rapid and accurate detection and assisting doctors in operation.

CN114663703BActive Publication Date: 2025-06-10ZIDONG INFORMATION TECH (SUZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210290114.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-09
Filing Date
2022-03-23
Publication Date
2025-06-10
Estimated Expiration
2042-03-23

AI Technical Summary

Technical Problem

During gastroscopy, doctors will lose their vision, which will lead to prolonged examination time and increased patient pain, and will bring noise to artificial intelligence image analysis algorithms and reduce work efficiency.

Method used

By preprocessing and dividing the gastroscopic image into multiple sub-pictures, the trained first classification model is used for label classification. If the field of view is lost, the second classification model will be entered, and the difference characteristics of the current frame and adjacent frames are classified to obtain the category of causes of the field of view loss.

Benefits of technology

It realizes rapid and accurate detection of pictures of the visual field loss and finds the cause of the visual field loss, assists the doctor in operating gastroscopy to reduce examination time and patient pain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114663703B_ABST
    Figure CN114663703B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for detecting lost-frame of gastroscope vision. The method includes the following steps: S1, preprocess the current-frame gastroscope image; S2, divide the preprocessed current-frame gastroscope image into multiple sub-images; S3, use the trained first classification model to classify the labels of the multiple sub-images one by one, and obtain the classification result through the sub-image dynamic decision rule. If the classification result is lost vision, go to step S4; if the classification result is no lost vision, return to step S1 and continue to detect the next-frame gastroscope image; S4, use the trained second classification model to classify the current-frame gastroscope image with a classification result of lost vision and its difference features from adjacent-frame gastroscope images, and obtain the cause category of the lost vision of the current-frame gastroscope image. The present invention can quickly and accurately detect the images with lost vision and find the cause of the lost vision to assist doctors in operating the gastroscope.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gastroscope image processing, and particularly relates to a method and system for detecting lost frames in the gastroscope field of view. Background Art

[0002] The gastroscope is a conventional examination method in digestive medicine. During the gastroscope examination, the doctor inserts a thin tube into the stomach to directly observe the lesions in parts such as the esophagus, stomach, and duodenum. The gastroscope examination can directly observe the real situation of the stomach and is the preferred examination method for upper digestive tract lesions.

[0003] Although there has already been a technology that uses an artificial intelligence gastroscope image analysis algorithm to analyze gastroscope videos. However, when doctors perform gastroscope operations, there will be a situation of lost field of view. The lost-field-of-view pictures are some pictures that do not contain any gastric mucosa information, and no information about the gastric part, diseases, etc. can be obtained from the lost-field-of-view pictures. According to data statistics, the video time of the lost field of view accounts for about 20% of the total gastroscope time. On the one hand, the lost field of view prolongs the gastroscope duration and increases the patient's pain; on the other hand, these lost-field-of-view video frames bring noise to the artificial intelligence gastroscope image analysis algorithm, increase the extra meaningless workload, and reduce the work efficiency. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for detecting lost frames in the gastroscope field of view with high feasibility and high detection accuracy.

[0005] To solve the above problems, the present invention provides a method for detecting lost frames in the gastroscope field of view, which includes the following steps:

[0006] S1. Preprocess the current-frame gastroscope picture;

[0007] S2. Cut the preprocessed current-frame gastroscope picture into multiple sub-pictures;

[0008] S3. Use the trained first classification model to classify the labels of the multiple sub-pictures one by one, and obtain the classification result through the sub-picture dynamic decision rule. If the classification result is a lost field of view, go to step S4; if the classification result is a non-lost field of view, return to step S1 and continue to detect the next-frame gastroscope picture;

[0009] S4. Use the trained second classification model to classify the current-frame gastroscope picture with a classification result of lost field of view and its difference features from adjacent-frame gastroscope pictures to obtain the cause category of the lost field of view of the current-frame gastroscope picture.

[0010] As a further improvement of the present invention, the second classification model includes a shared image representation layer and a classification layer, and the input of the shared image representation layer is (xt , x t+1 -x t ); where x t is the representation vector of the current frame gastroscope image, and x t+1 is the representation vector of the next frame gastroscope image. The intermediate representation vector z of the current frame gastroscope image is obtained through the shared image representation layer 1 and the intermediate representation vector z of the image difference 2 ;

[0011] The input of the classification layer is the intermediate representation vector z of the current frame gastroscope image 1 and the intermediate representation vector z of the image difference 2 The representation vector z after splicing t , and the cause category of the current frame field of view loss is obtained through the classification layer.

[0012] As a further improvement of the present invention, the classification layer includes a linear classification layer and a softmax layer. The intermediate representation vector z of the current frame gastroscope image 1 and the intermediate representation vector z of the image difference 2 The representation vector z after splicing t passes through the linear classification layer to obtain the prediction output v of each cause category i , and the prediction output v of each cause category i is input to the softmax layer to obtain the probability value of each cause category.

[0013] As a further improvement of the present invention, the softmax function of the softmax layer is as follows:

[0014]

[0015] where P i is the probability value of the i-th cause category, N is the total number of cause categories, and v i is the prediction output of the i-th cause category; v n is the prediction output of the n-th cause category; e is the base of the natural logarithm function.

[0016] As a further improvement of the present invention, in step S3, the classification result is obtained through the subgraph dynamic decision rule, including: if a subgraph decision is successful, the program is terminated in advance to obtain the classification result; otherwise, the next subgraph is continued to be tested. If all subgraphs fail to make a decision, the prediction probabilities of all subgraphs are fused to obtain the final classification result.

[0017] As a further improvement of the present invention, the subgraph dynamic decision rule adopts the following decision formula:

[0018] p(i) = max(xic ),m < n

[0019]

[0020] where c = 0, 1, 2......, C - 1;

[0021] When fusing the prediction probability results in the sub - graph dynamic decision rule, the following probability fusion formula is adopted:

[0022]

[0023]

[0024] where c = 0, 1, 2......, C - 1; m is the number of input pictures; n is the number of sub - graphs obtained by segmentation; c is the sub - script of the label; i is the sub - script of the input picture; C is the number of picture categories; λ is the decision threshold obtained by training; x ic represents the prediction probability of the c - th picture category of the i - th input picture; p(c) represents the fusion probability of the c - th picture category; label is the picture category information predicted by the model.

[0025] As a further improvement of the present invention, the pre - processing includes one or more of the following processes: scaling and cropping processing, random horizontal flipping processing, normalization processing, and picture cutting processing.

[0026] The present invention also provides a gastroscope field - of - view loss frame detection system, which includes the following modules:

[0027] A pre - processing module for pre - processing the current - frame gastroscope picture;

[0028] A picture segmentation module for segmenting the pre - processed current - frame gastroscope picture into multiple sub - graphs;

[0029] A first classification module for classifying the labels of the multiple sub - graphs one by one by using a trained first classification model, and obtaining a classification result through a sub - graph dynamic decision rule;

[0030] A judgment module for judging whether the classification result is a field - of - view loss. If so, it enters the second classification module; otherwise, it returns to the pre - processing module to continue detecting the next - frame gastroscope picture;

[0031] A second classification module for classifying the current - frame gastroscope picture with a classification result of field - of - view loss and the difference features between it and adjacent - frame gastroscope pictures by using a trained second classification model, and obtaining the cause category of the field - of - view loss of the current - frame gastroscope picture.

[0032] As a further improvement of the present invention, the second classification model includes a shared image representation layer and a classification layer, and the input of the shared image representation layer is (x t ,x t+1 -x t ), where x t is the representation vector of the gastroscopy image of the current frame, x t+1 is the representation vector of the next gastroscope image frame, and the intermediate representation vector z of the current gastroscope image frame is obtained through the shared image representation layer. 1 The intermediate representation vector z of the image difference 2 ;

[0033] The input of the classification layer is the intermediate representation vector z of the current frame gastroscope image 1 The intermediate representation vector z of the image difference 2 The representation vector z after splicing t , the cause category of the current frame field of view loss is obtained through the classification layer.

[0034] As a further improvement of the present invention, the classification layer includes a linear classification layer and a softmax layer, and the intermediate representation vector z of the current frame gastroscope image is 1 The intermediate representation vector z of the image difference 2 The representation vector z after splicing t After the linear classification layer, the predicted output v for each cause category is obtained. i , the predicted output v for each cause category i The input is sent to the softmax layer to obtain the probability value of each cause category.

[0035] Beneficial effects of the present invention:

[0036] The gastroscope field of view loss frame detection method of the present invention divides the gastroscope image into multiple sub-images, classifies the multiple sub-images by labels, and obtains the classification results; and classifies the difference features between the current frame gastroscope image and the adjacent frame gastroscope image to obtain the cause category of the current frame field of view loss. The present invention can quickly and accurately detect the image with field of view loss and find the cause of field of view loss to assist doctors in operating the gastroscope.

[0037] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the following specifically cites a preferred embodiment and describes it in detail with the accompanying drawings as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a flow chart of a method for detecting lost frames based on gastroscopic field of view in a preferred embodiment of the present invention. Specific Embodiment

[0039] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited are not intended to limit the present invention.

[0040] As Figure 1 shown, it is a method for detecting lost frames in the gastroscope field of view in a preferred embodiment of the present invention, including the following steps:

[0041] S1. Preprocess the current-frame gastroscope image;

[0042] Optionally, the preprocessing includes one or more of the following processes: scaling and cropping processing, random horizontal flipping processing, normalization processing, and image cutting processing.

[0043] Among them, the scaling and cropping processing is used to process the image into a fixed size. The normalization processing means subtracting the statistical average value of the data corresponding dimension in the RGB dimension of the image to eliminate the common part and highlight the features and differences between individuals. The random horizontal flipping processing is also for data augmentation to improve the generalization ability of the model. In this embodiment, the value of image scaling and cropping is not limited. Optionally, the sizes of different images are scaled to 640*640*3, and then cropped to 384*384*3, and the black redundant parts at the four corners of the image are cut off.

[0044] S2. Cut the preprocessed current-frame gastroscope image into multiple sub-images;

[0045] Optionally, the preprocessed gastroscope image is cut into n sub-images, where n is an integer multiple of 4. It can reduce the size of the image and expand the number of samples to n times the original.

[0046] S3. Use the trained first classification model to classify the labels of each of the multiple sub-images one by one, and obtain the classification result through the sub-image dynamic decision rule. If the classification result is a lost field of view, go to step S4; if the classification result is a non-lost field of view, return to step S1 and continue to detect the next-frame gastroscope image.

[0047] Optionally, the first classification model is the lightweight neural network model RseNet18. It has the advantages of smaller video memory, faster speed, and higher accuracy.

[0048] Furthermore, obtaining the classification result through the sub-image dynamic decision rule includes: if one sub-image decision is successful, the program ends in advance to obtain the classification result; otherwise, continue to test the next sub-image. If all sub-images fail to make a decision, the prediction probabilities of all sub-images are fused to obtain the final classification result.

[0049] Optionally, the sub - graph dynamic decision rule adopts the following decision formula:

[0050] p(i) = max(x ic ), m < n

[0051]

[0052] where c = 0, 1, 2......, C - 1;

[0053] When fusing the prediction probability results in the sub - graph dynamic decision rule, the following probability fusion formula is adopted:

[0054]

[0055]

[0056] where c = 0, 1, 2......, C - 1; m is the number of input pictures; n is the number of sub - graphs obtained by segmentation; c is the subscript of the label; i is the subscript of the input picture; C is the number of picture categories; λ is the decision threshold obtained by training; x ic represents the prediction probability of the c - th picture category of the i - th input picture; p(c) represents the fusion probability of the c - th picture category; label is the picture category information predicted by the model.

[0057] In this embodiment, C = 2, that is, the number of picture categories is 2; c = 0, 1, respectively representing vision loss and no vision loss. Specifically, it can be: c = 0 represents vision loss, c = 1 represents no vision loss, or c = 0 represents no vision loss, c = 1 represents vision loss;

[0058] When c = 0 represents vision loss and c = 1 represents no vision loss, the form of x ic is: x i0 is the probability of vision loss, and x i1 is the probability of no vision loss.

[0059] The maximum value of the probability is obtained through the above max() function: p(i) = max(x ic ), that is, take the maximum value of x i0 and x i1 .

[0060] Then judge whether p(i)>λ holds;

[0061] If it does not hold (i.e., the decision is unsuccessful): then continue to make a decision on the next sub - graph;

[0062] If it holds (i.e., the decision is successful): Then the predicted class label label is obtained through the above argmax() function, and the values of label are 0 and 1, representing visual field loss and no visual field loss respectively.

[0063] If all sub - graphs do not meet the decision conditions:

[0064]

[0065]

[0066] Then the probabilities corresponding to all sub - graphs are accumulated and averaged through the above p(c) formula, and then the predicted class label label is obtained through the above argmax() function. The values of label are 0 and 1, representing visual field loss and no visual field loss respectively.

[0067] Specifically, after the gastroscope image is segmented into n sub - graphs, m images are sequentially selected as the model input images in order. When the following conditions are met, the system makes a decision:

[0068] If m < n, when the predicted probability p(i) of the i - th sub - graph > λ, the decision is successful, and the predicted label is the subscript of the maximum probability of the current image, and the program ends prematurely; otherwise, the decision fails and the next sub - graph is continued to be tested.

[0069] If all sub - graphs have decision failures, the predicted probabilities of all sub - graphs are fused to obtain the final classification result.

[0070] Optionally, the following loss function is adopted when training the first classification model:

[0071]

[0072] Where N is the number of samples, i is the sample number, y is the sample label value, is the predicted probability, y(i) represents the label value of the i - th sample, is the predicted probability value of the i - th sample.

[0073] S4. Use the trained second classification model to classify the current frame gastroscope image with a classification result of visual field loss and its difference features from adjacent frame gastroscope images to obtain the cause category of visual field loss of the current frame gastroscope image. Among them, the cause categories of visual field loss include but are not limited to the following three: too fast moving speed, close to the digestive tract wall, others (light, noise, etc.).

[0074] Specifically, the second classification model includes a shared image representation layer and a classification layer. The input of the shared image representation layer is (x t ,x t+1 -x t); where x t is the representation vector of the current frame gastroscope image, and x t+1 is the representation vector of the next frame gastroscope image. The intermediate representation vector z 1 of the current frame gastroscope image and the intermediate representation vector z 2 of the image difference are obtained through the shared image representation layer;

[0075] Optionally, the shared image representation layer can be composed of Vision Transformer (ViT) or ResNet18. When it is ViT, the intermediate representation vector z 1 of the current frame gastroscope image and the intermediate representation vector z 2 of the image difference are obtained as follows:

[0076] z 1 , z 2 = f ViT (x t , x t+1 - x t )

[0077] where f ViT is the ViT network function.

[0078] The input of the classification layer is the representation vector z 1 of the intermediate representation of the current frame gastroscope image and the representation vector z 2 of the intermediate representation of the image difference after splicing. The cause category of the current frame field of view loss is obtained through the classification layer. t

[0079] Furthermore, the classification layer includes a linear classification layer and a softmax layer. The representation vector z 1 of the intermediate representation of the current frame gastroscope image and the representation vector z 2 of the intermediate representation of the image difference after splicing t pass through the linear classification layer to obtain the predicted output v i of each cause category. The predicted output v i of each cause category is input into the softmax layer to obtain the probability value of each cause category.

[0080] Among them, the representation vector z t is expressed as: z t = [z 1 , z 2 .

[0081] The softmax function of the softmax layer is as follows:

[0082] ​

[0083] Among them, P i is the probability value of the i-th cause category, N is the total number of cause categories, and v i is the predicted output of the i-th cause category; v n is the predicted output of the n-th cause category; e is the base of the natural logarithm function.

[0084] Optionally, both the first classification model and the second classification model are trained using the gradient descent algorithm.

[0085] The second classification model uses the following loss function during training:

[0086]

[0087] Among them, M represents the number of pictures in a video segment, N represents the total number of categories of gastroscope pictures, and y i represents the label value of the i-th picture, represents the predicted output of the i-th picture.

[0088] A preferred embodiment of the present invention also discloses a gastroscope field-of-view loss frame detection system, which includes the following modules:

[0089] A preprocessing module for preprocessing the current-frame gastroscope picture;

[0090] A picture segmentation module for segmenting the preprocessed current-frame gastroscope picture into multiple sub-pictures;

[0091] A first classification module for using the trained first classification model to classify the labels of the multiple sub-pictures one by one, and obtaining a classification result through a sub-picture dynamic decision rule;

[0092] A judgment module for judging whether the classification result is a field-of-view loss. If so, it enters the second classification module; otherwise, it returns to the preprocessing module to continue detecting the next-frame gastroscope picture;

[0093] A second classification module for using the trained second classification model to classify the current-frame gastroscope picture with a classification result of field-of-view loss and its difference features from adjacent-frame gastroscope pictures, and obtaining the cause category of the field-of-view loss of the current-frame gastroscope picture.

[0094] Specifically, the second classification model includes a shared image representation layer and a classification layer, and the input of the shared image representation layer is (x t , x t+1 -x t ); among them, x t is the representation vector of the current-frame gastroscope picture, and x t+1is the representation vector of the next frame of gastroscope image, and the intermediate representation vector z of the current frame of gastroscope image is obtained through the shared image representation layer 1 and the intermediate representation vector z of the image difference 2 ;

[0095] The input of the classification layer is the intermediate representation vector z of the current frame of gastroscope image 1 and the intermediate representation vector z of the image difference 2 The representation vector z after splicing t , and the cause category of the current frame of visual field loss is obtained through the classification layer

[0096] Optionally, the classification layer includes a linear classification layer and a softmax layer. The intermediate representation vector z of the current frame of gastroscope image 1 and the intermediate representation vector z of the image difference 2 The representation vector z after splicing t passes through the linear classification layer to obtain the predicted output vi of each cause category, and the predicted output vi of each cause category is input into the softmax layer to obtain the probability value of each cause category

[0097] Furthermore, the system further includes a quality monitoring result analysis module. The quality monitoring result analysis module displays the detection result of the current frame of gastroscope image in real time, and calculates the cumulative sum of the number of lost visual field frames and the cumulative lost frame time within every 1s based on the detection test results of all lost visual field frames at all previous moments up to the current time. The formulas are as follows:

[0098] K = K + k

[0099] t = k * m

[0100] where the interval between every two visual field frames is m milliseconds, k is the cumulative number of lost visual field frames within the current 1s, K is the cumulative number of lost visual field frames at all previous moments up to the current time, and t is the cumulative lost time of visual field frames within the current 1s. Whenever t = 1s, the corresponding chart is updated and displayed using the current total number of lost times K and the cumulative lost time t within the current 1s. By displaying the bar chart of the number of lost visual field frames, it can assist the doctor in observing the number of lost visual field frames up to the current moment; by displaying the line chart of the change in the lost duration of visual field frames, it can help the doctor accurately locate the lost time of visual field frames

[0101] The gastroscope visual field loss frame detection system in this embodiment is used to implement the aforementioned gastroscope visual field loss frame detection method. Therefore, the specific implementation of this system can be seen in the embodiment part of the gastroscope visual field loss frame detection method in the previous text. Therefore, its specific implementation can refer to the descriptions of the corresponding various part embodiments and will not be elaborated here

[0102] In addition, since the gastroscope vision loss frame detection system of this embodiment is used to implement the aforementioned gastroscope vision loss frame detection method, its function corresponds to that of the above method, and will not be elaborated here.

[0103] As shown in Table 1, it is the comparison results of the method of the present invention and the VisionTransformer (ViT) image classification algorithm, which is one of the advanced algorithms in the target image classification algorithm, in terms of speed, accuracy, video memory occupancy, and number of parameters.

[0104] ViT algorithm Method of the present invention Video memory occupancy 1.5G 1G Number of parameters 86M 11M FPS 112 181 Accuracy 90.69% 93.60%

[0105] Table 1

[0106] It can be seen from Table 1 that in terms of video memory occupancy, the video memory occupied by the ViT algorithm is 1.5G, while the video memory occupancy of the method of the present invention is only 1G, a reduction of 33.3%.

[0107] In terms of the number of parameters, the number of parameters of the ViT algorithm is 85M, while the number of parameters of the method of the present invention is only 11M, accounting for about 12.9% of the number of parameters of the ViT algorithm.

[0108] In terms of accuracy, the recognition accuracy of the trained model is about 3% higher than that of the ViT algorithm.

[0109] In terms of speed, the ViT algorithm can process 112 gastroscope pictures per second, while the trained model can process 181 gastroscope pictures, and the recognition speed has increased by 61.6%.

[0110] It can be seen that on the one hand, the present invention has a low hardware resource occupancy and few training parameters, and on the other hand, it has a high recognition accuracy and recognition speed.

[0111] As shown in Table 2, it is the comparison results of the traditional single-picture recognition model and the picture + (two consecutive frames) difference recognition model of the present application in terms of the F1 index.

[0112] F1-Score Single picture 0.88 Picture + difference 0.92

[0113] Table 2

[0114] By calculating the F1 metric for the prediction results, F1 is the equal-weighted harmonic mean of precision and recall (F1-Score), which is used to evaluate the recognition accuracy of the model. In the present invention, the detection method based on the loss of the gastroscope field of view frames has significantly improved accuracy compared to the traditional single-image recognition model (the F1-Score has increased by 4%). It can be seen that the detection method based on the loss of the gastroscope field of view frames in the present invention effectively solves the problem that the single-image input model cannot well predict the classification of the reasons for the loss of the field of view frames in the video. By combining the difference information of two consecutive frames to capture the relationship between the field of view of adjacent frame images, better performance in recognizing the reasons for the loss of the field of view is obtained.

[0115] The present invention can quickly and accurately detect the pictures with the loss of the field of view and find out the reasons for the loss of the field of view to assist doctors in operating the gastroscope.

[0116] The above embodiments are only preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention. The protection scope of the present invention shall be subject to the claims.

Claims

1. A method for detecting lost - field frames of a gastroscope, characterized in that, it includes the following steps: S1. Pre - process the current - frame gastroscope image; S2. Divide the pre - processed current - frame gastroscope image into multiple sub - images; S3. Use the trained first classification model to classify the labels of the multiple sub - images one by one, and obtain the classification result through the sub - image dynamic decision rule. If the classification result is a lost field, go to step S4; If the classification result is not a lost field, return to step S1 and continue to detect the next - frame gastroscope image; Among them, obtaining the classification result through the sub - image dynamic decision rule includes: if one sub - image decision is successful, end the program in advance to obtain the classification result; otherwise, continue to test the next sub - image. If all sub - images fail in decision, fuse the prediction probabilities of all sub - images to obtain the final classification result; S4. Use the trained second classification model to classify the current - frame gastroscope image with a classification result of a lost field and its difference features from adjacent - frame gastroscope images to obtain the cause category of the lost field of the current - frame gastroscope image.

2. The method for detecting lost - field frames of a gastroscope according to claim 1, characterized in that, The second classification model includes a shared image representation layer and a classification layer. The input of the shared image representation layer is (x t , x t+1 -x t ); where x t is the representation vector of the current-frame gastroscope image, x t+1 is the representation vector of the next-frame gastroscope image, and the intermediate representation vector z 1 of the current-frame gastroscope image and the intermediate representation vector z 2 of the image difference are obtained through the shared image representation layer; The input of the classification layer is the intermediate representation vector z of the current-frame gastroscope image 1 and the intermediate representation vector z of the image difference 2 The representation vector z after splicing t , and the cause category of the current-frame field of view loss is obtained through the classification layer 3. The method for detecting lost - field frames of a gastroscope according to claim 2, characterized in that, The classification layer includes a linear classification layer and a softmax layer, and the intermediate representation vector z of the current-frame gastroscope image 1 and the intermediate representation vector z of the image difference 2 The representation vector z after splicing t The predicted output v of each cause category is obtained through the linear classification layer i , and the predicted output v of each cause category i is input into the softmax layer to obtain the probability value of each cause category.

4. The method for detecting lost - field frames of a gastroscope according to claim 3, characterized in that, The softmax function of the softmax layer is as follows: where P i is the probability value of the i-th cause category, N is the total number of cause categories, v i is the predicted output of the i-th cause category; v n is the predicted output of the n-th cause category; e is the base of the natural logarithm function.

5. The method for detecting lost - field frames of a gastroscope according to claim 1, characterized in that, The sub - image dynamic decision rule adopts the following decision formula: p(i) = max(x ic ), m < n where c = 0, 1, 2......, C - 1; When fusing the prediction probability results in the sub - image dynamic decision rule, the following probability fusion formula is adopted: where c = 0, 1, 2......, C - 1; m is the number of input images; n is the number of sub-images obtained by segmentation; c is the subscript of the label; i is the subscript of the input image; C is the number of image categories; λ is the decision threshold obtained through training; x ic represents the predicted probability of the c-th image category of the i-th input image; p(c) represents the fusion probability of the c-th image category; label is the image category information predicted by the model.

6. The method for detecting lost - field frames of a gastroscope according to claim 1, characterized in that, The pre - processing includes one or more of the following processes: scaling and cropping processing, random horizontal flipping processing, standardization processing, and image cutting processing.

7. A system for detecting lost - field frames of a gastroscope, characterized in that, it includes the following modules: A pre - processing module for pre - processing the current - frame gastroscope image; An image - splitting module for dividing the pre - processed current - frame gastroscope image into multiple sub - images; A first classification module for using the trained first classification model to classify the labels of the multiple sub - images one by one and obtaining the classification result through the sub - image dynamic decision rule; A judgment module for judging whether the classification result is a lost field. If so, enter the second classification module; Otherwise, return to the pre - processing module and continue to detect the next - frame gastroscope image; Among them, obtaining the classification result through the sub - image dynamic decision rule includes: if one sub - image decision is successful, end the program in advance to obtain the classification result; otherwise, continue to test the next sub - image. If all sub - images fail in decision, fuse the prediction probabilities of all sub - images to obtain the final classification result; A second classification module, configured to classify the current frame gastroscope image with a classification result of visual field loss and its difference features from adjacent frame gastroscope images by using a trained second classification model, so as to obtain the cause category of the visual field loss of the current frame gastroscope image.

8. The gastroscope visual field loss frame detection system according to claim 7, wherein, The second classification model includes a shared image representation layer and a classification layer. The input of the shared image representation layer is (x t , x t+1 - x t ); where x t is the representation vector of the current frame gastroscope image, and x t+1 is the representation vector of the next frame gastroscope image. The intermediate representation vector z 1 of the current frame gastroscope image and the intermediate representation vector z 2 of the image difference are obtained through the shared image representation layer; The input of the classification layer is the intermediate representation vector z of the current-frame gastroscope image 1 and the intermediate representation vector z of the image difference 2 The representation vector z after concatenation t , and the category of the reason for the loss of the current-frame field of view is obtained through the classification layer.

9. The gastroscope visual field loss frame detection system according to claim 8, wherein, The classification layer includes a linear classification layer and a softmax layer, and the intermediate representation vector z of the current-frame gastroscope image 1 and the intermediate representation vector z of the image difference 2 The representation vector z after concatenation t Pass through the linear classification layer to obtain the prediction output v for each cause category i , the prediction output v for each cause category i Input to the softmax layer to obtain the probability value for each cause category.

Citation Information

Patent Citations

  • Method and system for achieving classification of pedestrians and vehicles based on neural network

    CN104504395A

  • Digestive endoscopy image abnormal feature real-time labeling system and method

    CN108852268A