Endoscopic target structure evaluation system, method, device, and storage medium
The endoscopic target structure evaluation system enhances gastrointestinal endoscopy by using a freeze screen detection module and a multi-scale fusion Transformer-based neural network to standardize and improve the accuracy of target size measurement, reducing errors and establishing a unified standard.
Patent Information
- Application Number
- JP2025545974
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-07
- Filing Date
- 2024-01-31
- Publication Date
- 2026-02-05
AI Technical Summary
Existing methods for estimating the size of gastrointestinal endoscopy targets, such as visual measurement and biopsy forceps comparison, are subjective and prone to errors due to physician variability and endoscope positioning, lacking a unified standard.
An endoscopic target structure evaluation system utilizing a freeze screen detection module and a gastrointestinal endoscopic target measurement module, combined with a multi-scale fusion Transformer-based neural network, to accurately measure target size by freezing images before and after scope retraction, correcting for distortion and establishing a standardized measurement.
Improves the accuracy of lesion size estimation by reducing subjective variability and establishing a unified standard, effectively addressing missed diagnoses and inaccurate measurements with conventional endoscopy.
Smart Images

Figure 2026504533000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to a Chinese patent application filed with the China Patent Office on February 7, 2023, bearing application number 202310069871.8 and entitled "Endoscopic target structure evaluation system, method, apparatus and storage medium," the entire contents of which are incorporated herein by reference. The present invention relates to the field of artificial intelligence, and in particular to a system, method, device and storage medium for evaluating a target structure under an endoscope. [Background technology]
[0002] Gastrointestinal endoscopy is widely used to observe the status of targets in the gastrointestinal tract, especially as a primary means of diagnosing and treating gastrointestinal pathologies. However, it has a high diagnostic miss rate for some gastrointestinal endoscopy targets (e.g., colon polyps, adenomas, etc.). With the assistance of artificial intelligence, this rate can be significantly reduced. After detecting a gastrointestinal endoscopy target, doctors need to observe the target's morphology and measure its size. The size of the gastrointestinal endoscopy target is one of the main criteria for determining the risk level of the gastrointestinal endoscopy target and selecting a treatment method. Summary of the Invention [Problem to be solved by the invention]
[0003] The main methods for estimating the size of gastrointestinal endoscopy targets are visual measurement and biopsy forceps comparison measurement. Visual measurement relies on physicians' own experience to estimate the size of gastrointestinal endoscopy targets, which is highly subjective. Due to differences in physician operating habits and experience, different physicians' estimates of the size of the same gastrointestinal endoscopy target vary widely. Biopsy forceps comparison measurement aims to assist in the measurement of the size of gastrointestinal endoscopy targets by inserting an endoscopic instrument, such as a biopsy forceps, during the endoscopic examination process and comparing it with the size of the gastrointestinal endoscopy target. This method requires the insertion of an endoscopic instrument during the examination process, but is not applicable to diagnostic (non-therapeutic) endoscopy, which does not require the insertion of an instrument. Furthermore, the position and angle of the instrument can cause significant errors in the measurement of the size of gastrointestinal endoscopy targets. To summarize the above two points, the biopsy forceps comparison measurement method has certain limitations.
[0004] Therefore, in the prior art, the operating habits and experience of the physician, and the shape and angle of the endoscope lens are all major factors that affect the measurement of the size of the lesion. How to improve the accuracy of the physician's estimation of the size of the lesion and establish a unified standard is urgently needed in the prior art. [Means for solving the problem]
[0005] SUMMARY OF THE INVENTION Embodiments of the present invention provide a system, method, apparatus, and storage medium for endoscopic target structure assessment that at least partially solve the above problems.
[0006] In the present invention, the target under gastrointestinal endoscopy is a target in a normal physiological state, such as, but not limited to, a tumor, a foreign body, a blood vessel, or a fecal matter, or may be a target in a pathological state, such as, but not limited to, a colon polyp or an adenoma.
[0007] In a first aspect, the present invention provides an endoscopic target structure evaluation system, the endoscopic target structure evaluation system including a freeze screen detection module and a gastrointestinal endoscopic target measurement module; the freeze screen detection module is used to freeze images of a target structure under gastrointestinal endoscopy from the endoscopic examination image before the scope return operation and the endoscopic examination image after the scope return operation, respectively, to obtain a first freeze image and a second freeze image, respectively; The gastrointestinal endoscopy target measurement module is used to determine the length of the gastrointestinal endoscopy target structure based on the first long axis distance of the first freeze image, the second long axis distance of the second freeze image, the scope return distance, the length occupied by a single pixel, and the number of pixels in the first long axis distance.
[0008] Preferably, the gastrointestinal endoscopic target measurement module determines a length B of the gastrointestinal endoscopic target structure using the following equation (1):
number
[0009] Preferably, the gastrointestinal endoscopy target identification module is further configured to call a gastrointestinal endoscopy target detection neural network based on a multi-scale fusion Transformer, which has been previously trained, to identify the first gastrointestinal endoscopy target freeze region in the first freeze-image and the second gastrointestinal endoscopy target freeze region in the second freeze-image, perform keypoint detection on the first gastrointestinal endoscopy target freeze region and the second gastrointestinal endoscopy target freeze region to obtain the first gastrointestinal endoscopy target freeze region and the second gastrointestinal endoscopy target freeze region that have been dewarped and aligned, detect the first long axis distance based on the dewarped and aligned first gastrointestinal endoscopy target freeze region, and detect the second long axis distance based on the dewarped and aligned second gastrointestinal endoscopy target freeze region.
[0010] Preferably, the scope return distance is determined by the length occupied by a single pixel.
[0011] Preferably, the endoscopic target structure evaluation system further comprises a client module; the client module is used to present a scope return operation when a target structure under gastrointestinal endoscopy is identified from the endoscopy image by the target identification module under gastrointestinal endoscopy; The gastrointestinal endoscopy target measurement module is further used to generate an auxiliary positioning frame by an image processing method, trigger the client module to draw and display the auxiliary positioning frame, and present an operation to move the gastrointestinal endoscopy target structure into the auxiliary positioning frame.
[0012] Preferably, when the gastrointestinal endoscopy target identification module identifies a gastrointestinal endoscopy target structure from the endoscopy video, it specifically calls a gastrointestinal endoscopy target detection neural network based on a multi-scale fusion Transformer, which has been pre-trained, to identify the gastrointestinal endoscopy target structure from the endoscopy video.
[0013] Preferably, the neural network for target detection under gastrointestinal endoscopy based on the multi-scale fusion Transformer includes a main feature extraction part, a feature fusion part, and a prediction head part; The main feature extraction part includes a normalization layer, a first feature extraction module in which three sets of standard residual convolution blocks are connected in series, a second feature extraction module in which six sets of standard residual convolution blocks are connected in series, a third feature extraction module in which nine sets of standard residual convolution blocks are connected in series, and a fourth feature extraction module in which a spatial pyramid pooling module and a Transformer module are connected in series; the normalization layer is used to scale the endoscopy video based on a preset scaling size; the first feature extraction module is used to compute the scaled endoscopy video to output a first feature map; the second feature extraction module is used to compute the first feature map to output a second feature map; the third feature extraction module is used to compute the second feature map to output a third feature map; and the fourth feature extraction module is used to compute the third feature map to obtain a fourth feature map; the feature fusion part is mainly used for performing attention weighting on the first, second, third and fourth feature maps to obtain first, second, third and fourth attention-weighted feature maps, respectively; calculating the first attention-weighted feature map using a first Transformer module to obtain a first fused feature map; calculating the second attention-weighted feature map and stitching it with the first attention-weighted feature map to obtain a second fused feature map using a second Transformer module; calculating the third attention-weighted feature map and stitching it with the second attention-weighted feature map to obtain a third fused feature map using a third Transformer module; and calculating the fourth attention-weighted feature map and stitching it with the third attention-weighted feature map to obtain a fourth fused feature map using a fourth Transformer module; The prediction head part is mainly used to perform grid search prediction using pre-set anchor boxes on the first, second, third and fourth fusion feature maps to identify the target structure under the gastrointestinal endoscopy.
[0014] In a second aspect, the present invention provides a method for endoscopic evaluation of a target structure, the method comprising: When a target structure under gastrointestinal endoscopy is identified from the endoscopy examination image, a suggestion to return the scope is made; Freezing images of the target structure under gastrointestinal endoscopy from the endoscopic examination image before the scope return operation and the endoscopic examination image after the scope return operation, respectively, to obtain a first frozen image and a second frozen image; and determining the length of the target structure under the gastrointestinal endoscope based on the first long axis distance of the first freeze image, the second long axis distance of the second freeze image, the scope return distance, the length occupied by a single pixel, and the number of pixels in the first long axis distance.
[0015] In a third aspect, the present invention provides an electronic device, the electronic device comprising: a memory; a controller; and a computer program stored in the memory and executable by the controller; When the computer program is executed by the controller, the steps of the method for evaluating a target structure under an endoscope described above are performed.
[0016] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein an evaluation program for an endoscopic target structure is stored in the computer-readable storage medium, and when the evaluation program for an endoscopic target structure is executed by a controller, the steps of the above-mentioned method for evaluating an endoscopic target structure are performed. [Effects of the Invention]
[0017] In each embodiment of the present invention, the images of the target structure under gastrointestinal endoscopy are frozen from the endoscopic examination image before and after the scope is retracted, respectively, to obtain a first frozen image and a second frozen image. The length of the target structure under gastrointestinal endoscopy is determined based on the first long axis distance of the first frozen image, the second long axis distance of the second frozen image, the scope retraction distance, the length occupied by a single pixel, and the number of pixels in the first long axis distance. This effectively solves the problems existing in the visual measurement method and the biopsy forceps comparative measurement method, improves the accuracy of doctors' estimation of the size of the lesion, and helps to establish a unified standard. It also effectively solves the problem of missed diagnoses when lesions are examined with the naked eye using a conventional endoscope, and avoids the problem of inaccurate measurement of the size of the lesion using an endoscope. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a block diagram of an endoscopic target structure evaluation system according to an embodiment of the present invention; [Figure 2] 1 is a flowchart of a method for evaluating a target structure under endoscopy in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0019] The present invention will be described in more detail below with reference to the drawings and specific examples. It should be understood that the specific examples described herein are merely for the purpose of interpreting the present invention and are not intended to limit the present invention.
[0020] In Example 1, an embodiment of the present invention provides an evaluation system for an endoscopic target structure. As shown in FIG. 1 , the evaluation system for an endoscopic target structure includes: a client module, a freeze screen detection module, and a gastrointestinal endoscopic target measurement module; the client module is used to present a scope return operation when a target structure under gastrointestinal endoscopy is identified from the endoscopy image by the target identification module under gastrointestinal endoscopy; the freeze screen detection module is used to freeze images of a target structure under gastrointestinal endoscopy from the endoscopic examination image before the scope return operation and the endoscopic examination image after the scope return operation, respectively, to obtain a first freeze image and a second freeze image, respectively; The gastrointestinal endoscopy target measurement module is used to determine the length of the gastrointestinal endoscopy target structure based on the first long axis distance of the first freeze image, the second long axis distance of the second freeze image, the scope return distance, the length occupied by a single pixel, and the number of pixels in the first long axis distance.
[0021] In one embodiment of the present invention, a system for evaluating a target structure under endoscopic surveillance includes a client module, a freeze-frame detection module, and a gastrointestinal endoscopy target measurement module. When a target structure under endoscopic surveillance is identified from an endoscopy video by the gastrointestinal endoscopy target identification module, the client module suggests a scope retraction operation. The freeze-frame detection module freezes images of the target structure under endoscopic surveillance from the endoscopy image before the scope retraction operation and the endoscopy image after the scope retraction operation to obtain a first freeze-frame image and a second freeze-frame image, respectively. The gastrointestinal endoscopy target measurement module determines the length of the target structure under endoscopic surveillance based on a first long axis distance of the first freeze-frame image, a second long axis distance of the second freeze-frame image, the scope retraction distance, the length occupied by a single pixel, and the number of pixels in the first long axis distance. This effectively solves the problems that exist in visual measurement and biopsy forceps comparative measurement, improves the accuracy of doctors' estimation of the size of the lesion site, and helps establish a unified standard. It also effectively solves the problem of missed diagnoses when examining lesions with the naked eye using a conventional endoscope, and avoids the problem of inaccurate measurement of the size of lesions using an endoscope.
[0022] In some embodiments, the endoscopic target structure evaluation system may further include a video acquisition module, and in detail, the endoscopic target structure evaluation system provided in this specific embodiment includes the following modules:
[0023] 1. S1 - Video Acquisition Module The module is used to read the video signal output from the endoscope device and obtain the video frame image. Specifically, it uses a video acquisition card to transmit digital / analog video signals such as HDMI (registered trademark), DVI, SDI, and S-Video into the computer, and then reads the video signal using OpenCV and transcodes it into an RGB image format frame by frame to obtain the endoscopy video.
[0024] 2. S2 - Target Identification Module for Gastrointestinal Endoscopy The module is mainly used to identify the target structure under gastrointestinal endoscopy from the endoscopy video, and to call a pre-trained gastrointestinal endoscopy target detection neural network based on a multi-scale fusion Transformer to identify the target freeze region under the first gastrointestinal endoscopy in the first freeze image and the target freeze region under the second gastrointestinal endoscopy in the second freeze image, perform keypoint detection on the target freeze region under the first gastrointestinal endoscopy and the target freeze region under the second gastrointestinal endoscopy to obtain the target freeze region under the first gastrointestinal endoscopy and the target freeze region under the second gastrointestinal endoscopy that have been dewarped and aligned, detect the first long axis distance based on the target freeze region under the first gastrointestinal endoscopy that has been dewarped and aligned, and detect the second long axis distance based on the target freeze region under the second gastrointestinal endoscopy that has been dewarped and aligned. (1) A neural network for detecting targets under gastrointestinal endoscopy based on artificial intelligence is constructed. In the embodiment of the present invention, a neural network for detecting targets under gastrointestinal endoscopy based on a multi-scale fusion Transformer is used.
[0025] i. A neural network architecture for target detection under gastrointestinal endoscopy is constructed, which includes a main feature extraction part, a feature fusion part and a prediction head part.
[0026] ii. First, the main feature extraction part uses a configuration including a cross-phase convolutional neural network, and includes a normalization layer, a first feature extraction module in which three sets of standard residual convolution blocks are connected in series, a second feature extraction module in which six sets of standard residual convolution blocks are connected in series, a third feature extraction module in which nine sets of standard residual convolution blocks are connected in series, and a fourth feature extraction module in which a spatial pyramid pooling module and a Transformer module are connected in series; Specifically, the structure-functions include: 1. A normalization layer is used to scale the endoscopy video based on a preset scaling size, for example, scaling the image input size to 640*640. 2. The first feature extraction module has three sets of standard residual convolution blocks connected in series, and outputs the first feature map through calculation. 3. The second feature extraction module has six sets of standard residual convolution blocks connected in series, and outputs the second feature map through calculation. 4. The third feature extraction module has nine sets of standard residual convolution blocks connected in series, and outputs the third feature map through calculation. 5. The fourth feature extraction module has one spatial pyramid pooling module and one transformer module connected in series inside it, which is used to highlight the dominant advanced features in the third feature map and obtain the fourth feature map through calculation.
[0027] iii. Then, the feature fusion part has a main functional structure including: 1. Attention weighting is performed on the first, second, third, and fourth feature maps. Specifically, three sets of residual convolution operations are first performed, and then spatial attention weighting is performed (using well-known spatial attention modules such as CBAM, ECA, RFB, and BAM). In this embodiment, the ECA module is selected and used to complete the spatial attention weighting and obtain the first, second, third, and fourth attention-weighted feature maps. 2. Use the first Transformer module to calculate the first attention weight feature map to obtain the first fused feature map. 3. Use the second Transformer module to calculate the second attention weight feature map, and stitch it with the first attention weight feature map to obtain the second fused feature map. 4. Use the third Transformer module to calculate the third attention weight feature map, and stitch it with the second attention weight feature map to obtain the third fused feature map. 5. Use the fourth Transformer module to calculate the fourth attention weight feature map, and stitch it with the third attention weight feature map to obtain the fourth fused feature map.
[0028] iv. Finally, the prediction head part, its main functional structure is as follows: 1. For the first, second, third, and fourth fusion feature maps, grid search prediction is performed using a pre-defined anchor box to identify the target structure under gastrointestinal endoscopy, and the prediction result includes center point coordinates, length, width, category ID, and category confidence. 2. Non-maximum suppression technology is used to remove duplicate results from the output predicted structure, and finally the center point coordinates, length, width, category ID, and category confidence of the target under gastrointestinal endoscopy predicted by the network are obtained.
[0029] (2) A neural network for detecting gastrointestinal endoscopy targets is trained, and the network model has the ability to identify gastrointestinal endoscopy targets.
[0030] i. Construct a gastrointestinal endoscopy target mark dataset and train a network model based on the preset dataset, the types of which include raised gastrointestinal endoscopy targets, flat gastrointestinal endoscopy targets, and depressed gastrointestinal endoscopy targets.
[0031] ii. The dataset image is input into the network model for each lot, and the error from the true data value is output using the cross-entropy loss calculation network model, and the parameters are updated based on back propagation.
[0032] iii. Repeat the above steps until the accuracy of the model reaches a preset value or the loss value reaches 1e-5, then stop training and obtain a pre-built artificial intelligence-based neural network for target detection under gastrointestinal endoscopy. c) The trained target detection neural network model for gastrointestinal endoscopy is used to identify each frame of the image, and finally, through a prediction layer, the identification result is obtained. The identification result is pushed to the client in coordinate format using websocket technology, and a rectangular identification frame is drawn using canvas drawing technology.
[0033] 3. S3 - Frozen Screen Detection Module The module is used to freeze images of a target structure under gastrointestinal endoscopy from the endoscopic examination image before the scope return operation and the endoscopic examination image after the scope return operation, respectively, to obtain a first frozen image and a second frozen image. a) When the gastrointestinal endoscopy target identification result appears in the video frame image, the user is prompted to turn on the gastrointestinal endoscopy target measurement function, and an auxiliary positioning frame is displayed in the center of the frame image screen using the client module, guiding the operation so that the user can move the gastrointestinal endoscopy target into the auxiliary positioning frame, and then the image is frozen. b) Using a matrix comparison method, identify the obtained frame image sequence to obtain a first freeze image, and based on the target identification result under the gastrointestinal endoscopy, obtain the detection coordinates of the target under the first gastrointestinal endoscopy, i.e., the target freeze area under the first gastrointestinal endoscopy. c) After the first frozen image is identified, the client module is used to prompt the user with the number of centimeters to return the scope for the first time (0.5 cm in this embodiment of the present invention), freeze the image again, and use a matrix comparison method to identify the obtained frame image sequence to obtain a second frozen image. Based on the target identification result under gastrointestinal endoscopy, the detection coordinates of the target under the second gastrointestinal endoscopy, i.e., the target freeze area under the second gastrointestinal endoscopy, are obtained.
[0034] 4. S4 - Target measurement module under gastrointestinal endoscopy The module is used to determine the length of the target structure under the gastrointestinal endoscopy based on the first long axis distance of the first freeze image, the second long axis distance of the second freeze image, the scope return distance, the length occupied by a single pixel, and the number of pixels in the first long axis distance. a) An auxiliary positioning frame is automatically generated using image processing technology, and the operation is guided so that the user can move the target under gastrointestinal endoscopy into the auxiliary positioning frame and freeze the image. b) When the freeze screen detection module detects a freeze and the gastrointestinal endoscopy target identification module simultaneously obtains an identification result, the first freeze image is automatically saved, and the detection coordinates of the first gastrointestinal endoscopy target are obtained based on the prediction result of the gastrointestinal endoscopy target identification module. c) Once the first frozen image is identified, the client module is used to present the user with the number of centimeters of the first scope return (in this embodiment, the distance of the scope return is determined by the length occupied by a single pixel. For example, if the measurement is performed 1 cm away from the caliper, δ is the pixel ratio at 1 cm, so the scope return L must also be 1 cm. If the measurement is performed 2 cm away from the caliper, the scope return range may be 0.4 to 1 cm, and 1 cm is preferred). Then, the image is frozen again, and a second frozen image is obtained using the frozen screen detection module. Based on the prediction result of the gastrointestinal endoscopy target identification module, the detection coordinates of the second gastrointestinal endoscopy target are obtained. d) Remove distortion of the target under gastrointestinal endoscopy due to the operation of "return the scope by how many centimeters." Specifically, using OBR keypoint detection technology, keypoints are extracted and matched for the detection coordinates of the target under the first gastrointestinal endoscopy and the detection coordinate area of the target under the second gastrointestinal endoscopy, RANSAC roughness removal is performed for the matched keypoints, and perspective transformation is further performed using the matched keypoints to obtain the aligned target coordinates of the first gastrointestinal endoscopy and the second gastrointestinal endoscopy with the distortion removed. e) Using the convex hull algorithm, the long axis distances of the target coordinates under the first and second gastrointestinal endoscopes are detected to obtain the first long axis distance B1 and the second long axis distance B2. f) Using the conversion formula (1), calculate the target length B under gastrointestinal endoscopy. The conversion formula (1) is as follows:
number
[0035] 5.S5 - Client module; The above modules are used as follows: a) When the gastrointestinal endoscopy target identification module identifies a gastrointestinal endoscopy target structure from the endoscopy image, a scope return operation is suggested. b) Based on the identification results of the target identification module under gastrointestinal endoscopy, a detection frame is drawn on the video frame image. c) Drawing an auxiliary positioning frame for a target under gastrointestinal endoscopy based on the target measurement module under gastrointestinal endoscopy. d) Displaying the size of the target under gastrointestinal endoscopy based on the detection result of the gastrointestinal endoscopy target measurement module.
[0036] This embodiment proposes a method for measuring gastrointestinal endoscopy target structures, particularly applicable to polyp detection. It utilizes a combination of techniques, including endoscopic lens pixel point mapping length measurement, freeze detection, AI-based gastrointestinal endoscopy target detection, and length transformation, to address the problem of accurately measuring clinical lesions. Furthermore, a gastrointestinal endoscopy target detection network based on a multi-scale fusion Transformer is proposed. This network incorporates four layers of scale fusion, simultaneously utilizing the functions of both Transformer and spatial attention, significantly improving gastrointestinal endoscopy target extraction capabilities and reducing false positives. This effectively solves the problems inherent in visual measurement and biopsy forceps comparison measurement, improving the accuracy of physicians' gastrointestinal endoscopy target size estimation and establishing a unified standard. It also effectively solves the problem of missed diagnoses when gastrointestinal endoscopy targets are inspected with the naked eye using conventional endoscopes, and avoids the inability to accurately measure gastrointestinal endoscopy target size.
[0037] In Example 2, an embodiment of the present invention provides an endoscopic target structure evaluation method, as shown in FIG. 2 , the endoscopic target structure evaluation method includes the following steps: In S101, when a target structure under gastrointestinal endoscopy is identified from the endoscopic examination image, a scope return operation is suggested, In S102, images of the target structure under the gastrointestinal endoscope are frozen from the endoscopic examination image before the scope return operation and the endoscopic examination image after the scope return operation, respectively, to obtain a first frozen image and a second frozen image; In S103, the length of the target structure under the gastrointestinal endoscope is determined based on the first long axis distance of the first freeze image, the second long axis distance of the second freeze image, the scope return distance, the length occupied by a single pixel, and the number of pixels in the first long axis distance.
[0038] In some embodiments, the length B of the target structure under gastrointestinal endoscopy is determined using the following equation (1):
number
[0039] In some embodiments, a pre-trained, multi-scale fusion Transformer-based target detection neural network for gastrointestinal endoscopy is invoked to identify a target freeze region for the first gastrointestinal endoscopy in the first freeze-image and a target freeze region for the second gastrointestinal endoscopy in the second freeze-image, keypoint detection is performed on the target freeze region for the first gastrointestinal endoscopy and the target freeze region for the second gastrointestinal endoscopy to obtain a dewarped and aligned target freeze region for the first gastrointestinal endoscopy and a target freeze region for the second gastrointestinal endoscopy, the first long axis distance is detected based on the dewarped and aligned target freeze region for the first gastrointestinal endoscopy, and the second long axis distance is detected based on the dewarped and aligned target freeze region for the second gastrointestinal endoscopy.
[0040] Preferably, the range of the scope return distance is 0.4 to 1 cm.
[0041] In some embodiments, the method for evaluating a target structure under endoscopy further includes generating an auxiliary positioning frame using an image processing method, triggering the client module to draw and display the auxiliary positioning frame, and presenting an operation to move the target structure under gastrointestinal endoscopy into the auxiliary positioning frame.
[0042] In some embodiments, identifying a gastrointestinal endoscopy target structure from the endoscopy video includes using a pre-trained gastrointestinal endoscopy target detection neural network based on a multi-scale fusion transformer to identify the gastrointestinal endoscopy target structure from the endoscopy video.
[0043] In Example 3, an embodiment of the present invention provides an electronic device, the electronic device including: a memory; a controller; and a computer program stored in the memory and executable by the controller; When the computer program is executed by the controller, the steps of the method for evaluating a target structure under endoscopy described in Example 2 are performed.
[0044] In Example 4, an embodiment of the present invention provides a computer-readable storage medium, in which an evaluation program for an endoscopic target structure is stored, and when the evaluation program for an endoscopic target structure is executed by a controller, the steps of the method for evaluating an endoscopic target structure described in Example 2 are performed.
[0045] In the concrete implementation process, Examples 2 to 4 can refer to Example 1, and have corresponding technical effects.
[0046] Although the embodiments of the present invention have been described above with reference to the drawings, the present invention is not limited to the above-mentioned specific embodiments, which are merely illustrative and not limiting. Those skilled in the art can implement many more forms under the guidance of the present invention without departing from the spirit of the present invention and the scope of protection of the claims, and all of these fall within the scope of protection of the present invention.
Claims
1. 1. An endoscopic target structure assessment system, comprising: The endoscopic target structure evaluation system includes a freeze screen detection module and a gastrointestinal endoscopic target measurement module; the freeze screen detection module is used to freeze images of a target structure under gastrointestinal endoscopy from the endoscopic examination image before the scope return operation and the endoscopic examination image after the scope return operation, respectively, to obtain a first freeze image and a second freeze image, respectively; The gastrointestinal endoscopy target measurement module is used to determine the length of the gastrointestinal endoscopy target structure based on the first long axis distance of the first freeze image, the second long axis distance of the second freeze image, the scope return distance, the length occupied by a single pixel, and the number of pixels in the first long axis distance.
2. The gastrointestinal endoscopic target measurement module determines a length B of the gastrointestinal endoscopic target structure using the following equation (1): [Equation 1] However, B 1 denotes the first long axis distance of the first freeze image, and B 2 denotes the second long axis distance of the second frozen image, L denotes the distance of the scope return, θ denotes the length occupied by a single pixel, and Pixs(B 1 2. The system for evaluating a target structure under endoscopy according to claim 1, wherein the first major axis distance is a pixel number.
3. 3. The system for evaluating a target structure under an endoscope according to claim 2, wherein the scope return distance is determined by the length occupied by a single pixel.
4. 2. The system for evaluating a target structure under endoscopy according to claim 1, wherein the gastrointestinal endoscopy target identification module is further configured to: invoke a gastrointestinal endoscopy target detection neural network based on a multi-scale fusion transformer, which has been previously trained, to identify a first gastrointestinal endoscopy target freeze region in the first freeze-image and a second gastrointestinal endoscopy target freeze region in the second freeze-image; perform keypoint detection on the first gastrointestinal endoscopy target freeze region and the second gastrointestinal endoscopy target freeze region to obtain a first gastrointestinal endoscopy target freeze region and a second gastrointestinal endoscopy target freeze region that have been dewarped and aligned; detect the first long axis distance based on the first gastrointestinal endoscopy target freeze region that have been dewarped and aligned; and detect the second long axis distance based on the second gastrointestinal endoscopy target freeze region that have been dewarped and aligned.
5. the endoscopic target structure evaluation system further includes a client module; the client module is used to present a scope return operation when a target structure under gastrointestinal endoscopy is identified from the endoscopy image by the target identification module under gastrointestinal endoscopy; The evaluation system for a target structure under an endoscope as described in claim 1, characterized in that the target measurement module under gastrointestinal endoscopy is further used to generate an auxiliary positioning frame using an image processing method, trigger the client module to draw and display the auxiliary positioning frame, and present an operation to move the target structure under gastrointestinal endoscopy into the auxiliary positioning frame.
6. The gastrointestinal endoscopy target identification module is specifically configured to call a gastrointestinal endoscopy target detection neural network based on a multi-scale fusion transformer, which has been trained in advance, when the gastrointestinal endoscopy target identification module identifies a gastrointestinal endoscopy target structure from the endoscopy video, to identify the gastrointestinal endoscopy target structure from the endoscopy video.
7. The neural network for target detection under gastrointestinal endoscopy based on the multi-scale fusion transformer includes a main feature extraction part, a feature fusion part, and a prediction head part; the main feature extraction part includes a normalization layer, a first feature extraction module in which three sets of standard residual convolution blocks are connected in series, a second feature extraction module in which six sets of standard residual convolution blocks are connected in series, a third feature extraction module in which nine sets of standard residual convolution blocks are connected in series, and a fourth feature extraction module in which a spatial pyramid pooling module and a transformer module are connected in series; the normalization layer is used to scale the endoscopy video based on a preset scaling size; the first feature extraction module is used to compute the scaled endoscopy video to output a first feature map; the second feature extraction module is used to compute the first feature map to output a second feature map; the third feature extraction module is used to compute the second feature map to output a third feature map; and the fourth feature extraction module is used to compute the third feature map to obtain a fourth feature map; the feature fusion part is mainly used for: performing attention weighting on the first, second, third, and fourth feature maps to obtain first, second, third, and fourth attention-weighted feature maps, respectively; calculating the first attention-weighted feature map using a first Transformer module to obtain a first fused feature map; calculating the second attention-weighted feature map and stitching it with the first attention-weighted feature map using a second Transformer module to obtain a second fused feature map; calculating the third attention-weighted feature map and stitching it with the second attention-weighted feature map using a third Transformer module to obtain a third fused feature map; and calculating the fourth attention-weighted feature map and stitching it with the third attention-weighted feature map using a fourth Transformer module to obtain a fourth fused feature map; The system for evaluating a target structure under endoscopy according to claim 6, characterized in that the prediction head part is mainly used to identify the target structure under endoscopy by performing grid search prediction using preset anchor boxes on the first, second, third and fourth fusion feature maps.
8. 1. A method for endoscopic evaluation of a target structure, comprising: The method for evaluating a target structure under an endoscope comprises: When a target structure under gastrointestinal endoscopy is identified from the endoscopy examination image, a suggestion to return the scope is made; acquiring a first frozen image and a second frozen image by freezing images of a target structure under gastrointestinal endoscopy from the endoscopic examination image before the scope return operation and the endoscopic examination image after the scope return operation, respectively; and determining the length of the target structure under gastrointestinal endoscopy based on a first long axis distance of the first freeze image, a second long axis distance of the second freeze image, a scope return distance, a length occupied by a single pixel, and the number of pixels at the first long axis distance.
9. An electronic device, the electronic device includes a memory, a controller, and a computer program stored in the memory and executable by the controller; 9. An electronic device, wherein, when the computer program is executed by the controller, steps of the method for evaluating a target structure under an endoscope according to claim 8 are performed.
10. A computer-readable storage medium, comprising: A computer-readable storage medium, characterized in that the steps of the method for evaluating a target structure under an endoscope described in claim 8 are performed, wherein an evaluation program for the target structure under an endoscope is stored in the computer-readable storage medium and the evaluation program for the target structure under an endoscope is executed by a controller.
Citation Information
Patent Citations
Endoscopic apparatus
JP1989250225A
Measuring endoscope device
JP1991080824A
Measuring method of object through endoscope
JP1994018219A
Endoscopic diagnostic device, image processing method, program, and recording medium
JP2016189861A
Blood vessel diameter measuring device and blood vessel diameter measuring method
JP2020185082A