Method, apparatus, readable medium, and electronic device for processing endoscopic images

The three-dimensional tissue image during the endoscopic examination is constructed through the three-dimensional reconstruction model and projected to the tissue template, which solves the problem of missing detection caused by blind spots in the endoscopic field of vision, and achieves the effect of timely response to the scope of inspection and ensuring the effectiveness of the inspection.

CN114332028BActive Publication Date: 2025-06-17XIAOHE MEDICAL EQUIP (HAINAN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111652171.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-06-17
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

During the examination of internal tissue of the human body, the blind spots in the field of vision due to peristalsis, wrinkles and operation, which may lead to missed inspection and affect the effectiveness of the inspection.

Method used

By acquiring tissue images collected by the current and historical endoscopy, the depth image and pose parameters are determined using a pre-trained 3D reconstruction model, and then the three-dimensional tissue images are constructed and projected to the tissue template to determine the visible area and blind area, and the blind area ratio is calculated.

Benefits of technology

Quickly obtain accurate three-dimensional tissue images, promptly respond to the inspection range, avoid missed inspections, and ensure the effectiveness of the inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332028B_ABST
    Figure CN114332028B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, apparatus, readable medium, and electronic device for processing endoscopic images, and relates to the technical field of image processing. The method includes: acquiring a tissue image collected by an endoscope at the current moment, determining a depth image corresponding to the tissue image, and pose parameters between the tissue image and a historical tissue image through a pre-trained three-dimensional reconstruction model according to the tissue image and the historical tissue image, and determining a three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image, and the pose parameters, where the historical tissue image is an image collected by the endoscope before the current moment, projecting the three-dimensional tissue image onto a tissue template to determine a visible region where the projection in the three-dimensional tissue image overlaps with the tissue template, and a blind area region where the projection in the three-dimensional tissue image does not overlap with the tissue template, and determining a blind area ratio during the endoscopic examination according to the visible region and the blind area region. The present disclosure can timely reflect the examination range to avoid missed detection and ensure the effectiveness of the examination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a method, device, readable medium, and electronic device for processing endoscopic images. Background Art

[0002] Endoscopes are provided with components such as optical lenses, image sensors, light sources, etc., and can enter the tissues inside the human body for examination, enabling doctors to directly observe the internal conditions of the human body and being widely used in the medical field. When the endoscope enters the tissues inside the human body for examination, due to reasons such as tissue peristalsis (e.g., intestines, stomach, etc.), the presence of tissue folds, or the actions of the examiner such as flushing and unlooping, it is easy to cause blind spots in the field of view of the endoscope, which may further lead to missed detections and cannot guarantee the effectiveness of the examination. Summary of the Invention

[0003] This Summary of the Invention section is provided to introduce concepts in a brief form, which will be described in detail in the subsequent Detailed Description section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] In a first aspect, the present disclosure provides a method for processing endoscopic images, the method including:

[0005] Obtain a tissue image collected by the endoscope at the current moment;

[0006] According to the tissue image and historical tissue images, determine, through a pre-trained three-dimensional reconstruction model, a depth image corresponding to the tissue image, pose parameters between the tissue image and the historical tissue images, and determine a three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image, and the pose parameters, where the historical tissue images are images collected by the endoscope before the current moment;

[0007] Project the three-dimensional tissue image onto a tissue template to determine a visible area where the projection in the three-dimensional tissue image overlaps with the tissue template, and a blind area where the projection in the three-dimensional tissue image does not overlap with the tissue template, where the tissue template is used to represent the overall shape of the tissue examined by the endoscope;

[0008] Determine a blind area ratio during the endoscope examination according to the visible area and the blind area.

[0009] In a second aspect, the present disclosure provides a device for processing endoscopic images, the device including:

[0010] An obtaining module, configured to obtain a tissue image collected by the endoscope at the current moment;

[0011] A reconstruction module, configured to determine a depth image corresponding to the tissue image, pose parameters between the tissue image and the historical tissue image according to the tissue image and the historical tissue image through a pre-trained three-dimensional reconstruction model, and determine a three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image, and the pose parameters, where the historical tissue image is an image acquired by the endoscope before the current moment;

[0012] A projection module, configured to project the three-dimensional tissue image onto a tissue template to determine a visible area where the projection in the three-dimensional tissue image overlaps with the tissue template, and a blind area where the projection in the three-dimensional tissue image does not overlap with the tissue template, where the tissue template is used to represent the overall shape of the tissue examined by the endoscope;

[0013] A processing module, configured to determine a blind area ratio during the endoscope examination according to the visible area and the blind area.

[0014] In a third aspect, the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of the method described in the first aspect of the present disclosure are implemented.

[0015] In a fourth aspect, the present disclosure provides an electronic device, including:

[0016] A storage device, on which a computer program is stored;

[0017] A processing device, configured to execute the computer program in the storage device to implement the steps of the method described in the first aspect of the present disclosure.

[0018] Through the above technical solutions, the present disclosure first acquires a tissue image acquired by an endoscope at the current moment, and then determines a depth image corresponding to the tissue image, pose parameters between the tissue image and the historical tissue image according to the tissue image and the historical tissue image through a pre-trained three-dimensional reconstruction model, and determines a three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image, and the pose parameters. Wherein, the historical tissue image is an image acquired by the endoscope before the current moment. Finally, the three-dimensional tissue image is projected onto a tissue template to determine a visible area where the projection in the three-dimensional tissue image overlaps with the tissue template, and a blind area where the projection in the three-dimensional tissue image does not overlap with the tissue template, and the blind area ratio during the endoscope examination is determined according to the visible area and the blind area. The present disclosure determines a depth image and pose parameters according to a tissue image through a three-dimensional reconstruction model, and performs three-dimensional reconstruction based on this, so as to quickly obtain an accurate three-dimensional tissue image. And based on the three-dimensional tissue image, the blind area ratio is determined in combination with the tissue template, which can timely reflect the examination range during the examination process, thereby avoiding missed inspections and ensuring the effectiveness of the examination.

[0019] Other features and advantages of the present disclosure will be described in detail in the following detailed description section. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale. In the drawings:

[0021] Figure 1 is a flowchart of a method for processing an endoscopic image shown according to an exemplary embodiment;

[0022] Figure 2 is a schematic diagram of a three-dimensional reconstruction model shown according to an exemplary embodiment;

[0023] Figure 3 is a flowchart of another method for processing an endoscopic image shown according to an exemplary embodiment;

[0024] Figure 4 is a flowchart of another method for processing an endoscopic image shown according to an exemplary embodiment;

[0025] Figure 5 is a flowchart of another method for processing an endoscopic image shown according to an exemplary embodiment;

[0026] Figure 6 is a schematic diagram of the connection relationship between a three-dimensional reconstruction model and an optical flow model shown according to an exemplary embodiment;

[0027] Figure 7 is a schematic diagram of jointly training a three-dimensional reconstruction model and an optical flow model shown according to an exemplary embodiment;

[0028] Figure 8 is a schematic diagram of another jointly training a three-dimensional reconstruction model and an optical flow model shown according to an exemplary embodiment;

[0029] Figure 9 is a flowchart of another method for processing an endoscopic image shown according to an exemplary embodiment;

[0030] Figure 10 is a block diagram of a device for processing an endoscopic image shown according to an exemplary embodiment;

[0031] Figure 11 is a block diagram of another device for processing an endoscopic image shown according to an exemplary embodiment;

[0032] Figure 12It is a block diagram of another endoscopic image processing device shown according to an exemplary embodiment;

[0033] Figure 13 It is a block diagram of another endoscopic image processing device shown according to an exemplary embodiment;

[0034] Figure 14 It is a block diagram of another endoscopic image processing device shown according to an exemplary embodiment;

[0035] Figure 15 It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0036] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0037] It should be understood that the various steps described in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0038] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0039] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0040] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0041] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0042] The tissues inside the human body are usually soft tissues with cavities. During the process of a doctor moving the endoscope, the soft tissues (such as the intestines, stomach, etc.) will peristalsis. And during the endoscopy examination, the doctor will perform operations such as flushing water and releasing loops, which may result in blind spots during the endoscopy examination. In addition, since there are usually folds in the soft tissues, some areas in the soft tissues may not appear in the field of view of the endoscope (i.e., there are blind spots). If the proportion of blind spots during the endoscopy examination is too large, it is very likely to lead to missed detections, such as not detecting small parts like polyps. Therefore, during the endoscopy examination, it is necessary to timely detect the proportion of blind spots to avoid missed detections and ensure the effectiveness of the examination.

[0043] Figure 1 is a flowchart of a method for processing endoscopic images shown according to an exemplary embodiment, as Figure 1 shown, the method may include the following steps:

[0044] Step 101, obtain the tissue image collected by the endoscope at the current moment.

[0045] For example, during endoscopy, the endoscope continuously collects images in the tissue according to a preset acquisition period. The tissue image in this embodiment can be understood as the image collected by the endoscope at the current moment. Correspondingly, the historical tissue image mentioned later can be understood as the image collected by the endoscope before the current moment, such as the image collected at the previous moment before the current moment. It should be noted that the endoscope described in the embodiments of the present disclosure can be, for example, a colonoscope, a gastroscope, etc. If the endoscope is a colonoscope, then the above tissue image is an intestinal image. If the endoscope is a gastroscope, then the above tissue image can be an esophageal image, a gastric image, or a duodenal image. The endoscope can also be used to collect images of other tissues, and the present disclosure does not make specific limitations on this.

[0046] During the endoscopy process, due to reasons such as unstable insertion techniques or inappropriate positions of the endoscope, many invalid images may be collected, such as images blocked by obstacles, overexposed images, and images with low clarity. These invalid images will interfere with the endoscopy examination results. Therefore, after obtaining the tissue image, it can be first determined whether the tissue image is valid. If the tissue image is an invalid image, the tissue image can be directly discarded, and wait for the tissue image collected by the endoscope at the next moment. If the tissue image is a valid image, then subsequent processing steps can be carried out, which can reduce unnecessary data processing and improve the processing speed. For example, a pre-trained recognition model can be used to recognize the tissue image to determine whether the tissue image is valid. The recognition model can be, for example, CNN (English: Convolutional Neural Networks, Chinese: Convolutional Neural Network) or LSTM (English: Long Short-Term Memory, Chinese: Long Short-Term Memory Network), or it can also be the Encoder in Transformer (such as Vision Transformer). The present disclosure does not make specific limitations in this regard. Further, preprocessing can also be performed on the tissue image, which can be understood as enhancing the data included in the tissue image. To ensure the quality of the tissue image, the preprocessing will not modify the blurriness or color of the tissue image. Therefore, the preprocessing can include: multi-crop processing, flipping processing (including: left-right flipping, up-down flipping, rotation, etc.), random affine transformation, size transformation (English: Resize), etc. The finally obtained preprocessed tissue image can be an image of a specified size (for example, it can be 384*384).

[0047] Step 102, according to the tissue image and the historical tissue image, determine the depth image corresponding to the tissue image, the pose parameters between the tissue image and the historical tissue image through a pre-trained three-dimensional reconstruction model, and determine the three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image, and the pose parameters. The historical tissue image is the image collected by the endoscope before the current moment.

[0048] Exemplarily, the preprocessed tissue image and the historical tissue image can be input into a pre-trained three-dimensional reconstruction model, so that the three-dimensional reconstruction model determines a depth image corresponding to the tissue image and pose parameters between the tissue image and the historical tissue image according to the tissue image and the historical tissue image. Among them, the depth image corresponding to the tissue image includes the depth (which can also be understood as the distance) of each pixel point in the tissue image. Therefore, the depth image can reflect the geometric shape of the visible surface in the tissue image and is not affected by the texture, color, etc. in the tissue image. That is to say, the structural information of the tissue corresponding to the tissue image can be characterized through the depth image. The pose parameters can characterize the movement process of the endoscope in the tissue and can include, for example, a rotation matrix and a translation vector.

[0049] After that, the three-dimensional reconstruction model can perform three-dimensional reconstruction according to the tissue image, the depth image, and the pose parameters to obtain a three-dimensional tissue image corresponding to the tissue image. The three-dimensional tissue image can characterize the three-dimensional structure of the tissue corresponding to the tissue image. The position of each pixel point in the tissue image can be understood as the two-dimensional coordinates of the pixel point. Correspondingly, the depth image corresponding to the tissue image includes the depth of each pixel point. The three-dimensional reconstruction model can fuse the tissue image with the corresponding depth image to obtain a three-dimensional tissue image including the three-dimensional coordinates of each pixel point. When performing the fusion, the pose parameters can also be used to remove distortion. Specifically, the SurfelMeshing fusion algorithm can be used to implement the three-dimensional reconstruction, or other fusion algorithms can be used to implement the three-dimensional reconstruction. The present disclosure does not make specific limitations on this.

[0050] Step 103: Project the three-dimensional tissue image onto the tissue template to determine a visible area where the projection in the three-dimensional tissue image overlaps with the tissue template and a blind area where the projection in the three-dimensional tissue image does not overlap with the tissue template. The tissue template is used to characterize the overall shape of the tissue examined by the endoscope.

[0051] Step 104: Determine the blind area ratio during the endoscope examination according to the visible area and the blind area.

[0052] Exemplarily, after obtaining the three-dimensional tissue image, the three-dimensional tissue image can be projected onto the tissue template, and based on the visible area where the projection overlaps with the tissue template and the blind area where the projection does not overlap with the tissue template, the blind area ratio during the endoscopy process can be determined. The blind area ratio can be understood as the ratio of the blind area (i.e., the part that cannot be observed in the field of view of the endoscope) to the total internal surface area of the tissue during the endoscopy process. The tissue template can be understood as a template that can reflect the overall shape of the tissue currently being examined by the endoscope. Taking the endoscope as a colonoscope as an example, the tissue to be examined is the intestine, and then the tissue template can be a distorted cylinder. In one implementation, the three-dimensional tissue image can be directly projected onto the tissue template to obtain the visible area where the projection in the three-dimensional tissue image overlaps with the tissue template and the blind area where the projection in the three-dimensional tissue image does not overlap with the tissue template, and then the blind area ratio can be determined according to the ratio of the blind area to the visible area.

[0053] In another implementation, since the three-dimensional tissue image represents the three-dimensional structure of the tissue corresponding to the tissue image, the field of view of a tissue image is limited. Taking the tissue image as an intestinal image as an example, the three-dimensional tissue image can only reflect the three-dimensional structure of a small section of the intestine. Therefore, multiple three-dimensional tissue images corresponding to multiple tissue images collected within a period of time (e.g., 15 s) can be stitched together to obtain a three-dimensional total image that can reflect the three-dimensional structure of a longer section of the intestine, and then the three-dimensional total image can be projected onto the tissue template to obtain the visible area where the projection overlaps with the tissue template and the blind area where the projection does not overlap with the tissue template, and then the blind area ratio can be determined according to the ratio of the blind area to the visible area.

[0054] The three-dimensional reconstruction model can determine the depth image corresponding to the tissue image, without the need to add a depth sensor during the endoscopy examination, which is convenient for operation and also saves costs. At the same time, the three-dimensional reconstruction model can determine the pose parameters to remove distortion during the three-dimensional reconstruction process, so as to quickly obtain an accurate three-dimensional tissue image. Further, determining the blind area ratio based on the three-dimensional tissue image in combination with the tissue template can timely reflect the examination range during the examination process, thereby avoiding missed examinations and ensuring the effectiveness of the examination.

[0055] In summary, the present disclosure first obtains a tissue image collected by an endoscope at the current moment, and then determines a depth image corresponding to the tissue image, pose parameters between the tissue image and the historical tissue image, based on the tissue image and the historical tissue image, through a pre-trained three-dimensional reconstruction model, and determines a three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image, and the pose parameters. The historical tissue image is an image collected by the endoscope before the current moment. Finally, the three-dimensional tissue image is projected onto a tissue template to determine a visible area where the projection in the three-dimensional tissue image overlaps with the tissue template, and a blind area where the projection in the three-dimensional tissue image does not overlap with the tissue template, and determines the blind area ratio during the endoscope examination according to the visible area and the blind area. The present disclosure determines the depth image and pose parameters according to the tissue image through the three-dimensional reconstruction model, and performs three-dimensional reconstruction based on this, so as to quickly obtain an accurate three-dimensional tissue image. And based on the three-dimensional tissue image, the blind area ratio is determined in combination with the tissue template, which can timely reflect the examination range during the examination, thereby avoiding missed detection and ensuring the effectiveness of the examination.

[0056] In one implementation, the structure of the three-dimensional reconstruction model can be as Figure 2 shown, which includes: a depth sub-model, a pose sub-model, and a fusion sub-model. The input of the depth sub-model and the input of the pose sub-model serve as the input of the three-dimensional reconstruction model, the output of the depth sub-model and the output of the pose sub-model together serve as the input of the fusion sub-model, and the output of the fusion sub-model serves as the output of the three-dimensional reconstruction model.

[0057] Figure 3 is a flowchart of another method for processing endoscopic images shown according to an exemplary embodiment, as Figure 3 shown, step 102 may include:

[0058] Step 1021, input the tissue image into the depth sub-model to obtain the depth image output by the depth sub-model.

[0059] Exemplarily, the tissue image can be used as the input of the depth sub-model, and the depth sub-model can output the corresponding depth image. The structure of the depth sub-model is as Figure 2 shown, and it can be a UNet structure, which includes multiple stride convolution layers (stride conv) to downsample the tissue image, for example, it can be downsampled to 1 / 8 of the resolution of the tissue image, and then multiple transpose convolution layers (transpose conv) are used for upsampling to the resolution of the tissue image to obtain the corresponding depth image.

[0060] Step 1022, input the tissue image and the historical tissue image into the pose sub-model to obtain the pose parameters output by the pose sub-model, and the pose parameters include a rotation matrix and a translation vector.

[0061] Exemplarily, the tissue image and the historical tissue image can be used as the input of the pose sub-model, and the pose sub-model can output the corresponding rotation matrix and translation vector. Specifically, the tissue image and the historical tissue image can be concatenated (English: Concat) so that the concatenated result is input into the pose sub-model. The structure of the pose sub-model is as Figure 2 shown, which can be a ResNet structure (for example, it can be ResNet34). The concatenated result of the tissue image and the historical tissue image is input into the initial convolutional pooling layer, passed through multiple intermediate residual blocks (English: Residual block), and finally the rotation matrix and translation vector are output by the fully connected layer.

[0062] Step 1023, perform three-dimensional fusion on the tissue image, the depth image, and the pose parameters through the fusion sub-model to obtain a three-dimensional tissue image.

[0063] Exemplarily, after obtaining the depth image and the pose parameters, the tissue image, the depth image, and the pose parameters can be input into the fusion sub-model. The fusion sub-model can fuse the tissue image with the corresponding depth image according to a preset fusion algorithm (such as the SurfelMeshing fusion algorithm), and at the same time use the pose parameters to remove distortion, so as to output a three-dimensional tissue image including the three-dimensional coordinates of each pixel point.

[0064] Figure 4 is a flowchart of another method for processing endoscopic images shown according to an exemplary embodiment. As Figure 4 shown, the method may further include:

[0065] Step 105, obtain the movement trajectory of the endoscope and smooth the movement trajectory.

[0066] Step 106, use the smoothed movement trajectory as the center line and establish a tissue template according to a preset template radius.

[0067] For example, the movement trajectory of the endoscope can be obtained in real time. Due to reasons such as unstable insertion technique or inappropriate position of the endoscope, the movement trajectory of the endoscope is relatively complex and has many bends. Therefore, the movement trajectory can be smoothed first. The smoothing process can be, for example, the Moving Average algorithm, or the Savitzky-Golay filtering algorithm, or the spline curve smoothing algorithm. The present disclosure does not make specific limitations in this regard. Taking the Moving Average algorithm as an example, the movement trajectory can be smoothed by formula 1 to obtain the smoothed movement trajectory:

[0068]

[0069] where y s(i) represents the coordinates of the i-th point in the motion trajectory after moving average, M represents the size of the moving window, and y i represents the coordinates of the i-th point in the motion trajectory obtained in real time.

[0070] After that, the smoothed motion trajectory can be used as the center line to establish a tissue template according to the preset template radius. Taking the tissue image as the intestinal image as an example, the tissue template is the intestinal template. A cylinder can be established with a radius of 1 using the smoothed motion trajectory as the center line as the tissue template.

[0071] Figure 5 is a flowchart of another endoscopic image processing method shown according to an exemplary embodiment. As Figure 5 shown, the implementation of step 103 may include:

[0072] Step 1031: Stitch the three-dimensional tissue images corresponding to each tissue image in multiple tissue images collected within a preset time period to obtain a three-dimensional total image.

[0073] Step 1032: Project the three-dimensional total image onto the tissue template to determine the visible area where the projection in the three-dimensional total image overlaps with the tissue template, and the blind area where the projection in the three-dimensional total image does not overlap with the tissue template.

[0074] Exemplarily, multiple tissue images collected within a preset time period can be obtained, and then the three-dimensional tissue images corresponding to each tissue image can be determined in turn in the manner of step 102. Then, the multiple three-dimensional tissue images are stitched into a three-dimensional total image according to the corresponding spatial position relationship. Similarly, taking the tissue image as the intestinal image as an example, each three-dimensional tissue image can be understood as the three-dimensional structure of a small section of the intestine, and the stitched three-dimensional total image can be understood as the three-dimensional structure of a longer section of the intestine.

[0075] After obtaining the three-dimensional total image, the three-dimensional total image can be projected onto the tissue template. The visible area where the projection in the three-dimensional total image overlaps with the tissue template is the range that the field of view can cover when the endoscope collects multiple tissue images, and the blind area where the projection in the three-dimensional total image does not overlap with the tissue template is the range that the field of view cannot cover (i.e., the blind area) when the endoscope collects multiple tissue images. Then, the blind area ratio = blind area / (blind area + visible area).

[0076] Since the projection of the three-dimensional total image on the tissue template is often irregular, the Monte Carlo method (English: Monte Carlo method) can be used to calculate the area of the blind area and the area of the visible area. For example, K test points (K≥100) can be evenly distributed on the tissue template, and then the number of test points Λ in the visible area and the number of test points Ω in the blind area are respectively counted. Then the blind area ratio

[0077] In one implementation, the three-dimensional reconstruction model can be jointly trained with the optical flow model, and the connection relationship between the three-dimensional reconstruction model and the optical flow model is as Figure 6 shown. Figure 7 FIG. is a schematic diagram showing the joint training of a three-dimensional reconstruction model and an optical flow model according to an exemplary embodiment. As Figure 7 shown, the three-dimensional reconstruction model is jointly trained with the optical flow model through the following steps:

[0078] Step A: Input the sample tissue image into the depth sub-model to obtain the sample depth image corresponding to the sample tissue image and the internal parameters of the endoscope that collected the sample tissue image, and input the historical sample tissue image into the depth sub-model to obtain the historical sample depth image corresponding to the historical sample tissue image. The historical sample tissue image is the image collected before the sample tissue image, and the internal parameters of the endoscope include the focal length and the translation size.

[0079] For example, taking the sample tissue image (denoted as I a ) as the input of the depth sub-model, the depth sub-model can output the sample depth image (denoted as D a ) corresponding to the sample tissue image and the internal parameters of the endoscope (denoted as K) that collected the sample tissue image. The internal parameters of the endoscope can include the focal length and the translation size. Similarly, taking the historical sample tissue image (denoted as I b ) as the input of the depth sub-model, the depth sub-model can output the historical sample depth image (denoted as D b ) corresponding to the historical sample tissue image. Among them, the sample tissue image can be extracted from the endoscope video. The endoscope video can be the video recorded during the previous endoscope examination and can be obtained by using different endoscopes to examine different users. Further, when extracting frames from the endoscope video, invalid images (such as images blocked by obstacles, overexposed images, and images with low clarity) can be filtered out. Correspondingly, the historical sample tissue image is the tissue image of the previous frame of the sample tissue image.

[0080] In the training stage, based on multiple stride convolutional layers and multiple transposed convolutional layers, the depth sub-model can also add a linear layer (denoted as linear), as Figure 6As shown. The linear layer can output the endoscopic internal parameters. The form of the endoscopic internal parameter K can be:

[0081]

[0082] where f x and f y respectively represent the focal lengths of the endoscope in the X and Y directions (in pixels), and c x and c y respectively represent the translation sizes of the origin in the X and Y directions (in pixels). The depth sub-model can obtain the endoscopic internal parameters while obtaining the sample depth image, without prior calibration of the endoscope, which is convenient for operation. At the same time, it can adapt to various different endoscopes, improving the applicable range of the depth sub-model.

[0083] Step B: Input the sample tissue image and the historical sample tissue image into the pose sub-model to obtain the sample pose parameters output by the pose sub-model for the sample tissue image and the historical sample tissue image.

[0084] Exemplarily, the sample tissue image and the historical sample tissue image can be used as the input of the pose sub-model. The pose sub-model can output the sample pose parameters between the sample tissue image and the historical sample tissue image. The sample pose parameters include the sample rotation matrix (denoted as R) and the sample translation vector (denoted as t). Specifically, the sample tissue image and the historical sample tissue image can be spliced and the spliced result can be input into the pose sub-model.

[0085] Step C: Input the sample tissue image and the historical sample tissue image into the optical flow model to obtain the sample optical flow map output by the optical flow model for the sample tissue image and the historical sample tissue image.

[0086] Exemplarily, the sample tissue image and the historical sample tissue image can be used as the input of the optical flow model. The optical flow model can output the sample optical flow map between the sample tissue image and the historical sample tissue image. The sample optical flow map includes the offsets of the pixel points at the same positions in the historical sample tissue image and the sample tissue image. That is to say, the sample optical flow map can characterize the movement speed of the observed surface (i.e., the surface of the tissue) during the process of the endoscope from collecting the historical sample tissue image to the sample tissue image. Specifically, the sample tissue image and the historical sample tissue image can be spliced and the spliced result can be input into the optical flow model. The structure of the optical flow model is as Figure 6As shown, it can also be a UNet structure. Compared with the depth sub-model, the optical flow model has fewer channels. The optical flow model includes multiple strided convolutional layers to downsample the sample tissue image. In order to capture the motion state at a relatively long distance, the optical flow model downsamples deeper (i.e., the number of strided convolutional layers in the optical flow model is greater than the number of strided convolutional layers in the depth sub-model). For example, it can be downsampled to 1 / 16 of the resolution of the sample tissue image, and then multiple transposed convolutional layers are used to upsample to the resolution of the sample tissue image (i.e., the number of transposed convolutional layers in the optical flow model is greater than the number of transposed convolutional layers in the depth sub-model) to obtain the corresponding sample optical flow map.

[0087] Step D: Determine the target loss according to the internal parameters of the endoscope, the sample depth image, the historical sample depth image, the sample pose parameters, and the sample optical flow map.

[0088] Step E: Aim to reduce the target loss and use the backpropagation algorithm to train the 3D reconstruction model and the optical flow model.

[0089] Exemplarily, the target loss can be determined according to the internal parameters of the endoscope, the sample depth image, the historical sample depth image, the sample pose parameters, and the sample optical flow map, and aim to reduce the target loss and use the backpropagation algorithm to train the 3D reconstruction model and the optical flow model. When training the 3D reconstruction model and the optical flow model, it is not necessary to perform pre-annotation, and the sample tissue image and the historical sample tissue image for training the 3D reconstruction model and the optical flow model can be obtained quickly. That is to say, the 3D reconstruction model and the optical flow model adopt an unsupervised learning training method.

[0090] Furthermore, the initial learning rate for training the 3D reconstruction model and the optical flow model can be set to: 1e-2, the Batchsize can be set to: 16*4, the optimizer can be selected as: SGD, the Epoch can be set to: 500, and the size of the sample tissue image can be: 384×384.

[0091] Figure 8 It is a schematic diagram showing another way of jointly training the 3D reconstruction model and the optical flow model according to an exemplary embodiment. As Figure 8 shown, the implementation of step D can include:

[0092] Step D1: Interpolate the historical sample tissue image according to the sample depth image, the sample pose parameters, and the internal parameters of the endoscope to obtain the interpolated tissue image.

[0093] Step D2: Determine the photometric loss according to the sample tissue image and the interpolated tissue image.

[0094] Exemplarily, using the sample depth image, the sample pose parameters, and the endoscopic internal parameters, a differentiable bilinear interpolation process can be performed on the historical sample tissue image to obtain an interpolated tissue image. Then, based on the sample tissue image and the interpolated tissue image, the photometric loss can be determined. The interpolated tissue image can be understood as an image obtained by observing the content in the sample tissue image from the perspective of the historical sample tissue image acquisition. According to the principle of bundle adjustment, the pixel gray value of the same spatial point should be fixed in each image. Therefore, when converting images acquired from different perspectives to another perspective, the pixels at the same position in the two images under the same perspective should be the same. Thus, the photometric loss can be understood as the difference between the sample tissue image and the interpolated tissue image. For example, the photometric loss can be determined by Equation 2:

[0095]

[0096] where L p represents the photometric loss, p represents a pixel point, N represents the valid pixel points in the sample tissue image, and |N| represents the number of valid pixel points. I a (p) represents the pixel value of p in the sample tissue image, and I' a (p) represents the pixel value of p in the interpolated tissue image. ||||1 represents the L1 norm, and the L1 norm is more robust for discrete points.

[0097] Step D3: Determine the smoothness loss based on the gradient of the sample depth image and the gradient of the sample tissue image.

[0098] Exemplarily, in the low-texture region of the sample tissue image (or the interpolated tissue image), since there is less image feature information, the manifestation of the photometric loss is weak. Therefore, the smoothness loss can be added as a regularization term to constrain the generated sample depth image. The smoothness loss can be determined based on the gradient of the sample depth image and the gradient of the sample tissue image. The smoothness loss can ensure that the sample depth image is generated under the guidance of the sample tissue image, so that the generated sample depth map can retain more gradient information at the edge, that is, the edge is more obvious and the detail information is more abundant. For example, the smoothness loss can be determined by Equation 3:

[0099]

[0100] where L s represents the smoothness loss, represents the gradient of p in the sample tissue image, represents the gradient of p in the sample depth image.

[0101] Step D4: Transform the sample depth image into a first depth image according to the sample pose parameters and the endoscopic internal parameters.

[0102] Step D5: Transform the historical sample depth image into a second depth image according to the sample optical flow map, the sample pose parameters, and the endoscopic internal parameters.

[0103] Step D6: Determine the consistency loss according to the first depth image and the second depth image.

[0104] Exemplarily, since the sample tissue image and the historical sample tissue image face the same three-dimensional space, there is spatial consistency between the sample depth image and the historical sample depth image. The sample depth image can be transformed into the first depth image (denoted as ) by using the sample pose parameters and the endoscopic internal parameters, and the historical sample depth image can be transformed into the second depth image (denoted as D′ b ) by using the sample optical flow map, the sample pose parameters, and the endoscopic internal parameters. Among them, the first depth image can be understood as the depth image obtained by transforming the sample depth image through pose transformation and observing the content in the sample tissue image from the perspective of collecting the historical sample tissue image. The second depth image can be understood as the depth image obtained by interpolating the historical sample depth image to obtain the content in the sample tissue image from the perspective of collecting the historical sample tissue image.

[0105] Then, determine the consistency loss according to the first depth image and the second depth image. That is to say, the consistency loss can reflect the difference between the first depth image and the second depth image. Through training, the consistency can be propagated to multiple sample depth images, which also ensures the scale consistency of multiple sample depth images. It is equivalent to smoothing multiple sample depth images and ensuring spatial consistency. For example, the consistency loss can be determined by Equation 4:

[0106]

[0107] where L G represents the consistency loss, represents the depth of p in the first depth image, and D′ b (p) represents the depth of p in the second depth image.

[0108] Step D7: Determine the target loss according to the photometric loss, the smoothness loss, and the consistency loss.

[0109] Exemplarily, the target loss can be determined according to the photometric loss, the smoothness loss, and the consistency loss. For example, the photometric loss, the smoothness loss, and the consistency loss can be weighted and summed by Equation 5 to obtain the target loss:

[0110] L = αL p + βL s + γL G Equation 5

[0111] Among them, α, β, and γ are the weights corresponding to the photometric loss, the smoothness loss, and the consistency loss respectively. Among them, α can be 0.7, β can be 0.7, and γ can be 0.3.

[0112] In one implementation, step D2 can be implemented in the following manner:

[0113] Determine the photometric loss according to the sample tissue image, the interpolated tissue image, and the structural similarity between the sample tissue image and the interpolated tissue image.

[0114] Exemplarily, when the endoscope acquires the sample tissue image and the historical sample tissue image, the illumination conditions may change. Therefore, SSIM (English: Structural Similarity, Chinese: Structural Similarity) can be introduced to determine the photometric loss to avoid the interference of illumination condition changes on the photometric loss. SSIM can reflect the similarity of local structures. The improved photometric loss can be determined by formula 6:

[0115]

[0116] Among them, λ1 and λ2 respectively represent preset weights, and SSIM(p) represents the per-pixel SSIM between the sample tissue image and the interpolated tissue image. Among them, λ1 can be 0.7 and λ2 can be 0.3.

[0117] In another implementation, step D2 may further include:

[0118] Step 1) Determine a mask matrix according to the difference degree between the first depth image and the second depth image. The mask matrix includes the weight corresponding to each pixel point in the sample tissue image.

[0119] Step 2) Correct the photometric loss according to the mask matrix.

[0120] Exemplarily, if the tissue corresponding to the sample tissue image undergoes peristalsis, it may affect the photometric loss. Therefore, the masking matrix can be determined based on the difference between the first depth image and the second depth image, where the masking matrix includes the weights corresponding to each pixel point in the sample tissue image, and the weights are negatively correlated with the difference between the first depth image and the second depth image. That is to say, if the difference between a certain pixel point in the first depth image and the second depth image is large, it indicates that the position corresponding to this pixel point may have undergone peristalsis when the endoscope collected the sample tissue image and the historical sample tissue image, which will interfere with the photometric loss. Therefore, the weight corresponding to this pixel point is small. If the difference between this pixel point in the first depth image and the second depth image is small, it indicates that the position corresponding to this pixel point has little change (close to the static state) when the endoscope collects the sample tissue image and the historical sample tissue image, and will not interfere with the photometric loss. Therefore, the weight corresponding to this pixel point is large. The masking matrix can be determined by Equation 7:

[0121]

[0122] where M represents the masking matrix, and M(p) represents the weight corresponding to p in the sample tissue image included in the masking matrix, represents the difference between the first depth image and the second depth image.

[0123] Then, the photometric loss is corrected according to the masking matrix. In this way, the corrected photometric loss can avoid the interference caused by tissue peristalsis. The corrected photometric loss can be determined by Equation 8:

[0124]

[0125] Figure 9 is a flowchart of another method for processing endoscopic images shown according to an exemplary embodiment. As Figure 9 shown, after step 103, the method may further include:

[0126] Step 107, output the blind area ratio, and issue a prompt message when the blind area ratio is greater than or equal to a preset ratio threshold, where the prompt message is used to indicate the risk of missed detection.

[0127] For example, after determining the blind area ratio, the blind area ratio can be output. For example, the blind area ratio can be displayed in real time on the display interface for displaying the tissue image, so as to display the inspection range during the endoscope inspection in real time. Further, when the blind area ratio is greater than or equal to a preset ratio threshold (for example, it can be 20%), a prompt message can be sent to prompt the doctor that there is a large blind area in the current field of view of the endoscope and there is a risk of missed detection. The presentation form of the prompt message can include at least one of the following: text form, image form, and sound form. For example, the prompt message can be a text prompt such as "High risk of missed detection currently", "Please recheck", "Please perform endoscope withdrawal", or an image prompt. The prompt message can also be a sound prompt such as voice, a beep of a specified frequency, or an alarm sound. In this way, the doctor can adjust the direction of the endoscope, or perform endoscope withdrawal, or recheck according to the prompt message. Thus, the blind area ratio can be monitored in real time during the doctor's endoscope inspection, and a prompt can be given when the blind area ratio is large, so as to avoid missed detection and ensure the effectiveness of the inspection.

[0128] In summary, the present disclosure first obtains the tissue image collected by the endoscope at the current moment, and then determines the depth image corresponding to the tissue image, the pose parameters between the tissue image and the historical tissue image through a pre-trained three-dimensional reconstruction model according to the tissue image and the historical tissue image, and determines the three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image, and the pose parameters. The historical tissue image is the image collected by the endoscope before the current moment. Finally, the three-dimensional tissue image is projected onto the tissue template to determine the visible area where the projection in the three-dimensional tissue image overlaps with the tissue template, and the blind area where the projection in the three-dimensional tissue image does not overlap with the tissue template, and the blind area ratio during the endoscope inspection is determined according to the visible area and the blind area. The present disclosure determines the depth image and the pose parameters according to the tissue image through the three-dimensional reconstruction model, and performs three-dimensional reconstruction based on this, so as to quickly obtain an accurate three-dimensional tissue image. And the blind area ratio is determined in combination with the tissue template based on the three-dimensional tissue image, which can timely reflect the inspection range during the inspection, so as to avoid missed detection and ensure the effectiveness of the inspection.

[0129] Figure 10 is a block diagram of a processing device for endoscope images shown according to an exemplary embodiment, as Figure 10 shown, the device 200 may include:

[0130] An acquisition module 201, configured to acquire a tissue image collected by the endoscope at the current moment.

[0131] The reconstruction module 202 is configured to determine a depth image corresponding to the tissue image, pose parameters between the tissue image and the historical tissue image based on the tissue image and the historical tissue image through a pre-trained three-dimensional reconstruction model, and determine a three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image, and the pose parameters. The historical tissue image is an image acquired by the endoscope before the current moment.

[0132] The projection module 203 is configured to project the three-dimensional tissue image onto the tissue template to determine a visible area where the projection in the three-dimensional tissue image overlaps with the tissue template, and a blind area where the projection in the three-dimensional tissue image does not overlap with the tissue template. The tissue template is used to represent the overall shape of the tissue examined by the endoscope.

[0133] The processing module 204 is configured to determine the blind area ratio during the endoscope examination according to the visible area and the blind area.

[0134] Figure 11 is a block diagram of another endoscopic image processing device shown according to an exemplary embodiment, as Figure 11 shown. The three-dimensional reconstruction model includes: a depth sub-model, a pose sub-model, and a fusion sub-model. The reconstruction module 202 may include:

[0135] The depth image determination sub-module 2021 is configured to input the tissue image into the depth sub-model to obtain the depth image output by the depth sub-model.

[0136] The pose determination sub-module 2022 is configured to input the tissue image and the historical tissue image into the pose sub-model to obtain the pose parameters output by the pose sub-model. The pose parameters include a rotation matrix and a translation vector.

[0137] The three-dimensional fusion sub-module 2023 is configured to perform three-dimensional fusion according to the tissue image, the depth image, and the pose parameters through the fusion sub-model to obtain a three-dimensional tissue image.

[0138] Figure 12 is a block diagram of another endoscopic image processing device shown according to an exemplary embodiment, as Figure 12 shown. The device 200 may further include:

[0139] The trajectory processing module 205 is configured to obtain the movement trajectory of the endoscope and perform smoothing processing on the movement trajectory.

[0140] The template establishment module 206 is configured to use the smoothed movement trajectory as the center line to establish a tissue template according to a preset template radius.

[0141] Figure 13 is a block diagram of another endoscopic image processing device shown according to an exemplary embodiment, as Figure 13As shown, the projection module 203 may include:

[0142] A stitching sub-module 2031 for stitching the three-dimensional tissue images corresponding to each tissue image among multiple tissue images collected within a preset time period to obtain a three-dimensional total image.

[0143] A projection sub-module 2032 for projecting the three-dimensional total image onto a tissue template to determine a visible area where the projection in the three-dimensional total image overlaps with the tissue template and a blind area where the projection in the three-dimensional total image does not overlap with the tissue template.

[0144] In one implementation, the three-dimensional reconstruction model is jointly trained with the optical flow model through the following steps:

[0145] Step A: Input a sample tissue image into the depth sub-model to obtain a sample depth image corresponding to the sample tissue image and the internal parameters of the endoscope that collected the sample tissue image, and input a historical sample tissue image into the depth sub-model to obtain a historical sample depth image corresponding to the historical sample tissue image. The historical sample tissue image is an image collected before the sample tissue image. The internal parameters of the endoscope include the focal length and the translation size.

[0146] Step B: Input the sample tissue image and the historical sample tissue image into the pose sub-model to obtain the sample pose parameters output by the pose sub-model between the sample tissue image and the historical sample tissue image.

[0147] Step C: Input the sample tissue image and the historical sample tissue image into the optical flow model to obtain the sample optical flow map output by the optical flow model between the sample tissue image and the historical sample tissue image.

[0148] Step D: Determine the target loss according to the internal parameters of the endoscope, the sample depth image, the historical sample depth image, the sample pose parameters, and the sample optical flow map.

[0149] Step E: Use the backpropagation algorithm to train the three-dimensional reconstruction model and the optical flow model with the goal of reducing the target loss.

[0150] In another implementation, step D may include:

[0151] Step D1: Interpolate the historical sample tissue image according to the sample depth image, the sample pose parameters, and the internal parameters of the endoscope to obtain an interpolated tissue image.

[0152] Step D2: Determine the photometric loss according to the sample tissue image and the interpolated tissue image.

[0153] Step D3: Determine the smoothness loss according to the gradient of the sample depth image and the gradient of the sample tissue image.

[0154] Step D4: Transform the sample depth image into a first depth image according to the sample pose parameters and the endoscope internal parameters.

[0155] Step D5: Transform the historical sample depth image into a second depth image according to the sample optical flow map, the sample pose parameters, and the endoscope internal parameters.

[0156] Step D6: Determine the consistency loss according to the first depth image and the second depth image.

[0157] Step D7: Determine the target loss according to the photometric loss, the smoothness loss, and the consistency loss.

[0158] In another implementation, step D2 can be implemented in the following way:

[0159] Determine the photometric loss according to the sample tissue image, the interpolated tissue image, and the structural similarity between the sample tissue image and the interpolated tissue image.

[0160] In another implementation, step D2 may further include:

[0161] Step 1): Determine a mask matrix according to the difference between the first depth image and the second depth image. The mask matrix includes the weight corresponding to each pixel point in the sample tissue image.

[0162] Step 2): Correct the photometric loss according to the mask matrix.

[0163] Figure 14 It is a block diagram of another endoscope image processing device shown according to an exemplary embodiment. As Figure 14 shown, the device 200 may further include:

[0164] An output module 207, configured to output the blind area ratio after determining the blind area ratio during the endoscope examination according to the three-dimensional tissue image and the tissue template, and send a prompt message when the blind area ratio is greater than or equal to a preset ratio threshold. The prompt message is used to indicate the risk of missed detection.

[0165] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0166] In summary, the present disclosure first obtains a tissue image collected by an endoscope at the current moment, and then, based on the tissue image and historical tissue images, determines a depth image corresponding to the tissue image, pose parameters between the tissue image and the historical tissue images through a pre-trained three-dimensional reconstruction model, and determines a three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image, and the pose parameters. Here, the historical tissue images are images collected by the endoscope before the current moment. Finally, the three-dimensional tissue image is projected onto a tissue template to determine a visible area where the projection in the three-dimensional tissue image overlaps with the tissue template, and a blind area where the projection in the three-dimensional tissue image does not overlap with the tissue template, and determines a blind area ratio during the endoscope examination according to the visible area and the blind area. The present disclosure determines the depth image and pose parameters based on the tissue image through the three-dimensional reconstruction model, and performs three-dimensional reconstruction therewith, so as to quickly obtain an accurate three-dimensional tissue image. And based on the three-dimensional tissue image, the blind area ratio is determined in combination with the tissue template, which can timely reflect the examination range during the examination, thereby avoiding missed detections and ensuring the effectiveness of the examination.

[0167] Reference is now made to Figure 15 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing the embodiments of the present disclosure (for example, the execution subject of the embodiments of the present disclosure, which may be a terminal device or a server). The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 15 The electronic device shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0168] As Figure 15 shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0169] Typically, the following devices can be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 can allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 15 an electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.

[0170] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the method of the embodiment of the present disclosure are executed.

[0171] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0172] In some embodiments, the terminal device and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0173] The above computer-readable medium can be included in the above electronic device; it can also exist separately without being assembled into the electronic device.

[0174] The above computer-readable medium carries one or more programs which, when executed by the electronic device, cause the electronic device to: obtain a tissue image collected by an endoscope at the current moment; determine, according to the tissue image and a historical tissue image, a depth image corresponding to the tissue image, pose parameters between the tissue image and the historical tissue image through a pre-trained three-dimensional reconstruction model, and determine a three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image and the pose parameters, the historical tissue image being an image collected by the endoscope before the current moment; project the three-dimensional tissue image onto a tissue template to determine a visible area where the projection in the three-dimensional tissue image overlaps with the tissue template and a blind area where the projection in the three-dimensional tissue image does not overlap with the tissue template, the tissue template being used to represent the overall shape of the tissue examined by the endoscope; and determine a blind area ratio during the endoscope examination according to the visible area and the blind area.

[0175] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0176] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0177] The modules described in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of a module does not constitute a limitation on the module itself. For example, the acquisition module can also be described as "the module for acquiring tissue images".

[0178] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0179] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0180] According to one or more embodiments of the present disclosure, Example 1 provides a method for processing endoscopic images, including: obtaining a tissue image collected by an endoscope at the current moment; determining, according to the tissue image and a historical tissue image, a depth image corresponding to the tissue image, pose parameters between the tissue image and the historical tissue image through a pre-trained three-dimensional reconstruction model, and determining a three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image, and the pose parameters, where the historical tissue image is an image collected by the endoscope before the current moment; projecting the three-dimensional tissue image onto a tissue template to determine a visible area where the projection in the three-dimensional tissue image overlaps with the tissue template, and a blind area where the projection in the three-dimensional tissue image does not overlap with the tissue template, where the tissue template is used to represent the overall shape of the tissue examined by the endoscope; and determining a blind area ratio during the endoscope examination according to the visible area and the blind area.

[0181] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, where the three-dimensional reconstruction model includes: a depth sub-model, a pose sub-model, and a fusion sub-model; the step of determining, according to the tissue image and a historical tissue image, a depth image corresponding to the tissue image, pose parameters between the tissue image and the historical tissue image through a pre-trained three-dimensional reconstruction model, and determining a three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image, and the pose parameters includes: inputting the tissue image into the depth sub-model to obtain the depth image output by the depth sub-model; inputting the tissue image and the historical tissue image into the pose sub-model to obtain the pose parameters output by the pose sub-model, where the pose parameters include a rotation matrix and a translation vector; and performing three-dimensional fusion through the fusion sub-model according to the tissue image, the depth image, and the pose parameters to obtain the three-dimensional tissue image.

[0182] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 1, and the method further includes: obtaining a movement trajectory of the endoscope and smoothing the movement trajectory; using the smoothed movement trajectory as a center line to establish the tissue template according to a preset template radius.

[0183] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 1. The method of projecting the three-dimensional tissue image onto the tissue template to determine the visible region where the projection in the three-dimensional image overlaps with the tissue template and the blind region where the projection in the three-dimensional image does not overlap with the tissue template includes: splicing the three-dimensional tissue images corresponding to each of the plurality of tissue images collected within a preset time period to obtain a three-dimensional total image; projecting the three-dimensional total image onto the tissue template to determine the visible region where the projection in the three-dimensional total image overlaps with the tissue template and the blind region where the projection in the three-dimensional total image does not overlap with the tissue template.

[0184] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 2. The three-dimensional reconstruction model is jointly trained with the optical flow model through the following steps: inputting the sample tissue image into the depth sub-model to obtain the sample depth image corresponding to the sample tissue image and the internal parameters of the endoscope that collected the sample tissue image, and inputting the historical sample tissue image into the depth sub-model to obtain the historical sample depth image corresponding to the historical sample tissue image, where the historical sample tissue image is an image collected before the sample tissue image, and the internal parameters of the endoscope include the focal length and the translation size; inputting the sample tissue image and the historical sample tissue image into the pose sub-model to obtain the sample pose parameters output by the pose sub-model, which are between the sample tissue image and the historical sample tissue image; inputting the sample tissue image and the historical sample tissue image into the optical flow model to obtain the sample optical flow map output by the optical flow model, which is between the sample tissue image and the historical sample tissue image; determining the target loss according to the internal parameters of the endoscope, the sample depth image, the historical sample depth image, the sample pose parameters, and the sample optical flow map; and training the three-dimensional reconstruction model and the optical flow model using the backpropagation algorithm with the goal of reducing the target loss.

[0185] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 5. Determining a target loss according to the endoscope internal parameters, the sample depth image, the historical sample depth image, the sample pose parameters, and the sample optical flow map includes: interpolating the historical sample tissue image according to the sample depth image, the sample pose parameters, and the endoscope internal parameters to obtain an interpolated tissue image; determining a photometric loss according to the sample tissue image and the interpolated tissue image; determining a smoothness loss according to the gradient of the sample depth image and the gradient of the sample tissue image; transforming the sample depth image into a first depth image according to the sample pose parameters and the endoscope internal parameters; transforming the historical sample depth image into a second depth image according to the sample optical flow map, the sample pose parameters, and the endoscope internal parameters; determining a consistency loss according to the first depth image and the second depth image; and determining the target loss according to the photometric loss, the smoothness loss, and the consistency loss.

[0186] According to one or more embodiments of the present disclosure, Example 7 provides the method of Example 6. Determining the photometric loss according to the sample tissue image and the interpolated tissue image includes: determining the photometric loss according to the sample tissue image, the interpolated tissue image, and the structural similarity between the sample tissue image and the interpolated tissue image.

[0187] According to one or more embodiments of the present disclosure, Example 8 provides the method of Example 7. Determining the photometric loss according to the sample tissue image and the interpolated tissue image further includes: determining a mask matrix according to the difference degree between the first depth image and the second depth image, where the mask matrix includes weights corresponding to each pixel point in the sample tissue image; and correcting the photometric loss according to the mask matrix.

[0188] According to one or more embodiments of the present disclosure, Example 9 provides the method of Examples 1 to 8. After determining the blind area ratio during the endoscopy according to the three-dimensional tissue image and the tissue template, the method further includes: outputting the blind area ratio, and issuing a prompt message when the blind area ratio is greater than or equal to a preset ratio threshold, where the prompt message is used to indicate the risk of missed detection.

[0189] According to one or more embodiments of the present disclosure, Example 10 provides a processing device for endoscopic images, including: an acquisition module, configured to acquire a tissue image collected by an endoscope at the current moment; a reconstruction module, configured to determine a depth image corresponding to the tissue image, pose parameters between the tissue image and a historical tissue image, and a three-dimensional tissue image corresponding to the tissue image according to the tissue image and the historical tissue image through a pre-trained three-dimensional reconstruction model, where the historical tissue image is an image collected by the endoscope before the current moment; a projection module, configured to project the three-dimensional tissue image onto a tissue template to determine a visible area where the projection in the three-dimensional tissue image overlaps with the tissue template, and a blind area where the projection in the three-dimensional tissue image does not overlap with the tissue template, where the tissue template is used to represent the overall shape of the tissue examined by the endoscope; and a processing module, configured to determine a blind area ratio during the endoscope examination according to the visible area and the blind area.

[0190] According to one or more embodiments of the present disclosure, Example 11 provides a computer-readable medium having a computer program stored thereon, and when the program is executed by a processing device, the steps of the methods described in Examples 1 to 9 are implemented.

[0191] According to one or more embodiments of the present disclosure, Example 12 provides an electronic device, including: a storage device having a computer program stored thereon; and a processing device, configured to execute the computer program in the storage device to implement the steps of the methods described in Examples 1 to 9.

[0192] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0193] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0194] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. With regard to the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein.

Claims

1. A method for processing endoscopic images, characterized in that, The method includes: Obtaining a tissue image collected by an endoscope at the current moment; According to the tissue image and the historical tissue image, determining a depth image corresponding to the tissue image, pose parameters between the tissue image and the historical tissue image through a pre-trained three-dimensional reconstruction model, and determining a three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image and the pose parameters, where the historical tissue image is an image collected by the endoscope before the current moment; Projecting the three-dimensional tissue image onto a tissue template to determine a visible region where the projection in the three-dimensional tissue image overlaps with the tissue template, and a blind area region where the projection in the three-dimensional tissue image does not overlap with the tissue template, where the tissue template is used to represent the overall shape of the tissue examined by the endoscope; Determining a blind area ratio during the endoscope examination according to the visible region and the blind area region; Wherein, the three-dimensional reconstruction model includes: a depth sub-model and a pose sub-model; The determining, according to the tissue image and the historical tissue image, a depth image corresponding to the tissue image and pose parameters between the tissue image and the historical tissue image through a pre-trained three-dimensional reconstruction model includes: Inputting the tissue image into the depth sub-model to obtain the depth image output by the depth sub-model; Inputting the tissue image and the historical tissue image into the pose sub-model to obtain the pose parameters output by the pose sub-model, where the pose parameters include a rotation matrix and a translation vector.

2. The method according to claim 1, characterized in that, The three-dimensional reconstruction model further includes a fusion sub-model; The determining the three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image and the pose parameters includes: Performing three-dimensional fusion according to the tissue image, the depth image and the pose parameters through the fusion sub-model to obtain the three-dimensional tissue image.

3. The method according to claim 1, characterized in that, The method further includes: Obtaining the movement trajectory of the endoscope and smoothing the movement trajectory; Taking the smoothed movement trajectory as a center line and establishing the tissue template according to a preset template radius.

4. The method according to claim 1, characterized in that, The projecting the three-dimensional tissue image onto the tissue template to determine a visible region where the projection in the three-dimensional tissue image overlaps with the tissue template, and a blind area region where the projection in the three-dimensional tissue image does not overlap with the tissue template includes: Stitching the three-dimensional tissue images corresponding to each of the tissue images collected within a preset time period to obtain a three-dimensional total image; Projecting the three-dimensional total image onto the tissue template to determine the visible region where the projection in the three-dimensional total image overlaps with the tissue template, and the blind area region where the projection in the three-dimensional total image does not overlap with the tissue template.

5. The method according to claim 2, characterized in that, The three-dimensional reconstruction model is jointly trained with an optical flow model through the following steps: Input the sample tissue image into the depth sub-model to obtain the sample depth image corresponding to the sample tissue image and the internal parameters of the endoscope that captured the sample tissue image, and input the historical sample tissue image into the depth sub-model to obtain the historical sample depth image corresponding to the historical sample tissue image. The historical sample tissue image is an image captured before the sample tissue image. The internal parameters of the endoscope include the focal length and the translation size; Input the sample tissue image and the historical sample tissue image into the pose sub-model to obtain the sample pose parameters output by the pose sub-model, which are between the sample tissue image and the historical sample tissue image; Input the sample tissue image and the historical sample tissue image into the optical flow model to obtain the sample optical flow map output by the optical flow model, which is between the sample tissue image and the historical sample tissue image; Determine the target loss according to the internal parameters of the endoscope, the sample depth image, the historical sample depth image, the sample pose parameters, and the sample optical flow map; Taking reducing the target loss as the goal, train the 3D reconstruction model and the optical flow model using the backpropagation algorithm.

6. The method according to claim 5, characterized in that, The determining the target loss according to the internal parameters of the endoscope, the sample depth image, the historical sample depth image, the sample pose parameters, and the sample optical flow map includes: Interpolate the historical sample tissue image according to the sample depth image, the sample pose parameters, and the internal parameters of the endoscope to obtain an interpolated tissue image; Determine the photometric loss according to the sample tissue image and the interpolated tissue image; Determine the smoothness loss according to the gradient of the sample depth image and the gradient of the sample tissue image; Transform the sample depth image into a first depth image according to the sample pose parameters and the internal parameters of the endoscope; Transform the historical sample depth image into a second depth image according to the sample optical flow map, the sample pose parameters, and the internal parameters of the endoscope; Determine the consistency loss according to the first depth image and the second depth image; Determine the target loss according to the photometric loss, the smoothness loss, and the consistency loss.

7. The method according to claim 6, characterized in that, The determining the photometric loss according to the sample tissue image and the interpolated tissue image includes: Determine the photometric loss according to the sample tissue image, the interpolated tissue image, and the structural similarity between the sample tissue image and the interpolated tissue image.

8. The method according to claim 7, wherein The determining the photometric loss according to the sample tissue image and the interpolated tissue image further includes: Determine a mask matrix according to the difference degree between the first depth image and the second depth image. The mask matrix includes the weights corresponding to each pixel point in the sample tissue image; Correct the photometric loss according to the mask matrix.

9. The method according to any one of claims 1-8, wherein After determining the blind area ratio during the endoscopy according to the 3D tissue image and the tissue template, the method further includes: Output the blind area ratio, and issue a prompt message when the blind area ratio is greater than or equal to a preset ratio threshold, where the prompt message is used to indicate the risk of missed detection.

10. A processing device for endoscopic images, wherein The device includes: An acquisition module, configured to acquire a tissue image collected by an endoscope at the current moment; A reconstruction module, configured to determine a depth image corresponding to the tissue image, pose parameters between the tissue image and the historical tissue image through a pre-trained three-dimensional reconstruction model according to the tissue image and the historical tissue image, and determine a three-dimensional tissue image corresponding to the tissue image according to the tissue image, the depth image, and the pose parameters, where the historical tissue image is an image collected by the endoscope before the current moment; A projection module, configured to project the three-dimensional tissue image onto a tissue template to determine a visible area where the projection in the three-dimensional tissue image overlaps with the tissue template, and a blind area where the projection in the three-dimensional tissue image does not overlap with the tissue template, where the tissue template is used to represent the overall shape of the tissue examined by the endoscope; A processing module, configured to determine the blind area ratio during the endoscope examination according to the visible area and the blind area; Wherein, the three-dimensional reconstruction model includes: a depth sub-model and a pose sub-model; The reconstruction module is configured to input the tissue image into the depth sub-model to obtain the depth image output by the depth sub-model; input the tissue image and the historical tissue image into the pose sub-model to obtain the pose parameters output by the pose sub-model, where the pose parameters include a rotation matrix and a translation vector.

11. A computer-readable medium having a computer program stored thereon, wherein When the program is executed by a processing device, it implements the steps of the method according to any one of claims 1-9.

12. An electronic device, wherein It includes: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Complex fabric surface three-dimensional reconstruction system and method under non-single visual angle

    CN110415332A

  • Endoscope image classification model training method, and image classification method and device

    CN113496489A