Video image generation method and device based on endoscopic imaging and endoscope system
By clarifying the image of the endoscope field of view and enhancing the field of view scene, the characteristics of suspended impurities are blurred, the problem of endoscopic imaging blur caused by suspended impurities is solved, the clarity of the video image is improved, and the surgical risk is reduced.
Patent Information
- Application Number
- CN202310762273.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-26
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2043-06-26
AI Technical Summary
Suspended impurities cause video images to become blurred during endoscopic imaging, affecting the accurate operation of surgical tools and increasing surgical risks.
By optimizing the endoscope field of view image, including image sharpening and field of view scene enhancement, the features of suspended impurities are blurred while preserving the features of the field of view scene, and neural networks or image processing algorithms are used for defogging and detail enhancement.
It improves the clarity of endoscopic video images, reduces the blurring interference of suspended impurities on imaging, and reduces the risk of surgical misinjury.
Smart Images

Figure CN116777786B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to medical imaging technology, and in particular to a method for generating a video image based on endoscopic imaging, a device for generating a video image based on endoscopic imaging, and an endoscope system. Background Art
[0002] Endoscopes are often used in surgical procedures such as minimally invasive surgery. By passing the endoscope through the human body cavity and into the lesion inside the body, the controlled operation process of the surgical tool at the lesion inside the body can be captured in real time.
[0003] In the case where there are hard growths such as stones at the lesion site in the body, the controlled operation process of the surgical tool at the lesion site includes crushing the hard growths. For example, the surgical tool may include a laser fiber head to crush the hard growths using the holmium laser generated by the laser fiber head.
[0004] However, the human body may contain suspended impurities that may affect the human eye's observation (e.g., debris or droplets in bodily fluids, water vapor in the body, and debris produced by the holmium laser from the laser fiber optic tip when hard growths such as stones at lesions are shattered). Regardless of the substance of the suspended impurities that affect the human eye's observation, if these suspended impurities appear within the field of view of the endoscope, the video stream containing the endoscopic field of view image will blur the visual field features of the operator's interest (e.g., human tissue and surgical tools such as the laser fiber optic tip), making it difficult for the operator to accurately understand the controlled operation of the surgical tools at the lesion in the body, thereby causing surgical risks (e.g., holmium laser light from the laser fiber optic tip that deviates from the hard growth causing thermal damage to human tissue).
[0005] It can be seen that how to reduce the blurring interference of suspended impurities on the video presentation effect based on endoscopic imaging has become a technical problem to be solved in the existing technology. Summary of the Invention
[0006] Embodiments of the present application provide a method for generating a video image based on endoscopic imaging, a device for generating a video image based on endoscopic imaging, and an endoscope system, which help reduce the blurring interference of suspended impurities on the video presentation effect based on endoscopic imaging.
[0007] In one embodiment of the present application, a method for generating a video image based on endoscopic imaging includes:
[0008] Acquiring an endoscopic field of view image based on endoscopic imaging, wherein the endoscopic field of view image includes field of view scene features within the field of view of the endoscope;
[0009] Based on the image optimization of the endoscopic field of view image, a video image for visual presentation is generated, wherein the image optimization is used to blur the suspended impurity features while preserving the fidelity of the field of view scene features when the endoscopic field of view image also includes suspended impurity features that cause blurring of the field of view scene features.
[0010] In some examples, optionally, the generation of a video image for visual presentation based on image optimization of the endoscopic field of view image includes: generating a sharpened image based on image sharpening processing of the endoscopic field of view image; generating the video image based on field of view scene enhancement processing of the sharpened image; wherein the image sharpening processing is used to enhance the local contrast of the image area where the suspended impurity feature is located, and the field of view scene enhancement processing is used to enhance the feature information of the field of view scene feature.
[0011] In some examples, optionally, it also includes: based on the acquired human-computer interaction instruction, determining the currently selected clarity level for the clarity processing among at least two preset clarity levels, and the currently selected enhancement level for the field of view scene enhancement processing among at least two preset enhancement levels.
[0012] In some examples, optionally, generating a cleared image based on image clearing processing of the endoscopic field of view image includes: performing defogging processing on the endoscopic field of view image using the suspended impurity feature as a fog feature to obtain the cleared image.
[0013] In some examples, optionally, the dehazing process is performed on the endoscopic field of view image using the suspended impurity features as the fog features to obtain the cleared image, comprising: inputting the endoscopic field of view image into a pre-trained neural network, wherein the training samples of the neural network include fog sample features having characteristic morphologies of the suspended impurity features, and the cleared image comprises the output image of the neural network.
[0014] In some examples, optionally, the defogging process is performed on the endoscopic field of view image using the suspended impurity feature as the fog feature to obtain the cleared image, including: performing optical enhancement process on the endoscopic field of view image, defogging the optical enhanced image obtained by the optical enhancement process based on a preset defogging algorithm, and performing inverse processing of the optical enhancement process on the defogged image obtained after the defogging process to obtain the cleared image.
[0015] In some examples, optionally, the generating of the video image based on the field of view scene enhancement processing of the sharpened image includes: performing global detail enhancement processing on the sharpened image to obtain a global detail enhanced image; determining matching feature information of the field of view scene features in the global detail enhanced image; obtaining a local detail enhanced image based on highlighting processing of the matching feature information in the global detail enhanced image; and generating the video image based on the sharpened image and the local detail enhanced image.
[0016] In some examples, optionally, it also includes: generating a saliency map with the field of view scene feature as the highlighted object based on the sharpened image, and / or obtaining feature identification information of the field of view scene feature in the sharpened image; determining the matching feature information of the field of view scene feature in the global detail enhanced image includes: using the saliency map and / or the feature identification information as indication information, determining the matching feature information of the field of view scene feature in the global detail enhanced image.
[0017] In some examples, optionally, generating a saliency map based on the sharpened image with the field of view scene features as the highlighted object includes: determining the local contrast of the sharpened image at each scale based on a mean filtered image obtained by performing multi-scale mean filtering on the sharpened image; and obtaining the saliency map based on weighted fusion of the local contrast of the sharpened image at each scale.
[0018] In some examples, optionally, generating the video image based on the sharpened image and the local detail enhanced image includes: performing brightness equalization processing on the sharpened image to obtain a brightness equalized image; and generating the video image based on image fusion of the brightness equalized image and the local detail enhanced image.
[0019] In some examples, optionally, performing brightness equalization processing on the cleared image to obtain a brightness-balanced image includes: performing illumination estimation on the cleared image; determining high-brightness areas and low-brightness areas in the cleared image divided by a brightness threshold based on illumination estimation data obtained by the illumination estimation; and obtaining the brightness-balanced image by performing brightness suppression on the high-brightness areas of the cleared image and brightness compensation on the low-brightness areas of the cleared image.
[0020] In some examples, optionally, performing illumination estimation on the sharpened image includes: generating illumination estimation data for the sharpened image based on a mean filtered image obtained by performing multi-scale mean filtering on the sharpened image.
[0021] In some examples, optionally, the generating of a video image for visual presentation based on image optimization of the endoscopic field of view image includes: inputting the endoscopic field of view image into an end-to-end neural network based on deep learning; wherein the end-to-end neural network is used to implement a mapping transformation from an input image to an output image, the mapping transformation includes feature blurring with the suspended impurity feature as a blurring target, and the video image is the output image of the end-to-end neural network.
[0022] In some examples, optionally, it also includes: training the end-to-end neural network using an image sample set, wherein: the image sample set includes multiple groups of image sample pairs, the first image sample and the second image sample in each group of the image sample pairs include the same suspended impurity sample feature, the suspended impurity sample feature is presented in the first image sample as a visible state higher than a preset recognition threshold, and the suspended impurity sample feature is presented in the second image sample as a blurred state lower than the preset recognition threshold; the training goal of the end-to-end neural network is set to: the conversion loss from the first image sample to the second image sample in each group of the image samples is lower than the target loss.
[0023] In some examples, optionally, the end-to-end neural network is a cyclic generative adversarial network, wherein: in the process of training the cyclic generative adversarial network using the image sample set, the cyclic generative adversarial network cyclically realizes bidirectional conversion between the first image sample and the second image sample in each group of the image samples; when the weighted value of the cycle consistency loss and the cycle perceptual consistency loss of the cyclic generative adversarial network reaches a preset threshold, the training process ends because the training goal is achieved.
[0024] In some examples, optionally, the generation of a video image for visual presentation based on image optimization of the endoscopic field of view image includes: determining an optimization mode for the image optimization based on an acquired human-computer interaction instruction, wherein the optimization mode includes a first optimization mode and a second optimization mode that are selectively enabled; in the first optimization mode, sequentially performing image sharpening processing and field of view scene enhancement processing on the endoscopic field of view image, the image sharpening processing is used to enhance the local contrast of the image area where the suspended impurity feature is located, and the field of view scene enhancement processing is used to enhance the feature information of the field of view scene feature; in the second optimization mode, inputting the endoscopic field of view image into an end-to-end neural network based on deep learning, the end-to-end neural network is used to realize a mapping transformation from an input image to an output image, the mapping transformation includes feature blurring with the suspended impurity feature as the blurring target, and the video image is the output image of the end-to-end neural network.
[0025] In another embodiment of the present application, a video image generating device based on endoscopic imaging includes:
[0026] An image acquisition module, configured to acquire an endoscopic field of view image based on endoscopic imaging, wherein the endoscopic field of view image includes field of view scene features within the field of view of the endoscope;
[0027] An image optimization module is used to generate a video image for visual presentation based on image optimization of the endoscopic field of view image, wherein the image optimization is used to blur the suspended impurity features while preserving the fidelity of the field of view scene features when the endoscopic field of view image also includes suspended impurity features that cause blurring of the field of view scene features.
[0028] In some examples, optionally, the image optimization module includes: an image sharpening submodule for generating a sharpened image based on image sharpening processing of the endoscopic field of view image, wherein the image sharpening processing is used to enhance the local contrast of the image area where the suspended impurity feature is located; and a field of view scene enhancement submodule for generating the video image based on field of view scene enhancement processing of the sharpened image, wherein the field of view scene enhancement processing is used to enhance the feature information of the field of view scene feature.
[0029] In some examples, optionally, the image optimization module also includes: a level configuration submodule, which is used to determine, based on the acquired human-computer interaction instructions, a currently selected clarity level for the clarity processing among at least two preset clarity levels, and a currently selected enhancement level for the field of view scene enhancement processing among at least two preset enhancement levels.
[0030] In some examples, optionally, the image clearing submodule is specifically configured to: perform defogging processing on the endoscopic field image using the suspended impurity feature as the fog feature to obtain the cleared image.
[0031] In some examples, optionally, the image clearing submodule is specifically configured to: input the endoscopic field of view image into a pre-trained neural network, wherein the training samples of the neural network include atomized sample features having a characteristic morphology of the suspended impurity features, and the cleared image includes the output image of the neural network.
[0032] In some examples, optionally, the image clearing submodule is specifically configured to: perform optical enhancement processing on the endoscopic field of view image, defog the optically enhanced image obtained by the optical enhancement processing based on a preset defogging algorithm, and perform inverse processing of the optical enhancement processing on the defogged image obtained after the defogging processing to obtain the cleared image.
[0033] In some examples, optionally, the field of view scene enhancement submodule is specifically configured to: perform global detail enhancement processing on the sharpened image to obtain a global detail enhanced image; determine matching feature information of the field of view scene features in the global detail enhanced image; obtain the local detail enhanced image based on highlighting processing of the matching feature information in the global detail enhanced image; and generate the video image based on the sharpened image and the local detail enhanced image.
[0034] In some examples, optionally, the field of view scene enhancement submodule is specifically configured to: generate a saliency map with the field of view scene feature as the highlighted object based on the sharpened image, and / or obtain feature identification information of the field of view scene feature in the sharpened image; and determine the matching feature information of the field of view scene feature in the global detail enhanced image using the saliency map and / or the feature identification information as indication information.
[0035] In some examples, optionally, the field of view scene enhancement submodule is specifically configured to: determine the local contrast of the sharpened image at each scale based on a mean filtered image obtained by performing multi-scale mean filtering on the sharpened image; and obtain the saliency map based on weighted fusion of the local contrast of the sharpened image at each scale.
[0036] In some examples, optionally, the field of view scene enhancement submodule is specifically configured to: perform brightness equalization processing on the sharpened image to obtain a brightness-balanced image; and generate the video image based on image fusion of the brightness-balanced image and the local detail enhanced image.
[0037] In some examples, optionally, the field of view scene enhancement submodule is specifically configured to: perform illumination estimation on the cleared image; determine the high-brightness area and the low-brightness area in the cleared image divided by a brightness threshold based on the illumination estimation data; obtain the brightness-balanced image by performing brightness suppression on the high-brightness area of the cleared image and performing brightness compensation on the low-brightness area of the cleared image.
[0038] In some examples, optionally, the field of view scene enhancement submodule is specifically configured to generate illumination estimation data for the sharpened image based on a mean filtered image obtained by performing multi-scale mean filtering on the sharpened image.
[0039] In some examples, optionally, the image optimization module includes: a model calling submodule, used to input the endoscopic field of view image into an end-to-end neural network based on deep learning, wherein the end-to-end neural network is used to implement a mapping transformation from the input image to the output image, the mapping transformation includes feature blurring with the suspended impurity feature as the blurring target, and the video image is the output image of the end-to-end neural network.
[0040] In some examples, optionally, it also includes: a model training module for training the end-to-end neural network using an image sample set, wherein: the image sample set includes multiple groups of image sample pairs, the first image sample and the second image sample in each group of the image sample pairs include the same suspended impurity sample feature, the suspended impurity sample feature is presented in the first image sample as a visible state higher than a preset recognition threshold, and the suspended impurity sample feature is presented in the second image sample as a blurred state lower than the preset recognition threshold; the training goal of the end-to-end neural network is set as: the conversion loss from the first image sample to the second image sample in each group of the image samples is lower than the target loss.
[0041] In some examples, optionally, the image optimization module includes: a mode enabling submodule, for determining an optimization mode of the image optimization based on an acquired human-computer interaction instruction, wherein the optimization mode includes a first optimization mode and a second optimization mode that are selectively enabled; an image clearing submodule and a field of view scene enhancement submodule enabled in the first optimization mode, for sequentially performing image clearing processing and field of view scene enhancement processing on the endoscopic field of view image, wherein the image clearing processing is used to enhance the local contrast of the image area where the suspended impurity feature is located, and the field of view scene enhancement processing is used to enhance the feature information of the field of view scene feature; a model calling submodule enabled in the second optimization mode, for inputting the endoscopic field of view image into an end-to-end neural network based on deep learning, wherein the end-to-end neural network is used to realize a mapping conversion from an input image to an output image, wherein the mapping conversion includes feature blurring with the suspended impurity feature as the blurring target, and the video image is the output image of the end-to-end neural network.
[0042] In another embodiment of the present application, an endoscope system includes:
[0043] An endoscopic imaging assembly, comprising an endoscope and a camera, wherein the camera is configured to generate an image of an endoscopic field of view imaged through the endoscope;
[0044] The processor component is used to execute the video image generation method in the above embodiment.
[0045] In another embodiment of the present application, a non-transitory computer-readable storage medium stores instructions, which, when executed by a processor, cause the processor to execute the video image generation method of the aforementioned embodiment.
[0046] Based on the above embodiment, the endoscopic field of view image based on endoscopic imaging can be optimized. If the endoscopic field of view image includes, in addition to the field of view scene features within the field of view of the endoscope, also includes suspended impurity features that cause blurring of the field of view scene features, then the image optimization of the endoscopic field of view image can be carried out by differentially processing the field of view scene features and the suspended impurity features. The suspended impurity features can be blurred on the basis of preserving the fidelity of the field of view scene features, so that the clarity of the field of view scene features in the video image obtained after image optimization is higher than that in the endoscopic field of view image, and further, it helps to reduce the blurring interference of suspended impurities on the video presentation effect based on endoscopic imaging, so as to reduce the risk of accidental injury to human tissue or even surgical accidents during surgery. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The following drawings are only provided for schematic illustration and explanation of the present application and do not limit the scope of the present application:
[0048] Figure 1 Schematic diagram of the system architecture of an endoscope system in one embodiment of the present application;
[0049] Figure 2 For example Figure 1 Schematic diagram of the working principle of the endoscope system shown;
[0050] Figure 3 For example Figure 1 A schematic diagram of a first optimization mode of image optimization implemented by the endoscope system;
[0051] Figure 4 For example Figure 3 An example schematic diagram of level adjustment in the first optimization mode is shown;
[0052] Figure 5 For example Figure 3 A schematic diagram of a basic example of field of view scene enhancement processing in the first optimization mode is shown;
[0053] Figure 6 For example Figure 3 A schematic diagram of an optimization example of the field of view scene enhancement processing in the first optimization mode is shown;
[0054] Figure 7 For example Figure 1 A schematic diagram of a second optimization mode of image optimization implemented by the endoscope system;
[0055] Figure 8 For example Figure 7 A schematic diagram of a training example of an end-to-end network used in the second optimization mode is shown;
[0056] Figure 9 For example Figure 3 The first optimization mode shown and Figure 7 A schematic diagram of an example in which the second optimization mode coexists in an alternative manner;
[0057] Figure 10 This is a schematic diagram of an exemplary flow chart of a method for generating a video image based on endoscopic imaging in another embodiment of the present application;
[0058] Figure 11 For example Figure 10 The schematic diagram of the principle flow of the video generation method shown in FIG. 1 when the first optimization mode is adopted;
[0059] Figure 12 For example Figure 10 The schematic diagram of the principle flow of the video generation method when the second optimization mode is adopted is shown;
[0060] Figure 13 For example Figure 10 The video generation method shown is a schematic diagram of an extended process for compatibility with the first optimization mode and the second optimization mode;
[0061] Figure 14 This is a schematic diagram of an exemplary structure of a video image generating device based on endoscopic imaging in another embodiment of the present application;
[0062] Figure 15 For example Figure 14 The optimized structural diagram of the video image generating device is shown. DETAILED DESCRIPTION
[0063] In order to make the objectives, technical solutions and advantages of this application more clear, the application is further described in detail below with reference to the accompanying drawings and examples.
[0064] Figure 1 Schematic diagram of the system architecture of an endoscope system in one embodiment of the present application. Figure 2 For example Figure 1 The working principle of the endoscope system is shown in the figure. Figure 1 and Figure 2In an embodiment of the present application, an endoscope system may include an endoscopic imaging component 10 and a host device 30 .
[0065] The endoscopic imaging component 10 includes an endoscope 11 and a camera 12. The endoscope 11 can enter the body through an operating channel, for example, from a human cavity such as a ureter into an organ (such as a bladder) of the human body, and the camera 12 can be used to generate an endoscopic field of view image imaged by the endoscope 11. For example, the camera 12 may include a photosensitive element such as a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS), and the endoscopic imaging component 10 may also include a light source such as a cold light source for supplementing the endoscopic field of view (not shown in the drawings), and an operable element such as a button (not shown in the drawings), which can be used to start or stop the imaging process of the endoscopic field of view image in response to the operator's external operation of the operable element. In addition, Figure 1 The endoscope 11 is illustrated as an electronic scope as an example, but it can be understood that such an illustration does not mean that unnecessary restrictions are placed on the type of the endoscope 11 in the embodiment of the present application.
[0066] The host device 30 includes a processor component 310 , a storage medium 320 , and a human-computer interaction component 330 .
[0067] The processor component 310 can be used to collect the endoscopic field of view image formed by the camera 12 through the endoscope 11, and realize image processing, video coding and intelligent processing (such as feature classification based on deep learning, feature recognition and coding indication in video coding) of the endoscopic field of view image. Therefore, Figure 2 The processor component 310 is shown as a plurality of functional units divided according to their functions: an image acquisition unit 311, an image processing unit 312, an intelligent processing unit 313, a video encoding unit 314, and a host control unit 315. In addition, the video images generated by the processor component 310 through image processing can be transmitted to the display device 40 for visual presentation after video encoding processing.
[0068] The processor component 310 that implements multiple functions may include electrical control elements such as a microcontroller unit (MCU), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA). It may also include data processing elements with data processing functions such as a central processing unit (CPU), an image signal processor (ISP), and a graphics processing unit (GPU).
[0069] The storage medium 320 may include non-volatile storage media such as Flash, and volatile storage media such as Random Access Memory (RAM) and its derivative storage media. The non-volatile storage medium may store instructions for the processor component 310 to implement the above functions, and the volatile storage medium may be configured as a memory for temporarily storing intermediate data.
[0070] The human-computer interaction component 330 may include at least one operable component such as a keyboard, a switch, a gear lever, a touch screen, etc. that can be used for contact-based human-computer interaction between the host device 30 and the operator (e.g., the operator or his assistant), and / or, the human-computer interaction component 330 may also include an interface component such as a Bluetooth interface that supports contactless human-computer interaction between the host device 30 and the operator (e.g., the operator or his assistant) based on a wireless communication connection mediated by a terminal device. Regardless of the components included in the human-computer interaction component 330, it can be used to obtain human-computer interaction instructions input by the operator, and the above-mentioned functions implemented by the processor component 310 can refer to the human-computer interaction instructions obtained by the human-computer interaction component 330, for example, Figure 2 The host control unit 315 shown in FIG can control other functional units based on human-computer interaction instructions.
[0071] In the embodiment of the present application, in order to reduce the blurring interference of suspended impurities on the video presentation effect based on endoscopic imaging, the image processing implemented by the processor component 310 can be improved, that is, Figure 2 The image processing unit 312 shown in FIG. 3 implements an improvement.
[0072] In this case, the image processing of the processor component 310 may specifically include:
[0073] Acquiring an endoscopic field of view image based on endoscopic imaging, wherein the endoscopic field of view image includes at least field of view scene features within the field of view of the endoscope (including but not limited to human target features such as human tissue and surgical auxiliary features such as surgical tools);
[0074] Based on image optimization of the endoscopic field of view image, a video image for visual presentation is generated, wherein the image optimization of the endoscopic field of view image is used to blur the suspended impurity features while maintaining the fidelity of the field of view scene features when the endoscopic field of view image also includes suspended impurity features that cause blurring of field of view scene features (such as at least one of debris or droplets in body fluids, water vapor in the body, and debris produced after hard growths are crushed).
[0075] Based on the above-mentioned image processing method of the processor component 310, the image of the endoscopic field of view image based on endoscopic imaging can be optimized. If the endoscopic field of view image includes, in addition to the field of view scene features within the field of view of the endoscope, also includes suspended impurity features that cause blurring of the field of view scene features, then the image optimization of the endoscopic field of view image can be carried out by differentially processing the field of view scene features and the suspended impurity features. On the basis of preserving the fidelity of the field of view scene features, the clarity of the field of view scene features in the video image obtained after image optimization is higher than that in the endoscopic field of view image. Furthermore, it helps to reduce the blurring interference of suspended impurities on the video presentation effect based on endoscopic imaging, so as to reduce the risk of accidental injury to human tissue or even surgical accidents during surgery.
[0076] In an embodiment of the present application, two optimization modes can be provided for image optimization of the endoscope field of view image, namely, a first optimization mode and a second optimization mode, and the processor component 310 can adopt either the first optimization mode or the second optimization mode when performing image optimization on the endoscope field of view image. That is, the first optimization mode and the second optimization mode can be selectively enabled by the processor component 310, wherein the "selective enabling" mentioned here can include two situations:
[0077] Case 1: Image processing function of the processor component 310 (e.g. Figure 2 The functional configuration of the image processing unit 312 in the image processing unit 312 only supports one of the first optimization mode and the second optimization mode, that is, the "selective activation" of the first optimization mode and the second optimization mode can be achieved through functional configuration;
[0078] Case 2: Image processing function of the processor component 310 (e.g. Figure 2The functional configuration of the image processing unit 312 in the image processing unit 312 may support the first optimization mode and the second optimization mode, and the “selection and activation” of the first optimization mode and the second optimization mode may be controlled by a human-computer interaction instruction.
[0079] Figure 3 For example Figure 1 Schematic diagram of the first optimization mode of image optimization implemented by the endoscope system shown. Figure 3 In an embodiment of the present application, the processor component 310, in the first optimization mode, may be configured to optimize the endoscope field of view image 510 in the following manner:
[0080] Based on the image sharpening process S_clr of the endoscope field image 510, a sharpened image 520 is generated;
[0081] Generate a video image 550 based on the visual field scene enhancement processing S_enh on the clear image 520;
[0082] Among them, the image clearing processing S_clr is used to enhance the local contrast of the image area where the suspended impurity features in the endoscopic field of view image 510 are located, so that the cleared image 520 has higher image transparency than the endoscopic field of view image 510; and the field of view scene enhancement processing S_enh is used to enhance the feature information of the field of view scene features in the cleared image 520.
[0083] For Figure 3 The first optimization mode shown, since the concentration of suspended impurity characteristics may vary due to pathological differences in the human body, and / or fluctuate as the surgical process progresses, the first optimization mode can support the clarity level adjustment of the image clarity processing S_clr and / or the enhancement level adjustment of the field of view scene enhancement processing S_enh, and the adjustment can be controlled by human-computer interaction instructions.
[0084] In this case, the first optimization mode can be regarded as a manual mode supporting human-computer interaction, and the processor component 310 can be further used to control the image optimization of the endoscope field of view image 510 based on the acquired human-computer interaction instruction, that is, to determine the currently selected clarity level for the clarity processing S_clr among the at least two preset clarity levels, and the currently selected enhancement level for the field of view scene enhancement processing S_enh among the at least two preset enhancement levels. The human-computer interaction instruction mentioned here may include:
[0085] a first gear selection instruction for determining a currently selected sharpening level for the sharpening process S_clr among at least two preset sharpening levels;
[0086] and / or,
[0087] The second gear selection instruction is used to determine a currently selected enhancement level for the field of view scene enhancement processing S_enh among at least two preset enhancement levels.
[0088] Accordingly, the processor component 310 can be specifically configured to: generate a sharpened image based on the image sharpening process S_clr of the currently selected sharpening level performed on the endoscope field of view image, and generate a video image based on the field of view scene enhancement process S_enh of the currently selected enhancement level performed on the sharpened image.
[0089] Figure 4 For example Figure 3 Schematic diagram of an example of level adjustment in the first optimization mode shown in FIG. Figure 4 Taking the example that the clarity processing S_clr has three available clarity levels L1, L2, and L3, and the field of view scene enhancement processing S_enh has three available enhancement levels S1, S2, and S3, through the combined constraints of the first gear selection instruction and the second gear selection instruction, the image optimization in the first optimization mode can have a total of nine optimization levels D1-D9.
[0090] Thus, the operator can make the presentation clarity of the visual field scene features of the video image 550 reach a recognizable level by inputting the first gear selection instruction and / or the second gear selection instruction based on the presentation effect of the video image 550 on the visual field scene features. Moreover, on the basis that the presentation clarity of the visual field scene features of the video image 550 reaches a recognizable level, the operator can further make the presentation effect of the visual field scene features of the video image match the operator's human eye observation comfort zone through the first gear selection instruction and / or the second gear selection instruction.
[0091] The embodiments of the present application do not attempt to impose unnecessary limitations on the sharpening process S_clr in the first optimization mode. For example, in the embodiments of the present application, the sharpening process S_clr in the first optimization mode may include dehazing, or other processing methods that can be used to improve image transparency by increasing the local contrast of the image area where the suspended impurity features are located.
[0092] If the sharpening process S_clr in the first optimization mode includes a dehazing process, the processor component 310 may be specifically configured to perform a dehazing process on the endoscope field of view image 510, using suspended impurity features as the hazing feature, to obtain a sharpened image 520. Furthermore, if at least two sharpening levels are required for the sharpening process S_clr, the sharpening level may be determined by the dehazing level of the dehazing process.
[0093] Among them, dehazing processing can be implemented by a neural network, or it can also be implemented using image processing methods.
[0094] For example, the processor component 310 may input the endoscope field of view image 510 into a pre-trained neural network, wherein the training samples of the neural network include haze sample features having characteristic morphologies characteristic of suspended impurities, and the cleared image 520 may include the output image of the neural network. The neural network used for dehazing may be a Retinex (retina + cortex) model based on trichromatic theory and color constancy, an atmospheric scattering model based on atmospheric scattering principles, or any other convolutional neural network.
[0095] For another example, the processor component 310 may also perform defogging on the endoscope field of view image 510 using any defogging algorithm such as homomorphic filtering, wavelet transform, histogram equalization, or a dark channel prior defogging algorithm.
[0096] Assuming that the processor component 310 uses a defogging algorithm (such as a dark channel prior defogging algorithm) to perform defogging on the endoscope field of view image 510, the defogging process may include:
[0097] Performing optical enhancement processing (such as inverse gamma transformation) on the endoscope field of view image 510 , and before performing the optical enhancement processing, performing image denoising processing such as Gaussian filtering on the endoscope field of view image 510 ;
[0098] Defogging the light-enhanced image obtained by the light-enhanced processing based on a preset defogging algorithm (such as a dark channel prior defogging algorithm);
[0099] The defogging image obtained after the defogging process is subjected to an inverse process of the optical enhancement process (eg, gamma transformation) to obtain a sharpened image 520 having higher image transparency than the endoscopic field image 510 .
[0100] The embodiments of the present application do not attempt to impose unnecessary limitations on the field of view scene enhancement process S_enh in the first optimization mode. For example, in the embodiments of the present application, the field of view scene enhancement process S_enh in the first optimization mode may include detail enhancement processing, and the detail enhancement processing may be implemented by combining global detail enhancement with enhanced detail screening.
[0101] Figure 5 For example Figure 3 Schematic diagram of a basic example of field of view scene enhancement processing in the first optimization mode shown. Figure 5The processor component 310 may be specifically configured to implement the field of view scene enhancement processing S_enh on the sharpened image 520 in the following manner:
[0102] Performing a global detail enhancement process S_enh_dtl (e.g., Gaussian high-pass filtering) on the sharpened image 520 to obtain a global detail enhanced image 531. If at least two enhancement levels are required for the field of view scene enhancement process S_enh, the enhancement levels may be determined by configuration parameters of the global detail enhancement process (e.g., Gaussian high-pass filtering).
[0103] Determining matching feature information of the field of view scene features in the global detail enhanced image 531;
[0104] Based on the highlighting processing S_enh_flt of the matching feature information in the global detail enhanced image 531, for example, weakening or deleting other feature information except the matching feature in the global detail enhanced image 531, a local detail enhanced image 532 that highlights the matching feature information of the field of view scene features is obtained;
[0105] The video image 550 is generated based on the sharpened image 520 and the local detail enhanced image 532 . For example, the video image 550 may be generated based on image fusion of the sharpened image 520 and the local detail enhanced image 532 .
[0106] In order to improve the accuracy of determining the above matching feature information, the processor component 310 may be further configured to:
[0107] A saliency map 533 is generated based on the sharpened image 520, with the visual field scene features as the highlighting objects. For example, the local contrast of the sharpened image 520 at each scale is determined based on a mean filtered image obtained by performing multi-scale mean filtering on the sharpened image 520. The saliency map 533 is then generated based on a weighted fusion of the local contrast of the sharpened image 520 at each scale.
[0108] and / or,
[0109] Obtain feature identification information 534 of the visual field scene features in the sharpened image 520, for example, Figure 2 The intelligent processing unit 315 shown in FIG can perform feature recognition on the visual field scene features including human tissue, surgical tools, etc., and the feature recognition information 534 generated by the intelligent processing unit 315 can include image position information such as the visual field scene features in the cleared image 520;
[0110] In this case, the processor component 310 can use the saliency map 533 and / or feature recognition information 534 as indication information to determine the matching feature information of the field of view scene features in the global detail enhanced image 531, and perform highlighting processing S_enh_flt on the matching feature information in the global detail enhanced image 531.
[0111] In an embodiment of the present application, in order to further improve the image quality of the video image 550 , brightness equalization may be further introduced into the field of view scene enhancement processing S_enh of the sharpened image 520 .
[0112] Figure 6 For example Figure 3 Schematic diagram of an optimization example of the field of view scene enhancement processing in the first optimization mode shown. Figure 6 The processor component 310 is specifically configured to further perform brightness equalization processing S_enh_bal on the clear image 520 during the field of view scene enhancement processing S_enh on the clear image 520 to obtain a brightness equalized image 535.
[0113] For example, the processor component 310 may be specifically configured to generate the brightness-balanced image 535 by:
[0114] Performing illumination estimation on the sharpened image 520 , for example, generating illumination estimation data for the sharpened image 520 based on the mean filtered image obtained by performing multi-scale mean filtering on the sharpened image 520 as mentioned above;
[0115] Based on illumination estimation data obtained by performing illumination estimation on the cleared image 520, determining a high-brightness area and a low-brightness area in the cleared image 520 divided by a brightness threshold. For example, the brightness threshold may be a dynamic threshold determined based on an average brightness Lm of the cleared image 520.
[0116] A brightness balanced image 535 is obtained by performing brightness suppression on the high brightness area of the cleared image 520 (e.g., a first nonlinear mapping for suppressing the highlights) and performing brightness compensation on the low brightness area of the cleared image 520 (e.g., a second nonlinear mapping for compensating the low brightness).
[0117] In this case, in the process of generating the video image 550 based on the sharpened image 520 and the local detail enhanced image 532 , the processor component 310 may replace the sharpened image 520 with the brightness-balanced image 535 for image fusion with the local detail enhanced image 532 , that is:
[0118] Based on the image fusion of the brightness-balanced image 535 and the local detail-enhanced image 532 , a video image 550 is generated.
[0119] Since the brightness balanced image 535 has a better brightness balanced characteristic than the sharpened image 520, in the process of generating the video image 550, the brightness balanced image 535 is used instead of the sharpened image 520 to be fused with the local detail enhanced image 532, which can suppress the brightness imbalance caused by the field of view scene enhancement processing S_enh (for example, caused by the global detail enhancement processing S_enh_dtl). Figure 6 The video image 550 generated by the optimization example shown can be generated by Figure 5 Basic example shown.
[0120] The above is an example description of the first optimization mode of image optimization in the embodiment of the present application. The second optimization mode of image optimization will be further introduced below.
[0121] Figure 7 For example Figure 1 Schematic diagram of the second optimization mode of image optimization implemented by the endoscope system shown. Figure 7 In an embodiment of the present application, the processor component 310 in the second optimization mode may be configured to optimize the endoscope field of view image 510 in the following manner:
[0122] Input the endoscope field of view image 510 into the end-to-end neural network 600 based on deep learning;
[0123] In which, the end-to-end neural network 600 is used to implement a mapping conversion from an input image to an output image. The mapping conversion implemented by the end-to-end neural network 600 includes feature blurring with suspended impurity features as the blurring target, that is, the mapping conversion is configured as a direct mapping conversion between two types of images, from blurred images to deblurred images, and the video image 550 is the output image of the end-to-end neural network 600.
[0124] The end-to-end neural network 600 may be implanted into the endoscope system after training is completed, or the end-to-end neural network 600 may be trained with the assistance of the processor component 310, that is, the processor component 310 may be further configured to train the end-to-end neural network 600 using an image sample set. Regardless of the execution entity and completion time of the training process of the end-to-end neural network 600, the image sample set used in the training process may include multiple sets of image sample pairs, wherein:
[0125] The first image sample and the second image sample in each set of image sample pairs include the same suspended impurity sample features;
[0126] The suspended impurity sample feature is presented in the first image sample as a visible state higher than a preset recognition threshold, that is, the first image sample belongs to the first image domain X representing the blurred image;
[0127] The suspended impurity sample feature is presented in the second image sample as a blurred state below a preset recognition threshold, that is, the second image sample belongs to the second image domain Y representing the deblurred image.
[0128] In this case, the training objective of the end-to-end neural network 600 is set as: the conversion loss of the mapping conversion from the first image sample (i.e., the first image domain X) to the second image sample (i.e., the second image domain Y) in each group of image samples is lower than the target loss, so that when the input image belongs to the first image domain X representing the blurred image, the mapping conversion can be adaptively initiated.
[0129] That is, the end-to-end neural network 600 has the ability to discriminate between blurred images and deblurred images, and can achieve a direct mapping conversion between the two types of images, from blurred images to deblurred images. Therefore:
[0130] When the endoscopic field of view image 510 does not include suspended impurity features, that is, when the endoscopic field of view image 510 belongs to the second image domain Y, the output image generated by the end-to-end neural network 600 also belongs to the second image domain Y and can be substantially the same as the input image. That is, the video image 550 at this time can be the same as the endoscopic field of view image 510, so as to avoid local distortion of the endoscopic field of view image 510 that does not include suspended impurity features due to additional processing aimed at blurring the suspended impurity features. That is, the endoscopic field of view image 510 that does not include suspended impurity features implements global fidelity.
[0131] When the endoscopic field of view image 510 includes suspended impurity features, that is, when the endoscopic field of view image 510 belongs to the first image domain X, the end-to-end neural network 600 can generate an input image belonging to the first image domain X, that is, a deblurred image after the above-mentioned mapping transformation, so as to blur the suspended impurity features while preserving the fidelity of the field of view scene features.
[0132] Compared with the first optimization mode, the second optimization mode can automatically identify whether the endoscopic field of view image 510 is a blurred image, and can adaptively initiate the above-mentioned mapping conversion without manual intervention when the endoscopic field of view image 510 is a blurred image. Therefore, the second optimization mode can be regarded as a time-adaptive optimization mode.
[0133] Figure 8 For example Figure 7 The diagram shows a training example of the end-to-end network used in the second optimization mode. Figure 8Taking the end-to-end neural network 600 as a CycleGAN (Cycle Generative Adversarial Networks) as an example, during the training process of the CycleGAN using an image sample set, the CycleGAN cycle implements a bidirectional conversion between the first image sample and the second image sample in each set of image samples. Figure 8 In the figure, the domain identifier "X" of the first image domain X represents the first image sample in each group of image samples, and the domain identifier "Y" of the second image domain Y represents the second image sample in each group of image samples. Therefore, the training process of CycleGAN can also be regarded as a training process of using the image sample set to cyclically realize the bidirectional image domain conversion between the first image domain X to which the first image sample belongs and the second image domain Y to which the second image sample belongs.
[0134] Specifically, if Figure 8 As shown, CycleGAN can specifically include two independent cyclic links. The first cyclic link includes a first generator G and a second generator F arranged in sequence, and the arrangement direction of the first generator G and the second generator F in the second cyclic link is opposite to that of the first cyclic link. The second cyclic link is provided with a first discriminator Dx at the connection node between the second generator F and the first generator G, and the first cyclic link is provided with a second discriminator Dy at the connection node between the first generator G and the second generator F.
[0135] Among them, the first generator G is the first image generator G used to realize the unidirectional conversion from the first image domain X to the second image domain Y; the second generator F is the second image generator F used to realize the unidirectional conversion from the second image domain Y to the first image domain X; the first discriminator Dx is used to identify whether the image generated by the second generator F of the second cyclic link comes from the second image domain Y or belongs to the first image domain X; the second discriminator Dy is used to identify whether the image generated by the first generator G of the first cyclic link comes from the first image domain X or belongs to the second image domain Y.
[0136] The cycle consistency loss and cycle perceptual consistency loss of CycleGAN can be determined by the identification results of the first discriminator Dx and the second discriminator Dy. When the weighted value of the cycle consistency loss and cycle perceptual consistency loss of CycleGAN reaches a preset threshold, it can be considered that the aforementioned training goal has been achieved and the training process is terminated. In addition, when using the end-to-end neural network 600 to optimize the endoscope field of view image 510:
[0137] If the endoscopic field of view image 510 does not include the suspended impurity feature, then the endoscopic field of view image 510 belonging to the second image domain Y is used as the input image and is transformed twice by the second generator F and the first generator G in the second loop chain to obtain an output image (i.e., video image 550) that belongs to the second image domain Y and is substantially identical to the input image (i.e., the endoscopic field of view image 510 that does not include the suspended impurity feature).
[0138] If the endoscopic field of view image 510 includes suspended impurity features, then the endoscopic field of view image 510 belonging to the first image domain X as an input image can be converted by the first generator G in the first loop link into an output image belonging to the second image domain Y (i.e., video image 550).
[0139] Figure 9 For example Figure 3 The first optimization mode shown and Figure 7 The diagram shows an example of the second optimization mode coexisting in an alternative manner. Figure 9 , the processor component 310 may be specifically configured to:
[0140] Based on the acquired human-computer interaction instruction, an optimization mode for optimizing the image of the endoscope field of view image 510 is determined, wherein the human-computer interaction instruction includes a mode enable instruction Sig_hc_mod for alternatively enabling a first optimization mode and a second optimization mode, and the image optimization of the endoscope field of view image 510 may include:
[0141] In the first optimization mode, the image sharpening process S_clr and the visual field scene enhancement process S_enh described above are sequentially performed on the endoscope visual field image 510;
[0142] In the second optimization mode, the endoscopic field of view image 510 is input into the end-to-end neural network 600 described above.
[0143] in addition, Figure 9 Also shown are a first gear selection instruction Sig_hc_cl for determining the sharpness level of the image sharpening process S_clr and a second gear selection instruction Sig_hc_el for determining the enhancement level of the visual field scene enhancement process S_enh.
[0144] Figure 10 This is an exemplary flow chart of a method for generating a video image based on endoscopic imaging in another embodiment of the present application. Figure 10 In an embodiment of the present application, a method for generating a video image based on endoscopic imaging may include:
[0145] S1010: Acquire an endoscopic field of view image based on endoscopic imaging, wherein the endoscopic field of view image includes field of view scene features within the field of view of the endoscope;
[0146] S1030, based on the image optimization of the endoscopic field of view image, a video image for visual presentation is generated, wherein the image optimization of the endoscopic field of view image is used to blur the suspended impurity features while preserving the fidelity of the field of view scene features when the endoscopic field of view image also includes suspended impurity features that cause blurring of field of view scene features.
[0147] Based on the above-mentioned video generation method, the endoscopic field of view image based on endoscopic imaging can be optimized. If the endoscopic field of view image includes, in addition to the field of view scene features within the field of view of the endoscope, also includes suspended impurity features that cause blurring of the field of view scene features, then the image optimization of the endoscopic field of view image can be carried out by differentially processing the field of view scene features and the suspended impurity features. The suspended impurity features can be blurred on the basis of preserving the fidelity of the field of view scene features, so that the clarity of the field of view scene features in the video image obtained after image optimization is higher than that in the endoscopic field of view image. Furthermore, it helps to reduce the blurring interference of suspended impurities on the video presentation effect based on endoscopic imaging, so as to reduce the risk of accidental injury to human tissue or even surgical accidents during surgery.
[0148] As described above, the embodiments of the present application can provide selectable first optimization mode and second optimization mode for image optimization of the endoscopic field of view image, that is, S1030 can implement image optimization of the endoscopic field of view image in the first optimization mode or the second optimization mode.
[0149] Figure 11 For example Figure 10 The schematic diagram of the principle flow of the video generation method when the first optimization mode is adopted is shown in FIG. Figure 11 As shown, if the first optimization mode is adopted for the image optimization of the endoscope field of view image, the video generation method in the embodiment of the present application may include:
[0150] S1010: Acquire an endoscopic field of view image based on endoscopic imaging, wherein the endoscopic field of view image includes field of view scene features within the field of view of the endoscope;
[0151] S1031: generating a sharpened image based on image sharpening processing of the endoscope field of view image, wherein the image sharpening processing is used to enhance the local contrast of the image region where the suspended impurity feature is located;
[0152] S1033: Generate a video image based on a field of view scene enhancement process performed on the sharpened image, wherein the field of view scene enhancement process is used to enhance feature information of field of view scene features.
[0153] In order to improve the image optimization in the first optimization mode so that it can better adapt to the different concentrations of suspended impurity features, so as to ensure that the clarity of the video image presentation reaches a recognizable level and even further matches the operator's human eye observation comfort zone, as a preferred extension method, before S1031 and S1033, the video generation method may also include: based on the acquired human-computer interaction instruction, determining the currently selected clarity level for the clarity processing from at least two pre-set clarity levels, and the currently selected enhancement level for the field of view scene enhancement processing from at least two pre-set enhancement levels. The human-computer interaction instructions for determining the clarity level and the enhancement level can be found in the description of the first gear selection instruction and the second gear selection instruction above, and will not be repeated here.
[0154] In this case, S1031 generates a sharpened image based on the image sharpening process of the currently selected sharpening level performed on the endoscope visual field image, and S1033 generates a video image based on the visual field scene enhancement process of the currently selected enhancement level performed on the sharpened image.
[0155] In the first optimization mode, the clarity processing of S1031 may include defogging processing, or other processing methods that can be used to improve the transparency of the image by enhancing the local contrast of the image area where the suspended impurity features are located. For example, S1031 may specifically include: performing defogging processing on the endoscopic field of view image with the suspended impurity features as the fogging features to obtain a clear image, wherein the specific processing process of the defogging processing performed by S1031 can be found in the relevant description in the previous text and will not be repeated here.
[0156] Moreover, in the first optimization mode, the field of view scene enhancement processing of S1033 may include detail enhancement processing, and the detail enhancement processing may be implemented by combining global detail enhancement with enhanced detail screening. For example, S1033 may specifically include:
[0157] Performing global detail enhancement processing on the sharpened image to obtain a global detail enhanced image;
[0158] Determining matching feature information of the visual field scene feature in the global detail enhanced image, for example, using the saliency map and / or feature recognition information described above as indication information, determining the matching feature information of the visual field scene feature in the global detail enhanced image, wherein the method for obtaining the saliency map and feature recognition information can be referred to the description above and will not be repeated here;
[0159] Based on the highlighting processing of the matching feature information in the global detail enhanced image, a local detail enhanced image is obtained;
[0160] A video image is generated based on the sharpened image and the local detail enhanced image. For example, the video image may be generated based on image fusion of the sharpened image and the local detail enhanced image.
[0161] In an embodiment of the present application, in order to further improve the image quality of the video image, S1033 can further introduce brightness equalization into the field of view scene enhancement processing of the cleared image, that is, S1033 can further include: performing brightness equalization processing on the cleared image to obtain a brightness equalized image.
[0162] For example, the process of obtaining a brightness-balanced image through equalization processing in S1033 may specifically include:
[0163] Performing illumination estimation on the sharpened image to generate illumination estimation data for the sharpened image. The specific implementation of illumination estimation can be found in the previous description and will not be repeated here.
[0164] Based on illumination estimation data obtained by performing illumination estimation on the sharpened image, determining a high-brightness area and a low-brightness area in the sharpened image divided by a brightness threshold, for example, the brightness threshold may be a dynamic threshold determined based on the average brightness of the sharpened image;
[0165] A brightness balanced image is obtained by performing brightness suppression on high brightness areas of the sharpened image (eg, a first nonlinear mapping for suppressing highlights) and performing brightness compensation on low brightness areas of the sharpened image (eg, a second nonlinear mapping for compensating low brightness).
[0166] Figure 12 For example Figure 10 The schematic diagram of the principle flow of the video generation method when the second optimization mode is adopted is shown. If the second optimization mode is adopted for the image optimization of the endoscope field of view image, the video generation method in the embodiment of the present application may include:
[0167] S1010: Acquire an endoscopic field of view image based on endoscopic imaging, wherein the endoscopic field of view image includes field of view scene features within the field of view of the endoscope.
[0168] S1035: Inputting the endoscope field image into an end-to-end neural network based on deep learning, wherein the end-to-end neural network is used to achieve a mapping transformation from the input image to the output image, the mapping transformation includes feature blurring with suspended impurity features as the blurring target, and the video image is the output image of the end-to-end neural network.
[0169] Among them, the training method and instance structure of the end-to-end neural network used by S1035 can be found in the previous description and will not be repeated here.
[0170] Figure 13 For example Figure 10 The video generation method shown is a schematic diagram of an extended process for compatibility with the first optimization mode and the second optimization mode. Figure 13 In an embodiment of the present application, if the switching between the first optimization mode and the second optimization mode based on human-computer interaction is supported, the video generation method based on endoscopic imaging may include:
[0171] S1010: Acquire an endoscopic field of view image based on endoscopic imaging, wherein the endoscopic field of view image includes field of view scene features within the field of view of the endoscope;
[0172] S1020: Detecting a current optimization mode of image optimization, wherein the optimization mode of image optimization is determined based on the acquired human-computer interaction instruction, and the optimization mode of image optimization includes a first optimization mode and a second optimization mode that are selectively enabled, and:
[0173] If the human-computer interaction instruction indicates the first optimization mode, then S1031 and S1033 are executed sequentially to sequentially perform image sharpening processing and field of view scene enhancement processing on the endoscope field of view image in the first optimization mode.
[0174] If the human-computer interaction instruction indicates the second optimization mode, then jump to S1035, that is, input the endoscope field of view image into the end-to-end neural network based on deep learning.
[0175] Figure 14 This is an exemplary structural diagram of a video image generation device based on endoscopic imaging in another embodiment of the present application. Figure 14 In an embodiment of the present application, a video image generating device based on endoscopic imaging may include:
[0176] An image acquisition module 1410 is configured to acquire an endoscopic field of view image based on endoscopic imaging, wherein the endoscopic field of view image includes field of view scene features within the field of view of the endoscope;
[0177] The image optimization module 1430 is used to generate a video image for visual presentation based on the image optimization of the endoscopic field of view image, wherein the image optimization of the endoscopic field of view image is used to blur the suspended impurity features while preserving the fidelity of the field of view scene features when the endoscopic field of view image also includes suspended impurity features that cause blurring of the field of view scene features.
[0178] Based on the above-mentioned video generation device, the image of the endoscopic field of view image based on endoscopic imaging can be optimized. If the endoscopic field of view image includes, in addition to the field of view scene features within the field of view of the endoscope, also includes suspended impurity features that cause blurring of the field of view scene features, then the image optimization of the endoscopic field of view image can be carried out by differentially processing the field of view scene features and the suspended impurity features. The suspended impurity features can be blurred on the basis of preserving the fidelity of the field of view scene features, so that the clarity of the field of view scene features in the video image obtained after image optimization is higher than that in the endoscopic field of view image, and thus it helps to reduce the blurring interference of suspended impurities on the video presentation effect based on endoscopic imaging, so as to reduce the risk of accidental injury to human tissue or even surgical accidents during surgery.
[0179] As described above, the embodiments of the present application can provide a selectable first optimization mode and a second optimization mode for image optimization of the endoscopic field of view image, that is, the image optimization module 1430 can implement image optimization of the endoscopic field of view image in the first optimization mode or the second optimization mode.
[0180] Figure 15 For example Figure 14 Schematic diagram of the optimized structure of the video image generating device shown in FIG. Figure 15 In an embodiment of the present application, the image optimization module 1430 may include: an image sharpening submodule 1431 and a field of view scene enhancement submodule 1433 , and / or a model calling submodule 1435 .
[0181] The image clearing submodule 1431 can be used to: generate a cleared image based on image clearing processing of the endoscopic field of view image, wherein the image clearing processing is used to enhance the local contrast of the image area where the suspended impurity features are located, and the specific implementation method of the image clearing processing can be found in the previous description and will not be repeated here.
[0182] The field of view scene enhancement submodule 1433 can be used to: generate a video image based on the field of view scene enhancement processing of the clear image, wherein the field of view scene enhancement processing is used to enhance the feature information of the field of view scene features, and the specific implementation method of the field of view scene enhancement processing can be found in the previous description and will not be repeated here.
[0183] The model calling submodule 1435 can be used to: input the endoscope field of view image into an end-to-end neural network based on deep learning, wherein the end-to-end neural network is used to realize the mapping transformation from the input image to the output image, the mapping transformation includes feature blurring with suspended impurity features as the blurring target, and the video image is the output image of the end-to-end neural network. In addition, the training method and instance structure of the end-to-end neural network used by the model calling submodule 1435 can be found in the previous description and will not be repeated here.
[0184] exist Figure 15 In the figure, the image optimization module 1430 includes an image clearing submodule 1431, a field of view scene enhancement submodule 1433, and a model calling submodule 1435 as an example for graphical expression. In this case, the image optimization module 1430 may also include a mode enabling submodule 1437 for determining the optimization mode when the image optimization module 1430 optimizes the endoscope field of view image based on the acquired human-computer interaction instructions, wherein the image clearing submodule 1431 and the field of view scene enhancement submodule 1433 are selectively enabled in the first optimization mode, and the model calling submodule 1435 is selectively enabled in the second optimization mode.
[0185] It is understandable that if the image optimization module 1430 only includes the image sharpening submodule 1431 and the field of view scene enhancement submodule 1433 , or only includes the model calling submodule 1435 , then the image optimization module 1430 may not include the mode enabling submodule 1437 .
[0186] In the case where the image optimization module 1430 includes only the image sharpening submodule 1431 and the visual field scene enhancement submodule 1433, or includes the image sharpening submodule 1431, the visual field scene enhancement submodule 1433, and the model calling submodule 1435, in order to improve the image optimization in the first optimization mode so as to better adapt to different concentrations of suspended impurity features, thereby ensuring that the video image presentation clarity reaches a recognizable level and even further matches the operator's human eye observation comfort zone, as a preferred extension method, the image optimization module 1430 may further include:
[0187] Level configuration submodule 1439 is configured to determine, based on the acquired human-computer interaction instruction, a currently selected clarity level for the clarity processing from among at least two preset clarity levels, and a currently selected enhancement level for the visual field scene enhancement processing from among at least two preset enhancement levels. The human-computer interaction instructions for determining the clarity level and enhancement level can be found in the description of the first gear selection instruction and the second gear selection instruction above, and are not further described here.
[0188] In addition, when the image optimization module 1430 only includes the model calling submodule 1435, or includes the image clearing submodule 1431, the field of view scene enhancement submodule 1433, and the model calling submodule 1435 at the same time, if the training of the end-to-end neural network is undertaken by the video generation device, then the video generation device may also include a model training module not shown in the figure, which is used to train the end-to-end neural network using an image sample set. The specific training method can be referred to the previous text and will not be repeated here.
[0189] In another embodiment of the present application, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores instructions. When the instructions are executed by a processor, the processor executes the video image generation method described above.
[0190] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for generating a video image based on endoscopic imaging, characterized in that: include: Acquiring an endoscopic field of view image based on endoscopic imaging, wherein the endoscopic field of view image includes field of view scene features within the field of view of the endoscope; generating a video image for visual presentation based on image optimization of the endoscopic field of view image, wherein the image optimization is used to blur the suspended impurity features while preserving the fidelity of the field of view scene features when the endoscopic field of view image also includes suspended impurity features that cause blurring of the field of view scene features; The step of generating a video image for visual presentation based on image optimization of the endoscopic field of view image comprises: generating a sharpened image based on image sharpening processing of the endoscope field of view image; generating the video image based on the visual field scene enhancement processing of the sharpened image; The image sharpening process is used to enhance the local contrast of the image region where the suspended impurity feature is located, the visual field scene enhancement process is used to enhance the feature information of the visual field scene feature, and the video image is generated based on the visual field scene enhancement process of the sharpened image, including: performing global detail enhancement processing on the sharpened image to obtain a global detail enhanced image; Determining matching feature information of the field of view scene feature in the global detail enhanced image; Obtaining a local detail enhanced image based on highlighting processing of the matching feature information in the global detail enhanced image; The video image is generated based on the sharpened image and the local detail enhanced image.
2. The video image generation method according to claim 1, characterized in that: The method further includes determining, based on the acquired human-computer interaction instruction, a currently selected clarity level for the clarity processing from among at least two preset clarity levels, and a currently selected enhancement level for the field of view scene enhancement processing from among at least two preset enhancement levels.
3. The video image generation method according to claim 1, wherein: The generating of a sharpened image based on the image sharpening processing of the endoscopic field image includes: The endoscope field image is subjected to defogging processing using the suspended impurity feature as the fog feature to obtain the cleared image.
4. The video image generation method according to claim 1, wherein: The method further includes: generating a saliency map with the visual field scene feature as a highlighted object based on the sharpened image, and / or obtaining feature recognition information of the visual field scene feature in the sharpened image; The determining of matching feature information of the field of view scene feature in the global detail enhanced image includes: determining the matching feature information of the field of view scene feature in the global detail enhanced image using the saliency map and / or the feature recognition information as indication information.
5. The video image generation method according to claim 1, characterized in that: The step of generating the video image based on the sharpened image and the local detail enhanced image includes: Performing brightness equalization processing on the sharpened image to obtain a brightness equalized image; The video image is generated based on image fusion of the brightness-balanced image and the local detail-enhanced image.
6. The video image generation method according to claim 5, characterized in that: The performing brightness equalization processing on the cleared image to obtain a brightness equalized image includes: performing illumination estimation on the sharpened image; Based on the illumination estimation data obtained by the illumination estimation, determining a high brightness area and a low brightness area in the sharpened image divided by a brightness threshold as a boundary; The brightness-balanced image is obtained by performing brightness suppression on a high-brightness area of the sharpened image and performing brightness compensation on a low-brightness area of the sharpened image.
7. The video image generation method according to claim 1, characterized in that: Also includes: Determining an optimization mode for the image optimization based on the acquired human-computer interaction instruction, wherein the optimization mode includes a first optimization mode and a second optimization mode that are selectively enabled, and sequentially performing the image sharpening process and the field of view scene enhancement process on the endoscopic field of view image in the first optimization mode; In the second optimization mode, the endoscope field of view image is input into an end-to-end neural network based on deep learning; The end-to-end neural network is used to implement a mapping conversion from an input image to an output image, wherein the mapping conversion includes feature blurring with the suspended impurity feature as the blurring target, and the video image is the output image of the end-to-end neural network.
8. The video image generation method according to claim 7, characterized in that: Also includes: The end-to-end neural network is trained using the image sample set, wherein: The image sample set includes a plurality of image sample pairs, wherein a first image sample and a second image sample in each of the image sample pairs include the same suspended impurity sample feature, wherein the suspended impurity sample feature is presented in a visible state higher than a preset recognition threshold in the first image sample, and wherein the suspended impurity sample feature is presented in a blurred state lower than the preset recognition threshold in the second image sample; The training objective of the end-to-end neural network is set as follows: the conversion loss from the first image sample to the second image sample in each group of the image samples is lower than the target loss.
9. The video image generation method according to claim 8, characterized in that: The end-to-end neural network is a recurrent generative adversarial network, where: In a process of training the cyclic generative adversarial network using the image sample set, the cyclic generative adversarial network cyclically implements a bidirectional conversion between the first image sample and the second image sample in each group of the image samples; When the weighted value of the cycle consistency loss and the cycle perceptual consistency loss of the cycle generative adversarial network reaches a preset threshold, the training process ends because the training goal is achieved.
10. A video image generation device based on endoscopic imaging, characterized in that: include: An image acquisition module, configured to acquire an endoscopic field of view image based on endoscopic imaging, wherein the endoscopic field of view image includes field of view scene features within the field of view of the endoscope; an image optimization module, configured to generate a video image for visual presentation based on image optimization of the endoscopic field of view image, wherein the image optimization module is configured to blur the suspended impurity features while preserving the fidelity of the field of view scene features when the endoscopic field of view image also includes suspended impurity features that cause blurring of the field of view scene features; The image optimization module includes: an image sharpening submodule, configured to generate a sharpened image based on image sharpening processing of the endoscopic field of view image, wherein the image sharpening processing is configured to enhance the local contrast of the image region where the suspended impurity feature is located; a visual field scene enhancement submodule, configured to generate the video image based on visual field scene enhancement processing of the sharpened image, wherein the visual field scene enhancement processing is configured to enhance feature information of the visual field scene features; The field of view scene enhancement submodule is specifically configured to: perform global detail enhancement processing on the sharpened image to obtain a global detail enhanced image; determine matching feature information of the field of view scene features in the global detail enhanced image; obtain a local detail enhanced image based on highlighting processing of the matching feature information in the global detail enhanced image; and generate the video image based on the sharpened image and the local detail enhanced image.
11. The video image generating device according to claim 10, characterized in that: The image optimization module also includes: a level configuration submodule, which is used to determine, based on the acquired human-computer interaction instruction, a currently selected clarity level for the clarity processing from at least two preset clarity levels, and a currently selected enhancement level for the field of view scene enhancement processing from at least two preset enhancement levels.
12. The video image generating device according to claim 10, wherein: The image optimization module also includes: a mode enabling submodule, configured to determine an optimization mode for the image optimization based on the acquired human-computer interaction instruction, wherein the optimization mode includes a first optimization mode and a second optimization mode that are selectively enabled, and the image sharpening submodule and the field of view scene enhancement submodule are enabled in the first optimization mode; The model calling submodule enabled in the second optimization mode is used to input the endoscopic field of view image into an end-to-end neural network based on deep learning, wherein the end-to-end neural network is used to achieve a mapping transformation from the input image to the output image, the mapping transformation includes feature blurring with the suspended impurity feature as the blurring target, and the video image is the output image of the end-to-end neural network.
13. An endoscope system, characterized in that: include: An endoscopic imaging assembly, comprising an endoscope and a camera, wherein the camera is configured to generate an image of an endoscopic field of view imaged through the endoscope; A processor component, configured to execute the video image generation method according to any one of claims 1 to 9.
14. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores instructions that, when executed by a processor, cause the processor to perform the video image generation method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Methods and systems for generating clarified and enhanced intraoperative imaging data
US20230122835A1
Image enhancement method and apparatus suitable for endoscope, and storage medium
WO2021031459A1