Software and hardware combined fluorescence endoscope automatic exposure method for shot object and self-group iteration upgrading device
By combining hardware and software, a fluorescence endoscope is used to automatically adjust the exposure time of the endoscope camera using a deep noise-reducing convolutional neural network and a multi-agent reinforcement learning model. This solves the problem of inappropriate exposure of fluorescence and white light images, improves equipment utilization, and reduces medical costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-01
- Publication Date
- 2026-04-07
AI Technical Summary
In existing medical fluorescence endoscopy technology, the observation area for fluorescence image capture is inappropriate, and the exposure time for white light image capture is also inappropriate, resulting in low equipment utilization and high medical costs.
By combining hardware and software approaches, a deep noise-reducing convolutional neural network and a multi-agent reinforcement learning model are used to automatically adjust the exposure time of the endoscope camera. Gaussian filtering and guided filtering techniques are used to denoise the fluorescence image, and a multi-agent reinforcement learning environment is constructed to achieve automatic exposure control of the endoscope camera system.
It improves the signal-to-noise ratio of lesion tissue and non-lesion tissue, accurately delineates regions of interest, increases the correction amount and accuracy of endoscopic camera exposure time, improves equipment utilization, and reduces medical costs.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of endoscopic image processing technology, specifically to a method for automatic exposure of a subject using a fluorescence endoscope that combines hardware and software, and a self-group iterative upgrade device. Background Technology
[0002] Optical imaging technology, a newly emerging molecular imaging technique in recent years, boasts advantages such as high sensitivity, non-ionizing radiation, and ease of operation. It is a technology capable of imaging biophysical processes within biological tissues at the cellular and molecular level. Furthermore, by equipping endoscopes with fluorescence cameras and laser light sources, optical imaging technology is increasingly being applied to fields such as clinical surgical navigation. By injecting fluorescent agents such as indocyanine green into the surgical subject, illuminating the observation area with a laser mounted on the endoscope, and capturing the fluorescence information of the observed area under laser illumination by a fluorescence camera, combined with the white light image information captured by the camera, the endoscope can fluorescently mark lesions in the white light image, thus aiding the surgeon in precise lesion removal. Since the endoscope provides real-time image information to the surgeon, it needs to correctly adjust the exposure according to changes in light and shadow within the surgical subject. Moreover, because the endoscope penetrates deep into the surgical subject, selecting appropriate and targeted automatic exposure methods is particularly important. Summary of the Invention
[0003] To address the shortcomings of existing medical fluorescence endoscopy technology, the present invention aims to provide a method for automatic exposure of a fluorescence endoscope for imaging subjects, which combines hardware and software. This method can effectively solve the problems of low signal-to-background ratio, unsuitable observation area for fluorescence image imaging, and unsuitable exposure time for white light image imaging, thereby further improving equipment utilization and reducing medical costs.
[0004] This invention achieves the above objective through the following technical solution—a method for automatic exposure of a subject using a fluorescence endoscope combining hardware and software, comprising the following steps:
[0005] Step S100: Obtain the undenoised fluorescence image output by the fluorescence camera and the fluorescence image after denoising processing using Gaussian filtering and guided filtering techniques.
[0006] In some preferred embodiments, the fluorescence image after noise reduction is calculated as follows:
[0007] Step S110: Acquire the undenoised raw fluorescence image directly output by the fluorescence camera;
[0008] Step S120: Apply Gaussian filtering to the original fluorescence image without noise reduction to obtain a Gaussian filtered image;
[0009] Step S130: Apply Gaussian filtering technology to the image to obtain a guided-filter image, which is the fluorescence image after noise reduction.
[0010] Step S200: Use the original fluorescence image without noise reduction and the fluorescence image after noise reduction as a training set to construct and train a deep noise-reducing convolutional neural network.
[0011] After the deep noise-reducing convolutional neural network is trained.
[0012] Step S300: Construct an information acquisition device for the fluorescence endoscope camera system as the physical hardware environment for the multi-agent reinforcement learning environment. The information acquisition device has two functions: 1. Reading the visible light images and fluorescence image data acquired by the fluorescence endoscope camera system; 2. Receiving system instructions to control the exposure time of the endoscope camera system camera.
[0013] Step S400: Build the reinforcement learning system software environment—the software driver system of the endoscope camera system. The system can achieve the following two main functions: 1. Receive visible light images and fluorescence image data acquired from the fluorescence endoscope camera system in real time; 2. Issue system commands and parameters to control the exposure time of the endoscope camera system camera.
[0014] Step S500: Based on the deep reinforcement learning environment of the fluorescence endoscope camera system, a multi-agent reinforcement learning model is built.
[0015] In some preferred embodiments, the multi-agent reinforcement learning model is constructed as follows:
[0016] Step S510: The image data acquired by the reinforcement learning environment through the fluorescence endoscope camera system is used as the original undenoised fluorescence image.
[0017] Step S520: Input the original undenoised fluorescence image into a deep denoising convolutional neural network, and output the denoised fluorescence image as the state of the reinforcement learning environment. .
[0018] Step S530: The exposure time of the acquisition camera is used as the action of the reinforcement learning environment. .
[0019] Step S540: Calculate the image quality score based on the denoised fluorescence image described in step S520. This leads to the creation of a reward system for reinforcement learning environments. The image quality score The formula can be expressed as
[0020]
[0021] in The weighting coefficient for the central region; The absolute center moment of the observation area, This is the absolute center moment of the non-observable region.
[0022] The central region weighting coefficient The formula is
[0023]
[0024] The absolute center moment of the observation area Absolute central moment of non-observed region The formulas can be expressed as follows:
[0025]
[0026] in, , These refer to the number of pixels in the observation area and the non-observation area, respectively. The grayscale value of a pixel; , These are the average gray values of pixels in the observation area and the non-observation area, respectively. The formula for the average gray values of pixels in the observation area and the non-observation area can be expressed as:
[0027]
[0028] Rewards to enhance the learning environment The calculation formula is:
[0029]
[0030] in, , This is a balancing constant term used to constrain rewards to a reasonable range; These are the optimal image quality parameters.
[0031] In step S550, two intelligent agents, Actor (policy network) and Critic (value network), are constructed respectively, with their basic structure being a U-shaped network structure based on convolutional neural networks.
[0032] The objective of the multi-agent reinforcement learning model is the image quality score of the image data acquired by the endoscopic camera system. Approximately optimal image quality parameters .
[0033] In step S600, the Deep Deterministic Policy Gradient Algorithm (DDPG) is used to solve the multi-agent reinforcement learning model proposed in step S500.
[0034] In some preferred embodiments, the following steps are included:
[0035] In step S610, the reinforcement learning environment acquires image data through an endoscopic camera system as the raw, undenoised fluorescence image.
[0036] Step S620: Input the original undenoised fluorescence image into a deep denoising convolutional neural network, and output the denoised fluorescence image as the state of the reinforcement learning environment. .
[0037] Step S630: Set the state of the reinforcement learning environment. The input is given to the Actor (policy network) agent, and the output is the exposure time. .
[0038] Step S640, the reinforcement learning environment in the observation area Below, based on exposure time Collect image data in different spectral bands.
[0039] Step S650: The image data obtained in step S620 is scored for image quality, and the environmental reward is calculated using the reward calculation formula. .
[0040] After the multi-agent reinforcement learning model has been trained:
[0041] In step S700, fluorescence image data is acquired through an endoscopic camera system as the original, undenoised fluorescence image.
[0042] Step S800: Input the original undenoised fluorescence image into a deep denoising convolutional neural network, and output the denoised fluorescence image as the state of the reinforcement learning environment. .
[0043] Step S900: Set the state of the reinforcement learning environment. The input is given to the Actor (policy network) agent, and the output is the exposure time of the fluorescence image. .
[0044] Step S1000, in the observation area Below, based on exposure time Fluorescence images were acquired in various spectral bands, and the observation area was redefined.
[0045] Step S1100: Visible light image data is acquired through an endoscopic camera system as the state of the reinforcement learning environment.
[0046] Step S1200: Input the state of the reinforcement learning environment into the Actor (policy network) agent to obtain the exposure time of visible light.
[0047] In step S1300, visible light image data is acquired again using a fluorescence endoscope camera system.
[0048] Step S1400: Input the reinforcement learning environment state, actions, rewards, and original undenoised fluorescence images into the self-group iterative upgrade device to correct the multi-agent reinforcement learning model and the deep denoised convolutional neural network.
[0049] The self-group iterative upgrade device includes a single-machine self-iterative upgrade module and a networked group iterative upgrade module. The difference between self-upgrade and group upgrade lies in whether the training information (dataset, dataset size, dataset quality, etc.) originates from a local single machine (self-upgrade) or a networked group (group upgrade). The group iterative upgrade module includes two functions: local group iterative upgrade and cloud group iterative upgrade. The difference between local group upgrade and cloud group upgrade lies in whether the training process is completed locally or in the cloud. The training device can be local, in the cloud, or both. The local group iterative upgrade includes two functions: local group single-machine training iterative upgrade and local group distributed training iterative upgrade. The local group single-machine training iterative upgrade includes two functions: local group single-machine training with shared dataset iterative upgrade (without sharing training weights) and local group single-machine training with shared dataset and weights iterative upgrade.
[0050] The deep learning-based self-iterative upgrade device includes:
[0051] The storage module is used to collect the reinforcement learning environment state, actions, rewards, and raw, undenoised fluorescence images input to the deep denoised convolutional neural network, which are input to the multi-agent reinforcement learning model.
[0052] The expert annotation module is used to annotate the original, undenoised fluorescence images of the storage module and generate corresponding labels;
[0053] The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model and the deep denoising convolutional neural network based on the training data collected by the storage module.
[0054] The self-iterative upgrade step based on the self-iterative upgrade device involves downloading the reinforcement learning environment state, action, reward, and original undenoised fluorescence image through a network module and inputting them into the multi-agent reinforcement learning model and the deep denoised convolutional neural network, respectively, for training.
[0055] The deep learning-based group iterative upgrade device includes:
[0056] The storage module is used to collect the reinforcement learning environment state, actions, rewards, and raw, undenoised fluorescence images input to the deep denoised convolutional neural network, which are input to the multi-agent reinforcement learning model.
[0057] The expert annotation module is used to annotate the original, undenoised fluorescence images of the storage module and generate corresponding labels;
[0058] The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model and the deep denoising convolutional neural network based on the training data collected by the storage module.
[0059] The network service module is used for communication and data transmission with the cloud control center.
[0060] The group-based iterative upgrade of the local single-machine training shared dataset based on the group iterative upgrade module includes:
[0061] The training dataset comes from all networked devices;
[0062] The reinforcement learning environment state, actions, rewards, and original undenoised fluorescence images shared by several other networked devices are downloaded via a network module and input into the current local device's multi-agent reinforcement learning model and deep denoised convolutional neural network, respectively, for training on the local device. The resulting corrected multi-agent reinforcement learning model and deep denoised convolutional neural network weights are then used by the current local device.
[0063] The group-based local single-machine training shared dataset and weight iterative upgrade based on the group iterative upgrade module includes:
[0064] The training dataset comes from all networked devices;
[0065] Prerequisite: All networked devices must use the same training model software;
[0066] The reinforcement learning environment state, actions, rewards, and original undenoised fluorescence images from several other networked devices are downloaded to the current local device via a network module and input into the current local device's multi-agent reinforcement learning model and deep denoised convolutional neural network, respectively. Then, training is performed to obtain the corrected weights of the multi-agent reinforcement learning model and deep denoised convolutional neural network. The weights of the current local device's multi-agent reinforcement learning model and deep denoised convolutional neural network are updated and shared in real time with all networked devices for their use.
[0067] The group-based local distributed training iterative upgrade based on the group iterative upgrade module includes:
[0068] Training information (dataset size, dataset quality, etc.) comes from all networked devices;
[0069] Prerequisites: Either all networked devices use the same training model software (which requires knowledge of the dataset size for each device), or the quantitative differences in the parameters of the training datasets generated by each device are known (which requires knowledge of the dataset quality for each device).
[0070] The reinforcement learning environment state, actions, rewards, and original undenoised fluorescence images of each local device are input into the local device's multi-agent reinforcement learning model and deep denoised convolutional neural network for training, resulting in a corrected multi-agent reinforcement learning model and deep denoised convolutional neural network weights. The corrected multi-agent reinforcement learning model and deep denoised convolutional neural network weights of several networked devices are uploaded to the cloud control center for global reduction (AllReduce). If the training model software of all networked devices is the same, global reduction is performed based on the size of each device's dataset. If the training model software of the networked devices is different and the quantitative differences in the dataset parameters of each device are known, global reduction is performed based on the quantitative differences in the quality of each device's dataset. This yields an updated multi-agent reinforcement learning model and deep denoised convolutional neural network weights.
[0071] The group cloud-based iterative upgrade based on the group iterative upgrade module includes:
[0072] Prerequisite: The training model software must be identical across all networked devices and cloud-based devices;
[0073] The reinforcement learning environment state, actions, rewards, and raw undenoised fluorescence images of several networked devices are transmitted to the cloud control center through the network service module; the cloud control center summarizes the received reinforcement learning environment state, actions, rewards, and raw undenoised fluorescence images, and trains a multi-agent reinforcement learning model and a deep denoised convolutional neural network; the local network service module downloads and synchronously updates the weights of the current device's multi-agent reinforcement learning model and deep denoised convolutional neural network.
[0074] Therefore, compared with the prior art, the method provided by the present invention has the following beneficial effects:
[0075] On the one hand, the present invention has a high signal-to-background ratio (signal to background signal) at both the lesion site and the non-lesion site, which can effectively delineate the region of interest at the lesion site on each spectral fluorescence image. Then, multiple regions of interest are merged to obtain a comprehensive region of interest. Using the comprehensive region of interest obtained from the fluorescence image as a standard, the present invention can extract the required image quality parameters at the corresponding position on the white light image, and then calculate the exposure time correction of the camera on the endoscope to achieve the purpose of automatic exposure.
[0076] On the other hand, the present invention performs deep noise reduction processing on each spectral fluorescence image through deep learning algorithms and models, which can further improve the signal-to-background ratio of lesion location and non-lesion location in lesion tissue, thus further improving the accuracy of the exposure time correction of the camera on the endoscope. Attached Figure Description
[0077] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0078] Figure 1 This is a flowchart of an embodiment of a method for automatic exposure of a photographed object using a fluorescence endoscope that combines hardware and software according to the present invention.
[0079] Figure 2 This is a network structure diagram of a deep noise reduction convolutional neural network in the automatic exposure method for photographing objects using a fluorescence endoscope that combines hardware and software according to the present invention.
[0080] Figure 3 This is a schematic diagram of the self-group iterative upgrade device based on deep learning according to the present invention.
[0081] Figure 4 This is a schematic diagram of the self-iterative upgrade process based on deep learning in this invention.
[0082] Figure 5 This is a logical block diagram of the self-iterative upgrade module based on deep learning in this invention.
[0083] Figure 6 This is a schematic diagram of the process of iterative upgrading of a shared dataset based on local single-machine training of a group using deep learning, as described in this invention.
[0084] Figure 7 This is a logical block diagram of the deep learning-based group local single-machine training shared dataset iterative upgrade module of the present invention.
[0085] Figure 8 This is a schematic diagram of the process of group local single-machine training and weight iterative upgrade based on deep learning in this invention.
[0086] Figure 9 This is a logical block diagram of the deep learning-based group local single-machine training shared dataset and weight iterative upgrade module of the present invention.
[0087] Figure 10This is a schematic diagram of the process of iterative upgrade of group local distributed training based on deep learning in this invention.
[0088] Figure 11 This is a logical block diagram of the group-local distributed iterative upgrade module based on deep learning in this invention.
[0089] Figure 12 This is a schematic diagram of the process of the cloud-based group iterative upgrade based on deep learning in this invention.
[0090] Figure 13 This is a logical block diagram of the cloud-based group iterative upgrade module based on deep learning, which is the subject of this invention.
[0091] Figure 14 This is a schematic diagram of the network structure for group local single-machine training and iterative upgrading based on deep learning, as described in this invention.
[0092] Figure 15 This is a schematic diagram of the network structure of the present invention, which is based on deep learning for group local distributed and group cloud iterative upgrade.
[0093] Figure 16 This is a flowchart of the automatic exposure of a fluorescent endoscope for a photographed object according to the present invention.
[0094] Figure 17 This is a schematic diagram of the structure of a computer system used to implement the embodiments of the methods, systems, and apparatus of the present invention.
[0095] It should be noted that, Figure 14 and Figure 15 For illustration purposes only. The number of devices is not limited to 5 or more. The number of devices for group iterative upgrade can be 2, 3 or more. Detailed Implementation
[0096] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention. It should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0097] See Figure 1 The first embodiment of the present invention provides a method for automatic exposure of a photographed object using a combined hardware and software fluorescence endoscope, comprising the following steps:
[0098] Step S100: Obtain the undenoised fluorescence image output by the fluorescence camera and the fluorescence image after denoising processing using Gaussian filtering and guided filtering techniques.
[0099] In some preferred embodiments, the fluorescence image after noise reduction is calculated as follows:
[0100] Step S110: Acquire the undenoised raw fluorescence image directly output by the fluorescence camera;
[0101] Step S120: Apply Gaussian filtering to the original fluorescence image without noise reduction to obtain a Gaussian filtered image;
[0102] Step S130: Apply Gaussian filtering technology to the image to obtain a guided-filter image, which is the fluorescence image after noise reduction.
[0103] Step S140: All the finally obtained images are randomly divided into training, validation, and test sets according to a reasonable ratio (here, 7:2:1) for model training. The training process uses a computing platform equipped with an NVIDIA graphics card (Tesla A800 with 80GB). The learning rate of the neural network is set to a reasonable value (here, it can be set to...). The optimization algorithm uses Adam, and the optimizer parameters are set to reasonable values (here, it can be set to...). , ).
[0104] In some preferred embodiments, see Figure 2 The deep denoising convolutional neural network is constructed as follows:
[0105] The deep denoising convolutional neural network consists of a first module, several second modules, and a third module connected sequentially. The third module is a single-layer convolutional layer, the first module consists of convolutional layers and activation function layers connected sequentially, and the second module consists of convolutional layers, batch normalization layers, and activation function layers connected sequentially.
[0106] The activation function layer employs a nonlinear activation function. The activation function performs non-linear activation, as shown in the following formula:
[0107]
[0108] in Indicates the first The first layer of the network One neuron, Indicates the first The first layer of the network One neuron, Indicates the first The connection weights of each neuron, where ReLU represents the activation function, are shown in the following formula:
[0109]
[0110] Step S200: Use the original fluorescence image without denoising and the fluorescence image after denoising as a training set to construct and train a deep denoising convolutional neural network.
[0111] Step S300: Construct an information acquisition device for the fluorescence endoscope camera system as the physical hardware environment for the multi-agent reinforcement learning environment. The information acquisition device has two functions: 1. Reading the visible light images and fluorescence image data acquired by the fluorescence endoscope camera system; 2. Receiving system instructions to control the exposure time of the endoscope camera system camera.
[0112] Step S400: Build the reinforcement learning system software environment and the software driver system for the endoscope camera system. The system can achieve the following two main functions: 1. Receive visible light images and fluorescence image data acquired from the fluorescence endoscope camera system in real time; 2. Issue system commands and parameters to control the exposure time of the endoscope camera system camera.
[0113] Step S500: Based on the deep reinforcement learning environment of the fluorescence endoscope camera system, a multi-agent reinforcement learning model is built.
[0114] In some preferred embodiments, the multi-agent reinforcement learning model is constructed as follows:
[0115] Step S510: The image data acquired by the reinforcement learning environment through the fluorescence endoscope camera system is used as the original undenoised fluorescence image.
[0116] Step S520: Input the original undenoised fluorescence image into a deep denoising convolutional neural network, and output the denoised fluorescence image as the state of the reinforcement learning environment. .
[0117] Step S530: The exposure time of the acquisition camera is used as the action of the reinforcement learning environment. .
[0118] Step S540: Calculate the image quality score based on the denoised fluorescence image described in step S520. This leads to the creation of a reward system for reinforcement learning environments. The image quality score The formula can be expressed as
[0119]
[0120] in The weighting coefficient for the central region; The absolute center moment of the observation area, This is the absolute center moment of the non-observable region.
[0121] The central region weighting coefficient The formula is
[0122]
[0123] The absolute center moment of the observation area Absolute central moment of non-observed region The formulas can be expressed as follows:
[0124]
[0125] in, , These refer to the number of pixels in the observation area and the non-observation area, respectively. The grayscale value of a pixel; , These are the average gray values of pixels in the observation area and the non-observation area, respectively. The formula for the average gray values of pixels in the observation area and the non-observation area can be expressed as:
[0126]
[0127] Rewards to enhance the learning environment The calculation formula is:
[0128]
[0129] in, , This is a balancing constant term used to constrain rewards to a reasonable range; These are the optimal image quality parameters.
[0130] In step S550, two intelligent agents, Actor (policy network) and Critic (value network), are constructed respectively, with their basic structure being a U-shaped network structure based on convolutional neural networks.
[0131] The objective of the multi-agent reinforcement learning model is the image quality score of the image data acquired by the endoscopic camera system. Approximately optimal image quality parameters .
[0132] In step S600, the Deep Deterministic Policy Gradient Algorithm (DDPG) is used to solve the multi-agent reinforcement learning model proposed in step S500.
[0133] In some preferred embodiments, the following steps are included:
[0134] In step S610, the reinforcement learning environment acquires image data through an endoscopic camera system as the raw, undenoised fluorescence image.
[0135] Step S620: Input the original undenoised fluorescence image into a deep denoising convolutional neural network, and output the denoised fluorescence image as the state of the reinforcement learning environment. .
[0136] Step S630: Set the state of the reinforcement learning environment. The input is given to the Actor (policy network) agent, and the output is the exposure time. .
[0137] Step S640, the reinforcement learning environment in the observation area Below, based on exposure time Collect image data in different spectral bands.
[0138] Step S650: The image data obtained in step S430 is scored for image quality, and the environmental reward is calculated using the reward calculation formula. .
[0139] After the multi-agent reinforcement learning model has been trained:
[0140] In step S700, fluorescence image data is acquired through an endoscopic camera system as the original, undenoised fluorescence image.
[0141] Step S800: Input the original undenoised fluorescence image into a deep denoising convolutional neural network, and output the denoised fluorescence image as the state of the reinforcement learning environment. .
[0142] Step S900: Set the state of the reinforcement learning environment. The input is given to the Actor (policy network) agent, and the output is the exposure time of the fluorescence image. .
[0143] Step S1000, in the observation area Below, based on exposure time Fluorescence images were acquired in various spectral bands, and the observation area was redefined.
[0144] Step S1100: Visible light image data is acquired through an endoscopic camera system as the state of the reinforcement learning environment.
[0145] Step S1200: Input the state of the reinforcement learning environment into the Actor (policy network) agent to obtain the exposure time of the fluorescence image.
[0146] In step S1300, visible light image data is acquired again using a fluorescence endoscope camera system.
[0147] Step S1400: Input the reinforcement learning environment state, actions, rewards, and original undenoised fluorescence images into the self-group iterative upgrade device to correct the multi-agent reinforcement learning model and the deep denoised convolutional neural network.
[0148] See Figure 3 The self-group iterative upgrade device includes a single-machine self-iterative upgrade module and a networked group iterative upgrade module. The difference between self-upgrade and group upgrade lies in whether the training information (dataset, dataset size, dataset quality, etc.) originates from a local single machine (self-upgrade) or a networked group (group upgrade). The group iterative upgrade module includes two functions: local group iterative upgrade and cloud group iterative upgrade. The difference between local group upgrade and cloud group upgrade lies in whether the training process is completed locally or in the cloud. The training device can be local, in the cloud, or both. The local group iterative upgrade includes two functions: local group single-machine training iterative upgrade and local group distributed training iterative upgrade. The local group single-machine training iterative upgrade includes two functions: local group single-machine training with shared dataset iterative upgrade (without sharing training weights) and local group single-machine training with shared dataset and weights iterative upgrade.
[0149] The deep learning-based self-iterative upgrade device includes:
[0150] The storage module is used to collect the reinforcement learning environment state, actions, rewards, and raw, undenoised fluorescence images input to the deep denoised convolutional neural network, which are input to the multi-agent reinforcement learning model.
[0151] The expert annotation module is used to annotate the original, undenoised fluorescence images of the storage module and generate corresponding labels;
[0152] The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model and the deep denoising convolutional neural network based on the training data collected by the storage module.
[0153] See Figure 4 , Figure 5 The self-iterative upgrade step based on the self-iterative upgrade device is as follows: the reinforcement learning environment state, action, reward and original undenoised fluorescence image are downloaded through the network module and respectively input into the multi-agent reinforcement learning model and the deep denoised convolutional neural network for training.
[0154] The deep learning-based group iterative upgrade device includes:
[0155] The storage module is used to collect the reinforcement learning environment state, actions, rewards, and raw, undenoised fluorescence images input to the deep denoised convolutional neural network, which are input to the multi-agent reinforcement learning model.
[0156] The expert annotation module is used to annotate the original, undenoised fluorescence images of the storage module and generate corresponding labels;
[0157] The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model and the deep denoising convolutional neural network based on the training data collected by the storage module.
[0158] The network service module is used for communication and data transmission with the cloud control center.
[0159] See Figure 6 , Figure 7 , Figure 14 The group-based local single-machine training shared dataset iterative upgrade based on the group iterative upgrade module includes:
[0160] The training dataset comes from all networked devices;
[0161] The reinforcement learning environment state, actions, rewards, and original undenoised fluorescence images shared by several other networked devices are downloaded via a network module and input into the current local device's multi-agent reinforcement learning model and deep denoised convolutional neural network, respectively, for training on the local device. The resulting corrected multi-agent reinforcement learning model and deep denoised convolutional neural network weights are then used by the current local device.
[0162] See Figure 8 , Figure 9 , Figure 14 The group-based local single-machine training shared dataset and weight iterative upgrade based on the group iterative upgrade module includes:
[0163] The training dataset comes from all networked devices;
[0164] Prerequisite: All networked devices must use the same training model software;
[0165] The reinforcement learning environment state, actions, rewards, and original undenoised fluorescence images from several other networked devices are downloaded to the current local device via a network module and input into the current local device's multi-agent reinforcement learning model and deep denoised convolutional neural network, respectively. Then, training is performed to obtain the corrected weights of the multi-agent reinforcement learning model and deep denoised convolutional neural network. The weights of the current local device's multi-agent reinforcement learning model and deep denoised convolutional neural network are updated and shared in real time with all networked devices for their use.
[0166] See Figure 10, Figure 11 , Figure 15 The group-based local distributed training iterative upgrade based on the group iterative upgrade module includes:
[0167] Training information (dataset size, dataset quality, etc.) comes from all networked devices;
[0168] Prerequisites: Either all networked devices use the same training model software (which requires knowledge of the dataset size for each device), or the quantitative differences in the parameters of the training datasets generated by each device are known (which requires knowledge of the dataset quality for each device).
[0169] The reinforcement learning environment state, actions, rewards, and original undenoised fluorescence images of each local device are input into the local device's multi-agent reinforcement learning model and deep denoised convolutional neural network for training, resulting in a corrected multi-agent reinforcement learning model and deep denoised convolutional neural network weights. The corrected multi-agent reinforcement learning model and deep denoised convolutional neural network weights of several networked devices are uploaded to the cloud control center for global reduction (AllReduce). If the training model software of all networked devices is the same, global reduction is performed based on the size of each device's dataset. If the training model software of the networked devices is different and the quantitative differences in the dataset parameters of each device are known, global reduction is performed based on the quantitative differences in the quality of each device's dataset. This yields an updated multi-agent reinforcement learning model and deep denoised convolutional neural network weights.
[0170] See Figure 12 , Figure 13 , Figure 15 The group cloud-based iterative upgrade based on the group iterative upgrade module includes:
[0171] Prerequisite: The training model software must be identical across all networked devices and cloud-based devices;
[0172] The reinforcement learning environment state, actions, rewards, and raw undenoised fluorescence images of several networked devices are transmitted to the cloud control center through the network service module; the cloud control center summarizes the received reinforcement learning environment state, actions, rewards, and raw undenoised fluorescence images, and trains a multi-agent reinforcement learning model and a deep denoised convolutional neural network; the local network service module downloads and synchronously updates the weights of the current device's multi-agent reinforcement learning model and deep denoised convolutional neural network.
[0173] The control flowchart of the automatic exposure method for the photographed object using a fluorescence endoscope based on a multi-agent reinforcement learning model according to the second embodiment of the present invention is shown below. Figure 16 :
[0174] Step S100: Initialize the fluorescence endoscope camera system, turn on the endoscope host and power supply, and switch to fluorescence acquisition mode.
[0175] Step S200: Fluorescence image data is acquired through an endoscopic camera system as the original, undenoised fluorescence image.
[0176] Step S300: Input the original undenoised fluorescence image into a deep denoising convolutional neural network, and output the denoised fluorescence image as the state of the reinforcement learning environment. At the same time, the observation area was divided. .
[0177] Step S400, the agent (policy network) reads the state. Exposure time for output fluorescence image This controls the exposure time of the fluorescence camera.
[0178] Step S500, in the observation area The fluorescence images of each spectral band acquired by the fluorescence endoscope imaging system are read, and the sharpness index corresponding to the image data is calculated.
[0179] Step S600: Redivide the observation area ;
[0180] In step S700, visible light images are acquired through an endoscopic camera system, input into a deep denoising convolutional neural network, and the denoised visible light images are output as the state of the reinforcement learning environment. .
[0181] Step S800: Set the state of the reinforcement learning environment. The exposure time of the fluorescence image is obtained by inputting it into the Actor (policy network) agent. .
[0182] In step S900, visible light image data is acquired again through the endoscopic camera system.
[0183] The third embodiment of the present invention is a hardware and software integrated automatic exposure control system for a fluorescence endoscope. The system is based on a hardware and software integrated method for automatic exposure of a fluorescence endoscope for a photographed object. The system includes: an acquisition module, a deep learning training module, and a reinforcement learning training module.
[0184] The acquisition module is configured to acquire fluorescence image data or visible light image data from the endoscopic camera system, input the data into a deep denoising convolutional neural network, and obtain denoised image data, which serves as the state of the reinforcement learning environment. The denoised image is then input into a trained multi-agent reinforcement learning model to obtain the exposure time. .
[0185] The training method for the deep denoising convolutional neural network is as follows:
[0186] The deep learning training module is configured to obtain denoised image data;
[0187] The original fluorescence image without denoising and the fluorescence image after denoising are used as the training set to construct and train a deep denoising convolutional neural network.
[0188] The training method for the multi-agent reinforcement learning model is as follows:
[0189] The reinforcement learning training module is configured to acquire exposure time. ;
[0190] By building an endoscope camera system information acquisition device and a software-driven system as a complete reinforcement learning environment, a multi-agent reinforcement learning model is constructed and trained.
[0191] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process and related descriptions of the storage device and processing device described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0192] It should be noted that the automatic control system for fluorescence endoscopy based on a multi-agent reinforcement learning model provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be merged into one module, or further divided into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing the various modules or steps and are not considered as an improper limitation of the present invention.
[0193] An electronic device according to a fourth embodiment of the present invention includes: at least one processor and a memory communicatively connected to at least one of the processors; wherein the memory stores instructions executable by the processor, the instructions being executed by the processor to implement the above-described method for automatic exposure of a photographed object by a fluorescence endoscope combining hardware and software.
[0194] A computer-readable storage medium according to a fifth embodiment of the present invention stores computer instructions, which are executed by the computer to implement the above-described method for automatic exposure of a photographed object by a fluorescence endoscope combining hardware and software.
[0195] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process and related descriptions of the storage device and processing device described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0196] Those skilled in the art will recognize that the modules and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. The programs corresponding to the software modules and method steps can be placed in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art. To clearly illustrate the interchangeability of electronic hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the invention.
[0197] The following is for reference. Figure 17 It shows a schematic diagram of the structure of a computer system for implementing embodiments of the methods, systems, and apparatus of the present invention. Figure 17 The server shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0198] like Figure 17 As shown, the computer system includes a Central Processing Unit (CPU) 501, which can perform various appropriate actions and processes based on programs stored in Read Only Memory (ROM) 502 or programs loaded from storage section 508 into Random Access Memory (RAM) 503. RAM 503 also stores various programs and data required for system operation. The CPU 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0199] The following components are connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. Removable media 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 510 as needed so that computer programs read from them can be installed into storage section 508 as needed.
[0200] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit (CPU) 501, it performs the functions defined in the methods of this invention. It should be noted that the computer-readable medium described above in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0201] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0202] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0203] The terms “first”, “second”, etc., are used to distinguish similar objects, not to describe or indicate a specific order or sequence.
[0204] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent in such process, method, article, or apparatus / device.
[0205] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A method for automatic exposure of an endoscopic subject using a fluorescence endoscope based on deep reinforcement learning, characterized in that, The method includes: Step S100: Obtain the undenoised fluorescence image output by the fluorescence camera and the fluorescence image after denoising by Gaussian filtering and mean filtering techniques; Step S200: Use the original fluorescence image without noise reduction and the fluorescence image after noise reduction as a training set to construct and train a deep noise-reducing convolutional neural network; Step S300: Construct an information acquisition device for the fluorescence endoscope camera system as the physical hardware environment for the multi-agent reinforcement learning environment. The information acquisition device has two functions:
1. Reading the visible light image and fluorescence image data acquired by the fluorescence endoscope camera system; 2. Receiving system instructions to control the exposure time of the endoscope camera system camera. Step S400: Build the reinforcement learning system software environment and the software driver system for the endoscope camera system. The system can achieve the following two main functions:
1. Receive visible light images and fluorescence image data acquired from the fluorescence endoscope camera system in real time; 2. Issue system commands and parameters to control the exposure time of the endoscope camera system camera. Step S500: Based on the deep reinforcement learning environment of the fluorescence endoscope camera system, build a multi-agent reinforcement learning model; In step S600, the deep deterministic policy gradient algorithm (DDPG) is used to solve the multi-agent reinforcement learning model proposed in step S300. After the multi-agent reinforcement learning model has been trained: Step S700: Fluorescence image data is acquired through the endoscopic camera system as the original undenoised fluorescence image; Step S800: Input the original undenoised fluorescence image into a deep denoising convolutional neural network, and output the denoised fluorescence image as the state of the reinforcement learning environment. ; Step S900: Set the state of the reinforcement learning environment. The input is given to the Actor agent, and the output is the exposure time of the fluorescence image. ; Step S1000, in the observation area Below, based on exposure time Fluorescence images were acquired in various spectral bands, and the observation area was redefined. Step S1100: Visible light image data is acquired through an endoscopic camera system as the state of the reinforcement learning environment; Step S1200: Input the state of the reinforcement learning environment into the Actor agent to obtain the exposure time of visible light; Step S1300: Visible light image data is acquired again using the fluorescence endoscope camera system; Step S1400: Input the reinforcement learning environment state, actions, rewards, and original undenoised fluorescence images into the self-group iterative upgrade device to correct the multi-agent reinforcement learning model and the deep denoised convolutional neural network.
2. A self-group iterative upgrade device based on deep learning, characterized in that, The self-group iterative upgrade device includes a single-machine self-iterative upgrade module and a networked group iterative upgrade module. The difference between self-upgrade and group upgrade lies in whether the training information (dataset, dataset size, dataset quality, etc.) originates from a local single machine (self-upgrade) or a networked group (group upgrade). The group iterative upgrade module includes two functions: local group iterative upgrade and cloud group iterative upgrade. The difference between local group upgrade and cloud group upgrade lies in whether the training process is completed locally or in the cloud. The training device can be local, in the cloud, or both. The local group iterative upgrade includes two functions: local group single-machine training iterative upgrade and local group distributed training iterative upgrade. The local group single-machine training iterative upgrade includes two functions: local group single-machine training with shared dataset iterative upgrade (without sharing training weights) and local group single-machine training with shared dataset and weights iterative upgrade.
3. The self-iterative upgrade module based on deep learning according to claim 2, characterized in that, The deep learning-based self-iterative upgrade module includes: The storage module is used to collect the reinforcement learning environment state, actions, rewards, and raw, undenoised fluorescence images input to the deep denoised convolutional neural network, which are input to the multi-agent reinforcement learning model. The expert annotation module is used to annotate the original, undenoised fluorescence images of the storage module and generate corresponding labels; The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model and the deep denoising convolutional neural network based on the training data collected by the storage module. The self-iterative upgrade step based on the self-iterative upgrade device involves downloading the reinforcement learning environment state, action, reward, and original undenoised fluorescence image through a network module and inputting them into the multi-agent reinforcement learning model and the deep denoised convolutional neural network, respectively, for training.
4. The deep learning-based population iterative upgrade module according to claim 2, characterized in that, The deep learning-based group iterative upgrade module includes: The storage module is used to collect the reinforcement learning environment state, actions, rewards, and raw, undenoised fluorescence images input to the deep denoised convolutional neural network, which are input to the multi-agent reinforcement learning model. The expert annotation module is used to annotate the original, undenoised fluorescence images of the storage module and generate corresponding labels; The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model and the deep denoising convolutional neural network based on the training data collected by the storage module. The network service module is used for communication and data transmission with the cloud control center.
5. The group local single-machine training shared dataset iterative upgrade function of the deep learning-based group iterative upgrade module according to claim 2, characterized in that, The iterative upgrade of the shared local single-machine training dataset based on the group iterative upgrade module includes: The training dataset comes from all networked devices; The reinforcement learning environment state, actions, rewards, and original undenoised fluorescence images shared by several other networked devices are downloaded via a network module and input into the current local device's multi-agent reinforcement learning model and deep denoised convolutional neural network, respectively, for training on the local device. The resulting corrected multi-agent reinforcement learning model and deep denoised convolutional neural network weights are then used by the current local device.
6. The group local single-machine training and weight iterative upgrade function of the deep learning-based group iterative upgrade module according to claim 2, characterized in that, The group-based local single-machine training shared dataset and weight iterative upgrade based on the group iterative upgrade module includes: The training dataset comes from all networked devices; Prerequisite: All networked devices must use the same training model software; The reinforcement learning environment state, actions, rewards, and original undenoised fluorescence images from several other networked devices are downloaded to the current local device via a network module and input into the current local device's multi-agent reinforcement learning model and deep denoised convolutional neural network, respectively. Then, training is performed to obtain the corrected weights of the multi-agent reinforcement learning model and deep denoised convolutional neural network. The weights of the current local device's multi-agent reinforcement learning model and deep denoised convolutional neural network are updated and shared in real time with all networked devices for their use.
7. The group-local distributed training iterative upgrade function of the group iterative upgrade module based on deep learning according to claim 2, characterized in that, The group-based local distributed training iterative upgrade based on the group iterative upgrade module includes: Training information (dataset size, dataset quality, etc.) comes from all networked devices; Prerequisites: Either all networked devices use the same training model software (which requires knowledge of the dataset size for each device), or the quantitative differences in the parameters of the training datasets generated by each device are known (which requires knowledge of the dataset quality for each device). The reinforcement learning environment state, actions, rewards, and original undenoised fluorescence images of each local device are input into the local device's multi-agent reinforcement learning model and deep denoised convolutional neural network for training, resulting in a corrected multi-agent reinforcement learning model and deep denoised convolutional neural network weights. The corrected multi-agent reinforcement learning model and deep denoised convolutional neural network weights of several networked devices are uploaded to the cloud control center for global reduction (AllReduce). If the training model software of all networked devices is the same, global reduction is performed based on the size of each device's dataset. If the training model software of the networked devices is different and the quantitative differences in the dataset parameters of each device are known, global reduction is performed based on the quantitative differences in the quality of each device's dataset. This yields an updated multi-agent reinforcement learning model and deep denoised convolutional neural network weights.
8. The group cloud-based iterative upgrade function of the group iterative upgrade module based on deep learning according to claim 2, characterized in that, The group cloud-based iterative upgrade based on the group iterative upgrade module includes: Prerequisite: The training model software must be identical across all networked devices and cloud-based devices; The reinforcement learning environment state, actions, rewards, and raw undenoised fluorescence images of several networked devices are transmitted to the cloud control center through the network service module; the cloud control center summarizes the received reinforcement learning environment state, actions, rewards, and raw undenoised fluorescence images, and trains a multi-agent reinforcement learning model and a deep denoised convolutional neural network; the local network service module downloads and synchronously updates the weights of the current device's multi-agent reinforcement learning model and deep denoised convolutional neural network.