Deep reinforcement learning-based parfocal method for visible light image and fluorescence image of three-dimensional fluorescence endoscope and self-group iteration upgrading device

By employing deep reinforcement learning technology, the visible light and fluorescence images of the three-dimensional fluorescence endoscope system are made cofocal, solving the problem of focal length differences in traditional endoscope systems, improving image clarity and lesion detection efficiency, and making it suitable for endoscope systems in the medical and industrial fields.

CN121053019APending Publication Date: 2025-12-02潘世遗
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511239254.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

In traditional endoscopic systems, visible light and fluorescence imaging are usually separated, resulting in focal length differences and the inability to image simultaneously, which affects image clarity and lesion detection efficiency.

Method used

A deep reinforcement learning-based approach is adopted to achieve co-focusing of visible light and fluorescence images of a 3D fluorescence endoscope by building a multi-agent reinforcement learning environment. This includes image denoising, uniformity compensation, brightness compensation, color correction, and fusion. The focus parameters are optimized using Actor and Critic networks, and the model is corrected and optimized by combining a self-group iterative upgrade device.

Benefits of technology

It improves image quality, enhances lesion detection capabilities, assists in navigation and localization, improves surgical outcomes, and reduces the risk of errors and complications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a deep reinforcement learning-based method for parfocal of a visible light image and a fluorescence image of a three-dimensional fluorescence endoscope and a self-group iteration upgrading device. The method comprises the following steps of: simultaneously acquiring a plurality of frames of visible light images and fluorescence images; preprocessing the plurality of frames of visible light images and fluorescence images; on the basis of carrying out definition evaluation on a fluorescence fusion image, through interactive training of an intelligent agent and a three-dimensional endoscope fluorescence imaging intelligent agent, a zero-sum game is formed, and the intelligent agent is gradually realized to rapidly and accurately operate a three-dimensional fluorescence endoscope camera system to carry out visible light and fluorescence image automatic focusing. The invention is used for solving the technical problem of incapability of simultaneous imaging due to focal length difference between visible light imaging and fluorescence imaging in the prior art, thereby achieving the purposes of improving the image quality and enhancing the detection capability on lesion. According to the self-group iteration upgrading device, model training iteration upgrading can be carried out in multiple modes, and the performance of the device is continuously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of endoscopic image focusing technology, specifically to a method and self-group iterative upgrading device for focusing visible light and fluorescence images of a three-dimensional fluorescence endoscope based on deep reinforcement learning. Background Technology

[0002] Currently, endoscopic imaging systems are widely used in medical, industrial, and scientific fields to observe and record images and videos of areas that are difficult to access directly. Fluorescence imaging technology, in particular, has important applications in the medical field, such as cancer detection and treatment. However, in traditional endoscopic imaging systems, visible light imaging and fluorescence imaging are usually separate processes. This means that doctors or operators need to switch between different modes during observation, leading to inconvenience and inefficiency. Simultaneous imaging of visible light and fluorescence can result in unclear or inaccurate fluorescence fusion images due to the focal length difference between the two imaging methods.

[0003] Deep Reinforcement Learning (DRL) is an artificial intelligence technique that combines neural networks and reinforcement learning to solve complex decision-making problems. In recent years, DRL has achieved significant results, as exemplified by AlphaGo and AlphaFold. Image recognition, an important branch of computer vision, aims to identify and classify objects by analyzing features in images. With the increasing volume of data and improved computing power, deep learning techniques have made significant progress in image recognition, as seen in technologies such as ResNet and Inception. Summary of the Invention

[0004] To overcome the shortcomings of existing technologies, the present invention aims to provide a method based on deep reinforcement learning algorithms for co-focusing visible light and fluorescence images in three-dimensional fluorescence endoscopy. This method addresses the technical problem that existing technologies cannot simultaneously image visible light and fluorescence images due to focal length differences, thereby improving image quality and enhancing the ability to detect lesions.

[0005] This invention achieves the above objective through the following technical solution—a method for co-focusing visible light and fluorescence images of a three-dimensional fluorescence endoscope based on deep reinforcement learning, the method comprising the following steps:

[0006] Step S100: Construct a data acquisition device for visible light and fluorescence images of a three-dimensional fluorescence endoscope system. This device serves as the physical hardware environment for a multi-agent reinforcement learning environment. The device has two functions: First, it simultaneously acquires several frames of visible light images and several frames of fluorescence images through the imaging equipment in the three-dimensional fluorescence endoscope camera system. Second, it receives system command parameters to enable the imaging equipment to automatically focus to a specified value.

[0007] Step S200: Set up the reinforcement learning software environment.

[0008] In some preferred embodiments, the reinforcement learning software environment is constructed as follows:

[0009] Step S210: Build the software driver system for the three-dimensional fluorescence endoscope camera system. The system can achieve the following two main functions: 1. Receive in real time several frames of visible light images and several frames of fluorescence images simultaneously acquired by the imaging device in the three-dimensional fluorescence endoscope camera system; 2. Issue focusing commands and parameters to enable the imaging device in the three-dimensional fluorescence endoscope camera system to automatically focus to the specified value.

[0010] Step S220: Build an image denoising model for a three-dimensional fluorescence endoscope camera system based on deep learning.

[0011] Step S230: Build an image uniformity compensation model for a three-dimensional fluorescence endoscope camera system based on deep learning.

[0012] Step S240: Build an image brightness compensation model for a three-dimensional fluorescence endoscope camera system based on deep learning.

[0013] Step S250: Build an image color correction model for a three-dimensional fluorescence endoscope camera system based on deep learning.

[0014] Step S260: Build a fusion model of visible light and fluorescence images of a three-dimensional fluorescence endoscope camera system based on deep learning.

[0015] Step S270: Construct an image sharpness evaluation module. This module uses the sum of the absolute values ​​of the gray-level differences in four neighboring regions as the evaluation function. Image region sharpness evaluation function Its formula can be expressed as:

[0016]

[0017] in The coordinates are ( The pixel value of ).

[0018] Step S300: Based on the reinforcement learning environment of the three-dimensional fluorescence endoscope camera system, a multi-agent reinforcement learning model is built.

[0019] In some preferred embodiments, the multi-agent reinforcement learning model is constructed as follows:

[0020] Step S310: The reinforcement learning environment acquires the original image through a three-dimensional fluorescence endoscope camera system. After noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, the output fluorescence fusion image is used as the state of the reinforcement learning environment. .

[0021] Step S320: The focusing parameters of the three-dimensional fluorescence endoscope camera system output by the intelligent agent are used as the actions of the reinforcement learning environment. .

[0022] Step S330: Based on the reinforcement learning environment, the original image is acquired through a three-dimensional fluorescence endoscope camera system. After noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, the output fluorescence fusion image is used to construct the reward of the reinforcement learning environment. Its formula can be expressed as

[0023]

[0024] in, This is a function for evaluating image sharpness. , This is a balancing constant term used to constrain rewards to a reasonable range.

[0025] In step S340, construct two agents, Actor (policy network) and Critic (value network), respectively. Their basic structure is a U-shaped network structure with 6 layers.

[0026] The objective of the multi-agent reinforcement learning model is to maximize the image sharpness evaluation function. .

[0027] Step S400: The Proximal Policy Optimization (PPO) algorithm is used to optimize the image sharpness evaluation function proposed in step S340. A multi-agent reinforcement learning model is used to solve the problem, which improves the robustness of the fluorescence imaging system and evaluates its impact on the fluorescence image fusion end.

[0028] In some preferred embodiments, the following steps are included:

[0029] Step S410: The reinforcement learning environment acquires the original image through a three-dimensional fluorescence endoscope camera system, and outputs a fluorescence fusion image after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion. .

[0030] Step S420: Input the fluorescence fusion image into the Actor (policy network) agent, and output the focusing parameters of the three-dimensional fluorescence endoscope imaging system. .

[0031] In step S430, the reinforcement learning environment focuses the three-dimensional fluorescence endoscope camera system according to the focusing parameters. Then, the reinforcement learning environment acquires the original image through the three-dimensional fluorescence endoscope camera system, and outputs a fluorescence fused image after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion. The reward for the reinforcement learning environment is obtained by evaluating the sharpness of the fused fluorescence image. .

[0032] After the multi-agent reinforcement learning model is trained:

[0033] In step S500, the original image is acquired through the three-dimensional fluorescence endoscope camera system. After noise reduction, uniformity compensation, brightness compensation, color correction and fusion, the output fluorescence fusion image is input into the Actor (policy network) intelligent agent to obtain the focusing parameters of the three-dimensional fluorescence endoscope camera system. The focusing of the three-dimensional fluorescence endoscope camera system can be completed by using the focusing parameters.

[0034] Step S600: Set the state of the reinforcement learning environment. , and the focusing parameters of the three-dimensional fluorescence endoscope imaging system The input is fed into the self-group iterative upgrade mechanism of the multi-agent reinforcement learning model to calculate the corresponding environmental reward. The multi-agent reinforcement learning model is modified.

[0035] The self-group iterative upgrade device includes a single-machine self-iterative upgrade module and a networked group iterative upgrade module. The difference between self-upgrade and group upgrade lies in whether the training information (dataset, dataset size, dataset quality, etc.) originates from a local single machine (self-upgrade) or a networked group (group upgrade). The group iterative upgrade module includes two functions: local group iterative upgrade and cloud group iterative upgrade. The difference between local group upgrade and cloud group upgrade lies in whether the training process is completed locally or in the cloud. The training device can be local, in the cloud, or both. The local group iterative upgrade includes two functions: local group single-machine training iterative upgrade and local group distributed training iterative upgrade. The local group single-machine training iterative upgrade includes two functions: local group single-machine training with shared dataset iterative upgrade (without sharing training weights) and local group single-machine training with shared dataset and weights iterative upgrade.

[0036] The self-iterative upgrade device based on deep reinforcement learning includes:

[0037] The environmental reward evaluation module is used to evaluate the state of the reinforcement learning environment. Corresponding environmental rewards .

[0038] The storage module is used to collect the reinforcement learning environment states input to the multi-agent reinforcement learning model. ;

[0039] The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model based on the data collected by the storage module;

[0040] The self-iterative upgrade step based on the self-iterative upgrade device involves using the fluorescence fusion image acquired by the three-dimensional fluorescence endoscope imaging system, after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The input is fed into the multi-agent reinforcement learning model for training.

[0041] The deep reinforcement learning-based population iterative upgrade device includes:

[0042] The environmental reward evaluation module is used to evaluate the state of the reinforcement learning environment. Corresponding environmental rewards .

[0043] The storage module is used to collect the reinforcement learning environment states input to the multi-agent reinforcement learning model. ;

[0044] The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model based on the images collected by the storage module;

[0045] The network service module is used for communication and data transmission with the cloud control center.

[0046] The iterative upgrade of the shared local single-machine training dataset based on the group iterative upgrade module includes:

[0047] The training dataset comes from all networked devices;

[0048] The fluorescence fusion image, acquired by a three-dimensional fluorescence endoscope camera system and shared by several other networked devices, after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, is used as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The multi-agent reinforcement learning model is downloaded and input into the current local device via the network module, trained locally, and the corrected weights of the multi-agent reinforcement learning model are obtained for use by the current local device.

[0049] The group-based local single-machine training shared dataset and weight iterative upgrade based on the group iterative upgrade module includes:

[0050] The training dataset comes from all networked devices;

[0051] Prerequisite: All networked devices must use the same training model software;

[0052] The fluorescence fusion image, acquired by several other networked devices through a 3D fluorescence endoscope camera system and processed through noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, is used as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The data is downloaded to the local device via the network module and input into the local device's multi-agent reinforcement learning model. The model is then trained to obtain the corrected weights of the multi-agent reinforcement learning model. The weights of the local device's multi-agent reinforcement learning model are then updated and shared in real time with all networked devices for their use.

[0053] The group-based local distributed training iterative upgrade based on the group iterative upgrade module includes:

[0054] Training information (dataset size, dataset quality, etc.) comes from all networked devices;

[0055] Prerequisites: Either all networked devices use the same training model software (which requires knowledge of the dataset size for each device), or the quantitative differences in the parameters of the training datasets generated by each device are known (which requires knowledge of the dataset quality for each device).

[0056] The fluorescence fusion images acquired by each current local device through a 3D fluorescence endoscope camera system, after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, are used as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The input is fed into the local device's multi-agent reinforcement learning model for training, resulting in the corrected multi-agent reinforcement learning model weights. The corrected multi-agent reinforcement learning model weights from several networked devices are then uploaded to the cloud control center for global reduction (AllReduce). If all networked devices use the same training model software, global reduction is performed based on the size of each device's dataset. If the networked devices use different training model software and the quantitative differences in the dataset parameters of each device are known, global reduction is performed based on the quantitative differences in the quality of each device's dataset. The updated multi-agent reinforcement learning model weights are then obtained.

[0057] The group cloud-based iterative upgrade based on the group iterative upgrade module includes:

[0058] Prerequisite: The training model software must be identical across all networked devices and cloud-based devices;

[0059] The fluorescence fusion image, acquired by several networked devices through a three-dimensional fluorescence endoscope camera system and processed after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, is used as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The data is transmitted to the cloud control center via the network service module; the cloud control center then summarizes the received reinforcement learning environment status. , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The system trains a multi-agent reinforcement learning model and downloads and updates the weights of the current device's multi-agent reinforcement learning model synchronously using the local network service module.

[0060] Therefore, compared with the prior art, the present invention has the following beneficial effects:

[0061] (1) Improve lesion detection rate: The three-dimensional fluorescence endoscope with visible light and fluorescence parfocal design can utilize visible light and fluorescence images simultaneously, improve image quality, and thus enhance the ability to detect lesions;

[0062] (2) Assisted navigation and positioning: This design can also be used to assist doctors in navigation and positioning, especially in minimally invasive surgery, where doctors can more accurately locate and orient the treatment area, reducing the risk of errors and damage to surrounding healthy tissues;

[0063] (3) Improve surgical outcomes: By providing clearer visualization and more accurate diagnosis of lesions, the visible light fluorescence parfocal design of the three-dimensional fluorescence endoscope helps to improve surgical outcomes, reduce the risk of surgical complications, and shorten recovery time. Attached Figure Description

[0064] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0065] Figure 1 This is a flowchart of an embodiment of the present invention, which is a method for co-focusing visible light and fluorescence images of a three-dimensional fluorescence endoscope based on deep reinforcement learning.

[0066] Figure 2 This is a schematic diagram of the network structure of Actor (policy network) and Critic (value network) in a method based on deep reinforcement learning for cofocalizing visible light and fluorescence images of a three-dimensional fluorescence endoscope according to the present invention.

[0067] Figure 3 This is a flowchart of an embodiment of the Proximal Policy Optimization (PPO) algorithm in a method for cofocalizing visible light and fluorescence images of a three-dimensional fluorescence endoscope based on deep reinforcement learning, according to the present invention.

[0068] Figure 4 This is a flowchart illustrating the implementation of image co-focusing in a method for co-focusing visible light and fluorescence images of a three-dimensional fluorescence endoscope based on deep reinforcement learning, as described in this invention.

[0069] Figure 5This is a schematic diagram of the self-group iterative upgrade device based on deep reinforcement learning according to the present invention.

[0070] Figure 6 This is a schematic diagram of the self-iterative upgrade process based on deep reinforcement learning in this invention.

[0071] Figure 7 This is a logical block diagram of the self-iterative upgrade module based on deep reinforcement learning in this invention.

[0072] Figure 8 This is a schematic diagram illustrating the process of iterative upgrading of a shared dataset for group local single-machine training based on deep reinforcement learning, as described in this invention.

[0073] Figure 9 This is a logical block diagram of the group local single-machine training shared dataset iterative upgrade module based on deep reinforcement learning in this invention.

[0074] Figure 10 This is a schematic diagram of the process of group local single-machine training and weight iterative upgrade based on deep reinforcement learning in this invention.

[0075] Figure 11 This is a logical block diagram of the shared dataset and weight iterative upgrade module for group local single-machine training based on deep reinforcement learning, which is based on the present invention.

[0076] Figure 12 This is a schematic diagram of the process of iterative upgrade of population local distributed training based on deep reinforcement learning in this invention.

[0077] Figure 13 This is a logical block diagram of the population-locally distributed iterative upgrade module based on deep reinforcement learning in this invention.

[0078] Figure 14 This is a schematic diagram of the process of the group cloud-based iterative upgrade based on deep reinforcement learning in this invention.

[0079] Figure 15 This is a logical block diagram of the group cloud-based iterative upgrade module based on deep reinforcement learning in this invention.

[0080] Figure 16 This is a schematic diagram of the network structure for group-based local single-machine training and iterative upgrading based on deep reinforcement learning, as described in this invention.

[0081] Figure 17 This is a schematic diagram of the network structure of the present invention based on deep reinforcement learning, which features local distributed and cloud-based iterative upgrades.

[0082] Figure 18 This is a schematic diagram of the structure of a computer system used to implement the embodiments of the methods, systems, and apparatus of the present invention.

[0083] It should be noted that, Figure 16 and Figure 17 For illustration purposes only. The number of devices is not limited to 5 or more. The number of devices for group iterative upgrade can be 2, 3 or more. Detailed Implementation

[0084] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0085] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0086] See Figure 1 The first embodiment of the present invention provides a method for confocalization of visible light and fluorescence images of a three-dimensional fluorescence endoscope based on deep reinforcement learning, comprising the following steps:

[0087] Step S100: Construct a data acquisition device for visible light and fluorescence images of a three-dimensional fluorescence endoscope system. This device serves as the physical hardware environment for a multi-agent reinforcement learning environment. The device has two functions: First, it simultaneously acquires several frames of visible light images and several frames of fluorescence images through the imaging equipment in the three-dimensional fluorescence endoscope camera system. Second, it receives instruction parameters to enable the imaging equipment to automatically focus to a specified value.

[0088] Step S200: Set up the reinforcement learning software environment.

[0089] In some preferred embodiments, the reinforcement learning software environment is constructed as follows:

[0090] Step S210: Build the software driver system for the three-dimensional fluorescence endoscope camera system. The system can achieve the following two main functions: 1. Receive in real time several frames of visible light images and several frames of fluorescence images simultaneously acquired by the imaging device in the three-dimensional fluorescence endoscope camera system; 2. Issue focusing commands and parameters to enable the imaging device in the three-dimensional fluorescence endoscope camera system to automatically focus to the specified value.

[0091] Step S220: Build an image denoising model for a three-dimensional fluorescence endoscope camera system based on deep learning.

[0092] Step S230: Build an image uniformity compensation model for a three-dimensional fluorescence endoscope camera system based on deep learning.

[0093] Step S240: Build an image brightness compensation model for a three-dimensional fluorescence endoscope camera system based on deep learning.

[0094] Step S250: Build an image color correction model for a three-dimensional fluorescence endoscope camera system based on deep learning.

[0095] Step S260: Build a fusion model of visible light and fluorescence images of a three-dimensional fluorescence endoscope camera system based on deep learning.

[0096] Step S270: Construct an image sharpness evaluation module, using the sum of the absolute values ​​of the gray-level differences in four neighboring regions as the evaluation function. Image region sharpness evaluation function Its formula can be expressed as:

[0097]

[0098] in The coordinates are ( The pixel value of ).

[0099] Step S300: Based on the reinforcement learning environment of the three-dimensional fluorescence endoscope camera system, a multi-agent reinforcement learning model is built.

[0100] In some preferred embodiments, the multi-agent reinforcement learning model is constructed as follows:

[0101] Step S310: The reinforcement learning environment acquires the original image through a three-dimensional fluorescence endoscope camera system. After noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, the output fluorescence fusion image is used as the state of the reinforcement learning environment. .

[0102] Step S320: The focusing parameters of the three-dimensional fluorescence endoscope camera system output by the intelligent agent are used as the actions of the reinforcement learning environment. .

[0103] Step S330: Based on the reinforcement learning environment, the original image is acquired through a three-dimensional fluorescence endoscope camera system. After noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, the output fluorescence fusion image is used to construct the reward of the reinforcement learning environment. Its formula can be expressed as

[0104]

[0105] in, This is a function for evaluating image sharpness. , This is a balancing constant term used to constrain rewards to a reasonable range.

[0106] Step S340: Construct two agents, Actor (policy network) and Critic (value network), respectively. Their basic structure is a U-shaped network structure. See [link / reference]. Figure 2 The network has 6 layers.

[0107] The objective of the multi-agent reinforcement learning model is to maximize the image sharpness evaluation function. .

[0108] Step S400: The Proximal Policy Optimization (PPO) algorithm is used to optimize the image sharpness evaluation function proposed in step S340. A multi-agent reinforcement learning model is used to solve the problem, which improves the robustness of the fluorescence imaging system and evaluates its impact on the fluorescence image fusion end.

[0109] PPO is a policy optimization algorithm based on offline learning, focusing on simplifying the training process and overcoming the computational complexity of traditional policy gradient methods (such as TRPO) while ensuring training effectiveness. The objective function of the PPO algorithm is:

[0110]

[0111] in, Indicates the state Select Action Behavioral value function In strategy Next state The value function of . Indicates the state According to the distribution Sampling is performed, and the expected value of the expression is obtained. This expected value represents the expression under a given strategy. Below, for behavioral value functions and value function The operation is performed and the expected value is obtained. The optimizer used in the PPO algorithm is the AdamW optimizer, and the learning rate is set to a reasonable value (here it can be set to...). ).

[0112] In the context of fluorescence compensation systems, see Figure 3N1 (Actor) is responsible for outputting the focusing parameters of the 3D fluorescence endoscope imaging system. It continuously tries different behaviors (i.e., different focusing parameters) and adjusts its behavior strategy based on the reward signals it receives. These reward signals can be the image sharpness or other evaluation metrics under different strategies.

[0113] N2 (i.e., Critic), as an adversarial agent, counters N1's policy by perturbing the observed state, aiming to obtain a better adversarial strategy. Its task is to evaluate the value of N1's policy in a specific state, thereby providing guidance or feedback to N1.

[0114] This PPO algorithm iteratively adjusts its strategy by continuously optimizing the objective function, namely maximizing the expected reward of N1. In a fluorescence imaging system, this means that N1 will continuously learn better parfocal strategies for visible and fluorescence images in a 3D fluorescence endoscope to adapt to different external environmental disturbances and maintain a more robust system operation under similar disturbances.

[0115] In some preferred embodiments, the following steps are included:

[0116] Step S410: The reinforcement learning environment acquires the original image through a three-dimensional fluorescence endoscope camera system, and outputs a fluorescence fusion image after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion. .

[0117] Step S420: Input the fluorescence fusion image into the Actor (policy network) agent, and output the focusing parameters of the three-dimensional fluorescence endoscope imaging system. .

[0118] In step S430, the reinforcement learning environment focuses the three-dimensional fluorescence endoscope camera system according to the focusing parameters. Then, the reinforcement learning environment acquires the original visible light image and the original fluorescence image through the three-dimensional fluorescence endoscope camera system. After noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, the output fluorescence fused image is generated. The reinforcement learning environment reward is obtained by evaluating the sharpness of the fluorescence fusion image. .

[0119] Step S500, see Figure 4 The original image is acquired by the three-dimensional fluorescence endoscope camera system. After noise reduction, uniformity compensation, brightness compensation, color correction and fusion, the output fluorescence fusion image is input into the Actor (policy network) intelligent agent to obtain the focusing parameters of the three-dimensional fluorescence endoscope camera system. The focusing of the three-dimensional fluorescence endoscope camera system can be completed by using the focusing parameters.

[0120] Step S600: Set the state of the reinforcement learning environment. , and the focusing parameters of the three-dimensional fluorescence endoscope imaging system The input is fed into the self-group iterative upgrade mechanism of the multi-agent reinforcement learning model to calculate the corresponding environmental reward. The multi-agent reinforcement learning model is modified.

[0121] See Figure 5 The self-group iterative upgrade device includes a single-machine self-iterative upgrade module and a networked group iterative upgrade module. The difference between self-upgrade and group upgrade lies in whether the training information (dataset, dataset size, dataset quality, etc.) originates from a local single machine (self-upgrade) or a networked group (group upgrade). The group iterative upgrade module includes two functions: local group iterative upgrade and cloud group iterative upgrade. The difference between local group upgrade and cloud group upgrade lies in whether the training process is completed locally or in the cloud. The training device can be local, in the cloud, or both. The local group iterative upgrade includes two functions: local group single-machine training iterative upgrade and local group distributed training iterative upgrade. The local group single-machine training iterative upgrade includes two functions: local group single-machine training with shared dataset iterative upgrade (without sharing training weights) and local group single-machine training with shared dataset and weights iterative upgrade.

[0122] The self-iterative upgrade device based on deep reinforcement learning includes:

[0123] The environmental reward evaluation module is used to evaluate the state of the reinforcement learning environment. Corresponding environmental rewards .

[0124] The storage module is used to collect the reinforcement learning environment states input to the multi-agent reinforcement learning model. ;

[0125] The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model based on the data collected by the storage module;

[0126] See Figure 6 , Figure 7 The self-iterative upgrade step based on the self-iterative upgrade device involves using the fluorescence fusion image acquired by the three-dimensional fluorescence endoscope camera system, after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The input is fed into the multi-agent reinforcement learning model for training.

[0127] The deep reinforcement learning-based population iterative upgrade device includes:

[0128] The environmental reward evaluation module is used to evaluate the state of the reinforcement learning environment. Corresponding environmental rewards .

[0129] The storage module is used to collect the reinforcement learning environment states input to the multi-agent reinforcement learning model. ;

[0130] The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model based on the images collected by the storage module;

[0131] The network service module is used for communication and data transmission with the cloud control center.

[0132] See Figure 8 , Figure 9 , Figure 16 The group-based local single-machine training shared dataset iterative upgrade based on the group iterative upgrade module includes:

[0133] The training dataset comes from all networked devices;

[0134] The fluorescence fusion image, acquired by a three-dimensional fluorescence endoscope camera system and shared by several other networked devices, after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, is used as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The multi-agent reinforcement learning model is downloaded and input into the current local device via the network module, trained locally, and the corrected weights of the multi-agent reinforcement learning model are obtained for use by the current local device.

[0135] See Figure 10 , Figure 11 , Figure 16 The group-based local single-machine training shared dataset and weight iterative upgrade based on the group iterative upgrade module includes:

[0136] The training dataset comes from all networked devices;

[0137] Prerequisite: All networked devices must use the same training model software;

[0138] The fluorescence fusion image, acquired by several other networked devices through a 3D fluorescence endoscope camera system and processed through noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, is used as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The data is downloaded to the local device via the network module and input into the local device's multi-agent reinforcement learning model. The model is then trained to obtain the corrected weights of the multi-agent reinforcement learning model. The weights of the local device's multi-agent reinforcement learning model are then updated and shared in real time with all networked devices for their use.

[0139] See Figure 12 , Figure 13 , Figure 17 The group-based local distributed training iterative upgrade based on the group iterative upgrade module includes:

[0140] Training information (dataset size, dataset quality, etc.) comes from all networked devices;

[0141] Prerequisites: Either all networked devices use the same training model software (which requires knowledge of the dataset size for each device), or the quantitative differences in the parameters of the training datasets generated by each device are known (which requires knowledge of the dataset quality for each device).

[0142] The fluorescence fusion images acquired by each current local device through a 3D fluorescence endoscope camera system, after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, are used as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The input is fed into the local device's multi-agent reinforcement learning model for training, resulting in the corrected multi-agent reinforcement learning model weights. The corrected multi-agent reinforcement learning model weights from several networked devices are then uploaded to the cloud control center for global reduction (AllReduce). If all networked devices use the same training model software, global reduction is performed based on the size of each device's dataset. If the networked devices use different training model software and the quantitative differences in the dataset parameters of each device are known, global reduction is performed based on the quantitative differences in the quality of each device's dataset. The updated multi-agent reinforcement learning model weights are then obtained.

[0143] See Figure 14 , Figure 15 , Figure 17The group cloud-based iterative upgrade based on the group iterative upgrade module includes:

[0144] Prerequisite: The training model software must be identical across all networked devices and cloud-based devices;

[0145] The fluorescence fusion image, acquired by several networked devices through a three-dimensional fluorescence endoscope camera system and processed after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, is used as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The data is transmitted to the cloud control center via the network service module; the cloud control center then summarizes the received reinforcement learning environment status. , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The system trains a multi-agent reinforcement learning model and downloads and updates the weights of the current device's multi-agent reinforcement learning model synchronously using the local network service module.

[0146] A second embodiment of the present invention provides a method for confocalization of visible light and fluorescence images of a three-dimensional fluorescence endoscope based on deep reinforcement learning. The system comprises: an acquisition module and a training module.

[0147] The acquisition module is configured to acquire visible light images and fluorescence images through a three-dimensional fluorescence endoscope camera system, and output a fluorescence fusion image after noise reduction, uniformity compensation, brightness compensation, color correction and fusion.

[0148] The training method for the intelligent agent used for cofocal imaging of visible light and fluorescence images in a three-dimensional fluorescence endoscope is as follows:

[0149] The training module is configured to acquire a fluorescence fusion image after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion.

[0150] Using the fluorescence fusion images and their corresponding sharpness evaluation functions as a dataset, a multi-agent reinforcement learning model for cofocal images of visible light and fluorescence images from a three-dimensional fluorescence endoscope was constructed and trained based on the proximal policy optimization algorithm (PPO).

[0151] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process and related descriptions of the storage device and processing device described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0152] It should be noted that the fluorescence imaging system based on deep reinforcement learning for three-dimensional fluorescence endoscopy with confocal visible and fluorescence images provided in the above embodiments is only an example illustrating the division of the functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be merged into one module, or further divided into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing the various modules or steps and are not considered as an improper limitation of the present invention.

[0153] An electronic device according to a third embodiment of the present invention includes: at least one processor and a memory communicatively connected to at least one of the processors; wherein the memory stores instructions executable by the processor, the instructions being executed by the processor to implement the above-described method for achieving confocal imaging of visible light and fluorescence images of a three-dimensional fluorescence endoscope based on deep reinforcement learning.

[0154] A fourth embodiment of the present invention provides a computer-readable storage medium storing computer instructions, which are executed by the computer to implement the above-described method for achieving co-focusing of visible light and fluorescence images of a three-dimensional fluorescence endoscope based on deep reinforcement learning.

[0155] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process and related descriptions of the storage device and processing device described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0156] Those skilled in the art will recognize that the modules and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. The programs corresponding to the software modules and method steps can be placed in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art. To clearly illustrate the interchangeability of electronic hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the invention.

[0157] The following is for reference. Figure 18It shows a schematic diagram of the structure of a computer system for implementing embodiments of the methods, systems, and apparatus of the present invention. Figure 18 The server shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0158] like Figure 18 As shown, the computer system includes a Central Processing Unit (CPU) 501, which can perform various appropriate actions and processes based on programs stored in Read Only Memory (ROM) 502 or programs loaded from storage section 508 into Random Access Memory (RAM) 503. RAM 503 also stores various programs and data required for system operation. The CPU 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0159] The following components are connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. Removable media 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 510 as needed so that computer programs read from them can be installed into storage section 508 as needed.

[0160] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit (CPU) 501, it performs the functions defined in the methods of this invention. It should be noted that the computer-readable medium described above in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0161] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0162] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0163] The terms “first”, “second”, etc., are used to distinguish similar objects, not to describe or indicate a specific order or sequence.

[0164] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent in such process, method, article, or apparatus / device.

[0165] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. A method based on deep reinforcement learning for cofocalizing visible light and fluorescence images from a three-dimensional fluorescence endoscope, characterized in that, The method includes: The original image is acquired by the three-dimensional fluorescence endoscope camera system. After noise reduction, uniformity compensation, brightness compensation, color correction and fusion, the output fluorescence fusion image is input into the Actor intelligent agent to obtain the focusing parameters of the three-dimensional fluorescence endoscope camera system. The focusing of the three-dimensional fluorescence endoscope camera system can be completed by the focusing parameters. The training method for the multi-agent reinforcement learning model is as follows: Step S100: Construct a data acquisition device for visible light and fluorescence images of a three-dimensional fluorescence endoscope system. This device serves as the physical hardware environment for a multi-agent reinforcement learning environment. The device has two functions: First, it simultaneously acquires several frames of visible light images and several frames of fluorescence images through the imaging equipment in the three-dimensional fluorescence endoscope camera system. Second, it receives instruction parameters to enable the imaging equipment to automatically focus to a specified value. Step S200: Set up the reinforcement learning system software environment; Step S300: Based on the reinforcement learning environment of the three-dimensional fluorescence endoscope camera system, build a multi-agent reinforcement learning model; Step S400: The Proximal Policy Optimization (PPO) algorithm is used to optimize the image sharpness evaluation function proposed in step S300. The multi-agent reinforcement learning model is used to solve the problem, which improves the robustness of the fluorescence imaging system and evaluates its impact on the fluorescence image fusion end. After the multi-agent reinforcement learning model is trained: Step S500: The original image is acquired by the three-dimensional fluorescence endoscope camera system. After noise reduction, uniformity compensation, brightness compensation, color correction and fusion, the output fluorescence fusion image is input into the Actor (policy network) intelligent agent to obtain the focusing parameters of the three-dimensional fluorescence endoscope camera system. The focusing of the three-dimensional fluorescence endoscope camera system can be completed by the focusing parameters. Step S600: Set the state of the reinforcement learning environment. , and the focusing parameters of the three-dimensional fluorescence endoscope imaging system The input is fed into the self-group iterative upgrade mechanism of the multi-agent reinforcement learning model to calculate the corresponding environmental reward. The multi-agent reinforcement learning model is modified.

2. The method for cofocalizing visible light and fluorescence images of a three-dimensional fluorescence endoscope based on deep reinforcement learning according to claim 1, characterized in that, The calculation method for the reinforcement learning system software environment is as follows: Step S210: Build the software driver system for the three-dimensional fluorescence endoscope camera system. The system can achieve the following two main functions:

1. Receive in real time several frames of visible light images and several frames of fluorescence images simultaneously acquired by the imaging device in the three-dimensional fluorescence endoscope camera system; 2. Issue focusing commands and parameters to enable the imaging device in the three-dimensional fluorescence endoscope camera system to automatically focus to the specified parameters. Step S220: Build an image denoising model for a three-dimensional fluorescence endoscope camera system based on deep learning; Step S230: Build an image uniformity compensation model for a three-dimensional fluorescence endoscope camera system based on deep learning; Step S240: Build an image brightness compensation model for a three-dimensional fluorescence endoscope camera system based on deep learning; Step S250: Build an image color correction model for a three-dimensional fluorescence endoscope camera system based on deep learning; Step S260: Build an image fusion model of visible light and fluorescence images of a three-dimensional fluorescence endoscope camera system based on deep learning; Step S270: Construct an image sharpness evaluation module, using the sum of the absolute values ​​of the gray-level differences in four neighboring regions as the evaluation function. The image region sharpness evaluation function can be expressed as follows: in The coordinates are ( The pixel value of ).

3. The method for cofocalizing visible light and fluorescence images of a three-dimensional fluorescence endoscope based on deep reinforcement learning according to claim 1, characterized in that, The construction method of the multi-agent reinforcement learning model is as follows: Step S310: The reinforcement learning environment acquires the original image through a three-dimensional fluorescence endoscope camera system. After noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, the output fluorescence fusion image is used as the state of the reinforcement learning environment. ; Step S320: The focusing parameters of the three-dimensional fluorescence endoscope camera system output by the intelligent agent are used as the actions of the reinforcement learning environment. ; Step S330: Based on the reinforcement learning environment, the original image is acquired through a three-dimensional fluorescence endoscope camera system. After noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, the output fluorescence fusion image is used to construct the reward of the reinforcement learning environment. Its formula can be expressed as in, This is a function for evaluating image sharpness. , This is a balancing constant term used to constrain rewards to a reasonable range.

4. The multi-agent reinforcement learning model according to claim 3, characterized in that, The network structure of the two agents, Actor and Critic, is as follows: The basic structure is a U-shaped convolutional network, as shown in Figure 3, with 6 layers.

5. The method for cofocalizing visible light and fluorescence images of a three-dimensional fluorescence endoscope based on deep reinforcement learning according to claim 1, characterized in that, The training process of the multi-agent reinforcement learning model is as follows: Step S410: The reinforcement learning environment acquires the original image through a three-dimensional fluorescence endoscope camera system, and outputs a fluorescence fusion image after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion. ; Step S420: Input the fused fluorescence image into the Actor agent and output the focusing parameters of the three-dimensional fluorescence endoscope imaging system. ; In step S430, the reinforcement learning environment focuses the three-dimensional fluorescence endoscope camera system according to the focusing parameters. Then, the reinforcement learning environment acquires the original image through the three-dimensional fluorescence endoscope camera system, and outputs a fluorescence fused image after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion. The reinforcement learning environment reward is obtained by evaluating the sharpness of the fused fluorescence image. .

6. The method for cofocalizing visible light and fluorescence images of a three-dimensional fluorescence endoscope based on deep reinforcement learning according to claim 1, characterized in that, The automatic focusing method for visible light and fluorescence images in a three-dimensional fluorescence endoscope has the following process: Referring to Figure 2, the original image is acquired by the three-dimensional fluorescence endoscope camera system. After noise reduction, uniformity compensation, brightness compensation, color correction and fusion, the output fluorescence fusion image is obtained. The fused fluorescence image is input into the Actor intelligent agent to obtain the focusing parameters of the three-dimensional fluorescence endoscope camera system. The focusing of the three-dimensional fluorescence endoscope camera system can be completed by using the focusing parameters.

7. A self-group iterative upgrade device based on deep reinforcement learning, characterized in that, The self-group iterative upgrade device includes a single-machine self-iterative upgrade module and a networked group iterative upgrade module. The difference between self-upgrade and group upgrade lies in whether the training information (dataset, dataset size, dataset quality, etc.) originates from a local single machine (self-upgrade) or a networked group (group upgrade). The group iterative upgrade module includes two functions: local group iterative upgrade and cloud group iterative upgrade. The difference between local group upgrade and cloud group upgrade lies in whether the training process is completed locally or in the cloud. The training device can be local, in the cloud, or both. The local group iterative upgrade includes two functions: local group single-machine training iterative upgrade and local group distributed training iterative upgrade. The local group single-machine training iterative upgrade includes two functions: local group single-machine training with shared dataset iterative upgrade (without sharing training weights) and local group single-machine training with shared dataset and weights iterative upgrade.

8. The self-iterative upgrade module based on deep reinforcement learning according to claim 7, characterized in that, The self-iterative upgrade module based on deep reinforcement learning includes: The environmental reward evaluation module is used to evaluate the state of the reinforcement learning environment. Corresponding environmental rewards ; The storage module is used to collect the reinforcement learning environment states input to the multi-agent reinforcement learning model. ; The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model based on the data collected by the storage module; The self-iterative upgrade step based on the self-iterative upgrade device involves using the fluorescence fusion image acquired by the three-dimensional fluorescence endoscope imaging system, after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The input is fed into the multi-agent reinforcement learning model for training.

9. The population iterative upgrade module based on deep reinforcement learning according to claim 7, characterized in that, The deep reinforcement learning-based population iterative upgrade module includes: The environmental reward evaluation module is used to evaluate the state of the reinforcement learning environment. Corresponding environmental rewards ; The storage module is used to collect the reinforcement learning environment states input to the multi-agent reinforcement learning model. ; The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model based on the images collected by the storage module; The network service module is used for communication and data transmission with the cloud control center.

10. The group local single-machine training shared dataset iterative upgrade function of the group iterative upgrade module based on deep reinforcement learning according to claim 7, characterized in that, The iterative upgrade of the shared local single-machine training dataset based on the group iterative upgrade module includes: The training dataset comes from all networked devices; The fluorescence fusion image, acquired by a three-dimensional fluorescence endoscope camera system and shared by several other networked devices, after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, is used as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The multi-agent reinforcement learning model is downloaded and input into the current local device via the network module, trained locally, and the corrected weights of the multi-agent reinforcement learning model are obtained for use by the current local device.

11. The group local single-machine training shared dataset and weight iterative upgrade function of the group iterative upgrade module based on deep reinforcement learning according to claim 7, characterized in that, The group-based local single-machine training shared dataset and weight iterative upgrade based on the group iterative upgrade module includes: The training dataset comes from all networked devices; Prerequisite: All networked devices must use the same training model software; The fluorescence fusion image, acquired by several other networked devices through a 3D fluorescence endoscope camera system and processed after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, is used as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The data is downloaded to the local device via the network module and input into the local device's multi-agent reinforcement learning model. The model is then trained to obtain the corrected weights of the multi-agent reinforcement learning model. The weights of the local device's multi-agent reinforcement learning model are then updated and shared in real time with all networked devices for their use.

12. The population-local distributed training iterative upgrade function of the population iterative upgrade module based on deep reinforcement learning according to claim 7, characterized in that, The group-based local distributed training iterative upgrade based on the group iterative upgrade module includes: Training information (dataset size, dataset quality, etc.) comes from all networked devices; Prerequisites: Either all networked devices use the same training model software (which requires knowledge of the dataset size for each device), or the quantitative differences in the parameters of the training datasets generated by each device are known (which requires knowledge of the dataset quality for each device). The fluorescence fusion images acquired by each current local device through a 3D fluorescence endoscope camera system, after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, are used as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The input is fed into the local device's multi-agent reinforcement learning model for training, resulting in the corrected multi-agent reinforcement learning model weights. The corrected multi-agent reinforcement learning model weights from several networked devices are then uploaded to the cloud control center for global reduction (AllReduce). If all networked devices use the same training model software, global reduction is performed based on the size of each device's dataset. If the networked devices use different training model software and the quantitative differences in the dataset parameters of each device are known, global reduction is performed based on the quantitative differences in the quality of each device's dataset. The updated multi-agent reinforcement learning model weights are then obtained.

13. The group cloud-based iterative upgrade function of the group iterative upgrade module based on deep reinforcement learning according to claim 7, characterized in that, The group cloud-based iterative upgrade based on the group iterative upgrade module includes: Prerequisite: The training model software must be identical across all networked devices and cloud-based devices; The fluorescence fusion image, acquired by several networked devices through a three-dimensional fluorescence endoscope camera system and processed after noise reduction, uniformity compensation, brightness compensation, color correction, and fusion, is used as the state of the reinforcement learning environment. The state of the reinforcement learning environment , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The data is transmitted to the cloud control center via the network service module; the cloud control center then summarizes the received reinforcement learning environment status. , And the focusing parameters (actions) of the three-dimensional fluorescence endoscope camera system. The system trains a multi-agent reinforcement learning model and downloads and updates the weights of the current device's multi-agent reinforcement learning model synchronously using the local network service module.