Method for realizing automatic control of endoscope light source by adjusting camera parameters based on deep reinforcement learning and self-group iteration upgrading device
By using a multi-agent model based on deep reinforcement learning and a self-group iterative upgrade device, the automatic control of the endoscope light source is achieved, solving the problem of low efficiency in light source brightness control in existing technologies and improving image quality and detection speed.
Patent Information
- Application Number
- CN202511231350.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-31
- Publication Date
- 2025-11-25
AI Technical Summary
Existing endoscopic light source brightness control technologies suffer from problems such as long data processing time, high computational resource consumption, and low efficiency, resulting in poor image quality and affecting the detection of lesions.
A multi-agent reinforcement learning model based on deep reinforcement learning is adopted to achieve automatic control of the endoscope light source by adjusting camera parameters. The camera shutter and gain parameters are optimized by using Actor and Critic networks, and the model is trained and upgraded by combining a self-group iterative upgrade device to achieve real-time adjustment of the light source power.
It improves the control efficiency and accuracy of endoscopic image brightness, reduces the impact on the performance of the endoscopic system, enhances detection speed and image quality, and is suitable for the limited computing resources of embedded devices.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of endoscopy technology, specifically to a method and a self-group iterative upgrade device for automatically controlling the endoscope light source by adjusting camera parameters based on deep reinforcement learning. Background Technology
[0002] With the rapid development of endoscopic surgery and minimally invasive surgery, endoscopes, which integrate knowledge from multiple fields such as mathematics, software technology, optics, and ergonomics, have become a widely used medical instrument. Through endoscopes, doctors can directly observe the morphology of internal organs and the condition of lesions, making diagnosis more convenient. Furthermore, the observed images can be output and stored for further diagnosis and treatment; their diagnostic and therapeutic advantages are widely recognized in the medical community.
[0003] Under normal circumstances, the brightness of the endoscope light source cannot be adaptively controlled. The brightness of the light source may be too strong or too weak. In order to ensure that there are no white high-brightness areas or too dark areas on the output image, some technologies for automatically controlling the light source have been developed.
[0004] However, existing automatic light source control technologies have the following drawbacks:
[0005] (1) Some brightness control systems achieve automatic control of the light source through software processing. However, software processing generally requires the processing and calculation of a large amount of data, such as image data processing and various environmental sensor data processing. Therefore, the data processing process involves a large amount of data, long calculation time, slow processing speed, and high resource consumption. It may also involve major hardware modifications, resulting in high complexity and high cost.
[0006] (2) Another part of the brightness control system uses the binary search algorithm to realize the automatic control of the light source, but it is inefficient and slow, and requires repeated attempts. Summary of the Invention
[0007] To overcome the shortcomings of existing technologies, the present invention aims to provide a method for automatically controlling the endoscope light source by adjusting camera parameters based on deep reinforcement learning. This method addresses the technical problems of existing technologies, such as large amounts of data, long computation time, slow processing speed, and low efficiency, thereby improving image quality and enhancing the ability to detect lesions. The method includes the following steps:
[0008] Step S100: Construct an information acquisition device for the endoscope camera system as the physical hardware environment for the multi-agent reinforcement learning environment. The information acquisition device has two functions: 1. Reading the camera shutter and gain parameters in the endoscope camera system; 2. Receiving a specified power to adjust and control the endoscope light source.
[0009] Step S200: Build a reinforcement learning software environment – a software driver system for the endoscope camera system. The system can perform the following two main functions: 1. Receive camera shutter and gain parameters from the endoscope camera system in real time; 2. Issue adjustment commands and parameters to adjust the endoscope light source in the endoscope camera system to a specified power.
[0010] Step S300: Based on the endoscopic camera system reinforcement learning environment, build a multi-agent reinforcement learning model.
[0011] In some preferred embodiments, the multi-agent reinforcement learning model is constructed as follows:
[0012] Step S310: The reinforcement learning environment is acquired by the endoscopic camera system, which collects camera parameters (shutter speed). and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment .
[0013] Step S320: The preset power of the endoscopic camera system light source output by the intelligent agent is used as the action of the reinforcement learning environment. .
[0014] Step S330: Based on the reinforcement learning environment, camera parameters (shutter speed) are acquired through the endoscopic camera system. and gain ), preset shutter threshold and the preset gain threshold Rewards for building reinforcement learning environments Its formula can be expressed as
[0015]
[0016] in, It is the balance constant term. This is used to constrain rewards to a reasonable range.
[0017] In step S340, two intelligent agents, Actor (policy network) and Critic (value network), are constructed respectively, with their basic structure being a fully connected network structure.
[0018] The goal of the multi-agent reinforcement learning model is to constrain the camera parameters of the endoscopic imaging system to a reasonable range.
[0019] In step S400, the proximal policy optimization algorithm (PPO) is used to solve the multi-agent reinforcement learning model proposed in step S300.
[0020] In some preferred embodiments, the following steps are included:
[0021] Step S410: The reinforcement learning environment is acquired using an endoscopic camera system to obtain camera parameters (shutter speed). and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment .
[0022] Step S420: Set the state of the reinforcement learning environment. The input is given to the Actor (policy network) agent, and the output is the preset power of the light source for the endoscopic camera system. .
[0023] Step S430: The reinforcement learning environment adjusts the power of the light source of the endoscope camera system according to the preset power of the light source, and then the reinforcement learning environment acquires camera parameters (shutter speed) through the endoscope camera system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment Environmental rewards are calculated. .
[0024] After the multi-agent reinforcement learning model is trained.
[0025] Step S500: Camera parameters (shutter speed) are acquired through the endoscopic imaging system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment .
[0026] Step S600: Set the state of the reinforcement learning environment. The input is fed into the Actor (policy network) agent to obtain the preset power of the light source for the endoscopic camera system. .
[0027] Step S700: Preset power of the light source Adjust the power of the light source of the endoscope camera system.
[0028] Step S800: Camera parameters (shutter speed) are acquired through the endoscopic imaging system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The input is fed into the self-group iterative upgrade mechanism of the multi-agent reinforcement learning model to calculate the corresponding environmental reward. The multi-agent reinforcement learning model is modified.
[0029] The self-group iterative upgrade device includes a single-machine self-iterative upgrade module and a networked group iterative upgrade module. The difference between self-upgrade and group upgrade lies in whether the training information (dataset, dataset size, dataset quality, etc.) originates from a local single machine (self-upgrade) or a networked group (group upgrade). The group iterative upgrade module includes two functions: local group iterative upgrade and cloud group iterative upgrade. The difference between local group upgrade and cloud group upgrade lies in whether the training process is completed locally or in the cloud. The training device can be local, in the cloud, or both. The local group iterative upgrade includes two functions: local group single-machine training iterative upgrade and local group distributed training iterative upgrade. The local group single-machine training iterative upgrade includes two functions: local group single-machine training with shared dataset iterative upgrade (without sharing training weights) and local group single-machine training with shared dataset and weights iterative upgrade.
[0030] The self-iterative upgrade device based on deep reinforcement learning includes:
[0031] The environmental reward evaluation module is used to evaluate the state of the reinforcement learning environment. Corresponding environmental rewards .
[0032] The storage module is used to collect the reinforcement learning environment states input to the multi-agent reinforcement learning model. ;
[0033] The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model based on the data collected by the storage module;
[0034] The self-iterative upgrade step based on the self-iterative upgrade device is to acquire camera parameters (shutter speed) from the endoscope imaging system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The input is fed into the multi-agent reinforcement learning model for training.
[0035] The deep reinforcement learning-based population iterative upgrade device includes:
[0036] The environmental reward evaluation module is used to evaluate the state of the reinforcement learning environment. Corresponding environmental rewards .
[0037] The storage module is used to collect the reinforcement learning environment states input to the multi-agent reinforcement learning model. ;
[0038] The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model based on the images collected by the storage module;
[0039] The network service module is used for communication and data transmission with the cloud control center.
[0040] The iterative upgrade of the shared local single-machine training dataset based on the group iterative upgrade module includes:
[0041] The training dataset comes from all networked devices;
[0042] Camera parameters (shutter speed) are acquired by an endoscope camera system shared by several other networked devices. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The multi-agent reinforcement learning model is downloaded and input into the current local device via the network module, trained locally, and the corrected weights of the multi-agent reinforcement learning model are obtained for use by the current local device.
[0043] The group-based local single-machine training shared dataset and weight iterative upgrade based on the group iterative upgrade module includes:
[0044] The training dataset comes from all networked devices;
[0045] Prerequisite: All networked devices must use the same training model software;
[0046] The camera parameters (shutter speed) are acquired by the endoscopic camera system from several other networked devices. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The data is downloaded to the local device via the network module and input into the local device's multi-agent reinforcement learning model. The model is then trained to obtain the corrected weights of the multi-agent reinforcement learning model. The weights of the local device's multi-agent reinforcement learning model are then updated and shared in real time with all networked devices for their use.
[0047] The group-based local distributed training iterative upgrade based on the group iterative upgrade module includes:
[0048] Training information (dataset size, dataset quality, etc.) comes from all networked devices;
[0049] Prerequisites: Either all networked devices use the same training model software (which requires knowledge of the dataset size for each device), or the quantitative differences in the parameters of the training datasets generated by each device are known (which requires knowledge of the dataset quality for each device).
[0050] The camera parameters (shutter speed) are acquired by the endoscope camera system of each current local device. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The input is fed into the local device's multi-agent reinforcement learning model for training, resulting in the corrected multi-agent reinforcement learning model weights. The corrected multi-agent reinforcement learning model weights from several networked devices are then uploaded to the cloud control center for global reduction (AllReduce). If all networked devices use the same training model software, global reduction is performed based on the size of each device's dataset. If the networked devices use different training model software and the quantitative differences in the dataset parameters of each device are known, global reduction is performed based on the quantitative differences in the quality of each device's dataset. The updated multi-agent reinforcement learning model weights are then obtained.
[0051] The group cloud-based iterative upgrade based on the group iterative upgrade module includes:
[0052] Prerequisite: The training model software must be identical across all networked devices and cloud-based devices;
[0053] The camera parameters (shutter speed) are acquired by an endoscope camera system connected to several networks. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The data is transmitted to the cloud control center via the network service module; the cloud control center then summarizes the received reinforcement learning environment status. , and the light source power of the endoscopic camera system The system trains a multi-agent reinforcement learning model and downloads and updates the weights of the current device's multi-agent reinforcement learning model synchronously using the local network service module.
[0054] Therefore, compared with the prior art, the present invention has the following beneficial effects:
[0055] (1) The automatic control method provided by the present invention does not require hardware modification and only needs to be based on simple camera parameters, which is simple and efficient;
[0056] (2) Endoscopic systems are embedded devices with limited computing resources, making them unsuitable for complex calculations and large-scale data processing. Currently, most devices are 2K / 4K, resulting in large data volumes and slow processing speeds. This invention improves upon the algorithm by eliminating the need for image data processing and requiring only the processing of a very small portion of camera parameter data. This reduces the demands and impact on the performance of the endoscopic system, significantly improving detection speed.
[0057] (3) The present invention detects camera parameter data and the software automatically adjusts the current light source power in real time, thereby avoiding excessively high or low brightness of the endoscope image display and controlling the brightness of the endoscope image to always be within a reasonable range, which is beneficial to the accuracy of the doctor's operation.
[0058] (4) This invention is based on a deep reinforcement learning algorithm, which has high operating efficiency and only requires a limited number of steps to keep the brightness of the endoscope image within a reasonable range for a short time. Attached Figure Description
[0059] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0060] Figure 1 This is a flowchart of an embodiment of the method for automatically controlling the light source of an endoscope by adjusting camera parameters based on deep reinforcement learning according to the present invention.
[0061] Figure 2 This is a schematic diagram of the self-group iterative upgrade device based on deep reinforcement learning according to the present invention.
[0062] Figure 3 This is a schematic diagram of the self-iterative upgrade process based on deep reinforcement learning in this invention.
[0063] Figure 4 This is a logical block diagram of the self-iterative upgrade module based on deep reinforcement learning in this invention.
[0064] Figure 5 This is a schematic diagram illustrating the process of iterative upgrading of a shared dataset for group local single-machine training based on deep reinforcement learning, as described in this invention.
[0065] Figure 6 This is a logical block diagram of the group local single-machine training shared dataset iterative upgrade module based on deep reinforcement learning in this invention.
[0066] Figure 7 This is a schematic diagram of the process of group local single-machine training and weight iterative upgrade based on deep reinforcement learning in this invention.
[0067] Figure 8 This is a logical block diagram of the shared dataset and weight iterative upgrade module for group local single-machine training based on deep reinforcement learning, which is based on the present invention.
[0068] Figure 9 This is a schematic diagram of the process of iterative upgrade of population local distributed training based on deep reinforcement learning in this invention.
[0069] Figure 10 This is a logical block diagram of the population-locally distributed iterative upgrade module based on deep reinforcement learning in this invention.
[0070] Figure 11 This is a schematic diagram of the process of the group cloud-based iterative upgrade based on deep reinforcement learning in this invention.
[0071] Figure 12 This is a logical block diagram of the group cloud-based iterative upgrade module based on deep reinforcement learning in this invention.
[0072] Figure 13 This is a schematic diagram of the network structure for group-based local single-machine training and iterative upgrading based on deep reinforcement learning, as described in this invention.
[0073] Figure 14 This is a schematic diagram of the network structure of the present invention based on deep reinforcement learning, which features local distributed and cloud-based iterative upgrades.
[0074] Figure 15 This is a flowchart of an invention that describes how to automatically control the light source of an endoscope by adjusting camera parameters.
[0075] Figure 16 This is a schematic diagram of the structure of a computer system used to implement the embodiments of the methods, systems, and apparatus of the present invention.
[0076] It should be noted that, Figure 13 and Figure 14 For illustration purposes only. The number of devices is not limited to 5 or more. The number of devices for group iterative upgrade can be 2, 3 or more. Detailed Implementation
[0077] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention. It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the drawings. Unless otherwise specified, the embodiments and features described herein can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0078] See Figure 1 The method for automatic control of endoscope light source by adjusting camera parameters based on deep reinforcement learning algorithm according to the first embodiment of the present invention includes the following steps:
[0079] Step S100: Construct an information acquisition device for the endoscope camera system as the physical hardware environment for the multi-agent reinforcement learning environment. The information acquisition device has two functions: 1. Reading the camera shutter and gain parameters in the endoscope camera system; 2. Receiving a specified power to adjust and control the endoscope light source.
[0080] Step S200: Build the reinforcement learning system software environment and the software driver system for the endoscope camera system. The system can achieve the following two main functions: 1. Receive camera shutter and gain parameters from the endoscope camera system in real time; 2. Issue adjustment commands and parameters to adjust the endoscope light source in the endoscope camera system to the specified power.
[0081] Step S300: Based on the endoscopic camera system reinforcement learning environment, build a multi-agent reinforcement learning model.
[0082] In some preferred embodiments, the multi-agent reinforcement learning model is constructed as follows:
[0083] Step S310: The reinforcement learning environment is acquired by the endoscopic camera system, which collects camera parameters (shutter speed). and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment .
[0084] Step S320: The preset power of the endoscopic camera system light source output by the intelligent agent is used as the action of the reinforcement learning environment. .
[0085] Step S330: Based on the reinforcement learning environment, camera parameters (shutter speed) are acquired through the endoscopic camera system. and gain ), preset shutter threshold and the preset gain threshold Rewards for building reinforcement learning environments Its formula can be expressed as
[0086]
[0087] in, For the equilibrium constant term, This is used to constrain rewards to a reasonable range.
[0088] In step S340, two intelligent agents, Actor (policy network) and Critic (value network), are constructed respectively, with their basic structure being a fully connected network structure.
[0089] The goal of the multi-agent reinforcement learning model is to constrain the camera parameters of the endoscopic imaging system to a reasonable range.
[0090] In step S400, the proximal policy optimization algorithm (PPO) is used to solve the multi-agent reinforcement learning model proposed in step S300.
[0091] In some preferred embodiments, the following steps are included:
[0092] Step S410, the reinforcement learning environment acquires camera parameters (shutter speed) through an endoscopic camera system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment .
[0093] Step S420: Set the state of the reinforcement learning environment. The input is given to the Actor (policy network) agent, and the output is the light source power of the endoscopic camera system. .
[0094] In step S430, the reinforcement learning environment adjusts the power of the light source of the endoscopic camera system according to the power of the light source, and then the reinforcement learning environment acquires camera parameters (shutter speed) through the endoscopic camera system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment Environmental rewards are calculated. .
[0095] Step S500: Camera parameters (shutter speed) are acquired through the endoscopic imaging system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment .
[0096] Step S600: Set the state of the reinforcement learning environment. The input is fed into the Actor (policy network) agent to obtain the light source power of the endoscopic camera system. .
[0097] Step S700: Adjust the power of the light source of the endoscope camera system by adjusting the power of the light source.
[0098] Step S800: Camera parameters (shutter speed) are acquired through the endoscopic imaging system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The input is fed into the self-group iterative upgrade mechanism of the multi-agent reinforcement learning model to calculate the corresponding environmental reward. The multi-agent reinforcement learning model is modified.
[0099] See Figure 2 The self-group iterative upgrade device includes a single-machine self-iterative upgrade module and a networked group iterative upgrade module. The difference between self-upgrade and group upgrade lies in whether the training information (dataset, dataset size, dataset quality, etc.) originates from a local single machine (self-upgrade) or a networked group (group upgrade). The group iterative upgrade module includes two functions: local group iterative upgrade and cloud group iterative upgrade. The difference between local group upgrade and cloud group upgrade lies in whether the training process is completed locally or in the cloud. The training device can be local, in the cloud, or both. The local group iterative upgrade includes two functions: local group single-machine training iterative upgrade and local group distributed training iterative upgrade. The local group single-machine training iterative upgrade includes two functions: local group single-machine training with shared dataset iterative upgrade (without sharing training weights) and local group single-machine training with shared dataset and weights iterative upgrade.
[0100] The self-iterative upgrade device based on deep reinforcement learning includes:
[0101] The environmental reward evaluation module is used to evaluate the state of the reinforcement learning environment. Corresponding environmental rewards ;
[0102] The storage module is used to collect the reinforcement learning environment states input to the multi-agent reinforcement learning model. ;
[0103] The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model based on the data collected by the storage module;
[0104] See Figure 3 , Figure 4 The self-iterative upgrade step based on the self-iterative upgrade device is to acquire camera parameters (shutter speed) from the endoscope camera system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The input is fed into the multi-agent reinforcement learning model for training.
[0105] The deep reinforcement learning-based population iterative upgrade device includes:
[0106] The environmental reward evaluation module is used to evaluate the state of the reinforcement learning environment. Corresponding environmental rewards .
[0107] The storage module is used to collect the reinforcement learning environment states input to the multi-agent reinforcement learning model. ;
[0108] The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model based on the images collected by the storage module;
[0109] The network service module is used for communication and data transmission with the cloud control center.
[0110] See Figure 5 , Figure 6 , Figure 13 The group-based local single-machine training shared dataset iterative upgrade based on the group iterative upgrade module includes:
[0111] The training dataset comes from all networked devices;
[0112] Camera parameters (shutter speed) are acquired by an endoscope camera system shared by several other networked devices. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The multi-agent reinforcement learning model is downloaded and input into the current local device via the network module, trained locally, and the corrected weights of the multi-agent reinforcement learning model are obtained for use by the current local device.
[0113] See Figure 7 , Figure 8 , Figure 13 The group-based local single-machine training shared dataset and weight iterative upgrade based on the group iterative upgrade module includes:
[0114] The training dataset comes from all networked devices;
[0115] Prerequisite: All networked devices must use the same training model software;
[0116] The camera parameters (shutter speed) are acquired by the endoscopic camera system from several other networked devices. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The data is downloaded to the local device via the network module and input into the local device's multi-agent reinforcement learning model. The model is then trained to obtain the corrected weights of the multi-agent reinforcement learning model. The weights of the local device's multi-agent reinforcement learning model are then updated and shared in real time with all networked devices for their use.
[0117] See Figure 9 , Figure 10 , Figure 14 The group-based local distributed training iterative upgrade based on the group iterative upgrade module includes:
[0118] Training information (dataset size, dataset quality, etc.) comes from all networked devices;
[0119] Prerequisites: Either all networked devices use the same training model software (which requires knowledge of the dataset size for each device), or the quantitative differences in the parameters of the training datasets generated by each device are known (which requires knowledge of the dataset quality for each device).
[0120] The camera parameters (shutter speed) are acquired by the endoscope camera system of each current local device. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The input is fed into the local device's multi-agent reinforcement learning model for training, resulting in the corrected multi-agent reinforcement learning model weights. The corrected multi-agent reinforcement learning model weights from several networked devices are then uploaded to the cloud control center for global reduction (AllReduce). If all networked devices use the same training model software, global reduction is performed based on the size of each device's dataset. If the networked devices use different training model software and the quantitative differences in the dataset parameters of each device are known, global reduction is performed based on the quantitative differences in the quality of each device's dataset. The updated multi-agent reinforcement learning model weights are then obtained.
[0121] See Figure 11 , Figure 12 , Figure 14 The group cloud-based iterative upgrade based on the group iterative upgrade module includes:
[0122] Prerequisite: The training model software must be identical across all networked devices and cloud-based devices;
[0123] The camera parameters (shutter speed) are acquired by an endoscope camera system connected to several networks. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The data is transmitted to the cloud control center via the network service module; the cloud control center then summarizes the received reinforcement learning environment status. , and the light source power of the endoscopic camera system The system trains a multi-agent reinforcement learning model and downloads and updates the weights of the current device's multi-agent reinforcement learning model synchronously using the local network service module.
[0124] The control flowchart of the method for automatic control of endoscope light source by adjusting camera parameters based on a multi-agent reinforcement learning model according to the second embodiment of the present invention is shown below. Figure 15 :
[0125] Step S100: Turn on the endoscope host and power supply to initialize the endoscope camera system.
[0126] Step S200: Activate the automatic exposure mode of the endoscope camera.
[0127] Step S300: Activate the automatic control function of the endoscope light source.
[0128] Step S400, the agent (policy network) reads the state. Output light source power And adjust the light source power of the endoscope camera system to .
[0129] Step S500: Read the camera parameters shutter speed and gain of the endoscope imaging system and compare them with the preset shutter speed threshold and gain threshold.
[0130] In step S600, determine whether the real-time shutter speed is greater than the shutter threshold limit. If yes, proceed to step S400; otherwise, proceed to step S700.
[0131] In step S700, determine whether the real-time gain is greater than the upper limit of the gain threshold. If yes, proceed to step S400; otherwise, proceed to step S800.
[0132] In step S800, determine whether the real-time shutter speed is less than the lower limit of the shutter speed threshold and whether the real-time gain is less than the lower limit of the gain threshold. If yes, proceed to step S400; otherwise, proceed to step S900.
[0133] Step S900: End this automatic control of the endoscope light source.
[0134] The third embodiment of the present invention provides an automatic control system for adjusting camera parameters to realize the endoscope light source based on a multi-agent reinforcement learning model. The system is based on a method for adjusting camera parameters to realize the automatic control of the endoscope light source using deep reinforcement learning. The system includes: an acquisition module and a training module.
[0135] The acquisition module is configured to acquire camera parameters (shutter speed) from the endoscopic camera system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The power of the light source in the endoscopic camera system is obtained by inputting it into a pre-trained multi-agent reinforcement learning model. .
[0136] The training method for the multi-agent reinforcement learning model is as follows:
[0137] The training module is configured to acquire the light source power of the endoscopic camera system. ;
[0138] By building an endoscope camera system information acquisition device and a software-driven system as a complete reinforcement learning environment, a multi-agent reinforcement learning model is constructed and trained.
[0139] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process and related descriptions of the storage device and processing device described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0140] It should be noted that the above embodiments of the automatic control system for adjusting camera parameters to realize the endoscope light source based on a multi-agent reinforcement learning model are only illustrative examples of the above functional module divisions. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be merged into one module, or further divided into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing the various modules or steps and are not considered as an improper limitation of the present invention.
[0141] An electronic device according to a fourth embodiment of the present invention includes: at least one processor and a memory communicatively connected to at least one of the processors; wherein the memory stores instructions executable by the processor, the instructions being executed by the processor to implement the above-described method for automatic control of endoscopic light source by adjusting camera parameters based on deep reinforcement learning.
[0142] A computer-readable storage medium according to a fifth embodiment of the present invention stores computer instructions, which are executed by the computer to implement the above-described method for automatic control of endoscope light source by adjusting camera parameters based on deep reinforcement learning.
[0143] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process and related descriptions of the storage device and processing device described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0144] Those skilled in the art will recognize that the modules and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. The programs corresponding to the software modules and method steps can be placed in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art. To clearly illustrate the interchangeability of electronic hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the invention.
[0145] The following is for reference. Figure 16It shows a schematic diagram of the structure of a computer system for implementing embodiments of the methods, systems, and apparatus of the present invention. Figure 16 The server shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0146] like Figure 16 As shown, the computer system includes a Central Processing Unit (CPU) 501, which can perform various appropriate actions and processes based on programs stored in Read Only Memory (ROM) 502 or programs loaded from storage section 508 into Random Access Memory (RAM) 503. RAM 503 also stores various programs and data required for system operation. The CPU 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0147] The following components are connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. Removable media 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 510 as needed so that computer programs read from them can be installed into storage section 508 as needed.
[0148] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit (CPU) 501, it performs the functions defined in the methods of this invention. It should be noted that the computer-readable medium described above in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0149] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0151] The terms “first”, “second”, etc., are used to distinguish similar objects, not to describe or indicate a specific order or sequence.
[0152] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent in such process, method, article, or apparatus / device.
[0153] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A method for automatically controlling an endoscope light source by adjusting camera parameters based on deep reinforcement learning, characterized in that, The method includes: Step S100: Build an information acquisition device for the endoscope camera system as the physical hardware environment for the multi-agent reinforcement learning environment. The information acquisition device has two functions:
1. Read the camera shutter and gain parameters in the endoscope camera system; 2. Receive specified power to adjust and control the endoscope light source. Step S200: Build a reinforcement learning software environment—the software driver system of the endoscope camera system. The system can achieve the following two main functions:
1. Receive camera shutter and gain parameters from the endoscope camera system in real time; 2. Issue adjustment commands and parameters to adjust the endoscope light source in the endoscope camera system to the specified power. Step S300: Based on the endoscopic camera system reinforcement learning environment, build a multi-agent reinforcement learning model; Step S400: The Proximal Policy Optimization (PPO) algorithm is used to solve the multi-agent reinforcement learning model proposed in step S300. Step S500: Camera parameters (shutter speed) are acquired through the endoscopic imaging system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment ; Step S600: Set the state of the reinforcement learning environment. The input is given to the Actor agent to obtain the preset power of the light source for the endoscopic camera system. ; Step S700: Adjust the power of the light source of the endoscope camera system by means of the preset power of the light source; Step S800: Camera parameters (shutter speed) are acquired through the endoscopic imaging system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The input is fed into the self-group iterative upgrade mechanism of the multi-agent reinforcement learning model to calculate the corresponding environmental reward. The multi-agent reinforcement learning model is modified.
2. The method for automatic control of endoscope light source by adjusting camera parameters based on deep reinforcement learning according to claim 1, characterized in that, The construction method of the multi-agent reinforcement learning model is as follows: Step S310: The reinforcement learning environment is acquired by the endoscopic camera system, which collects camera parameters (shutter speed). and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment ; Step S320: The preset power of the endoscopic camera system light source output by the intelligent agent is used as the action of the reinforcement learning environment. ; Step S330: Based on the reinforcement learning environment, camera parameters (shutter speed) are acquired through the endoscopic camera system. and gain ), preset shutter threshold and the preset gain threshold Rewards for building reinforcement learning environments Its formula can be expressed as in, For the equilibrium constant term, This is used to constrain rewards to a reasonable range.
3. The method for automatic control of endoscope light source by adjusting camera parameters based on deep reinforcement learning according to claim 1, characterized in that, The training method for the multi-agent reinforcement learning model is as follows: Step S410, the reinforcement learning environment acquires camera parameters (shutter speed) through an endoscopic camera system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment ; Step S420: Set the state of the reinforcement learning environment. The input is given to the Actor intelligent agent, and the output is the preset power of the light source of the endoscopic camera system. ; In step S430, the reinforcement learning environment adjusts the power of the light source of the endoscope camera system according to the preset power of the light source, and then the reinforcement learning environment acquires camera parameters (shutter speed) through the endoscope camera system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment Environmental rewards are calculated. .
4. The method for automatic control of endoscope light source by adjusting camera parameters based on deep reinforcement learning according to claim 1, characterized in that, The control flow is as follows: Step S100: Turn on the endoscope host and power supply to initialize the endoscope camera system; Step S200: Activate the automatic exposure mode of the endoscope camera; Step S300: Activate the automatic control function of the endoscope light source; Step S400, the intelligent agent (Actor) reads the state. Output light source preset power And adjust the preset power of the light source of the endoscope camera system to ; Step S500: Read the camera parameters shutter speed and gain of the endoscope imaging system and compare them with the preset shutter speed threshold and gain threshold. Step S600: Determine whether the real-time shutter speed is greater than the shutter threshold limit. If yes, proceed to step S400; otherwise, proceed to step S700. Step S700: Determine whether the real-time gain is greater than the upper limit of the gain threshold. If yes, proceed to step S400; otherwise, proceed to step S800. In step S800, determine whether the real-time shutter speed is less than the lower limit of the shutter speed threshold and whether the real-time gain is less than the lower limit of the gain threshold. If yes, proceed to step S400; otherwise, proceed to step S900. Step S900: End the automatic control process of the endoscope light source.
5. A self-group iterative upgrade device based on deep reinforcement learning, characterized in that, The self-group iterative upgrade device includes a single-machine self-iterative upgrade module and a networked group iterative upgrade module. The difference between self-upgrade and group upgrade lies in whether the training information (dataset, dataset size, dataset quality, etc.) originates from a local single machine (self-upgrade) or a networked group (group upgrade). The group iterative upgrade module includes two functions: local group iterative upgrade and cloud group iterative upgrade. The difference between local group upgrade and cloud group upgrade lies in whether the training process is completed locally or in the cloud. The training device can be local, in the cloud, or both. The local group iterative upgrade includes two functions: local group single-machine training iterative upgrade and local group distributed training iterative upgrade. The local group single-machine training iterative upgrade includes two functions: local group single-machine training with shared dataset iterative upgrade (without sharing training weights) and local group single-machine training with shared dataset and weights iterative upgrade.
6. The self-iterative upgrade module based on deep reinforcement learning according to claim 5, characterized in that, The self-iterative upgrade module based on deep reinforcement learning includes: The environmental reward evaluation module is used to evaluate the state of the reinforcement learning environment. Corresponding environmental rewards ; The storage module is used to collect the reinforcement learning environment states input to the multi-agent reinforcement learning model. ; The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model based on the data collected by the storage module; The self-iterative upgrade step based on the self-iterative upgrade device is to acquire camera parameters (shutter speed) from the endoscope imaging system. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The input is fed into the multi-agent reinforcement learning model for training.
7. The population iterative upgrade module based on deep reinforcement learning according to claim 5, characterized in that, The deep reinforcement learning-based population iterative upgrade module includes: The environmental reward evaluation module is used to evaluate the state of the reinforcement learning environment. Corresponding environmental rewards ; The storage module is used to collect the reinforcement learning environment states input to the multi-agent reinforcement learning model. ; The neural network fine-tuning module is used to correct and fine-tune the multi-agent reinforcement learning model based on the images collected by the storage module; The network service module is used for communication and data transmission with the cloud control center.
8. The group local single-machine training shared dataset iterative upgrade function of the group iterative upgrade module based on deep reinforcement learning according to claim 5, characterized in that, The iterative upgrade of the shared local single-machine training dataset based on the group iterative upgrade module includes: The training dataset comes from all networked devices; Camera parameters (shutter speed) are acquired by an endoscope camera system shared by several other networked devices. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The multi-agent reinforcement learning model is downloaded and input into the current local device via the network module, trained locally, and the corrected weights of the multi-agent reinforcement learning model are obtained for use by the current local device.
9. The group local single-machine training shared dataset and weight iterative upgrade function of the group iterative upgrade module based on deep reinforcement learning according to claim 5, characterized in that, The group-based local single-machine training shared dataset and weight iterative upgrade based on the group iterative upgrade module includes: The training dataset comes from all networked devices; Prerequisite: All networked devices must use the same training model software; The camera parameters (shutter speed) are acquired by the endoscopic camera system from several other networked devices. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The data is downloaded to the local device via the network module and input into the local device's multi-agent reinforcement learning model. The model is then trained to obtain the corrected weights of the multi-agent reinforcement learning model. The weights of the local device's multi-agent reinforcement learning model are then updated and shared in real time with all networked devices for their use.
10. The population-local distributed training iterative upgrade function of the population iterative upgrade module based on deep reinforcement learning according to claim 5, characterized in that, The group-based local distributed training iterative upgrade based on the group iterative upgrade module includes: Training information (dataset size, dataset quality, etc.) comes from all networked devices; Prerequisites: Either all networked devices use the same training model software (which requires knowledge of the dataset size for each device), or the quantitative differences in the parameters of the training datasets generated by each device are known (which requires knowledge of the dataset quality for each device). The camera parameters (shutter speed) are acquired by the endoscope camera system of each current local device. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The input is fed into the local device's multi-agent reinforcement learning model for training, resulting in the corrected multi-agent reinforcement learning model weights. The corrected multi-agent reinforcement learning model weights from several networked devices are then uploaded to the cloud control center for global reduction (AllReduce). If all networked devices use the same training model software, global reduction is performed based on the size of each device's dataset. If the networked devices use different training model software and the quantitative differences in the dataset parameters of each device are known, global reduction is performed based on the quantitative differences in the quality of each device's dataset. The updated multi-agent reinforcement learning model weights are then obtained.
11. The group cloud-based iterative upgrade function of the group iterative upgrade module based on deep reinforcement learning according to claim 5, characterized in that, The group cloud-based iterative upgrade based on the group iterative upgrade module includes: Prerequisite: The training model software must be identical across all networked devices and cloud-based devices; The camera parameters (shutter speed) are acquired by an endoscope camera system connected to several networks. and gain ), preset shutter threshold and the preset gain threshold As the state of the reinforcement learning environment The state of the reinforcement learning environment , and the light source power of the endoscopic camera system The data is transmitted to the cloud control center via the network service module; the cloud control center then summarizes the received reinforcement learning environment status. , and the light source power of the endoscopic camera system The system trains a multi-agent reinforcement learning model and downloads and updates the weights of the current device's multi-agent reinforcement learning model synchronously using the local network service module.