Image dynamic acquisition method based on imaging parameter optimization control
Through deep fusion of reinforcement learning and multimodal perception technology, a multi-domain joint optimization framework for space-time frequency is built, and imaging parameters are adaptively adjusted, solving the dynamic range bottleneck and motion blur problems of spatial-level sensors in high dynamic range, high-speed motion and microtexture imaging tasks, achieving efficient and accurate image acquisition, and improving imaging quality and system adaptability.
Patent Information
- Application Number
- CN202510413770.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
Existing spatial-level sensors have problems such as dynamic range bottlenecks, motion blur and dark area signal-to-noise ratio deterioration in high dynamic range, high-speed motion and microtexture imaging tasks, resulting in coexistence of high-light saturation, dark area noise change, and defocus blur in imaging results, seriously hindering fragment morphology reconstruction and orbital parameter inversion.
Using an image dynamic acquisition method based on deep fusion reinforcement learning and multimodal perception technology, the agent adaptively adjusts the imaging parameters under the framework of multi-domain joint optimization of space-time frequency, and uses the cross-modal interaction between historical fusion images and real-time environmental information to construct a comprehensive state representation including scene radiation characteristics, motion state and environmental parameters.
It effectively solves the problem of dynamic range expansion and motion artifact coupling in high-speed motion target imaging, significantly suppresses high-light saturation and dark area noise, ensures the integrity and consistency of image features under complex lighting conditions, and improves the scene adaptability of the imaging system.
Smart Images

Figure CN119942042A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of space debris detection and deep learning, and in particular to a method for dynamic image acquisition based on imaging parameter optimization control. Background Art
[0002] In the space debris monitoring scenario, the radiation brightness of the target object from the strong sunlight area to the shadow area of the earth can span up to 280dB, and it is accompanied by the compound challenges of high-speed motion and micro-scale structure. The current optoelectronic imaging system needs to meet multi-dimensional imaging indicators such as high dynamic range, motion blur suppression and micro-texture fidelity in a single frame, while commercial space-grade sensors are limited by hardware physical constraints: the charge storage capacity limitation leads to a 60dB dynamic range bottleneck, the shutter delay and readout noise cause motion blur and dark area signal-to-noise ratio degradation, and the fixed anti-aliasing filter causes high-frequency details to be lost. The coupling effect of such defects makes the imaging results show the degradation characteristics of high light saturation, dark area noise change, and out-of-focus blur, which seriously hinders the reconstruction of debris morphology and orbit parameter inversion.
[0003] The automatic exposure control method based on deep reinforcement learning proposed in Reference 1 achieves rapid convergence by adjusting the exposure time and gain through the intelligent agent. However, this method still has the following shortcomings:
[0004] First, this method only optimizes the exposure parameters (exposure time and gain) and does not consider the coordinated optimization of other imaging parameters such as aperture, which limits the applicability of the system in complex imaging tasks. Secondly, this method focuses on the optimal exposure adjustment of a single frame image and does not evaluate the adequacy of cross-frame information acquisition, which may lead to a lack of complementarity in the temporal or spatial domain of the image, affecting subsequent high-level visual tasks. Finally, its training environment relies on an LED-controlled darkroom. Although random cross-domain enhancement is used to enhance generalization capabilities, it may still perform unstably in complex real-world scenarios. Reference 1 is as follows:
[0005] "Lee K, Shin U, Lee B U. Learning to Control Camera Exposure via Reinforcement Learning[C] / / Proceedings of the IEEE / CVF Conference on ComputerVision and Pattern Recognition. 2024: 2975-2983." Summary of the invention
[0006] The purpose of the present invention is to provide an image dynamic acquisition method based on imaging parameter optimization control, which is used to jointly optimize multiple imaging parameters in the imaging control process, make full use of cross-frame information, and have stronger generalization ability.
[0007] In order to achieve the above tasks, the present invention adopts the following technical solutions:
[0008] A method for dynamic image acquisition based on imaging parameter optimization control, comprising:
[0009] The intelligent agent obtains the current moment image and the fused image of the previous moment by executing the imaging parameters of the current moment, and constructs the state value of the current moment by using the current moment image and the fused image; the imaging parameters and the state value are the values of the action space and the state space at the current moment respectively;
[0010] The feature extraction network is used to extract the features of the current state value to obtain the state features of the current state;
[0011] Based on the state characteristics of the current moment, the agent uses the strategy network to interact with the environment to obtain the state characteristics of the next moment;
[0012] Setting a dual value network, and training the dual value network based on the strategy network and the state characteristics of the next moment;
[0013] After each training of the dual value network training process, the strategy network is trained using a preset loss function; after the dual value network and the strategy network are trained, the trained strategy network is saved for optimization control during the dynamic acquisition of the intelligent body image.
[0014] Furthermore, the reward function used by the dual value network during training includes a metric based on space-time frequency complementarity and a metric based on entropy; in the metric based on space-time frequency complementarity, the space-time domain complementarity is calculated by setting a threshold and an intersection operation, and the amplitude complementarity and phase complementarity are determined by using the amplitude spectrum and phase of the current image and the corresponding historical fusion image, respectively, and then the time-frequency domain complementarity is obtained by weighting; in the metric based on entropy, the information entropy of the historical fusion image is determined; finally, the space-time domain complementarity, the time-frequency domain complementarity and the information entropy are weightedly combined to obtain the reward function.
[0015] Furthermore, the agent obtains the current moment image and the fused image of the previous moment by executing the imaging parameters of the current moment, and constructs the current moment state value using the current moment image and the fused image, including:
[0016] The agent executes the imaging parameters at the current moment , image the environment to obtain the current image ;
[0017] Read the fused image generated at the last moment from the agent's storage module ; The fused image generated at the last moment With the current moment image According to the preset image fusion rules, the fusion is performed to obtain the historical fusion image at the current moment ;
[0018] The agent obtains the state parameters of the current moment, and uses the state parameters and the historical fusion image Construct the state space and obtain the current state value based on the state space .
[0019] Furthermore, the action space is expressed as:
[0020] ;
[0021] in, is the action space, Indicates the exposure time, Indicates the gain value, Indicates the aperture size, Indicates the direction of the turntable;
[0022] The state space is represented as:
[0023] ;
[0024] in, is the state space, Represents the historical fusion image at the current moment, Represents the current moment image, Indicates the current trajectory information of the agent, Indicates the light intensity, Represents the agent posture information.
[0025] Furthermore, the feature extraction network is used to extract features of the state value at the current moment to obtain the state features at the current moment, including:
[0026] Merge the history of the current moment into the image With the current moment image Stitching along the channel dimension to get the stitched image ;
[0027] Stitching images Input to the feature extraction network to extract multi-scale semantic information and obtain the intermediate feature map , which is then flattened into image features in vector form ;
[0028] The current state parameters are concatenated into a vector and then input into the multi-layer perceptron for dimension mapping to obtain the vector feature. ;
[0029] Through a layer of linear transformation or direct alignment of image features and vector features The dimension is then concatenated on the sequence dimension or channel dimension to obtain the input feature ;
[0030] Input features Input into the transform encoder, through the processing of multi-head self-attention and feedforward network, the final output result is the state feature of the current moment .
[0031] Furthermore, based on the state characteristics of the current moment, the agent uses the strategy network to interact with the environment to obtain the state characteristics of the next moment, including:
[0032] The current state characteristics Input to the policy network Get the value of the next action space ;
[0033] The agent takes the value of the action space at the next moment As imaging parameters and execute, record the state value of the state space at the next moment after execution ;
[0034] The state value at the next moment Input into the feature extraction network to obtain the state features of the next moment .
[0035] Furthermore, when the dual value network is trained, the target reward value is calculated as follows:
[0036] ;
[0037] in, is the target reward value at the current time t, The reward function is preset The calculated reward value; is the decay hyperparameter; represents the jth value network; represents the policy network, is a learnable parameter;
[0038] The target reward value obtained in the above formula Dual Value Network The output is compared to minimize the preset error loss function, which is expressed as:
[0039] ;
[0040] in, represents the error loss function, They respectively represent the state characteristics and imaging parameters at the current moment.
[0041] Furthermore, the reward function construction process includes:
[0042] Computing complementarity in space and time :
[0043] ;
[0044] in, ,image The current moment image or a historical fusion image of the current moment ; is a binary mask, Representing images At pixel position The pixel value of is the preset low threshold, is the preset high threshold;
[0045] Computing time-frequency domain complementarity :
[0046] ;
[0047] in, and is the weight coefficient; is amplitude complementarity, is phase complementarity;
[0048] Then the reward function It is expressed as:
[0049] ;
[0050] in, , , is the weight coefficient, It is a historical fusion image Information entropy.
[0051] Furthermore, after each training of the dual value network training process, based on the following preset loss function Training the policy network :
[0052] ;
[0053] in, is the preset loss function, Indicates the state characteristics of the current moment, represents the jth value network, Indicates in status Strategy Network The value of the output action space.
[0054] Furthermore, the trained policy network is deployed to the agent to be optimized, and the agent obtains the current image by executing the imaging parameters at the current moment. And the fused image of the previous moment , combined with the state parameters of the current agent to construct the current state value , input it into the feature extraction network to obtain the state features of the current moment , the state characteristics Enter the policy network , through the policy network Output the value of the next action space As the imaging parameter of the next moment, it is used to optimize the intelligent agent.
[0055] A terminal device comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the image dynamic acquisition method based on imaging parameter optimization control is implemented.
[0056] A computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the image dynamic acquisition method based on imaging parameter optimization control is implemented.
[0057] Compared with the prior art, the present invention has the following technical features:
[0058] 1. The present invention constructs a closed-loop optimization system for dynamic imaging parameters with physical interpretability through deep fusion of reinforcement learning and multimodal sensing technology. Aiming at the core bottleneck of limited dynamic range of single-frame imaging in extreme radiation span scenarios, the imaging parameter control is innovatively modeled as a timing decision problem, and the reinforcement learning algorithm is used to achieve adaptive adjustment under the framework of multi-domain joint optimization of space, time and frequency. This method breaks through the traditional passive processing paradigm, through cross-modal interaction of historical fusion images and real-time environmental information, combined with a deep feature extraction network to construct a comprehensive state representation including scene radiation characteristics, motion state and environmental parameters, effectively solving the problem of dynamic range expansion and motion artifact coupling in high-speed moving target imaging.
[0059] 2. The reward mechanism designed based on the complementarity of multi-domain information guides the intelligent agent to achieve the global optimal balance in the control dimensions such as exposure time and gain parameters, significantly suppressing the saturation of highlights and the noise in dark areas, while ensuring the integrity and consistency of image features under complex lighting conditions. Without the need for hardware modification, the scene adaptability of the imaging system is improved to a new dimension, providing a breakthrough solution for extreme dynamic scenes such as space debris monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 A schematic diagram of a method in one embodiment of the present invention;
[0061] Figure 2 A schematic diagram of the application process of the present invention deployed on an imaging system;
[0062] Figure 3 A group of images acquired under different imaging parameters in one embodiment of the present invention;
[0063] Figure 4 This is an image obtained after optimizing the imaging system in one embodiment of the present invention. DETAILED DESCRIPTION
[0064] The present invention models the imaging parameter optimization control process as a sequential decision problem and uses reinforcement learning to solve it. In addition, in the design of the reward function, the present invention fully considers the complementarity of the multi-domain information of the photographed object, uses the multi-domain information of the observed image to evaluate the image quality, and guides the intelligent agent to autonomously optimize the imaging parameters.
[0065] The present invention provides a method for dynamically acquiring images based on imaging parameter optimization control, see Figure 1 , including the following steps:
[0066] Step 1: The intelligent agent obtains the current moment image and the fused image of the previous moment by executing the imaging parameters of the current moment, and uses the current moment image and the fused image to construct the current moment state value; the imaging parameters and state value are the values of the action space and state space at the current moment, respectively.
[0067] (1.1) The agent executes the imaging parameters at the current moment , image the environment to obtain the current image ; Imaging parameters Including exposure time , gain value , aperture size , Turntable direction ; In this solution, the intelligent agent is an imaging system used for space debris detection.
[0068] In order to realize intelligent imaging parameter optimization control, the intelligent agent needs to perceive the environmental state and perform corresponding actions, so it is necessary to reasonably define the state space and action space; the action space contains all optimizable imaging parameters to ensure that the control strategy has sufficient adjustment freedom.
[0069] In this scheme, the action space It is expressed as:
[0070] ;
[0071] The imaging parameters at the current moment are For action space The value of t at the current time.
[0072] (1.2) Read the fused image generated at the last moment from the storage module of the intelligent agent ; The fused image generated at the last moment With the current moment image According to the preset image fusion rules, the fusion is performed to obtain the historical fusion image at the current moment If the fused image generated at the last moment does not exist in the storage module during initialization, then set It is a fusion of two images taken in the default short exposure 1ms and long exposure 10ms shooting modes.
[0073] In this solution, the preset image fusion rule is expressed as:
[0074] ;
[0075] In the above formula, Indicates the process of fusion using the Mertens algorithm.
[0076] (1.3) The intelligent agent obtains the current state parameters through the imaging sensor and constructs the state space ; Wherein, the state parameters include the current trajectory information of the intelligent body , light intensity , Agent posture information .
[0077] The state space is used to abstractly represent the current shooting environment and related information to ensure that the agent uses sufficient information to make decisions; the state space is a representation of the current shooting environment and information. In order to fully consider the shooting environment information, the state space The definition is as follows:
[0078] ;
[0079] The current state value is is the state space The value of t at the current time.
[0080] Step 2: Use the feature extraction network to extract features of the current state value to obtain the current state features.
[0081] (2.1) Historical fusion image at the current moment With the current moment image Stitching along the channel dimension to get the stitched image :
[0082] ;
[0083] in is the number of input channels, and Represent the previous moment image The height and width of represents the real number space. In one embodiment of the present invention, .
[0084] (2.2) stitch the images Input to feature extraction network Extract multi-scale semantic information and obtain intermediate feature maps , which is then flattened into image features in vector form ; The feature extraction network Using convolutional network:
[0085] ;
[0086] ;
[0087] in represents convolutional network processing, For the flattening operation, is the intermediate feature map after the convolutional network The channel dimension, and They respectively represent the height and width of the features obtained after the image is processed by the convolutional network.
[0088] In one embodiment of the present invention, the convolutional network The number of layers is 4, and the size of the first convolution kernel is , the step size is 4, the padding is 3, and the output channel is 64; the convolution kernel size of the last three layers is , stride is 2, padding is 1, and output channels are 128.
[0089] (2.3) The current state parameters, including orbit information , light intensity , Agent posture information After being concatenated into vectors, they are input into a multi-layer perceptron for dimensional mapping to obtain vector features. :
[0090] ;
[0091] in, Represents the multi-layer perceptron processing process, Represents the dimension after vector feature mapping, and in one embodiment of the present invention, its value is 128.
[0092] (2.4) Through a layer of linear transformation or direct alignment of image features and vector features The dimension is then concatenated on the sequence dimension or channel dimension to form the input features that can be processed by the transform encoder. :
[0093] ;
[0094] in, Represents a splicing operation, represents a linear transformation, Represents the number of feature markers after splicing, and the image features can be regarded as feature markers, the vector feature can be regarded as a feature marker; in one embodiment of the present invention, The value of is 65.
[0095] (2.5) Input features Input into the transform encoder, through the processing of multi-head self-attention and feedforward network, fully integrate the image and non-image features and output the final result as the state feature of the current moment :
[0096] ;
[0097] in, Represents the processing of the transform encoder.
[0098] Step 3: Based on the state characteristics of the current moment, the agent uses the strategy network to interact with the environment to obtain the state characteristics of the next moment.
[0099] (3.1) The state characteristics of the current moment Input to the policy network Get the value of the next action space :
[0100] ;
[0101] (3.2) The agent will Action space values As imaging parameters and execute, record the state value of the state space at the next moment after execution .
[0102] (3.3) The state value at the next moment Input to the feature extraction network , get the state characteristics of the next moment .
[0103] Step 4, setting up a dual value network, and training the dual value network based on the strategy network and the state characteristics of the next moment; wherein the reward function used in the training includes a measure based on time-space-frequency complementarity and a measure based on entropy; in the measure based on time-space-frequency complementarity, the time-space domain complementarity is calculated by setting a threshold and an intersection operation, and the amplitude complementarity and phase complementarity are determined by using the amplitude spectrum and phase of the current moment image and the corresponding historical fusion image, respectively, and then the time-frequency domain complementarity is obtained by weighting; in the measure based on entropy, the information entropy of the historical fusion image is determined; finally, the time-space domain complementarity, the time-frequency domain complementarity and the information entropy are weightedly combined to obtain the reward function.
[0104] (4.1) Training of dual value networks.
[0105] Dual Value Network The input is the state feature at a certain moment and the value of the corresponding action space, and the output is the estimated target reward value; the use of two value networks is to alleviate the overestimation in reinforcement learning, and the minimum value of the dual value network can be used to calculate the target reward value; among them, the action prediction value at the next moment By Strategy Network Get; the target reward value is calculated as follows:
[0106] ;
[0107] in, is the target reward value at the current time t, The reward function is preset The calculated reward value; is the decay hyperparameter; represents the jth value network; represents the policy network, is a learnable parameter;
[0108] The target reward value obtained in the above formula Dual Value Network The output is compared to minimize the mean square error loss function. The process can be expressed as:
[0109] ;
[0110] in, represents the error loss function, They respectively represent the state characteristics and imaging parameters (values of the action space) at the current moment.
[0111] (4.2) Preset reward function .
[0112] In this scheme, the reward function is defined It is used to evaluate the effect of each action to obtain a reward signal, and optimize the decision-making strategy of the deep reinforcement learning model based on the obtained reward signal. The specific design process is as follows:
[0113] (a) Measurement based on space-time frequency complementarity.
[0114] The present invention adopts a complementary reward mechanism to optimize the exposure parameter adjustment during the image acquisition process. Specifically, for the shooting result of each frame of image, the contribution is evaluated by calculating the complementarity of the current image and the image fused with historical information, thereby guiding the reinforcement learning model to dynamically update the exposure parameters to maximize the complementarity of multi-frame image information.
[0115] Among them, the complementarity in time and space domains is as follows:
[0116] The present invention evaluates image complementarity by setting thresholds and intersection operations, aiming to calculate the spatiotemporal complementarity between images by dividing the effective and invalid areas of the image, thereby providing a reward signal for the optimization of exposure parameters. This method can effectively exclude background noise and overexposed areas, ensuring the diversity and effectiveness of image information. Specifically, a low threshold is set. and high threshold They are used to distinguish between valid areas and invalid areas, and are defined as follows:
[0117] ;
[0118] in, Representing images At pixel position The pixel value of is a binary mask representing the image effective area.
[0119] Computing complementarity in space and time It is used to quantify the spatial difference between the current image and the image fused with historical information. The specific formula is as follows:
[0120] ;
[0121] Among them, the complementarity in time and frequency domain is as follows:
[0122] Fourier transform converts images from the spatial domain to the frequency domain, which can reveal the different frequency components of the image. This method can effectively capture the frequency differences between images, thereby optimizing the exposure parameters of the image and improving the diversity of the image sequence.
[0123] In order to measure the time-frequency domain complementarity between images, the current image and the historical fusion image of the current moment The amplitude spectra after Fourier transformation are and ,in Represents the coordinates of the image in the frequency domain, amplitude complementarity It is expressed as:
[0124] ;
[0125] Current time image and the historical fusion image of the current moment The phases after Fourier transform are and , phase complementarity It is expressed as:
[0126] ;
[0127] By combining amplitude complementarity and phase complementarity in a weighted manner, we obtain complementarity in the time-frequency domain:
[0128] ;
[0129] in, , are weight coefficients used to control the contribution of amplitude and phase to the total complementarity.
[0130] (b) Entropy-based metrics.
[0131] Fusion of images through the history of the current moment Evaluate the overall information content of the entire image sequence to ensure that each frame of the image provides maximum unique information and avoid redundant acquisition. The information entropy of the historical fusion image is calculated using the entropy formula:
[0132] ;
[0133] in, Fusion images for history Pixel value The normalized histogram over the entire image, It is a historical fusion image The information entropy reflects the complexity and amount of information of the entire image sequence.
[0134] The present invention can comprehensively evaluate the overall information content of the image sequence. This method avoids the limitation of single image entropy calculation, so that each new image not only considers its own information content, but also comprehensively considers its contribution to the entire sequence information, thereby more effectively controlling the image acquisition process.
[0135] Based on the above metrics, the present invention designs the following comprehensive reward function:
[0136] ;
[0137] in, , , is a weight coefficient used to control the contribution of different metrics; in one embodiment of the present invention, the values are 0.3, 0.3 and 0.4 respectively.
[0138] By integrating the reward function of time-space-frequency complementarity and information entropy, the present invention can intelligently evaluate the quality of image sequences and optimize the acquisition of each frame of image. This method can not only improve the diversity and information content of image sequences, but also ensure the efficiency and accuracy of the image acquisition process, thereby maximizing the image quality. In addition, for the convenience of subsequent discussion, Time passes through the reward function The calculated value is recorded as .
[0139] Step 5, after each training of the dual value network training process, the policy network is trained using a preset loss function; after the dual value network and the policy network are trained, the trained policy network is saved for optimization control during the dynamic acquisition of the intelligent body image.
[0140] (5.1) Policy network training.
[0141] After each training of the dual value network training process, based on the following preset loss function Training the policy network :
[0142] ;
[0143] in, is the preset loss function, Indicates the state characteristics of the current moment, represents the jth value network, Indicates in status Strategy Network The value of the output action space.
[0144] During the training process, the dual value network training in step 4 and the policy network training in step 5 are performed in a loop:
[0145] The agent first obtains the current state value , and then update the policy network Dual Value Network , after updating the network; the agent obtains the value of the next action space through the updated strategy network , that is, the imaging parameters of the next moment are obtained; the agent uses the imaging parameters to obtain the state value, and then repeat the above process; when the loop iterates to the predetermined number of times, the agent can converge to a strategy with good performance, thereby completing the training and saving the trained strategy network and its corresponding network parameters.
[0146] (5.2) Deployment and application of the method.
[0147] like Figure 2 As shown, in actual application, the trained strategy network Deployed to the imaging system to be optimized for optimal control of the dynamic image acquisition process of the imaging system.
[0148] The imaging system obtains the current image by executing the imaging parameters at the current moment. And the fused image of the previous moment , combined with the state parameters of the current imaging system to construct the current state value , input it into the feature extraction network to obtain the state features of the current moment , the state characteristics Enter the policy network , through the policy network Output the value of the next action space As the imaging parameter at the next moment, it is used to optimize the imaging system.
[0149] The image sequence obtained by the above method has higher information content and is more complementary, which can achieve accurate imaging of targets in the fields of space debris monitoring. The dynamically acquired images can be used for subsequent deep processing such as recognition, detection or splicing, and can also be fed back to the imaging system in real time to continuously optimize the shooting strategy, realize adaptive and dynamic image acquisition, and effectively improve imaging efficiency and quality.
[0150] See also Figure 3 , is a set of images acquired using imaging parameters at different times; Figure 4As shown, the imaging system is continuously optimized using the method of the present invention, and the complementary information of the group of images is integrated, so that a clearer and higher-quality image is finally captured.
[0151] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for dynamic image acquisition based on imaging parameter optimization control, characterized in that: include: The intelligent agent obtains the current moment image and the fused image of the previous moment by executing the imaging parameters of the current moment, and constructs the state value of the current moment by using the current moment image and the fused image; The imaging parameters and state values are the values of the action space and state space at the current moment respectively; The feature extraction network is used to extract the features of the current state value to obtain the state features of the current state; Based on the state characteristics of the current moment, the agent uses the strategy network to interact with the environment to obtain the state characteristics of the next moment; Setting a dual value network, and training the dual value network based on the strategy network and the state characteristics of the next moment; After each training of the dual value network training process, the policy network is trained using the preset loss function; After the dual value network and the strategy network are trained, the trained strategy network is saved for optimization control during the dynamic acquisition process of the intelligent agent image.
2. The method for dynamic image acquisition based on imaging parameter optimization control according to claim 1, characterized in that: The reward function used by the dual value network during training includes a measure based on space-time frequency complementarity and a measure based on entropy. In the measure based on space-time frequency complementarity, space-time domain complementarity is calculated by setting a threshold and performing intersection operations. The amplitude complementarity and phase complementarity are determined by using the amplitude spectrum and phase of the current moment image and the corresponding historical fusion image, respectively, and then the time-frequency domain complementarity is obtained by weighting. In the measure based on entropy, the information entropy of the historical fusion image is determined. Finally, the space-time domain complementarity, the time-frequency domain complementarity and the information entropy are weightedly combined to obtain the reward function.
3. The method for dynamic image acquisition based on imaging parameter optimization control according to claim 1, characterized in that: The agent obtains the current moment image and the fused image of the previous moment by executing the imaging parameters of the current moment, and constructs the current moment state value using the current moment image and the fused image, including: The agent executes the imaging parameters at the current moment , image the environment to obtain the current image ; Read the fused image generated at the last moment from the agent's storage module ; The fused image generated at the last moment With the current moment image According to the preset image fusion rules, the fusion is performed to obtain the historical fusion image at the current moment ; The agent obtains the state parameters of the current moment, and uses the state parameters and the historical fusion image Construct the state space and obtain the current state value based on the state space .
4. The method for dynamic image acquisition based on imaging parameter optimization control according to claim 1, characterized in that: The action space is expressed as: ; in, is the action space, Indicates the exposure time, Indicates the gain value, Indicates the aperture size, Indicates the direction of the turntable; The state space is represented as: ; in, is the state space, Represents the historical fusion image at the current moment, Represents the current moment image, Indicates the current trajectory information of the agent, Indicates the light intensity, Represents the agent posture information.
5. The method for dynamic image acquisition based on imaging parameter optimization control according to claim 1, characterized in that: The feature extraction network is used to extract the features of the current state value to obtain the current state features, including: Merge the history of the current moment into the image With the current moment image Stitching along the channel dimension to get the stitched image ; Stitching images Input to the feature extraction network to extract multi-scale semantic information and obtain the intermediate feature map , which is then flattened into image features in vector form ; The current state parameters are concatenated into a vector and then input into the multi-layer perceptron for dimension mapping to obtain the vector feature. ; Through a layer of linear transformation or direct alignment of image features and vector features The dimension is then concatenated on the sequence dimension or channel dimension to obtain the input feature ; Input features Input into the transform encoder, through the processing of multi-head self-attention and feedforward network, the final output result is the state feature of the current moment .
6. The method for dynamic image acquisition based on imaging parameter optimization control according to claim 1, characterized in that: Based on the state characteristics of the current moment, the agent uses the policy network to interact with the environment to obtain the state characteristics of the next moment, including: The current state characteristics Input to the policy network Get the value of the next action space ; The agent takes the value of the action space at the next moment As imaging parameters and execute, record the state value of the state space at the next moment after execution ; The state value at the next moment Input into the feature extraction network to obtain the state features of the next moment .
7. The method for dynamic image acquisition based on imaging parameter optimization control according to claim 2, characterized in that: The reward function construction process includes: Computing complementarity in space and time : ; in, ,image The current moment image or a historical fusion image of the current moment ; is a binary mask, Representing images At pixel position The pixel value of is the preset low threshold, is the preset high threshold; Computing time-frequency domain complementarity : ; in, and is the weight coefficient; is amplitude complementarity, is phase complementarity; Then the reward function It is expressed as: ; in, , , is the weight coefficient, It is a historical fusion image Information entropy.
8. The method for dynamic image acquisition based on imaging parameter optimization control according to claim 1, characterized in that: After each training of the dual value network training process, based on the following preset loss function Training the policy network : ; in, is the preset loss function, Indicates the state characteristics of the current moment, represents the jth value network, Indicates in status Strategy Network The value of the output action space.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor; characterized in that: When the processor executes the computer program, the method for dynamic image acquisition based on imaging parameter optimization control according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium storing a computer program; characterized in that: When the computer program is executed by a processor, the method for dynamic image acquisition based on imaging parameter optimization control according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Method and system for generating high dynamic range video under guidance of event camera
CN116456183A
Automatic modulation identification method based on multi-modal fusion and multi-task optimization
CN118643460A
Space debris intelligent detection method and device based on spatial-temporal feature fusion
CN119180946A
Intelligent space debris detection method and device based on morphological feature difference learning
CN119180947A
Space debris observation method based on alternating exposure times of charge coupled device (CCD) camera
US20220345610A1