Dual-band automatic selection method and device for infrared target detection based on reinforcement learning
By automatically selecting the dual bands of the infrared detection system based on reinforcement learning, the problems of time-consuming and suboptimal solutions in traditional methods are solved, and fast and efficient optimal band selection is achieved, thereby improving detection performance.
Patent Information
- Application Number
- CN202510172954.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-02-17
AI Technical Summary
The existing infrared detection system takes a long time to select the band and relies on manual experience, which leads to suboptimal solutions and makes it difficult to select the optimal dual bands quickly and efficiently.
A reinforcement learning-based method is adopted to establish a reinforcement learning model by obtaining reference indicators of dual-band infrared target detection. The model is trained using the proximal policy gradient algorithm to automatically select the optimal dual-band.
It achieves fast and autonomous selection of the optimal dual bands, avoids the time-consuming and local optimal solution problems of traditional methods, and improves detection performance.
Smart Images

Figure CN120123672B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of infrared detection system design, and in particular to a reinforcement learning-based dual-band automatic selection method and device for infrared target detection. Background Art
[0002] The detection band selection of an infrared detection system directly affects the system's detection performance. To achieve better detection performance, the infrared detection system's band selection should cover the target's radiation characteristics as much as possible, thereby reducing the background radiation characteristics. In recent years, with the advancement of infrared detection system technology, the use of dual-band detection can better improve target detection performance, suppress background effects, and reduce the impact of interference. Therefore, the use of dual-band detection has received widespread attention both at home and abroad.
[0003] Currently, infrared detection systems typically select detection bands based on the infrared radiation characteristics of the target and background, combined with the basic parameters of the infrared detection system. The band selection process then determines the final band by analyzing indicators such as the contrast between the target and background within the selected band. However, this process requires repeated revisions of the selected band to determine the final band, requiring significant manual intervention. This makes the final band selection time-consuming and significantly dependent on the selector's experience, often resulting in suboptimal results and the occurrence of local optimal solutions. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a method and device for automatic dual-band selection of infrared target detection based on reinforcement learning, so as to solve the problems that the selection of the final band takes a long time and the local optimal solution appears in the selected final band.
[0005] To solve the above technical problems, the embodiments of the present invention provide the following technical solutions:
[0006] A first aspect of the present invention provides a dual-band automatic selection method for infrared target detection based on reinforcement learning, comprising:
[0007] Obtain reference indicators for dual-band infrared target detection;
[0008] Determine the expression of the reward function based on the reference indicator;
[0009] Establish a reinforcement learning model. This model uses the interaction between the agent and the environment to obtain rewards through the expression of the reward function, realizing the selection of dual-band infrared target detection.
[0010] The reinforcement learning model is trained using the proximal policy gradient algorithm to obtain a trained reinforcement learning model;
[0011] The initially selected bands are input into the trained reinforcement learning model, so that the trained reinforcement learning model is used to select the initially selected bands and output the final selected dual bands.
[0012] A second aspect of the present invention provides a dual-band automatic selection device for infrared target detection based on reinforcement learning, comprising:
[0013] Acquisition module, used to obtain reference indicators of dual-band infrared target detection;
[0014] A determination module, used to determine the expression of the reward function based on the reference indicator;
[0015] The dual-band selection module is used to input the initially selected band into the trained reinforcement learning model, so as to use the trained reinforcement learning model to select the initially selected band and output the final selected dual band. The trained reinforcement learning model is obtained by training the reinforcement learning model using the proximal policy gradient algorithm. The reinforcement learning model is a model that uses the interaction between the intelligent agent and the environment and obtains rewards through the expression of the reward function to realize the selection of dual bands for infrared target detection.
[0016] Compared to the prior art, the present invention provides a reinforcement learning-based automatic dual-band selection method and device for infrared target detection. The method obtains reference indicators for infrared target detection dual bands; determines an expression for a reward function based on the reference indicators; establishes a reinforcement learning model, which utilizes the interaction between an agent and an environment and obtains rewards through the expression of the reward function to select dual bands for infrared target detection; trains the reinforcement learning model using a proximal policy gradient algorithm to obtain a trained reinforcement learning model; inputs the initially selected bands into the trained reinforcement learning model, and uses the trained reinforcement learning model to select the initially selected bands and output the final selected dual bands. In this way, the selected dual bands fully utilize the autonomous learning capabilities of reinforcement learning to quickly search for the desired dual bands, avoiding the time-consuming band selection problem of traditional infrared detection systems and shortening the band selection process. Furthermore, by fully utilizing the autonomous learning capabilities of reinforcement learning, local optimal solutions can be avoided, and the trained reinforcement learning model can search for the optimal band selection result. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0018] Figure 1 The flowchart of the dual-band automatic selection method for infrared target detection based on reinforcement learning is schematically shown;
[0019] Figure 2 The following diagram schematically shows the change of the target's infrared radiation intensity with the wave number;
[0020] Figure 3 The diagram schematically shows how the total infrared radiation brightness of the background and atmosphere changes with the wave number;
[0021] Figure 4 The diagram schematically shows the change of atmospheric transmittance with wave number;
[0022] Figure 5 The diagram schematically shows how the atmospheric path radiance changes with wave number;
[0023] Figure 6 The final selected dual-band diagram is schematically shown;
[0024] Figure 7 The schematic diagram shows the state change of the agent selecting the current dual band;
[0025] Figure 8 The structure of the dual-band automatic selection device for infrared target detection based on reinforcement learning is schematically shown. DETAILED DESCRIPTION
[0026] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0027] It should be noted that, unless otherwise specified, the technical or scientific terms used in the present invention should have the common meanings understood by those skilled in the art to which the present invention belongs.
[0028] The method in the embodiment of the present invention is described in detail below.
[0029] Figure 1 The flowchart of the dual-band automatic selection method for infrared target detection based on reinforcement learning in an embodiment of the present invention is schematically shown. Figure 1 As shown, the dual-band automatic selection method for infrared target detection based on reinforcement learning may include:
[0030] S101. Obtain reference indicators for dual-band infrared target detection.
[0031] Among them, the reference indicators include the band range of the dual-band infrared target detection, the infrared radiation intensity of the target, the total infrared radiation brightness of the background and atmosphere, and the atmospheric transmittance.
[0032] Determine the band range of the infrared target detection dual band according to engineering requirements, and the starting wavelength of the infrared target detection dual band 2.99μm, the end wavelength of the dual-band infrared target detection The wavelength range of the infrared target detection dual-band is 14.29 μm, that is, the wavelength range of the infrared target detection dual-band can be [2.99, 14.29] μm. The wavelength range of the infrared target detection dual-band can be various, and the wavelength range of the infrared target detection dual-band is not specifically limited here.
[0033] Starting wavelength of dual-band infrared target detection The corresponding wave number is , the starting wavelength of the dual-band infrared target detection Corresponding wave number 700cm -1 , the end wavelength of the dual-band infrared target detection The corresponding wave number is , the end wavelength of the dual-band infrared target detection Corresponding wave number 3340cm -1 .
[0034] In order to obtain more detailed spectral information within the dual-band infrared target detection band, a smaller spectral interval can be determined , Take 10cm -1 , through spectral spacing Obtain reference indicators for dual-band infrared target detection.
[0035] Obtain infrared target detection dual-band band range, in the spectral interval Above, the infrared radiation intensity of the target. Figure 2 The schematic diagram shows the change of the target's infrared radiation intensity with the wave number, see Figure 2 As shown, the horizontal axis is the wave number, and the vertical axis is the infrared radiation intensity of the target. The wave number range is from 700cm -1 to 3340cm -1 , the infrared radiation intensity of the target is normalized so that the infrared radiation intensity of the target is normalized to between 0 and 1.
[0036] Obtain infrared target detection dual-band band range, in the spectral interval Above, the total infrared radiation brightness of the background and atmosphere. Figure 3Schematic diagram showing the variation of total infrared radiation brightness of background and atmosphere with wave number, see Figure 3 As shown, the horizontal axis is the wave number, and the vertical axis is the total infrared radiation brightness of the background and atmosphere. The wave number range is from 700cm -1 to 3340cm -1 , normalize the total infrared radiation brightness of the background and atmosphere so that the total infrared radiation brightness of the background and atmosphere is normalized to between 0 and 1.
[0037] Obtain infrared target detection dual-band band range, in the spectral interval Atmospheric transmittance. Figure 4 The diagram schematically shows how atmospheric transmittance varies with wave number, see Figure 4 As shown, the horizontal axis is the wave number, the vertical axis is the atmospheric transmittance, and the atmospheric transmittance is between 0 and 0.8.
[0038] Reference indicators also include the path radiance of the atmosphere, Figure 5 This diagram schematically shows how the atmospheric path radiance changes with wave number, see Figure 5 As shown, the horizontal axis is the wave number, and the vertical axis is the path radiation brightness of the atmosphere. The path radiation brightness of the atmosphere is between 0 and 0.7.
[0039] S102. Determine the expression of the reward function according to the reference indicator.
[0040] Specifically, the expression of the reward function is determined based on the reference indicator, including:
[0041] Step A1: Determine the radiation intensity of the target pixel and the radiation intensity of the background pixel based on the target's infrared radiation intensity, atmospheric transmittance, atmospheric path radiation brightness, the target's actual area, the area of the background corresponding to a pixel on the imaging surface of the infrared detection system, the total infrared radiation brightness of the background and atmosphere, and the penalty factor.
[0042] Specifically, the expression of the radiation intensity of the pixel where the target is located is:
[0043] ;
[0044] in, is the radiation intensity of the pixel where the target is located, is the infrared radiation intensity of the target, is the atmospheric transmittance, is the atmospheric path radiance, is the actual area of the target, is the total infrared radiation brightness of the background and atmosphere, is the area of the background corresponding to a pixel on the imaging surface of the infrared detection system, , is the distance from the target to the detection device, is the focal length of the detection device, is the wave number, The size of the detection device pixel. Since the target is too far away to fill a detection device pixel, the detection device pixel also needs to consider the contribution of the background. Subscript It is the first letter of the system in infrared detection system.
[0045] The expression of the radiation intensity of the background pixel is:
[0046] ;
[0047] in, is the radiation intensity of the background pixel, is the penalty factor. In the present invention, A smaller penalty factor can be used to make the subsequent reinforcement learning model converge better.
[0048] Step A2: Determine the signal-to-noise ratio based on the radiation intensity of the pixel where the target is located and the radiation intensity of the background pixel.
[0049] The expression of signal-to-noise ratio is:
[0050] ;
[0051] in, is the signal-to-noise ratio, is the radiation intensity of the background pixel, is the radiation intensity of the pixel where the target is located, It is the sum of the radiation intensities of the target pixels within a certain selected band. It is the sum of the radiation intensities of the background pixels within a certain selected band.
[0052] Step A3: Determine a preset reward function based on the signal-to-noise ratio, the radiation intensity of the pixel where the target is located, and the radiation intensity of the background pixel.
[0053] Specifically, the expression of the preset reward function is:
[0054] ;
[0055] in, For the The first iteration The preset reward function for each detection band, is the band number of the detection band, , is the signal-to-noise ratio, is the wave number, is the radiation intensity of the pixel where the target is located, is the radiation intensity of the background pixel. It is the sum of the radiation intensities of the target pixels within a certain selected band. It is the sum of the radiation intensities of the background pixels within a certain selected band.
[0056] Step A4: Determine the expression of the reward function according to the preset reward function.
[0057] Specifically, the reward function is expressed as:
[0058] ;
[0059] in, For the The reward for the iteration, For the The preset reward function for the first detection band of the iteration, For the The preset reward function for the second detection band of the iteration, For the Iterations.
[0060] The reward function of the reinforcement learning model of the present invention not only uses the band range of the dual-band infrared target detection, the infrared radiation intensity of the target, the total infrared radiation brightness of the background and atmosphere, and the atmospheric transmittance, but also uses basic parameters such as the actual area of the target, the area of the background corresponding to a pixel on the imaging surface of the infrared detection system, the distance from the target to the detection device, the size of the detection device pixel, and the focal length of the detection device. The basic parameters of the infrared detection system are coupled and considered to realize the multi-constraint solution of the dual-band selection of the infrared detection system.
[0061] S103. Establish a reinforcement learning model.
[0062] Among them, the reinforcement learning model utilizes the interaction between the intelligent agent and the environment and obtains rewards through the expression of the reward function to realize the selection of dual-band infrared target detection.
[0063] Specifically, a reinforcement learning model is established, including: using the current dual band selected within the infrared target detection dual band range as the state of the environment, using changes in the current dual band selection as the agent's action, using the expression of the reward function as the agent's reward, and using a policy network-value function network model as the agent's learning model. The policy network-value function network (actor-critic) model is composed of a policy (actor) network and a value function (critic) network. The current dual band includes a first current band and a second current band.
[0064] Specifically, both the actor network and the critic function network are implemented using neural networks. The actor network employs a fully connected neural network architecture, with states as input and actions as output. Based on the wavelength range and spectral spacing of the dual-band infrared target detection, the dimension of the state space (i.e., the set of all states encountered by the agent) is 265. The actor network has two hidden layers and 600 neurons in each.
[0065] The state of the actor network is described by the current dual band selected within the infrared target detection dual band band range. There are two current dual bands to select: the first current band, Band 1, and the second current band, Band 2. The wavenumbers covered by Band 1 can be marked as 1, the wavenumbers covered by Band 2 can be marked as -1, and the wavenumbers covered by the remaining unselected infrared target detection dual bands can be marked as 0.
[0066] In order to traverse the entire state space, the action should completely cover the entire state space. For this purpose, the dimension of the action space of the agent is set to 7, that is, there are 7 action settings for the agent. The actions of the 7 agents, that is, the changes in the current dual-band selection, are as follows:
[0067] Add one wave number interval to the left of the current dual band, subtract one wave number interval to the left of the current dual band, add one wave number interval to the right of the current dual band, subtract one wave number interval to the right of the current dual band, swap the first current band or the second current band, move the current dual band forward by one wave number interval, and move the current dual band backward by one wave number interval.
[0068] Exchanging the first current band or the second current band refers to exchanging the currently active first current band or the currently active second current band. For example, if the currently active first current band is the first current band, the currently active first current band is exchanged with the second current band, and vice versa.
[0069] When the first current band and the second current band overlap, the two bands are moved together, and when the change in the current dual-band selection, that is, the action of the agent reaches the boundary, the change in state is stopped.
[0070] The actor network uses the gradient-based optimization algorithm (Adaptive Moment Estimation, Adam) optimizer, the learning rate of the actor network uses an exponentially decaying learning rate, and the initial learning rate of the actor network is 1e-4.
[0071] The critic network uses a fully connected neural network structure. The input of the critic network is the state, and the output of the critic network is the value of the state. The critic network has two hidden layers, with 600 and 300 neurons respectively. The critic network uses the Adam optimizer, with an exponentially decaying learning rate and an initial learning rate of 1e-4.
[0072] S104. Use the proximal policy gradient algorithm to train the reinforcement learning model to obtain a trained reinforcement learning model.
[0073] Proximal Policy Optimization (PPO) stabilizes the training process by limiting the difference between the new policy (i.e., the second policy) and the old policy (i.e., the first policy). PPO avoids excessive policy updates, thereby reducing instability and sample complexity during training.
[0074] The reinforcement learning model is trained using the proximal policy gradient algorithm to make the reward value as large as possible.
[0075] Specifically, the PPO algorithm is used to train the reinforcement learning model to obtain a trained reinforcement learning model, including:
[0076] Step B1: Initialize the policy network and value function network.
[0077] Step B2: Set the current dual band as the current state Input into the agent so that the agent can output the current action And get the current rewards and the next state .
[0078] Among them, the current action is the action selected by the agent according to the initialized policy network, and the current reward and the next state It is the reward and state fed back by the environment based on the current action using the expression of the reward function.
[0079] Step B3: According to the first strategy and the second strategy , determine the current strategy ratio .
[0080] Among them, the current strategy ratio In the current state Take the current action The ratio of the first strategy The strategy that contains the current strategy network parameters, the second strategy is the policy that contains the network parameters for the next policy.
[0081] Indicates that the actor network takes the current policy network parameters , in the current state Take the current action The first strategy, Indicates that the actor network takes the next policy network parameter , in the current state Take the current action The second strategy. Current strategy network parameters Including the weights and biases of the current policy network, the parameters of the next policy network Includes the weights and biases of the next policy network.
[0082] Current Strategy Ratio The second strategy and the first strategy In the current state Take the current action The probability ratio.
[0083] when When it is larger, Restricted to a small area To prevent the policy update from being too large, It is the clip parameter.
[0084] Current Strategy Ratio The expression is:
[0085] ;
[0086] in, is the current strategy ratio, For the second strategy, The first strategy.
[0087] Step B4: Calculate the current loss of the initialized policy network based on the current policy ratio, current reward, discount factor, current value, next value, and interception parameter.
[0088] Among them, the current value is the value corresponding to the initialized value function network with the intermediate parameters and the current state, and the next value is the value corresponding to the initialized value function network with the intermediate parameters and the next state.
[0089] Specifically, the current loss expression of the initialized policy network is:
[0090] ;
[0091] in, is the current loss of the initialized policy network, is the next strategy network parameter, is the current strategy ratio, , For the second strategy, As the first strategy, To intercept parameters, To change the current strategy ratio The elements in the are restricted to the specified range Inside, is the advantage function, which indicates that Next, current action Advantages over average action, is the current state, For the current action, , For the The reward for the iteration, is the discount factor, For the next state, is the intermediate parameter, is the next value, that is, the next state and intermediate parameters The next value under is the current value, that is, the current state and intermediate parameters The current value of .
[0092] Intercept parameters Take a number between 0 and 1, for example, It can be taken as 0.2. Discount factor Take a number between 0 and 1, for example, You can take 0.95.
[0093] Step B5: Calculate the current loss of the initialized value function network based on the current reward, discount factor, current policy ratio, current value, and next value.
[0094] Specifically, the expression of the current loss of the initialized value function network is:
[0095] ;
[0096] in, is the current loss of the value function network after initialization.
[0097] Step B6: Calculate the cumulative loss based on the current loss of the initialized policy network, the current loss of the initialized value function network, the weight factor, the entropy regularization term under the second strategy, and the initial cumulative loss.
[0098] The weight factors include a first weight factor and a second weight factor.
[0099] Specifically, the expression of cumulative loss is:
[0100] ;
[0101] in, For cumulative losses, is the initial cumulative loss, is the first weight factor, is the second weight factor, is the entropy regularization term under the second strategy, that is, the actor network parameter in the next strategy and current status Take the current action under the effect is the entropy regularization term corresponding to the second strategy.
[0102] First weighting factor Can take 1, the second weight factor It can be taken as 0.01.
[0103] Step B7: Based on the accumulated loss, use the Adam optimizer and backpropagation algorithm to update the next policy network parameters in the reinforcement learning model and intermediate parameters .
[0104] Using the Adam optimizer to update the next policy network parameters in the reinforcement learning model and intermediate parameters , to minimize the cumulative loss .
[0105] Step B8: Input the next double band as the next state into the agent so that the agent outputs the next action and obtains the next reward and another state.
[0106] Here, the further state is the next state of the next state.
[0107] Step B9: Set the next state as the current state, the next action as the current action, the next reward as the current reward, and the next state as the current state, and return to the step of determining the current strategy ratio according to the first strategy and the second strategy until the reinforcement learning model converges to obtain a trained reinforcement learning model.
[0108] Specifically, the next state is used as the current state, the next action is used as the current action, the next reward is used as the current reward, and the next state is used as the current state, and the process returns to step B3 until the reinforcement learning model converges to obtain a trained reinforcement learning model.
[0109] The judgment condition for the convergence of the reinforcement learning model can be that when the value of the reward remains unchanged during the iteration process of the intelligent agent, the training of the reinforcement learning model is stopped to obtain a trained reinforcement learning model.
[0110] The present invention effectively leverages the automatic learning capabilities of reinforcement learning models, avoiding the extensive manual labor, complex calculations, and time-consuming nature of traditional infrared detection system band selection. By employing a reinforcement learning model to select the optimal dual-band result, the present invention addresses the difficulty of achieving optimal results with traditional band selection methods, thereby improving the detection performance of the infrared detection system. The present invention also employs a discretization method to discretize the band selection state into discrete states of the agent and discretize the transition from state to state into finite actions, effectively meeting the requirements of reinforcement learning models.
[0111] S105 , inputting the initially selected band into the trained reinforcement learning model, so as to select the initially selected band using the trained reinforcement learning model, and outputting the selected dual bands.
[0112] The selected dual bands output by the trained reinforcement learning model are infrared dual bands that meet the requirements.
[0113] Figure 6 The final selected dual-band diagram is shown schematically, with the horizontal axis being the wave number and the vertical axis being the infrared radiation intensity of the target and the background (i.e. the infrared radiation intensity of the target and the infrared radiation intensity of the background). Figure 6 The solid line in the middle is the target, and the dotted line is the background. The main infrared radiation characteristics of the target are selected into two bands, which shows that the reinforcement learning-based dual-band automatic selection method for infrared target detection of the present invention can well select the characteristics of the target.
[0114] Figure 7 The figure schematically shows the state changes of the intelligent agent selecting the current dual bands. The horizontal axis is the number of iterations, and the vertical axis is the band number, i.e. the selected dual bands. The top one is the state change of the first current band, i.e. the state change of band 1, and the bottom one is the state change of the second current band, i.e. the state change of band 2.
[0115] Based on the above Figure 1 As can be seen from the implementation method, the reinforcement learning-based automatic selection method for infrared target detection dual bands in an embodiment of the present invention includes obtaining reference indicators for infrared target detection dual bands; determining an expression for a reward function based on the reference indicators; establishing a reinforcement learning model, which utilizes the interaction between an agent and an environment and obtains rewards through the expression of the reward function to select the infrared target detection dual bands; training the reinforcement learning model using a proximal policy gradient algorithm to obtain a trained reinforcement learning model; inputting the initially selected band into the trained reinforcement learning model, so that the trained reinforcement learning model selects the initially selected band and outputs the selected dual bands. In this way, the selected dual bands fully utilize the autonomous learning capability of reinforcement learning to quickly search for the desired dual bands, avoiding the time-consuming band selection problem of traditional infrared detection systems and shortening the band selection process. Furthermore, fully utilizing the autonomous learning capability of reinforcement learning can avoid local optimal solutions, and the trained reinforcement learning model can search for the optimal result for band selection.
[0116] Based on the same inventive concept, as an implementation of the above-mentioned reinforcement learning-based dual-band automatic selection method for infrared target detection, an embodiment of the present invention also provides a reinforcement learning-based dual-band automatic selection device for infrared target detection. Figure 8 This is a structural diagram of the dual-band automatic selection device for infrared target detection based on reinforcement learning in an embodiment of the present invention, see Figure 8 As shown, the infrared target detection dual-band automatic selection device based on reinforcement learning may include:
[0117] An acquisition module 801 is used to obtain reference indicators for dual-band infrared target detection;
[0118] A determination module 802 is configured to determine an expression of a reward function according to a reference indicator;
[0119] The dual-band selection module 803 is used to input the initially selected band into the trained reinforcement learning model, use the trained reinforcement learning model to select the initially selected band, and output the selected dual band. The trained reinforcement learning model is obtained by training the reinforcement learning model using the proximal policy gradient algorithm. The reinforcement learning model is a model that uses the interaction between the intelligent agent and the environment and obtains rewards through the expression of the reward function to realize the selection of dual bands for infrared target detection.
[0120] Determination module 802 is specifically configured to determine the radiation intensity of the target pixel and the radiation intensity of the background pixel based on the target's infrared radiation intensity, atmospheric transmittance, atmospheric path radiation brightness, the target's actual area, the area of the background corresponding to a pixel on the imaging surface of the infrared detection system, the total infrared radiation brightness of the background and the atmosphere, and a penalty factor; determine a signal-to-noise ratio based on the radiation intensity of the target pixel and the radiation intensity of the background pixel; determine a preset reward function based on the signal-to-noise ratio, the radiation intensity of the target pixel and the radiation intensity of the background pixel; and determine an expression for the reward function based on the preset reward function.
[0121] In the determination module 802, the expression for the radiation intensity of the pixel where the infrared target detection dual band is located is:
[0122] ;
[0123] in, is the radiation intensity of the pixel where the target is located, is the infrared radiation intensity of the target, is the atmospheric transmittance, is the atmospheric path radiance, is the actual area of the target, is the total infrared radiation brightness of the background and atmosphere, is the area of the background corresponding to a pixel on the imaging surface of the infrared detection system, , is the distance from the target to the detection device, is the size of the detection device pixel, is the focal length of the detection device, is the wave number;
[0124] The expression of the radiation intensity of the background pixel is:
[0125] ;
[0126] in, is the radiation intensity of the background pixel, is the penalty factor.
[0127] In the determination module 802, the expression of the preset reward function is:
[0128] ;
[0129] in, For the The first iteration The preset reward function for each detection band, is the band number of the detection band, , is the signal-to-noise ratio, is the wave number, is the radiation intensity of the pixel where the target is located, is the radiation intensity of the background pixel.
[0130] In the determination module 802, the reward function is expressed as:
[0131] ;
[0132] in, For the The reward for the iteration, For the The preset reward function for the first detection band of the iteration, For the The preset reward function for the second detection band of the iteration, For the Iterations.
[0133] It should be noted that the above description of the embodiment of the dual-band automatic selection device for infrared target detection based on reinforcement learning is similar to the description of the embodiment of the dual-band automatic selection method for infrared target detection based on reinforcement learning, and has similar beneficial effects as the embodiment of the dual-band automatic selection method for infrared target detection based on reinforcement learning. For technical details not disclosed in the embodiment of the dual-band automatic selection device for infrared target detection based on reinforcement learning of the present invention, please refer to the description of the embodiment of the method of the present invention for understanding.
[0134] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A dual-band automatic selection method for infrared target detection based on reinforcement learning, characterized in that: include: Obtain reference indicators for dual-band infrared target detection; Determine an expression of a reward function according to the reference indicator; Establishing a reinforcement learning model, wherein the reinforcement learning model utilizes the interaction between the intelligent agent and the environment and obtains rewards through the expression of the reward function to achieve the selection of the infrared target detection dual band; Training the reinforcement learning model using a proximal policy gradient algorithm to obtain a trained reinforcement learning model; Inputting the initially selected band into the trained reinforcement learning model, so as to select the initially selected band using the trained reinforcement learning model and outputting a final selected dual band; The expression of the reward function determined according to the reference indicator includes: The radiation intensity of the target pixel and the radiation intensity of the background pixel are determined based on the target's infrared radiation intensity, atmospheric transmittance, atmospheric path radiation brightness, the target's actual area, the area of the background corresponding to a pixel on the imaging surface of the infrared detection system, the total infrared radiation brightness of the background and atmosphere, and the penalty factor. Determining a signal-to-noise ratio based on the radiation intensity of the pixel where the target is located and the radiation intensity of the background pixel where the target is located; Determining a preset reward function according to the signal-to-noise ratio, the radiation intensity of the pixel where the target is located, and the radiation intensity of the background pixel; Determining an expression of the reward function according to the preset reward function; The expression of the preset reward function is: ; in, For the The first iteration The preset reward function for each detection band, is the band number of the detection band, , is the signal-to-noise ratio, is the wave number, is the radiation intensity of the pixel where the target is located, is the radiation intensity of the background pixel; The expression of the reward function is: ; in, For the The reward for the iteration, For the The preset reward function for the first detection band of the iteration, For the The preset reward function for the second detection band of the iteration, For the Iterations.
2. The dual-band automatic selection method for infrared target detection based on reinforcement learning according to claim 1 is characterized in that: The reference indicators include the band range of the dual-band infrared target detection, the infrared radiation intensity of the target, the total infrared radiation brightness of the background and atmosphere, and the atmospheric transmittance.
3. The dual-band automatic selection method for infrared target detection based on reinforcement learning according to claim 2 is characterized in that: The expression of the radiation intensity of the pixel where the target is located is: ; in, is the radiation intensity of the pixel where the target is located, is the infrared radiation intensity of the target, is the atmospheric transmittance, is the path radiance of the atmosphere, is the actual area of the target, is the total infrared radiation brightness of the background and atmosphere, is the area of the background corresponding to a pixel on the imaging surface of the infrared detection system, , is the distance from the target to the detection device, is the size of the detection device pixel, is the focal length of the detection device, is the wave number; The expression of the radiation intensity of the background pixel is: ; in, is the radiation intensity of the background pixel, is the penalty factor.
4. The dual-band automatic selection method for infrared target detection based on reinforcement learning according to claim 2 is characterized in that: The establishment of the reinforcement learning model includes: The current dual band selected within the band range of the infrared target detection dual band is used as the state of the environment, the change in the current dual band selection is used as the action of the intelligent agent, the expression of the reward function is used as the reward of the intelligent agent, and the policy network-value function network model is used as the learning model of the intelligent agent. The policy network-value function network model is a model composed of a policy network and a value function network.
5. The dual-band automatic selection method for infrared target detection based on reinforcement learning according to claim 4 is characterized in that: The current dual band includes a first current band and a second current band, and the changes in the current dual band selection include adding a wave number interval to the left side of the current dual band, reducing one wave number interval to the left side of the current dual band, adding one wave number interval to the right side of the current dual band, reducing one wave number interval to the right side of the current dual band, exchanging the first current band or the second current band, moving the current dual band forward by one wave number interval, and moving the current dual band backward by one wave number interval.
6. The dual-band automatic selection method for infrared target detection based on reinforcement learning according to claim 4 is characterized in that: The method of training the reinforcement learning model using the proximal policy gradient algorithm to obtain a trained reinforcement learning model includes: Initializing the policy network and the value function network; Input the current dual-band as the current state into the agent, so that the agent outputs the current action and obtains the current reward and the next state, wherein the current action is the action selected by the agent according to the initialized policy network, and the current reward and the next state are the reward and state fed back by the environment according to the current action using the expression of the reward function; Determining a current policy ratio according to a first policy and a second policy, wherein the current policy ratio is a ratio of taking the current action in a current state, the first policy being a policy including current policy network parameters, and the second policy being a policy including next policy network parameters; Calculate the current loss of the initialized policy network based on the current policy ratio, the current reward, the discount factor, the current value, the next value, and the interception parameter, wherein the current value is the value corresponding to the initialized value function network under the intermediate parameters and the current state, and the next value is the value corresponding to the initialized value function network under the intermediate parameters and the next state; Calculating a current loss of the initialized value function network according to the current reward, the discount factor, the current policy ratio, the current value, and the next value; Calculate the cumulative loss based on the current loss of the initialized policy network, the current loss of the initialized value function network, the weight factor, the entropy regularization term under the second policy, and the initial cumulative loss; Based on the accumulated loss, the next policy network parameters and the intermediate parameters in the reinforcement learning model are updated using an Adam optimizer and a back-propagation algorithm; Inputting the next double band as the next state into the agent, so that the agent outputs the next action and obtains the next reward and another state, wherein the another state is the next state of the next state; The next state is used as the current state, the next action is used as the current action, the next reward is used as the current reward, the further state is used as the current state, and the process returns to the step of determining the current strategy ratio according to the first strategy and the second strategy until the reinforcement learning model converges to obtain the trained reinforcement learning model.
7. A dual-band automatic selection device for infrared target detection based on reinforcement learning, characterized in that: include: Acquisition module, used to obtain reference indicators of dual-band infrared target detection; A determination module, configured to determine an expression of a reward function according to the reference indicator; A dual-band selection module is configured to input the initially selected band into a trained reinforcement learning model, select the initially selected band using the trained reinforcement learning model, and output a final selected dual-band. The trained reinforcement learning model is obtained by training the reinforcement learning model using a proximal policy gradient algorithm. The reinforcement learning model utilizes interaction between an agent and an environment and obtains rewards through an expression of the reward function to implement selection of the dual-band infrared target detection model. The determination module is specifically configured to determine the radiation intensity of the pixel where the target is located and the radiation intensity of the background pixel based on the infrared radiation intensity of the target, the atmospheric transmittance, the path radiation brightness of the atmosphere, the actual area of the target, the area of the background corresponding to a pixel on the imaging surface of the infrared detection system, the total infrared radiation brightness of the background and the atmosphere, and a penalty factor; determine a signal-to-noise ratio based on the radiation intensity of the pixel where the target is located and the radiation intensity of the background pixel; determine a preset reward function based on the signal-to-noise ratio, the radiation intensity of the pixel where the target is located, and the radiation intensity of the background pixel; and determine an expression for the reward function based on the preset reward function; In the determination module, the expression of the preset reward function is: ; in, For the The first iteration The preset reward function for each detection band, is the band number of the detection band, , is the signal-to-noise ratio, is the wave number, is the radiation intensity of the pixel where the target is located, is the radiation intensity of the background pixel; In the determination module, the expression of the reward function is: ; in, For the The reward for the iteration, For the The preset reward function for the first detection band of the iteration, For the The preset reward function for the second detection band of the iteration, For the Iterations.