Infrared target detection dual-band automatic selection method and device based on reinforcement learning
Through the reinforcement learning method, the problem of long time-consuming band selection and local optimal solution of infrared detection system is solved, and fast and efficient dual-band selection is achieved, which improves detection performance.
Patent Information
- Application Number
- CN202510172954.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-17
AI Technical Summary
The infrared detection system takes a long time in the band selection process and is prone to the problem of local optimal solutions.
Using reinforcement learning-based method, we obtain the reference index of the infrared target detection dual-band, determine the expression of the reward function, establish a reinforcement learning model, and use the near-end strategy gradient algorithm to train the model to achieve automatic selection of dual-bands.
By strengthening the independent learning ability of the learning model, the expected dual bands are quickly searched, which avoids the time-consuming problem of traditional band selection, and avoids the local optimal solution, achieving the band selection of the optimal result.
Smart Images

Figure CN120123672A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of infrared detection system design, and particularly to a dual-band automatic selection method and device for infrared target detection based on reinforcement learning. Background Art
[0002] The detection band selection of an infrared detection system directly affects the detection performance of the system. In order to obtain better detection performance, the band selection of the infrared detection system should cover the radiation characteristics of the target as much as possible, so as to reduce the radiation characteristics of the background. In recent years, with the progress of infrared detection system technology, detecting with dual bands can better improve the performance of target detection, can suppress the influence of the background, and can also reduce the influence of interference at the same time. Therefore, detecting with dual bands has received extensive attention at home and abroad.
[0003] Currently, the selection method of the detection band of an infrared detection system is usually as follows: According to the infrared radiation characteristics of the target and the background, combined with the basic parameters of the infrared detection system, the detection band is selected. During the band selection process, by analyzing indexes such as the contrast between the target and the background within the range of the band to be selected, the final band is determined. However, during the band selection process, the band to be selected needs to be repeatedly modified to determine the final band, and a large amount of manual assistance is required during the entire band selection process, resulting in a long time-consuming for the selection of the final band; and the selected final band also greatly depends on the experience of the selector, so that the selected final band is often a sub-optimal result, leading to the emergence of local optimal solutions. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a dual-band automatic selection method and device for infrared target detection based on reinforcement learning, so as to solve the problems that the selection of the final band is time-consuming and the selected final band has local optimal solutions.
[0005] To solve the above technical problems, the embodiments of the present invention provide the following technical solutions: The first aspect of the present invention provides a dual-band automatic selection method for infrared target detection based on reinforcement learning, including: Obtain the reference indexes of the dual bands for infrared target detection; Determine the expression of the reward function according to the reference indexes; Establish a reinforcement learning model, where the reinforcement learning model is a model that uses an agent to interact with the environment and obtains rewards through the expression of the reward function to realize the selection of the dual bands for infrared target detection; Train the reinforcement learning model using the proximal policy gradient algorithm to obtain a trained reinforcement learning model; Input the initially selected band into the trained reinforcement learning model to use the trained reinforcement learning model to select the initially selected band and output the finally selected dual band.
[0006] The second aspect of the present invention provides a dual-band automatic selection device for infrared target detection based on reinforcement learning, including: An acquisition module for acquiring reference indicators of the dual band for infrared target detection; A determination module for determining the expression of the reward function according to the reference indicators; A dual-band selection module for inputting the initially selected band into the trained reinforcement learning model to use the trained reinforcement learning model to select the initially selected band and output the finally selected dual band, where the trained reinforcement learning model is obtained by training the reinforcement learning model using the proximal policy gradient algorithm, and the reinforcement learning model is a model that uses an agent to interact with the environment and obtains rewards through the expression of the reward function to realize the selection of the dual band for infrared target detection.
[0007] Compared with the prior art, the dual-band automatic selection method and device for infrared target detection based on reinforcement learning provided by the present invention acquire reference indicators of the dual band for infrared target detection; determine the expression of the reward function according to the reference indicators; establish a reinforcement learning model, where the reinforcement learning model is a model that uses an agent to interact with the environment and obtains rewards through the expression of the reward function to realize the selection of the dual band for infrared target detection; use the proximal policy gradient algorithm to train the reinforcement learning model to obtain a trained reinforcement learning model; input the initially selected band into the trained reinforcement learning model to use the trained reinforcement learning model to select the initially selected band and output the finally selected dual band. In this way, the selected dual band makes full use of the autonomous learning ability of reinforcement learning to quickly search for the desired dual band, avoiding the problem of long time consumption in the band selection of traditional infrared detection systems, and making the process of band selection time-consuming shorter; at the same time, making full use of the autonomous learning ability of reinforcement learning can avoid local optimal solutions, and the optimal result of band selection can be searched through the trained reinforcement learning model. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, where: Figure 1 Schematically shows a flowchart of a dual-band automatic selection method for infrared target detection based on reinforcement learning; Figure 2Schematically shows a schematic diagram of the change of the infrared radiation intensity of the target with the wave number; Figure 3 Schematically shows a schematic diagram of the change of the total infrared radiation brightness of the background and the atmosphere with the wave number; Figure 4 Schematically shows a schematic diagram of the change of the atmospheric transmittance with the wave number; Figure 5 Schematically shows a schematic diagram of the change of the path radiation brightness of the atmosphere with the wave number; Figure 6 Schematically shows a schematic diagram of the finally selected dual band; Figure 7 Schematically shows a schematic diagram of the state change of the agent selecting the current dual band; Figure 8 Schematically shows a structural diagram of a dual-band automatic selection device for infrared target detection based on reinforcement learning. Detailed implementation manners
[0009] Hereinafter, exemplary embodiments of the present invention will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.
[0010] It should be noted that: unless otherwise specified, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those skilled in the art to which the present invention belongs.
[0011] The method in the embodiments of the present invention will be described in detail below.
[0012] Figure 1 Schematically shows a flowchart of a dual-band automatic selection method for infrared target detection based on reinforcement learning in an embodiment of the present invention. Refer to Figure 1 As shown, the dual-band automatic selection method for infrared target detection based on reinforcement learning may include: S101. Obtain reference indicators for the dual band of infrared target detection.
[0013] Among them, the reference indicators include the band range of the dual band of infrared target detection, the infrared radiation intensity of the target, the total infrared radiation brightness of the background and the atmosphere, and the atmospheric transmittance.
[0014] Determine the band range of the dual band of infrared target detection according to engineering requirements. The starting wavelength of the dual band of infrared target detection is 2.99 μm, and the termination wavelength of the dual band of infrared target detection is 14.29 μm, that is to say, the band range of the dual-band infrared target detection can be [2.99, 14.29] μm. There can be various band ranges for the dual-band infrared target detection, and no specific limitation is made here for the band range of the dual-band infrared target detection.
[0015] The starting wavelength of the dual-band infrared target detection The corresponding wave number is , the starting wavelength of the dual-band infrared target detection The corresponding wave number is 700 cm -1 , the ending wavelength of the dual-band infrared target detection The corresponding wave number is , the ending wavelength of the dual-band infrared target detection The corresponding wave number is 3340 cm -1 .
[0016] In order to obtain more detailed spectral information within the band range of the dual-band infrared target detection, a smaller spectral interval can be determined , Take 10 cm -1 , through the spectral interval The reference indexes of the dual-band infrared target detection are obtained.
[0017] Obtain the infrared radiation intensity of the target within the band range of the dual-band infrared target detection at the spectral interval . Figure 2 Schematically shows the schematic diagram of the change of the infrared radiation intensity of the target with the wave number, see Figure 2 shown, the abscissa is the wave number, the ordinate is the infrared radiation intensity of the target, and the range of the wave number is from 700 cm -1 to 3340 cm -1 , and the infrared radiation intensity of the target is normalized so that the infrared radiation intensity of the target is normalized to between 0 and 1.
[0018] Obtain the total infrared radiation brightness of the background and the atmosphere within the band range of the dual-band infrared target detection at the spectral interval . Figure 3 Schematically shows the schematic diagram of the change of the total infrared radiation brightness of the background and the atmosphere with the wave number, see Figure 3 shown, the abscissa is the wave number, the ordinate is the total infrared radiation brightness of the background and the atmosphere, and the range of the wave number is from 700 cm -1 to 3340 cm -1 , and the total infrared radiation brightness of the background and the atmosphere is normalized so that the total infrared radiation brightness of the background and the atmosphere is normalized to between 0 and 1.
[0019] Obtain the atmospheric transmittance within the spectral interval of the dual bands for infrared target detection. Figure 4 Schematically shows a diagram of the variation of atmospheric transmittance with wave number. See Figure 4 As shown, the horizontal axis is the wave number and the vertical axis is the atmospheric transmittance, which ranges from 0 to 0.8.
[0020] The reference index also includes the path radiance of the atmosphere. Figure 5 Schematically shows a diagram of the variation of the path radiance of the atmosphere with wave number. See Figure 5 As shown, the horizontal axis is the wave number and the vertical axis is the path radiance of the atmosphere, which ranges from 0 to 0.7.
[0021] S102. Determine the expression of the reward function according to the reference index.
[0022] Specifically, determining the expression of the reward function according to the reference index includes: Step A1: Determine the radiation intensity of the pixel where the target is located and the radiation intensity of the background pixel according to the infrared radiation intensity of the target, the atmospheric transmittance, the path radiance of the atmosphere, the true area of the target, the area of the background corresponding to one pixel on the imaging surface of the infrared detection system, the total infrared radiance of the background and the atmosphere, and the penalty factor.
[0023] Specifically, the expression of the radiation intensity of the pixel where the target is located is: ; Where is the radiation intensity of the pixel where the target is located, is the infrared radiation intensity of the target, is the atmospheric transmittance, is the path radiance of the atmosphere, is the true area of the target, is the total infrared radiance of the background and the atmosphere, is the area of the background corresponding to one pixel on the imaging surface of the infrared detection system, , is the distance from the target to the detection device, is the focal length of the detection device, is the wave number, is the size of the pixel of the detection device. Since the target is far away and cannot fill one pixel of the detection device, the contribution of the background needs to be considered for the pixel of the detection device. The size of the subscript is the first letter of system in the infrared detection system. The expression for the radiation intensity of the background pixel is: ; Wherein, is the radiation intensity of the background pixel, is the penalty factor. In the present invention, a relatively small penalty factor can be adopted so that the subsequent reinforcement learning model converges better.
[0024] Step A2: Determine the signal-to-noise ratio according to the radiation intensity of the pixel where the target is located and the radiation intensity of the background pixel.
[0025] The expression for the signal-to-noise ratio is: ; Wherein, is the signal-to-noise ratio, is the radiation intensity of the background pixel, is the radiation intensity of the pixel where the target is located, is the sum of the radiation intensities of the pixels where the target is located within a selected wavelength band for a certain time, is the sum of the radiation intensities of the background pixels within a selected wavelength band for a certain time.
[0026] Step A3: Determine the preset reward function according to the signal-to-noise ratio, the radiation intensity of the pixel where the target is located, and the radiation intensity of the background pixel.
[0027] Specifically, the expression for the preset reward function is: ; Wherein, is the preset reward function for the th iteration and the th detection wavelength band, is the wavelength band serial number of the detection wavelength band, , is the signal-to-noise ratio, is the wave number, is the radiation intensity of the pixel where the target is located, is the radiation intensity of the background pixel. is the sum of the radiation intensities of the pixels where the target is located within a selected wavelength band for a certain time, is the sum of the radiation intensities of the background pixels within a selected wavelength band for a certain time.
[0028] Step A4: Determine the expression of the reward function according to the preset reward function.
[0029] Specifically, the expression for the reward function is: ; Among them, is the reward for the th iteration, is the preset reward function for the first detection band in the th iteration, is the preset reward function for the second detection band in the th iteration, is the th iteration.
[0030] In the reward function of the reinforcement learning model of the present invention, not only the band ranges of the dual bands for infrared target detection, the infrared radiation intensity of the target, the total infrared radiation luminance of the background and the atmosphere, and the atmospheric transmittance are used, but also the basic parameters including the true area of the target, the area of the background corresponding to one pixel on the imaging surface of the infrared detection system, the distance from the target to the detection device, the size of the pixels of the detection device, and the focal length of the detection device are used. By coupling and considering the basic parameters of the infrared detection system, the multi-constraint solution for the dual-band selection of the infrared detection system is realized.
[0031] S103. Establish a reinforcement learning model.
[0032] Among them, the reinforcement learning model is a model that uses an agent to interact with the environment and obtains a reward through the expression of the reward function to realize the selection of the dual bands for infrared target detection.
[0033] Specifically, establishing the reinforcement learning model includes: taking the current dual bands selected within the band ranges of the dual bands for infrared target detection as the state of the environment, taking the change situation of the current dual-band selection as the action of the agent, taking the expression of the reward function as the reward of the agent, and taking the policy network-value function network mode as the learning mode of the agent. The policy network-value function network (actor-critic) mode is a mode composed of a policy (actor) network and a value function (critic) network. The current dual bands include the first current band and the second current band.
[0034] Specifically, both the actor network and the critic function network are implemented using neural networks. The actor network adopts a fully connected neural network structure. The input of the actor network is the state, and the output of the actor network is the action. According to the band ranges and spectral intervals of the dual bands for infrared target detection, it can be known that the dimension of the state space (i.e., the set of all states encountered by the agent) is 265. There are 2 hidden layers in the actor network, and there are 600 neurons in the actor network respectively.
[0035] The state of the actor network is described by the current dual band selected within the band range of the dual band of infrared target detection. The two current dual bands to be selected are the first current band, i.e., band 1, and the second current band, i.e., band 2. The wave numbers covered by band 1 can be marked as 1, the wave numbers covered by band 2 can be marked as -1, and the wave numbers covered by the other unselected dual bands of infrared target detection are marked as 0.
[0036] To traverse the entire state space, the actions should fully cover the entire state space. For this purpose, the dimension of the action space of the agent is set to 7, that is, there are 7 actions set for the agent. The 7 actions of the agent, namely the changes in the current dual band selection, are specifically as follows: Adding a wave number interval to the left of the current dual band, reducing a wave number interval to the left of the current dual band, adding a wave number interval to the right of the current dual band, reducing a wave number interval to the right of the current dual band, swapping the first current band or the second current band, moving the current dual band forward by one wave number interval, and moving the current dual band backward by one wave number interval.
[0037] Swapping the first current band or the second current band means swapping the currently acting first current band or the currently acting second current band. For example, when the currently acting band is the first current band, the currently acting first current band is swapped to the second current band, and vice versa.
[0038] When the first current band and the second current band coincide, the two bands are moved together. When the changes in the current dual band selection, that is, the actions of the agent, reach the boundary, the change of the state stops.
[0039] The actor network uses an optimizer based on the gradient-based optimization algorithm (Adaptive Moment Estimation, Adam). The learning rate of the actor network uses an exponentially decaying learning rate, and the initial learning rate of the actor network is 1e-4.
[0040] The critic network adopts a fully connected neural network structure. The input of the critic network is the state, the output of the critic network is the value of the state, the hidden layer of the critic network is 2 layers, and the neurons in the critic network are 600 and 300 respectively. The critic network uses the Adam optimizer, the learning rate of the critic network uses an exponentially decaying learning rate, and the initial learning rate of the critic network is 1e-4.
[0041] S104. Use the proximal policy gradient algorithm to train the reinforcement learning model to obtain a trained reinforcement learning model.
[0042] The Proximal Policy Optimization (PPO) algorithm stabilizes the training process by restricting the difference between the new policy (i.e., the second policy) and the old policy (i.e., the first policy). The PPO algorithm can avoid excessive policy updates, thereby reducing the instability and sample complexity during the training process.
[0043] Use the proximal policy gradient algorithm to train the reinforcement learning model to make the value of the reward as large as possible.
[0044] Specifically, using the PPO algorithm to train the reinforcement learning model to obtain a trained reinforcement learning model, including: Step B1: Initialize the policy network and the value function network.
[0045] Step B2: Take the current dual-band as the current state and input it into the agent so that the agent outputs the current action and obtain the current reward and the next state .
[0046] Among them, the current action is the action selected by the agent according to the initialized policy network, and the current reward and the next state are the reward and state feedback by the environment according to the current action using the expression of the reward function.
[0047] Step B3: Determine the current policy ratio according to the first policy and the second policy .
[0048] Among them, the current policy ratio is the ratio of taking the current action in the current state . The first policy is the policy containing the current policy network parameters, and the second policy is the policy containing the next policy network parameters.
[0049] represents the first policy of the actor network taking the current policy network parameters and taking the current action in the current state . represents the second policy of the actor network taking the next policy network parameters and taking the current action in the current state . The current policy network parameters including the weights and biases of the current policy network, and the parameters of the next policy network including the weights and biases of the next policy network.
[0050] The current policy ratio is the second policy and the first policy at the current state when taking the current action the ratio of probabilities.
[0051] When is relatively large, clip to a small interval to prevent the policy update from being too large, where is the clip parameter.
[0052] The expression of the current policy ratio is as follows: ; where, is the current policy ratio, is the second policy, is the first policy.
[0053] Step B4: Calculate the current loss of the initialized policy network according to the current policy ratio, the current reward, the discount factor, the current value, the next value, and the clip parameter.
[0054] Among them, the current value is the value corresponding to the value function network initialized with the intermediate parameter and the current state, and the next value is the value corresponding to the value function network initialized with the intermediate parameter and the next state.
[0055] Specifically, the expression of the current loss of the initialized policy network is: ; where, is the current loss of the initialized policy network, is the parameter of the next policy network, is the current policy ratio, , is the second policy, is the first policy, is the clip parameter, is to limit the elements in the current policy ratio within the specified range inside, is the advantage function, and the advantage function represents at the current state when, the current action Advantages relative to the average action is the current state is the current action , is the reward for the th iteration is the discount factor is the next state is the intermediate parameter is the next value, i.e., the next value under the next state and the intermediate parameter under the next value is the current value, i.e., the current state and the intermediate parameter under the current value
[0056] Intercept parameter takes a number between 0 and 1. For example, can take 0.2. The discount factor takes a number between 0 and 1. For example, can take 0.95
[0057] Step B5: Calculate the current loss of the initialized value function network based on the current reward, discount factor, current policy ratio, current value, and next value
[0058] Specifically, the expression for the current loss of the initialized value function network is ; where is the current loss of the initialized value function network
[0059] Step B6: Calculate the cumulative loss based on the current loss of the initialized policy network, the current loss of the initialized value function network, the weight factor, the entropy regularization term under the second policy, and the initial cumulative loss
[0060] where the weight factor includes the first weight factor and the second weight factor
[0061] Specifically, the expression for the cumulative loss is ; where is the cumulative loss is the initial cumulative loss is the first weight factor is the second weight factor is the entropy regularization term under the second policy, i.e., the actor network takes the current action under the next policy network parameters and the current state action The entropy regularization term corresponding to the second strategy at this time.
[0062] The first weight factor can take 1, and the second weight factor can take 0.01.
[0063] Step B7: According to the cumulative loss, use the Adam optimizer and the backpropagation algorithm to update the parameters of the next policy network in the reinforcement learning model and the intermediate parameters .
[0064] Use the Adam optimizer to update the parameters of the next policy network in the reinforcement learning model and the intermediate parameters , to minimize the cumulative loss .
[0065] Step B8: Input the next pair of bands as the next state into the agent, so that the agent outputs the next action and obtains the next reward and another state.
[0066] Wherein, another state is the next state of the next state.
[0067] Step B9: Take the next state as the current state, the next action as the current action, the next reward as the current reward, and another state as the current state, and return to the step of determining the current policy ratio according to the first strategy and the second strategy until the reinforcement learning model converges, so as to obtain the trained reinforcement learning model.
[0068] Specifically, take the next state as the current state, the next action as the current action, the next reward as the current reward, and another state as the current state, and return to step B3 until the reinforcement learning model converges, so as to obtain the trained reinforcement learning model.
[0069] The judgment condition for the convergence of the reinforcement learning model can be that when the value of the reward remains unchanged during the iteration of the agent, stop training the reinforcement learning model to obtain the trained reinforcement learning model.
[0070] The present invention can well apply the automatic learning ability of the reinforcement learning model, avoiding the problems of a large amount of manual operation, complex calculation and long time consumption in the traditional infrared detection system band selection. The present invention adopts the optimal result of the dual bands selected by the reinforcement learning model, solves the problem that it is difficult to obtain the optimal result by the traditional band selection method, and improves the detection performance of the infrared detection system; the present invention adopts the discretization processing method, discretizes the state of the band selection into the discrete state of the agent, and discretizes the change from state to state into finite actions, which well meets the requirements of the reinforcement learning model.
[0071] S105. Input the initially selected band into the trained reinforcement learning model to use the trained reinforcement learning model to select the initially selected band and output the selected dual band.
[0072] The selected dual band output by the trained reinforcement learning model is the infrared dual band that meets the requirements.
[0073] Figure 6 Schematically shows the schematic diagram of the finally selected dual band. The abscissa is the wave number, and the ordinate is the infrared radiation intensity of the target and the background (i.e., the infrared radiation intensity of the target and the infrared radiation intensity of the background). Figure 6 The solid line in it is the target, and the dotted line is the background. The main infrared radiation characteristics of the target are selected into two bands, indicating that the dual-band automatic selection method for infrared target detection based on reinforcement learning of the present invention can well select the characteristics of the target.
[0074] Figure 7 Schematically shows the schematic diagram of the state change of the agent selecting the current dual band. The abscissa is the number of iterations, and the ordinate is the band number, that is, the selected dual band. The uppermost is the state change of the first current band, that is, the state change of band 1, and the lowermost is the state change of the second current band, that is, the state change of band 2.
[0075] Based on the above Figure 1 implementation method, it can be seen that the dual-band automatic selection method for infrared target detection based on reinforcement learning in the embodiment of the present invention includes obtaining the reference index of the dual band for infrared target detection; determining the expression of the reward function according to the reference index; establishing a reinforcement learning model, where the reinforcement learning model is a model that uses the agent and the environment to interact and obtains rewards through the expression of the reward function to realize the selection of the dual band for infrared target detection; using the proximal policy gradient algorithm to train the reinforcement learning model to obtain the trained reinforcement learning model; inputting the initially selected band into the trained reinforcement learning model to use the trained reinforcement learning model to select the initially selected band and output the selected dual band. In this way, the selected dual band makes full use of the autonomous learning ability of reinforcement learning to quickly search for the expected selected dual band, avoiding the problem of long time consumption in the band selection of traditional infrared detection systems, making the process of band selection time-consuming short; at the same time, making full use of the autonomous learning ability of reinforcement learning can avoid local optimal solutions, and the optimal result of band selection can be searched through the trained reinforcement learning model.
[0076] Based on the same inventive concept, as an implementation of the above dual-band automatic selection method for infrared target detection based on reinforcement learning, the embodiment of the present invention also provides a dual-band automatic selection device for infrared target detection based on reinforcement learning. Figure 8The following is a structural diagram of the dual-band automatic selection device for infrared target detection based on reinforcement learning in the embodiments of the present invention. Refer to Figure 8 As shown in the figure, the dual-band automatic selection device for infrared target detection based on reinforcement learning may include: An acquisition module 801, configured to acquire reference indicators of the dual bands for infrared target detection; A determination module 802, configured to determine an expression of a reward function according to the reference indicators; A dual-band selection module 803, configured to input the initially selected band into the trained reinforcement learning model, so as to use the trained reinforcement learning model to select the initially selected band and output the selected dual bands. The trained reinforcement learning model is obtained by training the reinforcement learning model using the proximal policy gradient algorithm. The reinforcement learning model is a model that realizes the selection of the dual bands for infrared target detection by interacting an intelligent agent with the environment and obtaining rewards through the expression of the reward function.
[0077] The determination module 802 is specifically configured to determine the radiation intensity of the pixel where the target is located and the radiation intensity of the background pixel according to the infrared radiation intensity of the target, the atmospheric transmittance, the path radiation brightness of the atmosphere, the true area of the target, the area of the background corresponding to one pixel on the imaging surface of the infrared detection system, the total infrared radiation brightness of the background and the atmosphere, and the penalty factor; determine the signal-to-noise ratio according to the radiation intensity of the pixel where the target is located and the radiation intensity of the background pixel; determine a preset reward function according to the signal-to-noise ratio, the radiation intensity of the pixel where the target is located and the radiation intensity of the background pixel; and determine an expression of the reward function according to the preset reward function.
[0078] In the determination module 802, the expression of the radiation intensity of the pixel where the dual bands for infrared target detection are located is: ; Wherein, is the radiation intensity of the pixel where the target is located, is the infrared radiation intensity of the target, is the atmospheric transmittance, is the path radiation brightness of the atmosphere, is the true area of the target, is the total infrared radiation brightness of the background and the atmosphere, is the area of the background corresponding to one pixel on the imaging surface of the infrared detection system, is, is the distance from the target to the detection device, is the size of the pixel of the detection device, is the focal length of the detection device, is the wave number; The expression of the radiation intensity of the background pixel is: ; Among them, is the radiation intensity where the background pixel is located, is the penalty factor.
[0079] In the determination module 802, the expression of the preset reward function is: ; Among them, is the preset reward function of the th iteration of the th detection band, is the band serial number of the detection band, , is the signal-to-noise ratio, is the wave number, is the radiation intensity of the pixel where the target is located, is the radiation intensity where the background pixel is located.
[0080] In the determination module 802, the expression of the reward function is: ; Among them, is the reward of the th iteration, is the preset reward function of the first detection band of the th iteration, is the preset reward function of the second detection band of the th iteration, is the th iteration.
[0081] It should be noted here that: The above description of the embodiment of the infrared target detection dual-band automatic selection device based on reinforcement learning is similar to the description of the above embodiment of the infrared target detection dual-band automatic selection method based on reinforcement learning, and has similar beneficial effects to the embodiment of the infrared target detection dual-band automatic selection method based on reinforcement learning. For the technical details not disclosed in the embodiment of the infrared target detection dual-band automatic selection device based on reinforcement learning of the present invention, please refer to the description of the method embodiment of the present invention for understanding.
[0082] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A dual-band automatic selection method for infrared target detection based on reinforcement learning, characterized in that: include: Obtain reference indicators for dual-band infrared target detection; Determine an expression of a reward function according to the reference indicator; Establishing a reinforcement learning model, wherein the reinforcement learning model utilizes the interaction between the intelligent agent and the environment and obtains rewards through the expression of the reward function to realize the selection of the infrared target detection dual band; The reinforcement learning model is trained using a proximal policy gradient algorithm to obtain a trained reinforcement learning model; The initially selected bands are input into the trained reinforcement learning model, so as to select the initially selected bands using the trained reinforcement learning model and output the final selected dual bands.
2. The dual-band automatic selection method for infrared target detection based on reinforcement learning according to claim 1 is characterized in that: The reference indicators include the band range of the dual-band infrared target detection, the infrared radiation intensity of the target, the total infrared radiation brightness of the background and atmosphere, and the atmospheric transmittance.
3. The dual-band automatic selection method for infrared target detection based on reinforcement learning according to claim 2 is characterized in that: The expression of determining the reward function according to the reference indicator comprises: Determine the radiation intensity of the pixel where the target is located and the radiation intensity of the background pixel where the background is located according to the infrared radiation intensity of the target, the atmospheric transmittance, the path radiation brightness of the atmosphere, the real area of the target, the area of the background corresponding to a pixel on the imaging surface of the infrared detection system, the total infrared radiation brightness of the background and the atmosphere, and the penalty factor; Determining a signal-to-noise ratio according to the radiation intensity of the pixel where the target is located and the radiation intensity of the background pixel where the background is located; Determining a preset reward function according to the signal-to-noise ratio, the radiation intensity of the pixel where the target is located, and the radiation intensity of the background pixel; An expression of the reward function is determined according to the preset reward function.
4. The dual-band automatic selection method for infrared target detection based on reinforcement learning according to claim 3 is characterized in that: The expression of the radiation intensity of the pixel where the target is located is: ; in, is the radiation intensity of the pixel where the target is located, is the infrared radiation intensity of the target, is the atmospheric transmittance, is the path radiance of the atmosphere, is the actual area of the target, is the total infrared radiation brightness of the background and atmosphere, is the area of the background corresponding to a pixel on the imaging surface of the infrared detection system, , is the distance from the target to the detection device, To detect the size of the device pixel, To detect the focal length of the device, is the wave number; The expression of the radiation intensity of the background pixel is: ; in, is the radiation intensity of the background pixel, is the penalty factor.
5. The method according to claim 4, characterized in that The expression of the preset reward function is: ; in, For the The iteration The preset reward function for each detection band, is the band number of the detection band, , is the signal-to-noise ratio, is the wave number, is the radiation intensity of the pixel where the target is located, is the radiation intensity of the background pixel.
6. The method according to claim 5, characterized in that The expression of the reward function is: ; in, For the The reward for the iteration, For the The preset reward function for the first detection band of the iteration, For the The preset reward function for the second detection band of the iteration, For the Iterations.
7. The dual-band automatic selection method for infrared target detection based on reinforcement learning according to claim 2 is characterized in that: The step of establishing a reinforcement learning model comprises: The current dual band selected within the band range of the infrared target detection dual band is taken as the state of the environment, the change of the current dual band selection is taken as the action of the intelligent agent, the expression of the reward function is taken as the reward of the intelligent agent, and the strategy network-value function network model is taken as the learning model of the intelligent agent. The strategy network-value function network model is a model composed of a strategy network and a value function network.
8. The dual-band automatic selection method for infrared target detection based on reinforcement learning according to claim 7 is characterized in that: The current dual band includes a first current band and a second current band, and changes in the current dual band selection include adding a wave number interval on the left side of the current dual band, reducing one wave number interval on the left side of the current dual band, adding one wave number interval on the right side of the current dual band, reducing one wave number interval on the right side of the current dual band, exchanging the first current band or the second current band, moving the current dual band forward by one wave number interval, and moving the current dual band backward by one wave number interval.
9. The dual-band automatic selection method for infrared target detection based on reinforcement learning according to claim 7 is characterized in that: The method of training the reinforcement learning model by using the proximal policy gradient algorithm to obtain a trained reinforcement learning model includes: Initializing the policy network and the value function network; Input the current dual-band as the current state into the agent, so that the agent outputs the current action and obtains the current reward and the next state, wherein the current action is the action selected by the agent according to the initialized policy network, and the current reward and the next state are the reward and the state fed back by the environment according to the current action using the expression of the reward function; Determine a current strategy ratio according to the first strategy and the second strategy, wherein the current strategy ratio is a ratio of taking the current action in the current state, the first strategy is a strategy including current strategy network parameters, and the second strategy is a strategy including next strategy network parameters; Calculate the current loss of the initialized policy network according to the current policy ratio, the current reward, the discount factor, the current value, the next value, and the interception parameter, wherein the current value is the value corresponding to the initialized value function network under the intermediate parameters and the current state, and the next value is the value corresponding to the initialized value function network under the intermediate parameters and the next state; Calculate the current loss of the initialized value function network according to the current reward, the discount factor, the current strategy ratio, the current value, and the next value; Calculate the cumulative loss according to the current loss of the initialized policy network, the current loss of the initialized value function network, the weight factor, the entropy regularization term under the second policy, and the initial cumulative loss; According to the accumulated loss, the next strategy network parameters and the intermediate parameters in the reinforcement learning model are updated using an Adam optimizer and a back-propagation algorithm; Inputting the next double band as the next state into the agent, so that the agent outputs the next action and obtains the next reward and another state, wherein the another state is the next state of the next state; The next state is used as the current state, the next action is used as the current action, the next reward is used as the current reward, the further state is used as the current state, and the step of determining the current strategy ratio according to the first strategy and the second strategy is returned until the reinforcement learning model converges to obtain the trained reinforcement learning model.
10. A dual-band automatic selection device for infrared target detection based on reinforcement learning, characterized in that: include: An acquisition module is used to obtain reference indicators for dual-band infrared target detection; A determination module, used to determine an expression of a reward function according to the reference indicator; The dual-band selection module is used to input the initially selected band into the trained reinforcement learning model, so as to select the initially selected band using the trained reinforcement learning model and output the final selected dual band, wherein the trained reinforcement learning model is obtained by training the reinforcement learning model using the proximal policy gradient algorithm, and the reinforcement learning model is a model that utilizes the interaction between the intelligent agent and the environment and obtains rewards through the expression of the reward function to realize the selection of the dual bands for infrared target detection.
Citation Information
Patent Citations
Behavior imitation training method for air intelligent game
CN113221444A
Automatic driving lane selection decision-making method and system based on inverse reinforcement learning
CN116890855A
Beam forming method, device and equipment for intelligent metasurface and storage medium
CN118054828A
Infrared small target detection method based on reinforcement learning
CN118781327A