A large surface source blackbody parameter-free temperature control method, device and medium

By modeling the temperature control problem of large-narrow source bold body as Markov decision-making process and combining with deep DQN network, parameterless temperature control is achieved, and the efficiency and accuracy problems of traditional temperature control methods under dynamic changes and nonlinear conditions are solved, significantly improving the temperature control response speed and control accuracy.

CN119536413BActive Publication Date: 2025-05-06HANGZHOU INST FOR ADVANCED STUDY UCAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510105108.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-06
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

When facing the temperature control needs of dynamic changes and nonlinear characteristics, the traditional large-narrow source black body temperature control method is inefficient, insufficient accuracy, and high regulation complexity, making it difficult to respond quickly to temperature changes, resulting in a long time-consuming test and calibration process and cannot meet the requirements of real-time and consistency.

Method used

The reinforcement learning method is adopted to model the temperature control problem as Markov decision-making process (MDP), and combined with the deep DQN network, the parameterless temperature control is achieved. Through the interaction between the agent and the environment, the power output is dynamically adjusted to achieve precise temperature control.

Benefits of technology

It significantly improves the response speed and control accuracy of temperature control, solves the problem of temperature control under dynamic nonlinear conditions, has good adaptability and scalability, and can quickly and accurately adjust the temperature within the minimum number of action steps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119536413B_ABST
    Figure CN119536413B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of temperature control technology, and discloses a large surface source black body parameter-free temperature control method, device and medium, including: S1, obtaining a data set; S2, problem modeling; S3, building a large surface source black body parameter-free temperature control model based on a deep DQN network, and then using a training data set to train the large surface source black body parameter-free temperature control model; S4, verification and adjustment, using a test data set to evaluate the large surface source black body parameter-free temperature control model, and obtain the optimal large surface source black body parameter-free temperature control model; S5, applying the optimal large surface source black body parameter-free temperature control model to a large surface source black body temperature control system to control the temperature of the large surface source black body. Based on a deep DQN network, the present invention can quickly and accurately adjust the temperature to a target value within the minimum number of action steps through precise state and action design and an optimized reward function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of temperature control technology, and in particular to a large surface source black body parameter-free temperature control method, device and medium. Background Art

[0002] As a calibration device with extremely high requirements for temperature control response rate and accuracy, large surface source blackbody is widely used in infrared detection equipment calibration, thermal imaging system testing and other scenarios. Traditional blackbody temperature control methods rely on mathematical modeling and parametric design, and usually require manual adjustment of multiple parameters to adapt to different experimental conditions. However, this method is prone to low efficiency, insufficient accuracy and high control complexity when faced with dynamic changes and nonlinear temperature control requirements. Especially in high-precision infrared detection and measurement, traditional methods are difficult to respond to temperature changes quickly, resulting in a long test and calibration process, which cannot meet the real-time and consistency requirements.

[0003] In response to the above problems, reinforcement learning, as a data-driven intelligent algorithm, has become an ideal tool for solving complex temperature control problems because it can perform autonomous learning and optimization without parameters. By modeling the temperature control problem as a Markov decision process (MDP) and combining it with a deep DQN network, the large-surface blackbody parameter-free temperature control method can sense the temperature state in real time and dynamically adjust the power output to achieve precise temperature control. This method abandons the traditional parameter dependence and optimizes the deep DQN network through the interaction between the intelligent agent and the environment, which significantly improves the response speed and control accuracy of the temperature control and solves the temperature control problem under dynamic nonlinear conditions. At the same time, the reinforcement learning method also has good adaptability and scalability, providing new research directions and application prospects for the development of temperature control technology. Summary of the invention

[0004] The object of the present invention is to provide a large surface source blackbody parameter-free temperature control method, device and medium to solve the problems raised in the above background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A large area source blackbody parameter-free temperature control method comprises the following steps:

[0007] S1, obtain a data set, generate a simulated temperature data set through blackbody system simulation, and build an actual temperature acquisition system to obtain a real temperature data set, mix the simulated temperature data set with the real temperature data set, and divide it into a training data set and a test data set in proportion;

[0008] S2, problem modeling, modeling the parameter-free temperature control problem as a Markov decision process, defining the "agent-environment" interaction mechanism, and clarifying the relationship between state, action, and reward;

[0009] S3, based on the deep DQN network, a large surface source blackbody parameter-free temperature control model is constructed, and then the large surface source blackbody parameter-free temperature control model is trained using the training data set; based on the deep DQN network, the intelligent agent uses a greedy strategy to accurately adjust the duty cycle and achieve efficient target temperature control within the minimum number of steps;

[0010] S4, verification and adjustment, use the test data set to evaluate the large surface source blackbody non-parameter temperature control model, compare the deviation between the predicted temperature and the target temperature, and optimize and adjust the parameters of the large surface source blackbody non-parameter temperature control model according to the verification results to obtain the optimal large surface source blackbody non-parameter temperature control model;

[0011] S5, applying the optimal large surface source black body parameter-free temperature control model to the large surface source black body temperature control system to control the temperature of the large surface source black body.

[0012] Furthermore, the S1 comprises the following steps:

[0013] S1.1, establish a blackbody system simulation model: simulate the heating process of the blackbody system through the blackbody system simulation model, and generate a simulation temperature data set covering different temperature ranges and heating rates;

[0014] Establishing an actual temperature data acquisition system: Build a temperature acquisition system in the experimental environment, use a high-precision sensor to collect actual blackbody temperature data, and form an actual temperature data set for comparison with the simulated temperature data set;

[0015] S1.2, the simulated temperature dataset and the real temperature dataset are mixed and then divided into a training dataset and a test dataset in proportion.

[0016] Further, the S2 comprises the following steps:

[0017] S2.1, define the state space: define the state vector including the current temperature T(t), the target temperature T goal 、Ambient temperature T E and the temperature change rate △T(t), that is, the state S=[T(t),T goal ,T E ,△T(t)];

[0018] S2.2, define the action space: action A represents the duty cycle value selected by the intelligent agent of the large surface source blackbody parameter-free temperature control algorithm, and the action space is A={10,20,30,...100};

[0019] S2.3, define the reward function, which takes into account temperature error, residual heat effect, control efficiency and overshoot penalty.

[0020] Furthermore, the S2.3 also includes: when the temperature control does not reach the target temperature, the reward function formula is:

[0021]

[0022] Where R represents the feedback reward of the environment for the current action, |T(t)-T goal ∣ represents the current temperature T (t) and the target temperature T goal The absolute deviation between step is the number of actions taken by the agent; T residual Indicates the additional temperature rise caused by the residual heat effect; λ, γ, δ are the weight coefficients of the balancing action steps, the residual heat effect, and the overshoot penalty intensity, respectively; max means taking the maximum value;

[0023] If the temperature control reaches the target temperature T goal , then the reward function formula is: R=R+C, where C is a fixed high reward value.

[0024] Furthermore, the working process of the large surface source blackbody parameter-free temperature control model includes:

[0025] S3.1, in the initialization stage of the large surface source blackbody parameter-free temperature control model, the parameters of the neural network are first randomly initialized; then, the experience replay buffer is initialized, and the experience replay buffer is used to store the four-tuple of state, action, reward and next state generated during the interaction between the intelligent agent and the environment;

[0026] S3.2, in the action selection stage of the large-surface blackbody parameter-free temperature control model, a greedy strategy is adopted. Under the greedy strategy, the agent selects the action based on probability. Randomly select an action a to explore; with probability Select the action A with the largest Q value in the current state, that is:

[0027]

[0028] S3.3, in the state update stage of the large surface source black body non-parameter temperature control model, the intelligent agent calculates the influence of the duty cycle on the large surface source black body non-parameter temperature control model according to the selected action A, and updates the state of the environment to obtain the next state S':

[0029]

[0030] Where T(t+1) is the updated temperature value, T E is the ambient temperature, △T(t+1) is the new temperature change; at the same time, the agent calculates the immediate reward R based on the feedback from the environment and the set reward function; the interaction data (S, A, R, S') is then stored in the experience replay buffer M;

[0031] S3.4, in the action update phase of the large surface source black body non-parameter temperature control model, the deep DQN network minimizes the error between the target Q value and the predicted Q value by learning from past experience, thereby optimizing the large surface source black body non-parameter temperature control model and maximizing the reward;

[0032] The optimized deep DQN network loss function Loss formula is as follows. This loss function is used to train the deep DQN network to minimize the predicted value Q(s t ,a t ;θ) and target value Q target (s t ,a t θ - ) is the mean square error between:

[0033]

[0034] Among them, Q(s t ,a t ; θ) is the current deep DQN network based on state s t and action a t The output of θ is the parameter of the deep DQN network; Q target (s t ,a t θ - ) is the target value calculated by the deep DQN network, θ - It is a delayed update version of the deep DQN network; the loss function is optimized by batch size N, and the average value of the mean squared error is calculated each time to adjust the parameters θ of the deep DQN network.

[0035] The present invention also provides a large surface source black body parameter-free temperature control device, comprising one or more processors, for implementing the large surface source black body parameter-free temperature control method as described above.

[0036] The present invention also provides a readable storage medium on which a program is stored. When the program is executed by a processor, the above-mentioned large surface source blackbody parameter-free temperature control method is implemented.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] The present invention is based on a deep DQN network. Through precise state and action design, and an optimized reward function, the temperature can be quickly and accurately adjusted to the target value within the minimum number of action steps. Compared with traditional temperature control methods, the deep DQN network does not rely on predefined control parameters, but adaptively adjusts the heating duty cycle through environmental feedback, thereby improving the response speed and accuracy of the temperature control process. In addition, through experience playback and DQN value updates, the deep DQN network can continuously optimize the control strategy, adapt to complex environmental changes, and ensure that the temperature control system always maintains efficient and stable operation. The present invention does not require any parameter input. The algorithm learns the thermodynamic properties of the device through three types of data sets: heating, waste heat, and cooling, thereby achieving temperature control. Overshoot penalties are added to the constraints of the temperature control algorithm reward conditions to avoid overshoot in the temperature control process, thereby speeding up the temperature control rate.

[0039] The present invention aims to solve the problems of cumbersome parameter adjustment, slow response speed, and insufficient precision of traditional control algorithms in complex temperature control scenarios, and proposes a large surface source blackbody parameter-free temperature control method. The method is simple to operate and easy to implement in practical applications, and can dynamically adjust the control strategy to significantly improve the temperature control rate. Its advantage is that there is no need to preset complex parameters, and it can adapt to different temperature control requirements only by intelligent learning. It has extremely broad application prospects and important practical significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a flow chart of the present invention.

[0041] Figure 2 It is the “agent-environment” interaction diagram in the Markov decision process of the present invention.

[0042] Figure 3 This is the deep DQN network structure diagram of the present invention.

[0043] Figure 4 It is a framework structure diagram of the large surface source blackbody parameter-free temperature control model of the present invention.

[0044] Figure 5 It is a structural schematic diagram of a large surface source blackbody parameter-free temperature control device of the present invention. DETAILED DESCRIPTION

[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] See also Figure 1-Figure 4, a large area source blackbody parameter-free temperature control method, comprising the following steps:

[0047] S1, obtain the data set, generate the simulated temperature data set through the blackbody system simulation, and build the actual temperature acquisition system to obtain the real temperature data set, mix the simulated temperature data set with the real temperature data set and divide it into training data set and test data set in proportion to provide data support for model training and verification. Specifically include:

[0048] S1.1, establish a blackbody system simulation model: simulate the heating process of the blackbody system through the blackbody system simulation model, generate a simulation temperature data set covering different temperature ranges and heating rates, and provide basic data for subsequent model training and verification.

[0049] Establish an actual temperature data acquisition system: Build a temperature acquisition system in the experimental environment and use a high-precision sensor to collect actual blackbody temperature data to form an actual temperature data set for comparison with the simulated temperature data set.

[0050] S1.2, the simulated temperature dataset and the real temperature dataset are mixed and then divided into a training dataset and a test dataset in proportion.

[0051] S2, Problem Modeling, Using reinforcement learning for parameter-free temperature control, we first define the control problem as an agent-environment interaction problem in a Markov decision process, i.e., MDP. The parameter-free temperature control method is regarded as the environment, and the deep DQN network is regarded as the agent. The design of the system state, control action, and reward function is as follows: Figure 2 shown.

[0052] S2.1, define state space: In the deep DQN network, the more comprehensive the influencing factors considered by the state, the more complete the information received by the agent, and the more accurate the prediction results of the deep DQN network. In the present invention, the input state fully describes the current state of the temperature control method. The state vector contains the current temperature T(t), the target temperature T goal 、Ambient temperature T E and the temperature change rate △T(t), that is, the state S=[T(t),T goal ,T E ,△T(t)], which fully describes the current operating status of the temperature control method, see Table 1.

[0053]

[0054] Table 1 System state variables of large surface blackbody parameter-free temperature control method

[0055] S2.2, define the action space: Action A represents the duty cycle value selected by the agent of the deep DQN network. Further, Figure 3As shown, the deep DQN network includes a data input layer, a fully connected layer FC1, a fully connected layer FC2, an activation function ReLU1 corresponding to the fully connected layer FC1, an activation function ReLU2 corresponding to the fully connected layer FC2, a fully connected layer FC3 and an output layer.

[0056] Data input layer: receives 128-dimensional input data.

[0057] Two fully connected layers (FC1 and FC2) and corresponding activation functions (ReLU1 and ReLU2): perform feature extraction and nonlinear transformation on the input data.

[0058] Fully connected layer FC3: further processes the outputs of the first two layers to obtain the final 1-dimensional output.

[0059] Output layer: Outputs the final 1D result.

[0060] The deep DQN network first receives 128-dimensional input data, and then extracts features and nonlinearly transforms the data through two fully connected layers (FC1 and FC2) and the corresponding ReLU activation function. After processing these two hidden layers, the data is converted into a 128-dimensional intermediate representation. Finally, the deep DQN network uses the third fully connected layer FC3 to further compress these features into a 1-dimensional output. The entire network structure follows the typical fully connected neural network architecture, and the required output results are finally obtained through multi-layer feature transformation. The action space is A={10,20,30,...100}, that is, each time the agent selects a fixed duty cycle value according to the environmental state to adjust the power output of the large surface source blackbody parameter-free temperature control model. Through this discretized action design, the agent can accurately control the temperature rise process of the system within a certain range, thereby achieving precise control of the target temperature.

[0061] S2.3, the goal of the deep DQN network is to enable the large surface source blackbody parameter-free temperature control model to learn to quickly and accurately approach the target temperature T in as few action steps as possible. goal Considering the residual heat effect and overshoot problem at the same time, this problem is expanded to a multi-objective optimal control problem. The large surface source blackbody parameter-free temperature control model combines temperature error, residual heat effect, control efficiency and overshoot penalty as a reward function.

[0062] When the temperature control does not reach the target temperature, the reward function is designed as:

[0063]

[0064] Where R represents the feedback reward of the environment for the current action, |T(t)-T goal ∣ represents the current temperature T (t) and the target temperature T goalThe absolute deviation between N and N is intended to penalize inaccuracies in temperature regulation; step is the number of action steps taken by the agent, which is used to punish excessive action times and encourage the agent to achieve accurate temperature regulation through as few action steps as possible; T residual Indicates the additional temperature rise caused by the residual heat effect, which is used to quantify the influence of the residual heat on the large surface source blackbody non-parameter temperature control model; λ, γ, δ are the weight coefficients of the balancing action steps, the residual heat influence, and the overshoot penalty intensity, respectively, and max represents the maximum value. Ensure that the large surface source blackbody non-parameter temperature control model not only pursues temperature accuracy, but also minimizes redundant actions and avoids the risk of overshoot.

[0065] If the temperature control reaches the target temperature T goal , then the reward function formula is: , where C is a fixed high reward value, indicating that the task is completed and the temperature accurately reaches the expected target, which motivates the agent to complete the task correctly. Through this reward function design, the large surface source blackbody parameter-free temperature control method not only needs to reduce the temperature error, but also needs to strictly avoid overshoot, further improving the efficiency of the large surface source blackbody parameter-free temperature control model, so as to achieve accurate adjustment of the target temperature in the least number of steps as possible, thereby effectively improving the control efficiency of the large surface source blackbody parameter-free temperature control model.

[0066] S3, based on the deep DQN network, builds a large surface source blackbody parameter-free temperature control model, and then uses the training data set to train the large surface source blackbody parameter-free temperature control model. Based on the deep DQN network, the intelligent agent uses a greedy strategy to accurately adjust the duty cycle and achieve efficient target temperature control within the minimum number of steps, including:

[0067] Agent decision-making and interaction with the environment: At each time step, the agent of the deep DQN network selects an action A based on the current state, adjusts the power of the large surface source blackbody parameter-free temperature control model based on the action, updates the environmental state, obtains the next state S′, and calculates the immediate reward R through feedback.

[0068] Experience replay and strategy optimization: Use the experience replay buffer to store interaction data (S, A, R, S'), where: S represents the current state (State), which is a comprehensive description of the current state of the large surface source black body parameter-free temperature control model, such as the current temperature, target temperature, ambient temperature and temperature change rate; A represents the action (Action) selected by the agent in the current state, such as the specific value of the duty cycle (such as 10%, 20%); R represents the feedback reward (Reward) of the environment for the current action, which is used to measure the quality of the action. The reward value may be related to factors such as temperature error, action steps, and overshoot penalty; S' represents the next state (Next State) after the action is executed, such as the updated temperature value and temperature change rate. By minimizing the error between the target Q value and the predicted Q value, the deep DQN network is optimized to achieve fast and accurate temperature regulation.

[0069] S3 specifically includes the following processes:

[0070] S3.1, in the initialization stage of the large surface source blackbody parameter-free temperature control model, the parameters of the neural network are first randomly initialized, which provides initial weights for the subsequent learning process; then, the experience replay buffer is initialized, which is used to store the four-tuple of state, action, reward and next state generated during the interaction between the intelligent agent and the environment; next, it enters the action selection stage, and the large surface source blackbody parameter-free temperature control selects actions in each state according to the greedy strategy.

[0071] S3.2, in the action selection stage of the large surface source blackbody parameter-free temperature control model, a greedy strategy is adopted. Under this strategy, the agent uses probability Randomly select an action a to explore, thereby trying different strategies; with probability Select the action A with the largest Q value in the current state, that is:

[0072]

[0073] At this point, the agent preferentially uses the learned greedy strategy to select actions to maximize the expected reward.

[0074] S3.3, in the state update stage of the large surface source blackbody parameter-free temperature control model, the agent calculates the influence of the duty cycle on the large surface source blackbody parameter-free temperature control according to the selected action A, and updates the state of the environment to obtain the next state S':

[0075]

[0076] Where T(t+1) is the updated temperature value, T Eis the ambient temperature, △T(t+1) is the new temperature change; at the same time, the agent calculates the immediate reward R according to the feedback from the environment and the set reward function, which reflects the effect of the current action on the large-surface blackbody parameter-free temperature control model; the interaction data (S, A, R, S') is then stored in the experience replay buffer M to provide data support for subsequent learning.

[0077] S3.4, in the action update stage of the large surface source blackbody parameter-free temperature control model, the deep DQN network optimizes the temperature control method and maximizes the reward by minimizing the error between the target Q value and the predicted Q value by learning from past experience.

[0078] The optimized deep DQN network loss function Loss formula is as follows. This loss function is used to train the deep DQN network to minimize the predicted value Q(s t ,a t ;θ) and target value Q target (s t ,a t θ - ) is the mean square error between:

[0079]

[0080] Among them, Q(s t ,a t ; θ) is the current deep DQN network based on state s t and action a t The output of θ is the parameter of the deep DQN network; Q target (s t ,a t θ - ) is the target value calculated by the deep DQN network, θ - It is a delayed update version of the deep DQN network. The loss function is optimized by the batch size N. The average value of the mean square error is calculated each time to adjust the parameters θ of the deep DQN network, thereby gradually approaching the true value of the state-action pair. During the training process, historical data is sampled through the experience replay mechanism to reduce the correlation between samples and improve the generalization ability of the method.

[0081] By continuously learning the heating, waste heat and cooling characteristics in simulation and real data sets, the deep DQN network can gradually master the optimal temperature control method and achieve fast and accurate temperature regulation in practical applications.

[0082] S4, verification and adjustment: In the verification stage, the effect of the trained large-surface blackbody parameter-free temperature control model in actual applications is evaluated through the test data set, the deviation between the predicted temperature and the target temperature is compared, and the method parameters are adjusted according to the results. By comparing the simulation data and the actual data, the generalization ability of the model and its adaptability in the real environment are verified. At the same time, error analysis is performed on the prediction deviation of the model, and the reward function and control strategy are adjusted to ensure that the model can operate stably under changing environmental conditions. Finally, by adjusting the strategy, the performance of the model is optimized so that the temperature control effect in actual applications meets the expectations.

[0083] S5, iterative optimization: Repeat S3 and S4 until the deep DQN network can accurately adjust the temperature to the target temperature T in the least number of steps goal , and maintain the stability and efficiency of the large surface source blackbody temperature control model. In the iterative optimization process, continuous training and verification are carried out, and the decision-making ability of the deep DQN network is improved through experience replay and greedy strategy to ensure the robustness and optimality of the temperature control model. Ultimately, the intelligent agent can quickly and accurately adjust the temperature according to different environmental changes and needs, and while ensuring control accuracy, it minimizes redundant operation steps and improves the control efficiency and response speed of the model.

[0084] Finally, the optimal large surface source blackbody non-parameter temperature control model is applied to the large surface source blackbody temperature control system to control the temperature of the large surface source blackbody.

[0085] See also Figure 5 The present invention provides a large surface source black body parameter-free temperature control device, which includes one or more processors for implementing a large surface source black body parameter-free temperature control method in the above embodiment.

[0086] An embodiment of a large surface source blackbody parameter-free temperature control device of the present invention can be applied to any device with data processing capability, and the device with data processing capability can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capability in which it is located reading the corresponding computer program instructions in the non-volatile memory into the internal memory for execution. From the hardware level, if Figure 5 As shown, it is a hardware structure diagram of any device with data processing capability where a large surface source blackbody parameter-free temperature control device of the present invention is located, except Figure 5 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in the embodiments may also include other hardware, usually based on the actual functions of the device with data processing capabilities, which will not be described in detail.

[0087] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0088] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0089] The present invention also provides a readable storage medium on which a program is stored. When the program is executed by a processor, a large surface source blackbody parameter-free temperature control method in the above embodiment is implemented.

[0090] The readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The readable storage medium may also be an external storage device, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the readable storage medium may also include both an internal storage unit of any device with data processing capability and an external storage device. The readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.

[0091] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A large surface source blackbody parameter-free temperature control method, characterized in that: The following steps are involved: S1, obtain a data set, generate a simulated temperature data set through blackbody system simulation, and build an actual temperature acquisition system to obtain a real temperature data set, mix the simulated temperature data set with the real temperature data set, and divide it into a training data set and a test data set in proportion; S2, problem modeling, modeling the parameter-free temperature control problem as a Markov decision process, defining the "agent-environment" interaction mechanism, and clarifying the relationship between state, action, and reward, including the following steps: S2.1, define the state space: define the state vector including the current temperature T(t), the target temperature T goal 、Ambient temperature T E and the temperature change rate △T(t), that is, the state S=[T(t),T goal ,T E ,△T(t)]; S2.2, define the action space: action A represents the duty cycle value selected by the intelligent agent of the large surface source blackbody parameter-free temperature control algorithm, and the action space is A={10,20,30,...100}; S2.3, define a reward function that takes into account temperature error, residual heat effect, control efficiency, and overshoot penalty; When the temperature control does not reach the target temperature, the reward function formula is: , Where R represents the feedback reward of the environment for the current action, |T(t)-T goal ∣ represents the current temperature T (t) and the target temperature T goal The absolute deviation between step is the number of actions taken by the agent; T residual Indicates the additional temperature rise caused by the residual heat effect; λ, γ, δ are the weight coefficients of the balancing action steps, the residual heat effect, and the overshoot penalty intensity, respectively; max means taking the maximum value; If the temperature control reaches the target temperature T goal , then the reward function formula is: R=R+C, where C is a fixed high reward value; S3, based on the deep DQN network, a large surface source blackbody parameter-free temperature control model is constructed, and then the large surface source blackbody parameter-free temperature control model is trained using the training data set; based on the deep DQN network, the intelligent agent uses a greedy strategy to accurately adjust the duty cycle and achieve efficient target temperature control within the minimum number of steps; The working process of the large surface source blackbody parameter-free temperature control model includes: S3.1, in the initialization stage of the large surface source blackbody parameter-free temperature control model, the parameters of the neural network are first randomly initialized; then, the experience replay buffer is initialized, and the experience replay buffer is used to store the four-tuple of state, action, reward and next state generated during the interaction between the intelligent agent and the environment; S3.2, in the action selection stage of the large-surface blackbody parameter-free temperature control model, a greedy strategy is adopted. Under the greedy strategy, the agent selects the action based on probability. Randomly select an action a to explore; with probability Select the action A with the largest Q value in the current state, that is: , S3.3, in the state update stage of the large surface source black body non-parameter temperature control model, the intelligent agent calculates the influence of the duty cycle on the large surface source black body non-parameter temperature control model according to the selected action A, and updates the state of the environment to obtain the next state S': , Where T(t+1) is the updated temperature value, T E is the ambient temperature, △T(t+1) is the new temperature change; at the same time, the agent calculates the immediate reward R based on the feedback from the environment and the set reward function; the interaction data (S, A, R, S') is then stored in the experience replay buffer M; S3.4, in the action update phase of the large surface source black body non-parameter temperature control model, the deep DQN network minimizes the error between the target Q value and the predicted Q value by learning from past experience, thereby optimizing the large surface source black body non-parameter temperature control model and maximizing the reward; The optimized deep DQN network loss function Loss formula is as follows. This loss function is used to train the deep DQN network to minimize the predicted value Q(s t ,a t ;θ) and target value Q target (s t ,a t θ - ) is the mean square error between: , Among them, Q(s t ,a t ; θ) is the current deep DQN network based on state s t and action a t The output of θ is the parameter of the deep DQN network; Q target (s t ,a t θ - ) is the target value calculated by the deep DQN network, θ - It is a delayed update version of the deep DQN network. The loss function is optimized by batch size N, and the average value of the mean square error is calculated each time to adjust the parameter θ of the deep DQN network. S4, verification and adjustment, use the test data set to evaluate the large surface source blackbody non-parameter temperature control model, compare the deviation between the predicted temperature and the target temperature, and optimize and adjust the parameters of the large surface source blackbody non-parameter temperature control model according to the verification results to obtain the optimal large surface source blackbody non-parameter temperature control model; S5, applying the optimal large surface source black body parameter-free temperature control model to the large surface source black body temperature control system to control the temperature of the large surface source black body.

2. A large surface source blackbody parameter-free temperature control method as claimed in claim 1, characterized in that: The S1 comprises the following steps: S1.1, establish a blackbody system simulation model: simulate the heating process of the blackbody system through the blackbody system simulation model, and generate a simulation temperature data set covering different temperature ranges and heating rates; Establishing an actual temperature data acquisition system: Build a temperature acquisition system in the experimental environment, use a high-precision sensor to collect actual blackbody temperature data, and form an actual temperature data set for comparison with the simulated temperature data set; S1.2, the simulated temperature dataset and the real temperature dataset are mixed and then divided into a training dataset and a test dataset in proportion.

3. A large surface source blackbody parameter-free temperature control device, characterized in that: It includes one or more processors for implementing the large surface source blackbody parameter-free temperature control method described in claim 1 or 2.

4. A readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, a large surface source blackbody parameter-free temperature control method as described in claim 1 or 2 is implemented.

Citation Information

Patent Citations

  • Multi-energy collaborative complementary optimization method and device based on target value competition and medium

    CN117856258A

  • Robust reinforcement learning reactive power optimization method suitable for unobservable power distribution network

    CN118199078A