Multi-microwave-source combined heating method for agent collaborative optimization under composite target consensus decision
By adopting the method of collaborative optimization of the agent under the consensus decision of composite targets in the microwave heating system, dynamically adjusting the power distribution of the microwave source, solving the problems of uneven resource allocation and multi-target conflict in traditional systems, and achieving efficient and uniform heating effects.
Patent Information
- Application Number
- CN202510327986.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-20
AI Technical Summary
Traditional microwave heating systems are difficult to coordinate power distribution and frequency selection in a multi-microwave environment, resulting in uneven resource allocation and multi-target conflicts, and cannot achieve efficient and uniform heating effects.
The multi-microwave source joint heating method with coordinated optimization of agents under composite target consensus decisions is adopted. The master-slave targets and constraints are set through the simulation platform, the agent is initialized, the conflict detection mechanism is established, and the power allocation of microwave sources is dynamically adjusted using the DDPG algorithm and the winner-take-all strategy to achieve multi-objective coordinated optimization.
It effectively improves the overall comprehensive benefits of microwave heating, achieves an effective balance between temperature uniformity and heating efficiency, and improves the stability and reliability of the system in complex situations.
Smart Images

Figure CN120180738A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-microwave source combined heating method for agent collaborative optimization under compound target consensus decision-making, belonging to the technical field of microwave heating. Background Art
[0002] Microwave heating technology is an efficient, fast and energy-saving heating method, which is widely used in industrial heating, food processing, medical disinfection, material processing and other fields. Compared with traditional heating methods, microwave heating can achieve simultaneous heating of the interior and surface of materials, and its high efficiency and uniformity have received wide attention. However, due to the complexity of the heating object and the need for multi-objective optimization, the traditional microwave heating system has the following problems:
[0003] Resource allocation problem: In a multi-microwave source environment, different microwave sources contribute differently to the heating area. How to coordinate the power allocation and frequency selection of multiple microwave sources has always been a difficult problem in the microwave heating process. Most of the existing technologies are based on simple allocation strategies with fixed power, lacking dynamic adaptability to complex environments.
[0004] Multi-objective conflict problem: The objectives in the microwave heating process usually include heating efficiency, temperature uniformity, etc., and there may be contradictions between these objectives. For example, improving heating efficiency often leads to uneven temperature distribution in the heating area, while optimizing temperature uniformity may reduce the overall heating efficiency. It is difficult for existing technologies to achieve dynamic balance among multiple objectives.
[0005] Insufficient intelligent optimization: Existing microwave heating systems mostly rely on traditional algorithms or manual adjustment, and cannot dynamically adjust strategies according to the real-time heating state, resulting in low heating process efficiency and difficulty in meeting the requirements of complex environments.
[0006] In response to the above problems, intelligent optimization algorithms (such as deep reinforcement learning) introduced in recent years have provided a new direction for solving complex optimization problems in microwave heating. Through the simulation platform combined with the optimization algorithm, breakthroughs can be achieved in multi-objective optimization and dynamic adaptability. For example, the Deep Deterministic Policy Gradient algorithm (DDPG) can generate efficient decisions in a continuous action space, while the Winner-Takes-All strategy can solve resource conflict problems in multi-agent systems. However, there is still a lack of a method in the existing technology that can comprehensively utilize the simulation platform, master-slave target optimization and intelligent algorithms to solve resource conflict and multi-objective collaborative optimization problems in the microwave heating process.
[0007] Therefore, there is an urgent need for a microwave heating method based on master-slave target optimization, which can dynamically adjust the power allocation and frequency selection of microwave sources on the simulation platform, solve multi-objective conflicts through intelligent algorithms, and achieve efficient and uniform heating effects. Summary of the Invention
[0008] The technical problem solved by the present invention is: the present invention provides a multi-microwave source joint heating method for collaborative optimization of agents under composite target consensus decision-making, which is used to solve the resource conflict and multi-objective collaborative optimization problems in the microwave heating process; the method of the present invention realizes multi-objective collaborative optimization and conflict resolution in the microwave heating process, and effectively improves the microwave heating performance.
[0009] The technical solution of the present invention is: a multi-microwave source joint heating method for collaborative optimization of agents under composite target consensus decision-making, the method comprising:
[0010] Step1. Construct a simulation platform for the microwave heating environment;
[0011] Step2. Set the optimization objectives and constraints; when setting the optimization objectives, include setting the main objective and the secondary objectives;
[0012] Step3. Initialize the agents; among them, assign global agents and local agents to the main objective and the secondary objectives;
[0013] Step4. Establish a conflict detection mechanism for judging the current execution state of the agents;
[0014] Step5. Design the DDPG algorithm corresponding to the execution actions of the global agent and the local agents according to the main objective and the secondary objectives;
[0015] Step6. According to the conflict detection mechanism, the global agent and the local agents will be divided into two states: non-conflict state and conflict state during the execution of the target tasks;
[0016] In the non-conflict state, the global agent and the local agents generate control actions according to the main objective and the secondary objectives through the DDPG algorithm;
[0017] When a conflict is triggered, adopt the winner-takes-all strategy to optimize the resource allocation between the main objective and the secondary objectives, and re-give the execution actions of the global agent and the local agents based on this resource allocation;
[0018] Step7. By dynamically adjusting the actions of the global agent and the local agents, form the action sets of different agents, conduct agent network training, and dynamically reconstruct the experience pool of the agents, so as to improve the training effect.
[0019] Further, the Step1 includes:
[0020] Step1.1. Use the simulation platform to simulate the microwave heating environment and generate a thermal effect model of a three-dimensional multi-microwave source; including determining the geometric structure of the cavity of the microwave heating system, the configuration of the microwave sources, the definition of the heating area, and determining the physical properties of the material to be heated;
[0021] The geometric structure of the cavity includes the cuboid structure of the cavity, with the width X, length Y, and height Z set.
[0022] The configuration of the microwave source includes the configuration of the number, type, and position of the microwave source.
[0023] The physical properties of the material to be heated include dielectric constant, magnetic permeability, conductivity, and heat capacity.
[0024] Step1.2. Collect simulation data, including the intensity and action range of the microwave source and the state parameters S of the heated object t ;
[0025] The thermal effect model of the three-dimensional multi-microwave source simulates the microwave power distribution and the material temperature distribution. Set the power and frequency parameters of each microwave source as optimization variables, and simulate the microwave heating process, which includes the distribution of the temperature field and the transfer of power.
[0026] Furthermore, the said Step2 includes:
[0027] Step2.1. Set the optimization objectives; take the temperature uniformity as the main objective and the heating efficiency as the secondary objective; define the main objective as the core of optimization and assign it a high priority; take the secondary objective as the auxiliary optimization objective and assign it a sub-optimal priority.
[0028] Setting the optimization objectives includes:
[0029] Step2.1.1. Set the initial temperature T0 and the target heating duration T of the heated material f , and take the coefficient of variation (COV) of the temperature uniformity evaluation index as the main objective. The calculation formula for the main objective is: where, T a is the average temperature, and T i is the temperature at the temperature sampling point;
[0030] Step2.1.2. Set the heating efficiency as the secondary objective;
[0031] The calculation formula for the heating efficiency is: is the energy absorbed by the material, E input is the total input energy, and E absorbed is the energy absorbed by the object to be heated;
[0032] Step2.2. Set the constraint conditions, including that the total power of the microwave source cannot exceed the set maximum power;
[0033] Among them, the constraint condition for the total power of the microwave source during the microwave heating process is: where, P0 is the total power of the microwave source, and P iis the power of the i-th microwave source, and n is the number of microwave sources.
[0034] Further, the said Step3 includes:
[0035] Step3.1. Each microwave source acts as an independent agent to control its power output; among which, it includes:
[0036] Step3.1.1. Set the agent matching the main target as the global agent, which is responsible for generating global priority actions;
[0037] Step3.1.2. Set the agent matching the subordinate target as the local agent, which is used to flexibly execute auxiliary tasks but is subject to the actions of the global agent;
[0038] Step3.2. The agent obtains the power state of the current microwave source and the temperature field data of the heated material in real time through the simulation platform.
[0039] Further, the said Step4 includes:
[0040] Step4.1. Detect target conflicts:
[0041] Based on the resource requirements of the main target and the subordinate target, calculate the conflict metric D between the main target and the subordinate target:
[0042] D = w1COV desire + w2U desire , where w1 and w2 are the main-subordinate target discrimination coefficients used to determine whether the corresponding target is the main target, COV desire is the ideal temperature uniformity, and U desire is the ideal heating efficiency; if D > δ, trigger the conflict resolution mechanism, where δ is the conflict threshold, otherwise it is regarded as no conflict occurred.
[0043] Further, the said Step5 includes:
[0044] Step5.1. Define the state space and action space of the global agent and the local agent;
[0045] Step5.2. Network structure design: Both the global agent and the local agent adopt Actor-Critic as the agent network structure;
[0046] Step5.3. Reward function design: Design the corresponding reward functions for the global agent and the local agent according to the main and subordinate targets.
[0047] Further, the said Step6 includes:
[0048] Judge whether the current execution state of the agent conflicts based on the conflict metric:
[0049] Non - conflict state D < δ:
[0050] When it is detected that there is no conflict between the main target and the secondary target, the global agent and the local agent respectively follow the action strategies given by the DDPG algorithm By dynamically adjusting the power distribution of the microwave source, the priority optimization of the main target is completed to ensure that the main target is preferentially achieved under the current state parameters S of the heated object t and as much as possible take into account the secondary optimization requirements of the secondary target;
[0051] Conflict state D > δ:
[0052] Use the winner - takes - all strategy to solve the target conflict, including:
[0053] Under conflict conditions, coordinate the resource allocation between the main target and the secondary target through the winner - takes - all strategy; specifically as follows:
[0054] In case of conflict, allocate power according to the priority of the main target while ensuring the minimum interference to the secondary target; the optimization problem is defined as:
[0055]
[0056] where: R t represents the sum of the episode rewards, represents the reward of the main target, represents the reward of the secondary target, represents the global agent of the DDPG algorithm performing actions with the main target as the object, represents the local agent of the DDPG algorithm performing actions with the secondary target; C sup (A) represents the interference cost of the secondary target, a1 and a2 respectively represent the weights of the main and secondary targets, and satisfy a1 + a2 = 1;
[0057] The power allocation scheme is:
[0058] Allocate the global power to the main target first to preferentially meet the power requirements of the main target; the process of allocating the global power to the main target is expressed as:
[0059]
[0060] where represents the power allocation of the global agent, that is, the action executed by the global agent in the conflict state, P0 represents the global power constraint. After allocating the power to the main target, allocate the remaining power to the secondary target; the process of allocating the remaining power to the secondary target is expressed as:
[0061]
[0062] Among them represents the power distribution of the local agent, that is, the execution action of the local agent in the conflict state.
[0063] Furthermore, the said Step7 includes:
[0064] The DDPG algorithm will use the conflict action set and the non-conflict action set to form the target action set as the initial action for the next round. In the DDPG algorithm, by continuously repeating the above operations with R t as the target, maximizing the function value, and adjusting the execution action to optimize the heating effect.
[0065] Furthermore, the said Step7 includes:
[0066] Establish a conflict and non-conflict motivation experience pool for storing the current state parameters S of the object to be heated t , the sum of round rewards R t and the state parameters S of the object to be heated after adjustment t+1 ;
[0067] During the execution of the DDPG algorithm, batch samples are proportionally drawn from the conflict experience pool B H and the non-conflict experience pool B′ H to form the experience pool Batch:
[0068]
[0069] where N1 and N2 are the sampling ratios;
[0070] The Critic network is updated in stages:
[0071]
[0072] Among them, represents the state-action set Q(S,A|θ Q ): The Critic network evaluates the quality of the state-action pair according to the Q value predicted by the current parameter θ Q ; y t is the target Q value, is the mathematical expectation, representing the average of all samples in the experience pool Batch;
[0073] The Actor network is updated:
[0074]
[0075] μ(S t |θ μ): The action generated by the Actor network according to the parameter θ μ The generated action;
[0076] q(S t , μ(S t )): The Q-value evaluation of the action generated by the Actor by the Critic network;
[0077] Soft update of the target network:
[0078] θ Q′ ← τθ Q + (1 - τ)θ Q′
[0079] θ μ′ ← τθ μ + (1 - τ)θ μ′
[0080] θ Q′ 、θ μ′ represent the parameters of the target Critic and Actor networks, and τ represents the soft update coefficient.
[0081] The beneficial effects of the present invention are:
[0082] 1. Synergy of multi-objective optimization: By setting the temperature uniformity as the main objective and the heating efficiency as the secondary objective, and assigning different priorities, combined with the collaborative optimization of the global agent and the local agent, on the basis of ensuring the priority achievement of temperature uniformity, the heating efficiency is improved as much as possible, achieving an effective balance and collaborative optimization between multiple objectives, and improving the overall comprehensive benefit of microwave heating.
[0083] 2. Effectiveness of agent division of labor and conflict resolution: Regarding the microwave source as an independent agent, clearly dividing the responsibilities of the global agent and the local agent, the global agent is responsible for generating global actions, and the local agent flexibly executes auxiliary tasks under the constraints of the global agent's actions. At the same time, by establishing conflict trigger conditions and the winner-takes-all strategy, the resource allocation conflict between the primary and secondary objectives can be detected and effectively resolved in a timely manner, ensuring that when a conflict occurs, the power is reasonably allocated according to the objective priority, guaranteeing the smooth achievement of the primary objective, and improving the stability and reliability of the system in complex situations.
[0084] 3. Precision and adaptability of the optimization algorithm: Using the DDPG algorithm to give the global agent's execution actions according to the primary objective, and dynamically adjusting the execution actions of the global agent and the local agent at different stages of conflict and non-conflict, and continuously adjusting the execution actions according to the target action set and the sum of episode rewards to achieve the optimization of the heating effect. Description of the Drawings
[0085] Figure 1 is the flowchart in the present invention;
[0086] Figure 2 This is the cavity model of the temperature field for the combined heating of multiple microwave sources in the embodiments of the present invention;
[0087] Figure 3 This is the temperature field slice diagram of the present invention; among them, (a) is the vertical slice diagram, and (b) is the horizontal slice diagram. Detailed implementation manners
[0088] Embodiment 1: As Figures 1 - 3 shown, a method for the combined heating of multiple microwave sources with intelligent agent collaborative optimization under composite target consensus decision-making, the method comprising:
[0089] Step1. Construct a simulation platform for the microwave heating environment;
[0090] Further, the Step1 includes:
[0091] Step1.1. Use the simulation platform to simulate the microwave heating environment, as Figure 2 shown, generate a thermal effect model of three-dimensional multiple microwave sources; including determining the geometric structure of the cavity of the microwave heating system, the configuration of the microwave sources, the definition of the heating area, and determining the physical properties of the material to be heated;
[0092] The geometric structure of the cavity includes the cuboid structure of the cavity, and the width X, length Y, and height Z are set;
[0093] The configuration of the microwave sources includes the number of microwave sources (such as Figure 2 the microwave sources 1-4 in), the type, and the configuration of the positions;
[0094] The physical properties of the material to be heated include dielectric constant, magnetic permeability, conductivity, and heat capacity;
[0095] Step1.2. Collect simulation data, including the intensity of the microwave sources, the action range, and the state parameters S of the object to be heated t ;
[0096] The thermal effect model of three-dimensional multiple microwave sources simulates the microwave power distribution and the material temperature distribution, sets the power and frequency parameters of each microwave source as optimization variables, and simulates the process of microwave heating, and the process of microwave heating includes the distribution of the temperature field and the transfer of power.
[0097] Step2. Set the optimization objectives and constraints; when setting the optimization objectives, include setting the main objective and the secondary objective;
[0098] Further, the Step2 includes:
[0099] Step 2.1. Set the optimization objectives; regard the temperature uniformity as the main objective and the heating efficiency as the secondary objective. Define the main objective as the core of optimization and assign it a high priority; regard the secondary objective as an auxiliary optimization objective and assign it a sub-optimal priority.
[0100] Setting the optimization objectives includes:
[0101] Step 2.1.1. Set the initial temperature T0 of the heated material and the target heating duration T f , and regard the coefficient of variation (COV) of the temperature uniformity evaluation index as the main objective. The calculation formula for the main objective is: where, T a is the average temperature, and T i is the temperature at the temperature sampling point;
[0102] Step 2.1.2. Set the heating efficiency as the secondary objective;
[0103] The calculation formula for the heating efficiency is: where, E input is the energy absorbed by the material, E absorbed is the total input energy, and E
[0104] Step 2.2. Set the constraint conditions, including that the total power of the microwave source cannot exceed the set maximum power;
[0105] Among them, the constraint condition for the total power of the microwave source during the microwave heating process is: where, P0 is the total power of the microwave source, P i is the power of the i-th microwave source, and n is the number of microwave sources.
[0106] Step 3. Initialize the agent; among them, assign global agents and local agents to the main objective and the secondary objective;
[0107] Furthermore, the said Step 3 includes:
[0108] Step 3.1. Each microwave source is regarded as an independent agent to control its power output; among them, it includes:
[0109] Step 3.1.1. Set the agent matching the main objective as the global agent, which is responsible for generating globally prioritized actions;
[0110] Step 3.1.2. Set the agent matching the secondary objective as the local agent, which is used to flexibly execute auxiliary tasks but is subject to the actions of the global agent;
[0111] Step3.2. The agent obtains the power status of the current microwave source and the temperature field data of the heated material in real time through the simulation platform. Perform moving average filtering, normalization, and denormalization on the collected temperature field data to convert the data into values within the range of [0, 1].
[0112] Step4. Set up a conflict detection mechanism to judge the current execution status of the agent;
[0113] Further, the Step4 includes:
[0114] Step4.1. Detect target conflicts:
[0115] Based on the resource requirements of the primary target and the secondary target, calculate the conflict metric D between the primary target and the secondary target:
[0116] D = w1COV desire + w2U desire , where w1 and w2 are the primary-secondary target discrimination coefficients used to judge whether the corresponding target is the primary target, COV desire is the ideal temperature uniformity, and U desire is the ideal heating efficiency; if D > δ, trigger the conflict resolution mechanism, where δ is the conflict threshold, otherwise it is regarded as no conflict occurred.
[0117] Step5. Design the DDPG algorithms corresponding to the execution actions of the global agent and the local agent according to the primary target and the secondary target;
[0118] Further, the Step5 includes:
[0119] Step5.1. Define the state space and action space of the global agent and the local agent;
[0120] Step5.2. Network structure design: Both the global agent and the local agent adopt Actor-Critic as the agent network structure;
[0121] Step5.3. Reward function design: Design the reward functions corresponding to the global agent and the local agent according to the primary and secondary targets.
[0122] Step6. According to the conflict detection mechanism, the global agent and the local agent will be divided into two states: non-conflict state and conflict state during the execution of the target task;
[0123] In the non-conflict state, the global agent and the local agent generate control actions according to the primary target and the secondary target through the DDPG algorithm;
[0124] When a conflict is triggered, a winner-takes-all strategy is adopted to optimize the resource allocation between the primary goal and the secondary goal, and based on this resource allocation, the execution actions of the global agent and the local agent are re-given.
[0125] Further, the Step6 includes:
[0126] Based on the conflict metric, determine whether the current execution state of the agent has a conflict:
[0127] Non-conflict state D < δ:
[0128] When it is detected that there is no conflict between the primary goal and the secondary goal, the global agent and the local agent respectively follow the action strategy given by the DDPG algorithm By dynamically adjusting the power allocation of the microwave source, the priority optimization of the primary goal is completed to ensure that the primary goal is preferentially achieved under the current state parameters S of the object to be heated t and as much as possible take into account the secondary optimization requirements of the secondary goal;
[0129] Conflict state D > δ:
[0130] Use the winner-takes-all strategy to solve the goal conflict, including:
[0131] Under conflict conditions, coordinate the resource allocation between the primary goal and the secondary goal through the winner-takes-all strategy; specifically as follows:
[0132] In case of conflict, allocate power according to the primary goal priority while ensuring the minimum interference to the secondary goal; the optimization problem is defined as:
[0133]
[0134] Where: R t represents the sum of the episode rewards, represents the reward of the primary goal, represents the reward of the secondary goal, represents the action executed by the global agent of the DDPG algorithm with the primary goal as the object, represents the action executed by the local agent of the DDPG algorithm with the secondary goal as the object; C sup (A) represents the interference cost of the secondary goal, and a1 and a2 respectively represent the weights of the primary and secondary goals, and satisfy a1 + a2 = 1;
[0135] The power allocation scheme is:
[0136] Allocate the global power to the primary goal first to preferentially meet the power demand of the primary goal; the process of allocating the global power to the primary goal is expressed as:
[0137]
[0138] Where It represents the power distribution of the global agent, that is, the execution action of the global agent in the conflict state. P0 represents the global power constraint. After the main target power is distributed, the remaining power is distributed to the subordinate targets. The process of distributing the remaining power to the subordinate targets is expressed as:
[0139]
[0140] Among them It represents the power distribution of the local agent, that is, the execution action of the local agent in the conflict state.
[0141] Step7. By dynamically adjusting the actions of the global agent and the local agent, an action set of different agents is formed for the training of the agent network. Through the experience pool that can be dynamically reconstructed for the agent, the training effect is improved.
[0142] Furthermore, the said Step7 includes:
[0143] The DDPG algorithm will use the conflict action set and the non-conflict action set to form the target action set as the initial action for the next round. In the DDPG algorithm, by continuously repeating the above operations with R t as the target, the function value is maximized, and the execution action is adjusted to optimize the heating effect.
[0144] Furthermore, the said Step7 includes:
[0145] Establish a conflict and non-conflict motivation experience pool for storing the current state parameters S of the object to be heated t , the sum of round rewards R t and the state parameters S of the object to be heated after adjustment t+1 ;
[0146] During the execution of the DDPG algorithm, batch samples are proportionally sampled from the conflict experience pool B H and the non-conflict experience pool B' H to form the experience pool Batch:
[0147]
[0148] where N1 and n2 are the sampling ratios;
[0149] The Critic network is updated in stages:
[0150]
[0151] Among them, represents the state-action set q(S,A|θQ ):The Critic network evaluates the quality of the state-action pair according to the current parameter θ Q using the predicted Q value; y t is the target Q value, is the mathematical expectation, representing the average of all samples in the experience pool Batch;
[0152] Actor network update:
[0153]
[0154] μ(S t |θ μ ):The Actor network generates an action according to the parameter θ μ ;
[0155] Q(S t ,μ(S t )): The Critic network evaluates the Q value of the action generated by the Actor;
[0156] Soft update of the target network:
[0157] θ Q′ ← τθ Q +(1 - τ)θ Q′
[0158] θ μ′ ← τθ μ +(1 - τ)θ μ′
[0159] θ Q′ 、θ μ′ represent the parameters of the target Critic and Actor networks, and τ represents the soft update coefficient.
[0160] Example 2: The method of this example is the same as that of Example 1, except that in this example, silicon carbide crystals are used as the heating material to verify the temperature uniformity optimization method;
[0161] In the cavity structure of this example, the cavity is a cuboid, the X-axis is the width direction, the Y-axis is the length direction, and the Z-axis is the height direction; an air domain is set in the cavity as a microwave energy exchange place;
[0162] The microwave source and boundary conditions of this example are: 4 microwave incident ports are set, and microwaves are injected into the cavity through the incident ports, and the cavity boundary is made of metal;
[0163] Simulation and optimization of this example:
[0164] A three-dimensional multi-microwave source thermal effect model is established through COMSOL software;
[0165] The DDPG algorithm is used in Matlab to iteratively optimize the microwave source power allocation;
[0166] The joint simulation platform updates the temperature field data after each round of iteration and adjusts the power allocation strategy for the next round based on the reward function;
[0167] The optimization results are shown in Tables 1 to 3. The results show that the obtained optimal power combination effectively improves the temperature uniformity COV and heating efficiency U, while meeting the total power constraint;
[0168] The quantitative analysis uses a hierarchical statistical method to calculate the temperature uniformity evaluation index for each cutting plane ( Given in Step 2.1.1), the smaller the COV value, the more uniform the temperature distribution. The uniformity evaluation slicing method is as follows Figure 3 As shown, the vertical section COV value distribution is shown in Table 1, the horizontal section COV value is summarized in Table 2, and the heating efficiency parameter U comparison data of each scheme is listed in Table 3. At the same time, the heating efficiency evaluation index is introduced The heating efficiency results are shown in Table 3;
[0169] Table 1 shows the vertical cross-section COV
[0170]
[0171] Table 2 shows the COV of horizontal sections.
[0172]
[0173] Table 3 shows the heating efficiency of different strategies
[0174]
[0175] As shown in Table 3, Experiment 1 is the traditional four-microwave source constant power heating, and Experiment 2 is the multi-microwave source joint heating optimization method proposed in this paper based on the global and local intelligent agent collaborative optimization under the composite multi-objective consensus decision. Compared with the traditional heating method, the method of the present invention improves the uniformity of the vertical section by 8.6%-51.7%, the heating efficiency of the horizontal section by 4.3%-36.22%, and the heating efficiency by 1.51%. By comparison, it can be seen that the present invention is far superior to the traditional heating effect in the regulation of multi-objective heating results.
[0176] Embodiment 3: This embodiment provides a multi-source microwave heating optimization device, which includes the following modules:
[0177] Processor module:
[0178] It includes a main processor and a parallel computing unit, which execute the optimization algorithm of the microwave heating system and control the power distribution of the microwave source in real time;
[0179] Run the Deep Deterministic Policy Gradient (DDPG) algorithm with value distribution and the winner-takes-all strategy to solve multi-objective conflicts and resource allocation problems;
[0180] Support high-frequency real-time data processing, be able to quickly respond to environmental changes, and update the decisions of the agent in real time.
[0181] Memory module:
[0182] It includes a cache, a main memory, and a long-term memory, which are used to store simulation data, model parameters, and optimization algorithms;
[0183] Store the cavity model of the microwave heating system, the parameters of the heated material (such as dielectric constant, magnetic permeability, conductivity, etc.), and the historical data of the heating process;
[0184] The experience replay buffer stores the quadruple (S t , A t , R t , S t+1 ) during the optimization process for the neural network training of the agent.
[0185] Simulation platform module:
[0186] Establish three-dimensional electromagnetic field and thermal models of the microwave heating system by the finite element method (FEM);
[0187] Dynamically simulate the influence of microwave source power distribution on the temperature distribution of the material and provide real-time feedback data;
[0188] Support multi-objective evaluation, including the coefficient of variation (COV) of temperature uniformity, energy utilization rate (U), and heating efficiency (E).
[0189] Data acquisition and sensing module:
[0190] Equipped with high-precision sensors to collect the real-time temperature distribution of the material in the heating cavity;
[0191] Transmit data to the processor module through wireless or wired communication interfaces, supporting real-time feedback and decision adjustment.
[0192] Control module:
[0193] According to the power distribution scheme generated by the optimization algorithm, control the power output and frequency of each microwave source;
[0194] Support fine-grained adjustment of power and frequency to ensure the accurate implementation of the optimization scheme.
[0195] User interaction module:
[0196] Provide a human-machine interaction interface that allows users to set heating targets (such as temperature uniformity and energy efficiency priorities), monitor the heating process, and view optimization results;
[0197] Support the visualization of optimization strategies, including temperature distribution, power allocation, and target achievement.
[0198] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. A multi-microwave source joint heating method with intelligent agent collaborative optimization under composite target consensus decision-making, characterized by: The method comprises: Step 1. Build a simulation platform for microwave heating environment; Step 2, set the optimization goal and constraints; setting the optimization goal includes setting the main goal and the secondary goal; Step 3, agent initialization; in which global agents and local agents are assigned to the main target and the sub-target; Step 4: Establish a conflict detection mechanism to determine the current execution status of the agent; Step 5. Design the DDPG algorithm corresponding to the execution actions of the global agent and the local agent based on the main goal and the sub-goal; Step 6: According to the conflict detection mechanism, the global agent and the local agent will be divided into two states: non-conflict state and conflict state in the process of executing the target task; In the non-conflict state, the global agent and the local agent generate control actions based on the main goal and the sub-goal through the DDPG algorithm; When a conflict is triggered, a winner-takes-all strategy is used to optimize the resource allocation between the primary and secondary goals, and based on this resource allocation, the execution actions of the global agent and the local agent are re-given; Step 7, by dynamically adjusting the actions of the global agent and the local agent, the action sets of different agents are formed, and the agent network training is carried out. The experience pool of the agent can be dynamically reconstructed to improve the training effect.
2. The multi-microwave source joint heating method with intelligent agent collaborative optimization under composite target consensus decision-making according to claim 1 is characterized by: The Step 1 includes: Step 1.
1. Use the simulation platform to simulate the microwave heating environment and generate a three-dimensional multi-microwave source thermal effect model, including determining the geometric structure of the microwave heating system cavity, the configuration of the microwave source, the definition of the heating area, and the physical properties of the heated material; The geometric structure of the cavity includes a rectangular parallelepiped structure of the cavity, and is set with a width X, a length Y, and a height Z; The configuration of the microwave sources includes the configuration of the quantity, type and position of the microwave sources; The physical properties of the heated material include dielectric constant, magnetic permeability, electrical conductivity, and heat capacity; Step 1.2, collect simulation data, including the intensity of the microwave source, the range of action and the state parameter S of the heated object t ; The three-dimensional multi-microwave source thermal effect model simulates the microwave power distribution and material temperature distribution, sets the power and frequency parameters of each microwave source as optimization variables, and simulates the microwave heating process, including the distribution of temperature field and power transfer.
3. The multi-microwave source joint heating method with intelligent agent collaborative optimization under composite target consensus decision-making according to claim 1 is characterized by: The Step 2 includes: Step 2.1, set the optimization goal; set temperature uniformity as the main goal and heating efficiency as the secondary goal; define the main goal as the optimization core and give it a high priority; and use the secondary goal as the auxiliary optimization goal and give it a suboptimal priority; Setting optimization goals includes: Step 2.1.1, set the initial temperature T0 of the heated material and the target heating time T f , taking the temperature uniformity evaluation index COV as the main target, the calculation formula of the main target is: Among them, T a is the average temperature, T i is the temperature of the temperature sampling point; Step2.1.2, set heating efficiency as the secondary target; The calculation formula for heating efficiency is: is the energy absorbed by the material, E input is the total input energy, E absorbed is the energy absorbed by the heated object; Step 2.2, set constraints, including that the total power of the microwave source cannot exceed the set maximum power; Among them, the constraint condition of the total power of the microwave source during microwave heating is: Among them, P0 is the total power of the microwave source, P i is the power of the ith microwave source, and n is the number of microwave sources.
4. The multi-microwave source joint heating method with intelligent agent collaborative optimization under composite target consensus decision-making according to claim 1 is characterized by: The Step 3 includes: Step 3.1, each microwave source acts as an independent intelligent entity to control its power output; including: Step 3.1.1, set the agent that matches the main goal as the global agent, which is responsible for generating global priority actions; Step 3.1.2, set the agent matched from the target as a local agent to flexibly perform auxiliary tasks, but it must be constrained by the actions of the global agent; Step 3.2, the intelligent agent obtains the current power state of the microwave source and the temperature field data of the heated material in real time through the simulation platform.
5. The multi-microwave source joint heating method with intelligent agent collaborative optimization under composite target consensus decision-making according to claim 1 is characterized by: The Step 4 includes: Step 4.1, Detect target conflict: Based on the resource requirements of the primary and secondary targets, the conflict measure D between the primary and secondary targets is calculated: D=w1COV desire +w2U desire , where w1 and w2 are the master-slave target discrimination coefficients used to determine whether the corresponding target is the master target, COV desire For ideal temperature uniformity, U desire is the ideal heating efficiency; if D>δ, the conflict resolution mechanism is triggered, δ is the conflict threshold, otherwise it is considered that no conflict occurs.
6. The multi-microwave source joint heating method with intelligent agent collaborative optimization under composite target consensus decision-making according to claim 1 is characterized by: The Step 5 includes: Step 5.1, define the state space and action space of the global agent and local agent; Step 5.2, Network structure design: Both global agent and local agent use Actor-Critic as the agent network structure; Step 5.3, Reward function design: Design the reward functions corresponding to the global agent and local agent based on the master-slave objectives.
7. The multi-microwave source joint heating method with intelligent agent collaborative optimization under composite target consensus decision-making according to claim 1 is characterized by: The Step 6 includes: Determine whether a conflict occurs in the current execution state of the agent based on the conflict metric: Non-conflict state D<δ: When it is detected that there is no conflict between the main target and the secondary target, the global agent and the local agent follow the action strategy given by the DDPG algorithm respectively. By dynamically adjusting the power distribution of the microwave source, the priority optimization of the main target is completed to ensure that the main target is in the current state parameter S of the heated object. t The following are given priority, and the secondary optimization needs of the goals are taken into account as much as possible; Conflict state D>δ: Resolve conflicting goals using a winner-take-all strategy, including: Under conflict conditions, the resource allocation between the primary goal and the secondary goal is coordinated through a winner-takes-all strategy; specifically, as follows: In case of conflict, power is allocated according to the priority of the primary target while ensuring the minimum interference of the secondary target; the optimization problem is defined as: Where: R t represents the sum of round rewards, represents the reward of the main goal, represents the reward from the goal, Indicates that the DDPG algorithm takes the main goal as the object of the global agent to perform actions, Indicates that the DDPG algorithm performs actions for local agents from the goal; C sup (A) represents the interference cost of the slave target, a1 and a2 represent the weights of the master and slave targets respectively, and satisfy a1+a2=1; The power distribution scheme is: The global power is allocated to the main target first, and the power demand of the main target is met first. The process of global power allocation to the main target is expressed as: in represents the power allocation of the global agent, that is, the execution action of the global agent under the conflict state, P0 represents the global power constraint, and after the power allocation of the main target, the remaining power is allocated to the slave target; the process of allocating the remaining power to the slave target is expressed as: in It represents the power allocation of local agents, that is, the execution actions of local agents under conflicting conditions.
8. The multi-microwave source joint heating method with intelligent agent collaborative optimization under composite target consensus decision-making according to claim 1 is characterized by: The Step 7 includes: The DDPG algorithm consists of a set of conflicting actions and non-conflicting action sets The target action set As the initial action of the next round, in the DDPG algorithm, the above operation is repeated continuously with R t The goal is to maximize the function value and adjust the execution actions to achieve the optimal heating effect.
9. The multi-microwave source joint heating method with intelligent agent collaborative optimization under composite target consensus decision-making according to claim 8 is characterized by: The Step 7 includes: Build a pool of conflict and non-conflict motivation experiences Used to store the current state parameter S of the heated object t , the sum of round rewards R t And the state parameter S of the heated object after adjustment t+1 ; During the execution of the DDPG algorithm, the conflict experience pool B H and non-conflicting experience pool B′ H The batch samples are drawn in proportion to form the experience pool Batch: Where N1 and N2 are sampling ratios; The Critic network will be updated in phases: in, Denotes the state-action set Q(S,A|θ Q ): Critic network according to the current parameter θ Q The predicted Q value evaluates the quality of the state-action pair; y t is the target Q value, It is the mathematical expectation, which represents the average of all samples in the experience pool Batch; Actor Network Updates: μ(S t |θ μ ): Actor network according to the parameter θ μ The actions generated; Q(S t ,μ(S t )): Critic network evaluates the Q value of the Actor's generated action; target network soft update: i Q′ ←tth Q +(1-τ)θ Q′ i μ′ ←tth μ +(1-τ)θ μ′ θ Q′ ,θ μ′ represents the parameters of the target Critic and Actor networks, and τ represents the soft update coefficient.
Citation Information
Cited By
Maintenance optimization method and system for heating tile in blister equipment
CN121073440A
Factory management and control system based on artificial intelligence and big data
CN121581515A