Task-adaptive stealth system and method based on meta-reinforcement learning
Through the task-adaptive stealth system based on meta-reinforcement learning, the reconfigurable metasurface and meta-reinforcement learning model are used to adjust the metasurface state in real time, which solves the problem of rapid adaptation of transparent stealth in complex electromagnetic environments in existing technologies and achieves a robust long-range stealth effect.
Patent Information
- Application Number
- CN202511099646.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Existing technologies make it difficult to achieve rapid adaptive transparent stealth in the face of objects of various shapes and dynamic changes and in complex electromagnetic environments. In particular, when the position, angle or speed of the object changes, deep learning models find it difficult to maintain a robust stealth effect, and the cost of sample collection is high.
A task-adaptive stealth system based on meta-reinforcement learning is adopted. Through the combination of reconfigurable metasurface, detector, controller and digital power supply, the meta-reinforcement learning model is used to adjust the metasurface state in real time, build a transparent stealth tunnel, and achieve precise control of electromagnetic wave scattering.
It achieves rapid adaptive long-range stealth in diverse and complex electromagnetic environments, can adapt to new tasks with very little new data, and improves the robustness and adaptability of the stealth effect.
Smart Images

Figure CN120601158B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electromagnetic metasurface design and intelligent stealth technology, and more particularly to a task-adaptive stealth system and method based on meta-reinforcement learning. Background Art
[0002] Invisibility—the superpower of making an object instantly invisible—has long been a dream pursued by mankind. Traditional invisibility cloaks achieve invisibility by wrapping objects, and their invisible area is defined by the space inside the cloak, so they usually take the form of a closed structure. In contrast, open invisibility cloaks aim to create an invisible area that does not rely on physical boundaries and allows objects to move freely. We call this form of invisibility "transparent invisibility." In theory, it can construct a magical transparent invisible tunnel (TCO). Imagine if a car or a train were placed in such a tunnel, it would be able to move freely while remaining invisible and unnoticed by the outside world. This is undoubtedly an exciting innovation. However, when faced with objects of various shapes and dynamic changes, as well as complex electromagnetic environments, creating such an intelligent TCO remains an extremely challenging task. Specifically, this task faces the following problems:
[0003] (1) Complexity: TCO needs to precisely adjust the phase response of each unit for objects of different shapes and positions to achieve effective control of electromagnetic wave scattering.
[0004] (2) Dynamicity: TCO must adjust its state in real time to cope with changes in the scattering field caused by object movement, deformation, and changes in the external environment.
[0005] (3) Diversity: Models that are effective in a specific environment are often difficult to quickly migrate to new environments.
[0006] Since Pendry et al. pioneered the theory of transformation optics in 2006, Smith et al. have achieved the first effective microwave-band stealth device by meticulously designing the constitutive parameters of the transforming medium, significantly stimulating the continued and extensive exploration of electromagnetic stealth technology. For stealth applications in diverse and complex electromagnetic environments, scientists have proposed a deep learning-based tunable metasurface strategy. This strategy leverages the mapping relationship between the electromagnetic properties of the environment and the target's stealth state to rapidly tune the metasurface's scattering properties. However, widespread deployment in complex, real-world scenarios still faces the bottleneck of high sample collection costs. In particular, deep learning often struggles to maintain robust stealth when stealth targets (such as vehicles) undergo changes in position, angle, or velocity. Deep reinforcement learning (DRL), which relies on continuous interaction with the environment, has achieved numerous breakthroughs in addressing continuous decision-making scenarios, and researchers have also successfully applied it to the stealth control of dynamic moving objects. However, reinforcement learning typically relies on large-scale environmental interaction data for training. Significant changes in external conditions can easily lead to a sharp drop in performance of previously learned models.
[0007] Therefore, faced with a series of stealth task distributions that are both interrelated and different from each other, learning a general strategy that can quickly adapt to any new task in the distribution with very little new data to solve the difficulties of existing technologies is an urgent problem that technicians in this field need to solve. Summary of the Invention
[0008] In view of this, the present invention provides a task-adaptive stealth system and method based on meta-reinforcement learning. By comprehensively considering the characteristics of different tasks and constructing a meta-reinforcement learning model, rapid adaptive long-range stealth can be achieved in diverse and complex electromagnetic environments.
[0009] In order to achieve the above object, the present invention adopts the following technical solutions:
[0010] In one aspect, the present invention provides a task-adaptive stealth system based on meta-reinforcement learning, comprising: a reconfigurable metasurface, a detector, a controller, and a digital power supply;
[0011] The detector is used to capture the environmental status outside the stealth channel area in real time;
[0012] The controller runs a meta-reinforcement learning stealth algorithm to generate a metasurface voltage decision based on the environment state and the task trajectory;
[0013] The digital power supply is a voltage control module that dynamically adjusts the voltage value of the reconfigurable metasurface according to the voltage decision output by the controller;
[0014] The reconfigurable metasurface is composed of an array of phase-adjustable units, each of which contains a varactor diode. By applying different levels of DC bias through the digital power supply, the propagation characteristics of incident electromagnetic waves are adjusted, thereby changing the fused scattering field of the reconfigurable metasurface and the object, and realizing electromagnetic invisibility.
[0015] Preferably, the phase-adjustable units of each column of the reconfigurable metasurface share the same bias voltage; the diodes of the phase-adjustable units are welded on the metal layer of the dielectric substrate.
[0016] Preferably, the controller is configured to:
[0017] The super network receives the meta features of a set of tasks m , and encodes the meta features of the tasks through a task meta feature encoder to obtain embedded encodings of the task features e T ;
[0018] The task trajectory encoder encodes the trajectory of the interaction of the agent with the environment , wherein the trajectory , represents the state of the environment , at time t t s t , a t the agent performs an action r t , and the agent obtains a reward
[0019] The embedded encodings of the task features e T and the trajectory encodings are input into a policy generation network through a fully connected layer to generate policy network parameters ;
[0020] The policy network outputs a current decision based on the current state of the environment s t a t .
[0021] Preferably, the meta-reinforcement learning invisibility algorithm maps data from a single task to a corresponding policy network ; the invisibility policy is represented as , and the policy network contains two parts of parameters, wherein represents the fixed parameters of the policy network, representing a super network according to tasks generated variable parameters.
[0022] Preferably, the goal of the meta-reinforcement learning is to optimize the meta-parameters to maximize the meta-episode cumulative reward over a set of task distributions, which is represented as:
[0023] ;
[0024] wherein, as an inner loop, for each trajectory generate specific reinforcement learning parameters , and in turn obtain a policy ; the outer loop is used to optimize the meta-parameters .
[0025] In another aspect, the present application provides a task-adaptive stealth method based on meta-reinforcement learning, comprising the following steps:
[0026] S1: the detector observes the electromagnetic environment of the current task, obtains the meta-features of the task m , and the state of the environment s t ;
[0027] S2: the super network receives a set of meta-features of the task m , and encodes them through a task meta-feature encoder to obtain the embedding encoding of the task features e T ;
[0028] S3: a task trajectory encoder performs trajectory encoding on the trajectory of the interaction between the agent and the environment , wherein the trajectory , , represents the corresponding environment state t at time t s t , a t the agent performs an action r t , and the agent obtains a reward
[0029] S4: the embedding encoding of the task features e T and the trajectory encoding are input to a policy generation network through a fully connected layer to generate policy network parameters ;
[0030] S5: The policy network is based on the current environment state s t , output the execution action of the current agent a t ;
[0031] S6: The digital power supply is based on the execution action of the current agent a t Change the output voltage of each column of phase-adjustable units, thereby changing the reconfigurable metasurface state, and achieving active regulation of electromagnetic scattering characteristics;
[0032] S7: Observe the environment and obtain the meta-features of the task m , and the state of the environment s t , based on the environment state s t Calculate the return r t ; if the return r t is less than the threshold value of invisibility r T , stop the loop; otherwise, add to the trajectory , and return to S2 for continuous execution.
[0033] Via the technical solution described above, compared with the prior art, the present disclosure provides a task-adaptive invisibility method system and method based on meta-reinforcement learning, which realizes remote regulation of the metasurface on the scatterer by continuously adjusting the metasurface state change using the meta-reinforcement learning model, and ultimately constructs a transparent invisibility channel, by real-time sensing of external electromagnetic signals. BRIEF DESCRIPTION OF DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present application, and those skilled in the art can obtain other drawings according to the provided drawings without creative labor.
[0035] Figure 1 A system structure diagram of task-adaptive invisibility based on meta-reinforcement learning is provided in the present application.
[0036] Figure 2(a) is a physical diagram of the metasurface designed in the present application, and Figure 2(b) is a schematic diagram of the metasurface unit structure.
[0037] Figure 3 An algorithm framework of meta-reinforcement learning invisibility is provided in the present application.
[0038] Figure 4 The present invention provides a process for implementing adaptive long-range stealth under different tasks. DETAILED DESCRIPTION
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0040] The embodiment of the present invention discloses a task-adaptive stealth system based on meta-reinforcement learning, such as Figure 1 As shown, it includes: reconfigurable metasurface, detector, controller, digital power supply;
[0041] The detector is used to capture the environmental status outside the stealth channel area in real time;
[0042] The controller runs a meta-reinforcement learning stealth algorithm to generate a metasurface voltage decision based on the environment state and the task trajectory;
[0043] The digital power supply is a voltage control module that dynamically adjusts the voltage value of the reconfigurable metasurface according to the voltage decision output by the controller;
[0044] The reconfigurable metasurface is composed of an array of phase-adjustable units. Each phase-adjustable unit contains a varactor diode. Different levels of DC bias are applied through a digital power supply to adjust the propagation characteristics of the incident electromagnetic wave, thereby changing the fused scattering field of the reconfigurable metasurface and the object to achieve electromagnetic stealth.
[0045] Furthermore, the phase-adjustable units in each column of the reconfigurable metasurface share the same bias voltage; the diodes of the phase-adjustable units are soldered to the metal layer of the dielectric substrate. The metasurface structure designed by the present invention is shown in Figure 2(a), and the phase-adjustable unit structure that constitutes the reconfigurable metasurface is shown in Figure 2(b).
[0046] In another embodiment, the controller is configured to:
[0047] The hypernetwork receives a set of meta-features of tasks m , and through the task meta-feature encoder Encode and obtain the embedded encoding of task features e T ;
[0048] Mission trajectory encoder Trajectory of the interaction between the agent and the environment Trajectory encoding , the trajectory , represents the time t , s t , a t performs an action for the agent, r t the reward obtained by the agent;
[0049] encodes the embedding of the task feature e T with the trajectory encoding input to the policy generation network through a fully connected layer to generate policy network parameters ;
[0050] The policy network outputs the current decision based on the current environment state s t a t .
[0051] Further, the meta-reinforcement learning cloaking algorithm maps data from a single task to the corresponding policy network ; the cloaking policy is represented as , and the policy network contains two parts of parameters, where represents the fixed parameters of the policy network, represents the hypernetwork generates variable parameters according to the task .
[0052] Further, the goal of meta-reinforcement learning is to optimize the meta-parameters to maximize the meta-episode cumulative reward over a series of task distributions, which is represented as:
[0053] ;
[0054] where, as an inner loop, generates specific reinforcement learning parameters for each trajectory , and in turn obtains the policy ; the outer loop is used to optimize the meta-parameters .
[0055] The cloaking behavior of an object in a specific state is considered as an independent task , and the process of realizing cloaking in the cloaking tunnel for the object is formalized as a Markov Decision Process (MDP) where S, A, R, and P represent state, action, reward, and transition probability, respectively is the discount factor. In each reinforcement learning task, the task label is omitted i .
[0056] For the cloaking task , a deep reinforcement learning is chosen to solve the corresponding Markov Decision Process (MDP) At each time step t , the electromagnetic environment is in state s t ∈ S, the agent outputs a metasurface voltage control decision ∈ A using a policy network a t The digital power source then executes the decision. The electromagnetic environment state s t is then transitioned to the next state according to the transition probability s t+1 , i.e., the fused scattering field undergoes a state transition. At this time, the agent obtains a reward defined by the reward function r t To maximize the expected discounted reward in the future, the agent samples actions according to the policy , whose objective can be expressed as:
[0057] ;
[0058] where represents a state-action-reward sequence generated by the agent in the MDP according to the policy , and the discount factor .
[0059] To achieve perfect cloaking on a wider set of tasks, a meta-reinforcement learning cloaking algorithm is introduced, which maps the data sampled from a single MDP to the policy parameters. Considering that reinforcement learning usually needs to go through multiple rounds of interaction to obtain a reasonable policy, the MRC adopts the complete trajectory as the conditional information, rather than single time step data. Here, represents a state-action-reward sequence spanning multiple episodes collected from the MDP , called a meta-episode. The cloaking policy is represented by the parameters , while the function f itself is uniquely characterized by the meta-parameters . Therefore, the goal of the MRC is to optimize the meta-parameters To maximize the total reward of the meta-round in the task distribution (MDPs), that is:
[0060] ;
[0061] in It is called the inner loop, which generates the MDP Corresponding reinforcement learning parameters , while the outer loop optimizes the meta-parameters .
[0062] On the other hand, the present invention provides a task-adaptive stealth method based on meta-reinforcement learning, such as Figure 3 As shown, the following steps are included:
[0063] S1: The detector observes the electromagnetic environment of the current task and obtains the meta-features of the task m , and the state of the environment s t ;
[0064] S2: The hypernetwork receives a set of meta-features for tasks m , and through the task meta-feature encoder Encode and obtain the embedded encoding of task features e T ;
[0065] S3: Mission trajectory encoder Trajectory of the interaction between the agent and the environment Trajectory encoding , the trajectory , Indicates the time t Corresponding environmental status s t , a t Execute actions for the agent, r t The reward obtained by the agent;
[0066] S4: Embedding the task features e T With trajectory encoding Input into the strategy generation network through the fully connected layer to generate the strategy network parameters ;
[0067] S5: Policy network based on the current environment state s t , output the current agent's execution action a t ;
[0068] S6: Digital power based on the execution action of the current agent a t Changing the output voltage of each column of phase-adjustable units changes the state of the reconfigurable metasurface, enabling active regulation of electromagnetic scattering properties.
[0069] S7: Observe the environment and obtain meta-features of the task m , and the state of the environment s t , based on the environmental state s t Calculating Returns r t If the return r t Less than the invisibility threshold r T , then stop the loop; otherwise, Add to track , return to S2 to continue execution.
[0070] HyperNetwork Utilizes Task Meta-Feature Encoders and mission trajectory encoder To encode task characteristics.
[0071] in, It is a fully connected network that encodes the normalized meta-features; It is a recurrent neural network that encodes the task trajectory.
[0072] The parameter set of the policy network is , the super network finally generates the parameter subset of the last two fully connected layers of the policy network , represents the fixed shared basis parameters of the policy network.
[0073] Figure 4 This example illustrates the process of achieving adaptive long-range cloaking in different missions provided by an embodiment of the present invention. In Mission 1, the vehicle, acting as the cloaked object, lies in close proximity to the left side of the metasurface. In Mission 2, the vehicle, acting as the cloaked object, is 100 mm away from the metasurface and close to the right side. Through algorithmic control and a series of iterative processes, the scattered fields surrounding the vehicle are reconstructed for each mission, achieving adaptive long-range cloaking.
[0074] Through the above embodiments, the present invention realizes the design of adaptive long-range stealth devices and provides a feasible solution for the practical application of task-adaptive stealth systems based on meta-reinforcement learning.
[0075] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0076] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A task-adaptive stealth method based on meta-reinforcement learning, characterized in that: The following steps are involved: S1: The detector observes the electromagnetic environment of the current task and obtains the meta-features of the task m , and the state of the environment s t ; S2: The hypernetwork receives a set of meta-features for tasks m , and through the task meta-feature encoder Encode and obtain the embedded encoding of task features e T ; S3: Mission trajectory encoder Trajectory of the interaction between the agent and the environment Trajectory encoding , the trajectory , Indicates the time t Corresponding environmental status s t , a t Execute actions for the agent, r t The reward obtained by the agent; S4: Embedding the task features e T With trajectory encoding Input into the policy generation network through the fully connected layer to generate policy network parameters; The policy generation network is stealth-based via a meta-reinforcement learning algorithm. , will be added from a single task Data Mapping to the corresponding policy network ; The stealth strategy is expressed as , policy network Contains two parts of parameters, represents the fixed parameters of the policy network, Represents a hypernetwork According to the task Generated variable parameters; S5: Policy network based on the current environment state s t , output the current agent's execution action a t ; S6: Digital power based on the execution action of the current agent a t Changing the output voltage of each column of phase-adjustable units changes the state of the reconfigurable metasurface, enabling active regulation of electromagnetic scattering properties. S7: Observe the environment and obtain meta-features of the task m , and the state of the environment s t , based on the environmental state s t Calculating Returns r t If the return r t Less than the invisibility threshold r T , then stop the loop; otherwise, Add to track Return to S2 to continue execution.
2. The task-adaptive stealth method based on meta-reinforcement learning according to claim 1, characterized in that: The goal of meta-reinforcement learning is to optimize the meta-parameters , to maximize the meta-round cumulative reward over a series of task distributions, the objective is expressed as: ; in, As an inner loop, for each trajectory Generate specific variable parameters , and then obtain the strategy ; The outer loop is used to optimize the meta parameters .
3. A task-adaptive stealth system based on meta-reinforcement learning, characterized in that: The system adopts a task-adaptive stealth method based on meta-reinforcement learning as described in any one of claims 1-2, comprising: a reconfigurable metasurface, a detector, a controller, and a digital power supply; The detector is used to capture the environmental status outside the stealth channel area in real time; The controller runs a meta-reinforcement learning stealth algorithm to generate a metasurface voltage decision based on the environment state and the task trajectory; The digital power supply is a voltage control module that dynamically adjusts the voltage value of the reconfigurable metasurface according to the voltage decision output by the controller; The reconfigurable metasurface is composed of an array of phase-adjustable units, each of which contains a varactor diode. Different levels of DC bias are applied by the digital power supply to adjust the propagation characteristics of the incident electromagnetic wave, thereby changing the fused scattering field of the reconfigurable metasurface and the object to achieve electromagnetic stealth.
4. The task-adaptive stealth system based on meta-reinforcement learning according to claim 3, characterized in that: The phase-adjustable units in each column of the reconfigurable metasurface share the same bias voltage; the diodes of the phase-adjustable units are welded on the metal layer of the dielectric substrate.
Citation Information
Patent Citations
Self-adapting super-surface electromagnetic invisibility cloak system and working method thereof
CN109489485A
Multi-region collaborative control method and device for high-precision pressure sensing array
CN119781417A