Intelligent sea floating type corner reflector centroid interference dynamic decision-making method

By enhancing the features and making dynamic decisions on the state information of the centroid interference of the sea-drifting corner reflector, and using the Actor network to output policy actions, the problem of centroid interference that cannot be dynamically adjusted in the existing technology is solved, thereby improving the threat avoidance capability of ships.

CN122043370APending Publication Date: 2026-05-15XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIDIAN UNIV
Filing Date
2026-02-12
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies cannot achieve dynamic changes in single-ship centroid interference methods in sea surface scenarios, and cannot provide an optimal joint strategy for corner reflector deployment and ship maneuvering, resulting in insufficient interference decision-making effectiveness.

Method used

By acquiring the state information of the target shipborne radar, performing state space mapping and feature enhancement, and using a pre-trained target Actor network to output jamming strategy actions, including corner reflector deployment positions and ship maneuvering directions, a jamming model is constructed for step-by-step decision-making.

Benefits of technology

It enables dynamic changes in jamming strategies, enhances the ship's ability to evade threats, and improves the success rate and effectiveness of center-of-gravity jamming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122043370A_ABST
    Figure CN122043370A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent sea floating type corner reflector centroid interference dynamic decision-making method. The method comprises the following steps: acquiring state information of a current moment detected by a target ship-borne radar; the state information is used for representing the states of the target ship, the seeker radar and the corner reflector; performing state space mapping and feature enhancement on the state information at the current moment to obtain enhanced state representation at the current moment; inputting the enhancement state representation at the current moment into a pre-trained target Actor network, and outputting a target strategy action at the current moment; the target strategy action comprises whether a corner reflector is deployed or not, the position of the corner reflector is deployed, and the maneuvering direction and the maneuvering speed of the target ship. According to the method, the deployment of the sea floating type corner reflector and the motion space of ship maneuvering are constructed, the mass center interference model is stepped through the step length, the interference strategy motion is output by observing the state information under different time steps, the dynamic change of the interference strategy is realized, and the threat avoidance capability of the ship is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent passive interference decision-making technology, specifically relating to an intelligent dynamic decision-making method for centroid interference of a sea-drifting corner reflector. Background Technology

[0002] Corner reflector centroid jamming is an important means for ships to avoid surface threats. Drifting corner reflector centroid jamming refers to deploying decoys within the radar's tracking range after the radar seeker has already tracked the ship. This forces the seeker radar to track the energy center of both the real and decoy targets, disrupting its stable tracking of the ship. Under certain conditions, this forces the radar to shift its tracking from the centroid to the decoy. This type of jamming is called centroid jamming. In addition to deploying corner reflectors, the ship's correct maneuvering is another crucial condition for the success of centroid jamming. After the corner reflectors are deployed, during the process of creating the centroid effect, the ship should choose an appropriate course and speed to evade the seeker radar's tracking as early as possible. The specific process is detailed in the attached diagram. Figure 1 As shown. In summary, the deployment method of corner reflectors and the ship's maneuvering methods together constitute a center-of-gravity jamming countermeasure strategy.

[0003] After the seeker radar tracks a ship, the ship should make timely and accurate center-of-gravity (CG) jamming decisions to avoid the threat. Because the time allotted for decision-making during CG jamming is very short, and the observation information is complex and constantly changing, making dynamic jamming decisions based on observation information within a very short time is extremely difficult. Therefore, a fast and effective CG jamming decision generation technology is urgently needed for intelligent dynamic CG jamming decision-making.

[0004] Existing methods for centroid interference of a single ship in maritime scenarios typically employ static actions performed at the start of centroid interference. These methods cannot dynamically adapt to different observation information during real-time interference. Furthermore, when analyzing the effectiveness of interference strategies, they can only provide the optimal ship maneuvering direction or the corner reflector deployment direction separately, and cannot provide the optimal combined strategy of corner reflector deployment and ship maneuvering. Summary of the Invention

[0005] To address the aforementioned problems in the existing technology, this invention provides an intelligent dynamic decision-making method for centroid interference of a sea-drifting corner reflector.

[0006] The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides an intelligent dynamic decision-making method for centroid interference of a sea-drifting corner reflector, comprising: Acquire the current state information detected by the target ship's onboard radar; the state information is used to characterize the state of the target ship, seeker radar, and corner reflector. The state information at the current moment is mapped to the state space and its features are enhanced to obtain the enhanced state representation at the current moment; The augmented state representation at the current moment is input into the pre-trained target Actor network, which outputs the target policy action at the current moment. The target policy action includes whether to deploy the corner reflector, the location of the corner reflector deployment, the target ship's maneuvering direction, and the maneuvering speed.

[0007] This invention provides an intelligent dynamic decision-making method for centroid interference of sea-drifting corner reflectors. By constructing the deployment space of sea-drifting corner reflectors and the maneuvering space of ships, the centroid interference model is stepped by step size. By observing the state information at different time steps, the interference strategy action is output, realizing the dynamic change of the interference strategy and improving the ship's ability to avoid threats.

[0008] The present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description

[0009] Figure 1 This is a schematic diagram of a scenario for an intelligent dynamic decision-making method for centroid interference of a drifting corner reflector provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating an intelligent dynamic decision-making method for centroid interference of a drifting corner reflector provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of a process for a dynamic decision-making method for centroid interference of a sea-drifting corner reflector provided in an embodiment of the present invention; Figures 4A to 4D This is a schematic diagram comparing the simulation results of an intelligent dynamic decision-making method for centroid interference of a drifting corner reflector provided in this embodiment of the invention with existing methods. Detailed Implementation

[0010] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0011] This invention provides an intelligent dynamic decision-making method for centroid interference of a drifting corner reflector. (See also...) Figure 2 The method includes the following steps: S10. Obtain the current status information detected by the target shipborne radar.

[0012] Among them, the status information is used to characterize the status of the target ship, the seeker radar, and the corner reflector.

[0013] For example, the status information includes threat information, corner reflector information, and target ship information. Threat information includes relative position, velocity direction, velocity magnitude, and seeker radar tracking time; corner reflector information includes the remaining number of corner reflectors, relative position of the corner reflectors, and corner reflector RCS value; target ship information includes target ship velocity magnitude and target ship velocity direction.

[0014] S20. Perform state-space mapping and feature enhancement on the current state information to obtain the enhanced state representation of the current time.

[0015] Optionally, step S20 may specifically include: S201. Map the current state information to a state space and divide the state space into multiple feature sub-blocks.

[0016] Among them, multiple feature sub-blocks include threat information feature sub-blocks, corner reflector information feature sub-blocks, and target ship information feature sub-blocks.

[0017] For example, state information is mapped to a state space, and then the state space is precisely divided according to semantic boundaries, resulting in three feature sub-blocks with clear physical meanings, as shown in the table below:

[0018] S202. Multiple feature sub-blocks are fused using a gated preprocessing network to obtain the enhanced state representation at the current time.

[0019] Optionally, step S202 may specifically include: A two-layer fully connected layer and a ReLU activation function are used to perform nonlinear feature extraction on multiple feature sub-blocks to obtain initial output features. The Sigmoid function is used to map the initial output features to the (0,1) interval to obtain the gating weight parameters corresponding to each feature sub-block. Multiple feature sub-blocks are processed through a normalization layer to obtain the normalized features corresponding to each feature sub-block. Based on the gating weight parameters corresponding to each feature sub-block, multiple feature sub-blocks and their corresponding normalized features are fused to obtain multiple fused features. The multiple fused features are concatenated to obtain the enhanced state representation at the current time step.

[0020] For example, since the data dimensions of multiple feature sub-blocks are large, in order to enable the network to dynamically focus on the most important information at present, this embodiment employs a gated preprocessing network to help the policy network preprocess multiple feature sub-blocks. The gated preprocessing network includes a two-layer fully connected layer, a ReLU activation function, a Sigmoid function, and a layer normalization layer.

[0021] Initial output features , represented as:

[0022] in, Indicates a two-layer fully connected layer. This represents all feature sub-blocks.

[0023] The gating weight parameters corresponding to each feature sub-block are expressed as follows:

[0024] in, This represents the gating weight parameters corresponding to the threat information feature sub-block. This represents the gating weight parameters corresponding to the corner reflector information feature sub-block. This represents the gating weight parameters corresponding to the target ship information feature sub-block. This represents the Sigmoid function.

[0025] To extract richer information from the features, each feature sub-block is processed through a normalization layer to obtain normalized features, represented as:

[0026] in, , and These represent the normalized features corresponding to the threat information feature block, the corner reflector information feature block, and the target ship information feature block, respectively. , and These represent the threat information feature sub-block, the corner reflector information feature sub-block, and the target ship information feature sub-block, respectively.

[0027] Then, using the learned gating weight vector, the original features and normalized features are complementaryly fused to obtain multiple fused features, represented as follows:

[0028] in, , and These represent the fused features corresponding to the threat information feature sub-block, the corner reflector information feature sub-block, and the target ship information feature sub-block, respectively.

[0029] Finally, the three fusion features are concatenated and integrated along the feature dimension to form the preprocessed enhanced state representation. , represented as:

[0030] S30. Input the enhanced state representation of the current time step into the pre-trained target Actor network and output the target policy action of the current time step.

[0031] The target strategy actions include whether to deploy corner reflectors, the location of the corner reflectors, the direction of maneuver of the target ship, and the speed of maneuver.

[0032] For example, the target actor network includes three policy action output heads, of which two discrete action output heads output whether to deploy a corner reflector and the location where the corner reflector is deployed, respectively, and a two-dimensional continuous action output head is used to output the maneuvering direction and speed of the target ship.

[0033] Optionally, the training process of the target Actor network includes: calculating the advantage function value based on pre-collected state information samples from multiple consecutive time steps, augmented state representation samples, policy action predictions output by the Actor network, reward, and state value predictions output by the Critic network; and updating the parameters of the Actor network and the Critic network based on the advantage function value, state value predictions, preset Actor network loss functions, and preset Critic network loss functions.

[0034] For example, the Actor network is a policy network, and the Critic network is a value network.

[0035] Optionally, the dominance function value is expressed as:

[0036] in, This represents the dominance function value at time step t. Indicates the state Take action below Expected returns This represents the state value prediction at time step t. This represents a sample of state information at time step t. This represents the prediction of the policy action at time step t. This represents the TD error (Temporal Difference Error) at time step t. Indicates the discount factor. Indicates GAE parameters, Indicates the end time. Indicates the time offset.

[0037] Optionally, the Actor network loss function is preset, expressed as:

[0038]

[0039] in, This indicates the preset loss function for the Actor network. This represents the loss function of the original Actor network. Represents the entropy regularization term. Expressing expectations, Indicates the policy update ratio, The parameters of the Actor network are represented. This represents the truncation function. This indicates the trimming parameters.

[0040] Optionally, the Critic network loss function is preset, expressed as:

[0041]

[0042] in, This indicates the preset loss function of the Critic network. Indicates state value prediction. Indicates the state value objective. This represents the estimate from the old Critic network. This represents the parameters of the Critic network.

[0043] Optionally, the process of pre-collecting state information samples from multiple consecutive time steps, augmented state representation samples, policy action predictions from the Actor network output, and state value predictions from the Critic network output includes: A1. Obtain a sample of the current time step's state information.

[0044] A2. Perform state space mapping and feature enhancement on the state information sample of the current time step to obtain the enhanced state representation sample of the current time step.

[0045] A3. Predict the augmented state representation sample of the current time step using the Actor network, and output the policy action prediction for the current time step; and determine the reward for the current time step based on the policy action prediction for the current time step and the preset segmented reward function.

[0046] Optionally, a pre-defined segmented reward function is provided, expressed as:

[0047] in, This represents the local reward function for each time step before the final time step. This represents the weighting coefficient of the local reward. The yaw angle represents the direction of tracking by the seeker radar. This represents the global reward function for the last time step. This represents the reward received when the centroid interference fails. This represents the weighting coefficient of the global reward. This indicates the distance between the seeker radar and the ship at the last time step. Indicates the number of missiles.

[0048] A4. Predict the state information sample at the current time step using the Critic network, and output the state value prediction for the current time step.

[0049] A5. Execute the policy action prediction for the current time step and obtain the state information sample for the next time step.

[0050] For example, the training process of the target Actor network is as follows:

[0051] This embodiment provides an intelligent dynamic decision-making method for centroid interference of sea-drifting corner reflectors. By constructing the deployment space of sea-drifting corner reflectors and the maneuvering space of ships, the centroid interference model is stepped by step size. By observing the state information at different time steps, the interference strategy action is output, realizing the dynamic change of the interference strategy and improving the ship's ability to avoid threats.

[0052] The following simulation experiment further illustrates the intelligent dynamic decision-making method for centroid interference of a drifting corner reflector provided by this invention.

[0053] (1) Simulation environment modeling To realize the interaction process between the two parties during the centroid interference, this embodiment uses the RCS (Radar Cross Section) data of a certain type of ship to construct an interaction model of centroid interference, in order to verify the effectiveness of the method of the present invention in a specific scenario.

[0054] Reference Figure 3 To address the issue of decision-making regarding ship center of gravity interference, it is necessary to comprehensively consider radar guidance models, ship maneuvering models, sea-drifting corner reflector motion models, and center of gravity interference models.

[0055] Center-of-gravity jamming, also known as sea-drift corner reflector jamming, refers to the use of decoys deployed around a ship within the radar's tracking range to exploit the radar's ability to track the target's energy center after the ship has been tracked by a seeker radar. The movement of the decoys and the ship causes the radar to track the energy centers of both the real and decoys, disrupting the radar's stable tracking of the target ship. Under certain conditions, the radar is eventually forced to switch from tracking the energy center to tracking the decoys.

[0056] The goal of center-of-mass interference decision-making is to rapidly and accurately adjust the ship's position and release the interference at the appropriate time and in the appropriate manner to ensure that the ship has a safe distance when the center-of-mass interference ends. Due to physical constraints during the interaction process, this problem can be modeled as follows:

[0057] in, The yaw angle of the seeker radar relative to the ship during tracking. This refers to the distance from the seeker radar to the ship at the end of the countermeasures. This represents the maximum number of corner reflectors that can be used. This is the ship's maximum speed. This refers to the range of the ship's maneuvering directions. This refers to the ship's maximum acceleration. This is the ship's maximum angular velocity of turning. This represents the maximum turning angular velocity of the seeker radar.

[0058] (II) Simulation conditions: This embodiment uses Python and is based on a simulation environment model. The number of ships is set to 1, and the number of seeker radars is set to 1-4. The wind direction is set to 90°, and the wind speed is set to 15 m / s. At the start of the simulation, the seeker radar begins to track the ship, and the ship enters an evasive state. In each simulation step, the ship can change its maneuvering actions and deploy corner reflectors as needed.

[0059] Using the simulation environment described above, and training the centroid interference strategy using the improved PPO algorithm, the network structure was built using the PyTorch deep learning framework. The hyperparameter settings for the experiment are shown in the table.

[0060]

[0061] The simulation software environment consists of an Intel(R) Core(TM) i9-12900H CPU @2500MHz, an NVIDIA GeForce RTX4090 GPU, and Python 3.11 running Windows 11 Ultimate 64-bit operating system.

[0062] (III) Simulation Content and Result Analysis To verify the effectiveness of the algorithm of this invention, relevant experiments were designed for illustration.

[0063] First, the algorithm of this invention and the traditional PPO algorithm are compared at different initial tracking distances of the seeker radar to test the effectiveness of centroid interference, demonstrating the performance advantages of the proposed algorithm.

[0064] Since the seeker radar may track the ship from different directions during actual interaction, a second experiment was designed to analyze the success probability of the algorithm of this invention in centroid interference under different initial tracking directions of the seeker radar.

[0065] To introduce more complex combat scenarios, the third experiment designed a scenario in which the ship was tracked by multiple seeker radars, and analyzed the advantages of the algorithm proposed in this invention.

[0066] Experiment 1: Existing solutions cannot provide a joint strategy for ship maneuvering and angle-counter deployment for specific scenarios. Therefore, this experiment compares the performance of the algorithm of this invention with the traditional PPO algorithm. (Refer to...) Figure 4A The results show a comparison between the algorithm of this invention and the traditional PPO algorithm. Compared with the traditional PPO, the method of this invention has a significant performance improvement. Figure 4A Simulations were conducted under different initial tracking distances. The blue line represents the algorithm of this invention. It can be observed that the method of this invention has a higher success rate under different initial tracking distances, which also indicates that the algorithm of this invention has better generalization ability.

[0067] Experiment 2: Figure 4B and Figure 4C After fixing the initial tracking distance of the seeker radar, centroid interference was applied under different tracking directions of the seeker radar. The success probability and actual deflection distance were analyzed for each initial tracking direction. It can be found that the algorithm of this invention has a high success probability for most initial tracking directions of the seeker radar, thus demonstrating strong application capabilities.

[0068] Experiment 3: Figure 4D This test examines the performance of the algorithm in complex scenarios. For a scenario where multiple seeker radars are tracking a ship, the simulation allows the ship to deploy two corner reflectors for interference, as shown in the figure. The algorithm of this invention demonstrates a higher success rate compared to the traditional PPO algorithm.

[0069] It should be noted that the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention.

[0070] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0071] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings and the disclosure in carrying out the claimed invention. In the description of the invention, the word "comprising" does not exclude other components or steps, "a" or "an" does not exclude a plurality, and "a plurality" means two or more, unless otherwise explicitly specified. Furthermore, while different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce good results.

[0072] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. An intelligent dynamic decision-making method for centroid interference of a sea-drifting corner reflector, characterized in that, include: Obtain the current status information detected by the target ship's onboard radar; The status information is used to characterize the status of the target ship, the seeker radar, and the corner reflector; The state information at the current moment is subjected to state space mapping and feature enhancement to obtain the enhanced state representation at the current moment; The enhanced state representation at the current moment is input into the pre-trained target Actor network, and the target policy action at the current moment is output; the target policy action includes whether to deploy corner reflectors, the location of the deployed corner reflectors, the maneuvering direction of the target ship, and the maneuvering speed.

2. The intelligent dynamic decision-making method for centroid interference of a drifting corner reflector according to claim 1, characterized in that, The step of performing state-space mapping and feature enhancement on the state information at the current moment to obtain the enhanced state representation at the current moment includes: The current state information is mapped to a state space, and the state space is divided into multiple feature sub-blocks, including threat information feature sub-blocks, corner reflector information feature sub-blocks, and target ship information feature sub-blocks. The multiple feature sub-blocks are fused using a gated preprocessing network to obtain the enhanced state representation at the current time.

3. The intelligent dynamic decision-making method for centroid interference of a drifting corner reflector according to claim 2, characterized in that, The process of fusing the multiple feature sub-blocks through a gated preprocessing network to obtain the enhanced state representation at the current time includes: A two-layer fully connected layer and a ReLU activation function are used to perform nonlinear feature extraction on the multiple feature sub-blocks to obtain the initial output features; The initial output features are mapped to the (0,1) interval using the Sigmoid function to obtain the gating weight parameters corresponding to each feature sub-block; The multiple feature sub-blocks are processed by a layer normalization layer to obtain the normalized features corresponding to each feature sub-block; Based on the gating weight parameters corresponding to each feature sub-block, the multiple feature sub-blocks and the normalized features corresponding to each feature sub-block are fused to obtain multiple fused features; The multiple fusion features are concatenated to obtain the enhanced state representation at the current moment.

4. The intelligent dynamic decision-making method for centroid interference of a drifting corner reflector according to claim 1, characterized in that, The training process of the target Actor network includes: The advantage function value is calculated based on the state information samples collected in multiple consecutive time steps, the enhanced state representation samples, the policy action predictions output by the Actor network, and the state value predictions output by the reward and Critic networks. Based on the advantage function value, state value prediction, preset Actor network loss function, and preset Critic network loss function, the parameters of the Actor network and the Critic network are updated.

5. The intelligent dynamic decision-making method for centroid interference of a drifting corner reflector according to claim 4, characterized in that, The process of pre-collecting state information samples from multiple consecutive time steps, augmented state representation samples, policy action predictions from the Actor network output, and state value predictions from the Critic network output includes: Obtain a sample of the current time step's state information; The state information sample at the current time step is subjected to state space mapping and feature enhancement to obtain the enhanced state representation sample at the current time step; The enhanced state representation sample at the current time step is predicted using an Actor network, and the policy action prediction for the current time step is output. The reward for the current time step is determined based on the policy action prediction and a preset segmented reward function. The Critic network is used to predict the state information sample at the current time step and output the state value prediction at the current time step. Execute the policy action prediction for the current time step and obtain a state information sample for the next time step.

6. The intelligent dynamic decision-making method for centroid interference of a drifting corner reflector according to claim 4, characterized in that, The dominant function value is expressed as: in, This represents the dominance function value at time step t. Indicates the state Take action below Expected returns This represents the state value prediction at time step t. This represents a sample of state information at time step t. This represents the prediction of the policy action at time step t. This represents the TD error at time step t. Indicates the discount factor. Indicates GAE parameters, Indicates the end time. Indicates the time offset.

7. The intelligent dynamic decision-making method for centroid interference of a drifting corner reflector according to claim 6, characterized in that, The preset Actor network loss function is expressed as: in, This indicates the preset loss function for the Actor network. This represents the loss function of the original Actor network. Represents the entropy regularization term. Expressing expectations, Indicates the policy update ratio, The parameters of the Actor network are represented. This represents the truncation function. This indicates the trimming parameters.

8. The intelligent dynamic decision-making method for centroid interference of a drifting corner reflector according to claim 7, characterized in that, The preset Critic network loss function is expressed as follows: in, This indicates the preset loss function of the Critic network. Indicates state value prediction. Indicates the state value objective. This represents the estimated value of the old Critic network. This represents the parameters of the Critic network.

9. The intelligent dynamic decision-making method for centroid interference of a drifting corner reflector according to claim 5, characterized in that, The preset segmented reward function is expressed as follows: in, This represents the local reward function for each time step before the final time step. This represents the weighting coefficient of the local reward. The yaw angle represents the direction of tracking by the seeker radar. This represents the global reward function for the last time step. This represents the reward received when the centroid interference fails. This represents the weighting coefficient of the global reward. This indicates the distance between the seeker radar and the ship at the last time step. Indicates the number of missiles.