Non-cooperative target control strategy identification method and system based on neural network
By constructing a non-cooperative target control strategy identification model based on neural networks, the difficult problem of identifying non-cooperative target control strategies in spatial game environments is solved, efficient identification and prediction are achieved in the absence of expert knowledge, and the game confrontation ability is improved.
Patent Information
- Application Number
- CN202210976947.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-15
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-08-15
AI Technical Summary
Existing technologies have difficulty effectively identifying control strategies for non-cooperative targets in spatial game environments, especially when domain expert knowledge is insufficient, resulting in inefficient game confrontation.
A neural network-based method is used to build a non-cooperative target control strategy identification model, generate a training data set using the LQR control strategy, establish and train a neural network, and identify the control strategy of non-cooperative targets.
In the absence of expert knowledge, it can effectively identify the control strategies of non-cooperative targets, improving the efficiency and accuracy of spatial game confrontation.
Smart Images

Figure CN115222023B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aerospace space safety warning technology, and in particular to a cooperative target control strategy identification method and system, and more particularly to a neural network-based non-cooperative target control strategy identification method and system. Background Art
[0002] With the continuous advancement of aerospace technology, the number of spacecraft continues to increase, mission capabilities are rapidly improving, and the space environment is becoming increasingly complex. Spacecraft need to possess corresponding space game and countermeasure capabilities to ensure they can effectively complete their missions. Spacecraft in this space game environment can optimize their game strategies and enhance their countermeasure capabilities by enhancing their intelligent perception capabilities, accurately identifying the control strategies of space threat targets, and predicting enemy intentions. However, the vast diversity in the types and capabilities of space threat targets, coupled with the high degree of uncertainty exhibited during the game and countermeasure process, pose challenges to the identification, location, and prediction of space threat targets, severely impacting the efficiency of the game and countermeasure.
[0003] Based on feature information, the control strategy identification, classification, inversion and intention prediction of space threat targets are mainly based on template matching, expert system, Bayesian network and neural network. (1)(2) Based on the knowledge of domain experts, a template library is constructed, and the characteristic information of threat targets or hostile targets is continuously collected. The DS evidence theory is used to infer the matching degree between the target characteristics and the template library to determine the opponent's control strategy or combat intention. (3) Based on the operational knowledge of domain experts, a knowledge base is constructed, and rules are used to express the correspondence between battlefield situation and operational intention. Finally, an inference engine is used to predict operational intention based on the obtained battlefield situation. (4)(5) Based on domain expert knowledge, a Bayesian network is constructed to explore the corresponding relationship between threat target characteristics, control strategies, and combat intent. While these methods can achieve a certain degree of enemy control strategy identification or intention prediction, they require a significant amount of expert prior knowledge. However, due to the complexity of the spatial game environment and the constant changes in the confrontation mode, it is difficult for domain experts to grasp comprehensive target information in a short period of time. Their prior knowledge is insufficient to accurately quantify the relationship between threat target attributes and combat intent.
[0004] References:
[0005] (1) Zhou Zhiqiang, Qian Jiangang, Yin Kangyin, et al. A target intention prediction method based on DS evidence theory[J]. Journal of Air Force Early Warning Academy, 2014, 000(002): 116-118.
[0006] (2) Li Man, Feng Xinxi, Zhang Wei. Template-based situation assessment reasoning model and algorithm[J]. Firepower and Command Control, 2010(06):64-66.
[0007] (3)Carling R L.Naval Situation Assessment Using A Realtime Knowledge-based System[J].Naval Engineers Journal,1999,111(3):173-187.
[0008] (4)Jin Q,Gou
[0009] (5) Sun Zhaolin. Research on situation estimation method based on Bayesian network[D]. National University of Defense Technology. Summary of the Invention
[0010] To address the above technical issues, the present invention provides a neural network-based method and system for identifying control strategies for non-cooperative targets. This method analyzes the dynamic characteristics exhibited by non-cooperative targets and predicts their intentions. The control law for the threatening target is designed using the minimum principle. Training and test datasets are established by setting different weight parameters. A neural network is constructed, and the training data is imported into the network for training. The trained network is then used to identify the control strategy of the non-cooperative target.
[0011] In order to achieve the above object, the present invention adopts the following technical solutions:
[0012] A non-cooperative target control strategy identification method based on a neural network comprises the following steps:
[0013] Obtaining non-cooperative target control strategy data;
[0014] The non-cooperative target control strategy identification model is used to identify the non-cooperative target control strategy data;
[0015] Complete identification and output non-cooperative target control strategy identification results;
[0016] The steps for constructing the non-cooperative target control strategy identification model are as follows:
[0017] S1: Establish a set of non-cooperative target control strategies for non-cooperative targets;
[0018] S2: Generate different trajectory data for different control strategies in the non-cooperative target control strategy set and set data labels for the trajectory data, and at the same time build a classification dataset of different control strategies for non-cooperative targets;
[0019] S3: Randomly select data from the classification dataset of different control strategies for non-cooperative targets to construct training sets and test sets;
[0020] S4: Normalize the data in the training set to obtain training data;
[0021] S5: Determine the information and number of the neural network's input nodes, hidden layer nodes, output nodes, and activation functions between layers;
[0022] S6: Perform neural network training on the training data and update neural network parameters;
[0023] S7: Test the neural network using the test set. If the number of tests reaches the preset number of tests, proceed to the next step; otherwise, update the number of iterations and jump to S6.
[0024] S8: Complete the training, solidify the network parameters, and obtain the non-cooperative target control strategy identification model.
[0025] Furthermore, in S1, the LQR control strategy is used to establish a non-cooperative target control strategy set for the non-cooperative target. The specific steps are as follows:
[0026] Set the number of control strategies N, and the value Q i and R i is the weight matrix, according to Q i and R i Solve and obtain the corresponding control law u i ;
[0027] Among them, the value of i ranges from 1 to N; select the spatial state Control quantity u=[u x ,u y ,u z ] T , and obtain the state space model of the relative motion control equation
[0028]
[0029] in,
[0030]
[0031]
[0032] I 3×3 is the identity matrix of dimension 3×3;
[0033] n represents the angular velocity of the central reference spacecraft moving around the Earth;
[0034] x represents the radial position component of the orbit in the relative coordinate system;
[0035] y represents the position component of the flight direction in the relative coordinate system;
[0036] z represents the position component of the direction of orbital angular momentum in the relative coordinate system;
[0037] It represents the velocity component of the orbital radial direction in the relative coordinate system;
[0038] Indicates the velocity component of the flight direction in the relative coordinate system;
[0039] It represents the velocity component of the direction of orbital angular momentum in the relative coordinate system;
[0040] u x Represents the control component of the orbital radial direction in the relative coordinate system;
[0041] u y Represents the control component of the flight direction in the relative coordinate system;
[0042] u z Represents the control component of the direction of orbital angular momentum in the relative coordinate system;
[0043] In the linear state equation of non-cooperative objectives, the control u(t) is not constrained, and the target control law is designed using the optimal control theory.
[0044] u i * (t) = K i (t)·X * (t)
[0045] Among them, K i (t) is the control law coefficient matrix, solving different control laws u i , construct the non-cooperative target control strategy set U=[u1,u2,…,u i …,u N ].
[0046] Furthermore, the specific steps of S2 are: Under the near-circular orbit, the non-cooperative target strategy set U=[u1,u2,…,u i …,u N] is the control quantity u of a certain control strategy i =[u ix ,u iy ,u iz ] T Introducing the CW equation, the relative motion control equation is obtained as follows:
[0047]
[0048] The motion state data of the non-cooperative target is obtained by performing orbit integration on the relative motion control equation, including relative position, relative velocity and relative acceleration. According to the data labels set for the trajectory data generated by each control strategy, a classification data set of different control strategies for the non-cooperative target is constructed.
[0049]
[0050] x represents the radial position component of the orbit in the relative coordinate system;
[0051] y represents the position component of the flight direction in the relative coordinate system;
[0052] z represents the position component of the direction of orbital angular momentum in the relative coordinate system;
[0053] It represents the velocity component of the orbital radial direction in the relative coordinate system;
[0054] Indicates the velocity component of the flight direction in the relative coordinate system;
[0055] It represents the velocity component of the direction of orbital angular momentum in the relative coordinate system;
[0056] It represents the acceleration component of the orbital radial direction in the relative coordinate system;
[0057] Indicates the acceleration component of the flight direction in the relative coordinate system;
[0058] It represents the acceleration component in the direction of orbital angular momentum in the relative coordinate system.
[0059] Furthermore, the specific steps of S4 are:
[0060] All data in the training set are converted into data between [0,1]. The data normalization method uses the mean variance method. The formula is as follows:
[0061]
[0062] Where x k is the input data, xmean is the mean of the data series, x var is the variance of the data.
[0063] Furthermore, the specific steps of S5 are:
[0064] The input nodes of the non-cooperative target control strategy identification model are determined to be relative position, relative velocity, and relative acceleration information. The number of nodes is 9 dimensions, the number of hidden layer nodes is 15, the number of output nodes is the same as the number of control strategy types, and the hidden layer activation function and the output layer activation function both use the Sigmiod function.
[0065] Furthermore, the specific steps of S6 are:
[0066] The neural network parameters are initialized, the training data is forward transmitted through the network structure to obtain the output value, the MSE loss function is used to calculate the error between the output value and the corresponding labeled data in the labeled data set, and then the chain rule is used to backpropagate the above error to the neural network. At the same time, the steepest gradient method is used to update the neural network weights and bias parameters.
[0067] Furthermore, in S3, the ratio of the training set to the test set is 3:1.
[0068] A non-cooperative target control strategy identification system based on neural network, comprising:
[0069] A data acquisition module, used to acquire non-cooperative target control strategy data;
[0070] an identification module for identifying non-cooperative target control strategy data using a non-cooperative target control strategy identification model;
[0071] Output module, used for outputting the non-cooperative target control strategy identification result;
[0072] Among them, a non-cooperative target control strategy identification model is used to identify the non-cooperative target control strategy data. The construction method of the non-cooperative target control strategy identification model is as follows:
[0073] S1: Establish a set of non-cooperative target control strategies for non-cooperative targets;
[0074] S2: Generate different trajectory data for different control strategies in the non-cooperative target control strategy set and set data labels for the trajectory data, and at the same time build a classification dataset of different control strategies for non-cooperative targets;
[0075] S3: Randomly select data from the classification dataset of different control strategies for non-cooperative targets to construct training sets and test sets;
[0076] S4: Normalize the data in the training set to obtain training data;
[0077] S5: Determine the information and number of the neural network's input nodes, hidden layer nodes, output nodes, and activation functions between layers;
[0078] S6: Perform neural network training on the training data and update neural network parameters;
[0079] S7: Test the neural network using the test set. If the number of tests reaches the preset number of tests, proceed to the next step; otherwise, update the number of iterations and jump to S6.
[0080] S8: Complete the training, solidify the network parameters, and obtain the non-cooperative target control strategy identification model.
[0081] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the above-mentioned neural network-based non-cooperative target control strategy identification method are implemented.
[0082] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above-mentioned neural network-based non-cooperative target control strategy identification method.
[0083] Compared with the prior art, the present invention has the following beneficial effects:
[0084] The present invention provides a non-cooperative target control strategy identification method and system based on a neural network. By constructing a non-cooperative target control strategy identification model, the model is used to identify the acquired non-cooperative target data and output the non-cooperative target control strategy identification result. The non-cooperative target control strategy identification method and system based on a neural network provided by the present invention can obtain the rules between characteristic states and control strategies through self-training under the condition of insufficient knowledge of domain experts, thereby completing the strategy identification of non-cooperative target control.
[0085] Preferably, the present invention establishes a non-cooperative target control strategy set by adopting the LQR control strategy and designs the control law of the non-cooperative target based on the minimum principle, so that the method can achieve better performance indicators and is simple, convenient and easy to implement.
[0086] Preferably, the present invention normalizes the training set, that is, converts all data into data between [0, 1], thereby eliminating the order of magnitude differences between data of each dimension and avoiding large network prediction errors caused by large order of magnitude differences in input data. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1Trajectory curve diagram of non-cooperative target control strategy 1 provided by an embodiment of the present invention;
[0088] Figure 2 Trajectory curve diagram of non-cooperative target control strategy 2 provided by an embodiment of the present invention;
[0089] Figure 3 A graph showing relative position data of non-cooperative targets provided by an embodiment of the present invention;
[0090] Among them, (a) is the relative position under the control of Q1 and R1, and (b) is the relative position under the control of Q2 and R2;
[0091] Figure 4 A graph showing relative speed data of a non-cooperative target provided by an embodiment of the present invention;
[0092] Among them, (a) is the relative speed under the control of Q1, R1, and (b) is the relative speed under the control of Q2, R2;
[0093] Figure 5 A graph showing relative acceleration data of a non-cooperative target provided by an embodiment of the present invention;
[0094] Where, (a) is the relative acceleration under the control of Q1, R1, and (b) is the relative acceleration under the control of Q2, R2;
[0095] Figure 6 A graph showing the loss of a neural network according to an embodiment of the present invention;
[0096] Figure 7 A neural network classification prediction error graph provided by an embodiment of the present invention;
[0097] Figure 8 A flowchart for constructing a non-cooperative target control strategy identification model provided by an embodiment of the present invention;
[0098] Figure 9 This is a flowchart of the non-cooperative target control strategy identification based on neural network provided by the present invention. DETAILED DESCRIPTION
[0099] The present invention provides a non-cooperative target control strategy identification method based on a neural network, comprising the following steps:
[0100] Obtaining non-cooperative target control strategy data;
[0101] The non-cooperative target control strategy identification model is used to identify the non-cooperative target control strategy data;
[0102] Complete identification and output non-cooperative target control strategy identification results;
[0103] Among them, the specific steps of constructing the non-cooperative target control strategy identification model are:
[0104] S1: Establish a set of non-cooperative target control strategies for non-cooperative targets;
[0105] The LQR control strategy is used to establish a non-cooperative target control strategy set for non-cooperative targets. The specific steps are as follows:
[0106] Set the number of control strategies N, and the value Q i and R i is the weight matrix, according to Q i and R i Solve the corresponding control law u i , where i ranges from 1 to N; select the spatial state Control quantity u=[u x ,u y ,u z ] T , then the state space model of the relative motion control equation can be obtained
[0107]
[0108] in,
[0109]
[0110]
[0111] I 3×3 is the identity matrix of dimension 3×3;
[0112] n represents the angular velocity of the central reference spacecraft moving around the Earth;
[0113] x represents the radial position component of the orbit in the relative coordinate system;
[0114] y represents the position component of the flight direction in the relative coordinate system;
[0115] z represents the position component of the direction of orbital angular momentum in the relative coordinate system;
[0116] It represents the velocity component of the orbital radial direction in the relative coordinate system;
[0117] Indicates the velocity component of the flight direction in the relative coordinate system;
[0118] It represents the velocity component of the direction of orbital angular momentum in the relative coordinate system;
[0119] u x Represents the control component of the orbital radial direction in the relative coordinate system;
[0120] u y Represents the control component of the flight direction in the relative coordinate system;
[0121] u z Represents the control component of the direction of orbital angular momentum in the relative coordinate system;
[0122] In the linear state equation of non-cooperative objectives, the control u(t) is not constrained, and the target control law is designed using the optimal control theory.
[0123] u i * (t) = K i (t)·X * (t)
[0124] Among them, K i (t) is the control law coefficient matrix, solving different control laws u i , construct the non-cooperative target control strategy set U=[u1,u2,…,u i …,u N ].
[0125] S2: Generate different trajectory data for different control strategies in the non-cooperative target control strategy set and set labels for the trajectory data, and at the same time build a classification dataset of different control strategies for non-cooperative targets;
[0126] The specific steps are as follows:
[0127] Under the near-circular orbit, the non-cooperative target strategy set U=[u1,u2,…,u i …,u N ] is the control quantity u of a certain control strategy i =[u ix ,u iy ,u iz ] T Introducing the CW equation, the relative motion control equation is obtained as follows:
[0128]
[0129] The motion state data of the non-cooperative target is obtained by performing orbit integration on the relative motion control equation, including relative position, relative velocity and relative acceleration. The trajectory data generated for each control strategy corresponds to the data label, and a classification data set of different control strategies for non-cooperative targets is constructed.
[0130]
[0131] x represents the radial position component of the orbit in the relative coordinate system;
[0132] y represents the position component of the flight direction in the relative coordinate system;
[0133] z represents the position component of the direction of orbital angular momentum in the relative coordinate system;
[0134] It represents the velocity component of the orbital radial direction in the relative coordinate system;
[0135] Indicates the velocity component of the flight direction in the relative coordinate system;
[0136] It represents the velocity component of the direction of orbital angular momentum in the relative coordinate system;
[0137] It represents the acceleration component of the orbital radial direction in the relative coordinate system;
[0138] Indicates the acceleration component of the flight direction in the relative coordinate system;
[0139] It represents the acceleration component in the direction of orbital angular momentum in the relative coordinate system.
[0140] S3: Randomly select data from the classification data set of different control strategies for non-cooperative targets to construct training sets and test sets, with the ratio of training sets to test sets being 3:1.
[0141] S4: Normalize the data in the training set to obtain training data;
[0142] The specific steps are:
[0143] All data in the training set are converted into data between [0,1]. The data normalization method uses the mean variance method. The formula is as follows:
[0144]
[0145] Where x k is the input data, x mean is the mean of the data series, x var is the variance of the data.
[0146] S5: Determine the information and quantity of the input nodes, hidden layer nodes, output nodes and activation functions between layers of the non-cooperative target control strategy identification model;
[0147] The specific steps are:
[0148] The input nodes of the neural network are determined to be relative position, relative velocity, and relative acceleration information. The number of nodes is 9-dimensional, the number of hidden layer nodes is 15, the number of output nodes is the same as the number of control strategy types, and the hidden layer activation function and the output layer activation function both use the Sigmiod function.
[0149] S6: Perform neural network training on the training data and update neural network parameters;
[0150] The specific steps are:
[0151] The neural network parameters are initialized, the training data is forward transmitted through the network structure to obtain the output value, the MSE loss function is used to calculate the error between the output value and the corresponding labeled data in the labeled data set, and then the chain rule is used to backpropagate the above error to the neural network. At the same time, the steepest gradient method is used to update the neural network weights and bias parameters.
[0152] S7: Use the test set to test the neural network. If the number of tests reaches the preset number of tests, proceed to the next step; otherwise, update the number of iterations and jump to S6.
[0153] S8: Complete the training, solidify the network parameters, and obtain the non-cooperative target control strategy identification model.
[0154] The present invention also provides a non-cooperative target control strategy identification system based on a neural network, comprising: a data acquisition module for acquiring non-cooperative target control strategy data; an identification module for identifying non-cooperative target control strategy data; and an output module for outputting non-cooperative target control strategy identification results.
[0155] The present invention also provides a computer device, comprising: a memory for storing a computer program; and a processor for implementing the steps of the above-mentioned neural network-based non-cooperative target control strategy identification method when executing the computer program.
[0156] When the processor executes the computer program, the steps of the neural network-based non-cooperative target control strategy identification method are implemented, such as: obtaining non-cooperative target control strategy data; identifying the non-cooperative target control strategy data using a non-cooperative target control strategy identification model; and outputting a non-cooperative target control strategy identification result.
[0157] Alternatively, when the processor executes the computer program, it implements the functions of each module in the above system, for example: a data acquisition module for acquiring non-cooperative target control strategy data; an identification module for identifying non-cooperative target control strategy data using a non-cooperative target control strategy identification model; and an output module for outputting non-cooperative target control strategy identification results.
[0158] Exemplarily, the computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments that can complete preset functions, and the instruction segments are used to describe the execution process of the computer program in the neural network-based non-cooperative target control strategy identification device. For example, the computer program can be divided into a data acquisition module, an identification module and an output module; the specific functions of each module are as follows: a data acquisition module for acquiring non-cooperative target control strategy data; an identification module for identifying non-cooperative target control strategy data using a non-cooperative target control strategy identification model; and an output module for outputting the non-cooperative target control strategy identification result.
[0159] The neural network-based non-cooperative target control strategy identification device can be a computing device such as a desktop computer, a notebook, a PDA, and a cloud server. The neural network-based non-cooperative target control strategy identification device can include, but is not limited to, a processor and a memory. Those skilled in the art will understand that the above is an example of a neural network-based non-cooperative target control strategy identification device and does not constitute a limitation on the neural network-based non-cooperative target control strategy identification. It can include more components than the above, or combine certain components, or different components. For example, the neural network-based non-cooperative target control strategy identification device can also include input and output devices, network access devices, buses, etc.
[0160] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc. The processor is the control center of the non-cooperative target control strategy recognition device based on a neural network, and utilizes various interfaces and lines to connect various parts of the entire non-cooperative target control strategy recognition device based on a neural network.
[0161] The memory can be used to store the computer program and / or module, and the processor realizes various functions of the neural network-based non-cooperative target control strategy identification device by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory.
[0162] The memory may mainly include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created based on the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory may include a high-speed random access memory and may also include a non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0163] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the neural network-based non-cooperative target control strategy identification method.
[0164] If the module / unit integrated in the neural network-based non-cooperative target control strategy identification system is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0165] Based on this understanding, the present invention implements all or part of the processes in the aforementioned neural network-based non-cooperative target control strategy identification method by using a computer program to instruct related hardware. The computer program may be stored in a computer-readable storage medium. When executed by a processor, the computer program may implement the steps of the aforementioned roundabout channelization and signal timing optimization method. The computer program includes computer program code, which may be in source code form, object code form, executable file, or a pre-defined intermediate form.
[0166] The computer-readable storage medium may include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0167] It should be noted that the content contained in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practices in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practices, computer-readable storage media do not include electrical carrier signals and telecommunication signals.
[0168] Example
[0169] This embodiment provides a non-cooperative target control strategy identification method based on a neural network, comprising the following steps:
[0170] Obtaining non-cooperative target control strategy data;
[0171] The non-cooperative target control strategy identification model is used to identify the non-cooperative target control strategy data;
[0172] Complete identification and output non-cooperative target control strategy identification results;
[0173] Among them, the construction steps of the non-cooperative target control strategy identification model are:
[0174] S1, adopt LQR control strategy for non-cooperative targets, set the number of control strategy types N = 2, and take the value Q i , R i Weight matrix, solve the corresponding control law u i , where i ranges from 1 to N, and a non-cooperative target control strategy set U is established;
[0175] Q1=I 6×6 ×5×10 -7 R1=I 6×6 ×10 4
[0176] as well as
[0177] Q2=I 6×6 ×2.5×10 -6 R2=I 6×6 ×2×10 4
[0178] Among them, I 6×6 is the identity matrix of dimension 6×6. According to the Riccati equation
[0179]
[0180] Calculating the solution P1 of the Riccati equation yields:
[0181]
[0182] Therefore, the control law coefficient matrix is
[0183] K1=-R1-1 (t)·B T P1(t)
[0184] Calculate K1 to get
[0185]
[0186] By the Riccati equation
[0187]
[0188] Calculating the solution P2 of the Riccati equation yields:
[0189]
[0190] Therefore, the control law coefficient matrix is
[0191] K2=-R2 -1 (t)·B T P2(t)
[0192] Calculate K2 to get
[0193]
[0194] therefore
[0195] u i (t) = K i (txt)
[0196] Available
[0197] u1(t)=K1(t)·X(t)
[0198] u2(t)=K2(t)·X(t)
[0199] So, for different Q i and R i Solve the unused control law u i , i ranges from 1 to 2, forming a non-cooperative target strategy set U = [u1,u2].
[0200] S2, based on different control strategies, generates different trajectory data, sets labels for the data, and constructs a classification dataset of different strategies for non-cooperative targets;
[0201] In the near-circular orbit, the control variable is introduced into the CW equation, and the relative motion control equation is obtained as follows:
[0202]
[0203] Control quantity u i =[u ix ,uiy ,u iz ] T For the non-cooperative target strategy set U=[u1,u2,…,u i …,u N ] is a control strategy.
[0204] because
[0205] u1(t)=K1(t)·X(t)
[0206] Calculated
[0207]
[0208] Assuming the initial state of the target is (7500, 0, 0, -13.2, 15000, 0), the orbital integral recursion is performed by combining u1(t) with the relative motion control equation to obtain the motion state data of the non-cooperative target, including the relative position x, y, z, and relative speed. and relative acceleration Generate trajectory curves such as Figure 1 shown.
[0209] because
[0210] u2(t)=K2(t)·X(t)
[0211] Calculated
[0212]
[0213] Assuming the initial state of the target is (7500, 0, 0, -13.2, 15000, 0), we can perform orbital integral recursion by combining u2(t) with the relative motion control equation to obtain the motion state data of the non-cooperative target, including the relative position x, y, z, and relative velocity. and relative acceleration Generate trajectory curves such as Figure 2 As shown, the relative position curve is as follows Figure 3 As shown, the relative speed curve is as follows Figure 4 As shown, the relative acceleration curve is as follows Figure 5 shown.
[0214] 1000 sets of data are extracted from the trajectory generated under each strategy to form a data set. The trajectory data generated by each control strategy corresponds to the data label Label i , set the labels of data generated by strategy u1(t) to "0", set the labels of data generated by strategy u2(t) to "2", and construct a classification dataset of different strategies for non-cooperative targets
[0215]
[0216] S3 uses a 3:1 ratio of training set to test set, and randomly selects data from the non-cooperative target different strategy classification dataset Dataset to construct the training dataset Train_dataset and the test dataset Test_dataset; there are a total of 2000 sets of track information, 1500 sets are randomly selected as training data for network training, and 500 sets are used as test data to test the network prediction ability.
[0217] S4, before network training, the training data needs to be normalized.
[0218] Data normalization is to convert all data into data between [0,1] to eliminate the magnitude difference between different dimensional data and avoid large network prediction errors caused by large magnitude differences in input data. The data normalization method uses the mean variance method, and the function is as follows
[0219]
[0220] Where x k is the input data, x mean is the mean of the data series, x var is the variance of the data.
[0221] S5, determine the information and quantity of the neural network model input nodes, hidden layer nodes, output nodes and the activation functions between each layer.
[0222] The neural network model input nodes are selected as "relative position," "relative velocity," and "relative acceleration." The number of nodes is 9, the number of hidden layer nodes is 15, and the number of output nodes is the same as the number of control strategy types N. The hidden layer activation function f(·) and the output layer activation function g(·) use the Sigmiod function.
[0223] S6, use the training data Train_dataset to perform model training and parameter update;
[0224] The parameters of the policy recognition neural network model are initialized, and the input data in the training data set is forward transmitted through the network structure to obtain the output value. The MSE loss function is used to calculate the error between the output value and the corresponding labeled data in the labeled data set labels. Then, the chain rule is used to backpropagate this error to the neural network, and the steepest gradient method is used to update the neural network weights and bias parameters to complete a round of learning.
[0225] S7, using the test set Test_dataset to test the neural network, and identify the control strategy of non-cooperative targets based on single-moment motion state information.
[0226] The parameters of the policy recognition neural network model are initialized, and the input data in the training data set is forward transmitted through the network structure to obtain the output value. The MSE loss function is used to calculate the error between the output value and the corresponding labeled data in the labeled data set labels. Then, the chain rule is used to backpropagate this error to the neural network, and the steepest gradient method is used to update the neural network weights and bias parameters to complete a round of learning. Figure 6 The graph shows the relationship between training loss, accuracy and number of iterations.
[0227] Table 1 Network prediction ability test
[0228] sequence Number of iterations Accuracy 1 50000 76.4% 2 80000 83.0% 3 100000 88.2% 4 200000 89.8% 5 500000 89.8%
[0229] As shown in the table above, with the continuous increase in the number of BP neural network training times, the prediction ability of the network has been improved within a certain range, and the final prediction accuracy can reach 89.8%. Figure 7 This is the error distribution diagram for 200,000 iterations. Points with values not equal to 0 indicate prediction errors. The preset number of tests is 500,000. After continuous testing, when the number of tests reaches the preset number of tests, proceed to the next step.
[0230] S8, complete the training, solidify the network parameters, and obtain the non-cooperative target control strategy identification model.
[0231] The above embodiment is only one of the implementation methods that can realize the technical solution of the present invention. The scope of protection claimed by the present invention is not limited only to this embodiment, but also includes changes, replacements and other implementation methods that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention.
Claims
1. A non-cooperative target control strategy identification method based on neural network, characterized in that: The steps include: Obtaining non-cooperative target control strategy data; The non-cooperative target control strategy identification model is used to identify the non-cooperative target control strategy data; Complete identification and output non-cooperative target control strategy identification results; The steps for constructing the non-cooperative target control strategy identification model are as follows: S1: Use the LQR control strategy to establish a non-cooperative target control strategy set for the non-cooperative target. The specific steps are as follows: Set the number of control strategies N, and the value Q i and R i is the weight matrix, according to Q i and R i Solve and obtain the corresponding control law u i ; Among them, the value of i ranges from 1 to N; select the spatial state , control quantity , and obtain the state space model of the relative motion control equation in, , , , , is the identity matrix of dimension 3×3; represents the angular velocity of the central reference spacecraft moving around the Earth; Represents the radial position component of the orbit in the relative coordinate system; The position component representing the flight direction in the relative coordinate system; Represents the position component of the direction of orbital angular momentum in the relative coordinate system; It represents the velocity component of the orbital radial direction in the relative coordinate system; Indicates the velocity component of the flight direction in the relative coordinate system; It represents the velocity component of the direction of orbital angular momentum in the relative coordinate system; Represents the control component of the orbital radial direction in the relative coordinate system; Represents the control component of the flight direction in the relative coordinate system; Represents the control component of the direction of orbital angular momentum in the relative coordinate system; In the linear state equation of non-cooperative objectives, the control u(t) is not constrained, and the target control law is designed using the optimal control theory. in, is the control law coefficient matrix, solve different control laws u i , construct a set of non-cooperative target control strategies ; Design target control laws for optimal control theory; The spatial state is poor; S2: Generate different trajectory data for different control strategies in the non-cooperative target control strategy set and set data labels for the trajectory data, and at the same time build a classification dataset of different control strategies for non-cooperative targets; S3: Randomly select data from the classification dataset of different control strategies for non-cooperative targets to construct training sets and test sets; S4: Normalize the data in the training set to obtain training data; S5: Determine the information and number of the neural network's input nodes, hidden layer nodes, output nodes, and activation functions between layers; S6: Perform neural network training on the training data and update neural network parameters; S7: Test the neural network using the test set. If the number of tests reaches the preset number of tests, proceed to the next step; otherwise, update the number of iterations and jump to S6. S8: Complete the training, solidify the network parameters, and obtain the non-cooperative target control strategy identification model.
2. The non-cooperative target control strategy identification method based on neural network according to claim 1 is characterized in that: The specific steps of S2 are: Under the near-circular orbit, the non-cooperative target strategy set The control quantity of a certain control strategy Introducing the CW equation, the relative motion control equation is obtained as follows: The motion state data of the non-cooperative target is obtained by performing orbit integration on the relative motion control equation, including relative position, relative velocity and relative acceleration. According to the data labels set for the trajectory data generated by each control strategy, a classification data set of different control strategies for the non-cooperative target is constructed. Represents the radial position component of the orbit in the relative coordinate system; The position component representing the flight direction in the relative coordinate system; Represents the position component of the direction of orbital angular momentum in the relative coordinate system; It represents the velocity component of the orbital radial direction in the relative coordinate system; Indicates the velocity component of the flight direction in the relative coordinate system; It represents the velocity component of the direction of orbital angular momentum in the relative coordinate system; It represents the acceleration component of the orbital radial direction in the relative coordinate system; Indicates the acceleration component of the flight direction in the relative coordinate system; It represents the acceleration component in the direction of orbital angular momentum in the relative coordinate system.
3. The method for identifying a non-cooperative target control strategy based on a neural network according to claim 1, characterized in that: The specific steps of S4 are: All data in the training set are converted into data between [0,1]. The data normalization method uses the mean variance method. The formula is as follows: Where, For input data, is the mean of the data series, is the variance of the data.
4. The method for identifying a non-cooperative target control strategy based on a neural network according to claim 1, characterized in that: The specific steps of S5 are: The input nodes of the non-cooperative target control strategy identification model are determined to be relative position, relative velocity, and relative acceleration information. The number of nodes is 9 dimensions, the number of hidden layer nodes is 15, the number of output nodes is the same as the number of control strategy types, and the hidden layer activation function and the output layer activation function both use the Sigmiod function.
5. The method for identifying a non-cooperative target control strategy based on a neural network according to claim 1, characterized in that: The specific steps of S6 are: The neural network parameters are initialized, the training data is forward transmitted through the network structure to obtain the output value, the MSE loss function is used to calculate the error between the output value and the corresponding labeled data in the labeled data set, and then the chain rule is used to backpropagate the above error to the neural network. At the same time, the steepest gradient method is used to update the neural network weights and bias parameters.
6. The method for identifying a non-cooperative target control strategy based on a neural network according to claim 1, characterized in that: In S3, the ratio of the training set to the test set is 3:
1.
7. A non-cooperative target control strategy identification system based on a neural network, used to implement the steps of the non-cooperative target control strategy identification method based on a neural network according to any one of claims 1 to 6, characterized in that: include: A data acquisition module, used to acquire non-cooperative target control strategy data; an identification module for identifying non-cooperative target control strategy data using a non-cooperative target control strategy identification model; Output module, used for outputting the non-cooperative target control strategy identification result; The steps for constructing the non-cooperative target control strategy identification model are as follows: S1: Use the LQR control strategy to establish a non-cooperative target control strategy set for the non-cooperative target. The specific steps are as follows: Set the number of control strategies N, and the value Q i and R i is the weight matrix, according to Q i and R i Solve and obtain the corresponding control law u i ; Among them, the value of i ranges from 1 to N; select the spatial state , control quantity , and obtain the state space model of the relative motion control equation in, , , , , is the identity matrix of dimension 3×3; represents the angular velocity of the central reference spacecraft moving around the Earth; Represents the radial position component of the orbit in the relative coordinate system; The position component representing the flight direction in the relative coordinate system; Represents the position component of the direction of orbital angular momentum in the relative coordinate system; It represents the velocity component of the orbital radial direction in the relative coordinate system; Indicates the velocity component of the flight direction in the relative coordinate system; It represents the velocity component of the direction of orbital angular momentum in the relative coordinate system; Represents the control component of the orbital radial direction in the relative coordinate system; Represents the control component of the flight direction in the relative coordinate system; Represents the control component of the direction of orbital angular momentum in the relative coordinate system; In the linear state equation of non-cooperative objectives, the control u(t) is not constrained, and the target control law is designed using the optimal control theory. in, is the control law coefficient matrix, solve different control laws u i , construct a set of non-cooperative target control strategies ; Design target control laws for optimal control theory; The spatial state is poor; S2: Generate different trajectory data for different control strategies in the non-cooperative target control strategy set and set data labels for the trajectory data, and at the same time build a classification dataset of different control strategies for non-cooperative targets; S3: Randomly select data from the classification dataset of different control strategies for non-cooperative targets to construct training sets and test sets; S4: Normalize the data in the training set to obtain training data; S5: Determine the information and number of the neural network's input nodes, hidden layer nodes, output nodes, and activation functions between layers; S6: Perform neural network training on the training data and update neural network parameters; S7: Test the neural network using the test set. If the number of tests reaches the preset number of tests, proceed to the next step; otherwise, update the number of iterations and jump to S6. S8: Complete the training, solidify the network parameters, and obtain the non-cooperative target control strategy identification model.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the neural network-based non-cooperative target control strategy identification method are implemented as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the neural network-based non-cooperative target control strategy identification method are implemented.
Citation Information
Patent Citations
Method and apparatus for preventing malicious attack
WO2020248687A1
Deep neural network hyperparameter optimization method, electronic device and storage medium
WO2021007812A1