A method for underwater communication confidentiality of dual unmanned ship system based on proximal strategy optimization
By adopting a dual unmanned ship system based on near-end strategy optimization in underwater wireless communication, the existing technology lack of defense and the complexity of traditional optimization algorithms is solved, and efficient communication confidentiality performance and system security are achieved.
Patent Information
- Application Number
- CN202410676273.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-05-29
AI Technical Summary
Existing underwater wireless communication security technologies are difficult to effectively resist active attacks without prior knowledge, and traditional optimization algorithms perform poorly in complex scenarios.
A dual unmanned ship system based on near-end strategy optimization is adopted to achieve autonomous motion and interference signal transmission of unmanned ships by obtaining system status information, constructing Markov decision-making processes, and using reinforcement learning to optimize action.
Without affecting legal communication, the confidentiality performance and system security of underwater wireless communication are improved, which is better than the single unmanned ship solution, and reduces the overall cost of building an underwater wireless communication system.
Smart Images

Figure CN118748807B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater wireless security communication, and in particular to an underwater communication confidentiality method for a dual unmanned ship system based on proximal strategy optimization. Background Art
[0002] Wireless communication security is crucial to protecting communication privacy and data integrity. Although traditional encryption and authentication technologies can prevent unauthorized access and data tampering, they are insufficient in resisting active attack methods. Even if the communication data is encrypted, attackers may still obtain sensitive information by stealing wireless signals and trying to decrypt them. In a man-in-the-middle attack, attackers may also forge identities to bypass authentication and control or interfere with communication links. Therefore, it is difficult to fully respond to various active attacks by relying solely on traditional passive protection, and more advanced and active security protection strategies are needed. The above problems also apply to underwater secure wireless communication scenarios. Due to the particularity of the underwater environment, wireless signal transmission is more difficult and faces more security threats, such as the deployment of eavesdropping devices and deliberate interference. Therefore, it is particularly necessary to adopt advanced active security protection strategies in underwater wireless communications to ensure the absolute security and reliability of communications.
[0003] In order to enhance the security of wireless communications, researchers have proposed active defense schemes including friendly interference. The working principle of friendly interference is to introduce additional noise signals into the communication link, which is "friendly" to legitimate users and aims to disrupt potential eavesdroppers, thereby blocking eavesdropping to a certain extent and providing protection for legitimate user communications. However, it requires certain prior knowledge to introduce noise in a targeted manner. The introduction of additional noise will also increase power consumption and complexity, and may reduce spectrum utilization and energy efficiency. Therefore, when using friendly interference, it is necessary to balance communication confidentiality and system efficiency, reasonably design noise introduction parameters, maximize the anti-eavesdropping effect and minimize the impact on system efficiency. Friendly interference can also be combined with other security technologies to give play to synergistic advantages and further enhance communication security and reliability.
[0004] Although friendly jamming provides a new active defense method for secure communications in theory, there are some limitations in existing research. On the one hand, friendly jamming requires certain prior knowledge, such as the location information of the eavesdropper, but such prior knowledge is often lacking in practical applications. The lack of prior knowledge will bring difficulties to the design and application of friendly jamming, and it is impossible to determine the specific location of the eavesdropper, which affects the formulation and effect of the strategy. Therefore, how to design an effective friendly jamming scheme in the absence of prior knowledge is an urgent problem to be solved. On the other hand, when establishing system models and solving optimization problems, researchers who use friendly jamming defense schemes generally face the challenge of non-convex constrained optimization. Non-convex optimization often requires more advanced methods, higher computational complexity, and higher requirements on computing power and latency. Therefore, when designing friendly jamming schemes, it is crucial to simplify model assumptions and the choice of solution algorithms, and it is necessary to find a balance between theoretical analysis and practicality.
[0005] Traditional iterative optimization algorithms have many defects, such as strict requirements on initial conditions and problem convexity, and easy to fall into local optimal solutions for high-dimensional complex problems. Although some intelligent optimization algorithms have global search capabilities, they often converge slowly and have low computational efficiency. In contrast, deep reinforcement learning (DRL) provides an end-to-end optimization paradigm based on trial-and-error learning, which can autonomously acquire experience and continuously adjust decisions to adapt to complex environments, breaking through many limitations of traditional algorithms. Introducing DRL into the field of wireless communication security optimization can not only get rid of the convexity assumption and initial point constraints, but also open up new ways and tools for the autonomous optimization design of active defense strategies in complex scenarios, which is expected to significantly improve the security and robustness of the system and has broad application prospects. Summary of the invention
[0006] The present invention aims to address the gaps in existing applications and provide a method for underwater communication confidentiality of a dual unmanned ship system based on proximal strategy optimization, aiming to solve the problems existing in practical applications.
[0007] To achieve the above object, the present invention provides the following technical solution: a method for underwater communication confidentiality of a dual unmanned ship system based on proximal strategy optimization, comprising the following steps:
[0008] S1: Each unmanned ship obtains the required system status information;
[0009] S2: Convert the system state information into the state space form of the constructed Markov decision process;
[0010] S3: Use the method based on proximal policy optimization to obtain the best action execution at the current moment and perform training at the same time;
[0011] S4: The unmanned ship continues to move, transmit and interfere;
[0012] S5: When the position of the unmanned ship and the confidentiality rate change, return to step S1.
[0013] As a preferred technical solution of the present invention: it also includes a system for implementing the communication security method, and the specific system structure is as follows:
[0014] P1: a water unmanned ship communication base station R-USV;
[0015] P2: A friendly jammer unmanned vessel, J-USV;
[0016] P3: Eve, an eavesdropper whose exact location is unknown;
[0017] P4: K underwater communication nodes that receive information transmission from the unmanned ship R-USV.
[0018] As a preferred technical solution of the present invention: the system implementing the communication confidentiality method deploys an unmanned ship R-USV on the water surface as a temporary mobile base station to communicate with several nodes of the underwater network to transmit confidential information; and also deploys a friendly interference unmanned ship J-USV to transmit hydroacoustic interference signals to interfere with the working frequency band of the eavesdropper Eve; the modeling of the system is as follows:
[0019] T1: The underwater scene is established in a three-dimensional Euclidean coordinate system. K underwater wireless network nodes are deployed on a plane with a depth of h. The underwater network nodes and the R-USV and J-USV on the water surface are represented as: q k =(x k ,y k ,h),q R [n]=(x R [n],y R [n],0) and q J [n]=(x J [n],y J [n],0), and use Indicates a collection of unmanned ships;
[0020] T2: Assuming the actual location of the eavesdropper is unknown, use (x e ,y e ,h e ) is the actual location of the eavesdropper, (x e0 ,y e0 ,h e0 ) as the estimated position, (Δx e ,Δy e ,Δh e ) is the estimation error, and the relationship between the three is expressed as:
[0021] x e =x e0 +Δx e,y e =y e0 +Δy e ,h e =h e0 +Δh e (1)
[0022]
[0023] Thus, the estimated position of the eavesdropper is limited to (x e0 ,y e0 ,h e0 ) is the center of the sphere, and the radius of the sphere is r e ;
[0024] T3: The task duration T is divided into N equal time slots d t , that is, T = Nd t , then when N is large enough, the navigation of each unmanned ship can be regarded as static;
[0025] T4: Communication link channel gain h from R-USV to underwater node k at time slot n Rk [n] Communication link with i-USV and Eve Re [n] are respectively represented as:
[0026]
[0027]
[0028] Where A0 is the unit normalization constant, f is the carrier frequency, λ is the energy diffusion factor that describes the propagation geometry. The absorption coefficient a(f) is approximated by the Thorp formula:
[0029]
[0030] T5: The reachability rate between underwater node k and eavesdropper Eve is expressed as:
[0031]
[0032]
[0033] in is additive Gaussian white noise, p R [n] and p J [n] are the transmission powers of R-USV and J-USV in time slot n, h Jk [n] and h Je [n] are the interference signal channels from J-USV to underwater nodes and eavesdroppers respectively;
[0034] T6: Let the unmanned ship trajectory be recorded as The transmitting power of the unmanned ship is recorded as The joint optimization problem is defined as:
[0035]
[0036] in Constraint C1 prevents collisions between unmanned ships, d RJ [n] is the distance between two unmanned ships, requiring the two unmanned ships to maintain a minimum distance Constraint C2 ensures that the speed of each unmanned ship is within a reasonable range, where V min With V max are the minimum and maximum speeds of the unmanned ship, respectively. Constraints C3 and C4 limit the maximum average transmission power and peak power of the unmanned ship. and are the average transmission power and peak transmission power upper limit of each unmanned ship, p i is the current transmission power of each unmanned ship, and the constraint C5 limits the steering angle of each unmanned ship, where is the steering angle of each unmanned boat to prevent rollover.
[0037] As a preferred technical solution of the present invention: the Markov decision process modeling steps in step S2 are as follows:
[0038] S2-1: State space s t Expressed as where {x i [t],y i [t]} is the horizontal position of i-USV, is the sum of the power generated by i-USV in the previous t time slots, d ik [t] is the distance between i-USV and underwater node, d ie [t] is the distance between i-USV and the eavesdropper, d RJ [t] is the distance between the two unmanned boats;
[0039] S2-2: The action space is designed as where p R [t] and p J [t], and v R [t] and v J [t] are the transmission power, steering angle and navigation speed of each unmanned boat;
[0040] S2-3: The reward function is represented by r t = r(s t ,a t) = r1[t] + αr2[t], where r1[t] is the sum of the average confidentiality rate and penalty in time slot t, and α is the weight factor, expressed as:
[0041]
[0042] Among them, L1(t) ensures that the i-USV navigates within the legal area, otherwise it is based on penalties; L2(t) is the power penalty term to ensure that the power constraint is met, and L3(t) is the collision penalty term to prevent collisions between i-USVs, which are defined as
[0043]
[0044] where x min , x max ,y min ,y max are the motion ranges of the unmanned ship on the x and y axes,
[0045]
[0046]
[0047] Among them, δ1, δ2, δ3 are penalty coefficients, and r2[t] is used to guide the USV to the end point, which is expressed as:
[0048]
[0049] Where dis(t)i is the distance from the i-USV to the end point at time slot t.
[0050] As a preferred technical solution of the present invention: the proximal strategy optimization training step in step S3:
[0051] S3-1: Initialization: experience playback D, actor network weights θ and θ', critic network weight ω;
[0052] S3-2: Get the observation state st of the initial environment;
[0053] S3-3: Select action at according to the current state st;
[0054] S3-4: Execute action at to obtain reward rt;
[0055] S3-5: Transfer to new state st +1 ;
[0056] S3-6: Convert the above tuple (st,at,rt,st +1 ) is stored in the experience playback D;
[0057] S3-7: Sample a small batch c from D;
[0058] S3-8: According to Calculation based on strategy π θ The action value function under the sample data of , where γ is the discount factor;
[0059] S3-9: Calculate the advantage function in is the mathematical expectation;
[0060] S3-10: Passed Perform gradient descent to update the network parameters, where and ε is the trust region control coefficient;
[0061] S3-11: If each USV reaches the end point, the training ends, otherwise return to step S3-3.
[0062] Compared with the prior art, the present invention has the following beneficial effects:
[0063] The method of the present invention achieves high confidentiality performance by actively moving to the estimated area of the eavesdropper and transmitting interference signals, moving to the network node to provide services, and controlling power to reduce interference and counteract eavesdropping without affecting legitimate communications. The proximal strategy optimization based on reinforcement learning enables the unmanned ship to continuously accumulate experience from the environment and optimize decisions, ultimately achieving excellent performance that is superior to the single unmanned ship solution. Although the introduction of active defense will cause energy loss, the experimental results fully prove the effectiveness and superiority of this method in enhancing the confidentiality performance of underwater wireless communications.
[0064] The method of the present invention not only improves the flexibility and adaptability of communication, avoids communication interruption and signal blocking caused by the impact of the marine environment on traditional communication nodes, but also ensures the communication performance of the system through a joint optimization algorithm, so that the unmanned ship can operate continuously for a long time in the marine environment, and reduces the overall cost of building an underwater wireless communication system. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 It is a schematic diagram of the process of the method proposed by the present invention;
[0066] Figure 2 It is a schematic diagram of an underwater communication model of a dual unmanned ship system based on proximal strategy optimization of the present invention;
[0067] Figure 3 The navigation track diagrams of the unmanned ship under two schemes of the embodiments of the present invention;
[0068] Figure 4The transmission power variation diagram of the unmanned ship under two schemes of the embodiment of the present invention;
[0069] Figure 5 The total energy variation diagram of the unmanned ship under two schemes of the embodiment of the present invention;
[0070] Figure 6 is a graph showing changes in average confidentiality rate over time under two solutions of an embodiment of the present invention;
[0071] Figure 7 Graph showing changes in algorithm cumulative rewards under two schemes of an embodiment of the present invention. DETAILED DESCRIPTION
[0072] Based on the attached schematic diagrams, several embodiments of the present invention are now described in detail to make the advantages and features of the present invention clearer and more specific, in order to provide those skilled in the art with a clear understanding of the protection scope of the claims of the present invention.
[0073] Embodiment: The underwater communication confidentiality method of the dual unmanned ship system based on proximal strategy optimization described in the present invention comprises: an unmanned ship communication base station on the water, a friendly interference unmanned ship, an eavesdropper with an unknown specific location, and k underwater communication nodes that receive information transmission from the unmanned ship;
[0074] like Figure 1-2 As shown, the present invention provides a method for underwater communication confidentiality of a dual unmanned ship system based on proximal strategy optimization, and the specific steps are as follows:
[0075] An unmanned ship R-USV is deployed on the water surface as a temporary mobile base station to communicate with several nodes of the underwater network to transmit confidential information; a friendly jamming unmanned ship J-USV is deployed to transmit hydroacoustic jamming signals to interfere with the working frequency band of the eavesdropper Eve; the modeling of this system is as follows:
[0076] The underwater scene is established in a three-dimensional Euclidean coordinate system, and K underwater wireless network nodes are deployed on a plane with a depth of h. The underwater network nodes and the R-USV and J-USV on the water surface are represented as: q k =(x k ,y k ,h),q R [n]=(x R [n],y R [n],0) and q J [n]=(x J [n],y J [n],0), and use Indicates a collection of unmanned ships;
[0077] Assuming the actual location of the eavesdropper is unknown, use (x e ,ye ,h e ) is the actual location of the eavesdropper, (x e0 ,y e0 ,h e0 ) as the estimated position, (Δx e ,Δy e ,Δh e ) is the estimation error, and the relationship between the three is expressed as:
[0078] x e =x e0 +Δx e ,y e =y e0 +Δy e ,h e =h e0 +Δh e (1)
[0079]
[0080] Thus, the estimated position of the eavesdropper is limited to (x e0 ,y e0 ,h e0 ) is the center of the sphere, and the radius of the sphere is r e ;
[0081] The task duration T is divided into N equal time slots d t , that is, T = Nd t . Then when N is large enough, the navigation of each unmanned ship can be regarded as static.
[0082] In time slot n, the communication link channel gain h from R-USV to underwater node k is Rk [n] Communication link with i-USV and Eve Re [n] are respectively represented as:
[0083]
[0084]
[0085] Where A0 is the unit normalization constant, f is the carrier frequency, λ is the energy diffusion factor that describes the propagation geometry. The absorption coefficient a(f) is approximated by the Thorp formula:
[0086]
[0087] The reachability rate between the underwater node k and the eavesdropper Eve is expressed as:
[0088]
[0089]
[0090] in is additive Gaussian white noise, p R [n] and p J [n] are the transmission powers of R-USV and J-USV in time slot n, h Jk [n] and h Je [n] are the interference signal channels from J-USV to underwater nodes and eavesdroppers respectively;
[0091] Let the unmanned ship trajectory be recorded as The transmitting power of the unmanned ship is recorded as The joint optimization problem is defined as:
[0092]
[0093] in Constraint C1 prevents collisions between unmanned ships, d RJ [n] is the distance between two unmanned ships, requiring the two unmanned ships to maintain a minimum distance Constraint C2 ensures that the speed of each unmanned ship is within a reasonable range, where V min With V max are the minimum and maximum speeds of the unmanned ship, respectively. Constraints C3 and C4 limit the maximum average transmission power and peak power of the unmanned ship. and are the average transmission power and peak transmission power upper limit of each unmanned ship, pi is the current transmission power of each unmanned ship, and the constraint condition C5 limits the steering angle of each unmanned ship, where is the steering angle of each unmanned boat to prevent rollover.
[0094] The steps to model the Markov decision process are as follows:
[0095] The state space st is expressed as where {xi[t],yi[t]} is the horizontal position of the i-USV, is the sum of the power generated by i-USV in the previous t time slots, dik[t] is the distance between i-USV and the underwater node, die[t] is the distance between i-USV and the eavesdropper, and d RJ [t] is the distance between the two unmanned boats;
[0096] The action space is designed as where p R [t] and p J [t], and v R [t] and v J[t] are the transmission power, steering angle and navigation speed of each unmanned boat;
[0097] The reward function is expressed as rt = r(st,at) = r1[t] + αr2[t], where r1[t] is the sum of the average confidentiality rate and penalty in time slot t, and α is the weight factor, expressed as:
[0098]
[0099] Among them, L1(t) ensures that the i-USV navigates within the legal area, otherwise it is based on penalties; L2(t) is the power penalty term to ensure that the power constraint is met, and L3(t) is the collision penalty term to prevent collisions between i-USVs, which are defined as
[0100]
[0101] Among them, xmin, xmax, ymin, and ymax are the motion ranges of the unmanned ship on the x and y axes respectively.
[0102]
[0103]
[0104] Among them, δ1, δ2, δ3 are penalty coefficients, and r2[t] is used to guide the USV to the end point, which is expressed as:
[0105]
[0106] where dis(t) i is the distance between i-USV and the end point at time slot t.
[0107] Proximal strategy optimization training steps:
[0108] 1): Initialization: experience replay D, actor network weights θ and θ', critic network weight ω;
[0109] 2): Get the observation state s of the initial environment t ;
[0110] 3): According to the current state s t Select action a t ;
[0111] 4): Execute action a t Get rewards t ;
[0112] 5): Transfer to new state s t+1 ;
[0113] 6): The above tuple (s t ,at ,r t ,s t+1 ) is stored in the experience playback D;
[0114] 7): Sample a small batch c from D;
[0115] 8): According to Calculation based on strategy π θ The action value function under the sample data of , where γ is the discount factor;
[0116] 9): Calculate the advantage function in E is the mathematical expectation;
[0117] 10): Pass Perform gradient descent to update the network parameters, where and ε is the trust region control coefficient;
[0118] 11): If each USV reaches the end point, the training ends, otherwise return to step 3;
[0119] Experimental results: Figure 3 As shown in the figure, the navigation trajectories under the two schemes are different. In the dual unmanned ship scheme, the J-USV will sail towards the estimated range of the eavesdropper without affecting the R-USV, and then circle around the estimated range of the eavesdropper at the lowest speed. During this period, the J-USV will transmit jamming signals to interfere with the underwater eavesdropper. At the same time, the R-USV will move toward the underwater wireless network node to provide reliable underwater communication services while avoiding the eavesdropper as much as possible. In contrast, when using a single unmanned ship scheme, the unmanned ship will move away from the eavesdropper to enhance the security of the system. However, this will cause the unmanned ship to be farther away from the node, resulting in a longer trajectory.
[0120] like Figure 4 As shown in the figure, the dual unmanned ship scheme increases the transmission power of the R-USV earlier than the single unmanned ship scheme to reduce the impact of friendly interference on the system. In addition, the J-USV will increase the transmission power near the estimated range of the underwater eavesdropper to effectively counter underwater eavesdropping attacks.
[0121] Figure 5 The display shows the overall energy changes of the system under two different schemes. From the results, it can be seen that by introducing active and friendly interference to counter eavesdropping attacks, the system will incur energy loss, resulting in the total system energy consumption of the dual unmanned ship scheme being higher and faster than that of the single unmanned ship scheme.
[0122] like Figure 6The figure shows the change of the system confidentiality rate over time T under the scheme. The confidentiality rate of both schemes increases over time. However, due to the effective interference of J-USV to eavesdroppers at unknown locations, the dual unmanned ship scheme is significantly better than the single unmanned ship scheme in terms of system security. In contrast, the single unmanned ship scheme performs poorly in terms of system security due to the lack of active interference to underwater eavesdroppers.
[0123] Figure 7 The trend of the algorithm reward in the two cases is shown. Since the unmanned ship has no knowledge of the environment in the initial training stage and the choice of action is almost random, the cumulative reward is at a relatively low value and fluctuates greatly at the beginning. Once the unmanned ship has accumulated enough samples over a period of time, it starts to use the accumulated samples to train the network, and the reward can eventually converge to a higher value, which can fully demonstrate the effectiveness of the proposed joint optimization algorithm. And from Figure 7 It can also be seen that the reward for the dual unmanned ship solution is higher than that for the single unmanned ship solution, which means that the dual unmanned ship solution is better for improving the confidentiality performance of the system.
[0124] This embodiment shows through experimental comparison that the underwater confidential communication method of the dual unmanned ship system based on proximal strategy optimization proposed by the present invention achieves higher confidentiality performance by actively moving to the estimated area of the eavesdropper and transmitting interference signals, moving to the network node to provide services, and power control to reduce interference and counter eavesdropping without affecting legitimate communications. The proximal strategy optimization based on reinforcement learning enables the unmanned ship to continuously accumulate experience from the environment and optimize decisions, and finally achieves excellent performance that is better than the single unmanned ship solution. Although the introduction of active defense will cause energy loss, the experimental results fully prove the effectiveness and superiority of this method in enhancing the confidentiality performance of underwater wireless communications.
[0125] The above examples are only specific implementation forms of the scheme of the present invention, and their description is intended to enhance understanding rather than limit the protection scope of the scheme of the present invention. It is worth pointing out that for professionals in this technical field, without departing from the overall concept and design principle of the scheme of the present invention, some equivalent conversions, modifications or extensions can be made to the scheme of the present invention for specific application scenarios or actual needs, and these variations and optimizations should be included in the protection scope of the scheme of the present invention. The protection scope of the scheme of the present invention shall be based on the statements made in the claims.
Claims
1. A method for underwater communication confidentiality of a dual unmanned ship system based on proximal strategy optimization, characterized in that: The method is performed by the following system, and the specific system composition is as follows: P1: a water unmanned ship communication base station R-USV; P2: A friendly jammer unmanned vessel, J-USV; P3: Eve, an eavesdropper whose exact location is unknown; P4: K underwater communication nodes that receive information transmission from the unmanned ship R-USV; The system implementing the communication confidentiality method deploys an unmanned ship R-USV on the water surface as a temporary mobile base station to communicate with several nodes of the underwater network to transmit confidential information; it also deploys a friendly jamming unmanned ship J-USV to transmit hydroacoustic jamming signals to interfere with the working frequency band of the eavesdropper Eve; the modeling of the system is as follows: T1: The underwater scene is established in a three-dimensional Euclidean coordinate system. K underwater wireless network nodes are deployed on a plane with a depth of h. The underwater network nodes and the R-USV and J-USV on the water surface are represented as: q k =(x k ,y k ,h),q R [n]=(x R [n],y R [n],0) and q J [n]=(x J [n],y J [n],0), and use Indicates a collection of unmanned ships; T2: Assuming the actual location of the eavesdropper is unknown, use (x e ,y e ,h e ) is the actual location of the eavesdropper, (x e0 ,y e0 ,h e0 ) as the estimated position, (Δx e ,Δy e ,Δh e ) is the estimation error, and the relationship between the three is expressed as: x e =x e0 +Δx e ,y e =y e0 +Δy e ,h e =h e0 +Δh e (1) Thus, the estimated position of the eavesdropper is limited to (x e0 ,y e0 ,h e0 ) is the center of the sphere, and the radius of the sphere is r e ; T3: The task duration T is divided into N equal time slots d t , that is, T = Nd t , then when N is large enough, the navigation of each unmanned ship can be regarded as static; T4: Communication link channel gain h from R-USV to underwater node k at time slot n Rk [n] Communication link with i-USV and Eve Re [n] are respectively represented as: Where A0 is the unit normalization constant, f is the carrier frequency, λ is the energy diffusion factor describing the propagation geometry, and the absorption coefficient a(f) is approximated by the Thorp formula: T5: The reachability rate between underwater node k and eavesdropper Eve is expressed as: in is additive Gaussian white noise, p R [n] and p J [n] are the transmission powers of R-USV and J-USV in time slot n, h Jk [n] and h Je [n] are the interference signal channels from J-USV to underwater nodes and eavesdroppers respectively; T6: Let the unmanned ship trajectory be recorded as The transmitting power of the unmanned ship is recorded as The joint optimization problem is defined as: in Constraint C1 prevents collisions between unmanned ships, d RJ [n] is the distance between two unmanned ships, requiring the two unmanned ships to maintain a minimum distance Constraint C2 ensures that the speed of each unmanned ship is within a reasonable range, where V min With V max are the minimum and maximum speeds of the unmanned ship, respectively. Constraints C3 and C4 limit the maximum average transmission power and peak power of the unmanned ship. and are the average transmission power and peak transmission power upper limit of each unmanned ship, p i is the current transmission power of each unmanned ship, and the constraint C5 limits the steering angle of each unmanned ship, where is the steering angle of each unmanned boat to prevent rollover; The specific steps of this method are as follows: S1: Each unmanned ship obtains the required system status information; S2: Convert the system state information into the state space form of the constructed Markov decision process; S3: Use the method based on proximal policy optimization to obtain the best action execution at the current moment and perform training at the same time; S4: The unmanned ship continues to move, transmit and interfere; S5: When the position of the unmanned ship and the confidentiality rate change, return to step S1.
2. According to claim 1, a method for underwater communication confidentiality of a dual unmanned ship system based on proximal strategy optimization is characterized in that: The steps of Markov decision process modeling in step S2 are as follows: S2-1: State space s t Expressed as where {x i [t],y i [t]} is the horizontal position of i-USV, is the sum of the power generated by i-USV in the previous t time slots, d ik [t] is the distance between i-USV and underwater node, d ie [t] is the distance between i-USV and the eavesdropper, d RJ [t] is the distance between the two unmanned boats; S2-2: The action space is designed as where p R [t] and p J [t], and v R [t] and v J [t] are the transmission power, steering angle and navigation speed of each unmanned boat; S2-3: The reward function is represented by r t = r(s t ,a t ) = r1[t] + αr2[t], where r1[t] is the sum of the average confidentiality rate and penalty in time slot t, and α is the weight factor, expressed as: Among them, L1(t) ensures that the i-USV navigates within the legal area, otherwise it is based on penalties; L2(t) is the power penalty term to ensure that the power constraint is met, and L3(t) is the collision penalty term to prevent collisions between i-USVs, which are defined as where x min , x max ,y min ,y max are the motion ranges of the unmanned ship on the x and y axes, Among them, δ1, δ2, δ3 are penalty coefficients, and r2[t] is used to guide the USV to the end point, which is expressed as: where dis(t) i is the distance between i-USV and the end point at time slot t.
3. According to claim 1, a method for underwater communication confidentiality of a dual unmanned ship system based on proximal strategy optimization is characterized in that: The proximal strategy optimization training steps described in step S3 are: S3-1: Initialization: experience playback D, actor network weights θ and θ', critic network weight ω; S3-2: Get the observation state s of the initial environment t ; S3-3: According to the current state s t Select action a t ; S3-4: Execute action a t Get rewards t ; S3-5: Transition to new state s t+1 ; S3-6: The above tuple (s t ,a t ,r t ,s t+1 ) is stored in the experience playback D; S3-7: Sample a small batch c from D; S3-8: According to Calculation based on strategy π θ The action value function under the sample data of , where γ is the discount factor; S3-9: Calculate the advantage function in is the mathematical expectation; S3-10: Passed Perform gradient descent to update the network parameters, where and ε is the trust region control coefficient; S3-11: If each USV reaches the end point, the training ends, otherwise return to step S3-3.
Citation Information
Patent Citations
Unmanned aerial vehicle intelligent reflecting surface secure transmission method based on deep reinforcement learning
CN115052285A
Communication game method for assisting UAV to resist dynamic interference based on RIS
CN117614509A