Cross-sea-air optical communication alignment method based on multi-agent deep reinforcement learning

Through the method based on the deep reinforcement learning of multiple agents, the cross-sea-optical communication alignment model is trained, which solves the problem that the central and cross-sea-optical communication alignment process of the existing technology is susceptible to environmental disturbances, and realizes stable alignment and efficient communication between aircraft and underwater vehicles.

CN120017160APending Publication Date: 2025-05-16SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510030289.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing cross-sea optical communication technology is susceptible to environmental disturbances during the alignment process, resulting in a decline in system performance and relies on the performance of hardware equipment, which is prone to communication instability problems.

Method used

A method based on deep reinforcement learning of multi-agents is adopted to form a three-dimensional dynamic sea surface model by applying the linear superposition method, and combined with a preset cross-sea air-optical communication channel model, the cross-sea air-optical communication alignment model is trained to achieve alignment between aircraft and underwater vehicles.

Benefits of technology

Without relying on sea surface relay, cross-sea air-optical communication alignment between the aircraft and autonomous underwater vehicles is achieved, improving the anti-interference ability and communication stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017160A_ABST
    Figure CN120017160A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-sea-air optical communication alignment method based on multi-agent deep reinforcement learning, and the method comprises the steps: superposing waveform data of a sea area where a target underwater vehicle is located together through employing a linear superposition method, and forming a three-dimensional dynamic sea surface model; combining the three-dimensional dynamic sea surface model and a preset cross-sea-air optical communication channel model, taking an aircraft as an intelligent agent, training a cross-sea-air optical communication alignment model based on a preset central light spot calculation method, a Koch snowflake formation structure and a mixed reward function through Markov decision, and obtaining a target cross-sea-air optical communication alignment model; and inputting the position information of the target underwater vehicle into the target cross-sea-air optical communication alignment model to obtain a target position of an aircraft corresponding to the target underwater vehicle, thereby realizing alignment of the aircraft and the underwater vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a cross-sea and air optical communication alignment method based on multi-agent deep reinforcement learning. Background Art

[0002] Traditional cross-sea and air medium communication mainly relies on the surface buoy relay system, and the aerial platform and the underwater platform communicate with the buoy through radio electromagnetic waves and acoustic waves respectively. Although this method solves the needs of cross-media communication to a certain extent, it has obvious defects. For example, the high bandwidth and low latency of air radio communication are seriously mismatched with the low rate and high latency of underwater acoustic wave communication, resulting in the limitation of the overall communication link efficiency. In addition, due to the existence of buoy relay, information transmission needs to be forwarded multiple times, which increases the delay. In order to overcome these problems, direct cross-sea and air optical communication has become a promising solution, especially blue-green laser communication, which shows significant advantages by utilizing the excellent penetration characteristics of the ocean optical window in water. Compared with traditional methods, blue-green laser communication can provide high-bandwidth communication rate, low latency and low interference, and is suitable for real-time communication needs in complex dynamic environments. However, due to the dynamic complexity of the sea surface environment and the high directivity of the laser, the alignment process of cross-sea and air optical communication faces significant challenges. Existing methods usually rely on the high performance of hardware equipment. However, when the environmental disturbance exceeds the response range of the hardware, the system performance may drop sharply. Summary of the invention

[0003] Based on this, it is necessary to propose a communication method and related equipment between an aircraft and an underwater vehicle to address the alignment difficulties and unstable communication problems of the existing cross-sea and air optical communication technology.

[0004] In a first aspect, a method for cross-sea and air optical communication alignment based on multi-agent deep reinforcement learning is provided, the method comprising:

[0005] The waveform data of the sea area where the target underwater vehicle is located is superimposed using the linear superposition method to form a three-dimensional dynamic sea surface model;

[0006] In combination with the three-dimensional dynamic sea surface model and the preset cross-sea and air optical communication channel model, the aircraft is used as an intelligent agent, and a cross-sea and air optical communication alignment model based on a preset central spot calculation method, a Koch snowflake formation structure and a hybrid reward function is trained through Markov decision training to obtain a target cross-sea and air optical communication alignment model;

[0007] The position information of the target underwater vehicle is input into the target cross-sea-air optical communication alignment model to obtain the target position of the aircraft corresponding to the target underwater vehicle, thereby achieving alignment between the aircraft and the underwater vehicle.

[0008] Optionally, the step of applying a linear superposition method to superimpose waveform data of the sea area where the target underwater vehicle is located to form a three-dimensional dynamic sea surface model includes:

[0009] The PM wave spectrum is used to obtain the frequency distribution of the waves in the sea area where the target underwater vehicle is located, and the specific shape of the spectrum is determined according to the environmental parameters to obtain the wave spectrum;

[0010] The wave spectrum is divided into a plurality of frequency bands by an equal energy frequency division method, and the energy of each frequency band is evenly distributed to obtain a sub-spectrum;

[0011] The ITTC directional distribution function is used to describe the energy distribution of waves in different directions on different sub-spectra, and the propagation direction of the waves is divided into multiple angle intervals.

[0012] The frequency, direction and amplitude of the sea wave are used as waveform data, and the sea wave waves with different waveform data are superimposed together using a linear superposition method to form a three-dimensional dynamic sea surface model.

[0013] Optionally, the preset cross-sea and air optical communication channel model is:

[0014]

[0015] Where K is the responsivity of the photodiode, τ is the transmittance, is the gain of the optical concentrator, η is the angle at which the light is incident on the receiving plane, and Ar is the effective receiving area of ​​the photodetector. dcc is the distance between the receiver and the center of the light spot, is the pointing angle of the transmitter relative to the optical link, and Pt is the average transmitted optical power.

[0016] Optionally, the preset central spot calculation method is:

[0017]

[0018] Among them, Pcenter is the center point of the light spot, X K is the X-axis coordinate of the center point of the light spot, Y k is the Y-axis coordinate of the center point of the light spot, ρ i is the number of points in the grid to which the center point of the light spot belongs.

[0019] Optionally, the hybrid reward function includes: a core drone reward for shortening the distance between the core aircraft and the center of the light spot, an auxiliary drone reward for maintaining the initial cluster structure of the aircraft, a collision reward function for avoiding collisions when multiple aircraft fly simultaneously, and a global reward for avoiding the negative impact of individual optimal behavior on the overall goal.

[0020] In a second aspect, a cross-sea and air optical communication alignment device based on multi-agent deep reinforcement learning is provided, the device comprising:

[0021] A data acquisition module, used to apply a linear superposition method to superimpose waveform data of the sea area where the target underwater vehicle is located to form a three-dimensional dynamic sea surface model;

[0022] A model training module is used to combine the three-dimensional dynamic sea surface model and the preset cross-sea and air optical communication channel model, take the aircraft as an intelligent agent, and train the cross-sea and air optical communication alignment model based on the preset central spot calculation method, the Koch snowflake formation structure and the hybrid reward function through Markov decision training to obtain the target cross-sea and air optical communication alignment model;

[0023] The alignment module is used to input the position information of the target underwater vehicle into the target cross-sea and air optical communication alignment model, obtain the target position of the aircraft corresponding to the target underwater vehicle, and realize the alignment of the aircraft and the underwater vehicle.

[0024] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the above-mentioned cross-sea and air optical communication alignment method based on multi-agent deep reinforcement learning are implemented.

[0025] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned cross-sea and air optical communication alignment method based on multi-agent deep reinforcement learning are implemented.

[0026] This application uses the linear superposition method to superimpose the waveform data of the sea area where the target underwater vehicle is located to form a three-dimensional dynamic sea surface model; combining the three-dimensional dynamic sea surface model and the preset cross-sea and air optical communication channel model, taking the aircraft as an intelligent body, and training the cross-sea and air optical communication alignment model based on the preset central spot calculation method, Koch snowflake formation structure and mixed reward function through Markov decision training, to obtain the target cross-sea and air optical communication alignment model; the position information of the target underwater vehicle is input into the target cross-sea and air optical communication alignment model to obtain the target position of the aircraft corresponding to the target underwater vehicle, so as to realize the alignment of the aircraft and the underwater vehicle. It is possible to realize the cross-sea and air optical communication alignment between the aircraft and the autonomous underwater vehicle without relying on the sea surface relay, which solves the problem of relying on the performance of hardware equipment and being susceptible to interference in the current cross-sea and air optical communication process. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0028] in:

[0029] Figure 1 A flowchart of a method for cross-sea and air optical communication alignment based on multi-agent deep reinforcement learning in one embodiment;

[0030] Figure 2 A three-dimensional dynamic sea surface model diagram in a cross-sea and air optical communication alignment method based on multi-agent deep reinforcement learning in one embodiment;

[0031] Figure 3 It is a structural block diagram of a cross-sea and air optical communication alignment device based on multi-agent deep reinforcement learning in one embodiment;

[0032] Figure 4 is a structural block diagram of a computer device in one embodiment;

[0033] Figure 5 It is a structural block diagram of a computer device in another embodiment. DETAILED DESCRIPTION

[0034] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0035] The present invention is described in detail below through specific embodiments.

[0036] See also Figure 1 As shown, Figure 1 A schematic flow chart of a cross-sea and air optical communication alignment method based on multi-agent deep reinforcement learning provided in an embodiment of the present invention includes the following steps:

[0037] S101, applying a linear superposition method to superimpose waveform data of the sea area where the target underwater vehicle is located to form a three-dimensional dynamic sea surface model;

[0038] In a possible implementation, the step of applying a linear superposition method to superimpose waveform data of the sea area where the target underwater vehicle is located to form a three-dimensional dynamic sea surface model includes:

[0039] The frequency distribution of the waves in the sea area where the target underwater vehicle is located is obtained by using the PM wave spectrum, and the specific shape of the spectrum is determined according to the environmental parameters to obtain the wave spectrum;

[0040] The wave spectrum is divided into a plurality of frequency bands by an equal energy frequency division method, and the energy of each frequency band is evenly distributed to obtain a sub-spectrum;

[0041] The ITTC directional distribution function is used to describe the energy distribution of waves in different directions on different sub-spectra, and the propagation direction of the waves is divided into multiple angle intervals.

[0042] The frequency, direction and amplitude of the sea wave are used as waveform data, and the sea wave waves with different waveform data are superimposed together using a linear superposition method to form a three-dimensional dynamic sea surface model.

[0043] Exemplarily, the frequency distribution of the waves is obtained using the PM wave spectrum, and the specific shape of the spectrum is determined according to parameters such as wind speed. Then, the spectrum is divided into multiple frequency bands by the equal energy frequency division method, and the energy of each frequency band is evenly distributed. At the same time, the ITTC directional distribution function is used to describe the energy distribution of the wave in different directions, and the propagation direction of the waves is divided into multiple angle intervals.

[0044] Finally, the linear superposition method is used to superimpose multiple waves of different frequencies, directions and amplitudes to form a three-dimensional dynamic sea surface model. The mathematical expression of the PM wave spectrum is:

[0045]

[0046] Where g is the acceleration due to gravity, U is the wind speed, ω is the wave angular frequency, and υ is a factor related to the wave energy density. is the proportionality coefficient.

[0047] The frequency of each harmonic ω i for:

[0048]

[0049] The amplitude of each harmonic a i Equal, according to the principle of equal energy division, if the simulation frequency band is divided into M parts, they are all:

[0050]

[0051] The ITTC directional distribution function is shown below:

[0052]

[0053] Where θ is the azimuth angle. Assume that the number of sampling points is N. If θ∈[-π, π] and the azimuth angle interval is Δθ, then the azimuth angle θ of each component wave is j for:

[0054]

[0055] Based on the above analysis, the mathematical model of three-dimensional waves is as follows:

[0056]

[0057] Among them, t0 represents a certain moment, k i is the wave number; ε is the initial phase, which is a random variable uniformly distributed on [0, 2π], and we get Figure 2 The three-dimensional dynamic sea surface model shown.

[0058] S102, combining the three-dimensional dynamic sea surface model and the preset cross-sea and air optical communication channel model, taking the aircraft as an intelligent agent, and training the cross-sea and air optical communication alignment model based on the preset center spot calculation method, the Koch snowflake formation structure and the hybrid reward function through Markov decision training, to obtain the target cross-sea and air optical communication alignment model;

[0059] For example, aircraft such as drones and aircraft, etc., abstract the cross-sea and air optical communication alignment process into a Markov decision process, that is, the Markov decision process is a mathematical framework used to simulate the situation in which a decision maker makes a decision in a series of states, and each action of the decision maker may lead to a change in state and some form of reward. It means finding an optimal strategy to maximize the long-term reward so that the cross-sea and air optical communication alignment between the drone and the autonomous underwater vehicle can be achieved.

[0060] For example, the stability of the UAV formation using the Koch snowflake structure can reduce the impact on the UAVs on the sea surface disturbed by wind speed and enhance the stability of the alignment process.

[0061] S103, inputting the position information of the target underwater vehicle into the target cross-sea-air optical communication alignment model, obtaining the target position of the aircraft corresponding to the target underwater vehicle, and realizing the alignment between the aircraft and the underwater vehicle.

[0062] The waveform data of the sea area where the target underwater vehicle is located is superimposed by applying the linear superposition method to form a three-dimensional dynamic sea surface model; the three-dimensional dynamic sea surface model and the preset cross-sea and air optical communication channel model are combined, and the aircraft is used as an intelligent body. Through Markov decision training, a cross-sea and air optical communication alignment model based on the preset central spot calculation method, Koch snowflake formation structure and mixed reward function is obtained to obtain a target cross-sea and air optical communication alignment model; the position information of the target underwater vehicle is input into the target cross-sea and air optical communication alignment model to obtain the target position of the aircraft corresponding to the target underwater vehicle, so as to achieve alignment between the aircraft and the underwater vehicle. It is possible to achieve cross-sea and air optical communication alignment between the aircraft and the autonomous underwater vehicle without relying on the sea surface relay, which solves the problem of relying on hardware equipment performance and being susceptible to interference in the current cross-sea and air optical communication process.

[0063] In a possible implementation manner, the preset cross-sea and air optical communication channel model is:

[0064]

[0065] Where K is the responsivity of the photodiode, τ is the transmittance, is the gain of the optical concentrator, η is the angle at which the light is incident on the receiving plane, and Ar is the effective receiving area of ​​the photodetector. dcc is the distance between the receiver and the center of the light spot, is the pointing angle of the transmitter relative to the optical link, and Pt is the average transmitted optical power.

[0066] Exemplarily, since the received signal strength I of the aircraft mainly depends on three factors: path loss, effective receiving area and light irradiance, the preset cross-sea and air optical communication channel model is as follows:

[0067]

[0068] Where K is the responsivity of the photodiode, τ is the transmittance, is the gain of the optical concentrator, η is the angle at which the light is incident on the receiving plane, A r is the effective receiving area of ​​the photodetector. cc is the distance between the receiver and the center of the light spot, is the pointing angle of the transmitter relative to the optical link, P t is the average transmitted optical power.

[0069] Light beams are affected by path loss when propagating in water and air. Path loss in air is mainly manifested as geometric loss. The attenuation of light intensity caused by the increase in distance during the propagation of the optical signal is shown below:

[0070]

[0071] Among them, d a represents the distance between the receiver and the incident point, and Ψ is the divergence angle of the light source. In addition, the path loss in water is mainly caused by scattering and absorption. Based on the exponential decay model, the attenuation of light in water can be expressed as To represent, a and b represent the attenuation coefficients of scattering and absorption phenomena respectively. The interference of ambient light noise is as follows:

[0072] P noise =p bg ΔλA r n 2

[0073] Among them, p bg is the spectral irradiance, Δλ is the spectral width, and n is the refractive index of the optical concentrator

[0074] In a possible implementation manner, the preset central spot calculation method is:

[0075]

[0076] Among them, Pcenter is the center point of the light spot, X K is the X-axis coordinate of the center point of the light spot, Y k is the Y-axis coordinate of the center point of the light spot, ρ i is the number of points in the grid to which the center point of the light spot belongs.

[0077] For example, assume that the autonomous underwater vehicle is at an underwater depth of h1, the receiving plane height of the core drone is h2, and the beam divergence half angle is θ half = 2°. According to the polar coordinates, the coordinates of the points on the x-axis and y-axis form the P×Q matrices X0 and Y0, respectively, as shown below:

[0078]

[0079] Substitute the coordinates of these points into the three-dimensional dynamic sea surface model at time t=t0 and solve the corresponding z coordinate matrix Z0, as shown below:

[0080] z0(p, q)=z[X0(p, q), Y0(p, q)]

[0081] The grid formed by the sampling points divides the light spot on the sea surface into PxQ small facets. Based on the above analysis, we can know that each wavefront ∑ ij The equation for is as follows:

[0082] z ij (x, y) = a ij cos(ω i tki xcosθ j -k i ysinθ j +ε ij )

[0083] At time t0, the normal equation of the sea surface at the incident point M0 (x0, y0, z0) is as follows:

[0084]

[0085] The laser transmitter is abstracted as a point light source with coordinates M(0, 0, -h1). The receiving plane can be expressed as follows:

[0086] Π1: z1=h2

[0087] The vector of the incident light ray can be expressed as M′0(x′0, y′0, z′0) is the intersection of the normal line and plane ∏1. By combining the normal line equation and the receiving plane, M′0 can be obtained. Then the vector of the normal line can be expressed as

[0088] Whether it is reflection or refraction, the incident light, normal and refracted light can jointly determine a plane, which is defined as the light action plane Π2. The plane Π2 can be represented by the coordinates of the three points M, M0 and M'0, as shown below:

[0089]

[0090] The vector of the refracted light can be expressed as Point M1 (x1, y1, z1) is on both plane ∏1 and plane Π2, so it satisfies the following equation:

[0091]

[0092] There are two unknowns in the above formula, x1 and y1, so another equation needs to be established. According to the law of refraction, the refraction angle α can be determined out With the incident angle α in The relationship is as follows:

[0093]

[0094] Angle of incidence α in for and The angle between them is as follows:

[0095]

[0096] Refraction angle α out yes and The angle between them is as follows:

[0097]

[0098] By combining the above formulas, we can get the coordinates of the refraction point M1 on the receiving plane. The calculated refraction points have a set of real solutions and a set of imaginary solutions. and The angle of refraction is also equal to the angle of refraction α out .Will and The angle between io ,and The angle between io Obviously, α io ≤α′ io , only when the incident angle α in =0, α io =α′ io Therefore, by comparing the angle α io and α′ io The cosine of is used to determine the correct solution.

[0099] According to the above analysis, the set of refraction points corresponding to all incident points is Since the refraction points are all on the receiving plane at a fixed height, the X and Y axes are divided into 100 equally spaced intervals, which define a 100×100 grid as shown below:

[0100]

[0101] The two-dimensional binning method is used to calculate the number of points in each grid cell. c(m,n) represents the number of points contained in the grid of the mth row and the nth column, as shown below:

[0102]

[0103] Among them, 1 is the indicator function, when the point (x i ,y i ) is in the grid (m,n), the value is 1, otherwise it is 0.

[0104] For each point (x i ,y i ), its density ρi is the number of points in the grid it belongs to, as follows:

[0105] ρ i =c(binX(i),binY(i))

[0106] Among them, binX(i) and binY(i) are points (x i ,y i) is the row and column index of the grid where . Finally, the point with the largest density P center As the center point of the light spot, as shown below:

[0107]

[0108] In one possible implementation, the hybrid reward function includes: a core drone reward for shortening the distance between the core aircraft and the center of the light spot, an auxiliary drone reward for maintaining the initial cluster structure of the aircraft, a collision reward function for avoiding collisions when multiple aircraft fly simultaneously, and a global reward for avoiding the negative impact of individual optimal behavior on the overall goal.

[0109] For example, the distance between drones i and j is d ij , the safety distance is d safe .

[0110] Core Drone Rewards: In order to achieve alignment quickly and accurately, the Core Drone needs to track the center of the light spot. Rewards are obtained by shortening the distance between the Core Drone and the center of the light spot:

[0111]

[0112] Among them, R cent represents the radius of the communication area around the center of the light spot, d cc Indicates the distance from the core drone to the center of the light spot, I c is the signal strength received by the core drone, Ω1, Ω2, Ω3 represent weight coefficients respectively. When the distance between the core drone and the center of the spot is less than this radius, the alignment is considered complete, and a variable reward Ω1·(R cent -d cc ). In addition, the closer the core drone is to the center of the light spot, the greater the reward it receives. Conversely, the farther the core drone is from the center of the light spot, the greater the reward it receives. c will decrease, so Ω2·I c -Ω3·d cc Will become smaller.

[0113] Auxiliary UAV Reward: When the core UAV is heading to the center of the light spot, the auxiliary UAV follows the core UAV to maintain the initial cluster structure, reducing the impact of wind speed on the movement of the UAV. When the auxiliary UAV is in the desired relative position area, a positive reward is given. On the contrary, a negative reward is given as a penalty, as shown below:

[0114]

[0115] Collision Reward: It is very important to consider the collision problem when multiple drones are flying at the same time. Therefore, a corresponding reward function is set to solve the collision problem when multiple drones are flying at the same time, as shown below:

[0116]

[0117] Global reward: In order to obtain the optimal solution during the alignment process between the core UAV and the autonomous underwater vehicle, a global reward is set to prevent the individual optimal behavior from having a negative impact on the overall goal. Therefore, the global reward of each UAV is as follows:

[0118]

[0119] Second, as Figure 3 As shown, a cross-sea and air optical communication alignment device based on multi-agent deep reinforcement learning is provided, and the device includes:

[0120] The data acquisition module 201 is used to apply the linear superposition method to superimpose the waveform data of the sea area where the target underwater vehicle is located to form a three-dimensional dynamic sea surface model;

[0121] The model training module 202 is used to combine the three-dimensional dynamic sea surface model and the preset cross-sea and air optical communication channel model, take the aircraft as an intelligent agent, and train the cross-sea and air optical communication alignment model based on the preset center spot calculation method, the Koch snowflake formation structure and the hybrid reward function through Markov decision training to obtain the target cross-sea and air optical communication alignment model;

[0122] The alignment module 203 is used to input the position information of the target underwater vehicle into the target cross-sea and air optical communication alignment model to obtain the target position of the aircraft corresponding to the target underwater vehicle, thereby achieving alignment between the aircraft and the underwater vehicle.

[0123] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a service-side method for cross-sea and air optical communication alignment based on multi-agent deep reinforcement learning.

[0124] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the client side of a cross-sea and air optical communication alignment method based on multi-agent deep reinforcement learning.

[0125] In one embodiment, a computer device is proposed, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program: applying a linear superposition method to superimpose waveform data of the sea area where a target underwater vehicle is located to form a three-dimensional dynamic sea surface model; combining the three-dimensional dynamic sea surface model and a preset cross-sea-air optical communication channel model, taking the aircraft as an intelligent agent, and training a cross-sea-air optical communication alignment model based on a preset central spot calculation method, a Koch snowflake formation structure, and a hybrid reward function through Markov decision training to obtain a target cross-sea-air optical communication alignment model; inputting the position information of the target underwater vehicle into the target cross-sea-air optical communication alignment model to obtain a target position of the aircraft corresponding to the target underwater vehicle, thereby achieving alignment between the aircraft and the underwater vehicle.

[0126] In one embodiment, a computer-readable storage medium is proposed, which stores a computer program. When the computer program is executed by a processor, the following steps are implemented: waveform data of the sea area where the target underwater vehicle is located are superimposed using a linear superposition method to form a three-dimensional dynamic sea surface model; combining the three-dimensional dynamic sea surface model and a preset cross-sea-to-air optical communication channel model, taking the aircraft as an intelligent agent, and training a cross-sea-to-air optical communication alignment model based on a preset central spot calculation method, a Koch snowflake formation structure, and a hybrid reward function through Markov decision training to obtain a target cross-sea-to-air optical communication alignment model; inputting the position information of the target underwater vehicle into the target cross-sea-to-air optical communication alignment model to obtain the target position of the aircraft corresponding to the target underwater vehicle, thereby achieving alignment between the aircraft and the underwater vehicle.

[0127] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0128] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0129] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0130] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A cross-sea and air optical communication alignment method based on multi-agent deep reinforcement learning, characterized in that: The method comprises: The waveform data of the sea area where the target underwater vehicle is located is superimposed using the linear superposition method to form a three-dimensional dynamic sea surface model; In combination with the three-dimensional dynamic sea surface model and the preset cross-sea and air optical communication channel model, the aircraft is used as an intelligent agent, and a cross-sea and air optical communication alignment model based on a preset central spot calculation method, a Koch snowflake formation structure and a hybrid reward function is trained through Markov decision training to obtain a target cross-sea and air optical communication alignment model; The position information of the target underwater vehicle is input into the target cross-sea-air optical communication alignment model to obtain the target position of the aircraft corresponding to the target underwater vehicle, thereby achieving alignment between the aircraft and the underwater vehicle.

2. According to claim 1, the cross-sea and air optical communication alignment method based on multi-agent deep reinforcement learning is characterized in that: The step of applying the linear superposition method to superimpose waveform data of the sea area where the target underwater vehicle is located to form a three-dimensional dynamic sea surface model includes: The frequency distribution of the waves in the sea area where the target underwater vehicle is located is obtained by using the PM wave spectrum, and the specific shape of the spectrum is determined according to the environmental parameters to obtain the wave spectrum; The wave spectrum is divided into a plurality of frequency bands by an equal energy frequency division method, and the energy of each frequency band is evenly distributed to obtain a sub-spectrum; The ITTC directional distribution function is used to describe the energy distribution of waves in different directions on different sub-spectra, and the propagation direction of the waves is divided into multiple angle intervals. The frequency, direction and amplitude of the sea wave are used as waveform data, and the sea wave waves with different waveform data are superimposed together using a linear superposition method to form a three-dimensional dynamic sea surface model.

3. The cross-sea and air optical communication alignment method based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: The preset cross-sea and air optical communication channel model is: Where K is the responsivity of the photodiode, τ is the transmittance, g(ζ) is the gain of the optical concentrator, η is the angle at which the light is incident on the receiving plane, Ar is the effective receiving area of ​​the photodetector, dcc is the distance between the receiver and the center of the light spot, is the pointing angle of the transmitter relative to the optical link, and Pt is the average transmitted optical power.

4. The cross-sea and air optical communication alignment method based on multi-agent deep reinforcement learning according to claim 3 is characterized in that: The calculation method of the preset central spot is: Among them, Pcenter is the center point of the light spot, X K is the X-axis coordinate of the center point of the light spot, Y k is the Y-axis coordinate of the center point of the light spot, ρ i is the number of points in the grid to which the center point of the light spot belongs.

5. The cross-sea and air optical communication alignment method based on multi-agent deep reinforcement learning according to claim 2 is characterized in that: The hybrid reward function includes: a core drone reward for shortening the distance between the core aircraft and the center of the light spot, an auxiliary drone reward for maintaining the initial cluster structure of the aircraft, a collision reward function for avoiding collisions when multiple aircraft fly simultaneously, and a global reward for avoiding the negative impact of individual optimal behavior on the overall goal.

6. A cross-sea and air optical communication alignment device based on multi-agent deep reinforcement learning, characterized in that: The device comprises: A data acquisition module, used to apply a linear superposition method to superimpose waveform data of the sea area where the target underwater vehicle is located to form a three-dimensional dynamic sea surface model; A model training module is used to combine the three-dimensional dynamic sea surface model and the preset cross-sea and air optical communication channel model, take the aircraft as an intelligent agent, and train the cross-sea and air optical communication alignment model based on the preset central spot calculation method, the Koch snowflake formation structure and the hybrid reward function through Markov decision training to obtain the target cross-sea and air optical communication alignment model; The alignment module is used to input the position information of the target underwater vehicle into the target cross-sea and air optical communication alignment model, obtain the target position of the aircraft corresponding to the target underwater vehicle, and realize the alignment of the aircraft and the underwater vehicle.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the cross-sea and air optical communication alignment method based on multi-agent deep reinforcement learning are implemented as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the cross-sea and air optical communication alignment method based on multi-agent deep reinforcement learning are implemented as described in any one of claims 1 to 5.