An intelligent underwater acoustic communication method and device based on hybrid training sequence strategy
Through an intelligent underwater acoustic communication method based on a hybrid training sequence strategy, the deep reinforcement learning DDPG model and the Markov decision model are used to adaptively adjust the OFDM training sequence carrier parameters, which solves the shortcomings of traditional underwater acoustic channel detection methods and achieves efficient underwater acoustic communication.
Patent Information
- Application Number
- CN202411543531.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2044-10-31
AI Technical Summary
Traditional underwater acoustic channel anomaly detection methods are difficult to effectively reflect the true characteristics of the channel and are difficult to adapt to the communication needs of different underwater environments and scenarios, resulting in high bit error rates, slow delayed feedback, and large additional overhead.
An intelligent underwater acoustic communication method based on a hybrid training sequence strategy is adopted. The deep reinforcement learning DDPG model and Markov decision model are used to adaptively adjust the OFDM training sequence carrier parameters. The self-attention mechanism and shared actor-critic network structure are combined to optimize the time-frequency resource consumption and communication performance of the underwater acoustic channel.
It realizes adaptive adjustment of OFDM parameters in unknown or dynamically changing underwater acoustic channel environments, improves communication efficiency, reduces computational complexity and communication overhead, and adapts to the communication needs of different underwater environments and scenarios.
Smart Images

Figure CN119628801B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of underwater acoustic signal communication, and more specifically, to an intelligent underwater acoustic communication method and device based on a hybrid training sequence strategy. Background Art
[0002] Underwater acoustic communication is a communication technology used in underwater environments. Its basic principle is to modulate the frequency, amplitude, and other parameters of underwater acoustic signals. The acoustic waves are then transmitted to the target location using the acoustic velocity and pressure differences created by the water medium. The received acoustic waves are then demodulated to recover the true information. Orthogonal frequency division multiplexing (OFDM) is a multi-carrier modulation technique that splits a single high-speed information data stream into multiple parallel, lower-speed data streams, modulating them on different orthogonal subcarrier channels and superimposing them to form an OFDM signal. OFDM has advantages such as robustness to inter-symbol interference, high spectral efficiency, and ease of implementation, making it widely used in underwater acoustic communication. However, the underwater acoustic channel is time-, space-, and frequency-varying, and is affected by numerous factors, such as water depth, water temperature, ocean currents, seabed topography, and marine life. These factors cause the underwater acoustic channel's characteristics, such as Doppler spread, delay spread, coherence bandwidth, and coherence time, to vary over time and space, thus affecting the transmission quality of the OFDM signal.
[0003] In order to adapt to the changes in the underwater acoustic channel, it is usually necessary to insert training sequence subcarriers (Training Sequence Symbol) known to both the sender and receiver into the signal, and dynamically adjust the parameters of the OFDM training sequence carrier to correctly estimate the current underwater acoustic channel state. However, most traditional methods use fixed training sequence subcarrier modulation, or use parameter adjustment based on channel feedback information. These methods have problems such as high bit error rate, slow delayed feedback, and large additional overhead. For example, this preset strategy is prone to cause too many modulated subcarriers under stable underwater acoustic conditions, resulting in excessive consumption of system time and frequency resources. On the contrary, under large-scale time-varying underwater acoustic conditions, it may be difficult to effectively reflect the true characteristics of the channel. Therefore, this type of traditional method is difficult to adapt to the communication needs of different underwater environments and scenarios. Summary of the Invention
[0004] In response to the defects of the existing technology, the purpose of this application is to provide an intelligent underwater acoustic communication method and device based on a hybrid training sequence strategy, aiming to solve the problems that traditional underwater acoustic channel anomaly detection methods cannot effectively reflect the true characteristics of the channel and are difficult to adapt to the communication needs of different underwater environments and scenarios.
[0005] To achieve the above objectives, the present application provides an intelligent underwater acoustic communication method based on a hybrid training sequence strategy, comprising the following steps:
[0006] S1 uses orthogonal frequency division multiplexing to divide the subcarrier into multiple subchannels, and adopts a mixed training sequence strategy to modulate the subcarrier parameters so that the multiple subchannels transmit underwater acoustic signals in parallel;
[0007] Based on the channel characteristics of underwater acoustic modulation communication environment, S2 models the subcarrier parameter adjustment problem of the mixed training sequence in the orthogonal frequency division multiplexing method as a Markov decision model;
[0008] S3 designs a reward function based on bit error rate and transmission performance planning based on the Markov decision model, and builds a deep reinforcement learning (DDPG) model. The deep reinforcement learning (DDPG) model includes a policy network structure, an internal network structure, and a value network structure.
[0009] S4 collects the output underwater acoustic signal and the received underwater acoustic signal as positive and negative samples of the underwater acoustic signal, uses the positive and negative samples of the underwater acoustic signal to train the deep reinforcement learning DDPG model, and uses the trained deep reinforcement learning DDPG model to restore the actual underwater acoustic signal.
[0010] This application proposes an intelligent underwater acoustic communication method based on hybrid training sequence carrier orthogonal frequency division multiplexing (TC-OFDM). This method can autonomously learn and select the optimal OFDM training sequence carrier parameters according to the real-time state of the underwater acoustic channel using deep deterministic policy gradient (DDPG). That is, the reinforcement learning DDPG model is used to set dynamically variable subcarriers for the underwater acoustic signal, thereby enhancing the OFDM signal quality and achieving the goal of maximizing communication efficiency.
[0011] Furthermore, the step of enabling the plurality of sub-channels to transmit underwater acoustic signals in parallel includes:
[0012] S101 divides the underwater acoustic channel according to the available idle bandwidth and caches the underwater acoustic signal data to be transmitted in segments;
[0013] S102 divides the cached multiple segments of underwater acoustic signal data into signal data queues, and adds the training sequence subcarriers to the signal data queues;
[0014] S103 uses multiple threads to perform an inverse fast Fourier transform on the underwater acoustic signal after the training sequence subcarrier is added, and serially transmits the underwater acoustic signal;
[0015] S104 repeats steps S101 to S103 until all the underwater acoustic signals to be transmitted are transmitted.
[0016] Furthermore, the specific steps of constructing the Markov decision model include:
[0017] S201 defines the state space of the Markov decision model: selects all underwater acoustic signal data segments with the same <hydrophone ID, hydrophone address, duration, modulation protocol> from the underwater acoustic signal output in step S1, and groups the underwater acoustic signal data segments;
[0018] S202 defines the action space ai of the Markov decision model as:
[0019] ai={A_a,A_t,A_f}
[0020] Where A_a is the training sequence type, A_t represents the set of interpolation positions of the training sequence in the time domain, and A_f represents the set of interpolation positions of the training sequence in the frequency domain.
[0021] S203 decomposes the vector of the action space into vectors including two dimensions of time domain and frequency domain, and converts the discrete action space parameters into a probability distribution vector.
[0022] Furthermore, in step S202, if the training sequence subcarrier interpolation is performed in the time domain, A_a is a normal training sequence, and its time domain and frequency domain positions are recorded; if the training sequence subcarrier interpolation is performed in both the time domain and the frequency domain, A_a is an orthogonal training sequence, and its time domain and frequency domain positions are recorded.
[0023] Furthermore, in step S3, the step of designing a reward function based on bit error rate and transmission performance planning includes:
[0024] S301 builds an average bit error rate calculation model:
[0025]
[0026] Where α and β are modulation coefficients, and F(x) is the cumulative distribution function of the underwater acoustic channel signal-to-noise ratio;
[0027] S302 constructs a signal power factor model of the underwater acoustic signal transmission process:
[0028]
[0029] Among them, R p is the signal power factor, p(x) is the OFDM signal power value transmitted in the underwater acoustic channel, x represents the xth underwater acoustic signal transmission sequence, and N is the total number of underwater acoustic signal transmissions;
[0030] S303 constructs the reward function based on the average bit error rate calculation model and the signal power factor model.
[0031] Furthermore, the reward function is expressed as:
[0032]
[0033] Among them, w k is the underwater acoustic signal weight, r i,j is the linear relationship between the average bit error rate and signal power during the system transmission process, b is a constant parameter, s t is the state space, a t is the action space, R(s t, a t ) represents the sum of the weighted benefits of each transmitted underwater acoustic signal, R D is the average bit error rate, R p is the signal power factor, k represents the kth underwater acoustic transmission sequence, and n is the total number of transmission sequence numbers.
[0034] Furthermore, attention modules are integrated into both the strategy network structure and the value network structure.
[0035] In a second aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the programs stored in the memory are executed, the processor is used to execute the intelligent underwater acoustic communication method described in the first aspect or any possible implementation of the first aspect.
[0036] In a third aspect, the present application provides a computer-readable storage medium storing a computer program. When the computer program runs on a processor, the processor executes the intelligent underwater acoustic communication method described in the first aspect or any possible implementation of the first aspect.
[0037] In a fourth aspect, the present application provides a computer program product. When the computer program product runs on a processor, it enables the processor to execute the intelligent underwater acoustic communication method described in the first aspect or any possible implementation of the first aspect.
[0038] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0039] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the existing technologies:
[0040] (1) The detection method of this application combines the deep reinforcement learning (DDPG) model technology, proposes an OFDM parameter optimization framework based on deep reinforcement learning, and designs a hybrid deep neural network structure based on a shared actor-critic (Co-Actor-Critic). This hybrid deep neural network structure combines the self-attention mechanism and can effectively extract the time-frequency characteristics and time-varying characteristics of the underwater acoustic channel. This hybrid deep neural network structure can effectively extract the time-frequency characteristics and time-varying characteristics of the underwater acoustic channel and output the optimal OFDM parameters; a reward function based on bit error rate and transmission power is designed. This reward function can comprehensively consider communication quality and energy efficiency, guide the hybrid deep neural network structure to learn and optimize, has good generalization ability and robustness, and can cope with unknown or dynamically changing underwater acoustic channel environments.
[0041] (2) This application models the training sequence carrier parameter adjustment problem in OFDM as a Markov decision process (MDP), where the state space is a vector composed of underwater acoustic channel characteristics, the action space is a set composed of optional OFDM parameters, and the reward function is a function defined by communication performance indicators (bit error rate, throughput, etc.); in addition, a deep deterministic policy gradient reinforcement learning algorithm (DDPG) is used to train an agent (Agent), and the training process consists of three deep networks, among which the internal network (Internal) is used to transfer gradient feature parameters between the actor network and the critic network, one is the policy network (Actor), which is used to generate OFDM training sequence subcarrier parameter actions; the other is the value network (Critic), which is used to evaluate the value of the action, so that it can select the optimal action according to the current state and update its strategy according to the feedback reward; in this way, the agent can adaptively learn the characteristics and dynamic change laws of the underwater acoustic channel without any prior knowledge or external guidance, thereby realizing the optimization of OFDM efficient parameter underwater acoustic communication.
[0042] (3) The present application can implement an adaptive adjustment strategy for the subcarrier parameters of the OFDM training sequence for underwater acoustic communication, and adapt to changes in the underwater acoustic channel by balancing the time-frequency resource consumption and channel communication performance of the OFDM underwater acoustic communication system. During the underwater acoustic communication process, there is no need to perform channel estimation or feedback information based on traditional preset parameters to avoid falling into a local optimal solution in a fixed environment. At the same time, the pre-trained model of deep reinforcement learning can also be deployed on the transducer client, reducing the computational complexity and communication overhead, and can effectively improve the efficiency of underwater acoustic communication in complex marine environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A flow chart of an intelligent underwater acoustic communication method based on a hybrid training sequence strategy provided in this application;
[0044] Figure 2 The OFDM underwater acoustic communication architecture diagram based on the hybrid training sequence strategy provided by this application;
[0045] Figure 3 OFDM channel modulation diagram of the hybrid training sequence strategy provided by this application;
[0046] Figure 4 A diagram of the deep learning model of OFDM underwater acoustic signals for the deterministic strategy provided in this application. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0048] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship, for example, A / B means either A or B.
[0049] The terms "first" and "second" in this specification and claims are used to distinguish different objects rather than to describe a specific order of objects. For example, "first response message" and "second response message" are used to distinguish different response messages rather than to describe a specific order of response messages.
[0050] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0051] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.
[0052] The technical solutions provided in the embodiments of this application are described in detail below.
[0053] This embodiment provides an intelligent underwater acoustic communication method based on a hybrid training sequence strategy. Figure 1 and 2 As shown, the intelligent underwater acoustic communication method includes the following steps:
[0054] S1 initializes the underwater acoustic modulation communication environment: the subcarrier of the signal is divided into multiple subchannels using the orthogonal frequency division multiplexing method, and the subcarrier parameters are modulated using a mixed training sequence strategy so that multiple subchannels can transmit underwater acoustic signals in parallel.
[0055] The specific steps to initialize the underwater acoustic modulation communication environment are as follows:
[0056] S101 divides the underwater acoustic channel according to the available idle bandwidth and caches the underwater acoustic signal data to be transmitted in segments;
[0057] S102 divides the cached multiple segments of underwater acoustic signal data into signal data queues, and adds the training sequence subcarriers to the signal data queues;
[0058] Specifically, at the underwater acoustic signal data transmitting end of the underwater acoustic signal communication system (hereinafter referred to as the system), the OFDM underwater acoustic signal is divided into underwater acoustic signal data queues, and the transmitted data is recorded as:
[0059]
[0060] The system sending module sets the parallel signal data queue size and divides the sending data into M parallel data blocks according to the number of parallel threads supported by the system, expressed as:
[0061] [D s ] M ={[D s ] 1 ,[D s ] 2 ,[D s ] 3 ,...,[D s ] M-1},
[0062] [D s ] k ∈D s and k={0,1,2,...,N}∈M; (2)
[0063] Among them, D s is the underwater acoustic signal data sent, M is the number of parallel data blocks, k is the total number of underwater acoustic signals sent, and n is the nth underwater acoustic signal data sent.
[0064] S103 uses multiple threads to perform inverse fast Fourier transform on the underwater acoustic signal after adding the training sequence subcarrier, and serializes the transmitted underwater acoustic signal;
[0065] Specifically, initialize the OFDM underwater acoustic communication bandwidth, record the system's dynamically allocable bandwidth as B, then divide the carrier into N bandwidths, and the bandwidth of each subcarrier is f = B / N. Record the phase factor of the wave modulation process as Inside, among them, is the sth subcarrier, j is the modulation amplitude, The inverse fast Fourier transform (IFFT) algorithm is used to process the underwater acoustic signal after adding the training sequence subcarrier. This can be parallelized in various ways, such as using OpenMP to create multiple threads and assigning different data blocks to each thread for IFFT calculations. Specifically, the system multiplies the scrambled phase sequence with the transmitted underwater acoustic data to add redundant information to the data signal, generating a set of underwater acoustic signal data to be transmitted at the system transmitter.
[0066] S104 repeats steps S101 to S103 until all OFDM underwater acoustic signals to be processed are successfully sent.
[0067] S2 builds a Markov decision model: Based on the channel characteristics of the underwater acoustic modulation communication environment, the subcarrier parameter adjustment problem of the mixed training sequence in the orthogonal frequency division multiplexing method is modeled as a Markov decision model. The specific steps include:
[0068] S201 defines the state space of the Markov decision model: select all underwater acoustic signal data segments with the same <hydrophone ID, hydrophone address, duration, modulation protocol> from the underwater acoustic signal output in step S1, and group the underwater acoustic signal data segments into {S_1, S_2, S_3, …, S_n}, where S_n represents the nth data segment;
[0069] S202 defines the action space of the Markov decision model as:
[0070] ai={A_a,A_t,A_f} (3)
[0071] Where A_a = {a1, a2} is the training sequence type, A_t = {t1, t2} represents the set of interpolation positions of the training sequence in the time domain, A_f = {f1, f2} represents the set of interpolation positions of the training sequence in the frequency domain, and ai represents the action space;
[0072] Specifically, such as Figure 3 As shown, if the training sequence subcarrier interpolation is performed in the time domain, then in ai={A_a, A_t, A_f}, A_a is a normal training sequence, and its time domain and frequency domain positions are recorded; if the training sequence subcarrier interpolation is performed in both the time domain and the frequency domain, then in ai={A_a, A_t, A_f}, A_a is an orthogonal training sequence, and its time domain and frequency domain positions are recorded.
[0073] S203 decomposes the action space ai vector into Figure 3 The vector shown in the figure contains two dimensions: time domain and frequency domain. t is the component of underwater acoustic signal in the time domain, n f The components of the underwater acoustic signal in the frequency domain are represented, and the Gumbel-softmax method (a reparameterization technique for processing discrete variables) is used to convert the discrete action space parameters into a probability distribution vector, which serves as the input of the deterministic gradient strategy DDPG neural network constructed subsequently.
[0074] Specifically, we take an n-dimensional vector v output in the action space and generate n independent samples {∈1,…,∈ n}, through G i =-log(-log(∈ i ))Calculate G i ,∈ n is the nth independent sample that follows the U(0,1) distribution, G i Gumbel distribution is a function that randomly samples discrete values.
[0075] The corresponding addition results in a new value vector:
[0076] v′=[v1+G1,v2+G2,...,v n +G n ] (4)
[0077] Among them, v' is the new value vector, v n is the nth original discrete vector, G n is the nth continuous distribution vector.
[0078] The selection probability of each category is calculated by the softmax function, and the value is gradually reduced in subsequent model training to approach the true discrete distribution:
[0079]
[0080] Among them, τ is the temperature parameter, v i ' and v j ' represents the unnormalized score, e is a natural constant; i and j represent the category numbers.
[0081] S3 designs a reward function based on bit error rate and transmission performance planning based on the Markov decision model. The specific steps include:
[0082] S301 builds an average bit error rate calculation model:
[0083]
[0084] Where α and β are modulation coefficients, F(x) is the cumulative distribution function of the underwater acoustic channel signal-to-noise ratio, x represents the xth underwater acoustic signal transmission sequence, and N is the total number of underwater acoustic signal transmissions;
[0085] S302 builds a signal power factor model for the underwater acoustic signal transmission process:
[0086]
[0087] Among them, R p is the signal power factor, p(x) is the OFDM signal power value transmitted in the underwater acoustic channel;
[0088] S303 constructs a reward function based on the average bit error rate calculation model and the signal power factor model. The reward function is expressed as:
[0089]
[0090] Among them, w k is the underwater acoustic signal weight, r i,j is the linear relationship between the average bit error rate and signal power during the system transmission process, b is a constant parameter, s t is the state space, a t is the action space, R(s t, a t ) represents the sum of the weighted benefits of each transmitted underwater acoustic signal, R D is the average bit error rate, R p is the signal power factor, k represents the kth underwater acoustic transmission sequence, and n is the total number of transmission sequence numbers.
[0091] This reward function provides a feedback mechanism for the agent, enabling it to adjust its behavior based on environmental feedback, thus enabling effective reinforcement learning. Specifically, the agent can learn how to take actions in an underwater acoustic communication environment to achieve its goals by maximizing cumulative rewards, demonstrating good generalization and robustness.
[0092] like Figure 4 As shown in the figure, the strategy network structure and value network structure of the DDPG model are constructed. Compared with the ordinary DDPG network model, this embodiment designs a hybrid network structure based on a shared actor-critic (Co-Actor-Critic) and combines the self-attention module transformer to obtain a deep reinforcement learning DDPG model. The specific steps include:
[0093] (1) Design a hybrid deterministic gradient strategy DDPG network structure based on a shared actor-critic (Co-Actor-Critic). The policy network Actor integrates an attention transformer structure. Actor outputs a deterministic action, which generates the network definition of this deterministic action. Actor also has a target network with the same structure but different parameters, which is used to update the value network Critic. The internal network Internal is a shared two-layer fully connected layer, which serves as the shared underlying network of the actor network and the critic network, ensuring that the gradient features can be passed to the corresponding network during training. The input of the value network Critic is the state-action pair. A transformer module is added to the deep neural network. The transformer module obtains the weights trained by the internal network as input, and then encodes the input features by stacking multiple encoder structures to obtain the encoded information matrix. Finally, the information is input into the stacked Decode structure to decode the encoded information matrix in space or channel, and the weighted global information is used as the output Q value, thereby utilizing the transformer's self-attention mechanism to enhance the state representation in DDPG and capture dependencies over a longer time range, thereby obtaining a deep reinforcement learning DDPG model.
[0094] (2) In the training of the deep reinforcement learning DDPG model, first, a sample (S, a, r, S′) is taken from the experience pool for training and the current critic network is updated. S and a in the experience pool (S, a, r, S′) are input into the current critic to obtain the current Q value. Then, the current actor network is updated, and the action corresponding to the state s is calculated by using the actor network. The Q value is given in the current critic, and the actor is updated so that the Q value output is maximized. Finally, in the internal network, the output of the underwater acoustic signal sample is divided into several two-dimensional matrices of the same size through the encoder and decoder structure. The shared fully connected layer is used for feature extraction according to mean pooling or maximum pooling. The obtained hidden features are concat-operated with the action action, and the parameters of the current network are used to update the target actor network and the target critic network.
[0095] S4 trains the deep reinforcement learning DDPG model (i.e., the deterministic gradient strategy DDPG neural network): collects the output underwater acoustic signal and the received underwater acoustic signal as positive and negative samples of the underwater acoustic signal, and uses the positive and negative samples of the underwater acoustic signal to train the deep reinforcement learning DDPG model. The specific training steps need to be combined with the aforementioned reward function. Since the model training method is a conventional technical means in this field and is not an innovation of this application, it will not be repeated here; then the trained deep reinforcement learning DDPG model is used to restore the underwater acoustic signal in the actual underwater acoustic communication process.
[0096] Specifically, at the receiving end of the system, the system tests and monitors the underwater acoustic communication link signal, using the underwater acoustic signal samples sent and received by the system as the input of the DDPG model. The policy network and evaluation network in the DDPG model are trained through the captured positive and negative samples of the underwater acoustic signals. Combined with the reward function, the system automatically adjusts the time-frequency parameters of the training sequence carrier according to the model decision results, and continuously monitors the bit error rate during the underwater acoustic communication process.
[0097] After obtaining the trained deep learning model, perform the following steps for practical application:
[0098] (1) Deploy the trained deep learning model on the target server.
[0099] (2) The system turns on the underwater acoustic receiving module, collects the OFDM underwater acoustic signal with underwater acoustic noise, records the OFDM underwater acoustic signal sent and received this time as a sample, and calculates the transmission characteristic information such as the channel bit error rate, communication power, and training sequence occupied bandwidth.
[0100] (3) The system starts the underwater acoustic signal transmission analysis module, loads the trained deep neural network model, and uses the channel estimation algorithm to calculate the channel characteristics, and then performs fast Fourier transform to restore the underwater acoustic signal.
[0101] This application proposes an intelligent underwater acoustic modulation detection method based on hybrid training sequence carrier orthogonal frequency division multiplexing (TC-OFDM). This method can autonomously learn and select the optimal OFDM training sequence carrier parameters according to the real-time state of the underwater acoustic channel using deep deterministic policy gradient (DDPG) to achieve the goal of maximizing communication efficiency.
[0102] This application combines the deep reinforcement learning (DDPG) model technology to propose a new OFDM parameter optimization framework based on deep reinforcement learning, and designs a hybrid deep neural network structure based on a shared actor-critic (Co-Actor-Critic). This hybrid deep neural network structure combines the self-attention mechanism, which can effectively extract the time-frequency characteristics and time-varying characteristics of the underwater acoustic channel and output the optimal OFDM parameters; it also designs a reward function based on bit error rate and transmission power. This reward function can comprehensively consider communication quality and energy efficiency, guide the deep neural network to learn and optimize, has good generalization ability and robustness, and can cope with unknown or dynamically changing underwater acoustic channel environments.
[0103] It is understandable that the detailed functional implementation of each of the above units / modules can be found in the introduction of the aforementioned method embodiment, and will not be repeated here.
[0104] It should be understood that the above-mentioned device is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program module in the device are similar to those described in the above-mentioned method. The working process of the device can refer to the corresponding process in the above-mentioned method and will not be repeated here.
[0105] Based on the method in the above embodiment, an embodiment of the present application provides an electronic device, which may include: a processor, a communications interface, a memory, and an underwater acoustic transducer, wherein the processor, the communications interface, and the memory communicate with each other via the underwater acoustic transducer, and the underwater acoustic transducer can convert underwater acoustic signals into electrical signals. The processor can call logic instructions in the memory to execute the method in the above embodiment.
[0106] In addition, the logical instructions in the above-mentioned memory can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.
[0107] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.
[0108] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.
[0109] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0110] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0111] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).
[0112] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.
[0113] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. An intelligent underwater acoustic communication method based on a hybrid training sequence strategy, characterized in that: The following steps are involved: S1 uses orthogonal frequency division multiplexing to divide the subcarrier into multiple subchannels, and adopts a mixed training sequence strategy to modulate the subcarrier parameters so that the multiple subchannels transmit underwater acoustic signals in parallel; Based on the channel characteristics of underwater acoustic modulation communication environment, S2 models the subcarrier parameter adjustment problem of the mixed training sequence in the orthogonal frequency division multiplexing method as a Markov decision model; S3 designs a reward function based on bit error rate and transmission performance planning based on the Markov decision model, and builds a deep reinforcement learning (DDPG) model. The deep reinforcement learning (DDPG) model includes a policy network structure, an internal network structure, and a value network structure. S4 collects the output underwater acoustic signal and the received underwater acoustic signal as positive and negative samples of the underwater acoustic signal, uses the positive and negative samples of the underwater acoustic signal to train the deep reinforcement learning DDPG model; and uses the trained deep reinforcement learning DDPG model to restore the actual underwater acoustic signal.
2. The intelligent underwater acoustic communication method based on a hybrid training sequence strategy according to claim 1, characterized in that: In step S1, multiple sub-channels are enabled to transmit underwater acoustic signals in parallel, which specifically includes the following steps: S101 divides the underwater acoustic channel according to the available idle bandwidth and caches the underwater acoustic signal data to be transmitted in segments; S102 divides the cached multiple segments of underwater acoustic signal data into signal data queues, and adds the training sequence subcarriers to the signal data queues; S103 uses multiple threads to perform an inverse fast Fourier transform on the underwater acoustic signal after the training sequence subcarrier is added, and serially transmits the underwater acoustic signal; S104 repeats steps S101 to S103 until all the underwater acoustic signals to be transmitted are transmitted.
3. The intelligent underwater acoustic communication method based on a hybrid training sequence strategy according to claim 1, characterized in that: Step S2 specifically includes the following sub-steps: S201 defines the state space of the Markov decision model: selects all underwater acoustic signal data segments with the same <hydrophone ID, hydrophone address, duration, modulation protocol> from the underwater acoustic signal output in step S1, and groups the underwater acoustic signal data segments; S202 defines the action space ai of the Markov decision model as: ai={A_a,A_t,A_f} Where A_a is the training sequence type, A_t represents the set of interpolation positions of the training sequence in the time domain, and A_f represents the set of interpolation positions of the training sequence in the frequency domain. S203 decomposes the vector of the action space into vectors including two dimensions of time domain and frequency domain, and converts the discrete action space parameters into a probability distribution vector.
4. The intelligent underwater acoustic communication method based on a hybrid training sequence strategy according to claim 3, characterized in that: In step S202, if the training sequence subcarrier interpolation is performed in the time domain, A_a is a normal training sequence, and its time domain and frequency domain positions are recorded; if the training sequence subcarrier interpolation is performed in both the time domain and the frequency domain, A_a is an orthogonal training sequence, and its time domain and frequency domain positions are recorded.
5. The intelligent underwater acoustic communication method based on a hybrid training sequence strategy according to claim 1, characterized in that: In step S3, the steps of designing a reward function based on bit error rate and transmission performance planning include: S301 builds an average bit error rate calculation model: Where α and β are modulation coefficients, F(x) is the cumulative distribution function of the underwater acoustic channel signal-to-noise ratio, x represents the xth underwater acoustic signal transmission sequence, and N is the total number of underwater acoustic signal transmissions; S302 constructs a signal power factor model of the underwater acoustic signal transmission process: Among them, R p is the signal power factor, p(x) is the OFDM signal power value transmitted in the underwater acoustic channel; S303 constructs the reward function based on the average bit error rate calculation model and the signal power factor model.
6. The intelligent underwater acoustic communication method based on a hybrid training sequence strategy according to claim 5, characterized in that: The reward function is expressed as: Among them, w k is the underwater acoustic signal weight, r i,j is the linear relationship between the average bit error rate and signal power during the system transmission process, b is a constant parameter, s t is the state space, a t is the action space, R(s t, a t ) represents the sum of the weighted benefits of each transmitted underwater acoustic signal, R D is the average bit error rate, R p is the signal power factor, k represents the kth underwater acoustic transmission sequence, and n is the total number of transmission sequence numbers.
7. The intelligent underwater acoustic communication method based on a hybrid training sequence strategy according to claim 1, characterized in that: Attention modules are integrated into both the strategy network structure and the value network structure.
8. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the intelligent underwater acoustic communication method according to any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program runs on a processor, the processor is enabled to execute the intelligent underwater acoustic communication method according to any one of claims 1 to 7.
10. A computer program product, characterized in that When the computer program product runs on a processor, the processor is enabled to execute the intelligent underwater acoustic communication method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Optimization method, device, equipment and medium of Massive MIMO
CN109379752A
Reinforcement learning-based AUV behavior planning and motion control method
CN110333739A