A brain-like obstacle avoidance decision method, device and storage medium for unmanned underwater vehicles based on pulse neural network
By combining spiking neural networks and deep reinforcement learning, a soft reset membrane potential update mechanism and a spiking encoder-decoder were designed, and a spiking neural network model was constructed. This solved the energy efficiency and computational resource problems of unmanned underwater vehicles in obstacle avoidance in complex underwater environments, and achieved low-energy, time-continuous, safe and reliable autonomous obstacle avoidance.
Patent Information
- Application Number
- CN202411438365.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-15
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-10-15
AI Technical Summary
Existing obstacle avoidance methods for unmanned underwater vehicles in complex and unknown underwater environments have high energy efficiency and computational resource requirements. Traditional methods cannot adapt to environmental changes, and deep reinforcement learning methods have high energy consumption and computational resource requirements, making them difficult to apply on a large scale.
A soft reset membrane potential update mechanism is designed by combining spiking neural networks and deep reinforcement learning. The soft reset pulse Actor network and deep Critic network are integrated, and the state information is transformed through pulse encoder and decoder. Low-energy neuromorphic computing is used to construct the spiking neural network model.
It achieves low-energy, time-continuous, safe and reliable autonomous obstacle avoidance exploration of unmanned underwater vehicles in complex and unknown underwater environments, reducing the energy consumption and computing requirements of each decision step.
Smart Images

Figure CN119418177B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of obstacle avoidance technology for unmanned underwater vehicles, specifically relating to a brain-like obstacle avoidance decision-making method, device, and storage medium for unmanned underwater vehicles based on spiking neural networks. Background Technology
[0002] With the development of marine resources and adhering to the strategic policy of developing the marine field, unmanned underwater vehicles (UUVs), as important equipment for in-depth exploration of the marine environment, have gradually become key marine equipment for achieving comprehensive marine resource exploration and demonstrating comprehensive deep-sea capabilities. Due to their advantages such as being unconstrained by cables, having a wide underwater operating range, high safety, and strong flexibility, they can autonomously navigate in marine environments that are difficult for humans to access or where dangers are unpredictable. They are widely used in civilian fields for tasks such as marine environmental monitoring, marine infrastructure inspection, and underwater emergency rescue, and have significant applications in marine surveying, fleet escort, and detection methods. Autonomous obstacle avoidance technology, as one of the core technologies for intelligent and safe navigation of UUVs, runs through the entire underwater navigation process. It requires UUVs to make timely and correct decision-making commands and accurately avoid obstacles in the face of complex and unknown underwater environments, forming an adaptive underwater autonomous obstacle avoidance capability. This is a research hotspot and technical challenge in the field of UUVs.
[0003] Brain-inspired obstacle avoidance is an important research direction in computational neuroscience and artificial intelligence. It aims to simulate the structure, function, and complex information processing mechanisms of the human brain, using techniques such as spiking neural networks and reinforcement learning to develop advanced decision-making capabilities similar to human thinking, enabling unmanned underwater vehicles (UUVs) to autonomously avoid obstacles and operate safely in complex and unknown marine environments. However, underwater operations for UUVs require temporal continuity, meaning the output at the current moment must be based on the state data from all previous moments. Each decision-making step by an UUV consumes energy, making energy efficiency and computational resource requirements crucial for long-duration cruises and large-scale search and rescue missions. Therefore, efficient, reliable, safe, and low-energy-consumption brain-inspired obstacle avoidance technology is of significant research importance for UUVs to achieve autonomous obstacle avoidance in complex underwater environments.
[0004] Traditional obstacle avoidance methods for unmanned underwater vehicles rely on scene maps, requiring the pre-construction of scene-based grid maps. These scenes are fixed and difficult to relocate, necessitating the reconstruction of scene information whenever the environment changes. Especially in unfamiliar environments, traditional algorithms cannot adapt and are only capable of handling obstacle avoidance tasks in simple conditions. Therefore, in complex and unknown underwater environments, traditional obstacle avoidance methods face challenges in terms of success rate, timely processing capability, and computational efficiency.
[0005] Reinforcement learning methods continuously interact with the environment, acquiring action instructions based on the current state to optimize decision-making strategies. Meanwhile, deep learning methods have advantages in handling high-dimensional information. Therefore, deep reinforcement learning, which combines reinforcement learning and deep learning, has gradually become a favored obstacle avoidance method for unmanned underwater vehicles (UUVs) by researchers both domestically and internationally, overcoming the limitations of traditional obstacle avoidance methods that are restricted by scene maps. However, each decision step of an UUV requires continuity in the temporal dimension. For long-duration underwater operations, the UUV's embedded resources are limited, and each underwater operation requires significant energy consumption and relies on high-performance computing equipment. Existing deep reinforcement learning methods have high energy consumption and computational resource requirements, making it difficult to apply these methods on a large scale in underwater environments to solve related practical problems.
[0006] Spiking neural networks (SNNs) are a new generation of neural network models inspired by biological neurons. They possess advanced levels of bio-neural simulation, mimicking the asynchronous and parallel information transmission in the human brain via pulses. They perform computations through pulse addition data streams, avoiding the numerous floating-point multiplication operations of traditional neural networks. SNNs offer advantages such as temporal continuity, rapid processing capabilities, and low energy consumption. Given these advantages, applying SNNs to obstacle avoidance tasks for unmanned underwater vehicles (UUVs) is both reasonable and feasible. Therefore, combining SNNs with deep reinforcement learning to fully leverage their respective strengths and construct a brain-inspired and biologically plausible SNN model is a crucial problem that urgently needs to be solved for UUVs to perform autonomous obstacle avoidance tasks in complex and unknown underwater environments. Summary of the Invention
[0007] The purpose of this invention is to provide a brain-like obstacle avoidance decision-making method, device and storage medium for unmanned underwater vehicles based on spiking neural networks. It adopts the soft reset membrane potential update mechanism of spiking neurons to better reflect the changes in neuronal membrane potential, and combines neuromorphic computing and deep reinforcement learning to realize autonomous obstacle avoidance exploration of unmanned underwater vehicles in complex and unknown underwater environments.
[0008] The objective of this invention is achieved through the following technical solution:
[0009] A brain-like obstacle avoidance decision-making method for unmanned underwater vehicles based on spiking neural networks includes the following steps:
[0010] Step 1: Interact with the underwater environment using an unmanned underwater vehicle to obtain raw state information;
[0011] Step 2: Design a soft reset membrane potential update mechanism for spiking neurons to reflect changes in neuronal membrane potential;
[0012] Step 3: Design a pulse encoder and pulse decoder to convert between continuous state information and pulse sequence information;
[0013] Step 4: Design a spiking neural network model that integrates a soft-reset spiking Actor network and a deep Critic network;
[0014] Step 5: Input the state information into the trained spiking neural network model, and perform autonomous obstacle avoidance exploration of the unmanned underwater vehicle in complex and unknown underwater environments based on the obstacle avoidance decisions output by the network.
[0015] Further, the original state information mentioned in step 1 includes the Euclidean distance and angular direction from the unmanned underwater vehicle to the target point, the linear velocity and angular velocity of the unmanned underwater vehicle, and the ranging information obtained by the unmanned underwater vehicle through sonar sensors; based on the interaction between the unmanned underwater vehicle and the environment, a reward function is set:
[0016]
[0017] Among them, R goal and R obs R represents the reward value for the unmanned underwater vehicle reaching the target point and the penalty value for a collision, respectively. dis This represents the change in distance between the unmanned underwater vehicle and the target point at the current moment and the previous moment. (D) th and O th These are the hyperparameters of the threshold, D dis O represents the distance between the center of the unmanned underwater vehicle and the target point. dis This represents the distance between the center of the unmanned underwater vehicle and the obstacle, where α is a non-zero constant.
[0018] Furthermore, the soft reset membrane potential update mechanism of the spiking neurons in step 2 can retain the portion of the membrane voltage exceeding the threshold when the membrane potential is reset, ensuring that the membrane potential data at the current moment changes based on the membrane potential data at the previous moment, corresponding to the temporal continuity of the spiking neural network at the biological level. Deploying the soft reset membrane potential update mechanism in each spiking neuron ensures that the decision made by the unmanned underwater vehicle in the underwater environment to avoid obstacles at the current moment is related to the decision made at the previous moment. The reset is only performed after each path is executed, which has temporal continuity and biological rationality.
[0019] Furthermore, the soft-reset membrane potential update mechanism is deployed in each spiking neuron, and the membrane current, membrane voltage, and pulse firing status are expressed by the following formulas:
[0020]
[0021] in, and η represents membrane current and membrane voltage, respectively. c and ηv W represents the attenuation factors of current and voltage, respectively. k Let b represent the synaptic weight matrix. k W represents the synaptic bias vector. k and b k These are the parameters that need to be updated during network training. Indicates the output pulse. Indicates pulse gating. U th This represents the membrane potential threshold. When U t >U th That is, the membrane potential U at the current moment. t If the threshold is exceeded, the soft-reset spiking neuron will fire a pulse signal.
[0022] Furthermore, the pulse encoder in step 3 uses Poisson coding to encode the original state information and generates a sequence of pulses as state input information within a given time window; the pulse decoder processes the pulse sequence of the output layer neurons into action information required for autonomous obstacle avoidance and exploration by averaging the pulses.
[0023] The Poisson encoding is:
[0024]
[0025] Where P(n) represents the probability that a soft-reset spiking neuron fires a pulse within a fixed time step T, n represents the number of pulses fired by the soft-reset spiking neuron, and λ represents the firing frequency.
[0026] Furthermore, the spiking neural network model in step 4 adopts an Actor-Critic architecture, which combines a soft-reset spiking Actor network and a deep Critic network based on the DDPG algorithm. The soft-reset spiking Actor network consists of four fully connected layers built by the soft-reset spiking neurons, ensuring that the autonomous obstacle avoidance exploration of the model has temporal continuity. The deep Critic network adopts an artificial neural network-based model, which consists of four fully connected layers built by traditional artificial neurons.
[0027] Furthermore, the training process of the spiking neural network model in step 5 includes:
[0028] The raw state information acquired by the unmanned underwater vehicle is input into the soft reset pulse Actor network composed of soft reset pulse neurons through the pulse encoder to obtain pulse output information, and then the pulse decoder is used to obtain the motion information of the unmanned underwater vehicle.
[0029] The reward value of the unmanned underwater vehicle is determined based on the underwater environment.
[0030] The current state information, the current action information, the next moment state information, and the reward value are stored as sequence information in the experience pool. A small batch of experience sequences are randomly selected for training the spiking neural network model and updating the relevant weight information.
[0031] Furthermore, the soft-reset pulse Actor network is trained using the STBP algorithm for backpropagation in both the time and spatial domains, while the deep Critic part is trained using the normal backpropagation algorithm. After the soft-reset pulse Actor network completes its forward propagation, the network output is fed into the last layer of the deep Critic network to generate the Q-value. The purpose of training the soft-reset pulse Actor network is to generate obstacle avoidance decisions for unmanned underwater vehicles with the maximum Q-value, as follows:
[0032] L TD (τ)=(R m +γQ m+1 -Q m ) 2
[0033] L Q (ξ)=-Q m
[0034] Among them, L TD (τ) represents the TD loss function of the deep Critic network, L Q (ξ) represents the loss function of the soft-reset pulse Actor network, R m Q represents the reward value obtained by the unmanned underwater vehicle, γ represents the discount factor, and Q represents the reward value obtained by the unmanned underwater vehicle. m Q represents the output Q-value of the deep Critic network. m+1 This represents the output Q-value of the Critic target network;
[0035] Repeat the above training steps, periodically copying the network parameters to the target network using a soft update method, until the unmanned underwater vehicle successfully avoids obstacles in the underwater environment. The parameter updates are as follows:
[0036] ξ'=μξ+(1-μ)ξ'
[0037] τ'=μτ+(1-μ)τ'
[0038] Where μ is a very small constant, ξ and τ represent the parameters of the soft reset pulse Actor network and the deep Critic network, respectively, and ξ' and τ' represent the parameters of the target network.
[0039] A computer device / equipment / system includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of a brain-like obstacle avoidance decision-making method for unmanned underwater vehicles based on a spiking neural network.
[0040] A computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the steps of a brain-like obstacle avoidance decision-making method for unmanned underwater vehicles based on a spiking neural network.
[0041] The beneficial effects of this invention are as follows:
[0042] 1. This invention designs a soft-reset membrane potential update mechanism to better reflect changes in neuronal membrane potential, designs a pulse encoder-decoder to convert continuous state information and pulse sequence information, and designs a spiking neural network model that integrates a soft-reset pulse Actor network and a deep Critic network. This avoids the high energy consumption of existing obstacle avoidance methods for unmanned underwater vehicles (UUVs) due to numerous floating-point multiplication operations and the discontinuous decision-making time at each underwater step, achieving efficient, sustainable, and safe autonomous obstacle avoidance for UUVs.
[0043] 2. This invention employs a combination of soft-reset pulse Actor networks and deep Critic networks as a brain-like intelligent obstacle avoidance decision-making strategy for unmanned underwater vehicles (UUVs). It is adapted to a specialized neuromorphic brain-like processing chip, possessing low-power consumption characteristics and spatiotemporal information processing capabilities. Therefore, UUVs can achieve autonomous obstacle avoidance exploration in complex and unknown underwater environments thanks to their low-power, time-continuous, and reliable obstacle avoidance capabilities, demonstrating significant application value. Attached Figure Description
[0044] Figure 1 This is a flowchart of the brain-like obstacle avoidance decision-making method for unmanned underwater vehicles based on spiking neural networks according to the present invention.
[0045] Figure 2 This is a structural diagram of the spiking neural network model of the present invention;
[0046] Figure 3 This is a schematic diagram of the soft-reset pulse neuron used in the soft-reset pulse Actor network of this invention;
[0047] Figure 4 This is a schematic diagram of the artificial neurons used in the deep Critic network of this invention;
[0048] Figure 5 This is a graph of the pseudo-gradient function of the present invention;
[0049] Figure 6 This is a schematic diagram of the back propagation process of the spiking neuron in this invention. Detailed Implementation
[0050] The present invention will now be further described with reference to the accompanying drawings.
[0051] Example 1
[0052] This embodiment provides a brain-like obstacle avoidance decision-making method for unmanned underwater vehicles based on spiking neural networks, such as... Figure 1 The diagram shown is an overall flowchart of the brain-like obstacle avoidance decision-making method of the present invention, which includes the following steps:
[0053] Step 1: The unmanned underwater vehicle interacts with the underwater environment to obtain raw state information;
[0054] The raw state information includes the Euclidean distance and angular direction from the unmanned underwater vehicle to the target point, the linear velocity and angular velocity of the unmanned underwater vehicle, and the ranging information obtained by the unmanned underwater vehicle through sonar sensors.
[0055] Based on the interaction between the unmanned underwater vehicle and the environment, a specific reward function is set as follows:
[0056]
[0057] Among them, R goal and R obs R represents the reward value for the unmanned underwater vehicle reaching the target point and the penalty value for a collision, respectively. dis This represents the change in distance between the unmanned underwater vehicle and the target point at the current moment and the previous moment. (D) th and O th These are the hyperparameters of the threshold, D dis O represents the distance between the center of the unmanned underwater vehicle and the target point. dis This represents the distance between the center of the unmanned underwater vehicle and the obstacle, where α is a non-zero constant.
[0058] Step 2: Design a soft reset membrane potential update mechanism for spiking neurons to better reflect changes in neuronal membrane potential;
[0059] The soft-reset membrane potential update mechanism retains the portion of the membrane voltage exceeding a threshold during membrane potential reset, ensuring that the current membrane potential data is based on the previous membrane potential data, corresponding to the temporal continuity of spiking neural networks at the biological level. Deploying the soft-reset membrane potential update mechanism in each spiking neuron ensures that the unmanned underwater vehicle's autonomous obstacle avoidance and exploration decisions in the underwater environment are relevant to previous decisions. Resetting only occurs after each path is executed, thus possessing both temporal continuity and biological rationality.
[0060] The specific formulas for membrane current, membrane voltage, and pulse firing in soft-reset spiking neurons are shown below:
[0061]
[0062] in, and η represents membrane current and membrane voltage, respectively. c and η v W represents the attenuation factors of current and voltage, respectively. k Let b represent the synaptic weight matrix. k W represents the synaptic bias vector. k and b k These are the parameters that need to be updated during network training. Indicates the output pulse. Indicates pulse gating. U th This represents the membrane potential threshold. When U t >U th That is, the membrane potential U at the current moment. t If the threshold is exceeded, the soft-reset spiking neuron will fire a pulse signal.
[0063] Step 3: Design a pulse encoder and pulse decoder to convert between continuous state information and pulse sequence information;
[0064] The raw, continuous state values obtained from the interaction between the unmanned underwater vehicle (UUV) and its environment include the Euclidean distance and angle from the UUV to the target point, the UUV's linear and angular velocities, and the ranging information obtained from the sonar sensors after data processing. Since the constructed network framework integrates a soft-reset pulse Actor network and a deep Critic network, and uses pulse sequences as the information carrier, it cannot be directly transmitted to the network for training. Therefore, a pulse encoder is constructed to convert the raw, continuous state information into a sequence of pulses.
[0065] The original continuous state information is first normalized to obtain pulse state probabilities with values in the range [0,1]. Then, it is sampled using Poisson encoding and iterated for time T times to process it into a state pulse sequence, which serves as the input to the soft-reset pulse Actor network. Since the output action information is also a pulse sequence, it is averaged over time using a pulse decoder and then converted into a data format acceptable to the unmanned underwater vehicle, serving as the output of the soft-reset pulse Actor network.
[0066] Poisson coding is specifically represented as:
[0067]
[0068] Where P(n) represents the probability that a soft-reset spiking neuron fires a pulse within a fixed time step T, n represents the number of pulses fired by the soft-reset spiking neuron, and λ represents the firing frequency.
[0069] Step 4: Design a spiking neural network model, introducing a soft-reset membrane potential update mechanism into each spiking neuron. This model integrates a soft-reset spiking Actor network and a deep Critic network. The spiking neural network model structure is as follows: Figure 2 The above;
[0070] The spiking neural network model designed in this invention is based on the reinforcement learning DDPG algorithm, introducing a soft-reset membrane potential update mechanism into each spiking neuron, and combining a soft-reset spiking Actor network and a deep Critic network to construct the model. Soft-reset LIF neurons are used as the basic neurons of the soft-reset spiking Actor network, containing four fully connected layers and 256 neurons in the intermediate hidden layers, thus forming the soft-reset spiking Actor network part. A schematic diagram of the soft-reset spiking neurons is shown below. Figure 3 As shown. The Critic part uses a four-layer fully connected network structure built with traditional artificial neurons. A schematic diagram of the artificial neurons used is shown below. Figure 4 As shown, ReLU is selected as the activation function for this part. The input to this part of the network is the pulse firing rate output by the soft reset pulse Actor network and the state data of the pulse encoder that failed. The network output is the Q value trained by the soft reset pulse Actor network.
[0071] Step 5: Input the state information into the trained spiking neural network model, train the network model and update the relevant weight information, and perform autonomous obstacle avoidance exploration of the unmanned underwater vehicle in complex and unknown underwater environments based on the obstacle avoidance decision output by the network.
[0072] The state information acquired by the unmanned underwater vehicle (UUV) in the simulation environment, the action output of the soft-reset spiking Actor network, the reward value, and the UUV's input state at the next moment are stored as a sequence in an experience pool. Small batches of experience are randomly selected for model training and related weight updates. The soft-reset spiking Actor network consists of multiple soft-reset spiking neurons. Due to the non-differentiable nature of spiking neural networks, a rectangular pseudo-gradient function is used to approximate the gradient of the spiking action. The pseudo-gradient function curve is shown in the figure. Figure 5 As shown. The soft reset pulse Actor network is trained using the STBP algorithm through backpropagation in both the time and spatial domains, as detailed below. Figure 6 As shown, the deep Critic part uses the normal backpropagation algorithm for network training. After the soft-reset pulse Actor network completes its forward propagation, the network output is fed into the last layer of the deep Critic network to generate the Q-value. The purpose of training the soft-reset pulse Actor network is to generate obstacle avoidance decisions for the unmanned underwater vehicle with the maximum Q-value, as detailed below:
[0073] LTD (τ)=(R m +γQ m+1 -Q m ) 2
[0074] L Q (ξ)=-Q m
[0075] Among them, L TD (τ) represents the TD loss function of the deep Critic network, L Q (ξ) represents the loss function of the soft-reset pulse Actor network, R m Q represents the reward value obtained by the unmanned underwater vehicle, γ represents the discount factor, and Q represents the reward value obtained by the unmanned underwater vehicle. m Q represents the output Q-value of the deep Critic network. m+1 This represents the output Q-value of the Critic target network.
[0076] Repeat the above training steps, periodically copying the network parameters to the target network using a soft update method, until the unmanned underwater vehicle successfully avoids obstacles in the underwater environment. The parameter update is detailed below:
[0077] ξ'=μξ+(1-μ)ξ'
[0078] τ'=μτ+(1-μ)τ'
[0079] Where μ is a very small constant, ξ and τ represent the parameters of the soft reset pulse Actor network and the deep Critic network, respectively, and ξ' and τ' represent the parameters of the target network.
[0080] In this embodiment of the invention, the construction and training of the spiking neural network model, the loading of the unmanned underwater vehicle (UUV), and the creation of the underwater environment are all performed in the Gazebo simulator within the ROS system. Underwater simulation training environments of varying complexity are set up to enable the UUV to explore unknown underwater environments and learn autonomous obstacle avoidance decisions. A curriculum-based learning method is used to train the spiking neural network model, encouraging the UUV to learn simple, brain-like obstacle avoidance decisions in simple underwater environments and gradually learn complex obstacle avoidance strategies in complex underwater environments. This facilitates faster convergence of the network model. The goal is to enable the UUV to successfully avoid obstacles and reach the target point based on the obstacle avoidance decisions output by the network on each initial path, demonstrating time continuity and low energy consumption underwater obstacle avoidance capabilities, thus completing the autonomous obstacle avoidance exploration of the UUV in complex and unknown underwater environments.
[0081] Example 2
[0082] This embodiment provides a computing device including a memory and a processor. The memory is used to store executable programs, and the processor is used to run the executable programs, thereby realizing the brain-like obstacle avoidance decision-making method for unmanned underwater vehicles based on spiking neural networks in Embodiment 1.
[0083] Example 3
[0084] This embodiment provides a computer-readable storage medium for storing a computer program. When the computer program is run by a processor, it executes the steps in the brain-like obstacle avoidance decision-making method for unmanned underwater vehicles based on spiking neural networks in Embodiment 1.
[0085] The above embodiments are all preferred embodiments of the present invention, and their protection scope is not limited to the above embodiments. All technical solutions based on the concept of the present invention fall within the protection scope of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the main concept and technical principles of the present invention should be considered within the protection scope defined by the present invention.
[0086] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A brain-like obstacle avoidance decision-making method for unmanned underwater vehicles based on spiking neural networks, characterized in that: Includes the following steps: Step 1: Interact with the underwater environment using an unmanned underwater vehicle to obtain raw state information; Step 2: Design a soft reset membrane potential update mechanism for spiking neurons to reflect changes in neuronal membrane potential; When resetting the membrane potential, the portion of the membrane voltage that exceeds the threshold can be retained, ensuring that the membrane potential data at the current moment changes based on the membrane potential data at the previous moment, which corresponds to the temporal continuity of the spiking neural network at the biological level. The soft reset membrane potential update mechanism is deployed in each spiking neuron to ensure that the decision made by the unmanned underwater vehicle in the underwater environment is related to the decision made in the current moment. The reset is only performed after each path is executed, which has temporal continuity and biological rationality. Step 3: Design a pulse encoder and pulse decoder to convert between continuous state information and pulse sequence information; The pulse encoder uses Poisson coding to encode the original state information and generates a sequence of pulses as state input information within a given time window. The pulse decoder processes the pulse sequence of the output layer neurons into action information required for autonomous obstacle avoidance and exploration by an unmanned underwater vehicle through average pulse accumulation. The Poisson encoding is: Where P(n) represents the probability that a soft-reset spiking neuron fires a pulse within a fixed time step T, n represents the number of pulses fired by the soft-reset spiking neuron, and λ represents the firing frequency. Step 4: Design a spiking neural network model that integrates a soft reset spiking Actor network and a deep Critic network. The soft reset spiking Actor network consists of four fully connected layers built by the soft reset spiking neurons, ensuring the temporal continuity of the model's autonomous obstacle avoidance exploration. The deep Critic network adopts an artificial neural network-based model, consisting of four fully connected layers built by traditional artificial neurons. Step 5: Input the state information into the trained spiking neural network model, and perform autonomous obstacle avoidance exploration of the unmanned underwater vehicle in complex and unknown underwater environments based on the obstacle avoidance decisions output by the network.
2. The brain-like obstacle avoidance decision-making method for unmanned underwater vehicles based on spiking neural networks according to claim 1, characterized in that: The raw state information mentioned in step 1 includes the Euclidean distance and angular direction from the unmanned underwater vehicle (UUV) to the target point, the linear velocity and angular velocity of the UUV, and the ranging information obtained by the UUV through sonar sensors; based on the interaction between the UUV and the environment, a reward function is set: Among them, R goal and R obs R represents the reward value for the unmanned underwater vehicle reaching the target point and the penalty value for a collision, respectively. dis This represents the change in distance between the unmanned underwater vehicle and the target point at the current moment and the previous moment; D th and O th These are the hyperparameters of the threshold, D dis O represents the distance between the center of the unmanned underwater vehicle and the target point. dis This represents the distance between the center of the unmanned underwater vehicle and the obstacle, where α is a non-zero constant.
3. The brain-like obstacle avoidance decision-making method for unmanned underwater vehicles based on spiking neural networks according to claim 1, characterized in that: The soft-reset membrane potential update mechanism is deployed in each spiking neuron, and the membrane current, membrane voltage, and pulse firing status are expressed by the following formulas: in, and η represents membrane current and membrane voltage, respectively. c and η v W represents the attenuation factors of current and voltage, respectively. k Let b represent the synaptic weight matrix. k W represents the synaptic bias vector. k and b k These are the parameters that need to be updated during network training; Indicates the output pulse. Indicates pulse gating; U th Indicates the membrane potential threshold; when U t >U th That is, the membrane potential U at the current moment. t If the threshold is exceeded, the soft-reset spiking neuron will fire a pulse signal.
4. The brain-like obstacle avoidance decision-making method for unmanned underwater vehicles based on spiking neural networks according to claim 1, characterized in that: The training process of the spiking neural network model in step 5 includes: The raw state information acquired by the unmanned underwater vehicle is input into the soft reset pulse Actor network composed of soft reset pulse neurons through the pulse encoder to obtain pulse output information, and then the pulse decoder is used to obtain the motion information of the unmanned underwater vehicle. The reward value of the unmanned underwater vehicle is determined based on the underwater environment. The current state information, current action information, next moment state information, and the reward value are stored as sequence information in the experience pool. A small batch of experience sequences are randomly selected for training the spiking neural network model and updating the relevant weight information.
5. The brain-like obstacle avoidance decision-making method for unmanned underwater vehicles based on spiking neural networks according to claim 4, characterized in that: The soft reset pulse Actor network is trained using the STBP algorithm for backpropagation in both the time and spatial domains, while the deep Critic part is trained using the normal backpropagation algorithm. After the forward propagation of the soft-reset pulse Actor network is completed, the network output is fed into the last layer of the deep Critic network to generate the Q-value. The purpose of training the soft-reset pulse Actor network is to generate obstacle avoidance decisions for unmanned underwater vehicles with the maximum Q-value, as follows: L TD (τ)=(R m +γQ m+1 -Q m ) 2 50 Q (ξ)=-Q m Among them, L TD (τ) represents the TD loss function of the deep Critic network, L Q (ξ) represents the loss function of the soft-reset pulse Actor network, R m Q represents the reward value obtained by the unmanned underwater vehicle, γ represents the discount factor, and Q represents the reward value obtained by the unmanned underwater vehicle. m Q represents the output Q-value of the deep Critic network. m+1 This represents the output Q-value of the Critic target network; Repeat the above training steps, periodically copying the network parameters to the target network using a soft update method, until the unmanned underwater vehicle successfully avoids obstacles in the underwater environment. The parameter updates are as follows: ξ′=μξ+(1-μ)ξ ′ τ′=μτ+(1-μ)τ' Where μ is a very small constant, ξ and τ represent the parameters of the soft reset pulse Actor network and the deep Critic network, respectively, and ξ′ and τ′ represent the parameters of the target network.
6. A computer device / equipment / system, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.
7. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that: When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Unmanned aerial vehicle brain-like obstacle avoidance method based on spiking neuron model and AC network
CN114859961A
Unmanned aerial vehicle brain-like obstacle avoidance method for avoiding collision with dynamic obstacles
CN114863268A